EDBT 2026 Demo / reviewers in the wild / expert
Gang Tan
dblp:91/6206
· DBLP profile ↗
111ranked-venue papers
10as first author
55since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 46 · 1 first-author · 19 since 2021Software engineering, systems software and programming languages · 33 · 5 first-author · 17 since 2021Systems, architecture and hardware · 10 · 5 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 6 since 2021Computer networks · 7 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language ModelsabstractLarge Language Models (LLMs) exhibit impressive capabilities, but remain susceptible to a growing spectrum of safety risks, including jailbreaks, toxic content, hallucinations, and bias. Existing defenses often address only a single threat type or resort to rigid outright rejection, sacrificing user experience and failing to generalize across diverse and novel attacks. This paper introduces Adversarial Scenario Extrapolation (ASE), a novel inference-time computation framework that leverages Chain-of-Thought (CoT) reasoning to simultaneously enhance LLM robustness and seamlessness. ASE guides the LLM through a self-generative process of contemplating potential adversarial scenarios and formulating defensive strategies before generating a response to the user query. Comprehensive evaluation on four adversarial benchmarks with four latest LLMs shows that ASE achieves near-zero jailbreak attack success rates and minimal toxicity, while slashing outright rejections to Md. Rafi Ur Rashid, Vishnu Asutosh Dasu, Ye Wang 0001, Gang Tan, Shagufta Mehnaz |
AAAI | 4 |
| 2026 | BinType: Type based Indirect Call Target Refinement on Binary ProgramsabstractConstructing precise and sound control flow graphs (CFGs) is critical for enforcing control flow integrity (CFI) defense against control-flow hijacking related exploits. One major challenge of constructing such CFGs is to infer the targets of indirect calls, which suffers extra difficulty on commercial off-the-shelf (COTS) binaries due to the absence of source-level information. One classic direction is to use points-to analysis, but it often suffers from scalability issues. Thus, signature-matching based approaches are employed to mitigate the scalability issues. However, existing binary-level signatures are coarse-grained and the paired matching policy must be conservative to pursue soundness, resulting in CFG precision decrease. In this paper, we present BinType, a new signature-matching approach that relies on type inference to improve the signature granularity. Methodology-wise, BinType identifies storage locations of high-confidence types, generates type equivalence relations between storage locations, and propagates types to callsite arguments and function parameters by following the type equivalence relations. Our evaluations show that BinType achieves >20% higher precision than the previous arity-based technique. Moreover, BinType shows comparable precision against the state-of-the-art points-to analysis based approach, while significantly improving the efficiency. Sun Hyoung Kim, Dongrui Zeng, Monika Santra, Gang Tan |
CODASPY | 4 |
| 2026 | Dynamic graph structure correction with nonadjacent correlations for multivariate time series forecasting
Dandan He, Chaoli Lou, Gang Tan, Qingyu Xiong, Guodong Sa |
Expert Syst. Appl. | 4 |
| 2026 | Automated design and performance prediction of micromixer chips via Convolutional Neural Network models for lab-on-a-chip detection
Binfeng Yin, Zhuoao Jiang, Gang Tan, Rashid Muhammad, Shiyu Zeng |
Expert Syst. Appl. | 3 |
| 2026 | Trajectory optimization for target tracking in UUV-based multistatic sonar systems with position offset correction
Shoude Jiang, Shefeng Yan, Linlin Mao, Chunjin Jiang, Gang Tan, Wei Wang 0499 |
Signal Process. | 5 |
| 2025 | Beyond Driver Isolation - Triaging Threats against Driver IsolationabstractDevice driver isolation aims to protect kernels from faulty/malicious drivers, yet its security guarantees are not fully understood. Compartment Interface Vulnerabilities (CIVs), known in userspace applications, also impact driver isolation, but this area is underexplored. This paper surveys existing driver isolation frameworks, systematizes CIV classifications, and evaluates them in the driver isolation context. Our analysis reveals CIV prevalence under a baseline threat model, with large drivers exhibiting over 100 CIV instances and an average of 33 across the studied drivers. Enforcing additional security properties like CFI reduces average CIVs to approximately 28. This work offers insights into driver isolation security, CIV prevalence, and guidance for future systems. Yongzhe Huang, Kaiming Huang, Matthew Ennis, Vikram Narayanan, Anton Burtsev, Trent Jaeger, Gang Tan |
ACSAC | 7 |
| 2025 | Disa: Accurate Learning-based Static Disassembly with AttentionsabstractFor reverse engineering related security domains, such as vulnerability detection, malware analysis, and binary hardening, disassembly is crucial yet challenging. The fundamental challenge of disassembly is to identify instruction and function boundaries. Classic approaches rely on file-format assumptions and architecture-specific heuristics to guess the boundaries, resulting in incomplete and incorrect disassembly, especially when the binary is obfuscated. Recent advancements of disassembly have demonstrated that deep learning can improve both the accuracy and efficiency of disassembly. In this paper, we propose Disa, a new learning-based disassembly approach that uses the information of superset instructions over the multi-head self-attention to learn the instructions' correlations, thus being able to infer function entry-points and instruction boundaries. Disa can further identify instructions relevant to memory block boundaries to facilitate an advanced block-memory model based value-set analysis for an accurate control flow graph (CFG) generation. Our experiments show that Disa outperforms prior deep-learning disassembly approaches in function entry-point identification, especially achieving 9.1% and 13.2% F1-score improvement on binaries respectively obfuscated by the disassembly desynchronization technique and popular source-level obfuscator. By achieving an 18.5% improvement in the memory block precision, Disa generates more accurate CFGs with a 4.4% reduction in Average Indirect Call Targets (AICT) compared with the state-of-the-art heuristic-based approach. Monika Santra, Cong Sun 0001, Dongrui Zeng, Gang Tan |
CCS | 6 |
| 2025 | Probabilistic Verification of Cybersickness in Virtual Reality Through Bayesian NetworksabstractCybersickness remains a major challenge in virtual and mixed reality (VR/MR), yet existing methods primarily focus on predicting its onset without offering formal guarantees regarding its occurrence or effective mitigation. As VR/MR applications expand into safety-critical domains like healthcare, defense, verifiable safety assurances become essential to protect users from adverse physiological and psychological effects. This paper introduces a probabilistic verification framework leveraging Bayesian Networks (BN) to explicitly model the interactions among system parameters, human physiological responses, and cybersickness severity. Unlike deep learning approaches that lack interpretability and formal verification capabilities, the proposed BN model explicitly captures how environmental and system-level factors (e.g., luminance, spectral entropy, and image gradient complexity via HoG features) influence physiological responses (e.g., heart rate, reaction time, eye tracking), ultimately affecting cybersickness severity. By learning the joint probability distribution of these factors, our approach provides rigorous formal guarantees on cybersickness risk under specified operational conditions. If these guarantees are not met, automated adaptive adjustments are recommended to restore safe conditions. Experimental validation involving physiological and systemlevel data demonstrates that Bayesian Networks provide an interpretable and efficient framework, uniquely enabling formal probabilistic verification of cybersickness risks. This capability makes the proposed approach particularly suitable for designing and deploying VR/MR systems with explicitly verified safety constraints. Peng Wu 0019, Nasim Ahmed, Abhiram Sarma, Kaiming Huang, Rifatul Islam, Bin Li 0014, Tian Lan 0001, Gang Tan, Mahdi Imani |
ISMAR | 8 |
| 2025 | Uncovering Discrimination Clusters: Quantifying and Explaining Systematic Fairness ViolationsabstractFairness in algorithmic decision-making is often framed in terms of individual fairness, which requires that similar individuals receive similar outcomes. A system violates individual fairness if there exists a pair of inputs differing only in protected attributes (such as race or gender) that lead to significantly different outcomes—for example, one favorable and the other unfavorable. While this notion highlights isolated instances of unfairness, it fails to capture broader patterns of clustered discrimination that may affect entire subgroups.We introduce and motivate the concept of discrimination clustering, a generalization of individual fairness violations. Rather than detecting single counterfactual disparities, we seek to uncover regions of the input space where small perturbations in protected features lead to k-significantly distinct clusters of outcomes. That is, for a given input, we identify a local neighborhood—differing only in protected attributes—whose members’ outputs separate into many distinct clusters. These clusters reveal significant arbitrariness in treatment solely based on protected attributes, exposing patterns of algorithmic bias that elude pairwise fairness checks.We present HyFair, a hybrid technique that combines formal symbolic analysis (via SMT and MILP solvers) to certify individual fairness with randomized search to discover discriminatory clusters. This combination enables both formal guarantees— when no counterexamples exist—and the detection of severe violations that are computationally challenging for symbolic methods alone. Given a set of inputs exhibiting high k-discrimination, we further introduce a novel explanation method that generates interpretable, decision-tree-style artifacts.Our experiments show that HyFair outperforms state-of-the-art fairness verification and local explanation methods. It reveals that some benchmarks exhibit substantial discrimination clustering, while others show limited or no disparities with respect to protected attributes. It also provides intuitive explanations that support understanding and mitigation of unfairness. Ranit Debnath Akash, Verya Monjezi, Ashutosh Trivedi 0001, Gang Tan, Saeid Tizpaz-Niari |
ASE | 5 |
| 2025 | Better Safe than Sorry: Preventing Policy Violations through Predictive Root-Cause-Analysis for IoT SystemsabstractIn an Internet of Things (IoT) environment, there are several way things can go wrong based on device activity. Poorly defined rules, conflicts between applications, physical interactions between devices, or unintentional interference by user behavior. Since these devices can have access to sensitive information or the capability to disrupt or harm physical elements in an environment, there is a strong motivation to protect confidentiality and integrity in IoT systems. In this paper we design IoTArmor, a novel Root-Cause-Analysis tool that uses machine learning models to select remediating actions that can prevent violations that would otherwise occur in the future. We assume violations have been predicted to occur and analyze the current system state to produce optimal fixes to prevent the violating behavior. Through this analysis, we can give accurate proposed fixes to prevent the violations, as well as detailed explanations to users as to why the fixes are effective. This methodology provides easily usable information to users about flaws in their environment, both in the current moment and in their overall application setup. Michael Norris, Syed Rafiul Hussain, Gang Tan |
ASE | 3 |
| 2025 | Demo: Perception Graph for Cognitive Attack Reasoning in Augmented RealityabstractAugmented reality (AR) systems are increasingly deployed in tactical environments, but their reliance on seamless human-computer interaction makes them vulnerable to cognitive attacks that manipulate a user's perception and severely compromise user decisionmaking. To address this challenge, we introduce the Perception Graph, a novel model designed to reason about human perception within these systems. Our model operates by first mimicking the human process of interpreting key information from an MR environment and then representing the outcomes using a semantically meaningful structure. We demonstrate how the model can compute a quantitative score that reflects the level of perception distortion, providing a robust and measurable method for detecting and analyzing the effects of such cognitive attacks. Shu Hong, Rifatul Islam, Mahdi Imani, Gang Tan, Tian Lan 0001 |
MobiHoc | 5 |
| 2025 | Poster: Time-Aware LSTM for Gaze Prediction in Mixed Reality Under Latency PerturbationsabstractCognitive attacks in mixed reality (MR), e.g., latency perturbations that induce frame-time jitter, can divert visual attention and degrade task performance. We study 2D gaze prediction under such disturbances and propose a time-aware sequence model that handles irregular sampling by supplying elapsed times Δt between observations and conditions on sparse event/object context available at prediction time via learned token embeddings. Using time-based windows, we evaluate within-user and cross-user temporal generalization on MR recordings spanning multiple attack intensities. Results indicate accurate, time-robust gaze regression under latency perturbations, supporting adaptive MR interfaces in adversarial settings. Shu Hong, Rifatul Islam, Mahdi Imani, Gang Tan, Tian Lan 0001 |
MobiHoc | 5 |
| 2025 | Validating Safety Guarantees of LSTM Models in MR ContextabstractEnsuring the safety of neural network (NN) models in mixed reality (MR) systems is challenging due to adversarial manipulation of system parameters. We present PolySafe, which extends DeepPoly and Prover to validate safety of LSTM-based MR models. PolySafe unrolls temporal dependencies, introduces multi-plane abstractions for tighter bounds, and establishes probabilistic safety guarantees. It further includes an adaptive search that identifies minimal sets of critical parameters required to be constrained for defense. Evaluation on an MR engagement prediction model shows that PolySafe provides rigorous and actionable safety assurances for deployment. Kaiming Huang, Peng Wu 0019, Mahdi Imani, Tian Lan 0001, Gang Tan |
MobiHoc | 5 |
| 2025 | Personalized Bayesian Networks for Cybersickness Prediction in Virtual RealityabstractPersonal characteristics fundamentally shape virtual reality (VR) experiences, yet their integration into predictive models remains underexplored. This paper studies how to incorporate personal attributes (age, gender, prior VR experience) into Bayesian networks for cybersickness prediction via: (i) direct inclusion as root nodes, (ii) a two-stage model that learns a susceptibility score from personal attributes, and (iii) a stratified model. Using 26,040 samples from VR maze-navigation experiments, direct inclusion attains 82.53% accuracy (+14.02 percentage points over a 68.51% no-personal baseline). The two-stage approach reaches 77.32% while supporting cold-start prediction for unseen users, and stratified models achieve 73.62%. Using participant-level cross-validation to avoid subject leakage, we find that personalization consistently improves cybersickness prediction. These results argue that personal attributes should be treated as first-class signals in cybersickness models, with clear design trade-offs between maximal accuracy and deployability for unseen users, informing personalized VR systems and adaptive content delivery. Peng Wu 0019, Nasim Ahmed, Kaiming Huang, Rifatul Islam, Tian Lan 0001, Gang Tan, Mahdi Imani |
MobiHoc | 6 |
| 2025 | SoK: Challenges and Paths Toward Memory Safety for eBPFabstractThe extended Berkeley Packet Filter (eBPF) subsystem in Linux enables the extension of kernel functionality without modifying kernel code. In addition to its use in networking, eBPF provides the flexibility to perform tracing, add security checks, etc. To ensure that eBPF does not enable attackers to compromise the kernel, eBPF includes a verifier to validate every eBPF program before its execution, which includes checks that aim to prevent eBPF programs from modifying kernel memory due to memory errors. However, numerous vulnerabilities have been identified in the eBPF subsystem, including the verifier itself, which greatly violate expectations, leading to concerns about the threats of memory safety brought by eBPF. This paper presents the first systematic analysis of the memory safety risks inherent in the eBPF ecosystem, focusing on the challenges faced by the limitations of the eBPF verifier and current kernel defenses. We then evaluate proposed research mitigation strategies that apply isolation techniques, runtime checks, and static validation, highlighting their contributions and gaps. Our study finds that only 1.62-3.74% (37–85) of the memory operations in public eBPF programs cannot be proven memory safe comprehensively, motivating actionable insights towards enforcing comprehensive memory safety while accounting for performance and compatibility. Kaiming Huang, Mathias Payer, Zhiyun Qian, Jack Sampson, Gang Tan, Trent Jaeger |
SP | 5 |
| 2025 | From a multi-period perspective: A periodic dynamics forecasting network for multivariate time series forecasting
Gang Tan, Ziyi Xiao, Dandan He, Guodong Sa |
Pattern Recognit. | 1 |
| 2025 | Sliver: A Scalable Slicing-Based Verification for Information Flow SecurityabstractStatic information flow analysis has been studied for a long time. It is usually considered more precise than dynamic taint analysis and more flexible and indispensable when running individual modules or the entire program is difficult. The state-of-the-art static information flow analyses are scalable on analyzing Java programs or mobile apps, and several type systems have enforced information flow security on different languages. However, static information-flow analyses have rarely scaled up to real-world C programs. This work presents Sliver, a slicing-based approach to verify information flow security on real-world C programs. The principle of Sliver is to convert the information-flow-involved parts of the original program into behavior-equivalent slices and use bounded model checking to enforce the end-to-end noninterference property or detect security violations on the slices after self-composition. We develop automated path-signature-guided slicing and adaptive self-composition approaches to ensure Sliver's efficacy and scalability. We also develop a consistency testing technique and metrics to estimate the correctness of slices generated by Sliver. The evaluations demonstrate Sliver's effectiveness, scalability, and the correctness of the generated slices. Xue Rao, Cong Sun 0001, Dongrui Zeng, Yongzhe Huang, Gang Tan |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | VIMU: Effective Physics-based Realtime Detection and Recovery against Stealthy Attacks on UAVsabstractSensor attacks on robotic vehicles have become pervasive and manipulative. Their latest advancements exploit sensor and detector characteristics to bypass detection. Recent security efforts have leveraged the physics-based model to detect or mitigate sensor attacks. However, these approaches are only resilient to a few sensor attacks and still need improvement in detection effectiveness. We present VIMU, an efficient sensor attack detection and resilience system for unmanned aerial vehicles. We propose a detection algorithm, CS-EMA, that leverages low-pass filtering to identify stealthy gyroscope attacks while achieving an overall effective sensor attack detection. We develop a fine-grained nonlinear physical model with precise aerodynamic and propulsion wrench modeling. We also augment the state estimation with a FIFO buffer safeguard to mitigate the impact of high-rate IMU attacks. The proposed physical model and buffer safeguard provide an effective system state recovery toward maintaining flight stability. We implement VIMU on PX4 autopilot. The evaluation results demonstrate the effectiveness of VIMU in detecting and mitigating various realistic sensor attacks, especially stealthy attacks. Yunbo Wang, Cong Sun 0001, Qiaosen Liu, Bingnan Su, Zongxu Zhang, Michael Norris, Gang Tan, Jianfeng Ma 0001 |
ACSAC | 7 |
| 2024 | Top of the Heap: Efficient Memory Error Protection of Safe Heap ObjectsabstractHeap memory errors remain a major source of software vulnerabilities. Existing memory safety defenses aim at protecting all objects, resulting in high performance cost and incomplete protection. Instead, we propose an approach that accurately identifies objects that are inexpensive to protect, and design a method to protect such objects comprehensively from all classes of memory errors. Towards this goal, we introduce the Uriah system that (1) statically identifies the heap objects whose accesses satisfy spatial and type safety, and (2) dynamically allocates such "safe" heap objects on an isolated safe heap to enforce a form of temporal safety while preserving spatial and type safety, called temporal allocated-type safety. Uriah finds 72.0% of heap allocation sites produce objects whose accesses always satisfy spatial and type safety in the SPEC CPU2006/2017 benchmarks, 5 server programs, and Firefox, which are then isolated on a safe heap using Uriah allocator to enforce temporal allocated-type safety. Uriah incurs only 2.9% and 2.6% runtime overhead, along with 9.3% and 5.4% memory overhead, on the SPEC CPU 2006 and 2017 benchmarks, while preventing exploits on all the heap memory errors in DARPA CGC binaries and 28 recent CVEs. Additionally, using existing defenses to enforce their memory safety guarantees on the unsafe heap objects significantly reduces overhead, enabling the protection of heap objects from all classes of memory errors at more practical costs. Kaiming Huang, Mathias Payer, Zhiyun Qian, Jack Sampson, Gang Tan, Trent Jaeger |
CCS | 5 |
| 2024 | NeuFair: Neural Network Fairness Repair with DropoutabstractThis paper investigates neuron dropout as a post-processing bias mitigation method for deep neural networks (DNNs). Neural-driven software solutions are increasingly applied in socially critical domains with significant fairness implications. While DNNs are exceptional at learning statistical patterns from data, they may encode and amplify historical biases. Existing bias mitigation algorithms often require modifying the input dataset or the learning algorithms. We posit that prevalent dropout methods may be an effective and less intrusive approach to improve fairness of pre-trained DNNs during inference. However, finding the ideal set of neurons to drop is a combinatorial problem. We propose NeuFair, a family of post-processing randomized algorithms that mitigate unfairness in pre-trained DNNs via dropouts during inference. Our randomized search is guided by an objective to minimize discrimination while maintaining the model’s utility. We show that NeuFair is efficient and effective in improving fairness (up to 69%) with minimal or no model performance degradation. We provide intuitive explanations of these phenomena and carefully examine the influence of various hyperparameters of NeuFair on the results. Finally, we empirically and conceptually compare NeuFair to different state-of-the-art bias mitigators. Vishnu Asutosh Dasu, Saeid Tizpaz-Niari, Gang Tan |
ISSTA | 4 |
| 2024 | Veiled Pathways: Investigating Covert and Side Channels Within GPU UncoreabstractWith the emergence of GPUs as first-class compute engines, more concentrated focus has been put into covert and side channel discovery in these architectures. However, most of the covert and side channels uncovered on GPUs to date are rooted in “GPU cores”, which include computational cores, cache and core interconnects, but they do not consider “GPU uncore”, which include non-computational engines, GPU DRAM, host-G PU links and inter-GPulinks. In this paper, we delve into the less-explored domains of GPU uncore, unveiling four novel leakage sources for covert and side channel exploitation: (1) GPU DRAM frequency scaling; (2) NVENC utilization; (3) NVDEC utilization; (4) NVJPEG utilization. What makes these covert and side channels interesting is that they all take effect under the GPU MPS mode - which fractionalizes GPU cores and GPU memory on both desktop-scale and server-scale GPUs. Furthermore, our study reevaluates PCI-e bandwidth allocation on GPUs. Notably, we have engineered covert and side channel capable of bypassing GPU MIG isolation - a mechanism implemented by NVIDIA to physically segregate hardware resources on server-scale GPUs. Our research showcases concrete examples of these covert and side channels, highlighting their potency in breaching system security, all achieved without necessitating root privileges. This underscores the practical implications and urgency of addressing these vulnerabilities in GPU architectures. Yuanqing Miao, Yingtian Zhang, Dinghao Wu, Danfeng Zhang, Gang Tan, Rui Zhang 0037, Mahmut T. Kandemir |
MICRO | 5 |
| 2024 | ORANalyst: Systematic Testing Framework for Open RAN Implementations
Tianchang Yang, Syed Md. Mukit Rashid, Ali Ranjbar, Gang Tan, Syed Rafiul Hussain |
USENIX Security Symposium | 4 |
| 2024 | SCML-GNN: A Graph Neural Network Model Leveraging Sensor Causality and Meta-Learning for Mechanical Fault ClassificationabstractFault classification of mechanical equipment is a vital issue in modern industrial production. Mechanical equipment typically relies on multiple sensors to collect the operational data, which is represented as multivariate time-series data. However, existing methods for analyzing mechanical faults often overlook the causal relationships between sensors and struggle with the scarcity of labeled training samples. To address these challenges, we propose a graph neural network model leveraging sensor causality and meta-learning for mechanical fault classification (SCML-GNN). Specifically, we use transfer entropy to represent multivariate time-series data as a graph, with each sensor as a node and their causal relationships as edges. We then extract the node features using temporal convolutional layers and apply a graph neural network to learn the low-dimensional features. Additionally, graph pooling methods are used to obtain global embeddings. To further tackle the issue of limited labeled training samples, we introduce a metric-based class prototype attention mechanism within SCML-GNN. Extensive experiments conducted on three real-world mechanical equipment datasets demonstrate the superior effectiveness and efficiency of SCML-GNN in mechanical fault classification compared to the other existing methods. Ziyi Xiao, Gang Tan, Dandan He, Wei Zhou 0028 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2024 | V-Star: Learning Visibly Pushdown Grammars from Program InputsabstractAccurate description of program inputs remains a critical challenge in the field of programming languages. Active learning, as a well-established field, achieves exact learning for regular languages. We offer an innovative grammar inference tool, V-Star, based on the active learning of visibly pushdown automata. V-Star deduces nesting structures of program input languages from sample inputs, employing a novel inference mechanism based on nested patterns. This mechanism identifies token boundaries and converts languages such as XML documents into VPLs. We then adapted Angluin’s L-Star, an exact learning algorithm, for VPA learning, which improves the precision of our tool. Our evaluation demonstrates that V-Star effectively and efficiently learns a variety of practical grammars, including S-Expressions, JSON, and XML, and outperforms other state-of-the-art tools. Xiaodong Jia 0004, Gang Tan |
Proc. ACM Program. Lang. | 2 |
| 2024 | psvCNN: A Zero-Knowledge CNN Prediction Integrity Verification StrategyabstractModel prediction based on machine learning is provided as a service in cloud environments, but how to verify that the model prediction service is entirely conducted becomes a critical challenge. Although zero-knowledge proof techniques potentially solve the integrity verification problem when applied to the prediction integrity of massive privacy-preserving Convolutional Neural Networks (CNNs), the significant proof burden results in low practicality. In this research, we present psvCNN (parallel splitting zero-knowledge technique for integrity verification). The psvCNN scheme effectively improves the utilization of computational resources in CNN prediction integrity, proving by an independent splitting design. Through a convolutional kernel-based model splitting design and an underlying zero-knowledge succinct non-interactive knowledge argument, our psvCNN develops parallelizable zero-knowledge proof circuits for CNN prediction. Furthermore, psvCNN presents an updated Freivalds algorithm for a faster integrity verification process. Experiments show that psvCNN is practical and efficient in terms of proof time and storage, generating a prediction integrity proof with a proof size of 1.2MB in 7.65s for the structurally complicated CNN model VGG16. psvCNN is 3765 times faster than the latest zk-SNARK-based non-interactive method vCNN, and 12 times faster than the latest sumcheck-based interactive technique zkCNN in terms of proving time. Yongkai Fan, Binyuan Xu, Linlin Zhang 0005, Gang Tan, Shui Yu 0001, Kuanching Li, Albert Y. Zomaya |
IEEE Trans. Cloud Comput. | 4 |
| 2024 | ValidCNN: A Large-Scale CNN Predictive Integrity Verification Scheme Based on zk-SNARKabstractThe integrity of cloud-based convolutional neural network (CNN) prediction services can be jeopardized by a malicious cloud server. Although zero-knowledge proof approaches can be used to verify integrity, they are difficult to use for larger CNN models like LeNet-5 and VGG16, due to the large cost (in terms of time and storage) of generating a proof. This paper proposes ValidCNN, which can efficiently generate integrity proofs based zk-SNARK. At the heart of ValidCNN, it is a novel usage of Freivald's concepts for circuit construction, and a more efficient way for verifying matrix multiplication. Our experimental results demonstrate that VaildCNN significantly outperforms the state-of-the-art approaches that are based on zk-SNARK. For example, compared with ZEN, VaildCNN achieves a 12-fold improvement in time and a 31-fold improvement in storage. Compared with vCNN, VaildCNN achieves a 195-fold and 279-fold improvement in time and storage respectively. Yongkai Fan, Kaile Ma, Linlin Zhang 0005, Guangquan Xu, Gang Tan |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2023 | Evolving Operating System Kernels Towards Secure Kernel-Driver InterfacesabstractOur work explores the challenge of developing secure kernel-driver interfaces designed to protect the kernel from isolated kernel extensions. We first analyze a range of possible attack vectors that exist in current isolation frameworks. Then, we suggest a new approach to building secure isolation boundaries centered around ideas that originate in safe operating systems: isolation of heaps and single ownership. Anton Burtsev, Vikram Narayanan, Yongzhe Huang, Kaiming Huang, Gang Tan, Trent Jaeger |
HotOS | 5 |
| 2023 | Information-Theoretic Testing and Debugging of Fairness Defects in Deep Neural NetworksabstractThe deep feedforward neural networks (DNNs) are increasingly deployed in socioeconomic critical decision support software systems. DNNs are exceptionally good at finding min-imal, sufficient statistical patterns within their training data. Consequently, DNNs may learn to encode decisions-amplifying existing biases or introducing new ones-that may disadvantage protected individuals/groups and may stand to violate legal protections. While the existing search based software testing approaches have been effective in discovering fairness defects, they do not supplement these defects with debugging aids-such as severity and causal explanations-crucial to help developers triage and decide on the next course of action. Can we measure the severity of fairness defects in DNNs? Are these defects symptomatic of improper training or they merely reflect biases present in the training data? To answer such questions, we present Dice: an information-theoretic testing and debugging framework to discover and localize fairness defects in DNNs. The key goal of Dice is to assist software developers in triaging fairness defects by ordering them by their severity. Towards this goal, we quantify fairness in terms of protected information (in bits) used in decision making. A quantitative view of fairness defects not only helps in ordering these defects, our empirical evaluation shows that it improves the search efficiency due to resulting smoothness of the search space. Guided by the quan-titative fairness, we present a causal debugging framework to localize inadequately trained layers and neurons responsible for fairness defects. Our experiments over ten DNNs, developed for socially critical tasks, show that Dice efficiently characterizes the amounts of discrimination, effectively generates discriminatory instances (vis-a-vis the state-of-the-art techniques), and localizes layers/neurons with significant biases. Verya Monjezi, Ashutosh Trivedi 0001, Gang Tan, Saeid Tizpaz-Niari |
ICSE | 3 |
| 2023 | Hardware Support for Constant-Time ProgrammingabstractSide-channel attacks are one of the rising security concerns in modern computing platforms. Observing this, researchers have proposed both hardware-based and software-based strategies to mitigate side-channel attacks, targeting not only on-chip caches but also other hardware components like memory controllers and on-chip networks. While hardware-based solutions to side-channel attacks are usually costly to implement as they require modifications to the underlying hardware, software-based solutions are more practical as they can work on unmodified hardware. One of the recent software-based solutions is constant-time programming, which tries to transform an input program to be protected against side-channel attacks such that an operation working on a data element/block to be protected would execute in an amount of time that is independent of the input. Unfortunately, while quite effective from a security angle, constant-time programming can lead to severe performance penalties. Yuanqing Miao, Mahmut T. Kandemir, Danfeng Zhang, Yingtian Zhang, Gang Tan, Dinghao Wu |
MICRO | 5 |
| 2023 | LibScan: Towards More Precise Third-Party Library Identification for Android Applications
Cong Sun 0001, Dongrui Zeng, Gang Tan, Siqi Ma 0001 |
USENIX Security Symposium | 4 |
| 2023 | CryptoEval: Evaluating the risk of cryptographic misuses in Android apps with data-flow analysisabstractAbstract The misunderstanding and incorrect configurations of cryptographic primitives have exposed severe security vulnerabilities to attackers. Due to the pervasiveness and diversity of cryptographic misuses, a comprehensive and accurate understanding of how cryptographic misuses can undermine the security of an Android app is critical to the subsequent mitigation strategies but also challenging. Although various approaches have been proposed to detect cryptographic misuse in Android apps, studies have yet to focus on estimating the security risks of cryptographic misuse. To address this problem, the authors present an extensible framework for deciding the threat level of cryptographic misuse in Android apps. Firstly, the authors propose a general and unified specification for representing cryptographic misuses to make our framework extensible and develop adapters to unify the detection results of the state‐of‐the‐art cryptographic misuse detectors, resulting in an adapter‐based detection tool chain for a more comprehensive list of cryptographic misuses. Secondly, the authors employ a misuse‐originating data‐flow analysis to connect each cryptographic misuse to a set of data‐flow sinks in an app, based on which the authors propose a quantitative data‐flow‐driven metric for assessing the overall risk of the app introduced by cryptographic misuses. To make the per‐app assessment more useful for app vetting at the app‐store level, the authors apply unsupervised learning to predict and classify the top risky threats to guide more efficient subsequent mitigation. In the experiments on an instantiated implementation of the framework, the authors evaluate the accuracy of our detection and the effect of data‐flow‐driven risk assessment of our framework. Our empirical study on over 40,000 apps, and the analysis of popular apps reveal important security observations on the real threats of cryptographic misuse in Android apps. Cong Sun 0001, Xinpeng Xu, Dongrui Zeng, Gang Tan, Siqi Ma 0001 |
IET Inf. Secur. | 5 |
| 2023 | ABSLearn: a GNN-based framework for aliasing and buffer-size information retrieval
Ke Liang 0006, Jim Tan, Dongrui Zeng, Yongzhe Huang, Gang Tan |
Pattern Anal. Appl. | 6 |
| 2023 | Quantifying and Mitigating Cache Side Channel Leakage with Differential SetabstractCache side-channel attacks leverage secret-dependent footprints in CPU cache to steal confidential information, such as encryption keys. Due to the lack of a proper abstraction for reasoning about cache side channels, existing static program analysis tools that can quantify or mitigate cache side channels are built on very different kinds of abstractions. As a consequence, it is hard to bridge advances in quantification and mitigation research. Moreover, existing abstractions lead to imprecise results. In this paper, we present a novel abstraction, called differential set, for analyzing cache side channels at compile time. A distinguishing feature of differential sets is that it allows compositional and precise reasoning about cache side channels. Moreover, it is the first abstraction that carries sufficient information for both side channel quantification and mitigation. Based on this new abstraction, we develop a static analysis tool DSA that automatically quantifies and mitigates cache side channel leakage at the same time. Experimental evaluation on a set of commonly used benchmarks shows that DSA can produce more precise leakage bound as well as mitigated code with fewer memory footprints, when compared with state-of-the-art tools that only quantify or mitigate cache side channel leakage. Cong Ma 0009, Dinghao Wu, Gang Tan, Mahmut T. Kandemir, Danfeng Zhang |
Proc. ACM Program. Lang. | 3 |
| 2023 | Interval Parsing Grammars for File Format ParsingabstractFile formats specify how data is encoded for persistent storage. They cannot be formalized as context-free grammars since their specifications include context-sensitive patterns such as the random access pattern and the type-length-value pattern. We propose a new grammar mechanism called Interval Parsing Grammars IPGs) for file format specifications. An IPG attaches to every nonterminal/terminal an interval, which specifies the range of input the nonterminal/terminal consumes. By connecting intervals and attributes, the context-sensitive patterns in file formats can be well handled. In this paper, we formalize IPGs' syntax as well as its semantics, and its semantics naturally leads to a parser generator that generates a recursive-descent parser from an IPG. In general, IPGs are declarative, modular, and enable termination checking. We have used IPGs to specify a number of file formats including ZIP, ELF, GIF, PE, and part of PDF; we have also evaluated the performance of the generated parsers. Jialun Zhang, J. Gregory Morrisett, Gang Tan |
Proc. ACM Program. Lang. | 3 |
| 2023 | μDep: Mutation-Based Dependency Generation for Precise Taint Analysis on Android Native CodeabstractThe existence of native code in Android apps plays an important role in triggering inconspicuous propagation of secrets and circumventing malware detection. However, the state-of-the-art information-flow analysis tools for Android apps all have limited capabilities of analyzing native code. Due to the complexity of binary-level static analysis, most static analyzers choose to build conservative models for a selected portion of native code. Though the recent inter-language analysis improves the capability of tracking information flow in native code, it is still far from attaining similar effectiveness of the state-of-the-art information-flow analyzers that focus on non-native Java methods. To overcome the above constraints, we propose a new analysis framework,$\mu$Dep, to detect sensitive information flows of the Android apps containing native code. In this framework, we combine a control-flow based static binary analysis with a mutation-based dynamic analysis to model the tainting behaviors of native code in the apps. Based on the result of the analyses,$\mu$Dep conducts a stub generation for the related native functions to facilitate the state-of-the-art analyzer DroidSafe with fine-grained tainting behavior summaries of native code. The experimental results show that our framework is competitive on the accuracy, and effective in analyzing the information flows in real-world apps and malware compared with the state-of-the-art inter-language static analysis. Cong Sun 0001, Yuwan Ma, Dongrui Zeng, Gang Tan, Siqi Ma 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | A Derivative-based Parser Generator for Visibly Pushdown GrammarsabstractIn this article, we present a derivative-based, functional recognizer and parser generator for visibly pushdown grammars. The generated parser accepts ambiguous grammars and produces a parse forest containing all valid parse trees for an input string in linear time. Each parse tree in the forest can then be extracted also in linear time. Besides the parser generator, to allow more flexible forms of the visibly pushdown grammars, we also present a translator that converts a tagged CFG to a visibly pushdown grammar in a sound way, and the parse trees of the tagged CFG are further produced by running the semantic actions embedded in the parse trees of the translated visibly pushdown grammar. The performance of the parser is compared with popular parsing tools, including ANTLR, GNU Bison, and other popular hand-crafted parsers. The correctness and the time complexity of the core parsing algorithm are formally verified in the proof assistant Coq. Xiaodong Jia 0004, Gang Tan |
ACM Trans. Program. Lang. Syst. | 3 |
| 2022 | BinPointer: towards precise, sound, and scalable binary-level pointer analysisabstractBinary-level pointer analysis is critical to binary-level applications such as reverse engineering and binary debloating. In this paper, we propose BinPointer, a new binary-level interprocedural pointer analysis that relies on an offset-sensitive value-tracking analysis to achieve high precision. We also propose a soundness and precision evaluation methodology based on runtime memory accesses triggered by reference input data. Our experimental results demonstrate that BinPointer has higher precision over prior work, while maintaining acceptable scalability. The soundness of BinPointer is also validated through runtime data. Sun Hyoung Kim, Dongrui Zeng, Cong Sun 0001, Gang Tan |
CC | 4 |
| 2022 | Fairness-aware Configuration of Machine Learning LibrariesabstractThis paper investigates the parameter space of machine learning (ML) algorithms in aggravating or mitigating fairness bugs. Data-driven software is increasingly applied in social-critical applications where ensuring fairness is of paramount importance. The existing approaches focus on addressing fairness bugs by either modifying the input dataset or modifying the learning algorithms. On the other hand, the selection of hyperparameters, which provide finer controls of ML algorithms, may enable a less intrusive approach to influence the fairness. Can hyperparameters amplify or suppress discrimination present in the input dataset? How can we help programmers in detecting, understanding, and exploiting the role of hyperparameters to improve the fairness? Saeid Tizpaz-Niari, Gang Tan, Ashutosh Trivedi 0001 |
ICSE | 3 |
| 2022 | ROS-SF: A Transparent and Efficient ROS Middleware using Serialization-Free MessageabstractIn recent years, ROS becomes the dominant middleware for robotic systems. The performance of its message-passing paradigm is crucial to the robot's reaction time. However, previous works only focus on efficiency, but ignore the requirement for transparency. We present ROS-SF framework, which can transparently eliminate serialization and de-serialization under the ROS APIs. The key contributions are a new serialization format called SFM and a life-cycle management method for serialization-free messages. Evaluation results show that our ROS-SF framework can improve the message-passing performance of ROS by up to 76.3\%. Application case study and applicability study show that our ROS-SF framework can be transparently applied to many existing ROS-based systems and packages. Even in the failure cases, our ROS-SF framework can provide modification guidance. Yu-Ping Wang 0001, Yue-Jiang Dong, Gang Tan |
Middleware | 3 |
| 2022 | The Taming of the Stack: Isolating Stack Data from Memory Errors
Kaiming Huang, Yongzhe Huang, Mathias Payer, Zhiyun Qian, Jack Sampson, Gang Tan, Trent Jaeger |
NDSS | 6 |
| 2022 | KSplit: Automating Device Driver Isolation
Yongzhe Huang, Vikram Narayanan, David Detweiler, Kaiming Huang, Gang Tan, Trent Jaeger, Anton Burtsev |
OSDI | 5 |
| 2022 | Building a Privacy-Preserving Smart Camera SystemabstractAbstract Millions of consumers depend on smart camera systems to remotely monitor their homes and businesses. However, the architecture and design of popular commercial systems require users to relinquish control of their data to untrusted third parties, such as service providers (e.g., the cloud). Third parties therefore can (and in some instances have) access the video footage without the users’ knowledge or consent—violating the core tenet of user privacy. In this paper, we present CaCTUs, a privacy-preserving smart Camera system Controlled Totally by Users. CaCTUs returns control to the user; the root of trust begins with the user and is maintained through a series of cryptographic protocols, designed to support popular features, such as sharing, deleting, and viewing videos live. We show that the system can support live streaming with a latency of 2 s at a frame rate of 10 fps and a resolution of 480 p. In so doing, we demonstrate that it is feasible to implement a performant smart-camera system that leverages the convenience of a cloud-based model while retaining the ability to control access to (private) data. Yohan Beugin, Quinn Burke 0002, Blaine Hoak, Ryan Sheatsley, Eric Pauley, Gang Tan, Syed Rafiul Hussain, Patrick D. McDaniel |
Proc. Priv. Enhancing Technol. | 6 |
| 2022 | TraceChain: A blockchain-based scheme to protect data confidentiality and traceabilityabstractSummary The risk of sharing data in cloud computing has gathered increasing attention. After the owner of some confidential data outsources the data to cloud storage services and shares it with others, the data owner lost the control to the data to a large extent. To achieve data sharing while keeping data confidentiality, attribute‐based encryption (ABE) can be employed by cloud storage services. However, ABE can only guarantee that outsourced data on the cloud is decrypted by attribute‐satisfying users but cannot restrict data from being accessed by dishonest users whose attributes also satisfy the access‐control policy. It is impossible for the data owner to control the shared data after it has been decrypted by dishonest users, especially when a set of attribute‐satisfying dishonest users may collude. To address this concern, we propose a traceable data sharing scheme called TraceChain. In TraceChain, data is encrypted over a new CP‐ABE scheme called E‐CP‐ABE. Furthermore, the system parameters for generating the private key in E‐CP‐ABE are uploaded to the private blockchain and transactions are performed on the chain. The data owner can obtain the identity of users by monitoring system parameters simultaneously and control data sharing on the blockchain. To prove the security of our scheme, the security analysis is given in this paper. Meanwhile, experimental results also show that our system is viable and efficient. Yongkai Fan, Wei Liang 0005, Gang Tan |
Softw. Pract. Exp. | 5 |
| 2022 | Artifact Suppression for the Joint Imaging of Primaries and Internal MultiplesabstractInternal multiples bring great challenges for conventional seismic data imaging and interpretation, in which primaries only are regarded as signals. It is of great importance to suppress the internal multiples before conventional imaging, which is computationally expensive. Considering that internal multiples are also real reflections from the underground interfaces and reverse time migration (RTM) can realize the imaging of multiples. Instead of suppression before imaging, we propose to do the joint imaging of primaries and internal multiples based on RTM. However, the artifacts caused by the internal multiples should be suppressed to guarantee the precision. First, the imaging condition based on up-going and down-going extrapolated wavefield separation is introduced to improve the imaging accuracy. Then, the generation mechanism of image artifacts is analyzed, it concludes that nonphysical wave paths caused by the missing boundary wavefields under the surface result in the infidelity reversed-time backward-extrapolation, which is the main reason for the artifacts. Next, we propose to apply the modeling boundary wavefields with a match based on pseudomultichannel matching filter to compensate the missing boundary wavefields, which can avoid the image artifacts and realize high-precision joint imaging. Finally, the feasibility and effectiveness of the proposed method are verified by numerical examples and field data. Zhina Li, Sikai Peng, Zhenchun Li, Yixuan Ding, Ning Qin, Gang Tan |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | IoTRepair: Flexible Fault Handling in Diverse IoT DeploymentsabstractIoT devices can be used to complete a wide array of physical tasks, but due to factors such as low computational resources and distributed physical deployment, they are susceptible to a wide array of faulty behaviors. Many devices deployed in homes, vehicles, industrial sites, and hospitals carry a great risk of damage to property, harm to a person, or breach of security if they behave faultily. We propose a general fault handling system named IoTRepair, which shows promising results for effectiveness with limited latency and power overhead in an IoT environment. IoTRepair dynamically organizes and customizes fault-handling techniques to address the unique problems associated with heterogeneous IoT deployments. We evaluate IoTRepair by creating a physical implementation mirroring a typical home environment to motivate the effectiveness of this system. Our evaluation showed that each of our fault-handling functions could be completed within 100 milliseconds after fault identification, which is a fraction of the time that state-of-the-art fault-identification methods take (measured in minutes). The power overhead is equally small, with the computation and device action consuming less than 30 milliwatts. This evaluation shows that IoTRepair not only can be deployed in a physical system, but offers significant benefits at a low overhead. Michael Norris, Z. Berkay Celik, Prasanna Venkatesh Rengasamy, Shulin Zhao 0001, Patrick D. McDaniel, Anand Sivasubramaniam, Gang Tan |
ACM Trans. Internet Things | 7 |
| 2021 | ReCFA: Resilient Control-Flow AttestationabstractRecent IoT applications gradually adapt more complicated end systems with commodity software. Ensuring the runtime integrity of these software is a challenging task for the remote controller or cloud services. Popular enforcement is the runtime remote attestation which requires the end system (prover) to generate evidence for its runtime behavior and a remote trusted verifier to attest the evidence. Control-flow attestation is a kind of runtime attestation that provides diagnoses towards the remote control-flow hijacking at the prover. Most of these attestation approaches focus on small or embedded software. The recent advance to attesting complicated software depends on the source code and CFG traversing to measure the checkpoint-separated subpaths, which may be unavailable for commodity software and cause possible context missing between consecutive subpaths in the measurements. Xinzhi Liu, Cong Sun 0001, Dongrui Zeng, Gang Tan, Xiao Kan, Siqi Ma 0001 |
ACSAC | 5 |
| 2021 | Ghost Thread: Effective User-Space Cache Side Channel ProtectionabstractCache-based side channel attacks pose a serious threat to computer security. Numerous cache attacks have been demonstrated, highlighting the need for effective and efficient defense mechanisms to shield systems from this threat. In this paper, we propose a novel application-level protection mechanism, called Ghost Thread. Ghost Thread is a flexible library that allows a user to protect cache accesses to a requested sensitive region to mitigate cache-based side channel attacks. This is accomplished by injecting random cache accesses to the sensitive cache region by separate threads. Compared with prior work that injects noise in a modified OS and hardware, our novel approach is applicable to commodity OS and hardware. Compared with other user-space mitigation mechanisms, our novel approach does not require any special hardware support, and it only requires slight code changes in the protected application making it readily deployable. Evaluation results on an Apache server show that Ghost Thread provides both strong protection and negligible overhead on real-world applications where only a fragment requires protection. In the worst-case scenario where the entire application requires protection, Ghost Thread still incurs negligible overhead when a system is under utilized, and moderate overhead when a system is fully utilized. Robert Brotzman, Danfeng Zhang, Mahmut T. Kandemir, Gang Tan |
CODASPY | 4 |
| 2021 | Refining Indirect Call Targets at the Binary Level
Sun Hyoung Kim, Cong Sun 0001, Dongrui Zeng, Gang Tan |
NDSS | 4 |
| 2021 | Sdft: A PDG-based Summarization for Efficient Dynamic Data Flow TrackingabstractDynamic taint analysis (DTA) has been widely used in various security-relevant scenarios that need to track the runtime information flow of programs. Dynamic binary instrumentation (DBI) is a prevalent technique in achieving effective dynamic taint tracking on commodity hardware and systems. However, the significant performance overhead incurred by dynamic taint analysis restricts its usage in production systems. Previous efforts on mitigating the performance penalty fall into two categories, parallelizing taint tracking from program execution and abstracting the tainting logic to a higher granularity. Both approaches have only met with limited success. In this work, we propose Sdft, an efficient approach that combines the precision of DBI-based instruction-level taint tracking and the efficiency of function-level abstract taint propagation. First, we build the library function summaries automatically with reachability analysis on the program dependency graph (PDG) to specify the control- and data dependencies between the input parameters, output parameters, and global variables of the target library. Then we derive the taint rules for the target library functions and develop taint tracking for library function that is tightly integrated into the state-of-the-art DTA framework Libdft. By applying our approach to the core C library functions of glibc, we report an average of 1.58x speed up of the tracking performance compared with Libdft64. We also validate the effectiveness of the hybrid taint tracking and the ability on detecting real-world vulnerabilities. Xiao Kan, Cong Sun 0001, Shen Liu 0002, Yongzhe Huang, Gang Tan, Siqi Ma 0001 |
QRS | 5 |
| 2021 | iTOP: Automating Counterfeit Object-Oriented Programming AttacksabstractExploiting a program requires a security analyst to manipulate data in program memory with the goal to obtain control over the program counter and to escalate privileges. However, this is a tedious and lengthy process as: (1) the analyst has to massage program data such that a logical reliable data passing chain can be established, and (2) depending on the attacker goal certain in-place fine-grained protection mechanisms need to be bypassed. Previous work has proposed various techniques to facilitate exploit development. Unfortunately, none of them can be easily used to address the given challenges. This is due to the fact that data in memory is difficult to be massaged by an analyst who does not know the peculiarities of the program as the attack specification is most of the time only textually available, and not automated at all. Paul Muntean 0001, Richard Viehoever, Zhiqiang Lin 0001, Gang Tan, Jens Grossklags, Claudia Eckert 0001 |
RAID | 4 |
| 2021 | MazeRunner: Evaluating the Attack Surface of Control-Flow Integrity PoliciesabstractControl-Flow Integrity (CFI) enforces a control-flow graph (CFG) to limit attackers' ability to manipulate runtime control flow. CFI variations, enforcing different CFGs, achieve different degrees of attack surface reduction. To compare the security strength of different CFI policies, measuring the remaining attack surface is critical but challenging. Therefore, we propose MazeRunner, a framework that quantitatively estimates the attack surface of a CFI-hardened program. Methodology-wise, it takes a program's CFG, an attack model, and a security-violation policy as input to discover risky program points by an attack-aware data dependency tracking algorithm. Risky program points and the CFG are used to compute a metric for the remaining attack surface. We evaluate MazeRunner with 3 CFG types, 3 attack models, and 4 security-violation policies against 13 realistic benchmarks, and demonstrate that the new metric achieves higher precision than traditional metrics while maintaining completeness. Dongrui Zeng, Ben Niu 0007, Gang Tan |
TrustCom | 3 |
| 2021 | PPMCK: Privacy-preserving multi-party computing for K-means clustering
Yongkai Fan, Jianrong Bai, Weiguo Lin, Guodong Wu, Jiaming Guo, Gang Tan |
J. Parallel Distributed Comput. | 8 |
| 2021 | SpecSafe: detecting cache side channels in a speculative worldabstractThe high-profile Spectre attack and its variants have revealed that speculative execution may leave secret-dependent footprints in the cache, allowing an attacker to learn confidential data. However, existing static side-channel detectors either ignore speculative execution, leading to false negatives, or lack a precise cache model, leading to false positives. In this paper, somewhat surprisingly, we show that it is challenging to develop a speculation-aware static analysis with precise cache models: a combination of existing works does not necessarily catch all cache side channels. Motivated by this observation, we present a new semantic definition of security against cache-based side-channel attacks, called Speculative-Aware noninterference (SANI), which is applicable to a variety of attacks and cache models. We also develop SpecSafe to detect the violations of SANI. Unlike other speculation-aware symbolic executors, SpecSafe employs a novel program transformation so that SANI can be soundly checked by speculation-unaware side-channel detectors. SpecSafe is shown to be both scalable and accurate on a set of moderately sized benchmarks, including commonly used cryptography libraries. Robert Brotzman, Danfeng Zhang, Mahmut T. Kandemir, Gang Tan |
Proc. ACM Program. Lang. | 4 |
| 2021 | A derivative-based parser generator for visibly Pushdown grammarsabstractIn this paper, we present a derivative-based, functional recognizer and parser generator for visibly pushdown grammars. The generated parser accepts ambiguous grammars and produces a parse forest containing all valid parse trees for an input string in linear time. Each parse tree in the forest can then be extracted also in linear time. Besides the parser generator, to allow more flexible forms of the visibly pushdown grammars, we also present a translator that converts a tagged CFG to a visibly pushdown grammar in a sound way, and the parse trees of the tagged CFG are further produced by running the semantic actions embedded in the parse trees of the translated visibly pushdown grammar. The performance of the parser is compared with a popular parsing tool ANTLR and other popular hand-crafted parsers. The correctness of the core parsing algorithm is formally verified in the proof assistant Coq. Xiaodong Jia 0004, Gang Tan |
Proc. ACM Program. Lang. | 3 |
| 2021 | SPX64: A Scratchpad Memory for General-purpose MicroprocessorsabstractGeneral-purpose computing systems employ memory hierarchies to provide the appearance of a single large, fast, coherent memory. In special-purpose CPUs, programmers manually manage distinct, non-coherent scratchpad memories. In this article, we combine these mechanisms by adding a virtually addressed, set-associative scratchpad to a general purpose CPU. Our scratchpad exists alongside a traditional cache and is able to avoid many of the programming challenges associated with traditional scratchpads without sacrificing generality (e.g., virtualization). Furthermore, our design delivers increased security and improves performance, especially for workloads with high locality or that interact with nonvolatile memory. Shail Dave, Pantea Zardoshti, Robert Brotzman, Chao Zhang 0039, Aviral Shrivastava, Gang Tan, Michael F. Spear |
ACM Trans. Archit. Code Optim. | 8 |
| 2020 | ρFEM: Efficient Backward-edge Protection Using Reversed Forward-edge MappingsabstractIn this paper, we propose reversed forward-edge mapper (ρFEM), a Clang/LLVM compiler-based tool, to protect the backward edges of a program’s control flow graph (CFG) against runtime control-flow hijacking (e.g., code reuse attacks). It protects backward-edge transfers in C/C++ originating from virtual and non-virtual functions by first statically constructing a precise virtual table hierarchy, with which to form a precise forward-edge mapping between callees and non-virtual calltargets based on precise function signatures, and then checks each instrumented callee return against the previously computed set at runtime. We have evaluated ρFEM using the Chrome browser, NodeJS, Nginx, Memcached, and the SPEC CPU2017 benchmark. Our results show that ρFEM enforces less than 2.77 return targets per callee in geomean, even for applications heavily relying on backward edges. ρFEM’s runtime overhead is less than 1% in geomean for the SPEC CPU2017 benchmark and 3.44% in geomean for the Chrome browser. Paul Muntean 0001, Matthias Neumayer, Zhiqiang Lin 0001, Gang Tan, Jens Grossklags, Claudia Eckert 0001 |
ACSAC | 4 |
| 2020 | Methodologies for Quantifying (Re-)randomization Security and Timing under JIT-ROPabstractJust-in-time return-oriented programming (JIT-ROP) allows one to dynamically discover instruction pages and launch code reuse attacks, effectively bypassing most fine-grained address space layout randomization (ASLR) protection. However, in-depth questions regarding the impact of code (re-)randomization on code reuse attacks have not been studied. For example, how would one compute the re-randomization interval effectively by considering the speed of gadget convergence to defeat JIT-ROP attacks? ; how do starting pointers in JIT-ROP impact gadget availability and gadget convergence time? ; what impact do fine-grained code randomizations have on the Turing-complete expressive power of JIT-ROP payloads? We conduct a comprehensive measurement study on the effectiveness of fine-grained code randomization schemes, with 5 tools, 20 applications including 6 browsers, 1 browser engine, and 25 dynamic libraries. We provide methodologies to measure JIT-ROP gadget availability, quality, and their Turing-complete expressiveness, as well as to empirically determine the upper bound of re-randomization intervals in re-randomization schemes using the Turing-complete (TC), priority, MOV TC, and payload gadget sets. Experiments show that the upper bound ranges from 1.5 to 3.5 seconds in our tested applications. Besides, our results show that locations of leaked pointers used in JIT-ROP attacks have no impacts on gadget availability but have an impact on how fast attackers find gadgets. Our results also show that instruction-level single-round randomization thwarts current gadget finding techniques under the JIT-ROP threat model. Salman Ahmed 0001, Ya Xiao 0002, Kevin Z. Snow, Gang Tan, Fabian Monrose, Danfeng Yao |
CCS | 4 |
| 2020 | Lightweight kernel isolation with virtualization and VM functionsabstractCommodity operating systems execute core kernel subsystems in a single address space along with hundreds of dynamically loaded extensions and device drivers. Lack of isolation within the kernel implies that a vulnerability in any of the kernel subsystems or device drivers opens a way to mount a successful attack on the entire kernel. Vikram Narayanan, Yongzhe Huang, Gang Tan, Trent Jaeger, Anton Burtsev |
VEE | 3 |
| 2020 | Prioritizing data flows and sinks for app security transformation
Ke Tian, Gang Tan, Barbara G. Ryder, Danfeng Yao |
Comput. Secur. | 2 |
| 2020 | Fine-grained access control based on Trusted Execution Environment
Yongkai Fan, Shengle Liu, Gang Tan, Fei Qiao |
Future Gener. Comput. Syst. | 3 |
| 2020 | Privacy preserving based logistic regression on big data
Yongkai Fan, Jianrong Bai, Yuqing Zhang 0001, Bin Zhang 0008, Kuanching Li, Gang Tan |
J. Netw. Comput. Appl. | 7 |
| 2020 | Detection of Repackaged Android Malware with Code-Heterogeneity FeaturesabstractDuring repackaging, malware writers statically inject malcode and modify the control flow to ensure its execution. Repackaged malware is difficult to detect by existing classification techniques, partly because of their behavioral similarities to benign apps. By exploring the app's internal different behaviors, we propose a new Android repackaged malware detection technique based on code heterogeneity analysis. Our solution strategically partitions the code structure of an app into multiple dependence-based regions (subsets of the code). Each region is independently classified on its behavioral features. We point out the security challenges and design choices for partitioning code structures at the class and method level graphs, and present a solution based on multiple dependence relations. We have performed experimental evaluation with over 7,542 Android apps. For repackaged malware, our partition-based detection reduces false negatives (i.e., missed detection) by 30-fold, when compared to the non-partition-based approach. Overall, our approach achieves a false negative rate of 0.35 percent and a false positive rate of 2.97 percent. Ke Tian, Danfeng Yao, Barbara G. Ryder, Gang Tan, Guojun Peng |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2019 | Analyzing control flow integrity with LLVM-CFIabstractControl-flow hijacking attacks are used to perform malicious computations. Current solutions for assessing the attack surface after a control flow integrity (CFI) policy was applied can measure only indirect transfer averages in the best case without providing any insights w.r.t. the absolute calltarget reduction per callsite, and gadget availability. Further, tool comparison is underdeveloped or not possible at all. CFI has proven to be one of the most promising protections against control flow hijacking attacks, thus many efforts have been made to improve CFI in various ways. However, there is a lack of systematic assessment of existing CFI protections. Paul Muntean 0001, Matthias Neumayer, Zhiqiang Lin 0001, Gang Tan, Jens Grossklags, Claudia Eckert 0001 |
ACSAC | 4 |
| 2019 | Program-mandering: Quantitative Privilege SeparationabstractPrivilege separation is an effective technique to improve software security. However, past partitioning systems do not allow programmers to make quantitative tradeoffs between security and performance. In this paper, we describe our toolchain called PM. It can automatically find the optimal boundary in program partitioning. This is achieved by solving an integer-programming model that optimizes for a user-chosen metric while satisfying the remaining security and performance constraints on other metrics. We choose security metrics to reason about how well computed partitions enforce information flow control to: (1) protect the program from low-integrity inputs or (2) prevent leakage of program secrets. As a result, functions in the sensitive module that fall on the optimal partition boundaries automatically identify where declassification is necessary. We used PM to experiment on a set of real-world programs to protect confidentiality and integrity; results show that, with moderate user guidance, PM can find partitions that have better balance between security and performance than partitions found by a previous tool that requires manual declassification. Shen Liu 0002, Dongrui Zeng, Yongzhe Huang, Frank Capobianco, Stephen McCamant, Trent Jaeger, Gang Tan |
CCS | 7 |
| 2019 | IoTGuard: Dynamic Enforcement of Security and Safety Policy in Commodity IoT
Z. Berkay Celik, Gang Tan, Patrick D. McDaniel |
NDSS | 2 |
| 2019 | Using Safety Properties to Generate Vulnerability PatchesabstractSecurity vulnerabilities are among the most critical software defects in existence. When identified, programmers aim to produce patches that prevent the vulnerability as quickly as possible, motivating the need for automatic program repair (APR) methods to generate patches automatically. Unfortunately, most current APR methods fall short because they approximate the properties necessary to prevent the vulnerability using examples. Approximations result in patches that either do not fix the vulnerability comprehensively, or may even introduce new bugs. Instead, we propose property-based APR, which uses human-specified, program-independent and vulnerability-specific safety properties to derive source code patches for security vulnerabilities. Unlike properties that are approximated by observing the execution of test cases, such safety properties are precise and complete. The primary challenge lies in mapping such safety properties into source code patches that can be instantiated into an existing program. To address these challenges, we propose Senx, which, given a set of safety properties and a single input that triggers the vulnerability, detects the safety property violated by the vulnerability input and generates a corresponding patch that enforces the safety property and thus, removes the vulnerability. Senx solves several challenges with property-based APR: it identifies the program expressions and variables that must be evaluated to check safety properties and identifies the program scopes where they can be evaluated, it generates new code to selectively compute the values it needs if calling existing program code would cause unwanted side effects, and it uses a novel access range analysis technique to avoid placing patches inside loops where it could incur performance overhead. Our evaluation shows that the patches generated by Senx successfully fix 32 of 42 real-world vulnerabilities from 11 applications including various tools or libraries for manipulating graphics/media files, a programming language interpreter, a relational database engine, a collection of programming tools for creating and managing binary programs, and a collection of basic file, shell, and text manipulation tools. Zhen Huang 0002, David Lie, Gang Tan, Trent Jaeger |
IEEE Symposium on Security and Privacy | 3 |
| 2019 | CaSym: Cache Aware Symbolic Execution for Side Channel Detection and MitigationabstractCache-based side channels are becoming an important attack vector through which secret information can be leaked to malicious parties. implementations and Previous work on cache-based side channel detection, however, suffers from the code coverage problem or does not provide diagnostic information that is crucial for applying mitigation techniques to vulnerable software. We propose CaSym, a cache-aware symbolic execution to identify and report precise information about where side channels occur in an input program. Compared with existing work, CaSym provides several unique features: (1) CaSym enables verification against various attack models and cache models, (2) unlike many symbolic-execution systems for bug finding, CaSym verifies all program execution paths in a sound way, (3) CaSym uses two novel abstract cache models that provide good balance between analysis scalability and precision, and (4) CaSym provides sufficient information on where and how to mitigate the identified side channels through techniques including preloading and pinning. Evaluation on a set of crypto and database benchmarks shows that CaSym is effective at identifying and mitigating side channels, with reasonable efficiency. Robert Brotzman, Shen Liu 0002, Danfeng Zhang, Gang Tan, Mahmut T. Kandemir |
IEEE Symposium on Security and Privacy | 4 |
| 2019 | A secure privacy preserving deduplication scheme for cloud computing
Yongkai Fan, Wei Liang 0005, Gang Tan, Priyadarsi Nanda |
Future Gener. Comput. Syst. | 4 |
| 2019 | One secure data integrity verification scheme for cloud storage
Yongkai Fan, Gang Tan, Yuqing Zhang 0001 |
Future Gener. Comput. Syst. | 3 |
| 2019 | SmartShell: Automated Shell Scripts Synthesis from Natural LanguageabstractModern shell scripts provide interfaces with rich functionality for system administration. However, it is not easy for end-users to write correct shell scripts; misusing commands may cause unpredictable results. In this paper, we present SmartShell, an automated function-based tool for shell script synthesis, which uses natural language descriptions as input. It can help the computer system to “understand” users’ intentions. SmartShell is based on two insights: (1) natural language descriptions for system objects (such as files and processes) and operations can be recognized by natural language processing tools; (2) system-administration tasks are often completed by short shell scripts that can be automatically synthesized from natural language descriptions. SmartShell synthesizes shell scripts in three steps: (1) using natural language processing tools to convert the description of a system-administration task into a syntax tree; (2) using program-synthesis techniques to construct a SmartShell intermediate-language script from the syntax tree; (3) translating the intermediate-language script into a shell script. Experimental results show that SmartShell can successfully synthesize 53.7% of tasks collected from shell-script helping forums. Hao Li 0056, Yu-Ping Wang 0001, Gang Tan |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2019 | IVT: an efficient method for sharing subtype polymorphic objectsabstractShared memory provides the fastest form of inter-process communication. Sharing polymorphic objects between different address spaces requires solving the issue of sharing pointers. In this paper, we propose a method, named Indexed Virtual Tables (IVT for short), to share polymorphic objects efficiently. On object construction, the virtual table pointers are replaced with indexes, which are used to find the actual virtual table pointers on dynamic dispatch. Only a few addition and load instructions are needed for both operations. Experimental results show that the IVT can outperform prior techniques on both object construction time and dynamic dispatch time. We also apply the proposed IVT technique to several practical scenarios, resulting the improvement of overall performance. Yu-Ping Wang 0001, Xu-Qiang Hu, Zixin Zou, Wende Tan, Gang Tan |
Proc. ACM Program. Lang. | 5 |
| 2019 | Debugopt: Debugging fully optimized natively compiled programs using multistage instrumentation
Gang Tan, Hao Li 0056, Xiaolong Bai, Yu-Ping Wang 0001, Shi-Min Hu 0001 |
Sci. Comput. Program. | 2 |
| 2018 | From Debugging-Information Based Binary-Level Type Inference to CFG GenerationabstractBinary-level Control-Flow Graph (CFG) construction is essential for applications such as control-flow integrity. There are two main approaches: the binary-analysis approach and the compiler-modification approach. The binary-analysis approach does not require source code, but it constructs low-precision CFGs. The compiler-modification approach requires source code and modifies compilers for CFG generation. We describe the design and implementation of an alternative system for high-precision CFG construction, which still assumes source code but does not modify compilers. Our approach makes use of standard compiler-generated meta-information, including symbol tables, relocation information, and debugging information. A key component in the system is a type-inference engine that infers types of low-level storage locations such as registers from types in debugging information. Inferred types enable a type-signature matching method for high-precision CFG construction. Dongrui Zeng, Gang Tan |
CODASPY | 2 |
| 2018 | τCFI: Type-Assisted Control Flow Integrity for x86-64 Binaries
Paul Muntean 0001, Gang Tan, Zhiqiang Lin 0001, Jens Grossklags, Claudia Eckert 0001 |
RAID | 3 |
| 2018 | Soteria: Automated IoT Safety and Security Analysis
Z. Berkay Celik, Patrick D. McDaniel, Gang Tan |
USENIX ATC | 3 |
| 2018 | Sensitive Information Tracking in Commodity IoT
Z. Berkay Celik, Leonardo Babun, Amit Kumar Sikder, Hidayet Aksu, Gang Tan, Patrick D. McDaniel, A. Selcuk Uluagac |
USENIX Security Symposium | 5 |
| 2018 | Bidirectional Grammars for Machine-Code Decoding and Encoding
Gang Tan, J. Gregory Morrisett |
J. Autom. Reason. | 1 |
| 2017 | PtrSplit: Supporting General Pointers in Automatic Program PartitioningabstractPartitioning a security-sensitive application into least-privileged components and putting each into a separate protection domain have long been a goal of security practitioners and researchers. However, a stumbling block to automatically partitioning C/C++ applications is the presence of pointers in these applications. Pointers make calculating data dependence, a key step in program partitioning, difficult and hard to scale; furthermore, C/C++ pointers do not carry bounds information, making it impossible to automatically marshall and unmarshall pointer data when they are sent across the boundary of partitions. In this paper, we propose a set of techniques for supporting general pointers in automatic program partitioning. Our system, called PtrSplit, constructs a Program Dependence Graph (PDG) for tracking data and control dependencies in the input program and employs a parameter-tree approach for representing data of pointer types; this approach is modular and avoids global pointer analysis. Furthermore, it performs selective pointer bounds tracking to enable automatic marshalling/unmarshalling of pointer data, even when there is circularity and arbitrary aliasing. As a result, PtrSplit can automatically generate executable partitions for C applications that contain arbitrary pointers. Shen Liu 0002, Gang Tan, Trent Jaeger |
CCS | 2 |
| 2015 | Per-Input Control-Flow IntegrityabstractControl-Flow Integrity (CFI) is an effective approach to mitigating control-flow hijacking attacks. Conventional CFI techniques statically extract a control-flow graph (CFG) from a program and instrument the program to enforce that CFG. The statically generated CFG includes all edges for all possible inputs; however, for a concrete input, the CFG may include many unnecessary edges. Ben Niu 0007, Gang Tan |
CCS | 2 |
| 2015 | WebC: toward a portable framework for deploying legacy code in web browsers
Gang Tan, Xiaolong Bai, Shi-Min Hu 0001 |
Sci. China Inf. Sci. | 2 |
| 2015 | JNI light: an operational model for the core JNIabstractThrough foreign function interfaces (FFIs), software components in different programming languages interact with each other in the same address space. Recent years have witnessed a number of systems that analyse FFIs for safety and reliability. However, lack of formal specifications of FFIs hampers progress in this endeavour. We present a formal operational model, Java Native Interface (JNI) light (JNIL), for a subset of a widely used FFI – the Java Native Interface (JNI). JNIL focuses on the core issues when a high-level garbage-collected language interacts with a low-level language. It proposes abstractions for handling a shared heap, cross-language method calls, cross-language exception handling, and garbage collection. JNIL can directly serve as a formal basis for JNI tools and systems. We demonstrate its utility by proving soundness of a system that checks native code in JNI programs for type-unsafe use of JNI functions. The abstractions in JNIL are also useful when modelling other FFIs, such as the Python/C interface and the OCaml/C interface. Gang Tan |
Math. Struct. Comput. Sci. | 1 |
| 2014 | RockJIT: Securing Just-In-Time Compilation Using Modular Control-Flow IntegrityabstractManaged languages such as JavaScript are popular. For performance, modern implementations of managed languages adopt Just-In-Time (JIT) compilation. The danger to a JIT compiler is that an attacker can often control the input program and use it to trigger a vulnerability in the JIT compiler to launch code injection or JIT spraying attacks. In this paper, we propose a general approach called RockJIT to securing JIT compilers through Control-Flow Integrity (CFI). RockJIT builds a fine-grained control-flow graph from the source code of the JIT compiler and dynamically updates the control-flow policy when new code is generated on the fly. Through evaluation on Google's V8 JavaScript engine, we demonstrate that RockJIT can enforce strong security on a JIT compiler, while incurring only modest performance overhead (14.6% on V8) and requiring a small amount of changes to V8's code. Key contributions of RockJIT are a general architecture for securing JIT compilers and a method for generating fine-grained control-flow graphs from C++ code. Ben Niu 0007, Gang Tan |
CCS | 2 |
| 2014 | Finding Reference-Counting Errors in Python/C Programs with Affine Analysis
Siliang Li, Gang Tan |
ECOOP | 2 |
| 2014 | Modular control-flow integrityabstractControl-Flow Integrity (CFI) is a software-hardening technique. It inlines checks into a program so that its execution always follows a predetermined Control-Flow Graph (CFG). As a result, CFI is effective at preventing control-flow hijacking attacks. However, past fine-grained CFI implementations do not support separate compilation, which hinders its adoption. Ben Niu 0007, Gang Tan |
PLDI | 2 |
| 2014 | NativeGuard: protecting android applications from third-party native librariesabstractAndroid applications often include third-party libraries written in native code. However, current native components are not well managed by Android's security architecture. We present NativeGuard, a security framework that isolates native libraries from other components in Android applications. Leveraging the process-based protection in Android, NativeGuard isolates native libraries of an Android application into a second application where unnecessary privileges are eliminated. NativeGuard requires neither modifications to Android nor access to the source code of an application. It addresses multiple technical issues to support various interfaces that Android provides to the native world. Experimental results demonstrate that our framework works well with a set of real-world applications, and incurs only modest overhead on benchmark programs. Mengtao Sun, Gang Tan |
WISEC | 2 |
| 2014 | Exception analysis in the Java Native Interface
Siliang Li, Gang Tan |
Sci. Comput. Program. | 2 |
| 2013 | Efficient user-space information flow controlabstractThe model of Decentralized Information Flow Control (DIFC) is effective at improving application security and can support rich confidentiality and integrity policies. We describe the design and implementation of duPro, an efficient user-space information flow control framework. duPro adopts Software-based Fault Isolation (SFI) to isolate protection domains within the same process. It controls the end-to-end information flow at the granularity of SFI domains. Being a user-space framework, duPro does not require any OS changes. Since SFI is more lightweight than hardware-based isolation (e.g., OS processes), the inter-domain communication and scheduling in duPro are more efficient than process-level DIFC systems. Finally, duPro supports a novel checkpointing-restoration mechanism for efficiently reusing protection domains. Experiments demonstrate applications can be ported to duPro with negligible overhead, enhanced security, and with tight control over information flow. Ben Niu 0007, Gang Tan |
AsiaCCS | 2 |
| 2013 | Monitor integrity protection with space efficiency and separate compilationabstractLow-level inlined reference monitors weave monitor code into a program for security. To ensure that monitor code cannot be bypassed by branching instructions, some form of control-flow integrity must be guaranteed. Past approaches to protecting monitor code either have high space overhead or do not support separate compilation. We present Monitor Integrity Protection (MIP), a form of coarse-grained control-flow integrity. The key idea of MIP is to arrange instructions in variable-sized chunks and dynamically restrict indirect branches to target only chunk beginnings. We show that this simple idea is effective in protecting monitor code integrity, enjoys low space and execution-time overhead, supports separate compilation, and is largely compatible with an existing compiler toolchain. We also show that MIP enables a separate verifier that completely disassembles a binary and verifies its security. MIP is designed to support inlined reference monitors. As a case study, we have implemented MIP-based Software-based Fault Isolation (SFI) on both x86-32 and x86-64. The evaluation shows that MIP-based SFI has competitive performance with other SFI implementations, while enjoying low space overhead. Ben Niu 0007, Gang Tan |
CCS | 2 |
| 2013 | Strato: A Retargetable Framework for Low-Level Inlined-Reference Monitors
Bin Zeng 0004, Gang Tan, Úlfar Erlingsson |
USENIX Security Symposium | 2 |
| 2013 | Bringing java's wild native world under controlabstractFor performance and for incorporating legacy libraries, many Java applications contain native-code components written in unsafe languages such as C and C++. Native-code components interoperate with Java components through the Java Native Interface (JNI). As native code is not regulated by Java's security model, it poses serious security threats to the managed Java world. We introduce a security framework that extends Java's security model and brings native code under control. Leveraging software-based fault isolation, the framework puts native code in a separate sandbox and allows the interaction between the native world and the Java world only through a carefully designed pathway. Two different implementations were built. In one implementation, the security framework is integrated into a Java Virtual Machine (JVM). In the second implementation, the framework is built outside of the JVM and takes advantage of JVM-independent interfaces. The second implementation provides JVM portability, at the expense of some performance degradation. Evaluation of our framework demonstrates that it incurs modest runtime overhead while significantly enhancing the security of Java applications. Mengtao Sun, Gang Tan, Joseph Siefers, Bin Zeng 0004, J. Gregory Morrisett |
ACM Trans. Inf. Syst. Secur. | 2 |
| 2012 | JATO: Native Code Atomicity for Java
Siliang Li, Yu David Liu, Gang Tan |
APLAS | 3 |
| 2012 | JVM-Portable Sandboxing of Java's Native Libraries
Mengtao Sun, Gang Tan |
ESORICS | 2 |
| 2012 | Smartphone Dual Defense Protection Framework: Detecting Malicious Applications in Android MarketsabstractIn this paper, we present a smart phone dual defense protection framework that allows Official and Alternative Android Markets to detect malicious applications among those new applications that are submitted for public release. Our framework consists of servers running on clouds where developers who wish to release their new applications can upload their software for verification purpose. The verification server first uses system call statistics to identify potential malicious applications. After verification, if the software is clean, the application will then be released to the relevant markets. To mitigate against false negative cases, users who run new applications can invoke our network traffic monitoring (NTM)tool which triggers network traffic capture upon detecting some suspicious behaviors e.g. detecting sensitive data being sent to output stream of an open socket. The network traffic will be analyzed to see if it matches network characteristics observed from malware applications. If suspicious network traffic is observed, the relevant Android markets will be notified tore move the application from the repository. We trained our system call and network traffic classifiers using 32 families of known Android malware families and some typical normal applications. Later, we evaluated our framework using other malware and normal applications that used in the training set. Our experimental results using 120 test applications (which consist of 50 malware and 70 normal applications) indicate that we can achieve a 94.2% and 99.2% accuracy with J.48 and Random forest classifier respectively using our framework. X. Su, Gang Tan |
MSN | 3 |
| 2012 | RockSalt: better, faster, stronger SFI for the x86abstractSoftware-based fault isolation (SFI), as used in Google's Native Client (NaCl), relies upon a conceptually simple machine-code analysis to enforce a security policy. But for complicated architectures such as the x86, it is all too easy to get the details of the analysis wrong. We have built a new checker that is smaller, faster, and has a much reduced trusted computing base when compared to Google's original analysis. The key to our approach is automatically generating the bulk of the analysis from a declarative description which we relate to a formal model of a subset of the x86 instruction set architecture. The x86 model, developed in Coq, is of independent interest and should be usable for a wide range of machine-level verification tasks. J. Gregory Morrisett, Gang Tan, Joseph Tassarotti, Jean-Baptiste Tristan, Edward Gan |
PLDI | 2 |
| 2011 | Detection and Classification of Different Botnet C&C Channels
Gregory Fedynyshyn, Mooi Choo Chuah, Gang Tan |
ATC | 3 |
| 2011 | Poster: uPro: a compartmentalization tool supporting fine-grained and flexible security configuration
Ben Niu 0007, Gang Tan |
CCS | 2 |
| 2011 | Combining control-flow integrity and static analysis for efficient and validated data sandboxingabstractIn many software attacks, inducing an illegal control-flow transfer in the target system is one common step. Control-Flow Integrity (CFI) protects a software system by enforcing a pre-determined control-flow graph. In addition to providing strong security, CFI enables static analysis on low-level code. This paper evaluates whether CFI-enabled static analysis can help build efficient and validated data sandboxing. Previous systems generally sandbox memory writes for integrity, but avoid protecting confidentiality due to the high overhead of sandboxing memory reads. To reduce overhead, we have implemented a series of optimizations that remove sandboxing instructions if they are proven unnecessary by static analysis. On top of CFI, our system adds only 2.7% runtime overhead on SPECint2000 for sandboxing memory writes and adds modest 19% for sandboxing both reads and writes. We have also built a principled data-sandboxing verifier based on range analysis. The verifier checks the safety of the results of the optimizer, which removes the need to trust the rewriter and optimizer. Our results show that the combination of CFI and static analysis has the potential of bringing down the cost of general inlined reference monitors, while maintaining strong security. Bin Zeng 0004, Gang Tan, J. Gregory Morrisett |
CCS | 2 |
| 2011 | JET: exception checking in the Java native interfaceabstractJava's type system enforces exception-checking rules that stipulate a checked exception thrown by a method must be declared in the throws clause of the method. Software written in Java often invokes native methods through the use of the Java Native Interface (JNI). Java's type system, however, cannot enforce the same exception-checking rules on Java exceptions raised in native methods. This gap makes Java software potentially buggy and often difficult to debug when an exception is raised in native code. In this paper, we propose a complete static-analysis framework called JET to extend exception-checking rules even on native code. The framework has a two-stage design where the first stage throws away a large portion of irrelevant code so that the second stage, a fine-grained analysis, can concentrate on a small set of code for accurate bug finding. This design achieves both high efficiency and accuracy. We have applied JET on a set of benchmark programs with a total over 227K lines of source code and identified 12 inconsistent native-method exception declarations. Siliang Li, Gang Tan |
OOPSLA | 2 |
| 2011 | Markup SVG - An Online Content-Aware Image Abstraction and Annotation ToolabstractSuppose you want to effectively search through millions of images, train an algorithm to perform image and video object recognition, or research the complex patterns and relationships that exist in our visual world. A common and essential component for any of these tasks is a large annotated image dataset. However, obtaining labeled image data is a complex and tedious task that requires methods for annotating and structuring content. Therefore, we developed a comprehensive online tool and data structure, Markup SVG, that simplifies the collection of annotated image data by leveraging state-of-the-art image processing techniques. As the core data structure of our tool, we adopt scalable vector graphics (SVG), an extensible and versatile language built upon XML. Given the extensibility of our framework, we are able to encode low-level image features, high-level semantics, and further define interactions with the data to assist the user with image annotation. We also demonstrate the ability to merge multiple online and offline datasets into our system in an effort to standardize image collection and its data representation. Lastly, we present our modular design; each component acts as a plug-in to our system. We developed several novel components and algorithms to highlight the possibilities of semi-supervised segmentation and automatic annotation within our proposed framework. Further, our modular design provides the necessary capabilities to incorporate future image features, methods, or algorithms. Our results show that our tool is able to greatly simplify the process of obtaining large annotated image collections in an online collaborative platform. Sharon X. Huang, Gang Tan |
IEEE Trans. Multim. | 3 |
| 2010 | JNI Light: An Operational Model for the Core JNI
Gang Tan |
APLAS | 1 |
| 2010 | Robusta: taming the native beast of the JVMabstractJava applications often need to incorporate native-code components for efficiency and for reusing legacy code. However, it is well known that the use of native code defeats Java's security model. We describe the design and implementation of Robusta, a complete framework that provides safety and security to native code in Java applications. Starting from software-based fault isolation (SFI), Robusta isolates native code into a sandbox where dynamic linking/loading of libraries in supported and unsafe system modification and confidentiality violations are prevented. It also mediates native system calls according to a security policy by connecting to Java's security manager. Our prototype implementation of Robusta is based onNative Client and OpenJDK. Experiments in this prototype demonstrate Robusta is effective and efficient, with modest runtime overhead on a set of JNI benchmark programs. Robusta can be used to sandbox native libraries used in Java's system classes to prevent attackers from exploiting bugs in the libraries. It can also enable trustworthy execution of mobile Java programs with native libraries. The design of Robusta should also be applicable when other type-safe languages (e.g., C#, Python) want to ensure safe interoperation with native libraries Joseph Siefers, Gang Tan, J. Gregory Morrisett |
CCS | 2 |
| 2010 | Semantic foundations for typed assembly languagesabstractTyped Assembly Languages (TALs) are used to validate the safety of machine-language programs. The Foundational Proof-Carrying Code project seeks to verify the soundness of TALs using the smallest possible set of axioms: the axioms of a suitably expressive logic plus a specification of machine semantics. This article proposes general semantic foundations that permit modular proofs of the soundness of TALs. These semantic foundations include Typed Machine Language (TML), a type theory for specifying properties of low-level data with powerful and orthogonal type constructors, and L c , a compositional logic for specifying properties of machine instructions with simplified reasoning about unstructured control flow. Both of these components, whose semantics we specify using higher-order logic, are useful for proving the soundness of TALs. We demonstrate this by using TML and L c to verify the soundness of a low-level, typed assembly language, LTAL, which is the target of our core-ML-to-sparc compiler. To prove the soundness of the TML type system we have successfully applied a new approach, that of step-indexed logical relations . This approach provides the first semantic model for a type system with updatable references to values of impredicative quantified types. Both impredicative polymorphism and mutable references are essential when representing function closures in compilers with typed closure conversion, or when compiling objects to simpler typed primitives. Amal Ahmed 0001, Andrew W. Appel, Christopher D. Richards, Kedar N. Swadi, Gang Tan, Daniel C. Wang |
ACM Trans. Program. Lang. Syst. | 5 |
| 2009 | Weak updates and separation logic
Gang Tan, Zhong Shao 0001, Xinyu Feng 0001, Hongxu Cai |
APLAS | 1 |
| 2009 | Finding bugs in exceptional situations of JNI programsabstractSoftware flaws in native methods may defeat Java's guarantees of safety and security. One common kind of flaws in native methods results from the discrepancy on how exceptions are handled in Java and in native methods. Unlike exceptions in Java, exceptions raised in the native code through the Java Native Interface (JNI) are not controlled by the Java Virtual Machine (JVM). Only after the native code finishes execution will the JVM's mechanism for exceptions take over. This discrepancy makes handling of JNI exceptions an error prone process and can cause serious security flaws in software written using the JNI. Siliang Li, Gang Tan |
CCS | 2 |
| 2009 | Document Analysis Support for the Manual Auditing of ElectionsabstractRecent developments have resulted in dramatic changes in the way elections are conducted, both in the United States and around the world. Well-publicized flaws in the security of electronic voting systems have led to a push for the use of verifiable paper records in the election process. In this paper, we describe the application of document analysis techniques to facilitate the manual auditing of elections,both to assure the reliability of the final outcome as well as to help reconcile the differences that may arise between repeated scans of the same ballot. We show how techniques developed for document duplicate detection can be applied to this problem, and present experimental results that demonstrate the efficacy of our approach. Related issues concerning machine support for the auditing of elections are also discussed. Daniel P. Lopresti, Xiang Sean Zhou, Sharon X. Huang, Gang Tan |
ICDAR | 4 |
| 2008 | An Empirical Security Study of the Native Code in the JDK
Gang Tan, Jason Croft |
USENIX Security Symposium | 1 |
| 2007 | Ilea: inter-language analysis across java and cabstractJava bug finders perform static analysis to find implementation mistakes that can lead to exploits and failures; Java compilers perform static analysis for optimization.allIf Java programs contain foreign function calls to C libraries, however, static analysis is forced to make either optimistic or pessimistic assumptions about the foreign function calls, since models of the C libraries are typically not available. Gang Tan, J. Gregory Morrisett |
OOPSLA | 1 |
| 2006 | A Compositional Logic for Control Flow
Gang Tan, Andrew W. Appel |
VMCAI | 1 |
| 2004 | EEG Source Localization Using Independent Residual Analysis
Gang Tan, Liqing Zhang 0001 |
ISNN (2) | 1 |
| 2004 | Construction of a Semantic Model for a Typed Assembly Language
Gang Tan, Andrew W. Appel, Kedar N. Swadi, Dinghao Wu |
VMCAI | 1 |
| 2002 | Braille to print translations for Chinese
Minghu Jiang, Georges Gielen, Elliott Drábek, Gang Tan, Ta Bao |
Inf. Softw. Technol. | 6 |