Han Qiu 0001

dblp:15/4507-1 · DBLP profile ↗
← Back
78ranked-venue papers
15as first author
69since 2021 · last 2026
0000-0003-2678-8070ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 30 since 2021Systems, architecture and hardware · 12 · 3 first-author · 10 since 2021Security and privacy · 11 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021Computer networks · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
abstract
Emojis are globally used non-verbal cues in digital communication, and extensive research has examined how large language models (LLMs) understand and utilize emojis across contexts. While usually associated with friendliness or playfulness, it is observed that emojis may trigger toxic content generation in LLMs. Motivated by such a observation, we aim to investigate: (1) whether emojis can clearly enhance the toxicity generation in LLMs and (2) how to interpret this phenomenon.* We begin with a comprehensive exploration of emoji-triggered LLM toxicity generation by automating the construction of prompts with emojis to subtly express toxic intent. Experiments across 5 mainstream languages on 7 famous LLMs along with jailbreak tasks demonstrate that prompts with emojis could easily induce toxicity generation. To understand this phenomenon, we conduct model-level interpretations spanning semantic cognition, sequence generation and tokenization, suggesting that emojis can act as a heterogeneous semantic channel to bypass the safety mechanisms. To pursue deeper insights, we further probe the pre-training corpus and uncover potential correlation between the emoji-related data polution with the toxicity generation behaviors.
Shiyao Cui, Xijia Feng, Yingkang Wang, Junxiao Yang, Zhexin Zhang, Biplab Sikdar 0001, Hongning Wang, Han Qiu 0001, Minlie Huang
AAAI8
2026 The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
abstract
Renmiao Chen, Yida Lu, Shiyao Cui, Xuan Ouyang, Victor Shea-Jay Huang, Shumin Zhang, Chengwei Pan, Han Qiu, Minlie Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Renmiao Chen, Yida Lu, Shiyao Cui, Xuan Ouyang, Victor Shea-Jay Huang, Chengwei Pan, Han Qiu 0001, Minlie Huang
ACL (1)8
2026 New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs
abstract
Shiyao Cui, QingLin Zhang, Di Wang, Yida Lu, Zhexin Zhang, Jinhua Gao, Jinglin Yang, Min He, Han Qiu, Minlie Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shiyao Cui, Yida Lu, Zhexin Zhang, Jinhua Gao, Jinglin Yang, Han Qiu 0001, Minlie Huang
ACL (1)9
2026 Revisiting the Reliability of Language Models in Instruction-Following
abstract
Advanced LLMs have achieved near-ceiling instruction-following accuracy on benchmarks such as IFEVAL.However, these impressive scores do not necessarily translate to reliable services in real-world use, where users often vary their phrasing, contextual framing, and task formulations.In this paper, we study nuance-oriented reliability: whether models exhibit consistent competence across cousin prompts that convey analogous user intents but with subtle nuances.To quantify this, we introduce a new metric, reliable@k, and develop an automated pipeline that generates high-quality cousin prompts via data augmentation.Building upon this, we construct IFE-VAL++ for systematic evaluation.Across 20 proprietary and 26 open-source LLMs, we find that current models exhibit substantial insufficiency in nuance-oriented reliability-their performance can drop by up to 61.8% with nuanced prompt modifications.What's more, we characterize it and explore three potential improvement recipes.Our findings highlight nuance-oriented reliability as a crucial yet underexplored next step toward more dependable and trustworthy LLM behavior.Our code and benchmark are accessible: https: //github.com/jianshuod/IFEval-pp.
Jianshuo Dong, Liu Yan, Zhenyu Zhong, Tao Wei 0002, Chao Zhang 0008, Han Qiu 0001
ACL (1)7
2026 LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
abstract
Junxiao Yang, Haoran Liu, Jinzhe Tu, Jiale Cheng, Zhexin Zhang, Shiyao Cui, Jiaqi Weng, Jialing Tao, Hui Xue, Hongning Wang, Han Qiu, Minlie Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junxiao Yang, Jinzhe Tu, Zhexin Zhang, Shiyao Cui, Jiaqi Weng, Jialing Tao, Hongning Wang, Han Qiu 0001, Minlie Huang
ACL (1)11
2026 DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
abstract
Despite the integration of safety alignment and external filters, text-to-image (T2I) generative systems are still susceptible to producing harmful content, such as sexual or violent imagery. This raises serious concerns about unintended exposure and potential misuse. Red teaming, which aims to proactively identify diverse prompts that can elicit unsafe outputs from the T2I system, is increasingly recognized as an essential method for assessing and improving safety before real-world deployment. However, existing automated red teaming approaches often treat prompt discovery as an isolated, prompt-level optimization task, which limits their scalability, diversity, and overall effectiveness. To bridge this gap, in this paper, we propose DREAM, a scalable red teaming framework to automatically uncover diverse problematic prompts from a given T2I system. Unlike prior work that optimizes prompts individually, DREAM directly models the probabilistic distribution of the target system's problematic prompts, which enables explicit optimization over both effectiveness and diversity, and allows efficient large-scale sampling after training. To achieve this without direct access to representative training samples, we draw inspiration from energy-based models and reformulate the objective into a simple and tractable form. We further introduce GC-SPSA, an efficient optimization algorithm that provides stable gradient estimates through the long and potentially non-differentiable T2I pipeline. During inference, we also propose a diversity-aware sampling strategy to enhance prompt variety. The effectiveness of DREAM is validated through extensive experiments, demonstrating state-of-the-art performance across a wide range of T2I models and safety filters in terms of both prompt success rate and diversity. Our code is available at https://github.com/AntigoneRandy/DREAM
Boheng Li, Junjie Wang 0007, Yiming Li 0004, Zhiyang Hu, Leyi Qi, Jianshuo Dong, Run Wang 0001, Han Qiu 0001, Zhan Qin, Tianwei Zhang 0004
SP8
2025 Understanding the Dark Side of LLMs' Intrinsic Self-Correction
abstract
Qingjie Zhang, Di Wang, Haoting Qian, Yiming Li, Tianwei Zhang, Minlie Huang, Ke Xu, Hewu Li, Liu Yan, Han Qiu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Haoting Qian, Yiming Li 0004, Tianwei Zhang 0004, Minlie Huang, Ke Xu 0002, Hewu Li, Liu Yan, Han Qiu 0001
ACL (1)10
2025 REMU: Memory-aware Radiation Emulation via Dual Addressing for In-orbit Deep Learning System
abstract
The deployment of commercial-off-the-shelf (COTS) GPUs in space has emerged as a promising approach for supporting inorbit deep neural network (DNN) inference. However, unlike terrestrial environments, understanding the impact of space radiation on COTS GPU-enabled DNNs is critical. This is challenging because existing methods, such as real-world radiation testing and software emulation, fail to link radiation-induced memory errors to runtime DNN behaviors. In this paper, we propose REMU, a memory-aware Radiation EMUlator to fill this gap. REMU introduces a dual addressing mechanism across virtual, physical, and DRAM memory spaces, enabling precise mapping and efficient injection of radiation-induced errors from DRAM to runtime DNN inference. Extensive evaluations across 10 well-known DNN models and 2 typical in-orbit computing tasks demonstrate the effectiveness of REMU, providing valuable insights for understanding the resilience of runtime DNN inferences on space radiations.
Longnv Xu, Han Qiu 0001, Jun Liu 0063, Yuanjie Li, Hewu Li
DAC3
2025 DCDiff: Enhancing JPEG Compression via Diffusion-based DC Coefficients Estimation
abstract
JPEG is the most widely-used image compression method on low-cost cameras which cannot support learning-based compressors. One promising approach to enhance JPEG aims to drop DC coefficients at the cameras’ ends (without extra computation) and reconstruct those DC coefficients after receiving them. They all face the challenge that their DC reconstruction relies on a statistical property, which will cause deviationintroduced errors and propagate. In this paper, we propose DCDiff, a novel end-to-end DC estimation method to tackle the above challenge. Instead of using statistical methods to recover DC coefficients and then fix errors, we directly leverage a generative model to estimate DC coefficients in an end-to-end manner. In the meantime, we generate masks to correct certain image locations that do not satisfy the statistical distribution to suppress error propagation. Extensive experiments show that DCDiff not only outperforms all baselines on compression performance but also introduces a tiny impact on downstream tasks and is fully compatible with 2 typical low-cost processors with JPEG support.
Han Qiu 0001, Tianwei Zhang 0004, Bin Chen 0011, Chao Zhang 0008
DAC2
2025 "I've Decided to Leak": Probing Internals Behind Prompt Leakage Intents
abstract
Large language models (LLMs) exhibit prompt leakage vulnerabilities, where they may be coaxed into revealing system prompts embedded in LLM services, raising intellectual property and confidentiality concerns.An intriguing question arises: Do LLMs genuinely internalize prompt leakage intents in their hidden states before generating tokens?In this work, we use probing techniques to capture LLMs' intent-related internal representations and confirm that the answer is yes.We start by comprehensively inducing prompt leakage behaviors across diverse system prompts, attack queries, and decoding methods.We develop a hybrid labeling pipeline, enabling the identification of broader prompt leakage behaviors beyond mere verbatim leaks.Our results show that a simple linear probe can predict prompt leakage risks from pre-generation hidden states without generating any tokens.Across all tested models, linear probes consistently achieve 90%+ AUROC, even when applied to new system prompts and attacks.Understanding the model internals behind prompt leakage drives practical applications, including intention-based detection of prompt leakage risks.
Jianshuo Dong, Liu Yan, Zhenyu Zhong, Tao Wei 0002, Ke Xu 0002, Minlie Huang, Chao Zhang 0008, Han Qiu 0001
EMNLP9
2025 When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
abstract
Large Audio-Language Models (LALMs) are enhanced with audio perception capabilities, enabling them to effectively process and understand multimodal inputs that combine audio and text.However, their performance in handling conflicting information between audio and text modalities remains largely unexamined.This paper introduces MCR-BENCH, the first comprehensive benchmark specifically designed to evaluate how LALMs prioritize information when presented with inconsistent audio-text pairs.Through extensive evaluation across diverse audio understanding tasks, we reveal a concerning phenomenon: when inconsistencies exist between modalities, LALMs display a significant bias toward textual input, frequently disregarding audio evidence.This tendency leads to substantial performance degradation in audio-centric tasks and raises important reliability concerns for real-world applications.We further investigate the influencing factors of text bias, and explore mitigation strategies through supervised finetuning, and analyze model confidence patterns that reveal persistent overconfidence even with contradictory inputs.These findings underscore the need for improved modality balance during training and more sophisticated fusion mechanisms to enhance the robustness when handling conflicting multi-modal inputs 1 .
Gelei Deng, Xianglin Yang, Han Qiu 0001, Tianwei Zhang 0004
EMNLP4
2025 Speculating LLMs' Chinese Training Data Pollution from Their Tokens
abstract
Qingjie Zhang, Di Wang, Haoting Qian, Liu Yan, Tianwei Zhang, Ke Xu, Qi Li, Minlie Huang, Hewu Li, Han Qiu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Haoting Qian, Liu Yan, Tianwei Zhang 0004, Ke Xu 0002, Qi Li 0002, Minlie Huang, Hewu Li, Han Qiu 0001
EMNLP10
2025 VISO: Accelerating In-Orbit Object Detection with Language-Guided Mask Learning and Sparse Inference
Han Qiu 0001
ICCV2
2025 Partitioning or Not? Hierarchical Task Offloading Optimization in Collaborative Satellite Edge Computing Networks
abstract
As a promising paradigm, Satellite Edge Computing (SEC) enables new opportunities for facilitating intelligent processing onboard, crucial for the timely execution of mission-critical tasks. These tasks typically involve high data capture rates and rely on compute-intensive Deep Neural Network (DNN) models. However, a single satellite struggles to handle these tasks promptly due to its limited computational capabilities. Thus, effective collaboration within the SEC network is urgently needed to adapt to diverse capture rates, optimize resource utilization, and ensure real-time responses. Motivated by the fact that partitioning a DNN model can accelerate task inference and make better use of idle resources by simultaneous sub-task execution and reduced transmitted data, we propose HiO2, a hierarchical task offloading framework that maximizes system throughput by effective collaboration among satellites and ground stations to process the partitioned sub-tasks. This highlights the challenge of designing effective task partitioning and offloading strategies in dynamic, resource-constrained networks. HiO2 addresses this challenge with two key methods. First, it adopts a distributed swarm-level task offloading strategy that assigns tasks to swarms based on their optimal quantity. Second, HiO2 introduces a distributed node-level partitioning and offloading scheme, which dynamically identifies efficient cut-points according to workload and network dynamics, then offloads sub-tasks by collaboration among nodes in each swarm. Extensive data-driven evaluations demonstrate that, compared to the state-of-the-art baselines, HiO2 improves throughput to 1.19×, reduces average task completion time to 69.7%, and consistently meets task deadlines.
Jun Liu 0063, Xiaolin Jia, Jiejie Zhao, Han Qiu 0001
ICDCS6
2025 An Engorgio Prompt Makes Large Language Model Babble on
abstract
Auto-regressive large language models (LLMs) have yielded impressive performance in many real-world tasks. However, the new paradigm of these LLMs also exposes novel threats. In this paper, we explore their vulnerability to inference cost attacks, where a malicious user crafts Engorgio prompts to intentionally increase the computation cost and latency of the inference process. We design Engorgio, a novel methodology, to efficiently generate adversarial Engorgio prompts to affect the target LLM's service availability. Engorgio has the following two technical contributions. (1) We employ a parameterized distribution to track LLMs' prediction trajectory. (2) Targeting the auto-regressive nature of LLMs' inference process, we propose novel loss functions to stably suppress the appearance of the <EOS> token, whose occurrence will interrupt the LLM's generation process. We conduct extensive experiments on 13 open-sourced LLMs with parameters ranging from 125M to 30B. The results show that Engorgio prompts can successfully induce LLMs to generate abnormally long outputs (i.e., roughly 2-13$\times$ longer to reach 90\%+ of the output length limit) in a white-box scenario and our real-world experiment demonstrates Engergio's threat to LLM service with limited computing resources. The code is released at https://github.com/jianshuod/Engorgio-prompt.
Jianshuo Dong, Tianwei Zhang 0004, Hao Wang 0003, Hewu Li, Qi Li 0002, Chao Zhang 0008, Ke Xu 0002, Han Qiu 0001
ICLR10
2025 VideoShield: Regulating Diffusion-based Video Generation Models via Watermarking
abstract
Artificial Intelligence Generated Content (AIGC) has advanced significantly, particularly with the development of video generation models such as text-to-video (T2V) models and image-to-video (I2V) models. However, like other AIGC types, video generation requires robust content control. A common approach is to embed watermarks, but most research has focused on images, with limited attention given to videos. Traditional methods, which embed watermarks frame-by-frame in a post-processing manner, often degrade video quality. In this paper, we propose VideoShield, a novel watermarking framework specifically designed for popular diffusion-based video generation models. Unlike post-processing methods, VideoShield embeds watermarks directly during video generation, eliminating the need for additional training. To ensure video integrity, we introduce a tamper localization feature that can detect changes both temporally (across frames) and spatially (within individual frames). Our method maps watermark bits to template bits, which are then used to generate watermarked noise during the denoising process. Using DDIM Inversion, we can reverse the video to its original watermarked noise, enabling straightforward watermark extraction. Additionally, template bits allow precise detection for potential spatial and temporal modification. Extensive experiments across various video models (both T2V and I2V models) demonstrate that our method effectively extracts watermarks and detects tamper without compromising video quality. Furthermore, we show that this approach is applicable to image generation models, enabling tamper detection in generated images as well. Codes and models are available at https://github.com/hurunyi/VideoShield.
Runyi Hu, Jie Zhang 0073, Yiming Li 0004, Jiwei Li 0001, Qing Guo 0005, Han Qiu 0001, Tianwei Zhang 0004
ICLR6
2025 A Benchmark for Semantic Sensitive Information in LLMs Outputs
abstract
Large language models (LLMs) can output sensitive information, which has emerged as a novel safety concern. Previous works focus on structured sensitive information (e.g. personal identifiable information). However, we notice that sensitive information can also be at semantic level, i.e. semantic sensitive information (SemSI). Particularly, *simple natural questions* can let state-of-the-art (SOTA) LLMs output SemSI. %which is hard to be detected compared with structured ones. Compared to previous work of structured sensitive information in LLM's outputs, SemSI are hard to define and are rarely studied. Therefore, we propose a novel and large-scale investigation on the existence of SemSI in SOTA LLMs induced by simple natural questions. First, we construct a comprehensive and labeled dataset of semantic sensitive information, SemSI-Set, by including three typical categories of SemSI. Then, we propose a large-scale benchmark, SemSI-Bench, to systematically evaluate semantic sensitive information in 25 SOTA LLMs. Our finding reveals that SemSI widely exists in SOTA LLMs' outputs by querying with simple natural questions. We open-source our project at https://semsi-project.github.io/.
Han Qiu 0001, Yiming Li 0004, Tianwei Zhang 0004, Wenyu Zhu, Haiqin Weng, Liu Yan, Chao Zhang 0008
ICLR2
2025 Cowpox: Towards the Immunity of VLM-based Multi-Agent Systems
abstract
Vision Language Model (VLM) Agents are stateful, autonomous entities capable of perceiving and interacting with their environments through vision and language. Multi-agent systems comprise specialized agents who collaborate to solve a (complex) task. A core security property is robustness, stating that the system maintains its integrity during adversarial attacks. Multi-agent systems lack robustness, as a successful exploit against one agent can spread and infect other agents to undermine the entire system’s integrity. We propose a defense Cowpox to provably enhance the robustness of a multi-agent system by a distributed mechanism that improves the recovery rate of agents by limiting the expected number of infections to other agents. The core idea is to generate and distribute a special cure sample that immunizes an agent against the attack before exposure. We demonstrate the effectiveness of Cowpox empirically and provide theoretical robustness guarantees.
Yutong Wu 0009, Jie Zhang 0073, Yiming Li 0004, Chao Zhang 0008, Qing Guo 0005, Han Qiu 0001, Nils Lukas, Tianwei Zhang 0004
ICML6
2025 ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs: ShieldVLM
abstract
Toxicity detection in multimodal text-image content faces growing challenges, especially with multimodal implicit toxicity, where each modality appears benign on its own but conveys hazard when combined. Multimodal implicit toxicity appears not only as formal statements in social platforms but also prompts that can lead to toxic dialogs from Large Vision-Language Models (LVLMs). Despite the success in unimodal text or image moderation, toxicity detection for multimodal content, particularly the multimodal implicit toxicity, remains underexplored. To fill this gap, we comprehensively build a taxonomy for multimodal implicit toxicity (MMIT) and introduce an MMIT-dataset, comprising 2,100 multimodal statements and prompts across 7 risk categories (31 sub-categories) and 5 typical cross-modal correlation modes. To advance the detection of multimodal implicit toxicity, we build ShieldVLM, a model which identifies implicit toxicity in multimodal statements, prompts and dialogs via deliberative cross-modal reasoning. Experiments show that ShieldVLM outperforms existing strong baselines in detecting both implicit and explicit toxicity. The model and dataset will be publicly available to support future researches (Warning: This paper contains potentially sensitive contents). Warning: This paper contains potentially sensitive contents.
Shiyao Cui, Xuan Ouyang, Renmiao Chen, Zhexin Zhang, Yida Lu, Hongning Wang, Han Qiu 0001, Minlie Huang
ACM Multimedia8
2025 Mask Image Watermarking
abstract
We present MaskWM, a simple, efficient, and flexible framework for image watermarking. MaskWM has two variants: (1) MaskWM-D, which supports global watermark embedding, watermark localization, and local watermark extraction for applications such as tamper detection; (2) MaskWM-ED, which focuses on local watermark embedding and extraction, offering enhanced robustness in small regions to support fine-grined image protection. MaskWM-D builds on the classical encoder-distortion layer-decoder training paradigm. In MaskWM-D, we introduce a simple masking mechanism during the decoding stage that enables both global and local watermark extraction. During training, the decoder is guided by various types of masks applied to watermarked images before extraction, helping it learn to localize watermarks and extract them from the corresponding local areas. MaskWM-ED extends this design by incorporating the mask into the encoding stage as well, guiding the encoder to embed the watermark in designated local regions, which improves robustness under regional attacks. Extensive experiments show that MaskWM achieves state-of-the-art performance in global and local watermark extraction, watermark localization, and multi-watermark embedding. It outperforms all existing baselines, including the recent leading model WAM for local watermarking, while preserving high visual quality of the watermarked images. In addition, MaskWM is highly efficient and adaptable. It requires only 20 hours of training on a single A6000 GPU, achieving 15× computational efficiency compared to WAM. By simply adjusting the distortion layer, MaskWM can be quickly fine-tuned to meet varying robustness requirements.
Runyi Hu, Jie Zhang 0073, Shiqian Zhao, Nils Lukas, Jiwei Li 0001, Qing Guo 0005, Han Qiu 0001, Tianwei Zhang 0004
NeurIPS7
2025 Image Compression for Resource-Constrained AIoT System With Compressed Sensing
abstract
In today’s big data era, a key requirement is to implement intelligent semantic analysis (such as image recognition) on data gathered from an extensive array of smart devices in Artificial Intelligence IoT (AIoT) scenarios, all of which is processed at central cloud service providers. Recent advancements in deep-learning-based image compression have fostered semantic compression between machines. However, the deployment of an overparameterized encoder on Internet of Things (IoT) devices remains a challenge due to their restricted computing and storage capabilities. To tackle this issue, we propose a novel approach named compressed sensing (CS)-based asymmetric semantic image compression (CS-ASIC), explicitly designed for resource-constrained AIoT systems. This asymmetric semantic compression scheme intends to surpass the limitations of IoT devices, thereby facilitating efficient semantic compression for machine vision tasks. CS-ASIC notably includes a lightweight front encoder founded on deep image CS techniques, which utilizes rich image priors to learn measurement matrices for sampling. In tandem, a deep iterative decoder is designed cooperatively with the linear encoder offloaded at the server to enhance image reconstruction and semantic analysis across various semantic analysis tasks. Furthermore, we introduce a groundbreaking lossy CS semantic rate-distortion theoretical framework that justifies a compromise in rate for extended semantic distortion. Extensive experimental results underscore the superiority of the proposed CS-ASIC concerning the signal-semantic rate-distortion tradeoff, and its lower encoding complexity over existing codecs in an AIoT simulation environment.
Bin Chen 0011, Yujun Huang, Han Qiu 0001, Shutao Xia, Wei Fei, Xuan Wang 0002, Meikang Qiu
IEEE Trans. Syst. Man Cybern. Syst.3
2024 The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
abstract
Rongwu Xu, Brian Lin, Shujian Yang, Tianqi Zhang, Weiyan Shi, Tianwei Zhang, Zhixuan Fang, Wei Xu, Han Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Rongwu Xu, Brian S. Lin, Shujian Yang, Weiyan Shi 0001, Tianwei Zhang 0004, Zhixuan Fang, Wei Xu 0039, Han Qiu 0001
ACL (1)9
2024 VisionGuard: Secure and Robust Visual Perception of Autonomous Vehicles in Practice
abstract
Modern Autonomous Vehicles (AVs) implement the Visual Perception Module (VPM) to perceive their surroundings. This VPM adopts various Deep Neural Network (DNN) models to process the data collected from cameras and LiDAR. Prior studies have shown that these models are vulnerable to physical adversarial examples (PAEs), which pose a critical safety risk to the autonomous driving task. While a few defense methods have been proposed to safeguard AVs, most of them only target a limited set of attack types and specific scenarios, making them impractical for real-world protection.
Xingshuo Han, Haozhao Wang, Kangqiao Zhao, Gelei Deng, Yuan Xu 0033, Hangcheng Liu, Han Qiu 0001, Tianwei Zhang 0004
CCS7
2024 PhyScout: Detecting Sensor Spoofing Attacks via Spatio-temporal Consistency
abstract
Existing defense approaches against sensor spoofing attacks suf- fer from the limitations of limited specific attack types, requiring GPU computation, exhibiting considerable detection latency and struggling with the interpretability of corner cases. We developed PhyScout, a holistic sensor spoofing defense framework to over- come the above limitations. Our framework capitalizes on the ob- servation that human drivers can rapidly and accurately identify spoofing attacks by performing spatio-temporal consistency checks of their environment. We commence by defining the generalized conflicts that different sensor spoofing attacks produce regarding the spatio-temporal consistency. These conflicts are subsequently unified and formalized through a least squares problem approach. This process is modeled using image-based feature point extrac- tion and matching techniques, followed by the design of a risk identification method for each conflict. We evaluate PhyScout across various environments, including simulators, datasets, and real-world scenarios. Compared to existing defense solutions, PhyScout offers rapid identification of sensor at- tacks (within 100ms) with low performance overhead (CPU-based), and conflict visualization. It demonstrates a fresh paradigm in au- tonomous vehicle security and presents new avenues for future research in robust and efficient defense mechanisms against sensor spoofing attacks. More video demos are at our anonymous website https://sites.google.com/view/physcout.
Yuan Xu 0033, Gelei Deng, Xingshuo Han, Han Qiu 0001, Tianwei Zhang 0004
CCS5
2024 Message from the Program Chairs; CSCloud2024
abstract
It is a great pleasure for us to welcome you on behalf of the conference committees to the 11th IEEE International Conference on Cyber Security and Cloud Computing (IEEE CSCloud 2024). I am glad that we can have this international conference at Shanghai, China in this summer.
Gérard Memmi, Han Qiu 0001, Zakirul Alam
CSCloud2
2024 Laser Shield: a Physical Defense with Polarizer against Laser Attacks on Autonomous Driving Systems
abstract
Autonomous driving systems (ADS) are boosted with deep neural networks (DNN) to perceive environments, while their security is doubted by DNN's vulnerability to adversarial attacks. Among them, a diversity of laser attacks emerges to be a new threat due to its minimal requirements and high attack success rate in the physical world. Nevertheless, current defense methods exhibit either a low defense success rate or a high computation cost against laser attacks. To fill this gap, we propose Laser Shield which leverages a polarizer along with a min-energy rotation mechanism to eliminate adversarial lasers from ADS scenes. We also provide a physical world dataset, LAPA, to evaluate its performance. Through exhaustive experiments with three baselines, four metrics, and three settings, Laser Shield is proved to surpass SOTA performance.
Lijun Chi, Mounira Msahli, Gérard Memmi, Tianwei Zhang 0004, Chao Zhang 0008, Han Qiu 0001
DAC8
2024 Protecting Confidential Virtual Machines from Hardware Performance Counter Side Channels
abstract
In modern cloud platforms, it is becoming more important to preserve the privacy of guest virtual machines (VMs) from the untrusted host. To this end, Secure Encrypted Virtualization (SEV) is developed as a hardware extension to protect VMs by encrypting their memory pages and register states. Unfortunately, such confidential VMs are still vulnerable to micro-architectural side channels, and Hardware Performance Counters (HPCs) are a prominent information leakage source. To make matters worse, currently there is no systematic defense against the HPC side channels. We introduce Aegis, a unified framework for demystifying the inherent relations between the instruction execution and HPC event statistics, and defending VMs against HPC side channels with provable privacy guarantee and minimal performance overhead. Aegis consists of three modules. Application Profiler profiles the application offline and adopts information theory to quantitatively estimate the vulnerability of HPC events. Event Fuzzer leverages the fuzzing technique to automatically generate interesting inputs, i.e., instruction sequences, that can effectively alter the HPC observations. Event Obfuscator injects noisy instructions into the protected VM based on the differential privacy mechanisms for high efficiency and privacy. We present three case studies to demonstrate that Aegis can defeat different types of HPC side-channel attacks (i.e., website fingerprinting, DNN model extraction, keystroke sniffing). Evaluations show that Aegis can effectively decrease the attack accuracy from 90% to 2%, with only 3% overhead on the application execution time and 7% overhead on the CPU usage.
Xiaoxuan Lou, Kangjie Chen, Guowen Xu, Han Qiu 0001, Shangwei Guo, Tianwei Zhang 0004
DSN4
2024 Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
abstract
The common toxicity and societal bias in contents generated by large language models (LLMs) necessitate strategies to reduce harm.Present solutions often demand whitebox access to the model or substantial training, which is impractical for cutting-edge commercial LLMs.Moreover, prevailing prompting methods depend on external tool feedback and fail to simultaneously lessen toxicity and bias.Motivated by social psychology principles, we propose a novel strategy named perspective-taking prompting (PET) that inspires LLMs to integrate diverse human perspectives and self-regulate their responses.This self-correction mechanism can significantly diminish toxicity (up to 89%) and bias (up to 73%) in LLMs' responses.Rigorous evaluations and ablation studies are conducted on two commercial LLMs (ChatGPT and GLM) and three open-source LLMs, revealing PET's superiority in producing less harmful responses, outperforming five strong baselines."Words kill, words give life; they're either poison or fruit-you choose."~Proverbs 18:21 (MSG)
Rongwu Xu, Zi'an Zhou, Tianwei Zhang 0004, Zehan Qi, Su Yao, Ke Xu 0002, Wei Xu 0039, Han Qiu 0001
EMNLP8
2024 Fingerprinting Image-to-Image Generative Adversarial Networks
abstract
Generative Adversarial Networks (GANs) have been widely used in various application scenarios. Since the production of a commercial GAN requires substantial computational and human resources, the copyright protection of GANs is urgently needed. This paper presents a novel finger-printing scheme for the Intellectual Property (IP) protection of image-to-image GANs based on a trusted third party. We break through the stealthiness and robustness bottlenecks suffered by previous fingerprinting methods for classification models being naively transferred to GANs. Specifically, we innovatively construct a composite deep learning model from the target GAN and a classifier. Then we generate fingerprint samples from this composite model, and embed them in the classifier for effective ownership verification. This scheme inspires some concrete methodologies to practically protect the modern image-to-image translation GANs. Theoretical analysis proves that these methods can satisfy different security requirements necessary for IP protection. We also conduct extensive experiments to show that our solutions outperform existing strategies.
Guowen Xu, Han Qiu 0001, Shangwei Guo, Run Wang 0001, Jiwei Li 0001, Tianwei Zhang 0004, Rongxing Lu
EuroS&P3
2024 UniGuard: A Unified Hardware-oriented Threat Detector for FPGA-based AI Accelerators
abstract
The proliferation of AI technology gives rise to a variety of security threats, significantly compromising the confidentiality and integrity of AI applications. Existing software-based solutions mainly target one specific attack, and require the implementation into the models, rendering them less practical. We design UniGuard, a novel unified and non-intrusive detection methodology to safeguard FPGA-based AI accelerators. The core idea of UniGuard is to harness power side-channel information generated during model inference to spot any anomaly. We employ a Time-to-Digital Converter to capture power fluctuations and train a supervised machine learning model to identify various types of threats. Evaluations demonstrate that UniGuard can achieve 94.0% attack detection accuracy, with high generalization over unknown or adaptive attacks and robustness against varied configurations (e.g., sensor frequency and location).
Xiaobei Yan, Han Qiu 0001, Tianwei Zhang 0004
FPL2
2024 You Only Query Once: An Efficient Label-Only Membership Inference Attack
abstract
As one of the privacy threats to machine learning models, the membership inference attack (MIA) tries to infer whether a given sample is in the original training set of a victim model by analyzing its outputs. Recent studies only use the predicted hard labels to achieve impressive membership inference accuracy. However, such label-only MIA approach requires very high query budgets to evaluate the distance of the target sample from the victim model's decision boundary. We propose YOQO, a novel label-only attack to overcome the above limitation.YOQO aims at identifying a special area (called improvement area) around the target sample and crafting a query sample, whose hard label from the victim model can reliably reflect the target sample's membership. YOQO can successfully reduce the query budget from more than 1,000 times to only ONCE. Experiments demonstrate that YOQO is not only as effective as SOTA attack methods, but also performs comparably or even more robustly against many sophisticated defenses.
Yutong Wu 0009, Han Qiu 0001, Shangwei Guo, Jiwei Li 0001, Tianwei Zhang 0004
ICLR2
2024 Purifying Quantization-conditioned Backdoors via Layer-wise Activation Correction with Distribution Approximation
abstract
Model quantization is a compression technique that converts a full-precision model to a more compact low-precision version for better storage. Despite the great success of quantization, recent studies revealed the feasibility of malicious exploiting model quantization via implanting quantization-conditioned backdoors (QCBs). These special backdoors remain dormant in full-precision models but are exposed upon quantization. Unfortunately, existing defenses have limited effects on mitigating QCBs. In this paper, we conduct an in-depth analysis of QCBs. We reveal an intriguing characteristic of QCBs, where activation of backdoor-related neurons on even benign samples enjoy a distribution drift after quantization, although this drift is more significant on poisoned samples. Motivated by this finding, we propose to purify the backdoor-exposed quantized model by aligning its layer-wise activation with its full-precision version. To further exploit the more pronounced activation drifts on poisoned samples, we design an additional module to layer-wisely approximate poisoned activation distribution based on batch normalization statistics of the full-precision model. Extensive experiments are conducted, verifying the effectiveness of our defense. Our code is publicly available.
Boheng Li, Yishuo Cai, Jisong Cai, Yiming Li 0004, Han Qiu 0001, Run Wang 0001, Tianwei Zhang 0004
ICML5
2024 CLAP: Learning Transferable Binary Code Representations with Natural Language Supervision
abstract
Binary code representation learning has shown significant performance in binary analysis tasks. But existing solutions often have poor transferability, particularly in few-shot and zero-shot scenarios where few or no training samples are available for the tasks. To address this problem, we present CLAP (Contrastive Language-Assembly Pre-training), which employs natural language supervision to learn better representations of binary code (i.e., assembly code) and get better transferability. At the core, our approach boosts superior transfer learning capabilities by effectively aligning binary code with their semantics explanations (in natural language), resulting a model able to generate better embeddings for binary code. To enable this alignment training, we then propose an efficient dataset engine that could automatically generate a large and diverse dataset comprising of binary code and corresponding natural language explanations. We have generated 195 million pairs of binary code and explanations and trained a prototype of CLAP. The evaluations of CLAP across various downstream tasks in binary analysis all demonstrate exceptional performance. Notably, without any task-specific training, CLAP is often competitive with a fully supervised baseline, showing excellent transferability.
Hao Wang 0226, Chao Zhang 0008, Zihan Sha, Yuchen Zhou 0007, Wenyu Zhu, Wenju Sun, Han Qiu 0001, Xi Xiao 0001
ISSTA9
2024 CEBin: A Cost-Effective Framework for Large-Scale Binary Code Similarity Detection
abstract
Binary code similarity detection (BCSD) is a fundamental technique for various applications. Many BCSD solutions have been proposed recently, which mostly are embedding-based, but have shown limited accuracy and efficiency especially when the volume of target binaries to search is large. To address this issue, we propose a cost-effective BCSD framework, CEBin, which fuses embedding-based and comparison-based approaches to significantly improve accuracy while minimizing overheads. Specifically, CEBin utilizes a refined embedding-based approach to extract features of target code, which efficiently narrows down the scope of candidate similar code and boosts performance. Then, it utilizes a comparison-based approach that performs a pairwise comparison on the candidates to capture more nuanced and complex relationships, which greatly improves the accuracy of similarity detection. By bridging the gap between embedding-based and comparison-based approaches, CEBin is able to provide an effective and efficient solution for detecting similar code (including vulnerable ones) in large-scale software ecosystems. Experimental results on three well-known datasets demonstrate the superiority of CEBin over existing state-of-the-art (SOTA) baselines. To further evaluate the usefulness of BCSD in real world, we construct a large-scale benchmark of vulnerability, offering the first precise evaluation scheme to assess BCSD methods for the 1-day vulnerability detection task. CEBin could identify the similar function from millions of candidate functions in just a few seconds and achieves an impressive recall rate of 85.46% on this more practical but challenging task, which are several order of magnitudes faster and 4.07× better than the best SOTA baseline.
Hao Wang 0226, Chao Zhang 0008, Yuchen Zhou 0007, Han Qiu 0001, Xi Xiao 0001
ISSTA6
2024 COSMIC: Compress Satellite Image Efficiently via Diffusion Compensation
abstract
With the rapidly increasing number of satellites in space and their enhanced capabilities, the amount of earth observation images collected by satellites is exceeding the transmission limits of satellite-to-ground links. Although existing learned image compression solutions achieve remarkable performance by using a sophisticated encoder to extract fruitful features as compression and using a decoder to reconstruct. It is still hard to directly deploy those complex encoders on current satellites' embedded GPUs with limited computing capability and power supply to compress images in orbit. In this paper, we propose COSMIC, a simple yet effective learned compression solution to transmit satellite images. We first design a lightweight encoder (i.e. reducing FLOPs by 2.5~5X) on satellite to achieve a high image compression ratio to save satellite-to-ground links. Then, for reconstructions on the ground, to deal with the feature extraction ability degradation due to simplifying encoders, we propose a diffusion-based model to compensate image details when decoding. Our insight is that satellite's earth observation photos are not just images but indeed multi-modal data with a nature of Text-to-Image pairing since they are collected with rich sensor data (e.g. coordinates, timestep, etc.) that can be used as the condition for diffusion generation. Extensive experiments show that COSMIC outperforms state-of-the-art baselines on both perceptual and distortion metrics.
Han Qiu 0001, Maosen Zhang, Jun Liu 0063, Bin Chen 0011, Tianwei Zhang 0004, Hewu Li
NeurIPS2
2024 Backdooring Multimodal Learning
abstract
Deep Neural Networks (DNNs) are vulnerable to backdoor attacks, which poison the training set to alter the model prediction over samples with a specific trigger. While existing efforts mainly focus on unimodal scenarios, modern AI systems usually employ multiple modalities to improve the model performance, making multimodal backdoor attacks more practical but structurally more complex due to inherent modality interactions, multiple attack surfaces, unbalanced modality contributions, etc. These factors affect the effectiveness of backdooring multimodal learning significantly but have not been fully investigated yet.To bridge this gap, we present the first data and computation efficient backdoor attacks towards multimodal learning. Our solution consists of two innovations. First, we propose a novel backdoor gradient-based score (BAGS), which can accurately quantify the contribution of each data sample to the backdoor learning at a very early training stage. Therefore, it can greatly save time and computational resources for the attacker. Second, we introduce a searching strategy with two attack modes to efficiently determine the optimal poisoning modalities and data samples.Our methodology leads to the following research outcomes. First, we comprehensively evaluate the proposed solution over state-of-the-art multimodal tasks, models, datasets and settings, to verify its effectiveness, efficiency and transferability. For instance, we only need to poison 0.005% of training samples to attack the Visual Question Answering task with the success rate of >96%. For the Audio Video Speech Recognition task, we poison 0.05% of samples to achieve the success rate of >93%. Second, we disclose several interesting findings during our experiments: (1) poisoning all modalities is not always better than individual ones, sometimes even making the attack worse; (2) modality competition and complementarity coexist in multimodal learning backdoor attacks; (3) A dominant modality in multimodal learning may not dominate the backdoor attacks. We hope this work will spur future research in improving the security of multimodal learning. Code is available at https://github.com/multimodalbags/BAGS_Multimodal.
Xingshuo Han, Yutong Wu 0009, Yuan Zhou 0005, Yuan Xu 0033, Han Qiu 0001, Guowen Xu, Tianwei Zhang 0004
SP6
2024 An Efficient Preprocessing-Based Approach to Mitigate Advanced Adversarial Attacks
abstract
Deep Neural Networks are well-known to be vulnerable to Adversarial Examples. Recently, advanced gradient-based attacks were proposed (e.g., BPDA and EOT), which can significantly increase the difficulty and complexity of designing effective defenses. In this paper, we present a study towards the opportunity of mitigating those powerful attacks with only pre-processing operations. We make the following two contributions. First, we perform an in-depth analysis of those attacks and summarize three fundamental properties that a good defense solution should have. Second, we design a lightweight preprocessing function with these properties and the capability of preserving the model's usability and robustness against these threats. Extensive evaluations indicate that our solutions can effectively mitigate all existing standard and advanced attack techniques, and beat 11 state-of-the-art defense solutions published in top-tier conferences over the past 2 years.
Han Qiu 0001, Yi Zeng 0005, Qinkai Zheng, Shangwei Guo, Tianwei Zhang 0004, Hewu Li
IEEE Trans. Computers1
2024 Incremental Learning, Incremental Backdoor Threats
abstract
Class incremental learning from a pre-trained DNN model is gaining lots of popularity. Unfortunately, the pre-trained model also introduces a new attack vector, which enables an adversary to inject a backdoor into it and further compromise the downstream models learned from it. Prior works proposed backdoor attacks against the pre-trained models in the transfer learning scenario. However, they become less effective when the adversary does not have the knowledge of the downstream tasks or new data, which is more practical and considered in this paper. To this end, we design the first latent backdoor attacks against incremental learning. We propose two novel techniques, which can effectively and stealthily embed a backdoor into the pre-trained model. Such backdoor can only be activated when the pre-trained model is extended to a downstream model with incremental learning. It has a very high attack success rate, and is able to bypass existing backdoor detection approaches. Extensive experiments confirm the effectiveness of our attacks over different datasets and incremental learning methods, as well as strong robustness against state-of-the-art backdoor defense mechanisms includingNeural Cleanse,Fine-PruningandSTRIP.
Wenbo Jiang 0001, Tianwei Zhang 0004, Han Qiu 0001, Hongwei Li 0001, Guowen Xu
IEEE Trans. Dependable Secur. Comput.3
2023 Public-attention-based Adversarial Attack on Traffic Sign Recognition
abstract
Autonomous driving systems (ADS) can instantaneously and accurately recognize traffic signs by using deep neural networks (DNNs). Although adversarial attacks are well-known to easily fool DNNs by adding tiny but malicious perturbations, most attack methods require sufficient information about the victim models (white-box) to perform. In this paper, we propose a black-box attack in the recognition system of ADS, Public Attention Attacks (PAA), that can attack a black-box model by collecting the generic attention patterns of other white-box DNNs to transfer the attack. Particularly, we select multiple dual or triple attention patterns of white-box model combinations to generate the transferable adversarial perturbations for PAA attacks. We perform the experimentation on four well-trained models in different adversarial settings separately. The results indicate that when more white-box models the adversary collects to perform PAA, the higher the attack success rate (ASR) he can achieve to attack the target black-box model.
Lijun Chi, Mounira Msahli, Gérard Memmi, Han Qiu 0001
CCNC4
2023 MPass: Bypassing Learning-based Static Malware Detectors
abstract
Machine learning (ML) based static malware detectors are widely deployed, but vulnerable to adversarial attacks. Unlike images or texts, tiny modifications to malware samples would significantly compromise their functionality. Consequently, existing attacks against images or texts will be significantly restricted when being deployed on malware detectors. In this work, we propose a hard-label black-box attack MPass against ML-based detectors. MPass employs a problem-space explainability method to locate critical positions of malware, applies adversarial modifications to such positions, and utilizes a runtime recovery technique to preserve the functionality. Experiments show MPass outperforms existing solutions and bypasses both state-of-the-art offline models and commercial ML-based antivirus products.
Jialai Wang, Wenjie Qu 0001, Han Qiu 0001, Qi Li 0002, Zongpeng Li, Chao Zhang 0008
DAC4
2023 One-bit Flip is All You Need: When Bit-flip Attack Meets Model Training
abstract
Deep neural networks (DNNs) are widely deployed on real-world devices. Concerns regarding their security have gained great attention from researchers. Recently, a new weight modification attack called bit flip attack (BFA) was proposed, which exploits memory fault inject techniques such as row hammer to attack quantized models in the deployment stage. With only a few bit flips, the target model can be rendered useless as a random guesser or even be implanted with malicious functionalities. In this work, we seek to further reduce the number of bit flips. We propose a training-assisted bit flip attack, in which the adversary is involved in the training stage to build a high-risk model to release. This high-risk model, obtained coupled with a corresponding malicious model, behaves normally and can escape various detection methods. The results on benchmark datasets show that an adversary can easily convert this high-risk but normal model to a malicious one on victim’s side by flipping only one critical bit on average in the deployment stage. Moreover, our attack still poses a significant threat even when defenses are employed. The codes for reproducing main experiments are available at https://github.com/jianshuod/TBA.
Jianshuo Dong, Han Qiu 0001, Yiming Li 0004, Tianwei Zhang 0004, Yuanjie Li, Zeqi Lai, Chao Zhang 0008, Shutao Xia
ICCV2
2023 Computation and Data Efficient Backdoor Attacks
abstract
Backdoor attacks against deep neural network (DNN) models have been widely studied. Various attack techniques have been proposed for different domains and paradigms, e.g., image, point cloud, natural language processing, transfer learning, etc. The most widely-used way to embed a backdoor into a DNN model is to poison the training data. They usually randomly select samples from the benign training set for poisoning, without considering the distinct contribution of each sample to the backdoor effectiveness, making the attack less optimal.A recent work [40] proposed to use the forgetting score to measure the importance of each poisoned sample and then filter out redundant data for effective backdoor training. However, this method is empirically designed without theoretical proofing. It is also very time-consuming as it needs to go through several training stages for data selection. To address such limitations, we propose a novel confidence-based scoring methodology, which can efficiently measure the contribution of each poisoning sample based on the distance posteriors. We further introduce a greedy search algorithm to find the most informative samples for backdoor injection more promptly. Experimental evaluations on both 2D image and 3D point cloud classification tasks show that our approach can achieve comparable performance or even surpass the forgetting score-based searching method while requiring only several extra epochs’ computation of a standard training process. Our code can be found at https://github.com/WU-YU-TONG/computational_efficient_backdoor
Yutong Wu 0009, Xingshuo Han, Han Qiu 0001, Tianwei Zhang 0004
ICCV3
2023 ATTA: Adversarial Task-transferable Attacks on Autonomous Driving Systems
abstract
Deep learning (DL) based perception models have enabled the possibility of current autonomous driving systems (ADS). However, various studies have pointed out that the DL models inside the ADS perception modules are vulnerable to adversarial attacks which can easily manipulate these DL models’ predictions. In this paper, we propose a more practical adversarial attack against the ADS perception module. Particularly, instead of targeting one of the DL models inside the ADS perception module, we propose to use one universal patch to mislead multiple DL models inside the ADS perception module simultaneously which leads to a higher chance of system-wide malfunction. We achieve such a goal by attacking the attention of DL models as a higher level of feature representation rather than traditional gradient-based attacks. We successfully generate a universal patch containing malicious perturbations that can attract multiple victim DL models’ attention to further induce their prediction errors. We verify our attack with extensive experiments on a typical ADS perception module structure with five famous datasets and also physical world scenes1.1We release our code at https://github.com/qingjiesjtu/ATTA
Maosen Zhang, Han Qiu 0001, Tianwei Zhang 0004, Mounira Msahli, Gérard Memmi
ICDM3
2023 Extracting Robust Models with Uncertain Examples
Guowen Xu, Shangwei Guo, Han Qiu 0001, Jiwei Li 0001, Tianwei Zhang 0004
ICLR4
2023 Mind Your Heart: Stealthy Backdoor Attack on Dynamic Deep Neural Network in Edge Computing
Han Qiu 0001, Tianwei Zhang 0004, Hewu Li, Terry Wang
INFOCOM3
2023 A Networking Perspective on Starlink's Self-Driving LEO Mega-Constellation
abstract
Low-earth-orbit (LEO) satellite mega-constellations, such as SpaceX Starlink, are under rocket-fast deployments and promise broadband Internet to remote areas that terrestrial networks cannot reach. For mission safety and sustainable uses of space, Starlink has adopted a proprietary onboard autonomous driving system for its extremely mobile LEO satellites. This paper demystifies and diagnoses its impacts on the LEO mega-constellation and satellite networks. We design a domain-specific method to characterize key components in Starlink's autonomous driving from various public space situational awareness datasets, including continuous orbit maintenance, collision avoidance, and maneuvers between orbital shells. Our analysis shows that, these operations have mixed impacts on the stability and performance of the entire mega-constellation, inter-satellite links, topology, and upper-layer network functions. To this end, we investigate and empirically assess the potential of networking-autonomous driving co-designs for the upcoming satellite networks.
Yuanjie Li, Hewu Li, Wei Liu 0192, Wei Zhao 0058, Yimei Chen, Qian Wu 0001, Jun Liu 0063, Zeqi Lai, Han Qiu 0001
MobiCom11
2023 Aegis: Mitigating Targeted Bit-flip Attacks against Deep Neural Networks
Jialai Wang, Han Qiu 0001, Tianwei Zhang 0004, Qi Li 0002, Zongpeng Li, Tao Wei 0002, Chao Zhang 0008
USENIX Security Symposium4
2023 DefQ: Defensive Quantization Against Inference Slow-Down Attack for Edge Computing
abstract
The novel multiexit deep neural network (DNN) architectures provide a new optimization solution for efficient model inference in edge systems. Inference of most samples can be completed within the first few layers on an edge device without the need to transmit them to a remote server. This can significantly increase the inference speed and system throughput, which is particularly beneficial to the resource-constrained scenarios. Unfortunately, researchers proposed an inference slow-down attack against this technique, where an external adversary can add imperceptible perturbations on clean samples to invalidate the multiexit mechanism. In this article, we propose a defensive quantization (DefQ) method as the first defense against the inference slow-down attack. It is designed to be lightweight and can be easily implemented in off-the-shelf camera sensors. Particularly,DefQintroduces a novel quantization operation to preprocess the input images. It is capable of removing the perturbations from the malicious samples and preserving the correct inference exit points and prediction accuracy. Meanwhile, it has little impact on the clean samples. Extensive evaluations show thatDefQcan effectively defeat the inference slow-down attack and well protect the efficiency of edge systems.
Han Qiu 0001, Tianwei Zhang 0004, Tianzhu Zhang 0002, Meikang Qiu
IEEE Internet Things J.1
2023 Wangiri Fraud: Pattern Analysis and Machine-Learning-Based Detection
abstract
The rapid growth of the telecommunication landscape leads to a rapid rise of frauds in such networks. In this article, Wangiri fraud in which users are deceived by being charged for services without their knowledge during a call is tackled. In fact, Wangiri fraud has significant negative financial and reputation consequences for the mobile service providers and also has a bad psychological impact on the victims. In order to identify this fraudulent behavior, three Wangiri fraud patterns are defined by analyzing call records of over a year. Then, the security and performance of unsupervised and supervised machine learning (ML) methods in detecting one Wangiri pattern are evaluated using a large real-world Call Detail Records (CDRs) data set. In the context of Wangiri fraud detection, classification algorithms outperformed the others based on the chosen security and performance metrics. Finally, the performance evaluation of these algorithms is extended in detecting the other two real-world Wangiri fraud patterns. This article provides a detailed definition of the Wangiri fraud patterns and outlines the implementation and evaluation of ML algorithms in the context of detecting Wangiri fraud. The security analysis and experimental results demonstrate that depending on fraud patterns the best ML algorithm to detect Wangiri fraud may also vary.
Akshaya Ravi, Mounira Msahli, Han Qiu 0001, Gérard Memmi, Albert Bifet, Meikang Qiu
IEEE Internet Things J.3
2023 Automatic Transformation Search Against Deep Leakage From Gradients
abstract
Collaborative learning has gained great popularity due to its benefit of data privacy protection: participants can jointly train a Deep Learning model without sharing their training sets. However, recent works discovered that an adversary can fully recover the sensitive training samples from the shared gradients. Such reconstruction attacks pose severe threats to collaborative learning. Hence, effective mitigation solutions are urgently desired. In this paper, we systematically analyze existing reconstruction attacks and propose to leverage data augmentation to defeat these attacks: by preprocessing sensitive images with carefully-selected transformation policies, it becomes infeasible for the adversary to extract training samples from the corresponding gradients. We first design two new metrics to quantify the impacts of transformations on data privacy and model usability. With the two metrics, we design a novel search method to automatically discover qualified policies from a given data augmentation library. Our defense method can be further combined with existing collaborative training systems without modifying the training protocols. We conduct comprehensive experiments on various system settings. Evaluation results demonstrate that the policies discovered by our method can defeat state-of-the-art reconstruction attacks in collaborative learning, with high efficiency and negligible impact on the model performance.
Wei Gao 0064, Shangwei Guo, Tianwei Zhang 0004, Tao Xiang 0001, Han Qiu 0001, Yonggang Wen 0001, Yang Liu 0003
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 ADS-Lead: Lifelong Anomaly Detection in Autonomous Driving Systems
abstract
Autonomous Vehicles (AVs) are closely connected in the Cooperative Intelligent Transportation System (C-ITS). They are equipped with various sensors and controlled by Autonomous Driving Systems (ADSs) to provide high-level autonomy. The vehicles exchange different types of real-time data with each other, which can help reduce traffic accidents and congestion, and improve the efficiency of transportation systems. However, when interacting with the environment, AVs suffer from a broad attack surface, and the sensory data are susceptible to anomalies caused by faults, sensor malfunctions, or attacks, which may jeopardize traffic safety and result in serious accidents. In this paper, we proposeADS-Lead, an efficient collaborative anomaly detection methodology to protect the lane-following mechanism of ADSs.ADS-Leadis equipped with a novel transformer-based one-class classification model to identify time series anomalies (GPS spoofing threat) and adversarial image examples (traffic sign and lane recognition attacks). Besides, AVs inside the C-ITS form a cognitive network, enabling us to apply the federated learning technology to our anomaly detection method, where the vehicles in the C-ITS jointly update the detection model with higher model generalization and data privacy. Experiments on Baidu Apollo and two public data sets (GTSRB and Tumsimple) indicate that our method can not only detect sensor anomalies effectively and efficiently but also outperform state-of-the-art anomaly detection methods.
Xingshuo Han, Yuan Zhou 0005, Kangjie Chen, Han Qiu 0001, Meikang Qiu, Yang Liu 0003, Tianwei Zhang 0004
IEEE Trans. Intell. Transp. Syst.4
2023 System Log Parsing: A Survey
abstract
Modern information and communication systems have become increasingly challenging to manage. The ubiquitous system logs contain plentiful information and are thus widely exploited as an alternative source for system management. As log files usually encompass large amounts of raw data, manually analyzing them is laborious and error-prone. Consequently, many research endeavors have been devoted to automatic log analysis. However, these works typically expect structured input and struggle with the heterogeneous nature of raw system logs. Log parsing closes this gap by converting the unstructured system logs to structured records. Many parsers were proposed during the last decades to accommodate various log analysis applications. However, due to the ample solution space and lack of systematic evaluation, it is not easy for practitioners to find ready-made solutions that fit their needs. This paper aims to provide a comprehensive survey on log parsing. We begin with an exhaustive taxonomy of existing log parsers. Then we empirically analyze the critical performance and operational features for 17 open-source solutions both quantitatively and qualitatively, and whenever applicable discuss the merits of alternative approaches. We also elaborate on future challenges and discuss the relevant research directions. We envision this survey as a helpful resource for system administrators and domain experts to choose the most desirable open-source solution or implement new ones based on application-specific requirements.
Tianzhu Zhang 0002, Han Qiu 0001, Gabriele Castellano, Myriana Rifai, Chung Shue Chen, Fabio Pianese
IEEE Trans. Knowl. Data Eng.2
2022 An MRC Framework for Semantic Role Labeling
abstract
Semantic Role Labeling (SRL) aims at recognizing the predicate-argument structure of a sentence and can be decomposed into two subtasks: predicate disambiguation and argument labeling. Prior work deals with these two tasks independently, which ignores the semantic connection between the two tasks. In this paper, we propose to use the machine reading comprehension (MRC) framework to bridge this gap. We formalize predicate disambiguation as multiple-choice machine reading comprehension, where the descriptions of candidate senses of a given predicate are used as options to select the correct sense. The chosen predicate sense is then used to determine the semantic roles for that predicate, and these semantic roles are used to construct the query for another MRC model for argument labeling. In this way, we are able to leverage both the predicate semantics and the semantic role semantics for argument labeling. We also propose to select a subset of all the possible semantic roles for computational efficiency. Experiments show that the proposed framework achieves state-of-the-art or comparable results to previous work.
Jiwei Li 0001, Yuxian Meng, Xiaofei Sun 0001, Han Qiu 0001, Guoyin Wang 0002, Jun He 0008
COLING5
2022 Compressive sensing based asymmetric semantic image compression for resource-constrained IoT system
abstract
The widespread application of Internet-of-Things (IoT) and deep learning have made machine-to-machine semantic communication possible. However, it remains challenging to deploy DNN model on IoT devices, due to their limited computing and storage capacity. In this paper, we propose Compressed Sensing based Asymmetric Semantic Image Compression (CS-ASIC) for resource-constrained IoT systems, which consists of a lightweight front encoder and a deep iterative decoder offloaded at the server. We further consider a task-oriented scenario and optimize CS-ASIC for the semantic recognition tasks. The experiment results demonstrate that CS-ASIC achieves considerable data-semantic rate-distortion trade-off, and low encoding complexity over prevailing codecs.
Yujun Huang, Bin Chen 0011, Jianghui Zhang, Han Qiu 0001, Shutao Xia
DAC4
2022 Improving Adversarial Robustness of 3D Point Cloud Classification Models
Guowen Xu, Han Qiu 0001, Ruan He, Jiwei Li 0001, Tianwei Zhang 0004
ECCV (4)3
2022 Improved DC Estimation for JPEG Compression Via Convex Relaxation
abstract
Mass image transmission has undergone an explosion of growth with the development of the internet, DCT-based lossy image compression like JPEG is pervasively conducted to save the transmission bandwidth. Recently, DCT-domain coefficient estimation approaches have been proposed to further improve the compression ratio by discarding DC coefficients at the sender’s end while recovering them at the receiver’s end via DC estimation. However, known DC estimation needs to enumerate all possible DC coefficients. Consequently, they are limited and resource-consuming due to the low delay requirements in real-time transmission. In this paper, we propose an improved DC estimation method via convex relaxation, which achieves state-of-the-art performance in terms of both recovery image quality and time complexity. Extensive experiments across various data sets demonstrate the advantages of our method.
Jianghui Zhang, Bin Chen 0011, Yujun Huang, Han Qiu 0001, Zhi Wang 0001, Shutao Xia
ICIP4
2022 BET: black-box efficient testing for convolutional neural networks
abstract
It is important to test convolutional neural networks (CNNs) to identify defects (e.g. error-inducing inputs) before deploying them in security-sensitive scenarios. Although existing white-box testing methods can effectively test CNN models with high neuron coverage, they are not applicable to privacy-sensitive scenarios where full knowledge of target CNN models is lacking. In this work, we propose a novel Black-box Efficient Testing (BET) method for CNN models. The core insight of BET is that CNNs are generally prone to be affected by continuous perturbations. Thus, by generating such continuous perturbations in a black-box manner, we design a tunable objective function to guide our testing process for thoroughly exploring defects in different decision boundaries of the target CNN models. We further design an efficiency-centric policy to find more error-inducing inputs within a fixed query budget. We conduct extensive evaluations with three well-known datasets and five popular CNN structures. The results show that BET significantly outperforms existing white-box and black-box testing methods considering the effective error-inducing inputs found in a fixed query/inference budget. We further show that the error-inducing inputs found by BET can be used to fine-tune the target model, improving its accuracy by up to 3%.
Jialai Wang, Han Qiu 0001, Hengkai Ye, Qi Li 0002, Zongpeng Li, Chao Zhang 0008
ISSTA2
2022 jTrans: jump-aware transformer for binary code similarity detection
abstract
Binary code similarity detection (BCSD) has important applications in various fields such as vulnerabilities detection, software component analysis, and reverse engineering. Recent studies have shown that deep neural networks (DNNs) can comprehend instructions or control-flow graphs (CFG) of binary code and support BCSD. In this study, we propose a novel Transformer-based approach, namely jTrans, to learn representations of binary code. It is the first solution that embeds control flow information of binary code into Transformer-based language models, by using a novel jump-aware representation of the analyzed binaries and a newly-designed pre-training task. Additionally, we release to the community a newly-created large dataset of binaries, BinaryCorp, which is the most diverse to date. Evaluation results show that jTrans outperforms state-of-the-art (SOTA) approaches on this more challenging dataset by 30.5% (i.e., from 32.0% to 62.5%). In a real-world task of known vulnerability searching, jTrans achieves a recall that is 2X higher than existing SOTA baselines.
Hao Wang 0226, Wenjie Qu 0001, Gilad Katz, Wenyu Zhu, Han Qiu 0001, Jianwei Zhuge, Chao Zhang 0008
ISSTA6
2022 Mitigating Targeted Bit-Flip Attacks via Data Augmentation: An Empirical Study
Wencheng Chen, Han Qiu 0001, Meikang Qiu
KSEM (3)4
2021 DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
abstract
Public resources and services (e.g., datasets, training platforms, pre-trained models) have been widely adopted to ease the development of Deep Learning-based applications. However, if the third-party providers are untrusted, they can inject poisoned samples into the datasets or embed backdoors in those models. Such an integrity breach can cause severe consequences, especially in safety- and security-critical applications. Various backdoor attack techniques have been proposed for higher effectiveness and stealthiness. Unfortunately, existing defense solutions are not practical to thwart those attacks in a comprehensive way.
Han Qiu 0001, Yi Zeng 0005, Shangwei Guo, Tianwei Zhang 0004, Meikang Qiu, Bhavani Thuraisingham
AsiaCCS1
2021 Privacy-Preserving Collaborative Learning With Automatic Transformation Search
abstract
Collaborative learning has gained great popularity due to its benefit of data privacy protection: participants can jointly train a Deep Learning model without sharing their training sets. However, recent works discovered that an adversary can fully recover the sensitive training samples from the shared gradients. Such reconstruction attacks pose severe threats to collaborative learning. Hence, effective mitigation solutions are urgently desired.In this paper, we propose to leverage data augmentation to defeat reconstruction attacks: by preprocessing sensitive images with carefully-selected transformation policies, it becomes infeasible for the adversary to extract any useful information from the corresponding gradients. We design a novel search method to automatically discover qualified policies. We adopt two new metrics to quantify the impacts of transformations on data privacy and model usability, which can significantly accelerate the search speed. Comprehensive evaluations demonstrate that the policies discovered by our method can defeat existing reconstruction attacks in collaborative learning, with high efficiency and negligible impact on the model performance.
Wei Gao 0064, Shangwei Guo, Tianwei Zhang 0004, Han Qiu 0001, Yonggang Wen 0001, Yang Liu 0003
CVPR4
2021 Fine-tuning Is Not Enough: A Simple yet Effective Watermark Removal Attack for DNN Models
abstract
Watermarking has become the tendency in protecting the intellectual property of DNN models. Recent works, from the adversary's perspective, attempted to subvert watermarking mechanisms by designing watermark removal attacks. However, these attacks mainly adopted sophisticated fine-tuning techniques, which have certain fatal drawbacks or unrealistic assumptions. In this paper, we propose a novel watermark removal attack from a different perspective. Instead of just fine-tuning the watermarked models, we design a simple yet powerful transformation algorithm by combining imperceptible pattern embedding and spatial-level transformations, which can effectively and blindly destroy the memorization of watermarked models to the watermark samples. We also introduce a lightweight fine-tuning strategy to preserve the model performance. Our solution requires much less resource or knowledge about the watermarking scheme than prior works. Extensive experimental results indicate that our attack can bypass state-of-the-art watermarking solutions with very high success rates. Based on our attack, we propose watermark augmentation techniques to enhance the robustness of existing watermarks.
Shangwei Guo, Tianwei Zhang 0004, Han Qiu 0001, Yi Zeng 0005, Tao Xiang 0001, Yang Liu 0003
IJCAI3
2021 Adversarial Attacks Against Network Intrusion Detection in IoT Systems
abstract
Deep learning (DL) has gained popularity in network intrusion detection, due to its strong capability of recognizing subtle differences between normal and malicious network activities. Although a variety of methods have been designed to leverage DL models for security protection, whether these systems are vulnerable to adversarial examples (AEs) is unknown. In this article, we design a novel adversarial attack against DL-based network intrusion detection systems (NIDSs) in the Internet-of-Things environment, with only black-box accesses to the DL model in such NIDS. We introduce two techniques: 1) model extraction is adopted to replicate the black-box model with a small amount of training data and 2) a saliency map is then used to disclose the impact of each packet attribute on the detection results, and the most critical features. This enables us to efficiently generate AEs using conventional methods. With these tehniques, we successfully compromise one state-of-the-art NIDS, Kitsune: the adversary only needs to modify less than 0.005% of bytes in the malicious packets to achieve an average 94.31% attack success rate.
Han Qiu 0001, Tianwei Zhang 0004, Gérard Memmi, Meikang Qiu
IEEE Internet Things J.1
2021 Toward Secure and Efficient Deep Learning Inference in Dependable IoT Systems
abstract
The rapid development of deep learning (DL) enables resource-constrained systems and devices [e.g., Internet of Things (IoT)] to perform sophisticated artificial intelligence (AI) applications. However, AI models, such as deep neural networks (DNNs), are known to be vulnerable to adversarial examples (AEs). Past works on defending against AEs require heavy computations in the model training or inference processes, making them impractical to be applied in IoT systems. In this article, we propose a novel method, Super-IoT, to enhance the security and efficiency of AI applications in distributed IoT systems. Specifically, Super-IoT utilizes a pixel drop operation to eliminate adversarial perturbations from the input and reduce network transmission throughput. Then, it adopts a sparse signal recovery method to reconstruct the dropped pixels and wavelet-based denoising method to reduce the artificial noise. Super-IoT is a lightweight method with negligible computation cost to IoT devices and little impact on the DNN model performance. Extensive evaluations show that it can outperform three existing AE defensive solutions against most of the AE attacks with better transmission efficiency.
Han Qiu 0001, Qinkai Zheng, Tianwei Zhang 0004, Meikang Qiu, Gérard Memmi
IEEE Internet Things J.1
2021 Novel denial-of-service attacks against cloud-based multi-robot systems
Yuan Xu 0034, Gelei Deng, Tianwei Zhang 0004, Han Qiu 0001, Yungang Bao
Inf. Sci.4
2021 A User-Centric Data Protection Method for Cloud Storage Based on Invertible DWT
abstract
Protection on end users’ data stored in Cloud servers becomes an important issue in today’s Cloud environments. In this paper, we present a novel data protection method combining Selective Encryption (SE) concept with fragmentation and dispersion on storage. Our method is based on the invertible Discrete Wavelet Transform (DWT) to divide agnostic data into three fragments with three different levels of protection. Then, these three fragments can be dispersed over different storage areas with different levels of trustworthiness to protect end users’ data by resisting possible leaks in Clouds. Thus, our method optimizes the storage cost by saving expensive, private, and secure storage spaces and utilizing cheap but low trustworthy storage space. We have intensive security analysis performed to verify the high protection level of our method. Additionally, the efficiency is proved by implementation of deploying tasks between CPU and General Purpose Graphic Processing Unit (GPGPU) in an optimized manner.
Han Qiu 0001, Hassan N. Noura, Meikang Qiu, Zhong Ming 0001, Gérard Memmi
IEEE Trans. Cloud Comput.1
2021 Deep Residual Learning-Based Enhanced JPEG Compression in the Internet of Things
abstract
With the development of big data and network technology, there are more use cases, such as edge computing, that require more secure and efficient multimedia big data transmission. Data compression methods can help achieving many tasks like providing data integrity, protection, as well as efficient transmission. Classical multimedia big data compression relies on methods like the spatial-frequency transformation for compressing with loss. Recent approaches use deep learning to further explore the limit of the data compression methods in communication constrained use cases like the Internet of Things (IoT). In this article, we propose a novel method to significantly enhance the transformation-based compression standards like JPEG by transmitting much fewer data of one image at the sender's end. At the receiver's end, we propose a two-step method by combining the state-of-the-art signal processing based recovery method with a deep residual learning model to recover the original data. Therefore, in the IoT use cases, the sender like edge device can transmit only 60% data of the original JPEG image without any additional calculation steps but the image quality can still be recovered at the receiver's end like cloud servers with peak signal-to-noise ratio over 31 dB.
Han Qiu 0001, Qinkai Zheng, Gérard Memmi, Meikang Qiu, Bhavani Thuraisingham
IEEE Trans. Ind. Informatics1
2021 Topological Graph Convolutional Network-Based Urban Traffic Flow and Density Prediction
abstract
With the development of modern Intelligent Transportation System (ITS), reliable and efficient transportation information sharing becomes more and more important. Although there are promising wireless communication schemes such as Vehicle-to-Everything (V2X) communication standards, information sharing in ITS still faces challenges such as the V2X communication overload when a large number of vehicles suddenly appeared in one area. This flash crowd situation is mainly due to the uncertainty of traffic especially in the urban areas during traffic rush hours and will significantly increase the V2X communication latency. In order to solve such flash crowd issues, we propose a novel system that can accurately predict the traffic flow and density in the urban area that can be used to avoid the V2X communication flash crowd situation. By combining the existing grid-based and graph-based traffic flow prediction methods, we use a Topological Graph Convolutional Network (ToGCN) followed with a Sequence-to-sequence (Seq2Seq) framework to predict future traffic flow and density with temporal correlations. The experimentation on a real-world taxi trajectory traffic data set is performed and the evaluation results prove the effectiveness of our method.
Han Qiu 0001, Qinkai Zheng, Mounira Msahli, Gérard Memmi, Meikang Qiu
IEEE Trans. Intell. Transp. Syst.1
2021 NFV Platforms: Taxonomy, Design Choices and Future Challenges
abstract
Due to the intrinsically inefficient service provisioning in traditional networks, Network Function Virtualization (NFV) keeps gaining attention from both industry and academia. By replacing the purpose-built, expensive, proprietary network equipment with software network functions consolidated on commodity hardware, NFV envisions a shift towards a more agile and open service provisioning paradigm. During the last few years, a large number of NFV platforms have been implemented in production environments that typically face critical challenges, including the development, deployment, and management of Virtual Network Functions (VNFs). Nonetheless, just like any complex system, such platforms commonly consist of abounding software and hardware components and usually incorporate disparate design choices based on distinct motivations or use cases. This broad collection of convoluted alternatives makes it extremely arduous for network operators to make proper choices. Although numerous efforts have been devoted to investigating different aspects of NFV, none of them specifically focused on NFV platforms or attempted to explore their design space. In this paper, we present a comprehensive survey on the NFV platform design. Our study solely targets existing NFV platform implementations. We begin with a top-down architectural view of the standard reference NFV platform and present our taxonomy of existing NFV platforms based on what features they provide in terms of a typical network function life cycle. Then we thoroughly explore the design space and elaborate on the implementation choices each platform opts for. We also envision future challenges for NFV platform design in the incoming 5G era. We believe that our study gives a detailed guideline for network operators or service providers to choose the most appropriate NFV platform based on their respective requirements. Our work also provides guidelines for implementing new NFV platforms.
Tianzhu Zhang 0002, Han Qiu 0001, Leonardo Linguaglossa, Walter Cerroni, Paolo Giaccone
IEEE Trans. Netw. Serv. Manag.2
2020 A Data Augmentation-Based Defense Method Against Adversarial Attacks in Neural Networks
Yi Zeng 0005, Han Qiu 0001, Gérard Memmi, Meikang Qiu
ICA3PP (2)2
2020 HAPE: A programmable big knowledge graph platform
Ruqian Lu, Chaoqun Fei, Chuanqing Wang, Shunfeng Gao, Han Qiu 0001, Songmao Zhang, Cun-gen Cao 0001
Inf. Sci.5
2020 Lightweight Selective Encryption for Social Data Protection Based on EBCOT Coding
abstract
Online social media today has a large number of users and has become a huge platform to collect and share the social data generated by the end users. In addition, based on the development of social applications, the social sensing system has been greatly developed to generate, transmit, and store, which is helping the prosperous of the social computing systems. However, the violation of the security of end users' social data stored and shared through the social computing system becomes a serious and urgent issue. The data protection on social media platforms is very different compared with the scenario of the traditional encryption algorithms, and most of the existing schemes are not suitable for data protection in the current social sensing and data-sharing system. In this article, we present a novel design based on the agnostic selective encryption concept to efficiently protect the social data based on the embedded block coding with optimized truncation system. By selectively encrypting only a small portion of the bitstreams in the middle layer of this coding system, a high level of protection and efficiency can both be achieved. We also experiment with our method on four common social data formats, and the security analysis tests are performed to verify the high protection level of our method.
Han Qiu 0001, Meikang Qiu, Meiqin Liu 0001, Zhong Ming 0001
IEEE Trans. Comput. Soc. Syst.1
2020 Secure Health Data Sharing for Medical Cyber-Physical Systems for the Healthcare 4.0
abstract
The recent spades of cyber attacks have compromised end-users' data security and privacy in Medical Cyber-Physical Systems (MCPS) in the era of Health 4.0. Traditional standard encryption algorithms for data protection are designed based on a viewpoint of system architecture rather than a viewpoint of end-users. As such encryption algorithms are transferring the protection on the data to the protection on the keys, data safety, and privacy will be compromised once the key is exposed. In this paper, we propose a secure data storage and sharing method consisted of a selective encryption algorithm combined with fragmentation and dispersion to protect the data safety and privacy even when both transmission media (e.g. cloud servers) and keys are compromised. This method is based on a user-centric design that protects the data on a trusted device such as the end-users' smartphone and lets the end-user control the access for data sharing. We also evaluate the performance of the algorithm on a smartphone platform to prove efficiency.
Han Qiu 0001, Meikang Qiu, Meiqin Liu 0001, Gérard Memmi
IEEE J. Biomed. Health Informatics1
2019 ChainIDE: A Cloud-Based Integrated Development Environment for Cross-Blockchain Smart Contracts
abstract
Blockchain has become novel solutions for many traditional issues in computer science and finance. Recently, with the release of Libra blockchain from Facebook, the decentralized finance concepts have another huge development. However, there are already many different blockchain systems existing now and the development of blockchain systems have become more and more complicated since each kind of blockchain development environment will consume time to be built. To make the programming on different blockchain systems more easily, we propose a novel cloud-based solution, namely ChainIDE, for the development of blockchain-based smart contracts on multiple kinds of blockchain systems. With chainIDE, cross-chain developing of smart contracts on different blockchain systems can be easily done without any time consumed by building the environment. Based on the operation statistics in this paper, we served more than 310,000 compilations in the past 30 days which makes us the most popular cloud-based cross-chain development platform in the world.
Han Qiu 0001, Victor C. M. Leung, Wei Cai 0002
CloudCom1
2019 DC coefficient recovery for JPEG images in ubiquitous communication systems
Han Qiu 0001, Gérard Memmi, Jian Xiong 0001
Future Gener. Comput. Syst.1
2019 All-Or-Nothing data protection for ubiquitous communication: Challenges and perspectives
Han Qiu 0001, Katarzyna Kapusta, Zhihui Lu 0002, Meikang Qiu, Gérard Memmi
Inf. Sci.1
2017 An Efficient Secure Storage Scheme Based on Information Fragmentation
abstract
In this paper, an efficient secure storage scheme is presented which aims to provide security to end-user's data while mostly storing it to public clouds. This proposed scheme is based on the invertible Discrete Wavelet Transform (DWT) to fragment data into two or three fragments with different levels of importance and protected accordingly. As a matter of fact, the most important fragment takes the smallest amount of storage space and can be stored in a user trusted area while the less important fragments take most of the storage space and are uploaded to public clouds. In order to reduce the required execution time, General Purpose Graphic Processing Unit (GPGPU) is employed for accelerating computation. Additionally, a benchmark was realized to compare between the proposed scheme and AES algorithm applied to the entire data.
Han Qiu 0001, Gérard Memmi, Hassan N. Noura
CSCloud1
2014 Fast Selective Encryption Method for Bitmaps Based on GPU Acceleration
abstract
In this paper, we are interested in image protection within limited calculation resources environment like a laptop with large amount images as input. Full traditional encryption of the data stream is not fast enough and takes too much CPU calculation resource in such an environment. We derive a new solution combined selective encryption with current GPGPU (General Purpose Graphic Process Unit) acceleration. After presenting related works, we introduce a new architecture and implementation of a selective encryption method by utilizing all calculation resources of a laptop including CPU and GPGPU. Then performance of our design is given and compared with traditional full encryption method.
Han Qiu 0001, Gérard Memmi
ISM1