VLDB 2026 Research / reviewers in the wild / expert
Jialing He
dblp:221/6024
· DBLP profile ↗
27ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0002-8643-0647ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 4 first-author · 6 since 2021Security and privacy · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MegaScale-Data: Scaling DataLoader for Multisource Large Foundation Model TrainingabstractModern frameworks for training large foundation models (LFMs) employ dataloaders in a data-parallel manner, with each loader processing a disjoint subset of training data. When preparing data for LFM training that originates from multiple, distinct sources, two fundamental challenges arise. First, due to the quadratic computational complexity of the attention operator, the non-uniform sample distribution over data-parallel ranks leads to significant workload imbalance among dataloaders, degrading the training efficiency. Second, supporting diverse data sources requires per-dataset file access states that are redundantly replicated across parallel loaders, consuming excessive memory. This also hinders dynamic data mixing (e.g., curriculum learning) and causes redundant access/memory overhead in hybrid parallelism. Juntao Zhao 0002, Borui Wan, Lei Zuo 0004, Junda Feng, Jianyu Jiang, Yangrui Chen, Shuaishuai Cao, Jialing He, Kaihua Jiang, Shibiao Nong, Yanghua Peng, Haibin Lin, Chuan Wu 0001 |
EuroSys | 10 |
| 2026 | HiMoE-CCTC: A Hierarchical Mixture-of-Experts Framework for Citizen Complaint Text Classification
Jialing He, Jiahe Liu, Yanming Gong |
PAKDD (2) | 1 |
| 2026 | Revolutionizing electricity theft detection: enhanced accuracy through NILM and multi-source data fusionabstractAbstract Electricity theft detection seeks to thwart the illegal use of electricity, thereby safeguarding the safety and stability of the power system. Traditional methods, which typically rely on aggregated household consumption data to identify theft, often overlook the fact that household consumption is vulnerable to fluctuations in normal user behavior. This results in high false positive and false negative rates. To refine the accuracy, we propose a novel electricity theft detection method based on Non-Intrusive Load Monitoring (NILM) and multi-source data fusion. Our approach employs advanced NILM algorithms to cost-effectively extract individual appliance consumption data from aggregated power signals. We then integrate this data with household aggregate consumption data through a multi-source data fusion architecture. By analyzing the unique consumption patterns of different types of appliances, our approach identifies theft behaviors that cannot be detected by aggregate consumption data alone. Experimental results across three real-world datasets demonstrate that our method significantly outperforms single-source data-based benchmarks, achieving up to a 7.92% gain in F1-score and a 12.6% gain in Precision. Moreover, our method exhibits strong generalization ability across a series of typical machine learning models. Zhiwei Deng, Junsen Feng, Jialing He, Guozhu Meng, Tao Xiang 0001 |
Cybersecur. | 3 |
| 2026 | HP2: Hybrid and precision-guided filter pruning for CNN compression
Shangwei Guo, Jialing He, Run Wang 0001, Tao Xiang 0001 |
Inf. Sci. | 3 |
| 2026 | Efficient Blockchain-Based Steganography via Backcalculating Generative Adversarial NetworkabstractBlockchain-based steganography enables data hiding via encoding the covert data into a specific blockchain transaction field. However, previous works focus on the specific field-embedding methods while lacking a consideration on required field-generation embedding. In this paper, we propose a generic blockchain-based steganography framework (GBSF). The sender generates the required fields such as amount and fees, where the additional covert data is embedded to enhance the channel capacity. Based on GBSF, we design a reversible generative adversarial network (R-GAN) that utilizes the generative adversarial network with a reversible generator to generate the required fields and encode additional covert data into the input noise of the reversible generator. We then explore the performance flaw of R-GAN. To further improve the performance, we propose R-GAN withCounter-intuitive data preprocessing andCustom activation functions, namelyCCR-GAN. The counter-intuitive data preprocessing (CIDP) mechanism is used to reduce decoding errors in covert data, while it incurs gradient explosion for model convergence. The custom activation function named ClipSigmoid is devised to overcome the problem. Theoretical justification for CIDP and ClipSigmoid is also provided. We also develop a mechanism named T2C, which balances capacity and concealment. We conduct experiments using the transaction amount of the Bitcoin mainnet as the required field to verify the feasibility. We then apply the proposed schemes to other transaction fields and blockchains to demonstrate the scalability. Finally, we evaluate capacity and concealment for various blockchains and transaction fields and explore the trade-off between capacity and concealment. Experimental results demonstrate that R-GAN and CCR-GAN are able to enhance the channel capacity effectively and outperform state-of-the-art works. Zhuo Chen 0001, Jialing He, Jiacheng Wang 0001, Zehui Xiong, Tao Xiang 0001, Liehuang Zhu, Dusit Niyato |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2026 | FastBOC: Toward Efficient Covert Communication Merging Blockchain and Onion NetworksabstractCovert communication over public blockchains has emerged as a promising approach for secret data transmission. However, existing methods often suffer from high communication costs, low communication efficiency, and the risk of permanent data exposure. To overcome the above challenges, we propose FastBOC, a hybrid covert communication framework that integrates blockchain and onion networks. In FastBOC, the blockchain is employed as a covert signal channel to transmit lightweight signals, while the onion network handles high-capacity secret data transmission. This decoupling significantly reduces communication costs and avoids permanent data exposure on the blockchain. We further design an address-based encoding scheme and a dynamic port activation mechanism to enhance concealment. We implement FastBOC on the Ethereum testnet and conduct experiments to evaluate its concealment, efficiency, and cost. The results demonstrate that FastBOC (1) achieves strong concealment and (2) can transmit 1-Megabyte (MB) data within 20.11 seconds and reduce communication cost by 5–7 orders of magnitude compared to prior blockchain-based covert communication schemes. Xiangbo Yuan, Zhuo Chen 0001, Jialing He, Tao Xiang 0001, Liehuang Zhu |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2026 | Safeguarding ISAC Performance in Low-Altitude Wireless Networks Under Channel Access AttackabstractThe increasing saturation of terrestrial resources has driven the exploration of low-altitude applications such as air taxis. Low altitude wireless networks (LAWNs) serve as the foundation for these applications, and integrated sensing and communication (ISAC) constitutes one of the core technologies within LAWNs. However, the open nature of low-altitude airspace makes LAWNs vulnerable to malicious channel access attacks, which degrade the ISAC performance. Therefore, this paper develops a game-based framework to mitigate the influence of the attacks on LAWNs. Concretely, we first derive expressions of communication data’s signal-to-interference-plus-noise ratio and the age of information of sensing data under attack conditions, which serve as quality of service metrics. Then, we formulate the ISAC performance optimization problem as a Stackelberg game, where the attacker acts as the leader, and the legitimate drone and the ground ISAC base station act as second and first followers, respectively. On this basis, we design a backward induction algorithm that achieves the Stackelberg equilibrium while maximizing the utilities of all participants, thereby mitigating the attack-induced degradation of ISAC performance in LAWNs. We further prove the existence of the equilibrium. Simulation results show that the proposed algorithm outperforms existing baselines and a static Nash equilibrium benchmark, ensuring that LAWNs can provide reliable service for low-altitude applications. Jiacheng Wang 0001, Jialing He, Geng Sun 0001, Zehui Xiong, Dusit Niyato, Shiwen Mao, Dong In Kim 0001, Tao Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | Security-Aware Joint Sensing, Communication, and Computing Optimization in Low Altitude Wireless NetworksabstractAs terrestrial resources become increasingly saturated, the developing attention is gradually shifting from the ground to the low-altitude airspace, which supports many emerging applications such as urban air taxis and aerial inspection. For these applications, low-altitude wireless networks (LAWNs) are the foundation, with integrated sensing, communications, and computing (ISCC) being one of the core parts. However, the openness of low-altitude airspace poses a serious threat to communications, degrading ISCC performance and ultimately compromising the reliability of applications supported by LAWNs. To address these challenges, this paper studies joint performance optimization of ISCC while considering security of the communications. Specifically, we derive beampattern error, secrecy rate, and age of information (AoI) as performance metrics for sensing, secure communication, and computing. Building on these metrics, we formulate a multi-objective optimization problem, which aims to balance sensing and computing performance while enhancing the secrecy rate of communications. We then propose a deep Q-network (DQN)-based multi-objective evolutionary algorithm, which adaptively selects evolutionary operators according to the evolving optimization objectives, thereby leading to more effective solutions. Extensive simulations show that the proposed method brings an average performance gain of about 14% compared to existing methods, thereby ensuring ISCC performance for applications supported by LAWNs. Jiacheng Wang 0001, Changyuan Zhao, Jialing He, Geng Sun 0001, Weijie Yuan 0001, Dusit Niyato, Liehuang Zhu, Tao Xiang 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | S-RAG: A Novel Audit Framework for Detecting Unauthorized Use of Personal Data in RAG SystemsabstractRetrieval-Augmented Generation (RAG) systems combine external data retrieval with text generation and have become essential in applications requiring accurate and context-specific responses. However, their reliance on external data raises critical concerns about unauthorized collection and usage of personal information. To ensure compliance with data protection regulations like GDPR and detect improper use of data, we propose the Shadow RAG Auditing Data Provenance (S-RAG) framework. S-RAG enables users to determine whether their textual data has been utilized in RAG systems, even in black-box settings with no prior system knowledge. It is effective across open-source and closed-source RAG systems and resilient to defense strategies. Experiments demonstrate that S-RAG achieves an improvement in Accuracy by 19.9% (compared to the best baseline), while maintaining strong performance under adversarial defenses. Furthermore, we analyze how the auditor’s knowledge of the target system affects performance, offering practical insights for privacy-preserving AI systems. Our code is open-sourced online. Zhirui Zeng, Jiamou Liu, Meng-Fen Chiang, Jialing He, Zijian Zhang 0001 |
ACL (1) | 4 |
| 2025 | Overcoming Data Mining in Blockchain-Based Covert Communication: Transaction Withdrawal and Multisig EmbeddingabstractBlockchain-based covert communication (BCC) provides high reliability and anonymity by embedding secret data into blockchain transactions. However, existing BCC approaches still face three fundamental limitations: (i) data mining risk, since transactions containing the secret data are permanently recorded on-chain and may be detected perpetually; (ii) limited efficiency, as only small payloads (e.g., 256 bits) can be carried per transaction; and (iii) private key leakage, where receivers often need access to the sender’s private key and may incur private key exposure. To address these issues, we propose a novel covert communication model with transaction withdrawal (BCC-TW) and a multisig-based data embedding scheme (MUL-DE). BCC-TW prevents covert transactions from being confirmed by constructing higher-fee double-spend transactions, thereby ensuring that secret data only exists temporarily in the mempool. MUL-DE encodes data into redundant public keys of Bitcoin multisig addresses, thus enabling higher efficiency and avoiding private key exposure. We implement a prototype on Bitcoin testnet and evaluate its concealment and efficiency. Experimental results demonstrate that the proposed approach achieves strong indistinguishability against statistical and deep-learning-based detectors, improves communication efficiency up to 251 bits per public key, and significantly reduces cost compared with state-of-the-art baselines. Jialing He, Zhuo Chen 0001, Yijing Lin, Jiacheng Wang 0001, Liehuang Zhu, Zhu Han 0001, Rahim Tafazolli, Tao Xiang 0001 |
TrustCom | 1 |
| 2025 | Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding IndistinguishabilityabstractPre-trained models (PTMs) are widely adopted across various downstream tasks in the machine learning supply chain. Adopting untrustworthy PTMs introduces significant security risks, where adversaries can poison the model supply chain by embedding hidden malicious behaviors (backdoors) into PTMs. However, existing backdoor attacks to PTMs can only achieve partially task-agnostic and the embedded backdoors are easily erased during the fine-tuning process. This makes it challenging for the backdoors to persist and propagate through the supply chain. In this paper, we propose a novel and severer backdoor attack, TransTroj, which enables the backdoors embedded in PTMs to efficiently transfer in the model supply chain. In particular, we first formalize this attack as an indistinguishability problem between poisoned and clean samples in the embedding space. We decompose embedding indistinguishability into pre- and post-indistinguishability, representing the similarity of the poisoned and reference embeddings before and after the attack. Then, we propose a two-stage optimization that separately optimizes triggers and victim PTMs to achieve embedding indistinguishability. We evaluate TransTroj on four PTMs and six downstream tasks. Experimental results show that our method significantly outperforms SOTA task-agnostic backdoor attacks -- achieving nearly 100% attack success rate on most downstream tasks -- and demonstrates robustness under various system settings. Our findings underscore the urgent need to secure the model supply chain against such transferable backdoor attacks. The code is available at https://github.com/haowang-cqu/TransTroj Hao Wang 0227, Shangwei Guo, Jialing He, Hangcheng Liu, Tianwei Zhang 0004, Tao Xiang 0001 |
WWW | 3 |
| 2025 | Maintaining Privacy in Smart Grid: Utilizing the Adversarial Attack Paradigm to Counter Nonintrusive Load Monitoring ModelsabstractThe nonintrusive load monitoring (NILM) technique, through its use of various deep neural networks (DNNs), is capable of learning residential appliances’ usage patterns from networked smart meters. However, such learned information may pose a serious privacy risk to users. In response to this privacy concern, in this article, we introduce an innovative adversarial attack. This attack can effectively restrict the NILM models’ ability to dissect power signals while maintaining accurate electricity charges for users. Given that previous adversarial attacks—which are designed for image classifiers and regressors with one-time output—cannot adequately handle NILM models and regressors with time-series output, we formally present the attack objective by leveraging the unique characteristics of regression and time-series data. Our proposed solution algorithms for this attack objective can generate imperceptible perturbations, effectively misleading the prediction of NILM models. To further ensure accurate billing calculation, we refine the attack objective to a practical version and propose a post-process that can iteratively remove the added perturbation in a certain period without compromising attack effectiveness. Experimental results on two real-world datasets, REDD and UK-DALE, demonstrate the effectiveness, transferability, and practicality of our proposed adversarial attack scheme. Jialing He, Tao Xiang 0001, Tianhao Wu 0017, Zhuo Chen 0001, Ning Wang 0003, Shangwei Guo |
IEEE Internet Things J. | 1 |
| 2025 | Accurate, Secure, and Efficient Semi-Constrained Navigation Over Encrypted City MapsabstractNavigation services enable users to find the shortest path from a starting point$S$to a destination$D$, reducing time, gas, and traffic congestion. Still, navigation users risk the exposure of their sensitive location data. Our motivation arises from how users can accurately, securely, and efficiently navigate from$S$to$D$while passing through$k$unordered stops, i.e., midway locations with a non-fixed visiting order. In this work, we formally define Semi-Constrained Navigation (SCN) and present a novel scheme Hermes to achieve accurate, secure, and efficient SCN. Specifically, we propose a divide-and-conquer approach to strike a good balance between accuracy and efficiency. It recursively depth-first-searches the whole area (a navigation tree) and invokes five carefully-crafted strategies stop-by-stop to compute three subpaths in three sequential subareas. We construct a path-distance oracle to encrypt the road graph and securely implement the strategies by using homomorphic encryption and garble circuits. We formally prove the security in the random oracle model and analyze the search complexity to be less than$O(k^{2})$. We experiment over a real-world city map and compare with six baselines. Results show that path search with$k=4$among$N=1000$intersections requires 5.58 seconds with a 3.2% distance deviation rate and an 82.5% path similarity. Meng Li 0006, Yifei Chen 0005, Jingyu Wu, Zijian Zhang 0001, Jialing He, Liehuang Zhu, Mauro Conti, Xiaodong Lin 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Semantic and Precise Trigger Inversion: Detecting Backdoored Language ModelsabstractBackdoor attacks pose a serious security threat to Natural Language Processing (NLP) models, allowing adversaries to manipulate model outputs through hidden triggers. Although backdoor detection methods have been developed to address this issue, existing approaches based on trigger inversion are effective only for simple, visible triggers. These methods struggle to handle semantically enhanced, invisible triggers and often fail to provide accurate backdoor determinations due to reliance on unreliable heuristics, making it difficult to reliably distinguish backdoored models from benign ones. This presents a critical gap in current detection techniques. To address these challenges, we propose a novel trigger inversionSemInvthat consists of two key contributions: consistent semantics inversion and identifiable condition inspection. Consistent semantics inversion introduces a new regularization technique into the trigger optimization process, enabling more effective inversion of semantically constrained triggers. Identifiable condition inspection assesses the attack performance margin across different identifiable conditions, providing robust evidence for distinguishing backdoored models from benign ones. We evaluateSemInvusing the TrojAI round 6–8 datasets and demonstrate that it significantly outperforms state-of-the-art approaches in both backdoor detection accuracy and trigger inversion performance. Our method also proves effective against models with stealthy triggers, advancing the field of NLP security by offering a more comprehensive solution for identifying backdoor attacks. The code repository is in https://github.com/Bluedask/SemInv. Chunlong Xie, Jialing He, Ying Yang 0019, Shangwei Guo, Tianwei Zhang 0004, Tao Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Preventing Non-Intrusive Load Monitoring Privacy Invasion: A Precise Adversarial Attack Scheme for Networked Smart MetersabstractSmart grid, through networked smart meters employing the non-intrusive load monitoring (NILM) technique, can considerably discern the usage patterns of residential appliances. However, this technique also incurs privacy leakage. To address this issue, we propose an innovative scheme based on adversarial attack in this paper. The scheme effectively prevents NILM models from violating appliance-level privacy, while also ensuring accurate billing calculation for users. To achieve this objective, we overcome two primary challenges. First, as NILM models fall under the category of time-series regression models, direct application of traditional adversarial attacks designed for classification tasks is not feasible. To tackle this issue, we formulate a novel adversarial attack problem tailored specifically for NILM and providing a theoretical foundation for utilizing the Jacobian of the NILM model to generate imperceptible perturbations. Leveraging the Jacobian, our scheme can produce perturbations, which effectively misleads the signal prediction of NILM models to safeguard users' appliance-level privacy. The second challenge pertains to fundamental utility requirements, where existing adversarial attack schemes struggle to achieve accurate billing calculation for users. To handle this problem, we introduce an additional constraint, mandating that the sum of added perturbations within a billing period must be precisely zero. Experimental validation on real-world power datasets REDD and U.K.-DALE demonstrates the efficacy of our proposed solutions, which can significantly amplify the discrepancy between the output of the targeted NILM model and the actual power signal of appliances, and enable accurate billing at the same time. Additionally, our solutions exhibit transferability, making the generated perturbation signal from one target model applicable to other diverse NILM models. Jialing He, Jiacheng Wang 0001, Ning Wang 0003, Shangwei Guo, Liehuang Zhu, Dusit Niyato, Tao Xiang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | A Generic Blockchain-based Steganography Framework with High Capacity via Reversible GANabstractBlockchain-based steganography enables data hiding via encoding the covert data into a specific blockchain transaction field. However, previous works focus on the specific field-embedding methods while lacking a consideration on required field-generation embedding. In this paper, we propose GBSF, a generic framework for blockchain-based steganography. The sender generates the required fields, where the additional covert data is embedded to enhance the channel capacity. Based on GBSF, we design R-GAN that utilizes the generative adversarial network (GAN) with a reversible generator to generate the required fields and encode additional covert data into the input noise of the reversible generator. We then explore the performance flaw of R-GAN and introduce CCR-GAN as an improvement. CCR-GAN employs a counter-intuitive data preprocessing mechanism to reduce decoding errors in covert data. It incurs gradient explosion for model convergence and we design a custom activation function. We conduct experiments using the transaction amount of the Bitcoin mainnet as the required field. The results demonstrate that R-GAN and CCR-GAN allow to embed 11-bit (embedding rate of 17.2%) and 24-bit (embedding rate of 37.5%) covert data within a transaction amount, and enhance the channel capacity of state-of-the-art works by 4.30% to 91.67% and 9.38% to 200.00%, respectively. Zhuo Chen 0001, Liehuang Zhu, Peng Jiang 0007, Jialing He, Zijian Zhang 0001 |
INFOCOM | 4 |
| 2024 | EvilEdit: Backdooring Text-to-Image Diffusion Models in One SecondabstractText-to-image (T2I) diffusion models enjoy great popularity and many individuals and companies build their applications based on publicly released T2I diffusion models. Previous studies have demonstrated that backdoor attacks can elicit T2I diffusion models to generate unsafe target images through textual triggers. However, existing backdoor attacks typically demand substantial tuning data for poisoning, limiting their practicality and potentially degrading the overall performance of T2I diffusion models. To address these issues, we propose EvilEdit, a training-free and data-free backdoor attack against T2I diffusion models. EvilEdit directly edits the projection matrices in the cross-attention layers to achieve projection alignment between a trigger and the corresponding backdoor target. We preserve the functionality of the backdoored model using a protected whitelist to ensure the semantic of non-trigger words is not accidentally altered by the backdoor. We also propose a visual target attack EvilEdit VTA, enabling adversaries to use specific images as backdoor targets. We conduct empirical experiments on Stable Diffusion and the results demonstrate that the EvilEdit can backdoor T2I diffusion models within one second with up to 100% success rate. Furthermore, our EvilEdit modifies only 2.2% of the parameters and maintains the model's performance on benign prompts. Our code is available at https://github.com/haowang-cqu/EvilEdit. Hao Wang 0227, Shangwei Guo, Jialing He, Kangjie Chen, Shudong Zhang, Tianwei Zhang 0004, Tao Xiang 0001 |
ACM Multimedia | 3 |
| 2024 | Contrast-Then-Approximate: Analyzing Keyword Leakage of Generative Language ModelsabstractThere is an increasing tendency to fine-tune large-scale pre-trained language models (LMs) using small private datasets to improve their capability for downstream applications. In this paper, we systematically analyze the pre-train and then fine-tune the process of generative LMs and show that the fine-tuned LMs would leak sensitive keywords of the private datasets even without any prior knowledge of the downstream tasks. Specifically, we propose a novel and efficient keyword inference attack framework to accurately and maximally recover sensitive keywords. Owing to the fine-tuning process, pre-trained and fine-tuned models might respond differently to identical input prefixes. To identify potential sensitive sentences for training the fine-tuend LM, we introduce a contrast difference score that assesses the response variations between a pre-trained LM and its corresponding fine-tuned LM. Following this, we iteratively fine-tune the pre-trained model using these sensitive sentences to minimize the disparity between the target model and the pre-trained model, thereby maximizing the number of inferred sensitive keywords. We implement two types of keyword inference attacks (i.e., domain and private) according to our framework and conduct comprehensive experiments on three downstream applications to evaluate the performance. The experimental results demonstrate that our domain keyword inference attack achieves a precision of 85%, while our private keyword inference attack can extract highly sensitive personal information for a significant number of individuals (approximately 0.3% of all customers in the private fine-tuning dataset, which contains 40,000 pieces of personal information). Zhirui Zeng, Tao Xiang 0001, Shangwei Guo, Jialing He, Qiao Zhang 0002, Guowen Xu, Tianwei Zhang 0004 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Anonymous, Secure, Traceable, and Efficient Decentralized Digital ForensicsabstractDigital forensics is crucial to fight crimes around the world. Decentralized Digital Forensics (DDF) promotes it to another level by channeling the power of blockchain into digital investigations. In this work, we focus on the privacy and security of DDF. Our motivations arise from (1) how to track an anonymous-and-malicious data user who leaks only a part of the previously requested data, (2) how to achieve access control while protecting data from untrusted data centers, and (3) how to enable efficient and secure search on the blockchain. To address these issues, we propose Themis: an anonymous and secure DDF scheme with traceable anonymity, private access control, and efficient search. Our framework is boosted by establishing a Trusted Execution Environment in each authority (blockchain node) for securing the uploading, requesting, and searching. To instantiate the framework, we design a secure and robust watermarking scheme in conjunction with decentralized anonymous authentication, a private and fine-grained access control scheme, and an efficient and secure search scheme based on a dynamically updated data structure. We formally define and prove the privacy and security of Themis. We build a prototype with Ethereum and Intel SGX2 to evaluate its performance, which supports processing data from a considerable number of data providers and investigators. Meng Li 0006, Yanzhe Shen, Guixin Ye, Jialing He, Zijian Zhang 0001, Liehuang Zhu, Mauro Conti |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | ABDP: Accurate Billing on Differentially Private Data Reporting for Smart GridsabstractWhile smart grid significantly facilitates energy efficiency by using users’ power consumption data, it poses privacy leakage risk for user personal behaviors. Differential privacy (DP) has emerged as a promising solution to address this issue. However, existing approaches suffer from severe data utility degradation due to the intensive noise introduced by DP. Additionally, some of these methods are vulnerable to security attacks. To bridge this gap, in this paper, we propose ABDP (accuratebilling-enableddifferentiallyprivate), a mechanism that achieves high-strength DP while ensuring accurate aggregation and billing operations without compromising security. In particular, we propose aggregated and individual noise cancellation algorithms to counteract the negative effects of noise on data utility. Specifically, our ABDP ensures precise aggregation and accurate billing calculations for the power grid and individual users, respectively Furthermore, we present a Blockchain smart contract exploiting the pseudo random function to enforce a fair and secure data reporting process. Theoretical analysis is provided to evaluate the privacy and security guarantees of ABDP. Experimental results on real-world datasets, namely NERL-DATA and REDD, demonstrate that ABDP achieves error-free aggregation and billing calculation, offers arbitrary intensity privacy protection against non-intrusive load monitoring and filtering attacks, and outperforms existing state-of-the-art approaches. Jialing He, Ning Wang 0003, Tao Xiang 0001, Yiqiao Wei, Zijian Zhang 0001, Meng Li 0006, Liehuang Zhu |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | MSDC: Exploiting Multi-State Power Consumption in Non-intrusive Load Monitoring Based on a Dual-CNN ModelabstractNon-intrusive load monitoring (NILM) aims to decompose aggregated electrical usage signal into appliance-specific power consumption and it amounts to a classical example of blind source separation tasks. Leveraging recent progress on deep learning techniques, we design a new neural NILM model {\em Multi-State Dual CNN} (MSDC). Different from previous models, MSDC explicitly extracts information about the appliance's multiple states and state transitions, which in turn regulates the prediction of signals for appliances. More specifically, we employ a dual-CNN architecture: one CNN for outputting state distributions and the other for predicting the power of each state. A new technique is invented that utilizes conditional random fields (CRF) to capture state transitions. Experiments on two real-world datasets REDD and UK-DALE demonstrate that our model significantly outperform state-of-the-art models while having good generalization capacity, achieving 6%-10% MAE gain and 33%-51% SAE gain to unseen appliances. Jialing He, Jiamou Liu, Zijian Zhang 0001, Yang Chen 0028, Bakhadyr Khoussainov, Liehuang Zhu |
AAAI | 1 |
| 2023 | Contrastive Fusion Representation: Mitigating Adversarial Attacks on VQA ModelsabstractVisual Question Answering (VQA) is the vision-language task of answering text-based questions presented in an image and has been advanced by the remarkable success of multimodal deep networks. Similar to unimodal networks, multimodal VQA models are also vulnerable to adversarial examples, which raises severe threats to the corresponding applications. Although several adversarial training methods have been proposed, most of them focus on improving the generalization ability of VQA models on clean samples instead of mitigating the adversarial attacks. In this paper, we systemically analyze the core structure of multimodal VQA networks and propose a novel adversarial training algorithm to mitigate adversarial attacks on VQA models. Specifically, our key component is a regularization term with our carefully designed Contrastive Fusion Representation (CFR), which can reduce the sensitivity of VQA models to adversarial perturbations of both the vision and language inputs. We further enhance the adversarial training with augmented CFRs. Comprehensive experimental results show that our method can mitigate adversarial attacks as well as preserve the generalization ability on clean samples under various system settings and outperforms other defense methods. Jialing He, Hangcheng Liu, Shangwei Guo, Biwen Chen, Ning Wang 0003, Tao Xiang 0001 |
ICME | 1 |
| 2022 | Practical Blockchain-Based Steganographic Communication Via Adversarial AI: A Case Study In BitcoinabstractAbstract With the development of 5G, the wireless Internet of Things (IoT) has become possible; how to provide privacy protections for the communication of IoT devices in a more vulnerable wireless transmission environment is a huge challenge. Thus, steganography is introduced as a safe and effective technology. Blockchain systems have been widely used in the area of steganography. Several works attempted to embed covert data into transactions in public blockchain systems such as Bitcoin, Ethereum and Monero. However, most of them merely focus on putting covert data into certain fields in transactions based on cryptographic algorithms. In this paper, a Covert Transaction Recognition (CTR) model is proposed by the Text Convolutional Neural Networks and Back Propagation Neural Networks. When utilizing the covert data-embedded field for recognizing, our CTR model can attain 0.79 precision and 0.83 recall on average for seven covert transaction construction schemes. The precision and recall can increase by at most 43 and 47%, respectively, if other unembedded fields were additionally exploited for recognition. We further propose a Practical Covert Transaction Construction (PCTC) model. This model fixes the contents in the embedded fields of the constructed transactions, and generates the contents in other fields using Generative Adversarial Networks. Experimental results demonstrated that the precision and recall are greatly decreased when identifying the covert transactions generated by our PCTC model. The data underlying this article are available in ‘covert-transaction-model’, at https://github.com/1997mint/covert-transaction-model. Minxian Wang, Zijian Zhang 0001, Jialing He, Feng Gao 0019, Meng Li 0006, Shubin Xu, Liehuang Zhu |
Comput. J. | 3 |
| 2022 | Proof of Continuous Work for Reliable Data Storage Over Permissionless BlockchainabstractBitcoin first proposed the Nakamoto consensus that applies proof of work into the blockchain structure to build a trustless append-only ledger. The Nakamoto consensus solves the distributed consistency problem in the public network but wastes too much computing power. Instead of consuming computing resources, many improved consensus schemes address this problem by leveraging miners’ storage resources. However, these schemes fail to let miners store data constantly and usually rely on a dealer to assign data, which is hard to build a reliable decentralized storage system. In this article, we first design a variant consensus algorithm named Proof of Continuous Work (PoCW) with a storage-related incentive mechanism. Miners can accumulate mining advantage by continuously submitting proofs of storage. Then, we present a hash ring-based data allocation algorithm using the blockchain’s state. Combined with both of them, we build a reliable blockchain-based storage system without relying on any third parties. The theoretical analysis and simulation results demonstrate that the proposed system has higher reliability than those existing systems, and we also give practical suggestions about system parameters. Finally, we discuss additional benefits that our system brings. Zijian Zhang 0001, Jialing He, Liran Ma, Liehuang Zhu, Meng Li 0006, Bakhadyr Khoussainov |
IEEE Internet Things J. | 3 |
| 2022 | Video Aficionado: We Know What You Are WatchingabstractUsers enjoy the convenience of watching videos on smart devices. However, video watching records can be exposed without users’ knowledge and be exploited to infer private information. In this paper, we design and implement a new side-channel attack system, namedvideo aficionado, which can identify video watching information without violating any access control policies on Android. Our system only needs to collect power consumption data of a video playing app, which does not require explicit user permission. The collected data is sent to a remote server, where noise is cleaned and identified by a multi-layer perceptron (MLP) trained classifier. We evaluate our proposed system through a set of carefully designed experiments. Experimental results demonstrate that our system can make an identification with 74.5 percent accuracy on average for each 20-second power measurement segment out of 3918 segments collected from 20 videos. To the best of our knowledge, video aficionado is the first real-time power consumption-based video identification system on smart devices. Jialing He, Zijian Zhang 0001, Liran Ma, Bakhadyr Khoussainov, Liehuang Zhu |
IEEE Trans. Mob. Comput. | 1 |
| 2019 | An Efficient and Accurate Nonintrusive Load Monitoring Scheme for Power ConsumptionabstractNonintrusive load monitoring (NILM) has attracted tremendous attention owing to its cost efficiency in electricity and sustainable development. NILM aims at acquiring individual appliance power consumption rates using an aggregated power smart meter reading. Each individual appliance's power consumption enables users to monitor their electricity usage habits for rational saving strategies. This is also a valuable tool for detecting failure in appliances. However, the major barriers facing NILM schemes are issues of accurately capturing the features of each appliance and decreasing the computing time. Motivated by these challenges, we propose a new, efficient, and accurate NILM scheme, consisting of a learning step and a decomposing step. In the learning step, we propose the fast search-and-find of density peaks (FSFDPs) clustering algorithm aimed at capturing the features of the power consumption patterns of appliances. In the decomposing step, we propose a genetic algorithm (GA)-based matching algorithm to estimate the power consumption of each individual appliance using the aggregated power reading. Using elitist and catastrophic strategies, this step reduces the searching space to achieve considerable efficiency. Experimental results using the reference energy disaggregation dataset (REDD) indicate that our proposed scheme promotes accuracy by 10% and reduces the decomposing time by half. Jialing He, Zijian Zhang 0001, Liehuang Zhu, Zhesi Zhu, Jiamou Liu, Keke Gai |
IEEE Internet Things J. | 1 |
| 2017 | An Efficient Sparse Coding-Based Data-Mining Scheme in Smart Grid
Dongshu Wang, Jialing He, Mussadiq Abdul Rahim, Zijian Zhang 0001, Liehuang Zhu |
MSN | 2 |