Yong Xiang 0001

dblp:98/2912-1 · DBLP profile ↗
← Back
228ranked-venue papers
6as first author
130since 2021 · last 2026
0000-0003-3545-7863ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 56 · 31 since 2021Artificial intelligence and machine learning · 54 · 4 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 12 since 2021Systems, architecture and hardware · 20 · 10 since 2021Databases, data management, data science and information retrieval · 20 · 9 since 2021Software engineering, systems software and programming languages · 19 · 18 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 13 since 2021Security and privacy · 18 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Theory of computation · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Fragment-energy audio watermarking resilient to de-synchronization attacks
Juan Zhao 0007, Tianrui Zong, Iynkaran Natgunanathan, Yong Xiang 0001, Guang Hua 0001, Longxiang Gao, Wanlei Zhou 0001
Expert Syst. Appl.4
2026 Sentinel: Dynamic Knowledge Distillation for Personalized Federated Intrusion Detection in Heterogeneous IoT Networks
abstract
Federated learning (FL) offers a privacy-preserving paradigm for distributed machine learning, but its application to intrusion detection systems (IDS) in IoT networks is hindered by severe class imbalance, highly non-IID data, and high communication overhead. These challenges severely degrade the performance of conventional FL methods in real-world network traffic classification. To overcome these limitations, we propose Sentinel, a personalized federated IDS (pFed-IDS) framework that incorporates a dual-model architecture on each client, consisting of a high-capacity personalized teacher and a lightweight globally shared student model. This design balances deep local adaptation with efficient global aggregation while preserving privacy and reducing communication overhead by transmitting only the compact student model. Sentinel integrates three key mechanisms to ensure robust performance: bidirectional knowledge distillation with adaptive temperature scheduling, lightweight multi-level feature alignment between teacher and student representations, and a class-balanced loss to handle highly skewed traffic. On the server side, normalized gradient aggregation with equal client weighting mitigates client drift and improves fairness across clients. Extensive experiments on the IoTID20 and 5GNIDD benchmark datasets demonstrate that Sentinel significantly outperforms state-of-the-art federated baselines under extreme data heterogeneity, while lowering communication overhead.
Keshav Sood, Pachamuthu Rajalakshmi, Yong Xiang 0001
IEEE Internet Things J.4
2026 Make Identity Indistinguishable: Utility-Preserving Face Dataset Publication With Provable Privacy Guarantees
abstract
With the popularity of personal devices, there are abundant valuable face image datasets in the industry, which provides opportunities for the development of visual models. However, privacy concerns related to identity sensitive information hinder face datasets sharing. Despite existing works dedicated to removing identity sensitive information from images, they either lack provable privacy guarantees or compromise crucial face dataset utilities, e.g., identity correlation and image naturalness. To overcome these weaknesses, we propose a novel face dataset publication scheme that protects face images by obfuscating face features. The obfuscated features still retain a certain level of correlation, allowing the protected dataset to be used for training. In the process of obfuscating the features, we design a novel metric differential privacy mechanism, which can enhance the correlation between features while ensuring privacy. Furthermore, we construct a latent diffusion model with identity and attribute as inputs to improve the naturalness of generated images. Extensive experimental results and theoretical analysis demonstrate our scheme significantly outperforms existing works in providing privacy protection while maintaining high dataset utility for downstream tasks.
Yushu Zhang 0001, Junhao Ji, Tao Wang 0084, Wenying Wen, Yong Xiang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Fast Convergent Federated Learning via Decaying SGD Updates
Md Palash Uddin, Yong Xiang 0001, Mahmudul Hasan 0018, Yao Zhao 0006, Youyang Qu, Longxiang Gao
IEEE Trans. Big Data2
2026 Collusion-Resistant and Time-Aware Co-Verification for Edge Data Integrity
abstract
MobileEdgeComputing (MEC) has incentivized App vendors to outsource various services and applications to distributed edge nodes for low access latency. However, the data cached on these nodes is vulnerable to both intentional and accidental corruption, necessitating periodic audits ofEdgeDataIntegrity (EDI). Existing solutions either rely on a “fully trustworthy”ThirdPartyAuditor (TPA) or leverage blockchain to enhance trust. However, they overlook the security risks brought by the use of blockchain, particularly collusion attacks. Furthermore, while they employ achallenge-responsemechanism to enhance efficiency by batch verification, they fail to account for the heterogeneity of edge nodes. To address these challenges, we propose$\mathtt {CTCV}$, aCollusion-resistant andTime-awareCollaborativeVerification framework.$\mathtt {CTCV}$aims to accommodate edge node heterogeneity while enabling public audits and batch verification without introducing additional security risks. Specifically, it incorporates blockchain to allow edge nodes to collaboratively verify EDI without trust dependencies, while mitigating collusion attacks through a carefully designed proof generation and verification approach. Considering the resource and state heterogeneity of edge nodes,$\mathtt {CTCV}$employs atime-constrained challenge-responsemechanism that sets a time threshold$\mathcal {T}$between the verification request issuance and the integrity proof inspection to avoid excessive delays. The selection guideline of$\mathcal {T}$, along with the correctness, efficiency, and collusion resistance of$\mathtt {CTCV}$, are rigorously analyzed. Extensive experiments validate that$\mathtt {CTCV}$is computationally and communicationally efficient compared to three baselines: EdgeWatch, EDI-S, and EDI-V. On average, given 10 edge nodes,$\mathtt {CTCV}$outperforms EdgeWatch, EDI-S, and EDI-V with computation efficiency improvements of 7.9, 9.0, and 5.0 times, and communication efficiency improvement of 2063.0, 4.8, and 2.6 times, respectively.
Yao Zhao 0006, Youyang Qu, Bo Li 0103, Lu Zhao 0001, Feifei Chen 0001, Yong Xiang 0001, Longxiang Gao
IEEE Trans. Dependable Secur. Comput.6
2026 Divergence-Regularized Federated GANs for Effective Cyber-Attack Detection on Non-IID and Unlabeled Edge Activity Data
abstract
Edge computing enables real-time Internet of Things data processing by bringing computation closer to data sources, but its distributed architecture creates cybersecurity vulnerabilities requiring privacy-preserving attack detection mechanisms capable of handling heterogeneous data distributions. This article proposes federated generative adversarial divergence (FedGAD), a plug-and-play modular framework that enhances existing federated learning methods through Jacobian-based regularization and dynamic complexity-aware weighting to address cyber-attack detection in non-independent and identically distributed (IID) and unlabeled edge data environments. Unlike existing approaches suffering from mode collapse and training instability, FedGAD maintains statistical consistency across distributed nodes through gradient-based stability mechanisms, supported by rigorous theoretical analysis establishing convergence guarantees and mode coverage properties. We conduct comprehensive experiments comparing FedGAD against four federated generative learning baselines federated trustworthy (FedTrust), anomaly detection generative adversarial network (ADGAN), federated generative adversarial network for intrusion detection system (FedGAN-IDS), and federated temporal sequential recurrent generative network (FedTSRGNet) and four regularization-based methods federated averaging (FedAvg), federated proximal (FedProx), learning with collaborative aggregation method (LeCam), and Jensen Shannon (JS) Divergence on telemetry data of networks - internet of things (ToN_IoT) and Communications Security Establishment in Canadian Institute for Cybersecurity - Intrusion Detection System (CSE_CIC_IDS) datasets, demonstrating FedGAD's superiority with accuracy improvements up to 3.5%, achieving 100% mode coverage compared to 25% for baseline methods while maintaining computational efficiency for resource-constrained edge deployments.
Zeseya Sharmin, Md Palash Uddin, Yong Xiang 0001, Feifei Chen 0001, Jine Tang, Yushu Zhang 0001
IEEE Trans. Ind. Informatics3
2026 Guest Editorial Introduction to the Special Issue on Federated Learning and Digital Twins for Intelligent Transportation System
Kuo-Hui Yeh, Yong Xiang 0001, Yingjiu Li, Chien-Ming Chen 0001
IEEE Trans. Intell. Transp. Syst.2
2026 Transfer Learning Assisted Detection of Anomalous Events With Insufficient Primary Attribute Data Samples in MEC Networks
abstract
Nowadays IoT devices in Mobile Edge Computing (MEC) networks have been deployed in large-scale quantities to guarantee sensing data collection for anomalous event detection as full as possible even if some devices are in fault. Some techniques, such as clustering and dimensionality reduction, are adopted to eliminate redundant sensing data collection in this large-scale deployment. However, they not only have high computational complexity and easily cause the loss of information on the primary sensing attributes for detection, but also bring certain errors to the detection because of their low sensitivity to data processed. In addition, insufficient collection of primary attribute data samples often results from physical or human factors, and blind imputation of large-scale data gaps without basis may lead to greater irreparable losses. To address the above challenges, we first complete the selection of optimal primary attribute device collection and aggregation (PADCA) path based on minimum spanning tree, reducing data communication cost for redundant primary attributes collection. Then, we propose an anomalous impact correlation search strategy to quickly locate all MEC servers whose management regions have cascading anomalous event and help determine the transferable source MEC servers. Leveraging this, we use transfer learning to help detect anomalous events in the management regions of the MEC servers with insufficient primary attribute data samples, where a particle swarm optimization based back-propagation (PSO-BP) neural network model is used to infer the fusion weight of each primary attribute. Experimental results show that our method achieves higher detection performance in terms of detection time, energy consumption, accuracy, and receiver operating characteristic (ROC) curve compared to the benchmarks by at least 24%, 34%, 0.5 and 0.05.
Jine Tang, Xiaotong Ma, Song Yang 0002, Yong Xiang 0001, Zhangbing Zhou
IEEE Trans. Mob. Comput.4
2026 User on-Demand Driven MEC Servers Deployment From Collaborative Device-Edge-Cloud Network
abstract
With the rapid development of 6 G communication technology and the Internet of Things (IoT), mobile edge computing (MEC) is regarded as an effective paradigm of providing low-delay, high-quality services to mobile users. In the IoT device-edge-cloud network, the optimal deployment of MEC servers is a prerequisite for a better task offloading, while the improved performance of mobile users task offloading also indicates the deployment scheme is optimal. Most of current MEC servers deployment studies focus on reducing delay and deployment costs, but ignore the offloading requirements of mobile users with similar task type and cooperative relationship arriving at the same community. In this paper, we study the MEC servers deployment driven by the task offloading requirements of community mobile users in current period by utilizing the stability of their social cooperative relationships to maximize the service satisfaction of all community mobile users in the future task offloading. First, the cooperative relationship strength between mobile users is measured to form a group of resource requesters based on interaction probability, movement trajectory and credit strength. Then, we implement the optimal search of base stations (BSs) using spatial index, followed by the one-to-many matching theory between BSs and community group resource requesters, to balance the load of BSs and reduce the communication delay between them. Finally, we use TD($\lambda$) algorithm and task similarity between cooperative users to deploy MEC servers with suitable resources around BSs so that the deployment scheme can significantly improve the future task offloading performance of all community mobile users. Based on the real data set provided by Shanghai Telecom, it is confirmed that the proposed scheme has significant advantages in improving all community mobile users service satisfaction, with an average improvement of 18.49% compared with the baselines.
Jine Tang, Jiahao Jin, Song Yang 0002, Yong Xiang 0001, Zhangbing Zhou
IEEE Trans. Serv. Comput.5
2025 FedAT - Federated Adversarial Training Framework for Insider Threat Detection
abstract
Insider threats pose significant security risks in distributed networks, because employees within the organisation may misuse their access to compromise systems. Centralised Machine Learning (ML) techniques are inappropriate in these situations due to privacy and data heterogeneity concerns. To address class imbalance and non-IID data, this study introduces FedAT, a Federated Adversarial Training that integrates federated learning (FL) with generative models to deliver privacy-preserving, multiclass Insider Threat Detection (ITD). FedAT outperforms centralized and conventional FL techniques in terms of scalability, privacy preservation, and detection accuracy, according to evaluations conducted on public CERT datasets.
R. G. Gayathri, Atul Sajjanhar, Md Palash Uddin, Yong Xiang 0001, Ying Zhao 0011
ICPADS4
2025 ContxE: Attention-based Context Aggregation for Temporal Knowledge Graph Completion
abstract
Knowledge graph completion (KGC) methods aim to predict missing links by learning from existing facts in a knowledge graph. Different from KGC, Temporal Knowledge Graph Completion (TKGC) further incorporates the time validity of facts (tagged timestamps) during the learning and inference to improve the completion accuracy. Many TKGC methods achieve this by projecting the static entity representations (time-invariant) of KGC embedding methods to time-dependent representations, which vary across timestamps. However, when measuring a fact, these TKGC methods only consider its subject/object entity representations corresponding to the tagged timestamp, but ignore their historical contexts that normally carry essential supportive information. With this observation, we propose a novel context aggregation (ContxE) method to include historical contexts of subject/object entities for TKGC. To achieve that, we propose a linear-rotary time embedding to obtain time-dependent entity representations that can preserve temporal relationships, and a relation-based attention to aggregate historical context for the score measurement. Comprehensive experiments on three temporal knowledge graph datasets show that the proposed ContxE achieves improved knowledge graph completion results compared to strong counterpart methods.
Borui Cai, Yong Xiang 0001, Longxiang Gao, Jiong Jin, Junfeng Wu 0010, Tom H. Luan
IJCNN2
2025 FedAdap: An Adaptive Federated Knowledge Graph Embedding Framework for Tackling KGs Heterogeneity via Partial Model Sharing
abstract
Knowledge Graph Embedding (KGE) is a technique used to capture structural information from Knowledge Graphs (KGs), enabling various downstream applications such as recommender system. KGE models trained on integrated KGs from multiple organizations tend to outperform those trained on a single KG, owing to the greater richness and diversity of information. Therefore, Federated Knowledge Graph Embedding (FKGE) emerges as a promising approach for privacy-preserving training of KGE models on KGs across organizations (clients). Existing FKGE framework learns a uniform global KGE model that achieves global optima by minimizing aggregated loss across clients. However, heterogeneity among KGs often leads to divergent local optima. This presents a fundamental trade-off: ensuring global optima can compromise local performance, while focusing on local optima can decrease global model utility. To overcome this, we propose Federated Local Adaptive Knowledge Graph Embedding (FedAdap) by drawing inspiration from partial federated learning. FedAdap employs a multilayer convolutional neural network, wherein its lower layers are shared across clients to learn shared information, it maps a seed KGE model into an alignment vector space representation. Its upper layers remain private, transforming the alignment vector space representation to an adaptive KGE model tailored to the local KG. Through this, FedAdap allows clients to leverage shared information while maintaining local adaptability and mitigating the impact of KGs heterogeneity. Experiments on data sets FB15k-237 and NELL-995 show that FedAdap outperforms its counterparts in link prediction tasks.
Borui Cai, Yong Xiang 0001, Keshav Sood
IJCNN3
2025 A Unified Solution to Diverse Heterogeneities in One-Shot Federated Learning
abstract
One-Shot Federated Learning (OSFL) restricts communication between the server and clients to a single round, significantly reducing communication costs and minimizing privacy leakage risks compared to traditional Federated Learning (FL), which requires multiple rounds of communication. However, existing OSFL frameworks remain vulnerable to distributional heterogeneity, as they primarily focus on model heterogeneity while neglecting data heterogeneity. To bridge this gap, we propose FedHydra, a unified, data-free, OSFL framework designed to effectively address both model and data heterogeneity. Unlike existing OSFL approaches, FedHydra introduces a novel two-stage learning mechanism. Specifically, it incorporates model stratification and heterogeneity-aware stratified aggregation to mitigate the challenges posed by both model and data heterogeneity. By this design, the data and model heterogeneity issues are simultaneously monitored from different aspects during learning. Consequently, FedHydra can effectively mitigate both issues by minimizing their inherent conflicts. We compared FedHydra with five SOTA baselines on four benchmark datasets. Experimental results show that our method outperforms the previous OSFL methods in both homogeneous and heterogeneous settings. The code is available at https://github.com/Jun-B0518/FedHydra.
Yiliao Song, Di Wu 0050, Atul Sajjanhar, Yong Xiang 0001, Wei Zhou 0044, Xiaohui Tao 0001, Yan Li 0002, Yue Li 0017
KDD (2)5
2025 Arms Race in Deep Learning: A Survey of Backdoor Defenses and Adaptive Attacks
Xiaoxing Mo, Nan Sun 0002, Leo Yu Zhang, Wei Luo 0001, Shang Gao 0003, Yong Xiang 0001
PAKDD (4)6
2025 Not All Edges are Equally Robust: Evaluating the Robustness of Ranking-Based Federated Learning
abstract
Federated Ranking Learning (FRL) is a state-of-the-art FL framework that stands out for its communication efficiency and resilience to poisoning attacks. It diverges from the traditional FL framework in two ways: 1) it leverages discrete rankings instead of gradient updates, significantly reducing communication costs and limiting the potential space for malicious updates, and 2) it uses majority voting on the server side to establish the global ranking, ensuring that individual updates have minimal influence since each client contributes only a single vote. These features enhance the system's scalability and position FRL as a promising paradigm for FL training. However, our analysis reveals that FRL is not inherently robust, as certain edges are particularly vulnerable to poisoning attacks. Through a theoretical investigation, we prove the existence of these vulnerable edges and establish a lower bound and an upper bound for identifying them in each layer. Based on this finding, we introduce a novel local model poisoning attack against FRL, namely Vulnerable Edge Manipulation (VEM) attack. The VEM attack focuses on identifying and perturbing the most vulnerable edges in each layer and leveraging an optimization-based approach to maximize the attack's impact. Through extensive experiments on benchmark datasets, we demonstrate that our attack achieves an overall 53.23 % attack impact and is 3.7× more impactful than existing methods. Our findings highlight significant vulnerabilities in ranking-based FL systems and underline the urgency for the development of new robust FL frameworks.
Zirui Gong, Yanjun Zhang 0002, Leo Yu Zhang, Zhaoxi Zhang 0001, Yong Xiang 0001, Shirui Pan
SP5
2025 Morality-Driven Mechanism Design: Application in Hierarchical Carbon Trading Markets
Ruhan Liu, Yao Zhang 0005, Youyang Qu, Longxiang Gao, Yong Xiang 0001, Shang Gao 0003, Tom H. Luan
IEEE Internet Things J.5
2025 AirDIV: Over-the-Air Cloud-Fog Data Integrity Verification Scheme for Industrial Cyber-Physical Systems
abstract
Industrial Cyber-Physical Systems (ICPSs) have been motivating various Industry 4.0 endeavours, particularly with the integration of fog computing. Cloud-fog data caching paradigms, as supportive elements of ICPSs, have been adopted to cache user data, catering to diverse ICPS requirements such as data sensitivity and reduced access latency. In this hierarchical caching context, ensuring Cloud-Fog Data Integrity (CFDI) is crucial for maintaining the consistent functionality of ICPSs. Existing solutions primarily focus on examining the integrity of data cached solely on either cloud or fog nodes. However, cloud-cached data and fog-cached data are tightly coupled and should be considered simultaneously when checking data integrity. In this work, we introduce an over-the-air CFDI verification scheme, namely AirDIV, with a high accuracy and security guarantee. Instead of aggregating integrity proofs after proof transmission, AirDIV completes proof aggregation and transmission over the air for efficiency improvement. To enhance practicability, we derive adjustable parameters and formulate an optimization problem to minimize over-the-air aggregation errors. Furthermore, with an effective proof generation method, AirDIV can defend against two common attacks, i.e., replay and forge attacks. We provide a theoretical analysis of AirDIV’s correctness, accuracy and security, while conducting extensive experiments on both simulated and real platforms to validate its efficiency.
Yao Zhao 0006, Yong Xiang 0001, Md Palash Uddin, Yushu Zhang 0001, Lu Liu 0001, Longxiang Gao
IEEE J. Sel. Areas Commun.2
2025 SPJEU: a self-sufficient plaintext-related JPEG image encryption scheme based on a unified key
abstract
In recent research on image encryption, many schemes associate the key generation mechanism with the plaintext to resist chosen plaintext attacks. However, when the sender encrypts many images, a large amount of additional data related to the plaintext need to be transmitted, which leads to problems such as high transmission costs, high requirements for key storage space, and complex key management. Therefore, in this paper, we propose a self-sufficient plaintext-related JPEG image encryption scheme based on a unified key (SPJEU). This scheme establishes the connection between the plaintext and the key by selecting the direct current (DC) coefficients in the JPEG image through a unified key. Homomorphic encryption is applied to the selected DC coefficients, allowing plaintext information to be decrypted directly from the ciphertext domain using a specific calculation method. The remaining DC coefficients are encrypted through group diffusion, and the alternating current (AC) coefficients are grouped and permuted based on the run length. Extensive experiments show that our scheme can resist chosen plaintext attacks, avoid transmitting plaintext-related additional data in the communication channel, and simplify key management. This scheme also ensures the security and format compatibility of the ciphertext image, and the file increment after encryption is very small.
Ming Li 0029, Mengdie Wang, Yushu Zhang 0001, Yong Xiang 0001
Frontiers Inf. Technol. Electron. Eng.5
2025 A federated compositional knowledge graph embedding for communication efficiency
abstract
Knowledge Graph Embedding (KGE), which automatically capture structural information from Knowledge Graphs (KGs), are essential for enhancing various downstream tasks, such as recommender systems. To further improve the effectiveness of KGE models, Federated Knowledge Graph Embedding (FKGE) has been introduced, enabling the privacy-preserving integration of KGs across multiple organizations. However, existing FKGE frameworks require aggregation of a large global KGE model (embeddings). resulting in significant communication overhead, thereby reducing the efficiency and utility of FKGE in practical scenarios. To address this challenge, we propose Federated Compositional Knowledge Graph Embedding (FedComp), which enhances communication efficiency by leveraging the compositional characteristics of KG entities. In FedComp, we design a lightweight global model that represents shareable latent features of entities. These global latent features are composed into personalized KGE models with local embedding generators on the clients, improving both local adaptability and performance. By this, FedComp can significantly reduce the number of parameters that need to be transmitted Experimental results show that FedComp outperforms state-of-the-art FKGE frameworks on link prediction accuracy, with only around 1.0% communication overhead compared to counterpart frameworks.
Borui Cai, Yong Xiang 0001, Yao Zhao 0006, Md Palash Uddin, Keshav Sood
Knowl. Based Syst.3
2025 PR3: Reversible and Usability-Enhanced Visual Privacy Protection via Thumbnail Preservation and Data Hiding
abstract
The image hosting platform is becoming increasingly popular due to its user-friendly features, but it is prone to causing privacy concerns. Only protecting privacy, in fact, can be easy to come true, but usability is frequently sacrificed. Visual privacy protection schemes aim to make a balance between privacy and usability, whereas they are often irreversible. Recently, some reversible visual privacy protection schemes have been proposed by preserving thumbnails (known as TPE). However, they either have excessive states in the Markov chain modeled by the scheme or cannot reverse losslessly. Meanwhile, images encrypted by existing TPE schemes can not embed additional information and thus the usability is limited to visual observation. In view of this, we pertinently propose a reversible and usability-enhanced visual privacy protection scheme (called PR3) based on thumbnail preservation and data hiding. In this scheme, we utilize the sum-preserving data embedding algorithm to substitute the the lowest seven bits of the image without changing the sum. Any data overflow resulting from the above process is stored in the vacated space of the most significant bits. The remaining space serves two purposes: embedding additional information and adjusting the image to approximate the thumbnail. Compared with existing TPE works, PR3 has fewer states in the Markov chain and supports lossless recovery of images. In addition, additional information can be embedded in the encrypted image to enhance usability.
Yushu Zhang 0001, Wenying Wen, Xinpeng Zhang 0001, Xiaochun Cao, Yong Xiang 0001
IEEE Trans. Big Data6
2025 Federated Learning With Adaptive Regularization for Efficient Edge Data Corruption Detection in Edge Intelligence
abstract
Edge intelligence is an emerging distributed computing paradigm that has been driven by the rapid proliferation of Internet of Things (IoT) devices, along with the advancements in edge computing and artificial intelligence. With latency-sensitive data commonly cached across multiple Edge Servers (ESs), efficient Edge Data Integrity Verification (EDIV) has become increasingly critical. Traditional ‘challenge-response’ EDIV methods incur substantial computation and communication costs by indiscriminately verifying all ESs, even though not all ESs may be simultaneously corrupted. A recent Federated Learning (FL)-based framework partially addressed this inefficiency by identifying potentially corrupted ESs early, considering only homogeneous ES activity data. However, due to heterogeneous activity data across diverse ESs, this approach suffers from reduced detection accuracy of potentially corrupted ESs, slower FL convergence, and unclear guidance for subsequent verification rounds, thus limiting the overall reduction in EDIV computation and communication costs. To that end, we proposeFederated learning withAdaptiveRegularizer-basedEdgeDataIntegrityVerification (FedAR-EDIV), which is an effective FL-based framework integrating an adaptive objective regularization strategy specifically designed to handle heterogeneous data distributions. FedAR-EDIV efficiently identifies potentially corrupted ESs during the FL process, achieves faster convergence, and significantly reduces computation and communication costs in the final EDIV procedure. It achieves up to 16× communication speedup and 9.1× computation cost reduction compared to baseline EDIV methods, and reaches FL-based detection accuracy exceeding 99.78% under heterogeneous conditions using KDD99 activity data. Additionally, FedAR-EDIV incorporates a dynamic reputation mechanism after each EDIV round to strategically guide subsequent verification rounds, ensuring fewer checks for trustworthy ESs and greater scrutiny for suspicious ones, thus further minimizing EDIV-related costs. We provide a theoretical analysis that demonstrates the convergence of FedAR-EDIV during FL training, as well as correctness, efficiency, and security during the EDIV process. Extensive experiments conducted on two different heterogeneous activity datasets validated that FedAR-EDIV substantially outperforms baseline methods in terms of corrupted ES detection accuracy, FL convergence speed, and overall EDIV computation and communication costs.
Md Palash Uddin, Yong Xiang 0001, Kuo-Hui Yeh, Lu Liu 0001, Jonathan Kua
IEEE Trans. Cloud Comput.3
2025 Trustworthy and Fair Federated Learning via Reputation-Based Consensus and Adaptive Incentives
abstract
Federated Learning (FL) allows collaborative training of a Machine Learning (ML) model while preserving data privacy across participating clients. Most existing studies consider FL clients to be proactive and completely honest in their participation. However, in reality, clients might lack the motivation to participate, and malicious behavior among some clients could negatively impact the interests of others. For these reasons, ensuring trust and fairness among FL clients is paramount but remains challenging due to limitations in FL consensus mechanisms and incentive strategies. To address these challenges, we introduce a Trustworthy and Fair FL (TFFL) framework that develops a reputation-based consensus mechanism called Dynamic Reputation Consensus (DRC), where clients’ reputations are dynamically assessed based on subjective opinions by evaluating real-time client behavior. We also incorporate time decay and temporal discounting of TFFL interactions along with the weighted measures of clients’ data quality, performance, and reliability to accurately reflect the evolving nature of client behavior over time. By adaptively adjusting clients’ incentives based on reputations and a cooperative game theory, DRC incentivizes honest participation and discourages malicious intent. In addition, we utilize blockchain and smart contracts to provide decentralized, regularized, and secure reputation management that is resistant to tampering and non-repudiation. Theoretical analysis and empirical results on widely used datasets (MNIST, CIFAR-10, and CIFAR-100) demonstrate the effectiveness of DRC in enhancing trust and fairness, improving performance, and providing robust security in FL settings. Results further exhibit that DRC offers superior performance in local model validation, consensus decision, and convergence time compared to related research approaches across various experimental settings.
Yong Xiang 0001, Md Palash Uddin, Jine Tang, Keshav Sood, Longxiang Gao
IEEE Trans. Inf. Forensics Secur.2
2025 Decentralized Data Integrity Auditing in Vehicular Cloud Computing
abstract
As Vehicular Cloud Computing (VCC) evolves, ensuring data integrity and availability becomes a critical challenge due to the vast amount of data being shared and stored. These properties are vital for preserving confidence in cloud services, guaranteeing that data is kept intact and readily available when needed. Traditional data auditing mechanisms, such as Proofs of Retrievability (PoR) and Provable Data Possession (PDP), are effective but often rely on centralized models that pose risks like single-point failures and susceptibility to collusion. To mitigate these risks, we introduce a blockchain-assisted protocol that leverages the decentralized and tamper-proof characteristics of blockchain to enhance the security of data auditing in VCC. Our approach incorporates a dynamic key update mechanism to counter key exposure issues prevalent in VCC and introduces a multi-replica mechanism to ensure data redundancy and reliability across different storage nodes. This feature significantly reduces the risk of data loss and improves trust in cloud services by distributing data storage responsibilities and preventing single-point failures. We conduct formal security analysis and implement a prototype of our protocol on the Ethereum blockchain. Experimental evaluations on both Ganache and Sepolia testnets validate its feasibility in decentralized environments. The results demonstrate that our scheme supports stable challenge-response latency, moderate gas consumption, and reliable multi-replica consistency—making it well-suited for VCC deployments with dynamic conditions and limited resources.
Tianang Yao, Hu Xiong, Kuo-Hui Yeh, Yong Xiang 0001, Changhai Nie
IEEE Trans. Intell. Transp. Syst.4
2025 LDGI: Location-Discriminative Geo-Indistinguishability for Location Privacy
abstract
Geo-Indistinguishability (GI) is a powerful privacy model that can effectively protect location information by limiting the ability of an attacker to infer a user's true location. In real life, locations usually have different sensitive levels in terms of privacy; for example, shopping malls might be low-sensitive while home addresses might be high-sensitive for users. But the GI model does not consider the various sensitive levels of locations, and implements the same perturbation on all locations to meet the highest privacy requirement. This would cause overprotection of low-sensitive locations and reduce data utility. To strike a good balance between privacy and utility, in this paper, we propose a novel privacy notion, termedLocation-DiscriminativeGeo-Indistinguishability (LDGI), which takes into account different sensitive levels of location privacy. With LDGI model, we then develop a perturbation scheme called EM-LDGI based on the exponential mechanism, and an advance scheme MinQL to further enhance data utility. To improve the efficiency of the proposed schemes, we design a scheme MinQL-S with the assistance of the spanner graph, at the cost of a slight utility degradation. We theoretically analyze that the proposed schemes satisfy LDGI and evaluate their performance by extensive experiments on both synthetic and real datasets. The comparison with GI mechanisms demonstrates the advantages of the LDGI model.
Youwen Zhu, Yuanyuan Hong, Qiao Xue, Xiao Lan, Yushu Zhang 0001, Yong Xiang 0001
IEEE Trans. Knowl. Data Eng.6
2025 Multiple Edge Data Integrity Verification With Multi-Vendors and Multi-Servers in Mobile Edge Computing
abstract
Ensuring Edge Data Integrity (EDI) is imperative in providing reliable and low-latency services in mobile edge computing. Existing EDI schemes typically address single-vendor (App Vendor, AV) single-server (Edge Server, ES), single-vendor multi-server, and multi-vendor multi-server scenarios, which consider a single data replica cached by an ES from the AVs. However, the most practical scenario of Multi-Vendors and Multi-Servers with Multiple Data (MVMS-MD) cached by an ES from different AVs remains unexplored. Current solutions struggle when applied to this scenario due to increased computation and communication costs in the verification process across all ESs using the classicalchallenge-response per-data multi-roundstrategy. To tackle this issue, we propose a Multiple EDI-Verification (MEDI-V) approach in this paper. In particular, our MEDI-V utilizes an adaptive Merkle Hash Tree (ad-MHT) to efficiently generate a tree of multiple data replicas within each AV. Next, the dynamic mechanism computes minimal verification information using ad-MHT to create achallengefor individual ESs to produce EDI proofs. The ES then leverages its ad-MHT and the ES's proof to send the reconstructed ad-MHT root to the AV for verification. Theoretical insights into MEDI-V's correctness, efficiency, security, and comprehensive evaluations demonstrate its superiority in addressing MEDI issues in the MVMS-MD scenario.
Yong Xiang 0001, Md Palash Uddin, Yao Zhao 0006, Jonathan Kua, Longxiang Gao
IEEE Trans. Mob. Comput.2
2025 Adaptive Search and Collaborative Offloading Under Device-to-Device Joint Edge Computing Network
abstract
Mobile Edge Computing (MEC) and Device-toDevice (D2D) peer offloading are two promising paradigms in the mobile Internet of Things (IoT). In this paper, we study the collaborative task offloading with redundant data and codes in large-scale IoT networks, where computing resource-starved IoT devices can offload their tasks to MEC servers via cellular links or to nearby peer devices (PDs) with idle resources through D2D links for execution. IoT tasks usually consist of a series of dependent and parallel subtasks, and the difficulties in current research are (i) how to eliminate redundancy in data or codes between subtasks, and (ii) how to leverage previous experience to adaptively search a set of collaborative MEC servers and PDs for matching offloading of dependent and parallel subtasks. From this, we propose a redundancy-aware adaptive search offloading (RASO) method based on the deep Q-network (DQN). Specifically, we first design a fine-grained task recombination scheme by judging the consistency of subtask data and codes. After that, we organize the global devices into a spatial index MP-tree to reduce the search solution space, and propose a fast adaptive search method based on the DQN combined with MP-tree, where optimal path-guiding parameters training of inner and outer layers is involved to efficiently help achieve collaborative devices to complete specific tasks with the same type. After finding the collaborative MEC servers and PDs along MP-tree for a certain task, a centralized stable matching algorithm is further developed to give a decision of offloading each of its divided dependent and parallel subtasks to the matched one, thereby optimizing offloading delay and energy consumption. Extensive simulation results show that compared to other counterpart solutions, our proposed method has improved task offloading performance in terms of delay and energy consumption.
Jine Tang, Jiahao Jin, Yong Xiang 0001, Xiaofei Wang 0001, Zhangbing Zhou
IEEE Trans. Mob. Comput.4
2025 Data Re-Outsourcing Detection With Latency-Constraint for Edge Storage
abstract
Edge storage has become a widely used solution for providing low-latency data access services, which motivates data owners to outsource data on geographically distributed edge nodes to deliver a positive user experience. Nevertheless, various security concerns raise in terms of data availability. Among them, edge data geo-location verification becomes a prominent concern when the data is out of owners' control, since outsourced data may be re-outsourced to other economical yet unknown third-party devices by dishonest edge nodes for saving storage space and pocketing the difference. Existing geo-localization approaches for cloud architectures can not be practically applied to identify such re-outsourcing behaviors due to the uniqueness of edge storage. To close this gap, we make the first attempt to investigate theedgedatare-outsourcingdetection (EDRD) problem, enabling the data owner to inspect if outsourced data is consistently cached on the rented edge nodes with agreed geo-location. We leverage timedChallenge-Responsemechanisms for data possession proof while measuring verification latency to detect re-outsourcing behaviors by comparing with re-outsourcing detection threshold$\mathbb {C}$. We prove that the edge node whose verification latency exceeds$\mathbb {C}$is dishonest. To obtain the optimal$\mathbb {C}$, we formulate thethresholddetermination (TD) problem and transform it to an easy-to-handle form for problem complexity reduction. Then, apreference-based approach named TD-P is developed to efficiently address the transformed TD problem. On top of that, we propose a$\mathbb {C}$-aware edge data re-outsourcing detection scheme entitled EDRD-$\mathbb {C}$to tackle the EDRD problem effectively. The efficiency and effectiveness of TD-P and EDRD-$\mathbb {C}$are verified by extensive theoretical analysis and experimental evaluations on both simulated and real platforms. Notably, EDRD-$\mathbb {C}$achieves 100% detection accuracy by sacrificing a reasonable amount of computing resources and 92.97% in the worst case.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao
IEEE Trans. Serv. Comput.3
2025 MM-SCS: Leveraging Multimodal Features to Enhance Smart Contract Code Search
abstract
Semantic code search technology allows searching for existing code snippets through natural language, which can greatly improve programming efficiency. Smart contracts, programs that run on the blockchain, have a code reuse rate of more than 79%, which means developers have a great demand for semantic code search tools. However, the existing code search models still have a semantic gap between code and query and perform poorly on specialized queries of smart contracts. In this paper, we propose a Multi-Modal Smart contract Code Search (MM-SCS) model. Specifically, we construct a Contract Elements Dependency Graph (CEDG) for MM-SCS as an additional modality to capture the data flow and control flow information of the code. To make the model more focused on the key contextual information, we use a multi-head attention network to generate embeddings for code features. In addition, we use a fine-tuned pretrained model to ensure the model's effectiveness when the training data is small. We compared MM-SCS with four state-of-the-art models on a dataset with 470K (code, docstring) pairs collected from Github and Etherscan. Experimental results show that MM-SCS achieves an MRR (Mean Reciprocal Rank) of 0.572, outperforming four state-of-the-art models UNIF, DeepCS, CARLCS-CNN, and TAB-CS by 34.2%, 59.3%, 36.8%, and 14.1%, respectively. Additionally, the search speed of MM-SCS is second only to UNIF, reaching 0.34s/query.
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Longxiang Gao
IEEE Trans. Software Eng.2
2024 FedInverse: Evaluating Privacy Leakage in Federated Learning
abstract
Federated Learning (FL) is a distributed machine learning technique where multiple devices (such as smartphones or IoT devices) train a shared global model by using their local data. FL claims that the data privacy of local participants is preserved well because local data will not be shared with either the server-side or other training participants. However, this paper discovers a pioneering finding that a model inversion (MI) attacker, who acts as a benign participant, can invert the shared global model and obtain the data belonging to other participants. This will lead to severe data-leakage risk in FL because it is difficult to identify attackers from benign participants. In addition, we found even the most advanced defense approaches could not effectively address this issue. Therefore, it is important to evaluate such data-leakage risks of an FL system before using it. To alleviate this issue, we propose FedInverse to evaluate whether the FL global model can be inverted by MI attackers. In particular, FedInverse can be optimized by leveraging the Hilbert-Schmidt independence criterion (HSIC) as a regularizer to adjust the diversity of the MI attack generator. We test FedInverse with three typical MI attackers, GMI, KED-MI, and VMI, and the experiments show our FedInverse method can successfully obtain the data belonging to other participants. The code of this work is available at https://github.com/Jun-B0518/FedInverse
Di Wu 0050, Yiliao Song, Wei Zhou 0044, Yong Xiang 0001, Atul Sajjanhar
ICLR6
2024 SE-shapelets: Semi-supervised Clustering of Time Series Using Representative Shapelets
abstract
Shapelets that discriminate time series using local features (subsequences) are promising for time series clustering. Existing time series clustering methods may fail to capture representative shapelets because they discover shapelets from a large pool of uninformative subsequences, and thus result in low clustering accuracy. This paper proposes a Semi-supervised Clustering of Time Series Using Representative Shapelets (SE-Shapelets) method, which utilizes a small number of labeled and propagated pseudo-labeled time series to help discover representative shapelets, thereby improving the clustering accuracy. In SE-Shapelets, we propose two techniques to discover representative shapelets for the effective clustering of time series. (1) A salient subsequence chain (SSC) that can extract salient subsequences (as candidate shapelets) of a labeled/pseudo-labeled time series, which helps remove massive uninformative subsequences from the pool. (2) A linear discriminant selection (LDS) algorithm to identify shapelets that can capture representative local features of time series in different classes, for convenient clustering. Experiments on UCR time series datasets demonstrate that SE-shapelets discovers representative shapelets and achieves higher clustering accuracy than counterpart semi-supervised time series clustering methods.
Borui Cai, Guangyan Huang, Shuiqiao Yang, Yong Xiang 0001, Chihung Chi
Expert Syst. Appl.4
2024 Hybrid deep learning model using SPCAGAN augmentation for insider threat analysis
R. G. Gayathri, Atul Sajjanhar, Yong Xiang 0001
Expert Syst. Appl.3
2024 FTPE-BC: Fast thumbnail-preserving image encryption using block-churning
Ming Li 0029, Qingchen Cui, Yushu Zhang 0001, Yong Xiang 0001
Expert Syst. Appl.5
2024 An Optimized Privacy-Protected Blockchain System for Supply Chain on Internet of Things
abstract
The consortium blockchain is being utilized in supply chains on the Internet of Things (IoT) for tracking and protecting supply chain data, such as manufacture, storage, and shipment. However, the supply chain data in a consortium blockchain is publicly accessible for all parties, which attracts widespread concerns about supply chain data privacy. Several existing attribute-based encryption (ABE)-based blockchain systems targeting to address the supply chain data privacy problem either bring about additional security problems or lack the feasibility analysis on IoTs. To address the aforementioned issues, in this article, a novel multiauthority ABE (MA-ABE)-based blockchain system is proposed to protect the data privacy for the supply chain on IoTs. Specifically, a four-way tradeoff optimization framework is designed so that the system decentralization, scalability, and storage consumption are not significantly affected by the improved privacy. The optimal attribute setting policies for different scale blockchain networks are dynamically generated by the nondominated sorting genetic algorithm II (NSGA-II). Extensive experiment results show that the proposed scheme remarkably improves data privacy protection for the supply chain without downgrading the other three key factors.
Chenhao Xu 0003, Youyang Qu, Yong Xiang 0001, Tom H. Luan, Longxiang Gao
IEEE Internet Things J.3
2024 A Learning-Based Hierarchical Edge Data Corruption Detection Framework in Edge Intelligence
abstract
Edge intelligence, an emerging distributed paradigm, is driven by the increasing number of Internet of Things devices and the development of edge computing and artificial intelligence. This paradigm revolutionizes the way of data caching by encouraging latency-sensitive data to be distributed across multiple edge nodes. In such data caching scenarios, ensuring the integrity of data stored at the edge nodes is critical for business continuity guarantee. Existing Edge Data Integrity (EDI) verification solutions rely on the interactive Challenge-Response mechanism. However, this mechanism imposes significant communication overhead on participants, leading to low verification efficiency. To address this challenge, we propose a Learning-based Hierarchical Edge Data Corruption Detection framework (LH-EDCD), aiming to enhance verification efficiency from a round perspective by reducing communication interaction between edge nodes and the data owner. LH-EDCD involves two layers of verification: internal and external. In the internal verification layer, each edge node self-inspects the cached data replica by running a corruption detection model distributedly trained by blockchain-based Federated Learning (FL). With such filtration, potential corruption can be efficiently identified without complex interaction. Considering the false positive existence in the model, in the external verification layer, LH-EDCD adopts a smart contract in blockchain to verify identified potentially corrupted data replicas for corruption confirmation, mitigating the trust concerns among edge nodes while reducing communication overhead on backbone networks. With the combination of these two layers, the overall EDI verification efficiency can be improved by reducing interaction verification time. Additionally, we make the first attempt to investigate the optimal verification time to improve the applicability and practicality of LH-EDCD. Extensive experimental results substantiate the advantages of employing FL in the first layer of LH-EDCD and demonstrate that LH-EDCD outperforms two state-of-the-art EDI approaches, i.e., EDI-S and EDI-V. Specifically, LH-EDCD achieves better model accuracy and convergence speed compared to centralized training, while exhibiting superior efficiency over EDI-S and EDI-V with 3.5 and 2.8 times performance improvements, respectively.
Yao Zhao 0006, Chenhao Xu 0003, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao
IEEE Internet Things J.4
2024 Representing Noisy Image Without Denoising
abstract
A long-standing topic in artificial intelligence is the effective recognition of patterns from noisy images. In this regard, the recent data-driven paradigm considers 1) improving the representation robustness by adding noisy samples in training phase (i.e., data augmentation) or 2) pre-processing the noisy image by learning to solve the inverse problem (i.e., image denoising). However, such methods generally exhibit inefficient process and unstable result, limiting their practical applications. In this paper, we explore a non-learning paradigm that aims to derive robust representation directly from noisy images, without the denoising as pre-processing. Here, the noise-robust representation is designed as Fractional-order Moments in Radon space (FMR), with also beneficial properties of orthogonality and rotation invariance. Unlike earlier integer-order methods, our work is a more generic design taking such classical methods as special cases, and the introduced fractional-order parameter offers time-frequency analysis capability that is not available in classical methods. Formally, both implicit and explicit paths for constructing the FMR are discussed in detail. Extensive simulation experiments and robust visual applications are provided to demonstrate the uniqueness and usefulness of our FMR, especially for noise robustness, rotation invariance, and time-frequency discriminability.
Yushu Zhang 0001, Chao Wang 0028, Tao Xiang 0001, Xiaochun Cao, Yong Xiang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Enhancing Robustness of Speech Watermarking Using a Transformer-Based Framework Exploiting Acoustic Features
abstract
Digital watermarking serves as an effective approach for safeguarding speech signal copyrights, achieved by the incorporation of ownership information into the original signal and its subsequent extraction from the watermarked signal. While traditional watermarking methods can embed and extract watermarks successfully when the watermarked signals are not exposed to severe alterations, these methods cannot withstand attacks such as de-synchronization. In this work, we introduce a novel transformer-based framework designed to enhance the imperceptibility and robustness of speech watermarking. This framework incorporates encoders and decoders built on multi-scale transformer blocks to effectively capture local and long-range features from inputs, such as acoustic features extracted by Short-Time Fourier Transformation (STFT). Further, a deep neural networks (DNNs) based generator, notably the Transformer architecture, is employed to adaptively embed imperceptible watermarks. These perturbations serve as a step for simulating noise, thereby bolstering the watermark robustness during the training phase. Experimental results show the superiority of our proposed framework in terms of watermark imperceptibility and robustness against various watermark attacks. When compared to the currently available related techniques, the framework exhibits an eightfold increase in embedding rate. Further, it also presents superior practicality with scalability and reduced inference time of DNN models.
Chuxuan Tong, Iynkaran Natgunanathan, Yong Xiang 0001, Jianhua Li 0002, Tianrui Zong, James Xi Zheng, Longxiang Gao
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 Context-Aware Consensus Algorithm for Blockchain-Empowered Federated Learning
abstract
Supported by cloud computing,FederatedLearning (FL) has experienced rapid advancement, as a promising technique to motivate clients to collaboratively train models without sharing local data. To improve the security and fairness of FL implementation, numerousBlockchain-empoweredFederatedLearning (BFL) frameworks have emerged accordingly. Among them, consensus algorithms play a pivotal role in determining the scalability, security, and consistency of BFL systems. Existing consensus solutions to block producer selection and reward allocation either focus on well-resourced scenarios or accommodate BFL based on clients' contributions to model training. However, these approaches limit consensus efficiency and undermine reward fairness, due to involving intricate consensus processes, disregarding clients' contributions during blockchain consensus, and failing to address lazy client problems (malicious clients plagiarizing local model updates from others to reap rewards). Given the aforementioned challenges, we make the first attempt to design a joint solution for efficient consensus and fair reward allocation in heterogeneous BFL systems with lazy clients. Specifically, we introduce a generalizable BFL workflow that can address lazy client problems well. Based on it, the global contribution of BFL clients is decoupled into five dominant metrics, and the block producer selection problem is formulated as a reward-constraint contribution maximization problem. By addressing this problem, the optimal block producer that maximizes global contribution can be identified to orchestrate consensus processes, and rewards are distributed to clients in proportion to their respective global contributions. To achieve it, we develop aContext-awareProof-of-Contribution consensus algorithm named CPoC to reach consensus and incentive simultaneously, followed by theoretical analysis of lazy client problems and privacy issues. Empirical results on widely-used datasets demonstrate the effectiveness of our design in improving consensus efficiency and maximizing global contribution.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao
IEEE Trans. Cloud Comput.3
2024 Design and Robust Evaluation of Next Generation Node Authentication Approach
abstract
The flexibility of 5G-NGNs makes them an ideal infrastructure for supporting mission-critical IoT applications that require low latency and high bandwidth. However, due to the rapid proliferation and the integration of IoTs with 5 G, the threat surface has considerably expanded. Hence the security of IoT devices is a big concern. Unfortunately, IoT devices have limited resources, and the traditional security approaches (authentication and intrusion detection approaches) of cryptography do not work effectively on 5G-IoT ecosystems. Motivated from this, we leverage the distinctive RF (Radio Frequency) fingerprinting signatures of IoT devices and used them to train a Deep learning model, Mahalanobis Distance theory in addition to the Chi-square distribution theory, to authenticate the IoT nodes. Under robust scenarios we have tested the approach shows detection accuracy (99.35%) as well as significant amount of reduction in model's training time as these two metrics are one of the primary key performance indicators (KPIs). In order to evaluate the effectiveness of the proposed method in real-time scenarios, we tested the proposed solution with a real RF dataset and the OSM-MANO 5 G platform. The model underwent formal verification using the Tamarin Prover tool, and the proposal was also compared with recent research works.
Dinh Duc Nha Nguyen, Keshav Sood, Yong Xiang 0001, Longxiang Gao, Lianhua Chi, Shui Yu 0001
IEEE Trans. Dependable Secur. Comput.3
2024 AgrAmplifier: Defending Federated Learning Against Poisoning Attacks Through Local Update Amplification
abstract
The collaborative nature of federated learning (FL) poses a major threat in the form of manipulation of local training data and local updates, known as the Byzantine poisoning attack. To address this issue, many Byzantine-robust aggregation rules (AGRs) have been proposed to filter out or moderate suspicious local updates uploaded by Byzantine participants. This paper introduces a novel approach called AGRAMPLIFIER, aiming to simultaneously improve robustness, fidelity, and efficiency of the existingAGRs. The core idea of AGRAMPLIFIER is to amplify the “morality” of local updates by identifying the most repressive features of each gradient update, which provides a clearer distinction between malicious and benign updates, consequently improving the detection effect. To achieve this objective, two approaches, namelyAGRMPandAGRXAI, are proposed.AGRMPorganizes local updates into patches and extracts the largest value from each patch, whileAGRXAIleverages explainable AI methods to extract the gradient of the most activated features. By equipping AGRAMPLIFIER with the existing Byzantine-robust mechanisms, we successfully enhance the model robustness, maintaining its fidelity and improving overall efficiency. AGRAMPLIFIER is universally compatible with the existing Byzantine-robust mechanisms. The paper demonstrates its effectiveness by integrating it with all mainstreamAGRmechanisms. Extensive evaluations conducted on seven datasets from diverse domains against seven representative poisoning attacks consistently show enhancements in robustness, fidelity, and efficiency, with average gains of 40.08%, 39.18%, and 10.68%, respectively.
Zirui Gong, Liyue Shen, Yanjun Zhang 0002, Leo Yu Zhang, Jingwei Wang 0003, Guangdong Bai, Yong Xiang 0001
IEEE Trans. Inf. Forensics Secur.7
2024 From Wide to Deep: Dimension Lifting Network for Parameter-Efficient Knowledge Graph Embedding
abstract
Knowledge graph embedding (KGE) that maps entities and relations into vector representations is essential for downstream applications. Conventional KGE methods require high-dimensional representations to learn the complex structure of knowledge graph, but lead to oversized model parameters. Recent advances reduce parameters by low-dimensional entity representations, while developing techniques (e.g., knowledge distillation or reinvented representation forms) to compensate for reduced dimension. However, such operations introduce complicated computations and model designs that may not benefit large knowledge graphs. To seek a simple strategy to improve the parameter efficiency of conventional KGE models, we take inspiration from that deeper neural networks require exponentially fewer parameters to achieve expressiveness comparable to wider networks for compositional structures. We view all entity representations as a single-layer embedding network, and conventional KGE methods that adopt high-dimensional entity representations equal widening the embedding network to gain expressiveness. To achieve parameter efficiency, we instead propose a deeper embedding network for entity representations, i.e., a narrow entity embedding layer plus a multi-layer dimension lifting network (LiftNet). Experiments on three public datasets show that by integrating LiftNet, four conventional KGE methods with 16-dimensional representations achieve comparable link prediction accuracy as original models that adopt 512-dimensional representations, saving 68.4% to 96.9% parameters.
Borui Cai, Yong Xiang 0001, Longxiang Gao, Di Wu 0050, He Zhang 0034, Jiong Jin, Tom H. Luan
IEEE Trans. Knowl. Data Eng.2
2024 Federated Learning-Assisted Task Offloading Based on Feature Matching and Caching in Collaborative Device-Edge-Cloud Networks
abstract
Mobile edge computing provides relatively rich computation resources for Internet-of-Things (IoT) task offloading at the edge of networks. As time goes on, user tasks present diverse requirements in function, type, dependency, urgency, etc., which makes edge servers take on dynamically diversified service features to adapt to the requirements of user tasks. Moreover, cache has been studied a lot in recent years for reducing the execution cost of related or dependent tasks. However, jointly considering which result data required to be cached and where to cache is still an intractable problem in task offloading due to dynamically diversified and sensitive features of task and edge servers for prediction. To provide more comprehensive consideration, we propose a multiple features matching scheme, coupled with federated learning-assisted collaborative caching, to enhance the efficiency of task offloading. Specifically, we first build a common features of historical tasks based FI-tree to help search for an edge server that best matches the requested task features. This helps to obtain optimal task allocation and improve offloading performance. Further, the results of tasks related to or dependent on cached results can be obtained directly through the collaborative edge cache prediction model trained by two-stage federated learning. In this way, the amount of data executed for offloaded tasks is reduced, thereby speeding up the return of final results as well as reducing the delay and energy of task execution. Meanwhile, it avoids massive transmission of task results correlated data and also protects the privacy of these data when training the prediction model. Experimental results show that our proposed method outperforms the benchmark approaches through reducing the time delay and energy consumption by at least 15.6% and 18.2%.
Jine Tang, Sen Wang 0011, Song Yang 0002, Yong Xiang 0001, Zhangbing Zhou
IEEE Trans. Mob. Comput.4
2024 ARFL: Adaptive and Robust Federated Learning
abstract
Federated Learning (FL) is a machine learning technique that enables multiple local clients holding individual datasets to collaboratively train a model, without exchanging the clients' datasets. Conventional FL approaches often assign a fixed workload (local epoch) and step size (learning rate) to the clients during the client-side local model training and utilize all collaborating trained models' parameters evenly during the server-side global model aggregation. Consequently, they frequently experience problems with data heterogeneity and high communication costs. In this paper, we propose a novel FL approach to mitigate the above problems. On the client side, we propose an adaptive model update approach that optimally allocates a needful number of local epochs and dynamically adjusts the learning rate to train the local model and regularizes the conventional objective function by adding a proximal term to it. On the server side, we propose a robust model aggregation strategy that potentially supplants the local outlier updates (models' weights) prior to the aggregation. We provide the theoretical convergence results and perform extensive experiments on different data setups over the MNIST, CIFAR-10, and Shakespeare datasets, which manifest that our FL scheme surpasses the baselines in terms of communication speedup, test-set performance, and global convergence.
Md Palash Uddin, Yong Xiang 0001, Borui Cai, Xuequan Lu, John Yearwood, Longxiang Gao
IEEE Trans. Mob. Comput.2
2024 SCEI: A Smart-Contract Driven Edge Intelligence Framework for IoT Systems
abstract
Federated learning (FL) enables collaborative training of a shared model on edge devices while maintaining data privacy. FL is effective when dealing with independent and identically distributed (iid) datasets, but struggles with non-iid datasets. Various personalized approaches have been proposed, but such approaches fail to handle underlying shifts in data distribution, such as data distribution skew commonly observed in real-world scenarios (e.g., driver behavior in smart transportation systems changing across time and location). Additionally, trust concerns among unacquainted devices and security concerns with the centralized aggregator pose additional challenges. To address these challenges, this paper presents a dynamically optimized personal deep learning scheme based on blockchain and federated learning. Specifically, the innovative smart contract implemented in the blockchain allows distributed edge devices to reach a consensus on the optimal weights of personalized models. Experimental evaluations using multiple models and real-world datasets demonstrate that the proposed scheme achieves higher accuracy and faster convergence compared to traditional federated and personalized learning approaches.
Chenhao Xu 0003, Jiaqi Ge, Longxiang Gao, Mengshi Zhang, Yong Xiang 0001, James Xi Zheng
IEEE Trans. Mob. Comput.7
2024 Data Integrity Verification in Mobile Edge Computing With Multi-Vendor and Multi-Server
abstract
The emergingMobileEdgeComputing (MEC) paradigm reforms the way of data caching by motivating App vendors to store latency-sensitive data on distributed edge servers. In volatile MEC environments, ensuringEdgeDataIntegrity (EDI) is a major concern for App vendors. Existing EDI solutions only consider the scenario with a single App vendor and multiple edge servers, neglecting more complex multi-vendor and multi-server cases. If multiple App vendors check their data replicas cached on the same edge server simultaneously, integrity verification efficiency will drop exponentially. To mitigate this challenge, we make the first attempt to develop aSmartInspectionAlgorithm (SIA) to pre-select unreliable data replicas for different App vendors in each verification round by jointly considering cache services' QoS (Quality-of-Service) and data replicas' unverified time. By implementing this approach, edge servers can merely verify the selected data replicas, greatly reducing computation and communication overheads in EDI verification. Theoretically, SIA can achieve$\mathcal {O}(n)$expected time complexity. Supported by SIA, we expand the EDI problem in multi-vendor and multi-server MEC environments (referred to as the MVMS-EDI problem) and propose a smart contract-based approach entitled MVMS-SC to tackle the problem efficiently and impartially. We provide a rigorous theoretical analysis of the correctness, security, and efficiency of MVMS-SC. Both large-scale and small-scale experiments with real-world datasets are correspondingly performed on a single machine and a real platform to validate the superiority of MVMS-SC in terms of computation and communication efficiencies.
Yao Zhao 0006, Youyang Qu, Feifei Chen 0001, Yong Xiang 0001, Longxiang Gao
IEEE Trans. Mob. Comput.4
2024 Long-Term Over One-Off: Heterogeneity-Oriented Dynamic Verification Assignment for Edge Data Integrity
abstract
EdgeIntelligence (EI), a burgeoning research area, motivates App vendors to cache data replicas on geographically distributed edge servers to deliver better services. On the downside, this benefit also incurs more data integrity audit overhead on App vendors, which calls for more efficientEdgeDataIntegrity (EDI) verification approaches. However, existing EDI solutions totally rely on an implicitresource homogeneity assumption-edge servers have identical resource availability throughout EDI inspection execution in each round-but it rarely holds in reality. The edge servers with insufficient computation and/or communication capacity greatly limit overall EDI verification efficiency from a round perspective. Thus, in this work, we release the identified impractical assumption and accordingly study the EDIDynamicVerificationAssignment (DVA) problem for the first time. The problem aims to maximize the number of data replicas being verified in the long term under the constraints of verification delay in resource-limited environments. In this way, App vendors merely need to check the integrity of selected data replicas in each round for efficiency improvement. Specifically, we first formalize the DVA problem as a delay-constrained long-term stochastic optimization problem and further prove its$\mathcal {NP}$-hardness. To resolve the problem efficiently, we decompose it to an easy-to-handle form and then develop a polynomial-timePriority-based approach named DVA-P with a theoretical analysis of its time complexity and performance bound. Finally, experimental evaluations validate that DVA-P can be seamlessly incorporated into existing EDI solutions to enhance overall verification efficiency while guaranteeing verification performance.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Chaochen Shi, Feifei Chen 0001, Longxiang Gao
IEEE Trans. Mob. Comput.3
2024 Prototype-Guided Memory Replay for Continual Learning
abstract
Continual learning (CL) is a machine learning paradigm that accumulates knowledge while learning sequentially. The main challenge in CL is catastrophic forgetting of previously seen tasks, which occurs due to shifts in the probability distribution. To retain knowledge, existing CL models often save some past examples and revisit them while learning new tasks. As a result, the size of saved samples dramatically increases as more samples are seen. To address this issue, we introduce an efficient CL method by storing only a few samples to achieve good performance. Specifically, we propose a dynamic prototype-guided memory replay (PMR) module, where synthetic prototypes serve as knowledge representations and guide the sample selection for memory replay. This module is integrated into an online meta-learning (OML) model for efficient knowledge transfer. We conduct extensive experiments on the CL benchmark text classification datasets and examine the effect of training set order on the performance of CL models. The experimental results demonstrate the superiority our approach in terms of accuracy and efficiency.
Stella Ho, Ming Liu 0028, Lan Du 0002, Longxiang Gao, Yong Xiang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Unambiguous and High-Fidelity Backdoor Watermarking for Deep Neural Networks
abstract
The unprecedented success of deep learning could not be achieved without the synergy of big data, computing power, and human knowledge, among which none is free. This calls for the copyright protection of deep neural networks (DNNs), which has been tackled via DNN watermarking. Due to the special structure of DNNs, backdoor watermarks have been one of the popular solutions. In this article, we first present a big picture of DNN watermarking scenarios with rigorous definitions unifying the black- and white-box concepts across watermark embedding, attack, and verification phases. Then, from the perspective of data diversity, especially adversarial and open set examples overlooked in the existing works, we rigorously reveal the vulnerability of backdoor watermarks against black-box ambiguity attacks. To solve this problem, we propose an unambiguous backdoor watermarking scheme via the design of deterministically dependent trigger samples and labels, showing that the cost of ambiguity attacks will increase from the existing linear complexity to exponential complexity. Furthermore, noting that the existing definition of backdoor fidelity is solely concerned with classification accuracy, we propose to more rigorously evaluate fidelity via examining training data feature distributions and decision boundaries before and after backdoor embedding. Incorporating the proposed prototype guided regularizer (PGR) and fine-tune all layers (FTAL) strategy, we show that backdoor fidelity can be substantially improved. Experimental results using two versions of the basic ResNet18, advanced wide residual network (WRN28_10) and EfficientNet-B0, on MNIST, CIFAR-10, CIFAR-100, and FOOD-101 classification tasks, respectively, illustrate the advantages of the proposed method.
Guang Hua 0001, Andrew Beng Jin Teoh, Yong Xiang 0001, Hao Jiang 0010
IEEE Trans. Neural Networks Learn. Syst.3
2024 PS-Net: A Learning Strategy for Accurately Exposing the Professional Photoshop Inpainting
abstract
Restoring missing areas without leaving visible traces has become a trivial task with Photoshop inpainting tools. However, such tools have potentially illegal or unethical uses, such as removing specific objects in images to deceive the public. Despite the emergence of many forensics methods of image inpainting, their detection ability is still insufficient when attending to professional Photoshop inpainting. Motivated by this, we propose a novel method termed primary-secondary network (PS-Net) to localize the Photoshop inpainted regions in images. To the best of our knowledge, this is the first forensic method devoted specifically to Photoshop inpainting. The PS-Net is designed to deal with the problems of delicate and professional inpainted images. It consists of two subnetworks: the primary network (P-Net) and the secondary network (S-Net). The P-Net aims at mining the frequency clues of subtle inpainting features through the convolutional network and further identifying the tampered region. The S-Net enables the model to mitigate compression and noise attacks to some extent by increasing the co-occurring feature weights and providing features that are not captured by the P-Net. Furthermore, the dense connection, Ghost modules, and channel attention blocks (C-A blocks) are adopted to further strengthen the localization ability of PS-Net. Extensive experimental results illustrate that PS-Net can successfully distinguish forged regions in elaborate inpainted images, outperforming several state-of-the-art solutions. The proposed PS-Net is also robust against some postprocessing operations commonly used in Photoshop.
Yushu Zhang 0001, Zhibin Fu, Mingfu Xue, Xiaochun Cao, Yong Xiang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Configurable Harris Hawks Optimisation for Application Placement in Space-Air-Ground Integrated Networks
abstract
Space-Air-Ground Integrated Network (SAGIN) has recently emerged as a viable solution for reliable transmission, high data rates, and seamless connectivity with extensive coverage. However, the characteristics of the computation and communication devices located at various levels of SAGIN make application placement within such environments a challenging task. Real-time service expectations and resource requirements of applications further intensify this issue, and push the domain to operate beyond its capacity, resulting in uneven delays and significant overhead. Taking these constraints into account, SAGIN’s application placement problem can be expressed as a multiobjective optimisation problem. This paper aims to solve such a problem using a Dynamic Weight-configurable Harris Hawks Optimisation (DW-HHO) algorithm, considering diverse application contexts such as deadlines, resource usage and the number of application activities. It simultaneously minimises application total service time and host resource overhead with a robust global search. The performance of the proposed solution is compared with benchmark metaheuristic solutions such as PSO, NSGA-II, Greedy and Random. Experimental results demonstrate that DW-HHO outperforms other benchmark metaheuristic solutions in optimising resource utilisation and service delivery time of applications in SAGIN environments. The proposed DW-HHO demonstrates notable improvements over existing methods. Specifically, when evaluating the total service time for PSO, NSGA-II, Greedy, and Random, DW-HHO outperforms these methods by 7.28%, 9.07%, 13.01%, and 14.97%, respectively.
Nasrin Akhter 0002, Md. Redowan Mahmud, Jiong Jin, Jason But, Iftekhar Ahmad, Yong Xiang 0001
IEEE Trans. Netw. Serv. Manag.6
2024 Evaluating Federated Learning-Based Intrusion Detection Scheme for Next Generation Networks
abstract
The proliferation of billions of heterogeneous Internet of Things (IoT) devices at a rapid pace has resulted in a marked expansion of attack surfaces. Numerous new attacks are constantly emerging to undermine the network’s availability, data confidentiality, and systems’ integrity due to inadequate security measures and resource limitations. Intrusion detection systems (IDSs) are used as the first line of defense to identify early instances of cyber-attacks targeting critical points. However, Next-Generation Networks (NGNs) with dense connectivity pose a challenge for traditional IDS approaches, as they raise concerns about users’ data privacy. Federated learning-based IDSs (Fed-IDSs) are an emerging and promising solution, as they permit the training of machine learning models on decentralized data stored on devices without compromising privacy. However, Fed-IDSs also have some unique issues. We identified that the existing Fed-IDSs have poor performance since the datasets used for evaluation, or the data in the real world, are highly imbalanced, and classes are not uniformly distributed. Motivated by this, we developed a novel IDS to effectively address the problem of class imbalance in federated learning at both the local and global levels. Following this, we evaluated the performance of our Fed-IDS under both independent and identically distributed (IID) and non-IID data settings and observed its generalizability to detect various attacks improved greatly. Extensive experiments are conducted to illustrate the effectiveness and benefits of this proposal.
Keshav Sood, Pachamuthu Rajalakshmi, Dinh Duc Nha Nguyen, Yong Xiang 0001
IEEE Trans. Netw. Serv. Manag.5
2024 E-TPE: Efficient Thumbnail-Preserving Encryption for Privacy Protection in Visual Sensor Networks
abstract
When visual sensor networks (VSNs) enter daily life in society, they not only bring great convenience but also cause people to worry about privacy. Traditional image encryption uses the snowflake effect to protect the privacy, but this can compromise the usability of directly browsing images captured by nodes in VSNs. Recently, thumbnail-preserving encryption (TPE) has been proposed to balance image usability and privacy. However, the Markov chain’s limited connectivity and high time cost, which result in low efficiency, may preclude its application in VSNs. Motivated by this, we propose a novel TPE scheme to protect the image collected in VSNs’ privacy efficiently. We first deeply study the bijective relationship between the three-pixel group and its rank, and then propose a new number method for both, based on which we design the efficient rank mapping. Subsequently, an efficient thumbnail-preserving image encryption via three-pixel rank mapping (E-TPE) is designed, achieving NR security and enhancing the Markov chain’s connectivity. The experiments demonstrate the scheme’s efficiency; that is, it is currently the quickest ideal-TPE scheme with NR security, and a good balance is achieved between usability and privacy.
Yushu Zhang 0001, Wenying Wen, Rushi Lan, Yong Xiang 0001
ACM Trans. Sens. Networks5
2024 Distributed Task Processing Platform for Infrastructure-Less IoT Networks: A Multi-Dimensional Optimization Approach
abstract
With the rapid development of artificial intelligence (AI) and the Internet of Things (IoT), intelligent information services have showcased unprecedented capabilities in acquiring and analysing information. The conventional task processing platforms rely on centralised Cloud processing, which encounters challenges in infrastructure-less environments with unstable or disrupted electrical grids and cellular networks. These challenges hinder the deployment of intelligent information services in such environments. To address these challenges, we propose a distributed task processing platform (${DTPP}$) designed to provide satisfactory performance for executing computationally intensive applications in infrastructure-less environments. This platform leverages numerous distributed homogeneous nodes to process the arriving task locally or collaboratively. Based on this platform, a distributed task allocation algorithm is developed to achieve high task processing performance with limited energy and bandwidth resources. To validate our approach,${DTPP}$has been tested in an experimental environment utilising real-world experimental data to simulate IoT network services in infrastructure-less environments. Extensive experiments demonstrate that our proposed solution surpasses comparative algorithms in key performance metrics, including task processing ratio, task processing accuracy, algorithm processing time, and energy consumption.
Qiushi Zheng, Jiong Jin, Zhishu Shen, Iftekhar Ahmad, Yong Xiang 0001
IEEE Trans. Parallel Distributed Syst.6
2024 FEUAGame: Fairness-Aware Edge User Allocation for App Vendors
abstract
Mobile edge computing (MEC) offers a new computing paradigm that turns computing and storage resources to the network edge to provide minimal service latency compared to cloud computing. Many research works have attempted to help app vendors allocate users to appropriate edge servers for high-performance service provisioning. However, existing edge user allocation (EUA) approaches have ignored fairness in users' data rates caused by interference, which is crucial in service provisioning in the MEC environment. To pursue fairness in EUA, edge users need to be assigned to edge servers so their quality of experience can be ensured at minimum costs without significant service performance differences among them. In this paper, we make the first attempt to address this fair edge user allocation (FEUA) problem. Specifically, we formulate the FEUA problem, prove its N P-hardness, and propose an optimal approach to solve small-scale FEUA problems. To accommodate large-scale FEUA scenarios, we propose a game-theoretic approach called FEUAGame that transforms the FEUA problem into a potential game that admits a Nash equilibrium. FEUA employs a decentralized algorithm to find the Nash equilibrium in the potential game as the solution to the FEUA problem. A widely-used real-world data set is utilised to experimentally compare the performance of FEUAGame to four representative approaches. The numerical outcomes show the effectiveness and efficiency of the proposed approaches in solving the FEUA problem
Feifei Chen 0001, Guangming Cui, Yong Xiang 0001, Qiang He 0001
IEEE Trans. Parallel Distributed Syst.4
2024 Blockchained Dual-Asynchronous Federated Learning Services for Digital Twin Empowered Edge-Cloud Continuum
abstract
The booming of learning-based Artificial Intelligence (AI) enables the integration of Big Data and emerging computing architectures, which facilitate the Edge-AI-as-a-Service (EAaaS) in the edge-cloud continuum. To meet the emerging demands, such as privacy preservation and autonomy, blockchain-enabled federated learning (B-FL) is proposed, which further provides decentralized processing, data falsification avoidance, and learning model reliability. However, synchronous global aggregation, which is deployed in most existing B-FL paradigms, is dragging down the performances due to the data and computing resources heterogeneity of diverse edge devices. In addition, the restricted resources of edge devices pose further challenges in executing learning tasks and blockchain-based consensus simultaneously. To solve these issues, we propose a blockchained dual-asynchronous federated learning (BAFL-DT) service model for EAaaS in the digital twin empowered edge-cloud continuum. In BAFL-DT, federated learning services are run on local edge devices, while the global aggregation is achieved by the consensus process of digital twins implemented in the cloud. Besides, dual-asynchronous FL allows both local training and global aggregation to be performed in an asynchronous manner, which is uniquely enabled by the proposed paradigm. Extensive evaluations of real-world datasets testify to the superior performances of EAaaS by improving accuracy and efficiency.
Youyang Qu, Shui Yu 0001, Longxiang Gao, Keshav Sood, Yong Xiang 0001
IEEE Trans. Serv. Comput.5
2024 Adaptive Regularization and Resilient Estimation in Federated Learning
abstract
Federated Learning (FL) is an emerging research area that produces a globally trained model using numerous local users' data and maintains their privacy. Heterogeneous or non-Independent and Identically Distributed ( non-IID) data affect the global model's convergence and, therefore, cause high communication costs. These are because traditional FL approaches often disregard an adaptive regularized objective for the user-side training and utilize conventional arithmetic mean on the locally trained models for the server-side aggregation. To alleviate these issues, we propose a novel FL scheme in this paper. In particular, we propose an adaptive regularization approach to add to the classical objective function of the users' local models during training and a resilient estimation approach to the locally trained models during aggregation. The adaptive regularization approach is derived using the users' local and global performance diversification while the resilient estimation scheme uses a modified geometric mean aggregation over the local models' parameters. We provide consolidated theoretical results and perform extensive experiments on the IID and non-IID settings of MNIST, CIFAR-10, and Shakespeare datasets with various deep networks. The results manifest that our FL scheme outperforms the state-of-the-art approaches in terms of communication speedup, test-set performance, training convergence stability, and resiliency against attacks.
Md Palash Uddin, Yong Xiang 0001, Yao Zhao 0006, Mumtaz Ali 0003, Yushu Zhang 0001, Longxiang Gao
IEEE Trans. Serv. Comput.2
2024 Long-Term Proof-of-Contribution: An Incentivized Consensus Algorithm for Blockchain-Enabled Federated Learning
abstract
The surge in data collected by local devices has given rise to a distributed machine learning architecture namedFederatedLearning (FL) for privacy-preserving model training. However, the security of centralized aggregation of local models becomes a primary concern, which can be mitigated byBlockchain-enabledFederatedLearning (BFL) to facilitate decentralized model aggregation. In BFL, consensus and incentive are two of the key components that impact the scalability, security, and consistency of the system. Existing joint solutions focus on selecting a block producer based on client contributions to model training but overlook contributions to blockchain consensus and lack consideration for correlations across communication rounds, inevitably affecting incentive performance. Motivated by these, we make the first attempt to achieve blockchain consensus with along-term incentive guaranteefor BFL systems. Following a generalizable BFL workflow, we decouple the global contribution of BFL clients into four rigorously modeled metrics, and formulate the block producer selection problem as a long-term total contribution maximization problem with reward constraints. ALong-termProof-of-Contribution algorithm named LPoC is developed to handle this problem efficiently. In each communication round, LPoC identifies an optimal block producer that can maximize total contributions from a long-term perspective while allocating rewards to continuously motivate clients to contribute to BFL. We provide a detailed analysis of time complexity and performance bounds, followed by extensive experimental evaluations. The results demonstrate the effectiveness of LPoC in maximizing long-term total contribution, improving consensus efficiency, and upgrading training performance.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Longxiang Gao
IEEE Trans. Serv. Comput.3
2024 Winning at the Starting Line: Unreliable Data Replica Selection for Edge Data Integrity Verification
abstract
MobileEdgeComputing (MEC) is an emerging technology, where App vendors are allowed to cache multiple data replicas on geographically distributed edge servers to serve adjacent mobile subscribers. However, this benefit introduces an extra workload for edge servers and App vendors, as they must audit the integrity of multiple data replicas periodically considering various threats caused by distributed and dynamic MEC environments. The large-scale growth of data replicas certainly is a challenge to design more efficientEdgeDataIntegrity (EDI) verification approaches. Existing solutions are mostly limited to improving efficiency by optimizing proof generation and verification methods, while the improvement is still far from satisfactory due to adopting indiscriminate inspection philosophy (checking all data replicas without discrimination). In this paper, we make the first attempt to abstract a pre-processing phase and correspondingly study theUnreliable dataReplicaSelection (URS) problem. It can be seamlessly integrated into existing EDI solutions by solving the URS problem at the start of each verification round. Such pre-selection can significantly enhance overall EDI verification efficiency by incorporating the cache serviceQualityofService (QoS) and verification success rate, especially in scenarios with a large number of data replicas. Specifically, we first formalize the URS problem as a constrained optimization problem, and further prove its$\mathcal {NP}$-hardness. To address the problem efficiently, we transform it into an easy-to-handle form and develop aPriority-based approach named URS-P. Both theoretical analysis and experimental evaluation validate the effectiveness and efficiency of our proposed solution.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Feifei Chen 0001, Md Palash Uddin, Longxiang Gao
IEEE Trans. Serv. Comput.3
2024 Cost-Effective Company Response Policy for Product Co-Creation in Company-Sponsored Online Community
abstract
Product co-creation based on company-sponsored online community has come to be a paradigm of developing new products collaboratively with customers. In such a product co-creation campaign, the sponsoring company needs to interact intensively with active community members about the design scheme of the product. We call the collection of the rates of the company’s response to active community members at all time in the co-creation campaign as a company response policy (CRP). This article addresses the problem of finding a cost-effective CRP (the CRP problem). First, we introduce a novel community state evolutionary model and, thereby, establish an optimal control model for the CRP problem (the CRP model). Second, based on the optimality system for the CRP model, we present an iterative algorithm for solving the CRP model (the CRP algorithm). Third, through extensive numerical experiments, we conclude that the CRP algorithm converges and the resulting CRP exhibits excellent cost benefit. Consequently, we recommend the resulting CRP to companies that embrace product co-creation. Next, we discuss how to implement the resulting CRP. Finally, we investigate the effect of some factors on the cost benefit of the resulting CRP. To our knowledge, this work is the first attempt to study value co-creation through optimal control theoretic approach.
Lu-Xing Yang, Xiaofan Yang 0001, Kaifan Huang, Gang Li 0009, Yong Xiang 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2023 Temporal Knowledge Graph Completion: A Survey
abstract
Knowledge graph completion (KGC) predicts missing links and is crucial for real-life knowledge graphs, which widely suffer from incompleteness. KGC methods assume a knowledge graph is static, but that may lead to inaccurate prediction results because many facts in the knowledge graphs change over time. Emerging methods have recently shown improved prediction results by further incorporating the temporal validity of facts; namely, temporal knowledge graph completion (TKGC). With this temporal information, TKGC methods explicitly learn the dynamic evolution of the knowledge graph that KGC methods fail to capture. In this paper, for the first time, we comprehensively summarize the recent advances in TKGC research. First, we detail the background of TKGC, including the preliminary knowledge, benchmark datasets, and evaluation metrics. Then, we summarize existing TKGC methods based on how the temporal validity of facts is used to capture the temporal dynamics. Finally, we conclude the paper and present future research directions of TKGC.
Borui Cai, Yong Xiang 0001, Longxiang Gao, He Zhang 0034, Jianxin Li 0001
IJCAI2
2023 Social tie-driven coupling propagation of user awareness and information in Device-to-Device communications
abstract
Regarding information dissemination in Device-to-Device (D2D) communications, most existing works only consider user awareness as the influencing factor, which cannot accurately reflect the interaction between users and devices. In fact, user awareness and information interaction are inseparable from the social tie that describes the users’ intimacy. In this paper, we propose a social tie-driven coupling propagation dynamical model, taking into account both user awareness diffusion and information transmission. This model includes interprocess interaction factors that fully describe the user-user and user-device interaction behaviors. Through intensive experiments on the real-world Peer-to-Peer (P2P) and Facebook datasets with a tailored propagation algorithm, we show that the proposed model has a higher propagation speed and covers a larger propagation range than that of an existing single propagation process model. Our work also reveals that a strong social tie between users will promote the transmission of device information, which further accelerates the diffusion of user awareness.
Chenquan Gan, Qingyi Zhu, Ye Zhu 0002, Yong Xiang 0001
Comput. Networks5
2023 DIVRS: Data integrity verification based on ring signature in cloud storage
Yushu Zhang 0001, Youwen Zhu, Liangmin Wang 0001, Yong Xiang 0001
Comput. Secur.5
2023 Design data decomposition-based reference evapotranspiration forecasting model: A soft feature filter based deep learning driven approach
Zihao Johnson Zheng, Mumtaz Ali 0003, Mehdi Jamei, Yong Xiang 0001, Masoud Karbasi, Zaher Mundher Yaseen, Aitazaz Ahsan Farooque
Eng. Appl. Artif. Intell.4
2023 Blockchain search engine: Its current research status and future prospect in Internet of Things network
Jine Tang, Xinming Lu, Yong Xiang 0001, Chaochen Shi, Junhua Gu
Future Gener. Comput. Syst.3
2023 Speech emotion recognition via multiple fusion under spatial-temporal parallel network
abstract
Speech, as a necessary way to express emotions, plays a vital role in human communication. With the continuous deepening of research on emotion recognition in human–computer interaction, speech emotion recognition (SER) has become an essential task to improve the human–computer interaction experience. When performing emotion feature extraction of speech, the method of cutting the speech spectrum will destroy the continuity of speech. Besides, the method of using the cascaded structure without cutting the speech spectrum cannot simultaneously extract speech spectrum information from both temporal and spatial domains. To this end, we propose a spatial–temporal parallel network for speech emotion recognition without cutting the speech spectrum. To further mix the temporal and spatial features, we design a novel fusion method (called multiple fusion) that combines the concatenate fusion and ensemble strategy. Finally, the experimental results on five datasets demonstrate that the proposed method outperforms state-of-the-art methods.
Chenquan Gan, Qingyi Zhu, Yong Xiang 0001, Deepak Kumar Jain 0001, Salvador García 0001
Neurocomputing4
2023 Machine translation-based fine-grained comments generation for solidity smart contracts
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Keshav Sood, Longxiang Gao
Inf. Softw. Technol.2
2023 Toward IoT Node Authentication Mechanism in Next Generation Networks
abstract
Although the next generation networks (5G-NGNs) provide a flexible infrastructure to support latency-sensitive and bandwidth-hungry mission-critical Internet of Things (IoT) applications, however, the 5G-IoT integration in NGNs has increased the threat surface. Unfortunately, IoT devices are resource constrained, and the traditional intrusion detection systems (IDS) approaches based on cryptography are not effective on 5G-IoT ecosystems. In this article, we propose an effective 5G-IoT node authentication approach that leverages unique radio frequency (RF) fingerprinting data to train the Deep learning model to detect legitimate and nonlegitimate IoT nodes. Our approach is based on Mahalanobis Distance theory and Chi-square distribution theories. The proposed approach achieves a higher detection accuracy (99.35%) as well as lower training time compared to other existing approaches which is a key benefit of our approach in NGNs. The experiments are conducted using ETSI-open source NFV management and orchestration (OSM-MANO) platform on Amazon Web Services (AWSs) cloud platform to verify how the proposed approach would fit in real-life scenarios. The method can be used as a standalone security system or as a part of multifactor authentication.
Dinh Duc Nha Nguyen, Keshav Sood, Yong Xiang 0001, Longxiang Gao, Lianhua Chi, Shui Yu 0001
IEEE Internet Things J.3
2023 Heterogeneous and Customized Cost-Efficient Reversible Image Degradation for Green IoT
abstract
With the large-scale deployment of the Internet of Things (IoT) in daily life, more and more privacy data are collected by IoT devices. These data are not directly physically controlled by users, which may cause privacy concerns. In fact, privacy has become one of the significant problems faced by IoT. In this article, we mainly study the protection of image privacy under the green IoT. We have conducted an in-depth analysis of the green IoT scenario and put forward the scope and corresponding goals that the scheme should have. Motivated by this, a novel image privacy protection scheme is proposed, i.e., heterogeneous and customized cost-efficient reversible image degradation for green IoT. This scheme fully considers the characteristics of privacy and the various users’ diverse requirements to achieve a heterogeneous and customized privacy protection. Meanwhile, cost effectiveness cannot be confined to the efficiency of the direct image processing at the expense of greatly increasing costs in other aspects, such as transmission and reversion. It is mitigated by preserving some visual content in the privacy-protected image. It also improves the image compression efficiency and ensures that the user can select the desired image according to the visual content for reversion. Some experiments have been carried out to demonstrate that this work has achieved the proposed scope and corresponding goals.
Yushu Zhang 0001, Rushi Lan, Zhongyun Hua, Yong Xiang 0001
IEEE Internet Things J.5
2023 Hybrid KD-NFT: A multi-layered NFT assisted robust Knowledge Distillation framework for Internet of Things
Nai Wang, Di Wu 0050, Wencheng Yang, Yong Xiang 0001, Atul Sajjanhar
J. Inf. Secur. Appl.5
2023 Hybrid variational autoencoder for time series forecasting
abstract
Variational autoencoders (VAE) are powerful generative models that learn the latent representations of input data as random variables. Recent studies show that VAE can flexibly learn the complex temporal dynamics of time series and achieve more promising forecasting results than deterministic models. However, a major limitation of existing works is that they fail to jointly learn the local patterns (e.g., seasonality and trend) and temporal dynamics of time series for forecasting. Accordingly, we propose a novel hybrid variational autoencoder (HyVAE) to integrate the learning of local patterns and temporal dynamics by variational inference for time series forecasting. Experimental results on four real-world datasets show that the proposed HyVAE achieves better forecasting results than various counterpart methods, as well as two HyVAE variants that only learn the local patterns or temporal dynamics of time series, respectively.
Borui Cai, Shuiqiao Yang, Longxiang Gao, Yong Xiang 0001
Knowl. Based Syst.4
2023 Statistically Dependent Blind Signal Separation Under Relaxed Sparsity
abstract
Most traditional blind source separation (BSS) methods are based on the basic assumption that source signals are mutually independent. However, many signals in the real world are often interrelated to each other. In this paper, a novel BSS method is proposed for dependent signals which only relies on the relaxed sparsity of sources. Under this constraint, the column vector of high-dimensional observed signals would be distributed in a specific subspace. According to the representation relation in time-frequency (TF) domain, each column vector of observed signal is uniquely clustered into the subspace, to which it belongs. Furthermore, the estimation of source signals is calculated by the inverse matrix or the pseudo-inverse matrix of the mixing matrix. It is also demonstrated mathematically and experimentally that under relaxed sparsity conditions, proposed subspace clustering methods is promising and hopeful to solve the dependent BSS (DBSS) problems.
Dezhong Peng, Yong Xiang 0001, Qingchuan Tao, Zhong Yuan
IEEE Signal Process. Lett.3
2023 Query-Efficient Black-Box Adversarial Attacks on Automatic Speech Recognition
abstract
The susceptibility of Deep Neural Networks (DNNs) to adversarial attacks has raised concerns regarding their practical applications in real-world scenarios. Although the vulnerability of DNNs to adversarial attacks has been extensively studied in the image domain, research in the audio domain, particularly in the black-box setting with Automatic Speech Recognition (ASR) models, remains limited. While various black-box attacks have been proposed for ASR models, such as transfer attacks, hardware attacks, and query-based attacks, this study concentrates on query-based black-box attacks. The article introduces a new gradient estimation technique, Temporal Natural Evolution Strategies (T-NES), to generate adversarial audio samples more efficiently than existing attacks. T-NES leverages the temporal correlation present in audio to speed up gradient estimation based on the probability scores returned by the target model. The empirical results on benchmark datasets, LibriSpeech and TEDLIUM, and two state-of-the-art ASR models, DeepSpeech2 and Wav2Letter, demonstrate that T-NES can generate successful attacks with up to 30% fewer queries than existing attacks within 500 queries. T-NES could provide a robust baseline for evaluating the black-box adversarial vulnerability of ASR systems.
Chuxuan Tong, James Xi Zheng, Jianhua Li 0002, Xingjun Ma, Longxiang Gao, Yong Xiang 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 SSVS-SSVD Based Desynchronization Attacks Resilient Watermarking Method for Stereo Signals
abstract
Most of the audio signals in real-world applications are stereo signals. However, the previous desynchronization attacks resilient watermarking methods cannot preserve perceptual quality or achieve robustness when constrained by high embedding rates and stereo host. In this paper, based on two novel features segmental singular values summation (SSVS) and segmental singular values difference (SSVD) that are generated using discrete cosine transform (DCT) and singular value decomposition (SVD), we present a robust watermarking method for stereo signals that not only is robust to desynchronization attacks and common signal processing attacks but also has a larger embedding rate compared with the previous methods. In the proposed method, we first apply DCT and SVD on each segment of the host signal to extract the SSVS feature and the SSVD feature. Then we generate the adaptive embedding parameters and embed watermark bits via optimized embedding strategies based on these features. Due to the use of the adaptive embedding parameters and the optimized embedding strategies, the proposed method significantly increases the embedding rate without compromising the robustness and perceptual quality. Analysis results show our proposed method outperforms the state-of-the-art methods by a large margin, where the perceptual quality improvement is over 14%, and the robustness against desynchronization attacks is improved by more than 49% when the embedding rate is 70 bps.
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Guang Hua 0001, Keshav Sood, Yushu Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Localization of Inpainting Forgery With Feature Enhancement Network
abstract
Inpainting the given region of an image is a typical requirement in computer vision. Conventional inpainting, through exemplar-based or diffusion-based strategies, can create realistic inpainted images at a very low cost. Also, such easy-to-use manipulation poses new security threats. Therefore, the detection of inpainting has attracted considerable attention from researchers. However, the existing methods are typically not suitable for the general detection of various inpainting algorithms. Motivated by this, in this work, an efficient feature enhancement network is proposed to locate the inpainted regions in the digital image. First, we design an artifact enhancement block to effectively capture the traces left by diffusion or exemplar-based inpainting. Then, the VGGNet is used as a feature extractor to describe advanced and low-resolution features. Finally, to take full advantage of enhanced features, we concatenate the features obtained by the feature extractor and the up-sampling operations. Extensive experimental evaluations, covering benchmarking, ablation, robustness, generalization, and efficiency studies, confirm the usefulness of the proposed method. This is especially true on the conventional inpainting dataset, our method obtains an average F1 score 7.63% higher than the second-best method. Theoretical and numerical analyses support the effectiveness of our feature enhancement network in representing the artifacts in inpainted images, exhibiting better potential for real-world forensics than various state-of-the-art strategies.
Yushu Zhang 0001, Zhibin Fu, Mingfu Xue, Zhongyun Hua, Yong Xiang 0001
IEEE Trans. Big Data6
2023 An Efficient Oblivious Random Data Access Scheme in Cloud Computing
abstract
With the development of cloud computing and cloud storage techniques, much attention has been focused on the privacy protection of outsourced data. Existing searchable encryption solutions can ensure the confidentiality and availability of data stored on the cloud. However, searchable encryption is vulnerable to statistical inference attacks, which exploit the disclosure of access patterns on encrypted indexes and encrypted file sets, which has become a potential way to reveal user privacy. Oblivious random access memory (ORAM) is an important means of concealing access patterns, yet its direct use in searchable encryptions is expensive. This paper presents a scheme for efficient and oblivious access to encrypted databases through encrypted indexes. This scheme is a hybrid ORAM scheme, which utilizes semi-homomorphic encryption to perform calculations in the ciphertext domain, overcoming the limitations of the huge overhead associated with Path-ORAM. For excessive amounts of data, semi-homomorphic encryption can significantly reduce communication and storage overhead. Our scheme can achieve high-security encrypted search and update operations at the same time. Moreover, the execution speed of ODS-Tree is 2-8x faster than that of ORAM-based schemes. In addition, the proposed scheme reduces the data block transmission and storage costs compared to existing frameworks.
Hong Liu 0025, Xiaojing Lu, Shengchen Duan, Yushu Zhang 0001, Yong Xiang 0001
IEEE Trans. Cloud Comput.5
2023 BeDCV: Blockchain-Enabled Decentralized Consistency Verification for Cross-Chain Calculation
abstract
With the increase of data stored on the blockchain, the efficiency of storage and calculation of blockchain has gradually become a bottleneck restricting the development of blockchain. By storing data on multiple chains, blockchains can request data from other chains for calculation and the storage pressure can be alleviated. But the transfer of a large amount of data between chains suffers from low transfer efficiency and poor security. A reasonable design is to perform the calculation on the data storage chain and only transfer the results across chains. However, since the calculation process is invisible, blockchains cannot judge the consistency of calculation results from other chains. In this paper, we provide a blockchain-enabled decentralized consistency verification scheme for cross-chain calculation (BeDCV). Considering the decentralized characteristic of blockchain, we adopt the blockchain calledsupervision chainfor decentralized auditing. We modify paillier homomorphic encryption to encrypt data involved in the calculation for correctness verification. Then, we aggregate the ciphertexts of data to generate the audit proof for integrity verification. Besides, we verify whether the data involved in the calculation are real-time by leveraging a counting bloom filter. The supervision chain can check the correctness, integrity, and real-time performance of cross-chain data calculation without revealing any original information about the data. The theoretical and experimental analysis demonstrates that BeDCV can verify the consistency of cross-chain data calculation result effectively, realizing secure and reliable expansion of blockchain.
Yushu Zhang 0001, Xuewen Dong, Liangmin Wang 0001, Yong Xiang 0001
IEEE Trans. Cloud Comput.5
2023 Frequency Spectrum Modification Process-Based Anti-Collusion Mechanism for Audio Signals
abstract
The collusion attack combines multiple multimedia files into one new file to erase the user identity information. The traditional anti-collusion methods (which aim to trace the traitors) can defend the collusion attack, but they cannot well defend some hybrid collusion attacks (e.g., a collusion attack combined with desynchronization attacks). To address this issue, we propose a frequency spectrum modification process (FSMP) to defend the collusion attack by significantly downgrading the perceptual quality of the colluded file. The severe perceptual quality degradation can demotivate the attackers from launching the collusion attack. Because FSMP is orthogonal to the existing traitor-trace-based methods, it can be combined with the existing methods to provide a double-layer protection against different attacks. In FSMP, after several signal processing procedures (e.g., uneven framing and smoothing), multiple signals (called FSMP signals) can be generated from the host signal. Launching collusion attack using the generated FSMP signals would lead to the energy disturbance and attenuation effect (EDAE) over the colluded signals. Due to the EDAE, FSMP can significantly degrade the perceptual quality of the colluded audio file, thereby thwarting the collusion attack. In addition, FSMP can well defend different hybrid collusion attacks. Theoretical analysis and experimental results confirm the validity of the proposed method.
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Guang Hua 0001, Longxiang Gao, Gleb Beliakov
IEEE Trans. Cybern.3
2023 Optimal Transport-Based Patch Matching for Image Style Transfer
abstract
State-of-the-art image style transfer methods have achieved impressive results by using neural networks. However, neural style transfer (NST) methods either ignore the local details of the style image by using the global statistics for style modeling or cannot fully use shallow features of neural networks, leading to the synthesized image having fewer details. In this study, we proposed a new patch-based style transfer method that directly operates in the image pixel domain without using any neural networks, achieving fascinating style transfer results with rich image details. The proposed method was derived from classic texture synthesis methods. Most previous methods rely on nearest neighbor search (NNS) for patch matching. However, this greedy strategy cannot guarantee the similarity of patch distributions between the synthesized image and the style image, which limits the expressiveness of textures. We solved this problem by proposing an optimal patch matching algorithm formed on the Optimal Transport (OT) theory, which theoretically guarantees the similarity of the patch distributions and gives a flexible style modeling method. Various qualitative and quantitative experiments demonstrated that the proposed method achieves better synthesized results than state-of-the-art style transfer methods, including NST and classic methods based on texture synthesis.
Jie Li 0023, Yong Xiang 0001, Hao Wu 0010, Shaowen Yao 0001, Dan Xu 0001
IEEE Trans. Multim.2
2023 Multi-Label Speech Emotion Recognition via Inter-Class Difference Loss Under Response Residual Network
abstract
Speech emotion recognition has always been a challenging task due to the difference in emotion expression and perception. Currently, in the supervised speech emotion recognition systems, the soft label overcomes the disadvantage of the hard label losing annotations variability and emotion perception subjectivity, but it only considers the emotion perceptions of a few annotators and thus still brings high statistical error. For this issue, this paper redefines the target and designs a novel loss function (denoted as inter-class difference loss), which enables the network to adaptively learn an emotion distribution in all utterances. This not only restricts the negative class probability less than the positive class probability, but also limits the negative class probability close to zero. To make the speech emotion recognition system more efficient, this paper proposes an end-to-end network, called response residual network (R-ResNet), which incorporates the ResNet for features extraction, together with the emotion response module for data augmentation and variable-length data processing. Finally, the experimental results not only demonstrate the advanced performance of our work, but also confirm that the ambiguous utterances contain emotional characteristics. In addition, another interesting finding is that, on the unbalanced dataset, the batch normalization (BN) after addition performs better than BN before addition.
Xiaoke Li, Zufan Zhang, Chenquan Gan, Yong Xiang 0001
IEEE Trans. Multim.4
2023 Joint Optimization of Coverage and Reliability for Application Placement in Mobile Edge Computing
abstract
Mobile edge computing (MEC) provides a new distributed computing paradigm that overcomes the inability of cloud computing to offer low end-to-end latency. In a MEC environment, app vendors can deliver lower-latency services to mobile app users by placing applications on edge servers in close proximity to app users. From an app vendor's perspective, an optimal edge application placement strategy under a budget ($k$) constraint aims to place application instances on$k$edge servers within a specific area to maximize user coverage. However, edge servers may be subject to failure due to multiple reasons, e.g., hardware faults, software exceptions, cyber-attacks, etc. App users served by failed edge servers need to access applications from remote cloud servers if they cannot access any other edge servers. This impacts app users’ quality of experience significantly. Thus, app vendors need to consider the reliability of the edge server network when choosing edge servers for placing their application instances. We make the first attempt in this paper to tackle this problem of joint optimization of coverage and reliability for edge application placement (EAP-CR). We formulate this problem as a constrained optimization problem and prove its$\mathcal {NP}$-hardness theoretically. Then, we propose an optimal approach to find the optimal solutions with the integer programming technique, an approximation approach is also proposed to find approximate solutions for large-scale EAP-CR problems. We evaluate EAP-$OPT$and EAP-$APX$against three relevant approaches through experiments conducted on a widely-used real-world data set and a synthetic data set. The results demonstrate that our proposed approaches can solve the EAP-CR problem effectively and efficiently.
Feifei Chen 0001, Xiaoyu Xia 0001, Yong Xiang 0001, Xuehong Tao, Qiang He 0001
IEEE Trans. Serv. Comput.4
2023 Optimization Search Strategy for Task Offloading From Collaborative Edge Computing
abstract
Edge computing is a popular paradigm in solving the problems of long time delay and high energy consumption in Internet of Things (IoT) network, which can effectively realize the IoT task offloading by collaboration of multiple edge servers. Nevertheless, how to choose the appropriate edge servers for offloading the dependent subtasks is still a big challenge, considering the limited resources and computing power of the edge servers as well as the start and end execution time of each subtask. These factors have a great impact on the execution efficiency of the whole task. At present, most of the research works focus on single-hop or multi-hop task offloading, where the edge servers farther away are not considered in the offloading decision. Such task offloading strategy is not optimal, and difficult to achieve high parallel execution of tasks, resulting in some delay-sensitive tasks not being completed within the specified time. In this paper, a two-stage optimization method is proposed to solve the resource allocation problem between edge servers and tasks. In the first stage, we group tasks according to their priorities, and the group with a higher priority is given the preference to resource allocation, thereby ensuring the timeliness of delay-sensitive tasks. Within the same group, resources are competed according to the game theory, and the total delay of all tasks is optimized. In the second stage, we aim to optimize the energy consumption of each task without increasing its completion time by allocating the computing resources to its subtasks based on their maximum completion time. For group resource allocation, we propose a spatial index tree to store the information of all edge servers for optimal server selection. During the selection process, an online learning based double prediction model is utilized to reduce the energy consumption caused by information transmission. We have evaluated the performance of the experiment on iFogSim simulator, and the experimental results show that our proposed method can achieve better performance in terms of time delay and energy consumption.
Jine Tang, Taishan Qin, Yong Xiang 0001, Zhangbing Zhou, Junhua Gu
IEEE Trans. Serv. Comput.3
2023 Federated Learning via Disentangled Information Bottleneck
abstract
Existing Federated Learning (FL) algorithms generally suffer from high communication costs and data heterogeneity due to the use of conventional loss function for local model update and the equal consideration of each local model for global model aggregation. In this paper, we propose a novel FL approach to address the above issues. For local model update, we propose a disentangled Information Bottleneck (IB) principle-based loss function. For global model aggregation, we suggest a model selection strategy based on Mutual Information (MI). Particularly, we design a Lagrangian-based loss function using the IB principle and “disentanglement” for maximizing MI between the ground truth and model prediction and minimizing MI between the intermediate representations. We calculate MI ratio between the ground truth and model prediction, and between the original input and ground truth to select the effective models for aggregation. We analyze the theoretical optimal cost of the loss function and manifest optimal convergence rate, and quantify the outlier robustness of the aggregation scheme. Experiments demonstrate the superiority of the proposed FL approach, in terms of testing performance and communication speedup (i.e., 3.00-14.88 times for IID MNIST, 2.5-50.75 times for non-IID MNIST, 1.87-18.40 times for IID CIFAR-10, and 1.24-2.10 times for non-IID MIMIC-III).
Md Palash Uddin, Yong Xiang 0001, Xuequan Lu, John Yearwood, Longxiang Gao
IEEE Trans. Serv. Comput.2
2023 Fair Outsourcing Paid in Fiat Money Using Blockchain
abstract
Seeking outsourcing from cloud service providers is common for resource-constrained users to complete complex computing. Conventional cloud computing outsourcing solutions focus on guiding users to verify the returned results, which is unfair since malicious users can refuse to pay by falsely claiming that the results are wrong. To address this problem, fair payment schemes, both blockchain-free and blockchain-based, have been proposed. However, the former is usually inefficient and the latter only supports payment through cryptocurrencies. The use of cryptocurrencies faces hurdles as it exposes cloud service providers to a significant risk of sharp currency price fluctuations, as well as vulnerability to government bans due to regulatory concerns. In contrast, fiat money payment is not only more in line with the business norm, but also naturally avoids the above troubles. Motivated by this, a blockchain-based fair outsourcing scheme to support payment in fiat money is proposed in this paper. The proposed scheme has broad compatibility and, as an example, is subsequently instantiated with a conventional outsourcing solution of eigen-decomposition. The performance of the scheme in fairness, privacy, and efficiency is verified by both theoretical and experimental evaluations.
Xiangli Xiao, Yushu Zhang 0001, Xuewen Dong, Liangmin Wang 0001, Yong Xiang 0001, Xiaochun Cao
IEEE Trans. Serv. Comput.5
2023 A Lightweight Model-Based Evolutionary Consensus Protocol in Blockchain as a Service for IoT
abstract
Internet of Things (IoT) is experiencing fast proliferation with emerging trends in autonomy and local decision-making to avoid the explosive burden on network infrastructure between cloud and edge. Thereby, blockchain as a Service (BaaS) for IoT, as an emerging distributed services computing paradigm, has drawn intense attention due to its decentralization, auditability, and tamper-resistance. However, the primary challenge is to design a tailor-made consensus protocol that is applicable to BaaS for IoT. Existing consensus protocols generally focus on power-intensive environments, which is not feasible for power-constrained BaaS-enabled IoT systems. In this article, to fully exploit BaaS's superiority (e.g., to sharing data securely), we propose a lightweight model-based evolutionary consensus protocol called Proof of Evolutionary Model (PoEM) that can improve the quality of BaaS in IoT environments. Beyond existing rule-based consensus protocols, PoEM iteratively trains a machine learning model to achieve consensus. In this way, PoEM enhances consensus efficiency and enables low-performance IoT devices to be involved. Moreover, considering IoT environments’ dynamics, a novel mechanism is designed to manage nodes joining and exiting dynamically. Extensive analytical and experimental results show PoEM's improved consensus efficiency and applicability in dynamic BaaS-based IoT environments while providing high-level security guarantees.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Yushu Zhang 0001, Longxiang Gao
IEEE Trans. Serv. Comput.3
2023 Data Caching Optimization With Fairness in Mobile Edge Computing
abstract
Mobile edge computing (MEC) provides a new computing paradigm that can overcome the inability of the traditional cloud computing paradigm to ensure low service latency by pushing computing power and resources to the network edge. Many studies have attempted to formulate edge data caching strategies for app vendors to optimize caching performance by caching the right data on the right edge servers. However, existing edge data caching approaches have unfortunately ignored fairness, which is an important issue from the app vendor's perspective. In general, an app vendor needs to cache data on edge servers to serve its users with insignificant latency differences at a minimum caching cost. In this paper, we make the first attempt to tackle the fair edge data caching (FEDC) problem. Specifically, we formulate the FEDC problem as a constraint optimization problem (COP) and prove its$\mathcal {NP}$-hardness. An optimal approach named FEDC-OPT is proposed to find optimal solutions to small-scale FEDC problems with integer programming technique. In addition, an approximate algorithm named FEDC-APX is proposed to find approximate solutions in large-scale FEDC problems. The performance of the proposed approaches is analyzed theoretically, and evaluated experimentally on a widely-used real-world data set against four representative approaches. The experimental results show that the proposed approaches can solve the FEDC problem efficiently and effectively.
Feifei Chen 0001, Qiang He 0001, Xiaoyu Xia 0001, Rui Wang 0008, Yong Xiang 0001
IEEE Trans. Serv. Comput.6
2023 CoSS: Leveraging Statement Semantics for Code Summarization
abstract
Automated code summarization tools allow generating descriptions for code snippets in natural language, which benefits software development and maintenance. Recent studies demonstrate that the quality of generated summaries can be improved by using additional code representations beyond token sequences. The majority of contemporary approaches mainly focus on extracting code syntactic and structural information from abstract syntax trees (ASTs). However, from the view of macro-structures, it is challenging to identify and capture semantically meaningful features due to fine-grained syntactic nodes involved in ASTs. To fill this gap, we investigate how to learn more code semantics and control flow features from the perspective of code statements. Accordingly, we propose a novel model entitled CoSS for code summarization. CoSS adopts a Transformer-based encoder and a graph attention network-based encoder to capture token-level and statement-level semantics from code token sequence and control flow graph, respectively. Then, after receiving two-level embeddings from encoders, a joint decoder with a multi-head attention mechanism predicts output sequences verbatim. Performance evaluations on Java, Python, and Solidity datasets validate that CoSS outperforms nine state-of-the-art (SOTA) neural code summarization models in effectiveness and is competitive in execution efficiency. Further, the ablation study reveals the contribution of each model component.
Chaochen Shi, Borui Cai, Yao Zhao 0006, Longxiang Gao, Keshav Sood, Yong Xiang 0001
IEEE Trans. Software Eng.6
2023 A novel feature-based framework enabling multi-type DDoS attacks detection
abstract
Abstract Distributed Denial of Service (DDoS) attacks are among the most severe threats in cyberspace. The existing methods are only designed to decide whether certain types of DDoS attacks are ongoing. As a result, they cannot detect other types of attacks, not to mention the even more challenging mixed DDoS attacks. In this paper, we comprehensively analyzed the characteristics of various types of DDoS attacks and innovatively proposed five new features from heterogeneous packets including entropy rate of IP source flow, entropy rate of flow, entropy of packet size, entropy rate of packet size, and number of ICMP destination unreachable packet to detect not only various types of DDoS attacks, but also the mixture of them. The experimental results show that the proposed fives features ranked at the top compared with other common features in terms of effectiveness. Besides, by using these features, our proposed framework outperforms the existing methods when detecting various DDoS attacks and mixed DDoS attacks. The detection accuracy improvements over the existing methods are between 21% and 53%.
Lu Zhou 0003, Ye Zhu 0002, Yong Xiang 0001, Tianrui Zong
World Wide Web (WWW)3
2022 BASS: Blockchain-Based Asynchronous SignSGD for Robust Collaborative Data Mining
abstract
Federated learning (FL) is a machine learning framework for collaborative data mining in many scenarios (e.g. Internet of Things) due to its privacy-preserving feature. However, various attacks arise security concerns of FL, such as poisoning, backdoor, and DDoS attacks. Several blockchain-based FL schemes strengthen credibility and security without considering the increased communication overhead. Some existing work compresses local updated gradients to sign vectors to lower communication overhead at the expense of model accuracy. To address the above concerns, this paper offers a blockchain-based asynchronous SignSGD (BASS) scheme. A novel asynchronous sign aggregation algorithm is introduced to ensure model accuracy even if the local updated gradients are compressed to sign vectors. Considering the unstable network connection on IoT, a consensus algorithm that elects multiple leader nodes enables reliable global model aggregation. The introduced blockchain improves credibility and security without downgrading efficiency. Empirical studies show that BASS outperforms other schemes in efficiency, model accuracy, and security.
Chenhao Xu 0003, Youyang Qu, Yong Xiang 0001, Longxiang Gao, David B. Smith 0001, Shui Yu 0001
DSAA3
2022 Impersonation Attack Detection in IoT Networks
abstract
The deployment of Internet of Things (IoT) networks is growing at an extraordinary speed from last decade and has expanded the interconnection of billions of nodes, providing a range of flexible communication and computing services, etc. We note that this significant expansion of the IoT surface has expanded the attack surfaces and is a danger to companies of every size from security aspects. The IoT devices are easy to compromise and therefore the attacker can easily act as an impersonator to impersonate other legitimate IoT nodes. This is known as impersonation attacks or spoofing attacks in wireless IoT networks. In this paper, we propose a new methodology to detect an impersonation attack in IoT networks. We use Mahalanobis Distance correlation theory based two-stage attack detection model to resist IoT node spoofing. The approach is evaluated on cloud platforms and is compared with the recent state-of-the-art literature. The proposal is deployed as a pluggable module in cloud networks. The key metrics of our evaluation and comparisons are accuracy with respect to the varying size of the IoT network, classification metrics, attack detection time, and CPU utilization.
Dinh Duc Nha Nguyen, Keshav Sood, Yong Xiang 0001, Longxiang Gao, Lianhua Chi
GLOBECOM3
2022 A Compressed Sensing Based Image Compression-Encryption Coding Scheme without Auxiliary Information Transmission
abstract
Facing the explosive growth of image data, how to realize low-cost compression coding and how to protect the confidentiality of image data have become two important research topics. In this work, a novel image coding scheme is proposed, which combines compression and encryption under the framework of compressed sensing (CS). Firstly, the signal is sampled using CS, then normalized and quantified. Next, according to the statistical characteristic of the measurements, the partially quantized measurements are selected for single-valued diffusion, and at the same time the auxiliary information, including the energy information and the quantization parameters, are automatically embedded. In contrast to other existing schemes, our scheme simultaneously achieves two security objectives: one is avoiding energy leakage in CS-based cryptosystem; the other is resisting Chosen Plaintext Attack (CPA). Furthermore, without transmitting the auxiliary information, the receiver can perform decryption correctly. Experimental results and analysis also demonstrate the effectiveness and security of the proposed scheme.
Di Xiao 0001, Yong Xiang 0001, Robin Doss
ICC3
2022 Attention Distraction: Watermark Removal Through Continual Learning with Selective Forgetting
abstract
Fine-tuning attacks are effective in removing the embedded watermarks in deep learning models. However, when the source data is unavailable, it is challenging to just erase the watermark without jeopardizing the model performance. In this context, we introduce Attention Distraction (AD), a novel source data-free watermark removal attack, to make the model selectively forget the embedded watermarks by customizing continual learning. In particular, AD first anchors the model's attention on the main task using some unlabeled data. Then, through continual learning, a small number of lures (randomly selected natural images) that are assigned a new label distract the model's attention away from the watermarks. Experimental results from different datasets and networks corroborate that AD can thoroughly remove the watermark with a small resource budget without compromising the model's performance on the main task, which outperforms the state-of-the-art works.
Leo Yu Zhang, Shengshan Hu, Longxiang Gao, Jun Zhang 0010, Yong Xiang 0001
ICME6
2022 FedEWA: Federated Learning with Elastic Weighted Averaging
abstract
Federated Learning (FL) offers a novel distributed machine learning context whereby a global model is collaboratively learned through edge devices without violating data privacy. However, intrinsic data heterogeneity in the federated network can induce model heterogeneity, thus posing a great challenge to the server-side model aggregation performance. Existing FL algorithms widely adopt model-wise weighted averaging for client models to generate the new global model, which emphasizes the importance of the holistic model but ignores the importance of distinctions between internal parameters of various client models. In this paper, we propose a novel parameter-wise elastic weighted averaging aggregation approach to realize the rapid fusion of heterogeneous client models. Specifically, each client evaluates the importance of model internal parameters in the model update and obtains the corresponding parameter importance coefficient vector; the server implements the parameter-wise weighted averaging for each parameter based on their importance coefficient vectors, thereby aggregating a new global model. Extensive experiments on MNIST and CIFAR-10 datasets with diverse network architectures and hyper-parameter combinations show that our proposed algorithm outperforms the existing state-of-the-art FL algorithms on the performance of heterogeneous model fusion.
Atul Sajjanhar, Yong Xiang 0001, Xiaojun Tong, Shan Zeng
IJCNN3
2022 Adversarial Training for Robust Insider Threat Detection
abstract
Insider threat analysis techniques based on machine learning provide convenient and effective automated detection of internally generated cyberattacks. When data are manipulated by adding slight perturbations, the threat intelligence models result in misclassifications of highly skewed class distribution with rare occurrences of events in insider threats. This paper proposes a generative model WGAN-GP conditioned by the class labels, referred to as CWGAN-GP, for insider threat analysis to create synthetic data samples for the rare malicious activities and shows that it generalizes well across different learning algorithms. Further, the robustness of the supervised algorithms to unknown inputs have not been investigated in any other works. This study explores how the synthetically created adversarial samples can increase the robustness of supervised models using adversarial training. We use a target classifier as threat model to generate one-step and iterative adversarial samples and perform a non-targeted test-time attack on the classifiers. We evaluate the robustness of various learning models against synthetic data from other data generation methods and demonstrate that the adversarial training using data generated from CWGAN-GP is less susceptible to adversarial attacks on insider threat classifiers using multiple versions of benchmark CMU CERT data set.
R. G. Gayathri, Atul Sajjanhar, Yong Xiang 0001
IJCNN3
2022 A Blockchain-based Multi-layer Decentralized Framework for Robust Federated Learning
abstract
With the expansion of the Internet of Things (IoT) development and application, federated learning has gained higher popularity in industrial researching fields. However, the security issues in federated learning have become hot-spots in the research area, such as privacy-preserving and poisoning attacks. This paper proposes a robust blockchained multi-layer decentralized federated learning (RBML-DFL) framework to ensure the federated learning's robustness. Firstly, by adopting the three-layered framework, the blockchain connects the federated learning components to secure the privacy and data safety of federated learning. Secondly, the proposed framework provides resilience on poisoning attacks to the central model compared to typical federated learning frameworks. Lastly, the decentralized structure associated with the blockchain tracing back mechanism can prevent the central server failure or mal-function compared to centralized federated learning. We evaluate and compare the proposed framework with other state-of-the-art federated learning frameworks on the accuracy, latency, and system robustness under poisoning attacks. The results show that the proposed RBML-DFL framework outperforms state-of-the-art baseline frameworks on all three metrics: accuracy, latency, and the robustness of the federated learning.
Di Wu 0050, Nai Wang, Jiale Zhang 0001, Yuan Zhang 0007, Yong Xiang 0001, Longxiang Gao
IJCNN5
2022 Towards Accurate Knowledge Transfer between Transformer-based Models for Code Summarization
abstract
Automatic code summarization generates high-level natural language descriptions of code snippets, which can benefit software maintenance and code comprehension.Recently, Transformer-based models achieved state-of-the-art performance on code summarization tasks.However, there are data gaps in neural model training for some programming languages.To fill this gap, we propose a novel transfer learning approach to accurately transfer knowledge between Transformer-based models.We train a discriminator to identify which heads of the multi-head attention module should be transferred.On this basis, we define a transfer strategy of parameter matrices.We evaluated the proposed transfer learning approach on four state-of-the-art Transformer-based code summarization models.Experimental results show that models with transferred knowledge outperform original models up to 10.70% in BLEU, 5.36% in ROUGE-L, and 4.34% in METEOR.
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Longxiang Gao
SEKE2
2022 A Bytecode-based Approach for Smart Contract Classification
abstract
With the development of blockchain technologies, the number of smart contracts deployed on blockchain platforms is growing exponentially, which makes it difficult for users to find desired services by manual screening. The automatic classification of smart contracts can provide blockchain users with keyword-based contract searching and helps to manage smart contracts effectively. Current research on smart contract classification focuses on Natural Language Processing (NLP) solutions which are based on contract source code. However, more than 94% of smart contracts are not open-source, so the application scenarios of NLP methods are very limited. Meanwhile, NLP models are vulnerable to adversarial attacks. This paper proposes a classification model based on features from contract bytecode instead of source code to solve these problems. We also use feature selection and ensemble learning to optimize the model. Our experimental studies on over 11K real-world Ethereum smart contracts show that our model can classify smart contracts without source code and has better performance than baseline models. Our model also has good resistance to adversarial attacks compared with NLP-based models. In addition, our analysis reveals that account features used in many smart contract classification models have little effect on classification and can be excluded.
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Longxiang Gao, Keshav Sood, Robin Doss
SANER2
2022 A novel privacy protection scheme for location-based services using collaborative caching
Nisha Nisha, Iynkaran Natgunanathan, Shang Gao 0003, Yong Xiang 0001
Comput. Networks4
2022 Privacy-preserving blockchain-enabled federated learning for B5G-Driven edge computing
Yichen Wan, Youyang Qu, Longxiang Gao, Yong Xiang 0001
Comput. Networks4
2022 A feature selection-based method for DDoS attack flow classification
Lu Zhou 0003, Ye Zhu 0002, Tianrui Zong, Yong Xiang 0001
Future Gener. Comput. Syst.4
2022 LSP: Lightweight Smart-Contract-Based Transaction Prioritization Scheme for Smart Healthcare
abstract
In recent years, several blockchain-based models have emerged to provide a secure way to store and access sensitive electronic medical records (EMRs) across the healthcare sector. These records are of different priorities and business requirements. From our comprehensive literature review, we observe that the existing models have no provision of prioritizing the EMR transactions. This critically affects the quick and streamline sharing of emergency EMRs in a smart healthcare environment. Furthermore, the lack of prioritization significantly restricts the optimal usage of the blockchain network. Motivated by this, we first propose a lightweight and deterministic method to prioritize the flow of emergency healthcare transactions through the smart contract. We also propose logical stateless transaction models for different entities involved in the system with varying levels of trust. Finally, the performance of the model based on the private Ethereum is verified and it outperforms the existing benchmark model in terms of usefulness in the healthcare setting and computation overhead with the use of a simple prioritizing algorithm. The obtained results demonstrate the feasibility of the proposed scheme in the real-time smart healthcare system.
Akanksha Saini, Dimaz Wijaya, Navneesh Kaur, Yong Xiang 0001, Longxiang Gao
IEEE Internet Things J.4
2022 A Lightweight and Attack-Proof Bidirectional Blockchain Paradigm for Internet of Things
abstract
Diverse technologies, such as machine learning and big data, have been driving the prosperity of the Internet of Things (IoT) and the ubiquitous proliferation of IoT devices. Consequently, it is natural that IoT becomes the driving force to meet the increasing demand for frictionless transactions. To secure transactions in IoT, blockchain is widely deployed since it can remove the necessity of a trusted central authority. However, the mainstream blockchain-based IoT payment platforms, dominated by Proof-of-Work (PoW) and Proof-of-Stake (PoS) consensus algorithms, face several major security and scalability challenges that result in system failures and financial loss. Among the three leading attacks in this scenario, double-spend attacks and long-range attacks threaten the tokens of blockchain users, while eclipse attacks target Denial of Service. To defeat these attacks, a novel bidirectional-linked blockchain (BLB) using chameleon hash functions is proposed, where bidirectional pointers are constructed between blocks. Furthermore, a new committee members auction (CMA) consensus algorithm is designed to improve the security and attack resistance of BLB while guaranteeing high scalability. In CMA, distributed blockchain nodes elect committee members through a verifiable random function. The smart contract uses Shamir’s secret-sharing scheme to distribute the trapdoor keys to committee members. To better investigate BLB’s resistance against double-spend attacks, an improved Nakamoto’s attack analysis is presented. In addition, a modified entropy metric is devised to measure eclipse attack resistance across different consensus algorithms. Extensive evaluation results show the superior resistance against attacks and demonstrate high scalability of BLB compared with current leading paradigms based on PoS and PoW.
Chenhao Xu 0003, Youyang Qu, Tom H. Luan, Peter W. Eklund, Yong Xiang 0001, Longxiang Gao
IEEE Internet Things J.5
2022 Multi-user image retrieval with suppression of search pattern leakage
Hong Liu 0025, Yushu Zhang 0001, Yong Xiang 0001, Bo Liu 0001, ErChuan Guo
Inf. Sci.3
2022 Noise-free thumbnail-preserving image encryption based on MSB prediction
Ye Zhu 0002, Yushu Zhang 0001, Xiangli Xiao, Rushi Lan, Yong Xiang 0001
Inf. Sci.6
2022 PRA-TPE: Perfectly Recoverable Approximate Thumbnail-Preserving Image Encryption
Xi Ye 0004, Yushu Zhang 0001, Rushi Lan, Yong Xiang 0001
J. Vis. Commun. Image Represent.5
2022 Robust Federated Averaging via Outlier Pruning
abstract
Federated Averaging (FedAvg) is the baseline Federated Learning (FL) algorithm that applies the stochastic gradient descent for local model training and the arithmetic averaging of the local models’ parameters for global model aggregation. Succeeding FL works commonly utilize the arithmetic averaging scheme of FedAvg for the aggregation. However, such arithmetic averaging is prone to the outlier model-updates, especially when the clients’ data are non-Independent and Identically Distributed (non-IID). As such, the classical aggregation approach suffers from the dominance of the outlier updates and, consequently, causes high communication costs towards producing a decent global model. In this letter, we propose a robust aggregation strategy to alleviate the above issues. In particular, we propose first pruning the node-wise outlier updates (weights) from the local trained models and then performing the aggregation on the selected effective weights-set at each node. We provide the theoretical result of our method and conduct extensive experiments on the MNIST, CIFAR-10, and Shakespeare datasets with IID and non-IID settings, which demonstrate that our aggregation approach outperforms the state-of-the-art methods in terms of communication speedup, test-set performance and training convergence.
Md Palash Uddin, Yong Xiang 0001, John Yearwood, Longxiang Gao
IEEE Signal Process. Lett.2
2022 Blockchain-Based Audio Watermarking Technique for Multimedia Copyright Protection in Distribution Networks
abstract
Copyright protection in multimedia protection distribution is a challenging problem. To protect multimedia data, many watermarking methods have been proposed in the literature. However, most of them cannot be used effectively in a multimedia distribution network (MDN) as they are not designed to support multi-layer watermark embedding. Multi-layer watermarking mechanisms were developed to protect multimedia data across different layers in an MDN. However, in those mechanisms, we need to trust the entities in the MDN, such as regional and country distributors. To overcome this potential drawback, in this article, we propose a novel privacy protection mechanism for MDNs by combining the advantages of both blockchain and watermarking technologies. A specifically designed watermarking algorithm is used to link the copyright information with the audio file, while a novel blockchain-based smart contract mechanism is developed to enforce the proper functioning of each entity in the distribution network. Moreover, the new audio mechanism is computationally efficient. Although audio signals are used to show the effectiveness of the proposed mechanism, the proposed approach can easily be extended to other multimedia objects, such as an image. The validity of the proposed mechanism is demonstrated by our simulation results. The proposed mechanism can benefit multimedia production companies and other entities in the MDN.
Iynkaran Natgunanathan, Purathani Praitheeshan, Longxiang Gao, Yong Xiang 0001, Lei Pan 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Cost-Friendly Differential Privacy of Smart Meters Using Energy Storage and Harvesting Devices
abstract
Cost-friendly differential privacy (CDP) of smart meters can be preserved by an appropriate charging and discharging mechanism that uses rechargeable batteries (RBs) to generate Laplace distributed random noise. However, the existing CDP methods have several issues. First, the maximum discharge rate of an RB requires to vary with the maximal consumption of houses. Second, the probability of an RB to charge/discharge depends on the demand, regardless of the state-of-charge (SoC) of an RB. Third, in extreme SoC (near-empty or almost fully charged) of an RB, no noise added to the demand. To overcome these, we propose a mechanism in which a novel probability density function is designed to generate near Laplace distributed random noise. We also utilize a renewable energy source with small storage in cascade with an RB to enhance performance. Both theoretical analysis and simulations are performed to demonstrate the effectiveness of our proposed method.
Mohammad Belayet Hossain, Iynkaran Natgunanathan, Yong Xiang 0001, Yushu Zhang 0001
IEEE Trans. Serv. Comput.3
2022 Desynchronization-attack-resilient audio watermarking mechanism for stereo signals using the linear correlation between channels
Tianrui Zong, Juan Zhao 0007, Yong Xiang 0001, Iynkaran Natgunanathan, Longxiang Gao, Wanlei Zhou 0001
World Wide Web3
2021 A Comprehensive Feature Importance Evaluation for DDoS Attacks Detection
Lu Zhou 0003, Ye Zhu 0002, Yong Xiang 0001
ADMA3
2021 Differentially Privacy-Preserving Federated Learning Using Wasserstein Generative Adversarial Network
abstract
Artificial intelligence (AI) requires a large amount of data to train high-quality machine learning (ML) models. However, due to privacy issues, individuals or organizations are not willing to share data with others, which results in “data islands”. This motivates the emergence of Federated Learning (FL), a novel ML framework allowing clients to exchange model parameters rather than the raw data. Unfortunately, the private data may be reconstructed by malicious participants by exploiting the context of model parameters in FL. This poses further challenges to privacy protection. To address this issue, we propose to integrate Wasserstein Generative Adversarial Network (WGAN) and differential privacy (DP) to protect the model parameters. WGAN is used to generate controllable random noise, which is then injected into model parameters. The new mechanism satisfies DP requirements while the data utility is highly improved. We experimentally demonstrate superior performances from aspects of convergence, accuracy, and data utility.
Yichen Wan, Youyang Qu, Longxiang Gao, Yong Xiang 0001
ISCC4
2021 BAFL: An Efficient Blockchain-Based Asynchronous Federated Learning Framework
abstract
With the widespread of 5G networks, the application of Federated Learning (FL) in Internet of Things (IoT) has become a trend. However, the trust problem caused by the centralized aggregation server, and the inefficiency problem caused by the low-performance devices, are still key challenges. Several studies involving asynchronous FL have been conducted to accelerate the training process, but they usually have a decreased model performance. In this paper, a blockchain-based asynchronous federated learning framework with a dynamic scaling factor is proposed. By adopting the blockchain, the trust problem among devices can be addressed. Meanwhile, the novel dynamic scaling factor is proposed to help improve the FL efficiency and accuracy. Extensive experiments are conducted on heterogeneous devices and the results show that the proposed framework mitigates the impact of low-performance devices while being as efficient as traditional FL with the extra benefit of alleviating the trust problem among IoT devices.
Chenhao Xu 0003, Youyang Qu, Peter W. Eklund, Yong Xiang 0001, Longxiang Gao
ISCC4
2021 Anomaly Detection for Scenario-based Insider Activities using CGAN Augmented Data
abstract
Insider threats are the cyber attacks from the trusted entities within an organization. An insider attack is hard to detect as it may not leave a footprint and potentially cause huge damage to organizations. Anomaly detection is the most common approach for insider threat detection. Lack of real-world data and the skewed class distribution in the datasets makes insider threat analysis an understudied research area. In this paper, we propose a Conditional Generative Adversarial Network (CGAN) to enrich under-represented minority class samples to provide meaningful and diverse data for anomaly detection from the original malicious scenarios. Comprehensive experiments performed on benchmark dataset demonstrates the effectiveness of using CGAN augmented data, and the capability of multi-class anomaly detection for insider activity analysis. Moreover, the method is compared with other existing methods against different parameters and performance metrics.
R. G. Gayathri, Atul Sajjanhar, Yong Xiang 0001, Xingjun Ma
TrustCom3
2021 Variational auto-encoder based Bayesian Poisson tensor factorization for sparse and imbalanced count data
Ming Liu 0028, Ruohua Xu, Lan Du 0002, Longxiang Gao, Yong Xiang 0001
Data Min. Knowl. Discov.7
2021 Scene image representation by foreground, background and hybrid features
Chiranjibi Sitaula, Yong Xiang 0001, Sunil Aryal, Xuequan Lu
Expert Syst. Appl.2
2021 Self-supervised cross-iterative clustering for unlabeled plant disease images
Uno Fang, Jianxin Li 0001, Xuequan Lu, Longxiang Gao, Mumtaz Ali 0003, Yong Xiang 0001
Neurocomputing6
2021 A Smart-Contract-Based Access Control Framework for Cloud Smart Healthcare System
abstract
In current healthcare systems, electronic medical records (EMRs) are always located in different hospitals and controlled by a centralized cloud provider. However, it leads to single point of failure as patients being the real owner lose track of their private and sensitive EMRs. Hence, this article aims to build an access control framework based on smart contract, which is built on the top of distributed ledger (blockchain), to secure the sharing of EMRs among different entities involved in the smart healthcare system. For this, we propose four forms of smart contracts for user verification, access authorization, misbehavior detection, and access revocation, respectively. In this framework, considering the block size of ledger and huge amount of patient data, the EMRs are stored in cloud after being encrypted through the cryptographic functions of elliptic curve cryptography (ECC) and Edwards-curve digital signature algorithm (EdDSA), while their corresponding hashes are packed into blockchain. The performance evaluation based on a private Ethereum system is used to verify the efficiency of proposed access control framework in the real-time smart healthcare system.
Akanksha Saini, Qingyi Zhu, Navneet Singh, Yong Xiang 0001, Longxiang Gao, Yushu Zhang 0001
IEEE Internet Things J.4
2021 Low-Cost and Confidentiality-Preserving Multi-Image Compressed Acquisition and Separate Reconstruction for Internet of Multimedia Things
abstract
Internet of Multimedia Things (IoMT) is facing how to achieve low-cost compression and acquisition of multimedia big data at the resource-constrained side while preserving the confidentiality of the data. In this article, we propose a low-cost and confidentiality-preserving multi-image compressed acquisition model and provide separate image reconstruction. We group a series of image sets and randomly sample them with compressive sensing in each group. It is noteworthy that we harness the sigmoid sequence to fuse the measurement of each group to alleviate the uneven distribution of the measurement. The experiment verifies that the proposed method avoids the energy information leakage of the original signal from the measurement. Subsequently, we assemble the sampling result of each group into a big master image with a suitable size and further encrypt it. The encrypted master image is transmitted to the cloud for storage and decryption service. The cloud performs a separate reconstruction of images. It provides different image reconstruction services for a fused measurement according to realistic demands, which reduces unnecessary resource consumption. Simulation results show the effectiveness and security of our acquisition scheme, which indicates the proposal can have a potential application to IoMT.
Mengdi Wang 0005, Di Xiao 0001, Yong Xiang 0001
IEEE Internet Things J.3
2021 DP-GAN: Differentially private consecutive data publishing using generative adversarial nets
Stella Ho, Youyang Qu, Bruce Gu, Longxiang Gao, Jianxin Li 0001, Yong Xiang 0001
J. Netw. Comput. Appl.6
2021 Reversible data hiding in encrypted color images using cross-channel correlations
Ming Li 0029, Hua Ren, Yong Xiang 0001, Yushu Zhang 0001
J. Vis. Commun. Image Represent.3
2021 Content and context features for scene image representation
Chiranjibi Sitaula, Sunil Aryal, Yong Xiang 0001, Anish Basnet, Xuequan Lu
Knowl. Based Syst.3
2021 Blockchain-based access control scheme with incentive mechanism for eHealth systems: patient as supervisor
Chenquan Gan, Akanksha Saini, Qingyi Zhu, Yong Xiang 0001, Zufan Zhang
Multim. Tools Appl.4
2021 Segmental DCT Coefficient Reversal Based Anti-Collusion Audio Fingerprinting Mechanism
abstract
Collusion attacks are challenging to tackle in audio fingerprinting. A new direction to resist collusion attacks is to degrade the perceptual quality of the colluded files so that these files cannot be reused. The existing method in this direction has low embedding capacity and limited anti-collusion performance when the number of colluders is odd. In this letter, we present an anti-collusion mechanism that has a higher embedding capacity and can significantly degrade the perceptual quality of the colluded files regardless of the number of colluders. In the proposed mechanism, we first segment the host audio file into frames and perform the discrete cosine transform (DCT) on each frame. Then multiple fingerprint bits are embedded into each frame by reversing the DCT coefficients of the corresponding frequency band. As a result, when a collusion attack occurs, our proposed embedding mechanism can introduce perceptibly annoying differences between frames in the colluded file, which leads to severe perceptual quality degradation. Theoretical analysis and experimental results validate the superiority of the proposed anti-collusion mechanism.
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Guang Hua 0001
IEEE Signal Process. Lett.3
2021 Desynchronization Attacks Resilient Watermarking Method Based on Frequency Singular Value Coefficient Modification
abstract
Desynchronization is a very challenging type of attack in audio watermarking. The traditional singular value decomposition (SVD) based audio watermarking methods embed the watermark information by modifying the singular value of individual segment, which have little resistance against desynchronization attacks. In this paper, we propose a novel frequency singular value coefficient (FSVC) feature, which reflects the ratio between the singular values of two consecutive segments and is insensitive to desynchronization attacks, to carry the watermark bits. To our best knowledge, it is the first time that the ratio between singular values is employed for audio watermarking. In the proposed method, the discrete cosine transform (DCT) is performed on two consecutive segments of the host audio signal and SVD is applied to the DCT coefficients of the mid frequency band of each segment to extract the FSVC. Then the watermark bits are embedded by adjusting the values of the FSVC. The watermark embedding procedure is optimized to minimize the perceptual quality degradation and an error buffer is created to enhance the robustness. As a result, the proposed method can achieve a much higher embedding capacity than the existing methods tackling desynchronization attacks. The impact of desynchronization and common signal processing attacks on the proposed watermarking method is mathematically modeled, and the effectiveness of the proposed method against these attacks is theoretically and experimentally validated.
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Wanlei Zhou 0001, Gleb Beliakov
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Non-Linear-Echo Based Anti-Collusion Mechanism for Audio Signals
abstract
Collusion attacks are considered to be challenging attacks in audio copyright protection. The traditional watermarking algorithms cannot identify the traitors when other attacks, such as desynchronization attacks, are applied with a collusion attack. Instead of tracing the traitors, in this paper we aim to tackle collusion attacks by removing the commercial value from the colluded copy, which will demotivate the attackers from launching collusion attacks. Since the commercial value of an audio signal is directly reflected by its perceptual quality, we propose a novel non-linear-echo generation (NLEG) based algorithm to significantly degrade the perceptual quality of the colluded copy by embedding a time delay sequence into the host signal. The proposed NLEG is also designed to be resilient to common signal processing attacks and desynchronization attacks. Furthermore, the proposed NLEG can be combined with other digital watermarking techniques to enhance its performance on protecting the copyright information. Experimental results show the validity of the proposed NLEG.
Tianrui Zong, Yong Xiang 0001, Iynkaran Natgunanathan, Longxiang Gao, Guang Hua 0001, Wanlei Zhou 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Computation Outsourcing Meets Lossy Channel: Secure Sparse Robustness Decoding Service in Multi-Clouds
abstract
This paper addresses the problem of lossy outsourcing, i.e., clients outsource computation needs to the cloud side through lossy channels, which is very common in practice. We focus on the case that the clients transmit 2D sparse signals to the semi-trusted clouds over packet-loss networks, and the clouds provide sparse robustness decoding service (SRDS) for the users. In order to achieve high level of efficiency and security, we propose to jointly exploit parallel compressive sensing for robust signal encoding and employ multiple cloud servers for SRDS. Specifically, prior to encoding, a signal is encrypted by only altering the indices and amplitudes of its non-zero entries. The encrypted signal is sensed using a Gaussian measurement matrix and the generated compressive measurements are then sent to multi-clouds for SRDS, along with the occurrence of packet loss. Each column in compressive measurements can be regarded as a packet and each description consists of a certain number of packets. Each description together with a small portion of support set is distributed to a cloud. When receiving the request from a user, each cloud performs SRDS using the acquired description, where the reconstructed signal is still in encrypted form so that the signal privacy is well preserved. After receiving the reconstructed signal, the user accomplishes the decryption operation. Experimental results show that the encryption algorithm improves compressibility and reconstruction performance compared with the case of no encryption, and the proposed privacy-assured outsourcing of SRDS is highly robust and efficient.
Yushu Zhang 0001, Jiantao Zhou 0001, Yong Xiang 0001, Leo Yu Zhang, Fei Chen 0003, Shaoning Pang 0001, Xiaofeng Liao 0001
IEEE Trans. Big Data3
2021 Guest Editorial: Advanced Complex Data Analytics for Smart City Industrial Environment
abstract
The papers in this special section focus on advanced complex data analytics for smart city industrial environments. With the continuous increase in size of populations living in cities, those residents face increasing environmental pressures and infrastructure needs. To deliver a better quality of life for these residents to address these demands at a sustainable cost, smart technologies can help cities meet these challenges, and have been become the next wave of public investment. It all starts with data to be generated by the residents living in the cities. This special section will focus on the ever-increasing challenges of big complex data like social network data, traffic network data, and IoT network data altogether in the industrial environment of smart city.
Jianxin Li 0001, Yong Xiang 0001, Timos K. Sellis, Jie Xiong 0001
IEEE Trans. Ind. Informatics2
2021 A Blockchained Federated Learning Framework for Cognitive Computing in Industry 4.0 Networks
abstract
Cognitive computing, a revolutionary AI concept emulating human brain's reasoning process, is progressively flourishing in the Industry 4.0 automation. With the advancement of various AI and machine learning technologies the evolution toward improved decision making as well as data-driven intelligent manufacturing has already been evident. However, several emerging issues, including the poisoning attacks, performance, and inadequate data resources, etc., have to be resolved. Recent research works studied the problem lightly, which often leads to unreliable performance, inefficiency, and privacy leakage. In this article, we developed a decentralized paradigm for big data-driven cognitive computing (D2C), using federated learning and blockchain jointly. Federated learning can solve the problem of “data island” with privacy protection and efficient processing while blockchain provides incentive mechanism, fully decentralized fashion, and robust against poisoning attacks. Using blockchain-enabled federated learning help quick convergence with advanced verifications and member selections. Extensive evaluation and assessment findings demonstrate D2C's effectiveness relative to existing leading designs and models.
Youyang Qu, Shiva Raj Pokhrel, Sahil Garg, Longxiang Gao, Yong Xiang 0001
IEEE Trans. Ind. Informatics5
2021 Privacy-Assured FogCS: Chaotic Compressive Sensing for Secure Industrial Big Image Data Processing in Fog Computing
abstract
In the age of the industrial big data, there are several significant problems such as high-overhead data acquisition, data privacy leakage, and data tampering. Fog computing capability is rapidly expanding to address not only network congestion issues but data security issues. This article presents a chaotic compressive sensing (CS) scheme for securely processing industrial big image data in the fog computing paradigm, called privacy-assured FogCS. Specially, the sine logistic modulation map is used to drive the privacy-assured, authenticated, and block CS for secure image data collection in the sensor nodes. After sampling, the measurements are normalized in the fog nodes. The normalized measurements can achieve the perfect secrecy and their energy values are further masked through the proposed permutation-diffusion architecture. Finally, these relevant data are transmitted to the clouds (data centers) for storage, reconstruction, and authentication if required. In addition, a hardware implementation reference on a field programmable gate array is designed. Simulation analyses show the feasibility and efficiency of the privacy-assured FogCS scheme.
Yushu Zhang 0001, Ping Wang 0029, Hui Huang 0008, Youwen Zhu, Di Xiao 0001, Yong Xiang 0001
IEEE Trans. Ind. Informatics6
2021 Robust Coding of Encrypted Images via 2D Compressed Sensing
abstract
In many practical scenarios, image encryption should be implemented before image compression. This leads to the requirement of compressing encrypted images. Compressed sensing (CS), a breakthrough in signal processing, has been demonstrated to be an effective method for compressing encrypted images with robustness. However, for the exiting CS-based image encryption-then-compression (ETC) systems, image encryption is usually performed by using linear operations. When linear operations are used, we cannot achieve low computational complexity and high security in the meantime. To solve this problem, a novel 2D CS (2DCS) based ETC (2DCS-ETC) scheme is proposed in this paper. First, two nonlinear operations, including global random permutation (GRP) and negative-positive transformation (NPT), are utilized to encrypt the original image for high security purpose. Second, the encrypted image is compressed by using 2DCS for low computational complexity purpose. Furthermore, a gray mapping operation is embedded prior to CS encoding. Since gray mapping strategy can reduce the dynamic range of the CS samples, this strategy is also helpful for the rate distortion (R-D) performance improvement. Third, a 2D projected gradient with embedding decryption (2DPG-ED) algorithm is proposed, which can be utilized for the original image reconstruction even if the encrypted image is not sparse anymore. Compared with the previous CS-based ETC methods, the proposed approach can simultaneously achieve high security and low computational complexity with better robustness.
Bo Zhang 0030, Di Xiao 0001, Yong Xiang 0001
IEEE Trans. Multim.3
2021 Mutual Information Driven Federated Learning
abstract
Federated Learning (FL) is an emerging research field that yields a global trained model from different local clients without violating data privacy. Existing FL techniques often ignore the effective distinction between local models and the aggregated global model when doing the client-side weight update, as well as the distinction of local models for the server-side aggregation. In this article, we propose a novel FL approach with resorting to mutual information (MI). Specifically, in client-side, the weight update is reformulated through minimizing the MI between local and aggregated models and employing Negative Correlation Learning (NCL) strategy. In server-side, we select top effective models for aggregation based on the MI between an individual local model and its previous aggregated model. We also theoretically prove the convergence of our algorithm. Experiments conducted on MNIST, CIFAR-10, ImageNet, and the clinical MIMIC-III datasets manifest that our method outperforms the state-of-the-art techniques in terms of both communication and testing performance.
Md Palash Uddin, Yong Xiang 0001, Xuequan Lu, John Yearwood, Longxiang Gao
IEEE Trans. Parallel Distributed Syst.2
2021 Effective Quarantine and Recovery Scheme Against Advanced Persistent Threat
abstract
Advanced persistent threat (APT) for cyber espionage poses a great threat to modern organizations. In order to mitigate the impact of APT on an organization, all the compromised systems in the organization must be quarantined and recovered in a timely and effective way. This article focuses on the problem of customizing a dynamic quarantine and recovery (QAR) scheme for an organization so that the APT impact is minimized. Based on a novel node-level epidemic model characterizing the effect of the QAR scheme on the expected state of the underlying network, we estimate the expected impact of APT under a QAR scheme. On this basis, we model the original problem as an optimal control problem. By use of optimal control theory, we derive the optimality system for the optimal control problem and thereby introduce the concept of normal potential optimal (NPO) control. Next, through comparative experiments, we find that the NPO control outperforms a set of heuristic controls. Hence, the QAR scheme associated with the NPO control is satisfactory in terms of the effectiveness of defending against APT. Finally, we examine the effect of some factors on the expected APT impact under the NPO control. This article would be helpful to the defense against APT for cyber espionage.
Lu-Xing Yang, Pengdeng Li, Xiaofan Yang 0001, Yong Xiang 0001, Frank Jiang 0001, Wanlei Zhou 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2020 Reliable Customized Privacy-Preserving in Fog Computing
abstract
Fog computing is an emergent computing paradigm that extends the cloud paradigm to the edge. With the explosive growth of smart devices and massive data generated everyday, cloud computing no longer matches the requirements of the Internet of Things (IoT) era, such as low latency, uninterrupted service and location awareness. Thus, fog computing has been introduced as a complement of the current cloud computing model to meet the requirements in IoT. Fog computing is a relatively new networking paradigm and considered as a promising solution to support IoT scenarios. On the one hand, fog computing inherits many features from cloud; on the other hand, fog computing also inherits some challenges and issues from cloud computing: privacy issue is one of them. In this paper, we propose a personalized differential privacy model based on the distance between two fog nodes in a fog network. We also identify the collusion attack in differential privacy framework which compromised the personalized Laplace function. Based on that, we develop a personalized differential privacy model, which not only eliminate this particular attack but also optimize the trade-off between privacy preserving and data utility.
Xiaodong Wang 0017, Bruce Gu, Youyang Qu, Yongli Ren, Yong Xiang 0001, Longxiang Gao
ICC5
2020 HDF: Hybrid Deep Features for Scene Image Representation
abstract
Nowadays it is prevalent to take features extracted from pre-trained deep learning models as image representations which have achieved promising classification performance. Existing methods usually consider either object-based features or scene-based features only. However, both types of features are important for complex images like scene images, as they can complement each other. In this paper, we propose a novel type of features - hybrid deep features, for scene images. Specifically, we exploit both object-based and scene-based features at two levels: part image level (i.e., parts of an image) and whole image level (i.e., a whole image), which produces a total number of four types of deep features. Regarding the part image level, we also propose two new slicing techniques to extract part based features. Finally, we aggregate these four types of deep features via the concatenation operator. We demonstrate the effectiveness of our hybrid deep features on three commonly used scene datasets (MIT-67, Scene-15, and Event-8), in terms of the scene image classification task. Extensive comparisons show that our introduced features can produce state-of-the-art classification accuracies which are more consistent and stable than the results of existing features across all datasets.
Chiranjibi Sitaula, Yong Xiang 0001, Anish Basnet, Sunil Aryal, Xuequan Lu
IJCNN2
2020 A Privacy Preserving Aggregation Scheme for Fog-Based Recommender System
Xiaodong Wang 0017, Bruce Gu, Youyang Qu, Yongli Ren, Yong Xiang 0001, Longxiang Gao
NSS5
2020 Protecting IP of Deep Neural Networks with Watermarking: A New Label Helps
Leo Yu Zhang, Jun Zhang 0010, Longxiang Gao, Yong Xiang 0001
PAKDD (2)5
2020 Protecting the Intellectual Property of Deep Neural Networks with Watermarking: The Frequency Domain Approach
abstract
Similar to other digital assets, deep neural network (DNN) models could suffer from piracy threat initiated by insider and/or outsider adversaries due to their inherent commercial value. DNN watermarking is a promising technique to mitigate this threat to intellectual property. This work focuses on black-box DNN watermarking, with which an owner can only verify his ownership by issuing special trigger queries to a remote suspicious model. However, informed attackers, who are aware of the watermark and somehow obtain the triggers, could forge fake triggers to claim their ownerships since the poor robustness of triggers and the lack of correlation between the model and the owner identity. This consideration calls for new watermarking methods that can achieve better trade-off for addressing the discrepancy. In this paper, we exploit frequency domain image watermarking to generate triggers and build our DNN watermarking algorithm accordingly. Since watermarking in the frequency domain is high concealment and robust to signal processing operation, the proposed algorithm is superior to existing schemes in resisting fraudulent claim attack. Besides, extensive experimental results on 3 datasets and 8 neural networks demonstrate that the proposed DNN watermarking algorithm achieves similar performance on functionality metrics and better performance on security metrics when compared with existing algorithms.
Leo Yu Zhang, Yajuan Du, Jun Zhang 0010, Yong Xiang 0001
TrustCom6
2020 Clustering Hashtags Using Temporal Patterns
Borui Cai, Guangyan Huang, Shuiqiao Yang, Yong Xiang 0001, Chihung Chi
WISE (1)4
2020 Robust Blockchain-Based Cross-Platform Audio Copyright Protection System Using Content-Based Fingerprint
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Gleb Beliakov
WISE (2)3
2020 Channel Correlation Based Robust Audio Watermarking Mechanism for Stereo Signals
Tianrui Zong, Yong Xiang 0001, Iynkaran Natgunanathan, Longxiang Gao, Wanlei Zhou 0001
WISE (2)2
2020 A technical survey on statistical modelling and design methods for crowdsourcing quality control
Mark J. Carman, Ye Zhu 0002, Yong Xiang 0001
Artif. Intell.4
2020 Decentralized Privacy Using Blockchain-Enabled Federated Learning in Fog Computing
abstract
As the extension of cloud computing and a foundation of IoT, fog computing is experiencing fast prosperity because of its potential to mitigate some troublesome issues, such as network congestion, latency, and local autonomy. However, privacy issues and the subsequent inefficiency are dragging down the performances of fog computing. The majority of existing works hardly consider a reasonable balance between them while suffering from poisoning attacks. To address the aforementioned issues, we propose a novel blockchain-enabled federated learning (FL-Block) scheme to close the gap. FL-Block allows local learning updates of end devices exchanges with a blockchain-based global learning model, which is verified by miners. Built upon this, FL-Block enables the autonomous machine learning without any centralized authority to maintain the global model and coordinates by using a Proof-of-Work consensus mechanism of the blockchain. Furthermore, we analyze the latency performance of FL-Block and further derive the optimal block generation rate by taking communication, consensus delays, and computation cost into consideration. Extensive evaluation results show the superior performances of FL-Block from the aspects of privacy protection, efficiency, and resistance to the poisoning attack.
Youyang Qu, Longxiang Gao, Tom H. Luan, Yong Xiang 0001, Shui Yu 0001, Gavin Zheng
IEEE Internet Things J.4
2020 Alleviating Heterogeneity in SDN-IoT Networks to Maintain QoS and Enhance Security
abstract
Software-defined networks (SDNs) offer unique and attractive solutions to solve challenging management issues in Internet of Things (IoT)-based large-scale multi-technological networks. SDN-IoT network collaboration is innovative and attractive but expected to be extremely heterogeneous in future generation IoT systems. For example, multi-technology network, network externality, and nodes heterogeneity in SDN-IoT may seriously affect the flow or application-specific quality-of-service (QoS) requirements. Furthermore, it highly influences security adoption in a network of interconnected IoT nodes. We observe that both QoS and security are interdependent and nonnegligible factors, thus we emphasize that in order to alleviate heterogeneity it is inevitable to study both these factors hand to hand (or vice versa). With this aim, first, we discuss significant and reasonable cases to encourage researchers to study QoS and security integrally in order to alleviate heterogeneity at SDN-IoT control plane. Second, we propose a framework which successfully transforms the m heterogeneous controllers to n homogeneous controller groups. The key metric of our observation and analysis is the SDN controller's response time. Following this, to validate our approach, we use the mathematical model and a proof of concept (PoC) in a virtual SDN ecosystem is demonstrated. From performance evaluation, we observe that the proposed framework significantly alleviates heterogeneity which helps to maintain QoS and enhance security. This fundamental analysis will enable network security individuals to deal heterogeneity, QoS, and security, of SDN-IoT, in more successful and promising ways.
Keshav Sood, Kallol Krishna Karmakar, Shui Yu 0001, Vijay Varadharajan, Shiva Raj Pokhrel, Yong Xiang 0001
IEEE Internet Things J.6
2020 A Fog-Based Recommender System
abstract
Fog computing is an emergent computing paradigm that extends the cloud paradigm. With the explosive growth of smart devices and mobile users, cloud computing no longer matches the requirements of the Internet of Things (IoT) era. Fog computing is a promising solution to satisfying these new requirements, such as low latency, uninterrupted service, and location awareness. As a typical new computing paradigm and network architecture, fog computing raises new challenges, such as privacy, data management, data analytics, information overload, and participatory sensing. In this article, we present a fog-based hybrid recommender system to address the issue of information overload in fog computing. Our proposed system not only abstracts useful information from the fog environment but can also be considered as an optimization tool due to its ability to provide suggestions to improve system performance. In particular, we demonstrate that the proposed system provides personalized and localized recommendations to users, and also advise the system itself to precache the content to optimize the storage capacity of the fog server.
Xiaodong Wang 0017, Bruce Gu, Yongli Ren, Shui Yu 0001, Yong Xiang 0001, Longxiang Gao
IEEE Internet Things J.6
2020 Cloud-assisted privacy-conscious large-scale Markowitz portfolio
Yushu Zhang 0001, Yong Xiang 0001, Ye Zhu 0002, Liangtian Wan, Xiyuan Xie
Inf. Sci.3
2020 Efficient authentication protocol with anonymity and key protection for mobile Internet users
Yan Jiang 0002, Youwen Zhu, Jian Wang 0038, Yong Xiang 0001
J. Parallel Distributed Comput.4
2020 Informed Histogram-Based Watermarking
abstract
The existing works on histogram-based watermark embedding share the common notions of non-informed random host sample selection and uniform embedding, resulting in limited and yet similar performances. This letter develops two methods to improve imperceptibility and robustness of histogram-based embedding. We first propose a content-aware method which performs uniform embedding in ranked energy-significant regions of the host signal. We show that the embedded ternary watermark signal is more likely to be masked by the strong host signal component, thus improving imperceptibility without compromising robustness. Further, a non-uniform embedding method is proposed, which modifies host samples according to their relationships with the center of the target histogram bin. It ensures minimum host sample modifications without reducing the payload size. Meanwhile, watermark robustness could also be improved since it becomes the most difficult to move samples at bin centers into other bins via attacks. The proposed methods are supported by extensive experimental results.
Guang Hua 0001, Yong Xiang 0001, Leo Yu Zhang
IEEE Signal Process. Lett.2
2020 A Novel Anti-Collusion Audio Fingerprinting Scheme Based on Fourier Coefficients Reversing
abstract
Most anti-collusion audio fingerprinting schemes are aiming at finding colluders from the illegal redistributed audio copies. However, the loss caused by the redistributed versions is inevitable. In this letter, a novel fingerprinting scheme is proposed to eliminate the motivation of collusion attack. The audio signal is transformed to the frequency domain by the Fourier transform, and the coefficients in frequency domain are reversed in different degrees according to the fingerprint sequence. Different from other fingerprinting schemes, the coefficients of the host media are excessively modified by the proposed method in order to reduce the quality of the colluded version significantly, but the imperceptibility is well preserved. Experiments show that the colluded audio cannot be reused because of the poor quality. In addition, the proposed method can also resist other common attacks. Various kinds of copyright risks and losses caused by the illegal redistribution are effectively avoided, which is significant for protecting the copyright of audio.
Ming Li 0029, Huimin Chang, Yong Xiang 0001, Dezhi An
IEEE Signal Process. Lett.3
2020 Achieving Secure and Efficient Dynamic Searchable Symmetric Encryption over Medical Cloud Data
abstract
In medical cloud computing, a patient can remotely outsource her medical data to the cloud server. In this case, only authorized doctors are allowed to access the data since the medical data is highly sensitive. Before outsourcing, the data is commonly encrypted, where the corresponding secret key is sent to authorized doctors. However, performing searches on encrypted medical data is difficult without decryption. In this paper, we propose two Secure and Efficient Dynamic Searchable Symmetric Encryption (SEDSSE) schemes over medical cloud data. First, we utilize the secure k-Nearest Neighbor (kNN) and Attribute-Based Encryption (ABE) techniques to construct a dynamic searchable symmetric encryption scheme, which can achieve forward privacy and backward privacy simultaneously. These tow security properties are vital and very challenging in the area of dynamic searchable symmetric encryption. Then, we propose an enhanced scheme to solve the key sharing problem which widely exists in the kNN based searchable encryption scheme. Compared with existing proposals, our schemes are better in terms of storage, search and updating complexity. Extensive experiments demonstrate the efficiency of our schemes on storage overhead, index building, trapdoor generating and query.
Hongwei Li 0001, Yi Yang 0027, Yuan-Shun Dai, Shui Yu 0001, Yong Xiang 0001
IEEE Trans. Cloud Comput.5
2020 Secure and Efficient Outsourcing of PCA-Based Face Recognition
abstract
Face recognition has become increasingly popular in recent years. However, in some special cases, many face recognition calculations cannot be performed effectively due to the lack of sufficient computing power of the terminal, which poses a challenge to the practical application of face recognition technology. Cloud computing provides a good platform for solving this problem due to its abundant computing resources. However, cloud computing poses new challenges, such as how to protect clients' data privacy without reducing efficiency. In this paper, we review some of the results of previous research and analyze an outsourcing protocol for eigen decomposition and singular value decomposition. On this basis, we propose a secure and efficient outsourcing protocol for face recognition through principal component analysis. In the proposed protocol, information privacy is well protected, and computational resources are saved by means of conversions of the original image information. In addition, local verification is supported to cope with the laziness of the cloud. We show the feasibility and advancement of our protocol from both theoretical and experimental perspectives.
Yushu Zhang 0001, Xiangli Xiao, Lu-Xing Yang, Yong Xiang 0001, Sheng Zhong 0002
IEEE Trans. Inf. Forensics Secur.4
2020 A Shared Bus Profiling Scheme for Smart Cities Based on Heterogeneous Mobile Crowdsourced Data
abstract
Mobile crowdsourcing (MCS), as an effective and crucial technique of Industrial Internet of Things, is enabling smart city initiatives in the real world. It aims at incorporating the intelligence of dynamic crowds to collect and compute decentralized ubiquitous sensing data that can be used to solve major urbanization problems such as traffic congestion. The shared bus, as a neotype transportation mode, aims at improving the resource utilization rate and maintaining the advantages of convenience and economy. In this article, we provide a scheme to profile shared buses through heterogeneous mobile crowdsourced data (TRProfiling). First, we design an MCS-based shared bus data generation and collection solution to overcome the aforementioned data scarcity issue. Then, we propose a travel profiling to profile resident travel and design a method called multiconstraint evolution algorithm to optimize the routes. Experimental results demonstrate that TRProfiling has an excellent performance in satisfying passengers’ travel requirements.
Xiangjie Kong 0001, Feng Xia 0001, Jianxin Li 0001, Mingliang Hou, Yong Xiang 0001
IEEE Trans. Ind. Informatics6
2020 Guest Editorial: AI and Machine Learning Solution Cyber Intelligence Technologies: New Methodologies and Applications
abstract
The papers in this special section focus on new methodologies and applications in artificial intelligence (AI) and machine learning (ML). With the recent development of machine learning (ML), artificial intelligence (AI) and cyber technologies in the field of industrial informatics, it is important to migrate the traditional businesses and services in the physical world to the digital cyber-enabled world. Cyber intelligence technologies, such as the fifth-generation (5G) mobile communication network, big data, Internet of Things (IoT), cloud computing, cognitive computing, ubiquitous computing and blockchains, enable goods, houses, services, information, and capitalization to be shared through the Internet of the Appendix]. Industrial applications, including various mechanical systems, utilities, supply chains, energy systems, power grids, infrastructures, manufactures, traffics, healthcare, and environmental issues, are partially operated or managed remotely with the influence of AI and cyber intelligence technology.
Ke Yan 0001, Lu Liu 0001, Yong Xiang 0001, Qun Jin
IEEE Trans. Ind. Informatics3
2020 A Low-Overhead, Confidentiality-Assured, and Authenticated Data Acquisition Framework for IoT
abstract
In the presence of several critical issues during data acquisition in industrial-informatics-based applications, like Internet of Things (IoT) and smart grid, this article proposes a novel framework based on compressive sensing (CS) and a cascade chaotic system (CCS). This framework can ensure low overhead, confidentiality, and authentication. Based on CS and the CCS, three technologies, including CCS-driven CS, CCS-driven local perturbation, and authentication mechanism, are introduced in the proposed data acquisition framework in this article. CCS-driven CS generates the measurement matrix with chaotic initial conditions and avoids the transmission of a large-size measurement matrix. CCS-driven local perturbation only perturbs a small number of elements in the original measurement matrix for each sampling and avoids the regeneration of the large-size measurement matrix. The authentication mechanism employs the authentication password and the access password to deal with the passive tampering attack and the active tampering attack, respectively. Moreover, the permutation-diffusion structure is used to encrypt the obtained measurements to enhance the security. Both theoretical and experimental analyses validate low overhead, confidentiality, and effective authentication of the proposed data acquisition framework for a number of industrial-informatics-based applications, such as IoT.
Yushu Zhang 0001, Guo Chen 0002, Xinpeng Zhang 0001, Yong Xiang 0001
IEEE Trans. Ind. Informatics5
2020 Non-Negative Matrix Factorization With Dual Constraints for Image Clustering
abstract
How to learn dimension-reduced representations of image data for clustering has been attracting much attention. Motivated by that the clustering accuracy is affected by both the prior-known label information of some of the images and the sparsity feature of the representations, we propose a non-negative matrix factorization (NMF) method with dual constraints in this paper. In our model, one constraint is used to keep the label feature and the other constraint is utilized to enhance the sparsity of the representations. Notably that these two constraints are embedded naturally into the traditional NMF model, refraining from the usage of the balance parameters which are hard to choose. Meantime, for solving the proposed model, the alternative iteration scheme is employed, and an efficient algorithm based on convex optimization is designed to conduct each iteration operation. It is proved that this algorithm achieves a nonlinear convergence rate, much faster than existing methods with linear rate. Simulation results demonstrate the advantages of the proposed method.
Zuyuan Yang, Yu Zhang 0009, Yong Xiang 0001, Wei Yan 0009, Shengli Xie 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2019 Community Enhanced Record Linkage Method for Vehicle Insurance System
Christian Lu, Guangyan Huang, Yong Xiang 0001
ADMA3
2019 Generative Adversarial Nets Enhanced Continual Data Release Using Differential Privacy
Stella Ho, Youyang Qu, Longxiang Gao, Jianxin Li 0001, Yong Xiang 0001
ICA3PP (2)5
2019 Context-Aware Privacy Preservation in a Hierarchical Fog Computing System
abstract
Fog computing faces various security and privacy threats. Internet of Things (IoTs) devices have limited computing, storage, and other resources. They are vulnerable to attack by adversaries. Although the existing privacy-preserving solutions in fog computing can be migrated to address some privacy issues, specific privacy challenges still exist because of the unique features of fog computing, such as the decentralized and hierarchical infrastructure, mobility, location and content-aware applications. Unfortunately, privacy-preserving issues and resources in fog computing have not been systematically identified, especially the privacy preservation in multiple fog node communication with end users. In this paper, we propose a dynamic MDP-based privacy-preserving model in zero-sum game to identify the efficiency of the privacy loss and payoff changes to preserve sensitive content in a fog computing environment. First, we develop a new dynamic model with MDP-based comprehensive algorithms. Then, extensive experimental results identify the significance of the proposed model compared with others in more effectively and feasibly solving the discussed issues.
Bruce Gu, Xiaodong Wang 0017, Youyang Qu, Jiong Jin, Yong Xiang 0001, Longxiang Gao
ICC5
2019 When Geo-Text Meets Security: Privacy-Preserving Boolean Spatial Keyword Queries
abstract
In recent years, spatial keyword query has attracted wide-spread research attention due to the popularity of the location-based services. To efficiently support the online spatial keyword query processing, the data owners need to outsource their data and the query processing service to cloud platforms. However, the outsourcing services may raise privacy leaking issues because the cloud server on the platforms may not be trusted for both data owners and query users. Therefore, in this work, we first propose and formalize the problem of privacy-preserving boolean spatial keyword query under the widely accepted Known Background Thread Model. And then, we devise a novel privacy-preserving spatial-textual Bloom Filter encoding structure and an encrypted R-tree index. They can maintain both spatial and text information together in a secure way while answering the encrypted spatial keyword queries without the need for data decryption. To further accelerate the query processing, a compressed encrypted index is provided to deal with the challenges of the large dimension expansion and the expensive space consumption in the encrypted R-tree index. In addition, we develop the corresponding algorithms based on the designed index, and present the in-depth security analysis to show our work's satisfaction meeting the strong secure scheme. Finally, we demonstrate the performance of our proposed index and algorithms by conducting extensive experiments on four datasets under various system settings.
Ningning Cui, Jianxin Li 0001, Xiaochun Yang 0001, Bin Wang 0015, Mark Reynolds 0001, Yong Xiang 0001
ICDE6
2019 Tag-Based Semantic Features for Scene Image Classification
Chiranjibi Sitaula, Yong Xiang 0001, Anish Basnet, Sunil Aryal, Xuequan Lu
ICONIP (3)2
2019 Pre-adjustment Based Anti-collusion Mechanism for Audio Signals
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Gleb Beliakov
NSS3
2019 Unsupervised Deep Features for Privacy Image Classification
Chiranjibi Sitaula, Yong Xiang 0001, Sunil Aryal, Xuequan Lu
PSIVT2
2019 Deep gene selection method to select genes from microarray datasets for cancer classification
abstract
BACKGROUND: Microarray datasets consist of complex and high-dimensional samples and genes, and generally the number of samples is much smaller than the number of genes. Due to this data imbalance, gene selection is a demanding task for microarray expression data analysis. RESULTS: The gene set selected by DGS has shown its superior performances in cancer classification. DGS has a high capability of reducing the number of genes in the original microarray datasets. The experimental comparisons with other representative and state-of-the-art gene selection methods also showed that DGS achieved the best performance in terms of the number of selected genes, classification accuracy, and computational cost. CONCLUSIONS: We provide an efficient gene selection algorithm can select relevant genes which are significantly sensitive to the samples' classes. With the few discriminative genes and less cost time by the proposed algorithm achieved much high prediction accuracy on several public microarray data, which in turn verifies the efficiency and effectiveness of the proposed gene selection method.
Russul Alanni, Jingyu Hou 0001, Hasseeb Azzawi, Yong Xiang 0001
BMC Bioinform.4
2019 Adversaries or allies? Privacy and deep learning in big data era
abstract
Summary Deep learning methods have become the basis of new AI‐based services on the Internet in big data era because of their unprecedented accuracy. Meanwhile, it raises obvious privacy issues. The deep learning–assisted privacy attack can extract sensitive personal information not only from the text but also from unstructured data such as images and videos. In this paper, we proposed a framework to protect image privacy against deep learning tools, along with two new metrics that measure image privacy. Moreover, we propose two different image privacy protection schemes based on the two metrics, utilizing the adversarial example idea. The performance of our solution is validated by simulations on two different datasets. Our research shows that we can protect the image privacy by adding a small amount of noise that has a humanly imperceptible impact on the image quality, especially for images of complex structures and textures.
Bo Liu 0001, Ming Ding 0001, Tianqing Zhu, Yong Xiang 0001, Wanlei Zhou 0001
Concurr. Comput. Pract. Exp.4
2019 Enhanced Smart Meter Privacy Protection Using Rechargeable Batteries
abstract
Due to the rapid growth of smart grids, use of smart meters (SMs) have increased in the recent days. The main problem with the use of SMs is that by observing the SMs reading, it is possible to infer the daily activities of the consumers. Therefore, protection of privacy is a major concern related to SMs. Using rechargeable batteries (RBs) is a popular method in protecting the privacy in SMs as these methods do not tamper with SM readings. The major problem in RB-based mechanism is that the energy management unit (EMU) cannot protect privacy, if the demand is lower or higher for a longer period. To overcome this problem, in this paper a heuristic method has been proposed by considering time varying target output load based on the three major properties of artificial fish swarm optimization algorithm. For the optimal choice of the time varying target output load, RB constraints as well as reduction of the average cost of energy have been considered in our proposed method. We have proposed two privacy preserving mechanisms for both offline and online scenarios. The proposed method preserves privacy while reducing the cost of energy. Simulation results show that the proposed method is able to provide privacy by overcoming the problem identified in the existing methods.
Mohammad Belayet Hossain, Iynkaran Natgunanathan, Yong Xiang 0001, Lu-Xing Yang, Guangyan Huang
IEEE Internet Things J.3
2019 Sustainability Analysis for Fog Nodes With Renewable Energy Supplies
abstract
There is a growing interest in the use of renewable energy sources to power fog networks in order to mitigate the detrimental effects of conventional energy production. However, renewable energy sources, such as solar and wind, are by nature unstable in their availability and capacity. The dynamics of energy supply hence impose new challenges for network planning and resource management. In this paper, the sustainable performance of a fog node powered by renewable energy sources is studied. We develop a generic analytical model to study the energy sustainability of fog nodes powered by renewable energy sources, by generalizing the leaky bucket model to shape and police traffic source for rate-based congestion control in high-speed fog networks. Based on the closed-form solutions of energy buffer analysis, i.e., the energy depletion probability and mean energy length, we study the energy sustainability in two special but real-happening scenarios. The experimental results show that with proper design the leaky bucket model effectively reflects the energy sustainability of data traffic in fog networks. Numerical results also reveal that the model performance is sensitive to certain traffic source characteristics in fog networks.
Jiaojiao Jiang 0001, Longxiang Gao, Jiong Jin, Tom H. Luan, Shui Yu 0001, Yong Xiang 0001, Saurabh Kumar Garg 0001
IEEE Internet Things J.6
2019 Progressive Average-Based Smart Meter Privacy Enhancement Using Rechargeable Batteries
abstract
Usage of smart meters (SMs) have significantly increased in the recent days due to the advantages they offer. However, it is possible for an adversary to extract private information about a consumer by observing the SM readings. Therefore, it is important to protect the privacy of consumers using SMs. Among the SM privacy protection mechanisms, rechargeable battery (RB)-based mechanisms are preferred as they do not alter SM readings. The existing mechanisms cannot protect the privacy when the consumer energy usage is either low or high for a longer period. Furthermore, these mechanisms do not perform well in online scenario where the energy management unit (EMU) only knows the current and past consumer energy demands. To solve this problem, in this article, we proposed a novel online privacy protection mechanism to protect the privacy of SMs using a progressive average-based algorithm (PABA). The proposed PABA uses two uniquely designed algorithms to ensure protection of privacy and energy cost reduction during peak and off-peak periods. Moreover, an adaptive output smoothing technique is used to further enhance the privacy. Compared with the privacy protection mechanisms designed for tackling SM privacy, the proposed mechanism achieves higher amount of privacy while reducing the energy cost. The validity of the proposed privacy enhancement mechanism is demonstrated by simulation results.
Iynkaran Natgunanathan, Mohammad Belayet Hossain, Yong Xiang 0001, Longxiang Gao, Dezhong Peng, Jianxin Li 0001
IEEE Internet Things J.3
2019 Location Privacy Protection in Smart Health Care System
abstract
In a smart health system, patients' location information is periodically sent to hospitals and this information helps hospitals to provide improved health care services. The location information together with time stamp alone can reveal a patient's private information, such as person's life style, places frequently visited by the person, and personal interests. Thus, it is important to protect the location privacy of a patient. In the existing privacy protection mechanisms, trusted third party (TTP) and location perturbation techniques are used. However, in TTP-based mechanism, an adversary who illegally gets access to TTP server will have access to the private location information. On the other hand, in location perturbation technique, utility of the location information is significantly compromised. In this paper, we propose a location privacy protection mechanism in which location privacy is protected while maintaining the utility of the location data. In the proposed mechanism, a main processing unit attached to a patient's body generates the perturbed location by considering the distance between the patient's location and the preidentified patient's sensitive locations. This adaptive generation of perturbed location, removes the necessity to trust other parties while preserving the privacy and utility of the location data. The validity of the proposed mechanism is demonstrated by simulation results.
Iynkaran Natgunanathan, Abid Mehmood, Yong Xiang 0001, Longxiang Gao, Shui Yu 0001
IEEE Internet Things J.3
2019 Efficiently and securely outsourcing compressed sensing reconstruction to a cloud
Yushu Zhang 0001, Yong Xiang 0001, Leo Yu Zhang, Lu-Xing Yang, Jiantao Zhou 0001
Inf. Sci.2
2019 A reliable reputation computation framework for online items in E-commerce
Lichuan Ma, Qingqi Pei, Yong Xiang 0001, Lina Yao 0001, Shui Yu 0001
J. Netw. Comput. Appl.3
2019 Multimodal adversarial network for cross-modal retrieval
Peng Hu 0002, Dezhong Peng, Xu Wang 0028, Yong Xiang 0001
Knowl. Based Syst.4
2019 Random Matching Pursuit for Image Watermarking
abstract
The classical solution to an underdetermined system of linear equations mainly has two opposite directions, which lead to either a large ℓ2-norm sparse solution or a non-sparse minimum ℓ2-norm solution. In this paper, we systematically show that by modifying the well-known basic matching pursuit algorithm originally proposed to identify the sparse solution, an alternative solution between the two classical ones could be obtained. The modified algorithm, termed as random matching pursuit (RMP), is then used to create a novel image watermarking framework. Compared to conventional systems, the security is substantially improved by the use of random over-complete dictionaries and the order parameter of RMP. Capacity can also be increased thanks to the transform with over-complete dictionaries that could expand signal dimension. Meanwhile, imperceptibility and robustness properties of the proposed design framework are not compromised. The classical spread spectrum and improved spread spectrum techniques are applied to the proposed framework for practical implementations. The novelty and effectiveness of the proposed systems are supported by rigorous performance analysis and experimental results using an image data set. This paper reveals the potential of using over-complete dictionaries in multimedia watermarking systems, which theoretically leads to the exploration of alternative candidates among the infinite solutions to underdetermined linear systems other than minimum ℓ2-norm and sparse ones.
Guang Hua 0001, Lifan Zhao, Guoan Bi, Yong Xiang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2019 Effective Repair Strategy Against Advanced Persistent Threat: A Differential Game Approach
abstract
Advanced persistent threat (APT) is a new kind of cyberattack that poses a serious threat to modern society. When an APT campaign on an organization has been identified, the available repair resources must be reasonably allocated to the potentially insecure hosts to mitigate the potential loss of the organization. We refer to the feasible repair resource allocation strategies as repair strategies. This paper focuses on the APT repair problem, i.e., the problem of developing effective repair strategies for organizations. First, for an organization with time-varying communication relationship, we establish an evolution model of the organization's expected state, in which the impact of lateral movement of APT is accommodated. On this basis, we model the APT repair problem as a differential Nash game problem (the APT repair game) in which the attacker attempts to maximize his potential benefit, and the organization manages to minimize its potential loss. Second, we derive a system (the potential system) for calculating a potential Nash equilibrium of an APT repair game, and we examine the structure of the potential attack and repair strategies in a potential Nash equilibrium. Next, we solve some potential systems to get the corresponding potential Nash equilibria. Finally, by comparison with a large number of randomly generated attack and repair strategies, we conclude that the potential Nash equilibrium of each APT repair game is a Nash equilibrium of the game. Therefore, we recommend to organizations their respective potential repair strategies. Our findings help to better understand and effectively defend against APT.
Lu-Xing Yang, Pengdeng Li, Yushu Zhang 0001, Xiaofan Yang 0001, Yong Xiang 0001, Wanlei Zhou 0001
IEEE Trans. Inf. Forensics Secur.5
2019 Compressed Sensing Based Selective Encryption With Data Hiding Capability
abstract
This paper proposes a joint selective encryption and data hiding scheme based on compressed sensing (CS), with a focus to its application in secure imaging. Specifically, working with a semantic-secure stream cipher, we suggest to selectively encrypt the sign bits of the CS measurements during its quantization stage and insert the authentication information using a nonseparable histogram-shifting based data hiding scheme. The rationale behind the sign encryption is that CS measurements, when measured by random subspace projection, is random in nature and thus, from both theoretical and experimental points of view, the mean squared errors associated with authorized users and attackers are significant. Due to the indistinguishability of the output ciphertext and the nonlinearity of the CS decoder, it is robust, when comparing with the existing selective encryption system of multimedia data, against known error concealment attacks. When applied in imaging, we demonstrate that it could effectively degrade the visual quality level while saving the computation load by at least 90%. Moreover, we further show that a state-of-the-art data hiding system can be seamlessly incorporated into the sign encryption, thus allowing soft data authentication without heavy computation. The proposed scheme is expected to strengthen the security of applications in the field where both energy and privacy are the concerns, such as sensitive information protection for multimedia data in wireless sensor networks.
Jia Wang 0008, Leo Yu Zhang, Junxin Chen 0001, Guang Hua 0001, Yushu Zhang 0001, Yong Xiang 0001
IEEE Trans. Ind. Informatics6
2019 Multi-View Linear Discriminant Analysis Network
abstract
In many real-world applications, an object can be described from multiple views or styles, leading to the emerging multi-view analysis. To eliminate the complicated (usually highly nonlinear) view discrepancy for favorable cross-view recognition and retrieval, we propose a Multi-view Linear Discriminant Analysis Network (MvLDAN) by seeking a nonlinear discriminant and view-invariant representation shared among multiple views. Unlike existing multi-view methods which directly learn a common space to reduce the view gap, our MvLDAN employs multiple feedforward neural networks (one for each view) and a novel eigenvalue-based multi-view objective function to encapsulate as much discriminative variance as possible into all the available common feature dimensions. With the proposed objective function, the MvLDAN could produce representations possessing: 1) low variance within the same class regardless of view discrepancy, 2) high variance between different classes regardless of view discrepancy, and 3) high covariance between any two views. In brief, in the learned multi-view space, the obtained deep features can be projected into a latent common space in which the samples from the same class are as close to each other as possible (even though they are from different views), and the samples from different classes are as far from each other as possible (even though they are from the same view). The effectiveness of the proposed method is verified by extensive experiments carried out on five databases, in comparison with the 19 state-of-the-art approaches.
Peng Hu 0002, Dezhong Peng, Yongsheng Sang, Yong Xiang 0001
IEEE Trans. Image Process.4
2019 Improving the Visual Quality of Size-Invariant Visual Cryptography for Grayscale Images: An Analysis-by-Synthesis (AbS) Approach
abstract
In visual cryptography (VC) for grayscale image, size reduction leads to bad perceptual quality to the reconstructed secret image. To improve the quality, the current efforts are limited to the design of VC algorithm for binary image, and measuring the quality with metrics that are not directly related to how the human visual system (HVS) perceives halftone images. We propose an analysis-by-synthesis (AbS) framework to integrate the halftoning process and the VC encoding: the secret pixel/block is reconstructed from the shares in the encoder and the error between the reconstructed secret and the original secret images is fed back and compensated concurrently by the error diffusion process. In doing so, the error between the reconstructed secret and original secret is pushed to high frequency band, thus producing visually pleasing reconstructed secret image. This framework is simple and flexible in that it can be combined with many existing size-invariant VC algorithms, including probabilistic VC, random grid VC and vector/block VC. More importantly, it is proved that this AbS framework is as secure as the traditional VC algorithms. Experimental results demonstrate the effectiveness of the proposed AbS framework.
Bin Yan 0001, Yong Xiang 0001, Guang Hua 0001
IEEE Trans. Image Process.2
2019 Non-Local Texture Optimization With Wasserstein Regularization Under Convolutional Neural Network
abstract
Example-based texture synthesis aims to generate a new texture from an exemplar texture and has long been drawing attention in the fields of computer graphics, computer vision, and image processing. Nevertheless, synthesizing structured textures remains a challenging task. Most previous methods rely on additional guidance channels, which encode the structured features of textures. However, estimating the guidance channel is very difficult, and often fails when a texture has unpronounced features. In this paper, we propose a novel texture synthesis method, based on non-local operators, which captures the long-range structure of a texture without the additional guidance channel. The synthesized texture is generated by minimizing non-local texture energy through an expectation-maximization like optimization algorithm. A statistical constraint based on the Wasserstein distance is also proposed to ensure that the synthesized texture preserves the global statistics of the exemplar texture. Extensive experiments show that the proposed method can stably handle textures with different scale structures.
Jie Li 0023, Yong Xiang 0001, Jingyu Hou 0001, Dan Xu 0001
IEEE Trans. Multim.2
2019 Privacy-Preserving Reputation Management for Edge Computing Enhanced Mobile Crowdsensing
abstract
Mobile crowdsensing (MCS) has gained popularity for its potential to leverage individual mobile devices to sense, collect, and analyze data instead of deploying sensors. As the sensing data become increasingly fine-grained and complicated, there is a tendency to enhance MCS with the edge computing paradigm to reduce time delays and high bandwidth costs. The sensing data may reveal personal information, and thus it is of great significance to preserve the privacy of the participants. However, preserving privacy may hinder the process of handling malicious participants. In this paper, we propose two privacy preserving reputation management schemes for edge computing enhanced MCS to simultaneously preserve privacy and deal with malicious participants. In the basic scheme, a novel reputation value updating method is designed based on the deviations of the encrypted sensing data from the final aggregating result. The basic scheme is efficient at the expense of revealing the deviation value of each participant to the reputation manager. To conquer this drawback, we propose an advanced scheme by updating the reputation values utilizing the rank of deviations. Extensive experiments demonstrate that both these two schemes have high cost efficiency and are effective to deal with malicious participants.
Lichuan Ma, Xuefeng Liu 0002, Qingqi Pei, Yong Xiang 0001
IEEE Trans. Serv. Comput.4
2018 SBC: A New Strategy for Multiclass Lung Cancer Classification Based on Tumour Structural Information and Microarray Data
abstract
Lung cancer has different subtypes which are different in cell size and growth pattern. Correctly classifying subtypes of lung cancer can help design specific treatments to increase patient survival rate. In this work, we propose an innovative Structural Binary Classification (SBC) strategy for classifying lung cancer subtypes using microarray data. The strategy is based on Gene Expression Programming (GEP) algorithm. Classification performance evaluations and comparisons between our GEP based model and common binary decomposition strategies, as well as three representative machine learning methods, support vector machine, neural network and C4.5, were conducted thoroughly on real microarray lung cancer datasets. Reliability was assessed by the cross-data set validation. The experimental results showed that GEP model with our strategy outperformed other models in terms of accuracy, standard deviation and area under the receiver operating characteristic curve. The work provides a useful tool for lung cancer classification based on tumour structural information.
Hasseeb Azzawi, Jingyu Hou 0001, Russul Alanni, Yong Xiang 0001
ICIS4
2018 Using Adversarial Noises to Protect Privacy in Deep Learning Era
abstract
The unprecedented accuracy of deep learning methods has earned themselves as the foundation of new AI-based services on the Internet. At the same time, it presents obvious privacy issues. The deep learning aided privacy attack can dig out sensitive personal information not only from the text but also from unstructured data such as images and videos. In this paper, we proposed a framework to protect image privacy against the deep learning tools. We also propose two new metrics to measure the image privacy. Moreover, we propose two different image privacy protection schemes based on the two metrics, utilizing the adversarial example idea. The performance of our schemes is validated by simulation on a large-scale dataset. Our study shows that we can protect the image privacy by adding a small amount of noise, while the added noise has a humanly imperceptible impact on the image quality.
Bo Liu 0001, Ming Ding 0001, Tianqing Zhu, Yong Xiang 0001, Wanlei Zhou 0001
GLOBECOM4
2018 SPARSE: Privacy-Aware and Collusion Resistant Location Proof Generation and Verification
abstract
Recently, there has been an increase in the number of location-based services and applications. It is common for these applications to provide facilities or rewards for users who visit specific venues frequently. This creates the incentive for dishonest users to lie about their location and submit fake check-ins by changing their GPS data. To solve this issue, different distributed location proof schemes have been proposed to generate location proofs for mobile users. However, these schemes have some drawbacks: (1) they are vulnerable to either Prover-Prover or Prover-Witness collusions, (2) the location proof generation process is slow when users adopt a long private key, and (3) their implementation requires some hardware changes on mobile devices. To address these issues, we propose the Secure, Privacy-Aware and collusion Resistant poSition vErification (SPARSE) scheme to generate private location proofs for mobile users. SPARSE has a distributed architecture designed for ad-hoc scenarios in which mobile users generate location proofs for each other. Since we do not integrate any distance bounding protocol into SPARSE, it becomes an easy-to-implement scheme in which the location proof generation process is independent of the length of the users' private key. We provide a comprehensive security analysis and simulation which show that SPARSE provides privacy protection as well as security properties for users including integrity, unforgeability and non-transferability of the location proofs. Moreover, it achieves a highly reliable performance against collusions.
Mohammad Reza Nosouhi, Shui Yu 0001, Marthie Grobler, Yong Xiang 0001, Zuqing Zhu
GLOBECOM4
2018 Clustering of Multiple Density Peaks
Borui Cai, Guangyan Huang, Yong Xiang 0001, Jing He 0004, Guang-Li Huang, Xiangmin Zhou
PAKDD (3)3
2018 Low-Cost and Confidentiality-Preserving Data Acquisition for Internet of Multimedia Things
abstract
Internet of Multimedia Things (IoMT) faces the challenge of how to realize low-cost data acquisition while still preserve data confidentiality. In this paper, we present a low-cost and confidentiality-preserving data acquisition framework for IoMT. First, we harness chaotic convolution and random subsampling to capture multiple image signals. The measurement matrix is under the control of chaos, ensuring the security of the sampling process. Next, we assemble these sampled images into a big master image, and then encrypt this master image based on Arnold transform and single value diffusion. The computation of these two transforms only requires some low-complexity operations. Finally, the encrypted image is delivered to cloud servers for storage and decryption service. Experimental results demonstrate the security and effectiveness of the proposed framework.
Yushu Zhang 0001, Yong Xiang 0001, Leo Yu Zhang, Bo Liu 0001, Junxin Chen 0001, Yiyuan Xie
IEEE Internet Things J.3
2018 An overview of protection of privacy in multibiometrics
Iynkaran Natgunanathan, Abid Mehmood, Yong Xiang 0001, Guang Hua 0001, Gang Li 0009, Shaun Bangay
Multim. Tools Appl.3
2018 A compression-diffusion-permutation strategy for securing image
Hui Huang 0008, Xing He 0001, Yong Xiang 0001, Wenying Wen, Yushu Zhang 0001
Signal Process.3
2018 Spread Spectrum Audio Watermarking Using Multiple Orthogonal PN Sequences and Variable Embedding Strengths and Polarities
abstract
Copyright protection of audio data is a serious problem and spread spectrum (SS) based audio watermarking is a promising technology to tackle this problem. Although a number of SS-based audio watermarking methods have been reported in the literature, they cannot achieve high robustness and embedding capacity at the same time. In this paper, we propose a novel SS-based audio watermarking method that can embed a large number of watermark bits into an audio signal without compromising the robustness against common attacks. Compared with the existing audio watermarking methods, the proposed one is especially robust against severe noise addition and compression attacks, while achieving high embedding capacity. Moreover, the new audio watermarking method is computationally efficient. The validity of the proposed SS-based audio watermarking method is demonstrated by simulation results.
Yong Xiang 0001, Iynkaran Natgunanathan, Dezhong Peng, Guang Hua 0001, Bo Liu 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Multiclass Lung Cancer Diagnosis by Gene Expression Programming and Microarray Datasets
Hasseeb Azzawi, Jingyu Hou 0001, Russul Alanni, Yong Xiang 0001, Rana Abdu-Aljabar, Ali Azzawi
ADMA4
2017 A Hybrid Location Privacy Protection Scheme in Big Data Environment
abstract
Location privacy has become a significant challenge of big data. Particularly, by the advantage of big data handling tools availability, huge location data can be managed and processed easily by an adversary to obtain user private information from Location-Based Services (LBS). So far, many methods have been proposed to preserve user location privacy for these services. Among them, dummy-based methods have various advantages in terms of implementation and low computation costs. However, they suffer from the spatiotemporal correlation issue when users submit consecutive requests. To solve this problem, a practical hybrid location privacy protection scheme is presented in this paper. The proposed method filters out the correlated fake location data (dummies) before submissions. Therefore, the adversary can not identify the user's real location. Evaluations and experiments show that our proposed filtering technique significantly improves the performance of existing dummy-based methods and enables them to effectively protect the user's location privacy in the environment of big data.
Mohammad Reza Nosouhi, Vu Viet Hoang Pham, Shui Yu 0001, Yong Xiang 0001, Matthew J. Warren
GLOBECOM4
2017 Home Location Protection in Mobile Social Networks: A Community Based Method (Short Paper)
Bo Liu 0001, Wanlei Zhou 0001, Shui Yu 0001, Kun Wang 0005, Yu Wang 0017, Yong Xiang 0001, Jin Li 0002
ISPEC6
2017 Towards an Analysis of Traffic Shaping and Policing in Fog Networks Using Stochastic Fluid Models
abstract
This paper gives models and analytic techniques for studying shaping and policing data traffic in fog networks. The traffic in these networks is expected to be highly diverse and bursty, and regulation will be required as an integral part of congestion control. We generalize the Leaky Bucket model to shape and police traffic source for rate-based congestion control in high-speed fog networks. In particular, the Markov modulated fluid sources reflect the bursty characteristics of data traffic. To measure the performance of the model in shaping and policing traffic, we derive four performance metrics. The experimental results show that with proper design the Leaky Bucket model effectively controls a 4-way trade-off between throughput, loss probability, delay and burstiness of data traffic. Numerical results also reveal that the model performance is sensitive to certain traffic source characteristics.
Jiaojiao Jiang 0001, Longxiang Gao, Jiong Jin, Tom H. Luan, Shui Yu 0001, Dong Yuan 0001, Yong Xiang 0001, Dongfeng Yuan
MobiQuitous7
2017 Harnessing the Hybrid Cloud for Secure Big Image Data Service
abstract
Various kinds of image sensors capture a large number of images in Internet of Things (IoT) every day. It is increasingly concerned how to securely store and share these big image data from IoT. In this paper, we harness the hybrid cloud to provide secure big image data storage and share service for users. The basic idea is to partition each image into a small set of sensitive data and a large set of insensitive data, which are securely stored in the private cloud and the public cloud, respectively. Specially, the private cloud divides each image into the sensitive data (80%) based on sensitivity identification approaches like Sobel edge detector. The sensitive data are encrypted in parallel at a counter mode and then stored in the private cloud. The insensitive data are encrypted-then-subsampled and then placed in the public cloud, in which the encryption employs the permutation-diffusion architecture and the subsampling utilizes compressed sampling technique. The keystreams used in encryption operations are managed by the tent-logistic system with high initial value sensitivity. Once users make a request for an image, the public cloud provides a privacy-guaranteed insensitive data reconstruction service, and the private cloud decrypts the sensitive and insensitive data and regroups them into a complete image. Experimental results demonstrate that the proposed framework can provide secure big image data service.
Yushu Zhang 0001, Hui Huang 0008, Yong Xiang 0001, Leo Yu Zhang, Xing He 0001
IEEE Internet Things J.3
2017 Patchwork-Based Multilayer Audio Watermarking
abstract
A multilayer watermarking system is a system that is able to embed watermarks to a host media signal repeatedly in an overlaying manner, without incurring troubles in extracting the watermarks in each layer. In this paper, we present a novel patchwork-based audio watermarking algorithm that can embed and extract watermark bits successfully in such a multilayer framework. In the proposed method, a new watermark embedding algorithm is designed to ensure that the embedded watermarks in a certain layer do not affect the detection of watermarks in other layers. Adding multiple layers of watermark bits inevitably reduces the perceptual quality. However, to minimize the perceptual quality degradation in multilayer watermarking, the audio fragments for watermark embedding are selected from a set of specially arranged discrete cosine transform coefficients of the host audio signal. Watermark embedding is achieved by modifying the mean values of selected sample fragments. With the use of an embedding error buffer, the proposed system can withstand a wide range of common attacks. To maintain the balance between the perceptual quality and robustness, watermark embedding strength is adjusted according to the specific layer used. The proposed multilayer scheme ensures the independence of the processing in different layers. The effectiveness of the proposed system is demonstrated and verified by extensive simulation results.
Iynkaran Natgunanathan, Yong Xiang 0001, Guang Hua 0001, Gleb Beliakov, John Yearwood
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Adaptive Method for Nonsmooth Nonnegative Matrix Factorization
abstract
Nonnegative matrix factorization (NMF) is an emerging tool for meaningful low-rank matrix representation. In NMF, explicit constraints are usually required, such that NMF generates desired products (or factorizations), especially when the products have significant sparseness features. It is known that the ability of NMF in learning sparse representation can be improved by embedding a smoothness factor between the products. Motivated by this result, we propose an adaptive nonsmooth NMF (Ans-NMF) method in this paper. In our method, the embedded factor is obtained by using a data-related approach, so it matches well with the underlying products, implying a superior faithfulness of the representations. Besides, due to the usage of an adaptive selection scheme to this factor, the sparseness of the products can be separately constrained, leading to wider applicability and interpretability. Furthermore, since the adaptive selection scheme is processed through solving a series of typical linear programming problems, it can be easily implemented. Simulations using computer-generated data and real-world data show the advantages of the proposed Ans-NMF method over the state-of-the-art methods.
Zuyuan Yang, Yong Xiang 0001, Kan Xie 0002, Yue Lai
IEEE Trans. Neural Networks Learn. Syst.2
2017 Underdetermined Blind Source Separation Using Sparse Coding
abstract
In an underdetermined mixture system with unknown sources, it is a challenging task to separate these sources from their observed mixture signals, where . By exploiting the technique of sparse coding, we propose an effective approach to discover some 1-D subspaces from the set consisting of all the time-frequency (TF) representation vectors of observed mixture signals. We show that these 1-D subspaces are associated with TF points where only single source possesses dominant energy. By grouping the vectors in these subspaces via hierarchical clustering algorithm, we obtain the estimation of the mixing matrix. Finally, the source signals could be recovered by solving a series of least squares problems. Since the sparse coding strategy considers the linear representation relations among all the TF representation vectors of mixing signals, the proposed algorithm can provide an accurate estimation of the mixing matrix and is robust to the noises compared with the existing underdetermined blind source separation approaches. Theoretical analysis and experimental results demonstrate the effectiveness of the proposed method.
Liangli Zhen, Dezhong Peng, Zhang Yi 0001, Yong Xiang 0001, Peng Chen 0007
IEEE Trans. Neural Networks Learn. Syst.4
2016 Software-Defined Wireless Networking Opportunities and Challenges for Internet-of-Things: A Review
abstract
With the emergence of Internet-of-Things (IoT), there is now growing interest to simplify wireless network controls. This is a very challenging task, comprising information acquisition, information analysis, decision-making, and action implementation on large scale IoT networks. Resulting in research to explore the integration of software-defined networking (SDN) and IoT for a simpler, easier, and strain less network control. SDN is a promising novel paradigm shift which has the capability to enable a simplified and robust programmable wireless network serving an array of physical objects and applications. This paper starts with the emergence of SDN and then highlights recent significant developments in the wireless and optical domains with the aim of integrating SDN and IoT. Challenges in SDN and IoT integration are also discussed from both security and scalability perspectives.
Keshav Sood, Shui Yu 0001, Yong Xiang 0001
IEEE Internet Things J.3
2016 When Compressive Sensing Meets Data Hiding
abstract
We present a novel framework of performing multimedia data hiding using an over-complete dictionary, which brings compressive sensing to the application of data hiding. Unlike the conventional orthonormal full-space dictionary, the over-complete dictionary produces an underdetermined system with infinite transform results. We first discuss the minimum norm formulation (ℓ2-norm) which yields a closed-form solution and the concept of watermark projection, so that higher embedding capacity and an additional privacy preserving feature can be obtained. Furthermore, we study the sparse formulation (ℓ2-norm) and illustrate that as long as the ℓ0-norm of the sparse representation of the host signal is less than the signal's dimension in the original domain, an informed sparse domain data hiding system can be established by modifying the coefficients of the atoms that have not participated in representing the host signal. A single support modification-based data hiding system is then proposed and analyzed as an example. Several potential research directions are discussed for further studies. More generally, apart from the ℓ2- and ℓ0-norm constraints, other conditions for reliable detection performance are worth of future investigation.
Guang Hua 0001, Yong Xiang 0001, Guoan Bi
IEEE Signal Process. Lett.2
2016 A General Communication Cost Optimization Framework for Big Data Stream Processing in Geo-Distributed Data Centers
abstract
With the explosion of big data, processing large numbers of continuous data streams, i.e., big data stream processing (BDSP), has become a crucial requirement for many scientific and industrial applications in recent years. By offering a pool of computation, communication and storage resources, public clouds, like Amazon's EC2, are undoubtedly the most efficient platforms to meet the ever-growing needs of BDSP. Public cloud service providers usually operate a number of geo-distributed datacenters across the globe. Different datacenter pairs are with different inter-datacenter network costs charged by Internet Service Providers (ISPs). While, inter-datacenter traffic in BDSP constitutes a large portion of a cloud provider's traffic demand over the Internet and incurs substantial communication cost, which may even become the dominant operational expenditure factor. As the datacenter resources are provided in a virtualized way, the virtual machines (VMs) for stream processing tasks can be freely deployed onto any datacenters, provided that the Service Level Agreement (SLA, e.g., quality-of-information) is obeyed. This raises the opportunity, but also a challenge, to explore the inter-datacenter network cost diversities to optimize both VM placement and load balancing towards network cost minimization with guaranteed SLA. In this paper, we first propose a general modeling framework that describes all representative inter-task relationship semantics in BDSP. Based on our novel framework, we then formulate the communication cost minimization problem for BDSP into a mixed-integer linear programming (MILP) problem and prove it to be NP-hard. We then propose a computation-efficient solution based on MILP. The high efficiency of our proposal is validated by extensive simulation based studies.
Lin Gu 0002, Deze Zeng, Song Guo 0001, Yong Xiang 0001, Jiankun Hu
IEEE Trans. Computers4
2015 Spread Spectrum-Based High Embedding Capacity Watermarking Method for Audio Signals
abstract
Audio watermarking is a promising technology for copyright protection of audio data. Built upon the concept of spread spectrum (SS), many SS-based audio watermarking methods have been developed, where a pseudonoise (PN) sequence is usually used to introduce security. A major drawback of the existing SS-based audio watermarking methods is their low embedding capacity. In this paper, we propose a new SS-based audio watermarking method which possesses much higher embedding capacity while ensuring satisfactory imperceptibility and robustness. The high embedding capacity is achieved through a set of mechanisms: embedding multiple watermark bits in one audio segment, reducing host signal interference on watermark extraction, and adaptively adjusting PN sequence amplitude in watermark embedding based on the property of audio segments. The effectiveness of the proposed audio watermarking method is demonstrated by simulation examples.
Yong Xiang 0001, Iynkaran Natgunanathan, Yue Rong, Song Guo 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Energy Minimization in Multi-Task Software-Defined Sensor Networks
abstract
After a decade of extensive research on application-specific wireless sensor networks (WSNs), the recent development of information and communication technologies makes it practical to realize the software-defined sensor networks (SDSNs), which are able to adapt to various application requirements and to fully explore the resources of WSNs. A sensor node in SDSN is able to conduct multiple tasks with different sensing targets simultaneously. A given sensing task usually involves multiple sensors to achieve a certain quality-of-sensing, e.g., coverage ratio. It is significant to design an energy-efficient sensor scheduling and management strategy with guaranteed quality-of-sensing for all tasks. To this end, three issues are investigated in this paper: 1) the subset of sensor nodes that shall be activated, i.e., sensor activation, 2) the task that each sensor node shall be assigned, i.e., task mapping, and 3) the sampling rate on a sensor for a target, i.e., sensing scheduling. They are jointly considered and formulated as a mixed-integer with quadratic constraints programming (MIQP) problem, which is then reformulated into a mixed-integer linear programming (MILP) formulation with low computation complexity via linearization. To deal with dynamic events such as sensor node participation and departure, during SDSN operations, an efficient online algorithm using local optimization is developed. Simulation results show that our proposed online algorithm approaches the globally optimized network energy efficiency with much lower rescheduling time and control overhead.
Deze Zeng, Peng Li 0017, Song Guo 0001, Toshiaki Miyazaki, Jiankun Hu, Yong Xiang 0001
IEEE Trans. Computers6
2015 Robust Histogram Shape-Based Method for Image Watermarking
abstract
Cropping and random bending are two common attacks in image watermarking. In this paper we propose a novel image-watermarking method to deal with these attacks, as well as other common attacks. In the embedding process, we first preprocess the host image by a Gaussian low-pass filter. Then, a secret key is used to randomly select a number of gray levels and the histogram of the filtered image with respect to these selected gray levels is constructed. After that, a histogram-shape-related index is introduced to choose the pixel groups with the highest number of pixels and a safe band is built between the chosen and nonchosen pixel groups. A watermark-embedding scheme is proposed to insert watermarks into the chosen pixel groups. The usage of the histogram-shape-related index and safe band results in good robustness. Moreover, a novel high-frequency component modification mechanism is also utilized in the embedding scheme to further improve robustness. At the decoding end, based on the available secret key, the watermarked pixel groups are identified and watermarks are extracted from them. The effectiveness of the proposed image-watermarking method is demonstrated by simulation examples.
Tianrui Zong, Yong Xiang 0001, Iynkaran Natgunanathan, Song Guo 0001, Wanlei Zhou 0001, Gleb Beliakov
IEEE Trans. Circuits Syst. Video Technol.2
2015 A Convex Geometry-Based Blind Source Separation Method for Separating Nonnegative Sources
abstract
This paper presents a convex geometry (CG)-based method for blind separation of nonnegative sources. First, the unaccessible source matrix is normalized to be column-sum-to-one by mapping the available observation matrix. Then, its zero-samples are found by searching the facets of the convex hull spanned by the mapped observations. Considering these zero-samples, a quadratic cost function with respect to each row of the unmixing matrix, together with a linear constraint in relation to the involved variables, is proposed. Upon which, an algorithm is presented to estimate the unmixing matrix by solving a classical convex optimization problem. Unlike the traditional blind source separation (BSS) methods, the CG-based method does not require the independence assumption, nor the uncorrelation assumption. Compared with the BSS methods that are specifically designed to distinguish between nonnegative sources, the proposed method requires a weaker sparsity condition. Provided simulation results illustrate the performance of our method.
Zuyuan Yang, Yong Xiang 0001, Yue Rong, Kan Xie 0002
IEEE Trans. Neural Networks Learn. Syst.2
2015 Channel Estimation for Two-Way MIMO Relay Systems in Frequency-Selective Fading Environments
abstract
In this paper, we investigate the channel estimation problem for two-way multiple-input multiple-output (MIMO) relay communication systems in frequency-selective fading environments. We apply the method of superimposed channel training to estimate the individual channel state information (CSI) of the first-hop and second-hop links for two-way MIMO relay systems with frequency-selective fading channels. In this algorithm, a relay training sequence is superimposed on the received signals at the relay node to assist the estimation of the second-hop channel matrices. The optimal structure of the source and relay training sequences is derived to minimize the mean-squared error (MSE) of channel estimation. Moreover, the optimal power allocation between the source and relay training sequences is derived to improve the performance of channel estimation. Numerical examples are shown to demonstrate the performance of the proposed superimposed channel training algorithm for two-way MIMO relay systems in frequency-selective fading environments.
Choo W. R. Chiong, Yue Rong, Yong Xiang 0001
IEEE Trans. Wirel. Commun.3
2015 Channel Estimation for Time-Varying MIMO Relay Systems
abstract
In this paper, we investigate the channel estimation problem for multiple-input multiple-output (MIMO) relay communication systems with time-varying channels. The time-varying characteristic of the channels is described by the complex-exponential basis expansion model (CE-BEM). We propose a superimposed channel training algorithm to estimate the individual first-hop and second-hop time-varying channel matrices for MIMO relay systems. In particular, the estimation of the second-hop time-varying channel matrix is performed by exploiting the superimposed training sequence at the relay node, while the first-hop time-varying channel matrix is estimated through the source node training sequence and the estimated second-hop channel. To improve the performance of channel estimation, we derive the optimal structure of the source and relay training sequences that minimize the mean-squared error (MSE) of channel estimation. We also optimize the relay amplification factor that governs the power allocation between the source and relay training sequences. Numerical simulations demonstrate that the proposed superimposed channel training algorithm for MIMO relay systems with time-varying channels outperforms the conventional two-stage channel estimation scheme.
Choo W. R. Chiong, Yue Rong, Yong Xiang 0001
IEEE Trans. Wirel. Commun.3
2014 A new interpolation error expansion based reversible watermarking algorithm considering the human visual system
abstract
Reversible watermarking has merged over the past few years as a promising solution for copyright protection, especially for applications like remote sensing, medical imaging and military applications which require lossless recovery of the host media. In this paper, we aim to extend the additive interpolation error expansion technique in [16]. We will consider the human visual system (HVS) to improve the embedding rate while maintaining the image visual quality. To this end, the just noticeable difference (JND) is used to embed more watermark bits. The experimental results show that the proposed algorithm can improve the embedding rate while preserving the image visual quality.
Suzan Elbadry, Yong Xiang 0001, Tianrui Zong, Iynkaran Natgunanathan
ICC2
2014 Robustness enhancement of quantization based audio watermarking method using adaptive safe-band
abstract
This paper presents a novel adaptive safe-band for quantization based audio watermarking methods, aiming to improve robustness. Considerable number of audio watermarking methods have been developed using quantization based techniques. These techniques are generally vulnerable to signal processing attacks. For these conventional quantization based techniques, robustness can be marginally improved by choosing larger step sizes at the cost of significant perceptual quality degradation. We first introduce fixed size safe-band between two quantization steps to improve robustness. This safe-band will act as a buffer to withstand certain types of attacks. Then we further improve the robustness by adaptively changing the size of the safe-band based on the audio signal feature used for watermarking. Compared with conventional quantization based method and the fixed size safe-band based method, the proposed adaptive safe-band based quantization method is more robust to attacks. The effectiveness of the proposed technique is demonstrated by simulation results.
Iynkaran Natgunanathan, Yong Xiang 0001, Tianrui Zong, Yang Xiang 0001
ICC2
2014 Histogram shape-based robust image watermarking method
abstract
Developing a watermarking method that is robust to cropping attack and random bending attacks (RBAs) is a challenging task in image watermarking. In this paper, we propose a histogram-based image watermarking method to tackle with both cropping attack and RBAs. In this method first the gray levels are divided into groups. Secondly the groups for watermark embedding are selected according to the number of pixels in them, which makes this method fully based on the histogram shape of the original image and adaptive to different images. Then the watermark bits are embedded by modifying the histogram of the selected groups. Since histogram shape is insensitive to cropping and independent from pixel positions, the proposed method is robust to cropping attack and RBAs. Besides, it also has high robustness against other common attacks. Experimental results demonstrate the effectiveness of the proposed method.
Tianrui Zong, Yong Xiang 0001, Iynkaran Natgunanathan
ICC2
2014 Image encryption based on compressed sensing and blind source separation
abstract
A novel image encryption scheme based on compressed sensing and blind source separation is proposed in this work, where there is no statistical requirement to plaintexts. In the proposed method, for encryption, the plaintexts and keys are mixed with each other using a underdetermined matrix first, and then compressed under a project matrix. As a result, it forms a difficult underdetermined blind source separation (UBSS) problem without statistical features of sources. Regarding the decryption, given the keys, a new model will be constructed, which is solvable under compressed sensing (CS) frame. Due to the usage of CS technology, the plaintexts are compressed into the data with smaller size when they are encrypted. Meanwhile, they can be decrypted from parts of the received data packets and thus allows to lose some packets. This is beneficial for the proposed encryption method to suit practical communication systems. Simulations are given to illustrate the availability and the superiority of our method.
Zuyuan Yang, Yong Xiang 0001, Chuan Lu
IJCNN2
2014 Channel estimation for frequency-selective two-way MIMO relay systems
Choo W. R. Chiong, Yue Rong, Yong Xiang 0001
ISITA3
2014 Robust patchwork-based watermarking method for stereo audio signals
Iynkaran Natgunanathan, Yong Xiang 0001, Yue Rong, Dezhong Peng
Multim. Tools Appl.2
2014 Patchwork-Based Audio Watermarking Method Robust to De-synchronization Attacks
abstract
This paper presents a patchwork-based audio watermarking method to resist de-synchronization attacks such as pitch-scaling, time-scaling, and jitter attacks. At the embedding stage, the watermarks are embedded into the host audio signal in the discrete cosine transform (DCT) domain. Then, a set of synchronization bits are implanted into the watermarked signal in the logarithmic DCT (LDCT) domain. At the decoding stage, we analyze the received audio signal in the LDCT domain to find the scaling factor imposed by an attack. Then, we modify the received signal to remove the scaling effect, together with the embedded synchronization bits. After that, watermarks are extracted from the modified signal. Simulation results show that at the embedding rate of 10 bps, the proposed method achieves 98.9% detection rate on average under the considered de-synchronization attacks. At the embedding rate of 16 bps, it can still obtain 94.7% detection rate on average. So, the proposed method is much more robust to de-synchronization attacks than other patchwork watermarking methods. Compared with the audio watermarking methods designed for tackling de-synchronization attacks, our method has much higher embedding capacity.
Yong Xiang 0001, Iynkaran Natgunanathan, Song Guo 0001, Wanlei Zhou 0001, Saeid Nahavandi
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 On the Multicast Lifetime of WANETs with Multibeam Antennas: Formulation, Algorithms, and Analysis
abstract
We explore the multicast lifetime capacity of energy-limited wireless ad hoc networks using directional multibeam antennas by formulating and solving the corresponding optimization problem. In such networks, each node is equipped with a practical smart antenna array that can be configured to support multiple beams with adjustable orientation and beamwidth. The special case of this optimization problem in networks with single beams have been extensively studied and shown to be NP-hard. In this paper, we provide a globally optimal solution to this problem by developing a general MILP formulation that can apply to various configurable antenna models, many of which are not supported by the existing formulations. In order to study the multicast lifetime capacity of large-scale networks, we also propose an efficient heuristic algorithm with guaranteed theoretical performance. In particular, we provide a sufficient condition to determine if its performance reaches optimum based on the analysis of its approximation ratio. These results are validated by experiments as well. The multicast lifetime capacity is then quantitatively studied by evaluating the proposed exact and heuristic algorithms using simulations. The experimental results also show that using two-beam antennas can exploit most lifetime capacity of the networks for multicast communications.
Song Guo 0001, Minyi Guo, Victor C. M. Leung, Shui Yu 0001, Yong Xiang 0001
IEEE Trans. Computers5
2014 On the Throughput of Two-Way Relay Networks Using Network Coding
abstract
Network coding has shown the promise of significant throughput improvement. In this paper, we study the network throughput using network coding and explore how the maximum throughput can be achieved in a two-way relay wireless network. Unlike previous studies, we consider a more general network with arbitrary structure of overhearing status between receivers and transmitters. To efficiently utilize the coding opportunities, we invent the concept of network coding cliques (NCCs), upon which a formal analysis on the network throughput using network coding is elaborated. In particular, we derive the closed-form expression of the network throughput under certain traffic load in a slotted ALOHA network with basic medium access control. Furthermore, the maximum throughput as well as optimal medium access probability at each node is studied under various network settings. Our theoretical findings have been validated by simulation as well.
Deze Zeng, Song Guo 0001, Yong Xiang 0001, Hai Jin 0001
IEEE Trans. Parallel Distributed Syst.3
2013 Internet Traffic Classification by Aggregating Correlated Naive Bayes Predictions
abstract
This paper presents a novel traffic classification scheme to improve classification performance when few training data are available. In the proposed scheme, traffic flows are described using the discretized statistical features and flow correlation information is modeled by bag-of-flow (BoF). We solve the BoF-based traffic classification in a classifier combination framework and theoretically analyze the performance benefit. Furthermore, a new BoF-based traffic classification method is proposed to aggregate the naive Bayes (NB) predictions of the correlated flows. We also present an analysis on prediction error sensitivity of the aggregation strategies. Finally, a large number of experiments are carried out on two large-scale real-world traffic datasets to evaluate the proposed scheme. The experimental results show that the proposed scheme can achieve much better classification performance than existing state-of-the-art traffic classification methods.
Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001, Wanlei Zhou 0001, Yong Xiang 0001
IEEE Trans. Inf. Forensics Secur.5
2013 Novel Z-Domain Precoding Method for Blind Separation of Spatially Correlated Signals
abstract
In this paper, we address the problem of blind separation of spatially correlated signals, which is encountered in some emerging applications, e.g., distributed wireless sensor networks and wireless surveillance systems. We preprocess the source signals in transmitters prior to transmission. Specifically, the source signals are first filtered by a set of properly designed precoders and then the coded signals are transmitted. On the receiving side, the Z-domain features of the precoders are exploited to separate the coded signals, from which the source signals are recovered. Based on the proposed precoders, a closed-form algorithm is derived to estimate the coded signals and the source signals. Unlike traditional blind source separation approaches, the proposed method does not require the source signals to be uncorrelated, sparse, or nonnegative. Compared with the existing precoder-based approach, the new method uses precoders with much lower order, which reduces the delay in data transmission and is easier to implement in practice.
Yong Xiang 0001, Dezhong Peng, Yang Xiang 0001, Song Guo 0001
IEEE Trans. Neural Networks Learn. Syst.1
2013 Projection-Pursuit-Based Method for Blind Separation of Nonnegative Sources
abstract
This paper presents a projection pursuit (PP) based method for blind separation of nonnegative sources. First, the available observation matrix is mapped to construct a new mixing model, in which the inaccessible source matrix is normalized to be column-sum-to-1. Then, the PP method is proposed to solve this new model, where the mixing matrix is estimated column by column through tracing the projections to the mapped observations in specified directions, which leads to the recovery of the sources. The proposed method is much faster than Chan's method, which has similar assumptions to ours, due to the usage of optimal projection. It is also more advantageous in separating cross-correlated sources than the independence- and uncorrelation-based methods, as it does not employ any statistical information of the sources. Furthermore, the new method does not require the mixing matrix to be nonnegative. Simulation results demonstrate the superior performance of our method.
Zuyuan Yang, Yong Xiang 0001, Yue Rong, Shengli Xie 0001
IEEE Trans. Neural Networks Learn. Syst.2
2013 Network Traffic Classification Using Correlation Information
abstract
Traffic classification has wide applications in network management, from security monitoring to quality of service measurements. Recent research tends to apply machine learning techniques to flow statistical feature based classification methods. The nearest neighbor (NN)-based method has exhibited superior classification performance. It also has several important advantages, such as no requirements of training procedure, no risk of overfitting of parameters, and naturally being able to handle a huge number of classes. However, the performance of NN classifier can be severely affected if the size of training data is small. In this paper, we propose a novel nonparametric approach for traffic classification, which can improve the classification performance effectively by incorporating correlated information into the classification process. We analyze the new classification approach and its performance benefit from both theoretical and empirical perspectives. A large number of experiments are carried out on two real-world traffic data sets to validate the proposed approach. The results show the traffic classification performance can be improved significantly even under the extreme difficult circumstance of very few training samples.
Jun Zhang 0010, Yang Xiang 0001, Yu Wang 0017, Wanlei Zhou 0001, Yong Xiang 0001
IEEE Trans. Parallel Distributed Syst.5
2012 Robust channel estimation algorithm for dual-hop MIMO relay channels
abstract
In conventional two-phase channel estimation algorithms for dual-hop multiple-input multiple-output (MIMO) relay systems, the relay-destination channel estimated in the first phase is used for the source-relay channel estimation in the second phase. For these algorithms, the mismatch between the estimated and the true relay-destination channel affects the accuracy of the source-relay channel estimation. In this paper, we investigate the impact of such channel state information (CSI) mismatch on the performance of the two-phase channel estimation algorithm. By explicitly taking into account the CSI mismatch, we develop a robust algorithm to estimate the source-relay channel. Numerical examples demonstrate the improved performance of the proposed algorithm.
Choo W. R. Chiong, Yue Rong, Yong Xiang 0001
PIMRC3
2012 Robust Patchwork-Based Embedding and Decoding Scheme for Digital Audio Watermarking
abstract
This paper presents a novel patchwork-based embedding and decoding scheme for digital audio watermarking. At the embedding stage, an audio segment is divided into two subsegments and the discrete cosine transform (DCT) coefficients of the subsegments are computed. The DCT coefficients related to a specified frequency region are then partitioned into a number of frame pairs. The DCT frame pairs suitable for watermark embedding are chosen by a selection criterion and watermarks are embedded into the selected DCT frame pairs by modifying their coefficients, controlled by a secret key. The modifications are conducted in such a way that the selection criterion used at the embedding stage can be applied at the decoding stage to identify the watermarked DCT frame pairs. At the decoding stage, the secret key is utilized to extract watermarks from the watermarked DCT frame pairs. Compared with existing patchwork watermarking methods, the proposed scheme does not require information of which frame pairs of the watermarked audio signal enclose watermarks and is more robust to conventional attacks.
Iynkaran Natgunanathan, Yong Xiang 0001, Yue Rong, Wanlei Zhou 0001, Song Guo 0001
IEEE Trans. Speech Audio Process.2
2012 A Dual-Channel Time-Spread Echo Method for Audio Watermarking
abstract
This work proposes a novel dual-channel time-spread echo method for audio watermarking, aiming to improve robustness and perceptual quality. At the embedding stage, the host audio signal is divided into two subsignals, which are considered to be signals obtained from two virtual audio channels. The watermarks are implanted into the two subsignals simultaneously. Then the subsignals embedded with watermarks are combined to form the watermarked signal. At the decoding stage, the watermarked signal is split up into two watermarked subsignals. The similarity of the cepstra corresponding to the watermarked subsignals is exploited to extract the embedded watermarks. Moreover, if a properly designed colored pseudonoise sequence is used, the large peaks of its auto-correlation function can be utilized to further enhance the performance of watermark extraction. Compared with the existing time-spread echo-based schemes, the proposed method is more robust to attacks and has higher imperceptibility. The effectiveness of our method is demonstrated by simulation results.
Yong Xiang 0001, Iynkaran Natgunanathan, Dezhong Peng, Wanlei Zhou 0001, Shui Yu 0001
IEEE Trans. Inf. Forensics Secur.1
2012 Document Clustering in Correlation Similarity Measure Space
abstract
This paper presents a new spectral clustering method called correlation preserving indexing (CPI), which is performed in the correlation similarity measure space. In this framework, the documents are projected into a low-dimensional semantic space in which the correlations between the documents in the local patches are maximized while the correlations between the documents outside these patches are minimized simultaneously. Since the intrinsic geometrical structure of the document space is often embedded in the similarities between the documents, correlation as a similarity measure is more suitable for detecting the intrinsic geometrical structure of the document space than euclidean distance. Consequently, the proposed CPI method can effectively discover the intrinsic structures embedded in high-dimensional document space. The effectiveness of the new method is demonstrated by extensive experiments conducted on various data sets and by comparison with existing document clustering methods.
Taiping Zhang, Yuan Yan Tang, Bin Fang 0001, Yong Xiang 0001
IEEE Trans. Knowl. Data Eng.4
2012 A Globally Convergent MC Algorithm With an Adaptive Learning Rate
abstract
This brief deals with the problem of minor component analysis (MCA). Artificial neural networks can be exploited to achieve the task of MCA. Recent research works show that convergence of neural networks based MCA algorithms can be guaranteed if the learning rates are less than certain thresholds. However, the computation of these thresholds needs information about the eigenvalues of the autocorrelation matrix of data set, which is unavailable in online extraction of minor component from input data stream. In this correspondence, we introduce an adaptive learning rate into the OJAn MCA algorithm, such that its convergence condition does not depend on any unobtainable information, and can be easily satisfied in practical applications.
Dezhong Peng, Zhang Yi 0001, Yong Xiang 0001, Haixian Zhang
IEEE Trans. Neural Networks Learn. Syst.3
2012 Time-Frequency Approach to Underdetermined Blind Source Separation
abstract
This paper presents a new time-frequency (TF) underdetermined blind source separation approach based on Wigner-Ville distribution (WVD) and Khatri-Rao product to separate N non-stationary sources from M(M <; N) mixtures. First, an improved method is proposed for estimating the mixing matrix, where the negative value of the auto WVD of the sources is fully considered. Then after extracting all the auto-term TF points, the auto WVD value of the sources at every auto-term TF point can be found out exactly with the proposed approach no matter how many active sources there are as long as N ≤ 2M-1. Further discussion about the extraction of auto-term TF points is made and finally the numerical simulation results are presented to show the superiority of the proposed algorithm by comparing it with the existing ones.
Shengli Xie 0001, Liu Yang 0002, Jun-Mei Yang, Guoxu Zhou, Yong Xiang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2012 Nonnegative Blind Source Separation by Sparse Component Analysis Based on Determinant Measure
abstract
The problem of nonnegative blind source separation (NBSS) is addressed in this paper, where both the sources and the mixing matrix are nonnegative. Because many real-world signals are sparse, we deal with NBSS by sparse component analysis. First, a determinant-based sparseness measure, named D-measure, is introduced to gauge the temporal and spatial sparseness of signals. Based on this measure, a new NBSS model is derived, and an iterative sparseness maximization (ISM) approach is proposed to solve this model. In the ISM approach, the NBSS problem can be cast into row-to-row optimizations with respect to the unmixing matrix, and then the quadratic programming (QP) technique is used to optimize each row. Furthermore, we analyze the source identifiability and the computational complexity of the proposed ISM-QP method. The new method requires relatively weak conditions on the sources and the mixing matrix, has high computational efficiency, and is easy to implement. Simulation results demonstrate the effectiveness of our method.
Zuyuan Yang, Yong Xiang 0001, Shengli Xie 0001, Shuxue Ding, Yue Rong
IEEE Trans. Neural Networks Learn. Syst.2
2012 Discriminating DDoS Attacks from Flash Crowds Using Flow Correlation Coefficient
abstract
Distributed Denial of Service (DDoS) attack is a critical threat to the Internet, and botnets are usually the engines behind them. Sophisticated botmasters attempt to disable detectors by mimicking the traffic patterns of flash crowds. This poses a critical challenge to those who defend against DDoS attacks. In our deep study of the size and organization of current botnets, we found that the current attack flows are usually more similar to each other compared to the flows of flash crowds. Based on this, we proposed a discrimination algorithm using the flow correlation coefficient as a similarity metric among suspicious flows. We formulated the problem, and presented theoretical proofs for the feasibility of the proposed discrimination method in theory. Our extensive experiments confirmed the theoretical analysis and demonstrated the effectiveness of the proposed method in practice.
Shui Yu 0001, Wanlei Zhou 0001, Weijia Jia 0001, Song Guo 0001, Yong Xiang 0001, Feilong Tang 0001
IEEE Trans. Parallel Distributed Syst.5
2012 Channel Estimation of Dual-Hop MIMO Relay System via Parallel Factor Analysis
abstract
The optimal source precoding matrix and relay amplifying matrix have been developed in recent works on multiple-input multiple-output (MIMO) relay communication systems assuming that the instantaneous channel state information (CSI) is available. However, in practical relay communication systems, the instantaneous CSI is unknown, and therefore, has to be estimated at the destination node. In this paper, we develop a novel channel estimation algorithm for two-hop MIMO relay systems using the parallel factor (PARAFAC) analysis. The proposed algorithm provides the destination node with full knowledge of all channel matrices involved in the communication. Compared with existing approaches, the proposed algorithm requires less number of training data blocks, yields smaller channel estimation error, and is applicable for both one-way and two-way MIMO relay systems with single or multiple relay nodes. Numerical examples demonstrate the effectiveness of the PARAFAC-based channel estimation algorithm.
Yue Rong, Muhammad R. A. Khandaker, Yong Xiang 0001
IEEE Trans. Wirel. Commun.3
2011 Echo hiding based stereo audio watermarking against pitch-scaling attacks
abstract
In audio watermarking, the robustness against pitch-scaling attack, is one of the most challenging problems. In this paper, we propose an algorithm, based on traditional time-spread(TS) echo hiding based audio watermarking to solve this problem. In TS echo hiding based watermarking, pitch-scaling attack shifts the location of pseudonoise (PN) sequence which appears in the cepstrum domain. Thus, position of the peak, which occurs after correlating with PN-sequence changes by an un-known amount and that causes the error. In the proposed scheme, we replace PN-sequence with unit-sample sequence and modify the decoding algorithm in such a way it will not depend on a particular point in cepstrum domain for extraction of watermark. Moreover proposed algorithm is applied to stereo audio signals to further improve the robustness. Experimental results illustrate the effectiveness of the proposed algorithm against pitch-scaling attacks compared to existing methods. In addition to that proposed algorithm also gives better robustness against other conventional signal processing attacks.
Iynkaran Natgunanathan, Yong Xiang 0001
SIN2
2011 Effective Pseudonoise Sequence and Decoding Function for Imperceptibility and Robustness Enhancement in Time-Spread Echo-Based Audio Watermarking
abstract
This paper proposes an effective pseudonoise (PN) sequence and the corresponding decoding function for time-spread echo-based audio watermarking. Different from the traditional PN sequence used in time-spread echo hiding, the proposed PN sequence has two features. Firstly, the echo kernel resulting from the new PN sequence has frequency characteristics with smaller magnitudes in perceptually significant region. This leads to higher perceptual quality. Secondly, the correlation function of the new PN sequence has three times more large peaks than that of the existing PN sequence. Based on this feature, we propose a new decoding function to improve the robustness of time-spread echo-based audio watermarking. The effectiveness of the proposed PN sequence and decoding function is illustrated by theoretical analysis, simulation examples, and listening test.
Yong Xiang 0001, Dezhong Peng, Iynkaran Natgunanathan, Wanlei Zhou 0001
IEEE Trans. Multim.1
2011 Multiuser Multi-Hop MIMO Relay Systems with Correlated Fading Channels
abstract
In this letter, we address multiuser multi-hop multiple-input multiple-output (MIMO) relay communication systems with correlated MIMO fading channels. In particular, we consider the practical scenario where the channel fading is fast and thus the instantaneous channel state information (CSI) is only available at the destination node, but unknown at all users and all relay nodes. We derive the structure of the optimal user precoding matrices and relay amplifying matrices that maximizes the users-destination ergodic sum mutual information. Compared with existing works, our results are more general, since we address multiuser scenarios, consider MIMO relays with a finite dimension, and take into account the noise vector at each relay node.
Yue Rong, Yong Xiang 0001
IEEE Trans. Wirel. Commun.2
2009 A Novel Pseudonoise Sequence for Time-Spread Echo Based Audio Watermarking
abstract
This paper deals with the problem of digital audio watermarking using echo hiding. Compared to many other methods for audio watermarking, echo hiding techniques exhibit advantages in terms of relatively simple encoding and decoding, and robustness against common attacks. The low security issue existing in most echo hiding techniques is overcome in the time-spread echo method by using pseudonoise (PN) sequence as a secret key. In this paper, we propose a novel sequence, in conjunction with a new decoding function, to improve the imperceptibility and the robustness of time-spread echo based audio watermarking. Theoretical analysis and simulation examples illustrate the effectiveness of the proposed sequence and decoding function.
Iynkaran Natgunanathan, Yong Xiang 0001
GLOBECOM2
2008 A neural networks learning algorithm for minor component analysis and its convergence analysis
Dezhong Peng, Zhang Yi 0001, Jiancheng Lv 0001, Yong Xiang 0001
Neurocomputing4
2006 A New Blind Signal Separation Algorithm for Instantaneous MIMO System
abstract
We address the problem of adaptive blind source separation (BSS) from instantaneous multi-input multi- output (MIMO) channels. In this paper, we propose a new constant modulus (CM)-based algorithm which employ nonlinear function as the de-correlation term. Moreover, it is shown by theoretical analysis that the proposed algorithm has less mean square error (MSE), i.e., better separation performance, in steady state than the cross-correlation and constant modulus algorithm (CC-CMA). Numerical simulations show the effectiveness of the proposed result.
Nong Gu, Zhenying Guan, Saeid Nahavandi, Yong Xiang 0001
GLOBECOM4