EDBT 2026 Demo / reviewers in the wild / expert
Weidong Qiu
dblp:26/221
· DBLP profile ↗
78ranked-venue papers
5as first author
49since 2021 · last 2026
0000-0001-6428-1655ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 24 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 19 · 12 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 4 since 2021Computer networks · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedTopo: Topology-Informed Representation Alignment in Federated Learning Under Non-I.I.D. ConditionsabstractCurrent federated-learning models deteriorate under heterogeneous (non-I.I.D.) client data, as their feature representations diverge and pixel- or patch-level objectives fail to capture the global topology which is essential for high-dimensional visual tasks. We propose FedTopo, a framework that integrates Topological-Guided Block Screening (TGBS) and Topological Embedding (TE) to leverage topological information, yielding coherently aligned cross-client representations by Topological Alignment Loss (TAL). First, Topology-Guided Block Screening (TGBS) automatically selects the most topology-informative block, i.e., the one with maximal topological separability, whose persistence-based signatures best distinguish within- versus between-class pairs, ensuring that subsequent analysis focuses on topology-rich features. Next, this block yields a compact Topological Embedding, which quantifies the topological information for each client. Finally, a Topological Alignment Loss (TAL) guides clients to maintain topological consistency with the global model during optimization, reducing representation drift across rounds. Experiments on Fashion-MNIST, CIFAR-10, and CIFAR-100 under four non-I.I.D. partitions show that FedTopo accelerates convergence and improves accuracy over strong baselines. Liyao Xiang, Peng Tang 0002, Weidong Qiu |
AAAI | 4 |
| 2026 | Behind the Meme: Understanding User Experiences with Memes on Social MediaabstractWhile memes enhance social interaction on social media, they can raise privacy and security concerns. Despite research on overtly toxic or unsafe memes, little attention has been given to users’ experiences with seemingly safe memes and how contextual factors trigger privacy concerns. This study explores users’ comfort levels, influencing factors, underlying reasons for discomfort, and unmet needs when engaging with such memes. We first collected and analyzed 2,317 Reddit posts describing real-world meme experiences, then conducted an online survey with 324 participants to evaluate comfort across curated scenarios. Our findings reveal that perceived-safe memes can cause harm when shared inappropriately, with comfort shaped by content and context. Privacy concerns intensify with deeper involvement, strangers, and sensitive meme topics. We identified users’ desire for consent and control in meme interactions. Based on our study, we make recommendations for users, developers of social media platforms and policymakers to address meme-related privacy and contextual concerns. Yuqi Niu, Dilara Keküllüoglu, Weidong Qiu, Nadin Kökciyan |
CHI | 3 |
| 2026 | Faster Than Ever: A New Lightweight Private Set Intersection and Its Variants
Guowei Ling, Peng Tang 0002, Jinyong Shan, Liyao Xiang, Weidong Qiu |
NDSS | 5 |
| 2026 | Blockchain-Based Privacy-Preserving Alternative Credit Data SharingabstractIn comparison to the lending data submitted by banks to credit bureaus under the traditional credit scoring paradigm, alternative credit data (such as social media activities and e-commerce consumption records) has increasingly demonstrated its significance in enhancing the accuracy of credit scores and addressing the issue of credit-invisible individuals in recent years. However, credit scoring model based on alternative credit data typically necessitates large-scale data circulation and may involve sensitive information, thereby raising concerns related to data security, user privacy, and data rights. Traditional cryptographic methods often encounter limitations in functionality, efficiency, flexibility, and traceability when addressing these issues. This article initially proposes a novel credit data sharing framework based on an alternative data cloud platform. Subsequently, based on this framework, a blockchain-based privacy-preserving alternative credit data sharing scheme is constructed. This scheme achieves efficient, privacy-preserving, and wildcard-supported attribute-based encryption (ABE) scheme through inner product operations, and implements a “two-level” access control by designing a keyword search mechanism in conjunction with the aforementioned scheme. Furthermore, a hybrid encryption mechanism is introduced to further enhance efficiency and security under high-frequency access scenarios. Security analysis and rigorous formal security reductions have been conducted to demonstrate the security of the proposed scheme. Comparative experimental results also indicate that the proposed scheme exhibits significant advantages in practicality compared with related schemes. Yangyang Bao, Jianfei Sun, Xiaochun Cheng, Weidong Qiu, Liming Nie |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2026 | A Comprehensive Study on GDPR-Oriented Analysis of Privacy Policies: Taxonomy, Corpus and GDPR Concept ClassifiersabstractMachine learning (ML) based classifiers that take a privacy policy as the input and predict relevant concepts are useful in different applications such as (semi-)automated compliance analysis against requirements of a specific data protection law such as the EU GDPR. Although many researchers have studied ML-based privacy policy concept classifiers, we observed multiple research gaps, e.g., the lack of a more complete GDPR taxonomy and the less consideration of hierarchical information in privacy policies. To fill such research gaps, we produced a more complete GDPR-oriented privacy policy concept taxonomy, constructed the first privacy policy corpus with explicitly hierarchical information at three levels, and conducted the most comprehensive performance evaluation study of GDPR concept classifiers for privacy policies, cover many aspects that have not been studied systematically. Our work led to multiple findings and insights, including the usefulness of considering hierarchical contextual features and different hierarchical structures, the observation that a “one size fits all” approach may not work, the reduced performance of such classifiers on our newly constructed corpus especially after the first level, and the necessity to split the training and testing sets by documents. Peng Tang 0002, Weidong Qiu, Haochen Mei, Allison Holmes, Fenghua Li 0001, Shujun Li 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | Efficient Updatable PSI From Asymmetric PSI and PSUabstractPrivate Set Intersection (PSI) allows two mutually untrusted parties to compute the intersection of their private sets without revealing additional information. In general, PSI operates in a static setting, where the computation is performed only once on the input sets of both parties. Badrinarayanan et al. initiated the study of Updatable PSI (UPSI), which extends this capability to dynamically updating sets, enabling both parties to securely compute the intersection as their sets are modified while incurring significantly less overhead than re-executing a conventional PSI. However, existing UPSI protocols either do not support arbitrary deletion of elements or incur high computational and communication overhead. This work combines asymmetric PSI with Private Set Union (PSU) to present a novel UPSI protocol, which supports arbitrary additions and deletions of elements, offering a flexible approach to update sets. Furthermore, we design a primitive called multi-round OPRF to satisfy the forward security (IEEE TIFS 2024). Our protocol enjoys efficient performance compared to previous work. Specifically, we implement our protocol and compare it against state-of-the-art conventional PSI and UPSI protocols. Experimental results demonstrate that our UPSI protocol achieves up to three orders of magnitude reduction in computational overhead and incurs 190 ∼ 707×less communication overhead than the state-of-the-art UPSI protocol (ASIACRYPT 2024) that supports arbitrary additions and deletions. Guowei Ling, Peng Tang 0002, Shifeng Sun 0001, Weidong Qiu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | "I am not the primary focus" - Understanding the Perspectives of Bystanders in Photos Shared OnlineabstractWhen taking photos in a crowd, unintended individuals, such as bystanders, are often captured alongside the main subject(s). In an effort to protect bystanders' privacy, existing methods have been developed to automatically detect bystanders. However, inconsistent definitions of who qualifies as a bystander limit their effectiveness. To better understand bystanders' perceptions, we conducted an online survey with 486 participants, analyzing their responses to 864 image-based scenarios and their comfort with sharing these images online. Our results revealed no significant correlation between comfort with public photo sharing and bystander status. We identified limitations in current bystander detection methodologies, as they often fail to recognize bystanders who are not clearly in the background, hence missing individuals with privacy concerns. Moreover, comfort with public sharing varied significantly depending on the image context. Our findings highlight the importance of considering the context of captured images to address privacy concerns in image sharing. Yuqi Niu, Nicole Meng 0001, Weidong Qiu, Nadin Kökciyan |
CHI | 3 |
| 2025 | FedDEAP: Adaptive Dual-Prompt Tuning for Multi-Domain Federated LearningabstractFederated learning (FL) enables multiple clients to collaboratively train machine learning models without exposing local data, balancing performance and privacy. However, domain shift and label heterogeneity across clients often hinder the generalization of the aggregated global model. Recently, large-scale vision-language models like CLIP have shown strong zero-shot classification capabilities, raising the question of how to effectively fine-tune CLIP across domains in a federated setting. In this work, we propose an adaptive federated prompt tuning framework, FedDEAP, to enhance CLIP's generalization in multi-domain scenarios. Our method includes the following three key components: (1) To mitigate the loss of domain-specific information caused by label-supervised tuning, we disentangle semantic and domain-specific features in images by using semantic and domain transformation networks with unbiased mappings; (2) To preserve domain-specific knowledge during global prompt aggregation, we introduce a dual-prompt design with a global semantic prompt and a local domain prompt to balance shared and personalized information; (3) To maximize the inclusion of semantic and domain information from images in the generated text features, we align textual and visual representations under the two learned transformations to preserve semantic and domain consistency. Theoretical analysis and extensive experiments on four datasets demonstrate the effectiveness of our method in enhancing the generalization of CLIP for federated image recognition across multiple domains. Yubin Zheng, Pak-Hei Yeung, Tianjie Ju, Peng Tang 0002, Weidong Qiu, Jagath C. Rajapakse |
ACM Multimedia | 6 |
| 2025 | Practical Collision Attacks on Reduced-Round Xoodyak Hash Mode
Huina Li, Weidong Qiu |
SAC | 3 |
| 2025 | Fed-SHA: An Efficient Hyperparameter Optimization Approach for Federated LearningabstractFederated Learning (FL) enables collaborative machine learning without exposing raw data, effectively addressing privacy concerns. However, FL still faces major challenges, particularly in communication efficiency. Among them, federated hyperparameter optimization emerges as a critical yet underexplored problem. Properly initialized hyperparameters play a vital role in accelerating convergence and significantly reducing communication overhead in FL settings. To address this, we propose a communication-efficient algorithm, Fed-SHA, based on continuous halving strategies. This work systematically analyzes the unique challenges of hyperparameter optimization in FL and introduces an alternative optimization objective for client-side tuning, which can be solved independently of federated training. Inspired by multi-fidelity optimization, Fed-SHA combines local search with global halving to improve efficiency. The algorithm leverages local computation to explore hyperparameter configurations and uses a central server to approximate federated loss and progressively reduce the search space. The experimental results show that Fed-SHA significantly reduces communication rounds and costs, while achieving better performance than existing baseline methods. Yubin Zheng, Peng Tang 0002, Xiheng Zhang, Yijie Hong, Weidong Qiu |
SMC | 5 |
| 2025 | Everyone's Privacy Matters! An Analysis of Privacy Leakage from Real-World Facial Images on Twitter and Associated User BehaviorsabstractOnline users often post facial images of themselves and other people on online social networks (OSNs) and other Web 2.0 platforms, which can lead to potential privacy leakage of people whose faces are included in such images. There is limited research on understanding face privacy in social media while considering user behavior. It is crucial to consider privacy of subjects and bystanders separately. This calls for the development of privacy-aware face detection classifiers that can distinguish between subjects and bystanders automatically. This paper introduces such a classifier trained on face-based features, which outperforms the two state-of-the-art methods with a significant margin (by 13.1% and 3.1% for OSN images, and by 17.9% and 5.9% for non-OSN images). We developed a semi-automated framework for conducting a large-scale analysis of the face privacy problem by using our novel bystander-subject classifier. We collected 27,800 images, each including at least one face, shared by 6,423 Twitter users. We then applied our framework to analyze this dataset thoroughly. Our analysis reveals eight key findings of different aspects of Twitter users' real-world behaviors on face privacy, and we provide quantitative and qualitative results to better explain these findings. We share the practical implications of our study to empower online platforms and users in addressing the face privacy problem efficiently. Yuqi Niu, Weidong Qiu, Peng Tang 0002, Lifan Wang, Shujun Li 0001, Nadin Kökciyan, Ben Niu 0001 |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2025 | Exploring Efficient Hardware Accelerator for Learning-Based Image CompressionabstractRecently, learning-based image compression (LIC) methods have surpassed manually designed approaches in both compression quality and bitrate. However, increasing computational demands and insufficient optimizations in codec performance have hindered the advancement of LIC acceleration. Most researches focus on optimizing specific components, often neglecting the sources of underutilization during the execution of LIC models. Generally, efficient LIC acceleration encounters three primary challenges: 1) extra overheads introduced by individual optimizations; 2) load and computation imbalances in small kernels; and 3) mismatches between hardware configurations and the LIC models. To address these challenges, we propose a framework named extensive accelerator for LIC (X-LIC) for efficiently exploring the design space under constrained resources. First, we quantitatively characterize a representative LIC model, including its latency, computation size, and temporal utilization across various accelerators. We design a hardware-optimized quantization method to compensate for the lack of LIC-oriented research, particularly regarding data precision, distortion, and resource consumption. Additionally, we propose a parameterized LIC accelerator architecture that integrates seamlessly with existing loop optimization models and supports various LIC operators. Two optimization schemes are proposed for redundant computation in transposed convolution and load and computation imbalance in small kernels. Experimental results show that our framework demonstrates significant flexibility across a broad design space, achieving an average of 78%–95% of the theoretical peak performance and up to 688.2/759.1 GOP/s en/de-coder performance with INT8 precision. As a result, the en/de-coder performance can reach up to 33/36 FPS in 720P resolution. An FPGA demo of X-LIC is available athttps://github.com/sjtu-tcloud/X-LIC. Chen Chen 0067, Kaicheng Guo, Xingzi Yu, Weidong Qiu, Zhengwei Qi, Haibing Guan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Analyzing Social Media Comments to Understand and Detect Privacy ViolationsabstractSocial media users generate vast amounts of content that may contain sensitive information, posing threats to online and real-world privacy and security. Existing automated classifiers focus primarily on preserving privacy in original posts while neglecting the potential privacy risks in comments. This article addresses this gap by presenting a comprehensive study on privacy leaks in textual comments. We curate a real-world dataset of 1250 tweet-comment interactions from Twitter and also introduce features to detect privacy leaks in comments. Using these features, we train various classifiers that achieve an average F-score of 0.86 on the Twitter dataset. We then use our methods to detect privacy leaks on another social media platform, namely Reddit, and we get an average F-score of 0.90, which demonstrates the adaptability of our methods. This research shows the significance of addressing privacy leaks in comments. We also show how our approach could work across two different social media platforms. Yuqi Niu, Nadin Kökciyan, Weidong Qiu |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | Ultra-Fast Private Set Intersection From Efficient Oblivious Key-Value StoresabstractPrivate Set Intersection (PSI) enables us to compute the intersection of private sets without leaking additional data. The state-of-the-art PSI protocol$\mathsf {RR22}$(CCS 2022) is derived from an Oblivious Pseudo-Random Function (OPRF) protocol based on Oblivious Key-Value Stores (OKVS). However, the existing OKVS suffers either low computation efficiency or high encoding redundancy. In this work, we propose a new efficient bucket-based OKVS with only 1% redundancy. The encoding algorithm of our OKVS is 4 to 15 times faster than the recent state-of-the-art OKVS (USENIX Security 2023). Specifically, our OKVS can encode$2^{24}$key-value pairs in only 2.1 to 8.5 seconds, corresponding to 30% to 1% redundancy, while the latter takes about 30 seconds with at least 3%. We can then obtain a new ultra-fast PSI protocol with lower communication from our OKVS in both semi-honest and malicious settings. Furthermore, we implemented our PSI protocol and conducted an extensive evaluation, which shows that it outperforms the existing PSI protocols, such as$\mathsf {KKRT16}$(CCS 2016),$\mathsf {CM20}$(Crypto 2020),$\mathsf {RS21}$(EuroCrypt 2021),$\mathsf {RR22}$(CCS 2022), and$\mathsf {KBM23}$(NDSS 2023). Since our PSI features an ultra-low communication overhead, it has overall advantages for the network environment with a small bandwidth. For example, our PSI takes only about 468 and 476 seconds in semi-honest and malicious settings with the input size of$2^{24}$when the bandwidth is 10 Mbps, while the state-of-the-art$\mathsf {RR22}$requires about 541 and 625 seconds. Our implementation is available onhttps://github.com/ShallMate/fastpsi. Guowei Ling, Peng Tang 0002, Fei Tang 0001, Shifeng Sun 0001, Shouling Ji, Weidong Qiu |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Privacy-Preserving Authorized Set Matching via Dishonest Majority Multiparty ComputationabstractPrivate Set Intersection (PSI) enables each party with a private set to compute the intersection without disclosing other information. However, even in maliciously secure PSI, it does not guarantee input authenticity and output integrity, which becomes problematic in certain scenarios. For instance, in Web 3.0, one of the essential requirements is to find common certifiers among the parties. However, if certifier identities are meant to be protected, some parties may attempt to forge certifier identities or intentionally exclude a particular certifier during protocol execution. Recently, the Private Certifier Intersection (PCI), a variant of PSI, has been proposed to address this problem. Nevertheless, it incurs significantly high computational and communication overhead. This work proposes thePrivate Identity Intersection(PII), which takes private identifiers and corresponding anonymous signatures from mutually distrusting parties as input, verifies them, and delivers the intersection of the successfully verified identifiers to all parties while ensuring the integrity of the output. Furthermore, PII can naturally extend from two to multiple-party settings while resisting the collusion attack. To achieve the ideal functionality of PII, we implement a user-friendly MPC framework called$\mathsf {Oryx}$without third-party libraries. Based on$\mathsf {Oryx}$, we instantiate PII with two digital signature schemes, one proposed in this paper. Compared to existing work, our PII protocols reduce the computation overhead by up to$163\times$and the communication overhead by up to$190\times$, representing an improvement of two orders of magnitude. To demonstrate the practicality of our work, we evaluate its performance in WAN environments with bandwidths of 100 Mbps and 500 Mbps, under a fixed latency of 20 ms. Guowei Ling, Peng Tang 0002, Fei Tang 0001, Shifeng Sun 0001, Jinyong Shan, Liyao Xiang, Weidong Qiu |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2025 | Privacy-Preserving Fine-Grained Data Sharing With Dynamic Service for the Cloud-Edge IoTabstractThe cloud-edge computing model has been expected to play a revolutionary role in promoting the quality of future generation large-scale Internet of Things (IoT) services. However, security and privacy in data sharing remain crucial issues hindering the success of cloud-edge IoT services. While some solutions based on attribute-based encryption (ABE) have been proposed to address these issues, they still face practical challenges such as attribute privacy leakage, resource-constrained devices, dynamic user groups, inflexible and inefficient service response. To address these challenges, this paper proposes a privacy-preserving fine-grained data sharing scheme with dynamic service (PF2DS), which implements access control by calculating the inner product between an attribute vector and an access vector. PF2DS is also capable of providing dynamic user group services through an efficient and indirect user revocation mechanism that periodically updates the key-embedded leaf nodes. Building on PF2DS, edge-assisted PF2DS (EPF2DS) delegates most of the operations to the edge device, which facilitates the performance of resource-constrained IoT devices. EPF2DS also supports efficient and asynchronous keyword search over the ciphertexts stored in the cloud. We demonstrate the security by the rigorous security proof. Both theoretical comparisons and experimental simulations demonstrate the practicality and superiority of our schemes over existing works. Jianfei Sun, Yangyang Bao, Weidong Qiu, Rongxing Lu, Songnian Zhang, Yunguo Guan, Xiaochun Cheng |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | FLAD: Byzantine-Robust Federated Learning Based on Gradient Feature Anomaly DetectionabstractFederated Learning (FL) has gained significant attention due to its ability to jointly train global models by exchanging local gradients instead of raw local datasets. However, poisoning attacks have emerged as a severe threat to FL security, where malicious clients submit crafted gradients to compromise the integrity and availability of the model. Although researchers have worked on countering these attacks to achieve Byzantine-robust FL, it remains challenging to balance high accuracy, robustness, and efficiency simultaneously. We propose FLAD, a novel Byzantine-robust FL approach based on gradient feature anomaly detection, which is the first approach that uses neural networks to adaptively learn gradient features and measure feature similarity to counteract various types of poisoning attacks. Specifically, FLAD employs a small clean dataset to bootstrap trust and trains Feature Extraction Models (FEM). With FEM and DBSCAN clustering, abnormal gradients from malicious clients are detected and eliminated. Extensive experiments on both Non-IID and IID datasets demonstrate that FLAD achieves superior accuracy, robustness, efficiency, and generalizability compared to state-of-the-art approaches. Additionally, we implement privacy-preserving FLAD (PFLAD) using CKKS and Random Permutation techniques to ensure transmitted gradient privacy. Peng Tang 0002, Weidong Qiu, Zhenyu Mu, Shujun Li 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | SE#PCFG: Semantically Enhanced PCFG for Password Analysis and CrackingabstractMuch research has been done on user-generated textual passwords. Surprisingly, semantic information in such passwords remain under-investigated, with passwords created by English- and/or Chinese-speaking users being more studied with limited semantics. This article fills this gap by proposing ageneral frameworkbased onsemantically enhancedPCFG (probabilistic context-free grammars) named SE#PCFG. It allowed us to consider 43 types of semantic information, the richest set considered so far, for password analysis. Applying SE#PCFG to 17 large leaked password databases of user speaking four languages (English, Chinese, German and French), we demonstrate its usefulness and report a wide range of new insights about password semantics at different levels such as cross-website password correlations. Furthermore, based on SE#PCFG and a new systematic smoothing method, we proposed the Semantically Enhanced Password Cracking Architecture (SEPCA), and compared its performance against three SOTA (state-of-the-art) benchmarks in terms of the password coverage rate: two other PCFG variants and neural network. Our experimental results showed that SEPCA outperformed all the three benchmarks consistently and significantly across 52 test cases, by up to 21.53%, 52.55% and 7.86%, respectively, at the user-level (with duplicate passwords). At the level of unique passwords, SEPCA also beats the three counterparts by up to 43.83%, 94.11% and 11.16%, respectively. Yangde Wang, Weidong Qiu, Peng Tang 0002, Shujun Li 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | What Makes a Good Exchange? Privacy-Preserving and Fair Contract Agreement in Data TradingabstractExchange-assisted data trading (EADT) has become an essential paradigm in current data marketplaces. With data exchanges, sellers and buyers can trade data in an efficient and convenient way. However, existing EADT systems are vulnerable to privacy violations. Sensitive information about the data owned by sellers (manifested as attributes of the data) and the purchasing requirements of buyers (manifested as interests) are highly susceptible to leakage. On the one hand, buyers and sellers have direct access to the type of data supplied or desired before the data transaction is established. On the other hand, the information about transactions between the seller and buyer is transparent to the exchange, including the content of the transaction contract. In addition, the participants are likely to repudiate the content of previously accepted contracts or trigger a bidding war by contract first authorized by others, which raises threats towards authenticity and fairness. In this paper, we investigate the contract agreement in actual EADT systems, enumerate the inherent requirements of secrecy and fairness, and formally define them. Then we propose a privacy-preserving and fair contract agreement framework, dubbed PFCA, which consists of order-matching, negotiation, and authorization. We further propose a practical instantiation of PFCA, dubbed BestPFCA, utilizing efficient private set intersection (PSI), secure messaging (SM), and three-party signature (TPS). In addition, we also implement a BestPFCA prototype and conduct a comprehensive performance evaluation, which demonstrates the efficiency and practicality of BestPFCA. Yuan Zhang 0006, Yaqing Song, Weidong Qiu, Hongwei Li 0001, Qiang Tang 0005 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | More Efficient, Privacy-Enhanced, and Powerful Privacy-Preserving Feature Retrieval Private Set IntersectionabstractPrivate Set Intersection (PSI) allows two parties, the sender and the receiver, each possessing a private set, to compute the intersection of their sets, with only the receiver learning the intersection and without revealing any additional information. Privacy-Preserving Feature Retrieval PSI (P2FRPSI) is a variant of PSI. In P2FRPSI, the receiver designs a predicate and obtains the intersection of private sets that satisfy this predicate, while the sender learns nothing about the predicate. However, the existing two PRFPSI protocols (TIFS 2024), based respectively on the DH key agreement and Oblivious Pseudo- Random Function (OPRF), are not highly efficient due to their reliance on expensive homomorphic encryption. Moreover, the existing DH-based P2FRPSI protocol reveals the output size and the original intersection size to the sender. We also observed that the existing P2FRPSI protocols do not support threshold retrieval and the logical connective OR and can only work when feature values of the sender have very low dimensionality. This paper also proposes two new P2FRPSI protocols, one based on DH key agreement and the other based on OPRF, to fully address the issues present in existing P2FRPSI protocols. Our DH-based P2FRPSI is 30× faster than the existing DH-based protocol, with only a 36% increase in communication overhead. Furthermore, our OPRF-based P2FRPSI protocol is 2× as fast as existing OPRF-based protocol and reduces communication overhead by a factor of 4.6. Our DH-based P2FRPSI protocol completely eliminates the leakage of the original intersection size and the output size. Meanwhile, our protocols support the logical connective OR for linking sub-predicates and also enable threshold-based retrieval. They are proven to be secure in the semi-honest model. Our open-source implementations can be found at https://github.com/ShallMate/pfrpsi, which can help readers understand our protocols and reproduce the experiments. Guowei Ling, Peng Tang 0002, Jinyong Shan, Fei Tang 0001, Weidong Qiu |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | PassTSL: Modeling Human-Created Passwords Through Two-Stage Learning
Haozhang Li, Yangde Wang, Weidong Qiu, Shujun Li 0001, Peng Tang 0002 |
ACISP (3) | 3 |
| 2024 | AlgSAT - A SAT Method for Verification of Differential Trails from an Algebraic Perspective
Huina Li, Guozhen Liu, Weidong Qiu |
ACISP (1) | 5 |
| 2024 | Federated Semi-supervised Learning for Medical Image Segmentation with Intra-client and Inter-client ConsistencyabstractMedical image segmentation plays a vital role in medical image analysis. However, it is impractical to build a large-scale centralized segmentation dataset due to the privacy of medical images. Federated learning (FL) aims to train a shared model of isolated clients without local data exchange which aligns well with the scarcity and privacy characteristics of medical images. Moreover, there is a large amount of unlabeled data in clients due to the difficulty in annotating medical images. Federated semi-supervised learning (FSSL) can leverage the unlabeled data of clients to improve the performance of the global model. Many existing FSSL methods apply the complicated semi-supervised learning protocols and some of them neglect the problem of data heterogeneity in FL. In this paper, we propose a novel federated semi-supervised learning framework for medical image segmentation incorporating intra-client and inter-client consistency learning. The intra-client consistency learning can introduce global data noise in data augmentation which can improve the generalization ability of the model and reduce the impact of data heterogeneity. The inter-client consistency learning is proposed to expand the feature search space and learn the ensemble knowledge of different clients. The two consistency learning mechanisms are achieved with the assistance of a Variational Autoencoder (VAE) trained collaboratively by clients. The experimental results illustrate that our method outperforms the state-of-the-art methods under different FSSL settings. The code is available at https://github.com/zyb98/FV2IC. Yubin Zheng, Peng Tang 0002, Tianjie Ju, Weidong Qiu, Jagath C. Rajapakse |
BIBM | 5 |
| 2024 | ADP-VFL: An Adaptive Differential Privacy Scheme for VPP Based on Federated LearningabstractIn recent years, with the remarkable development of Virtual Power Plants (VPP) and the surge in the number of Electric Vehicles (EVs), the issue of data privacy leakage has become increasingly prominent. The effectiveness of existing federated learning schemes in mitigating data privacy leakage, it still faces potential threats such as inference attacks and user and server collusion. To protect the privacy of federated learning, some schemes have introduced differential privacy(DP). Nevertheless, applying DP will inevitably affect the accuracy to some extent. In this paper, we propose an adaptive differential privacy scheme for VPP based on federated learning, named ADP-VFL. Our ADP-VFL scheme can achieve data privacy preservation by transmitting the noise-added data as a chain, defend against inference attacks by innovating offset noise mechanism and a parallel transmission scheme. The performance evaluation results demonstrate that the proposed scheme can improve the aggregation accuracy and reduces the communication overhead. Mi Wen, Weiwei Li 0007, Ben Niu 0001, Weidong Qiu, Fenghua Li 0001 |
ICC | 5 |
| 2024 | Feature Norm Regularized Federated Learning: Utilizing Data Disparities for Model Performance Gains
Liyao Xiang, Peng Tang 0002, Weidong Qiu |
IJCAI | 4 |
| 2024 | Fine-Grained Contrastive Learning for Pulmonary Nodule ClassificationabstractLung cancer is one of the most threatening human diseases which develops from malignant pulmonary nodules. The accurate classification of benign and malignant pulmonary nodules is important for formulating treatment plans and improving the survival rate of lung cancer patients. However, detecting and classifying pulmonary nodules pose a challenge due to their small region of interest (ROI) and diverse patterns. Many deep learning methods struggle to extract the valid feature representations of small pulmonary nodules which results in poor classification performance. In this work, we propose a novel fine-grained contrastive learning method for pulmonary nodule classification. Traditional contrastive learning utilizes positive and negative sample pairs to learn good representations of images. However, different pulmonary nodule patterns contain similar features that are vital for nodule classification. We propose using attributes which are categorial features labeled by experts to adjust the importance of sample pairs in contrastive learning. The method can guide the model to understand the degree of similarity and difference of pulmonary nodules and obtain their meaningful representations. Due to the small ROI of pulmonary nodules, we discard CNN backbone and use Vision Transformer (ViT) to learn the correlation of adjacent small-size slices. Compared with other deep learning methods, our method achieves the state-of-the-art performance in five metrics. In addition, we transfer the representations of pulmonary nodules learned by our model to a new dataset. The model obtains competitive performance without additional domain expertise which proves the prospect of our model in transfer learning. Yubin Zheng, Peng Tang 0002, Tianjie Ju, Weidong Qiu |
IJCNN | 4 |
| 2024 | FLKT: Improving the Fidelity and Robustness of Federated Learning Aggregation Rules via the Key-Data and Trap-Model
Peng Tang 0002, Weidong Qiu |
SecureComm (4) | 4 |
| 2024 | When graph convolution meets double attention: online privacy disclosure detection with multi-label text classificationabstractAbstract With the rise of Web 2.0 platforms such as online social media, people’s private information, such as their location, occupation and even family information, is often inadvertently disclosed through online discussions. Therefore, it is important to detect such unwanted privacy disclosures to help alert people affected and the online platform. In this paper, privacy disclosure detection is modeled as a multi-label text classification (MLTC) problem, and a new privacy disclosure detection model is proposed to construct an MLTC classifier for detecting online privacy disclosures. This classifier takes an online post as the input and outputs multiple labels, each reflecting a possible privacy disclosure. The proposed presentation method combines three different sources of information, the input text itself, the label-to-text correlation and the label-to-label correlation. A double-attention mechanism is used to combine the first two sources of information, and a graph convolutional network is employed to extract the third source of information that is then used to help fuse features extracted from the first two sources of information. Our extensive experimental results, obtained on a public dataset of privacy-disclosing posts on Twitter, demonstrated that our proposed privacy disclosure detection method significantly and consistently outperformed other state-of-the-art methods in terms of all key performance indicators. Zhanbo Liang, Jie Guo 0011, Weidong Qiu, Shujun Li 0001 |
Data Min. Knowl. Discov. | 3 |
| 2024 | P²FRPSI: Privacy-Preserving Feature Retrieved Private Set IntersectionabstractPrivate Set Intersection (PSI) protocols can securely compute the intersection of the private sets on the server and the client without revealing additional data. This work introduces the concept of Privacy-Preserving Feature Retrieved Private Set Intersection ($\mathsf {P^{2}FRPSI}$). In$\mathsf {P^{2}FRPSI}$protocols, the client can obtain the intersection that satisfies a given predicate without revealing the predicate and additional data. We formally define the$\mathsf {P^{2}FRPSI}$protocol, including its inputs, outputs, functionality, and security. To achieve the privacy guarantee in$\mathsf {P^{2}FRPSI}$protocols, a new two-party protocol is designed, namely Secure Secret Shared Retrieval ($\mathsf {S^{3}R}$), which can be used to securely determine whether each item on the server satisfies the predicate. We construct an$\mathsf {S^{3}R}$protocol and prove its security in the semi-honest model. On the basis of this, we design an efficient OT-based$\mathsf {P^{2}FRPSI}$protocol and an easy-to-implement DH-based$\mathsf {P^{2}FRPSI}$protocol and prove that they are secure in the semi-honest model. Our implementation shows that the OT-based$\mathsf {P^{2}FRPSI}$protocol can perform the matching for about 1000K items in 3.8 seconds with a single thread. Moreover, the DH-based$\mathsf {P^{2}FRPSI}$can perform the matching for about 7000K items in one hour with four threads, with communication totaling 1456 MB, while the OT-based$\mathsf {P^{2}FRPSI}$protocol requires 1673 MB. Guowei Ling, Fei Tang 0001, Chaochao Cai, Jinyong Shan, Haiyang Xue, Wulu Li, Peng Tang 0002, Xinyi Huang 0001, Weidong Qiu |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2023 | Automatic Search of Linear Structure: Applications to Keccak and Ascon
Huina Li, Guozhen Liu, Peng Tang 0002, Weidong Qiu |
Inscrypt (2) | 5 |
| 2023 | A Novel Deep Learning Framework for Interpretable Drug-Target Interaction Prediction with Attention and Multi-task Mechanism
Yubin Zheng, Peng Tang 0002, Weidong Qiu, Jie Guo 0011 |
DASFAA (4) | 3 |
| 2023 | Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and DisinformationabstractMis-and disinformation online have become a major societal problem as major sources of online harms of different kinds.One common form of mis-and disinformation is outof-context (OOC) information, where different pieces of information are falsely associated, e.g., a real image combined with a false textual caption or a misleading textual description.Although some past studies have attempted to defend against OOC mis-and disinformation through external evidence, they tend to disregard the role of different pieces of evidence with different stances.Motivated by the intuition that the stance of evidence represents a bias towards different detection results, we propose a stance extraction network (SEN) that can extract the stances of different pieces of multi-modal evidence in a unified framework.Moreover, we introduce a support-refutation score calculated based on the co-occurrence relations of named entities into the textual SEN.Extensive experiments on a public large-scale dataset demonstrated that our proposed method outperformed the state-ofthe-art baselines, with the best model achieving a performance gain of 3.2% in accuracy.* Corresponding co-authors 1 In the literature the terms "misinformation" and "disinformation" often have inconsistent definitions.In our work, we adopt the more established definitions by the United Nations (https://www.undp.org/eurasia/dis/ misinformation): misinformation refers to information that is false but not created with the intention of causing harm and disinformation to information that is false and deliberately created to cause harm.Our work can be applied to both mis-and disinformation, so we will mostly use the term "mis-/disinformation". Xin Yuan 0011, Jie Guo 0011, Weidong Qiu, Shujun Li 0001 |
EMNLP | 3 |
| 2023 | GenTC: Generative Transformer via Contrastive Learning for Receipt Information Extraction
Xinrui Deng, Kefan Ma, Kai Chen 0006, Jie Guo 0011, Weidong Qiu |
ICANN (6) | 6 |
| 2023 | RRecT: Chinese Text Recognition with Radical-Enhanced Recognition Transformer
Xinrui Deng, Kefan Ma, Kai Chen 0006, Jie Guo 0011, Weidong Qiu |
ICANN (6) | 6 |
| 2023 | LED: Label Correlation Enhanced Decoder for Multi-Label Text ClassificationabstractMulti-label text classification, which aims to predict the relevant labels for each given document, is one of the fundamental tasks of natural language processing. Recent studies have utilized Transformer, which embeds texts and class labels into a joint space to capture the label correlation. However, existing methods tend to take up extra input length and ignore the significance of taxonomic hierarchy. For this reason, we introduce a label correlation enhanced decoder (LED) for multi-label text classification. LED predicts the presence or absence of class labels in parallel with label representation and captures label correlation through multi-task learning. In addition, we propose a hierarchy-aware mask to capture the hierarchical dependency between labels. Comprehensive experiments on four benchmark datasets show that LED outperforms the state-of-the-art baselines. Detailed analysis validates the effectiveness of our proposed method. Kefan Ma, Xinrui Deng, Jie Guo 0011, Weidong Qiu |
ICASSP | 5 |
| 2023 | TDAE: Text Detection with Affinity Areas and Evolution Strategies
Kefan Ma, Kai Chen 0006, Jie Guo 0011, Weidong Qiu |
ICDAR (6) | 6 |
| 2023 | A Multi-Modal Approach for the Detection of Account Anonymity on Social Media PlatformsabstractAnonymous users on Twitter are highly related to network hazards such as cyberbullying, fraud, rumor spread, etc. Given the inefficient and expensive manual review approach, an automated method to predict account anonymity is urgently needed. However, there are three challenges when employing deep learning algorithms to detect account anonymity: the lack of open-source public datasets, tedious manual collection and labeling for training data, and abundant types of information that the existing single-modal methods are not competent for. Our approach includes an automated labeling system for data collection and a multi-modal method for account anonymity detection. The labeling system is based on name extraction and identity verification and, to ensure the quality of the dataset, web mining and face similarity measurement are applied. We establish a dataset containing 20133 accounts and make it available to the public. The multi-modal method exploits the visual, textual, and numerical features extracted with ViT, BERT, and MLP, and uses the Transformer and MLP for feature fusion. Experimental results prove the effectiveness of the feature fusion and the accuracy of classification with accuracy of 86.21% and 92.46%. Jie Guo 0011, Weidong Qiu |
IJCNN | 4 |
| 2023 | Privacy-preserving and fine-grained data sharing for resource-constrained healthcare CPS devicesabstractAbstract Medical cyber‐physical systems (CPS) provide the possibility for real‐time health monitoring of patients and flexible diagnostic services based on expert systems by collaboratively integrating and connecting various physical devices including sensors, terminals, and cloud infrastructure. However, the ubiquitous security threats in cyberspace have raised concerns about data security and user privacy. Although related works propose to protect data security and user privacy with cryptographic protocols, their heavy computational and storage overheads incur performance and battery life challenges for resource‐constrained devices in the healthcare CPS. This article proposes an energy‐saving and privacy‐preserving data sharing (ESPPDS) scheme to address the challenge. ESPPDS inherits the anonymous fine‐grained access control from attribute‐based encryption (ABE) while protecting data integrity and supporting efficient user revocation. We also eliminate the repetitive computations of ciphertext components by utilizing the online/ offline encryption technology, and design a subtle and secure trick to delegate the decryption operations to the edge device, thereby reducing the computational overheads of the resource‐constrained devices. We then show the security proof, and discuss the construction in the untrusted/ compromised server setting. The comparison and experiment indicate that ESPPDS is practical and more efficient than related schemes. Yangyang Bao, Weidong Qiu, Xiaochun Cheng |
Expert Syst. J. Knowl. Eng. | 2 |
| 2023 | Correction of whitespace and word segmentation in noisy Pashto text using CRF
Ijazul Haq, Weidong Qiu, Jie Guo 0011, Peng Tang 0002 |
Speech Commun. | 2 |
| 2023 | Solving Small Exponential ECDLP in EC-Based Additively Homomorphic Encryption and ApplicationsabstractAdditively Homomorphic Encryption (AHE) has been widely used in various applications, such as federated learning, blockchain, and online auctions. Elliptic Curve (EC) based AHE has the advantages of efficient encryption, homomorphic addition, scalar multiplication algorithms, and short ciphertext length. However, EC-based AHE schemes require solving a small exponential Elliptic Curve Discrete Logarithm Problem (ECDLP) when running the decryption algorithm, i.e., recovering the plaintext$m\in \{0,1\}^{\ell} $from$m \ast G$. Therefore, the decryption of EC-based AHE schemes is inefficient when the plaintext length$\ell > 32$. This leads to people being more inclined to use RSA-based AHE schemes rather than EC-based ones. This paper proposes an efficient algorithm called$\mathsf {FastECDLP}$for solving the small exponential ECDLP at 128-bit security level. We perform a series of deep optimizations from two points: computation and memory overhead. These optimizations ensure efficient decryption when the plaintext length$\ell $is as long as possible in practice. Moreover, we also provide a concrete implementation and apply$\mathsf {FastECDLP}$to some specific applications. Experimental results show that$\mathsf {FastECDLP}$is far faster than the previous works. For example, the decryption can be done in 0.35 ms with a single thread when$\ell = 40$, which is about 30 times faster than that of Paillier. Furthermore, we experiment with$\ell $from 27 to 54, and the existing works generally only consider$\ell \leq 32$. The decryption only requires 1 second with 16 threads when$\ell = 54$. In the practical applications, we can speed up model training of existing vertical federated learning frameworks by 4 to 14 times. At the same time, the decryption efficiency is accelerated by about 140 times in a blockchain financial system (ESORICS 2021) with the same memory overhead. Fei Tang 0001, Guowei Ling, Chaochao Cai, Jinyong Shan, Xuanqi Liu, Peng Tang 0002, Weidong Qiu |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2023 | Fine-Grained Data Sharing With Enhanced Privacy Protection and Dynamic Users Group Service for the IoVabstractThe Internet of Vehicles (IoV) is expected to play a revolutionary role in improving users’ driving experience and urban traffic governance. By widely absorbing emerging technologies including cloud computing, the future IoV evolution is leading towards providing more flexible and diversified data services. However, the publicly accessible IoV environment arouses the user’s concerns about the leakage of data and personal privacy. Despite some cryptographic solutions have been proposed, they still raise challenges on privacy, efficiency and usability. To cope with these challenges, this paper first presents an efficient scheme PH-ABE-DS, which attains the full policy hiding by implementing the access control with the inner product. Besides, we design an efficient indirect revocation mechanism, to enable the cloud and users to update the ciphertext and user secret key with slight storage and computational overheads. On this basis, we then present the EA-PH-ABE-DS scheme, by resorting to edge computing, it further reduces the overheads of resource-constrained devices. We design a deployment model for EA-PH-ABE-DS in IoV to discuss its usability. Rigorous security proof and security properties analysis show that our proposal is secure and reliable. Finally, through detailed comparisons on theoretical and experimental, both our two schemes show their superiority over the latest related works in terms of functionality and performance. The simulation evaluates and demonstrates the practicality of our solutions in practical IoT scenarios. Yangyang Bao, Weidong Qiu, Xiaochun Cheng, Jianfei Sun |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | GTRNet: a graph-based table reconstructed networkabstractTabular data, with an exceedingly effective data structure, can give us a more intuitive visual display. To under-stand well and make use of the spatial and logic dependencies of it, we propose an end-to-end, graph-based table reconstructed network, namely GTRNet, in this paper. Our model works differently from most existing models which treat tables as either a markup sequence problem or a graph structure of rows and columns. It can utilize a table as input and extract its features in the text, image and position coordinate to predict the dependencies of the text instances and well distinguish the spatial relationship to infer whether multiple text segments belong to the same merged cell. Optimized along with this network, we can then restore the structure of this table. Moreover, we also create a new Chinese benchmark dataset GraphTable for this task to tackle complex challenges on the table. The competitive results on ICDAR-2013, GraphTable, SciTSR and FinTab benchmarks further confirm the great effectiveness of GTRN et. Jie Guo 0011, Weidong Qiu |
ICTAI | 4 |
| 2022 | "Comments Matter and The More The Better!": Improving Rumor Detection with User CommentsabstractWhile many online platforms bring great benefits to their users by allowing user-generated content, they have also facilitated generation and spreading of harmful content such as rumors. Researcher have proposed different rumor detection methods based on features extracted from the original post and/or associated comments, but how comments affect the performance of such methods remains largely less understood. In this paper, we first propose a new BERT-based rumor detection method that can outperform other state-of-the-art methods, and then used it to study the role of comments in rumor detection. Our proposed method concatenates the original post and associated comments to form a single long text, which is then segmented into shorter chunks more suitable for BERT-based vectorization. Features extracted from all trunks are fed into a classifier based on an LSTM network or a transformer layer for the classification task. The experimental results on the PHEME and Ma-Weibo datasets proved the superior performance of our method. We conducted additional experiments on different settings of our proposed method to study different aspects of the role comments play in the rumor detection task. These additional experiments led to some very interesting findings, including the surprising result that fixed-length segmentation is better than natural segmentation, and the observation that including more comments can help improve the rumor detector’s performance. Some of these findings have profound operational implications for online platforms, e.g., commentators can contribute to rumor detection positively so online platforms can leverage the crowd intelligence to detect online rumors more effectively without applying overstrict content consensus policies. Jie Guo 0011, Weidong Qiu, Enes Altuncu, Shujun Li 0001 |
TrustCom | 3 |
| 2022 | Secure and Lightweight Fine-Grained Searchable Data Sharing for IoT-Oriented and Cloud-Assisted Smart Healthcare SystemabstractIt is a new trend in healthcare informatization construction to build the smart healthcare system by using the Internet of Things (IoT) and cloud. This IoT-oriented and cloud-assisted healthcare system enables the doctor to monitor the patient’s health state to respond to the paroxysmal diseases in real time. Considering the sensitivity of the patient’s privacy, it is necessary to encrypt the cloud-stored health data to prevent the semi-trusted cloud and unauthorized users from accessing them. However, the encrypted health data stored in the cloud brings inconvenience to the retrieval for the data user. In addition, the expensive computational consumption also raises the challenge to the resource-constrained devices in the patient and doctor sides. To support efficient ciphertext retrieval and cope with the performance challenge, in this article we propose a lightweight attribute-based searchable encryption (LABSE) scheme, which realizes fine-grained access control and keyword search, while reducing the computational overhead for the resource-constrained devices. We rigorously prove the semantic security of the proposed LABSE scheme, and analyze other security properties to response the security requirements under the healthcare scenario. Subsequently, we construct a concrete deployment model for LABSE under the healthcare system. We also compare LABSE with the state-of-the-art related schemes in terms of functionality and complexity. Finally, we demonstrate the practicality and performance advantages by the experiment. Yangyang Bao, Weidong Qiu, Xiaochun Cheng |
IEEE Internet Things J. | 2 |
| 2022 | Efficient, Revocable, and Privacy-Preserving Fine-Grained Data Sharing With Keyword Search for the Cloud-Assisted Medical IoT SystemabstractThe cloud-assisted medical Internet of Things (MIoT) has played a revolutionary role in promoting the quality of public medical services. However, the practical deployment of cloud-assisted MIoT in an open healthcare scenario raises the concern on data security and user's privacy. Despite endeavors by academic and industrial community to eliminate this concern by cryptographic methods, resource-constrained devices in MIoT may be subject to the heavy computational overheads of cryptographic computations. To address this issue, this paper proposes an efficient, revocable, privacy-preserving fine-grained data sharing with keyword search (ERPF-DS-KS) scheme, which realizes the efficient and fine-grained access control and ciphertext keyword search, and enables the flexible indirect revocation to malicious data users. A pseudo identity-based signature mechanism is designed to provide the data authenticity. We analyze the security properties of our proposed scheme, and via the theoretical comparison and experimental results we demonstrate that for the resource-constrained devices in the patient and doctor side of MIoT, in comparison with other related schemes, ERPF-DS-KS just consumes the lightweight and constant size communication/storage as well as computational time cost. For the keyword search, compared with related schemes, the cloud can quickly check whether a ciphertext contains the specified keyword with slight computations in the online phase. This further demonstrates that ERPF-DS-KS is efficient and practical in the cloud-assisted MIoT scenario. Yangyang Bao, Weidong Qiu, Peng Tang 0002, Xiaochun Cheng |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Scale Invariant Domain Generalization Image Recapture Detection
Jinian Luo, Jie Guo 0011, Weidong Qiu, Hong Hui |
ICONIP (4) | 3 |
| 2021 | Terroristic Content Detection using a Multi-scene classification systemabstractThe spread of terroristic images on the World Wide Web will bring significant risks to social security. Terroristic image detection technology can help image filtering. Otherwise, due to the lack of high-quality terrorist image dataset, deep learning based recognition methods have not been popularized. Furthermore, characteristics of occlusions and diversity of scenes, automatic approaches to terroristic content detection need to be well-designed. In this paper, a multi-model system is intended to detect for various types of terroristic content. For an input image, the system will locate and identify the terrorist organization leader or flag, determine whether the text information in the image belongs to terrorist slogan, and distinguish a terroristic picture from an ordinary one. For a terroristic image, the system will further detect sensitive objects, e.g. guns or flames, and subdivide the scene of brutal terrorist events. We build a terroristic dataset containing over 41,000 images for experimentation. The experiment result shows the accuracy of the multi-model terroristic content detection system. Jie Guo 0011, Weidong Qiu |
ICTAI | 3 |
| 2021 | Unsupervised Anomaly Detection for Time Series with Outlier ExposureabstractIt is of great practical significance to accurately model and analyze abnormal events in time series. For example, the identification of anomaly patterns on infrastructure sensor curves helps locate equipment failures. In this paper, we propose an unsupervised anomaly detection approach for time series, which can comprehensively consider both point anomalies and subsequence anomalies. We innovatively introduce RNN into the architecture of Adversarial Autoencoder to better analyze anomaly events based on overall relationship of time series. In addition, we innovatively apply the Outlier Exposure technique for the performance optimization of anomaly detector. Meanwhile, a WGAN-based method is utilized to generate anomaly datasets through normal distribution learning. Finally, we apply the proposed method for fraud detection on a financial statement dataset and intrusion detection on a network traffic dataset. Experimental results demonstrates that our model can comprehensively consider different anomaly types in time series, and achieve promising detection performance overall. In the experiment of fraud detection, the LSTM integrated AAE model achieves an F1 score of 0.810, while the Outlier Exposure enhanced model achieves an F1 score of 0.894. This indicates that our method can improve the performance of current audit systems and facilitate discovering malicious behaviors. Jiaming Feng, Jie Guo 0011, Weidong Qiu |
SSDBM | 4 |
| 2021 | Efficient and Fine-Grained Signature for IIoT With Resistance to Key ExposureabstractAttribute-based signature (ABS) can provide fine-grained and anonymous authentication for the Industrial Internet of Things (IIoT). However, key exposure and the computational performance bottleneck of resource-constrained devices raise a challenge to the IIoT-oriented ABS schemes. To confront the above challenge, this article proposes an intrusion-resilient server-aided ABS (IR-SA-ABS) scheme. This scheme designs to periodically update the signing key with the assistance of a helper device, and the signing key is refreshed for multiple rounds aperiodically within each time period under the supervision of the helper device. In this way, the system remains secure even if both the helper device and the signing device are compromised. During this process, we significantly reduce the computational overheads of resource-constrained devices by delegating all of the key update and refresh operations and most of the signature generation and verification operations to a powerful server. Subsequently, we present the rigorous security proof of IR-SA-ABS. The comparison and experiment demonstrate the efficiency and practicality of IR-SA-ABS in the IIoT scenario. Yangyang Bao, Weidong Qiu, Xiaochun Cheng |
IEEE Internet Things J. | 2 |
| 2020 | RLST: A Reinforcement Learning Approach to Scene Text Detection RefinementabstractWithin the research of scene text detection, some previous work has already achieved significant accuracy and efficiency. However, most of the work was generally done without considering about the implicit relationship between detection and eye movements. In this paper, we propose a new method for scene text detection especially for its refinement based on reinforcement learning. The idea of this method is inspired by Saccadic Eye Movements and Peripheral Vision. A saccade makes it possible for humans to orient the gaze to the location where a visual object has appeared. Peripheral vision gathers visual information of surroundings which provides supplement to foveal vision during gazing. We propose a simple pipeline, imitating the way human eyes do a saccade and collect peripheral information, to locate scene text roughly and to refine multi-scale vision field iteratively using reinforcement learning. For both training and evaluation, we use ICDAR2015 Challenge 4 dataset as a base and design several criteria to measure the feasibility of our work. Kai Chen 0006, Jie Guo 0011, Weidong Qiu |
ICPR | 5 |
| 2020 | Privacy-preserving spatial query protocol based on the Moore curve for location-based service
Huijuan Lian, Weidong Qiu, Jie Guo 0011, Peng Tang 0002 |
Comput. Secur. | 2 |
| 2020 | Anomaly detection in electronic invoice systems based on machine learning
Peng Tang 0002, Weidong Qiu, Huijuan Lian |
Inf. Sci. | 2 |
| 2020 | Detection of SQL injection based on artificial neural network
Peng Tang 0002, Weidong Qiu, Huijuan Lian, Guozhen Liu |
Knowl. Based Syst. | 2 |
| 2020 | Efficient and secure k-nearest neighbor query on outsourced data
Huijuan Lian, Weidong Qiu, Peng Tang 0002 |
Peer-to-Peer Netw. Appl. | 2 |
| 2019 | Integrating Coordinates with Context for Information Extraction in Document ImagesabstractInformation extraction from document collections is a fundamental and important step to understand, structure and analyze data. Many approaches with rules and deep learning based techniques have been applied on plain text, however, when it comes to document images, such demand still exists but becomes quite challenging without linguistic knowledge. In this paper, we propose an approach to extract required named entities (NEs) from document images by integrating the coordinate information from the detection and recognition stage into the contextual information of the BiLSTM-CRF model with an attention mechanism. We test this method on two real-world datasets. One is a Contract Dataset of Listed Companies, and the other is an Insurance Policy Dataset of our own. The result shows a combination of coordinates and context with attention leverages extraction in document images, opening up potential applications on such tasks. Yunrui Lian, Jie Guo 0011, Weidong Qiu |
ICDAR | 5 |
| 2019 | Wacnet: Word Segmentation Guided Characters Aggregation Net for Scene Text Spotting With Arbitrary ShapesabstractIn this paper, we propose an end-to-end trainable framework for scene text spotting which can handle text with arbitrary shapes. The proposed framework is called Word Segmentation Guided Characters Aggregation Net (WAC-Net), which consists of a shared convolutional backbone and two task-specific subnetworks. One subnetwork does word-level instance-aware segmentation (WSN) and the other does char-level detection and recognition (CDRN). The entire framework segments each word instance while detects and recognizes each character in one single forward pass. These two subnetworks are jointly trained by multi-task learning. At the inference stage, characters are aggregated into words guided by word instance segmentation results. Experiments are conducted on two datasets with arbitrary shapes, and the results demonstrate the effectiveness of the proposed method. Yuchen Dai, Kai Chen 0006, Jie Guo 0011, Weidong Qiu |
ICIP | 6 |
| 2019 | Encrypted data indexing for the secure outsourcing of spectral clustering
Bozhong Liu, Ling Chen 0006, Xingquan Zhu 0001, Weidong Qiu |
Knowl. Inf. Syst. | 4 |
| 2018 | SQL Injection Behavior Mining Based Deep Learning
Peng Tang 0002, Weidong Qiu, Huijuan Lian, Guozhen Liu |
ADMA | 2 |
| 2018 | Fused Text Segmentation Networks for Multi-oriented Scene Text DetectionabstractIn this paper, we introduce a novel end-end framework for multi-oriented scene text detection from an instance-aware semantic segmentation perspective. We present Fused Text Segmentation Networks, which combine multi-level features during the feature extracting as text instance may rely on finer feature expression compared to general objects. It detects and segments the text instance jointly and simultaneously, leveraging merits from both semantic segmentation task and region proposal based object detection task. Not involving any extra pipelines, our approach surpasses the current state of the art on multi-oriented scene text detection benchmarks: ICDAR2015 Incidental Scene Text and MSRA-TD500 reaching Hmean 84.1 % and 82.0 % respectively. Morever, we report a baseline on total-text containing curved text which suggests effectiveness of the proposed approach. Yuchen Dai, Youxuan Xu, Kai Chen 0006, Jie Guo 0011, Weidong Qiu |
ICPR | 7 |
| 2018 | Sparse coding with cross-view invariant dictionaries for person re-identification
Yunlu Xu, Jie Guo 0011, Weidong Qiu |
Multim. Tools Appl. | 4 |
| 2018 | Protecting lightweight block cipher implementation in mobile big data computing - A GPU-based approach
Weidong Qiu, Bozhong Liu, Can Ge, Lingzhi Xu, Xiaoming Tang, Guozhen Liu |
Peer-to-Peer Netw. Appl. | 1 |
| 2018 | A generic optimization method of multivariate systems on graphic processing units
Guohong Liao, Weidong Qiu |
Soft Comput. | 4 |
| 2017 | Protecting Location Privacy in Spatial Crowdsourcing using Encrypted Dataabstract© 2017, Copyright is with the authors. In spatial crowdsourcing, spatial tasks are outsourced to a set of workers in proximity of the task locations for efficient assignment. It usually requires workers to disclose their locations, which inevitably raises security concerns about the privacy of the workers’ locations. In this paper, we propose a secure SC framework based on encryption, which ensures that workers’ location information is never released to any party, yet the system can still assign tasks to workers situated in proximity of each task’s location. We solve the challenge of assigning tasks based on encrypted data using homomorphic encryption. Moreover, to overcome the efficiency issue, we propose a novel secure indexing technique with a newly devised SKD-tree to index encrypted worker locations. Experiments on real-world data evaluate various aspects of the performance of the proposed SC platform. Bozhong Liu, Ling Chen 0006, Xingquan Zhu 0001, Ying Zhang 0001, Chengqi Zhang, Weidong Qiu |
EDBT | 6 |
| 2016 | Mining Co-locations from Continuously Distributed Uncertain Spatial Data
Bozhong Liu, Ling Chen 0006, Chengqi Zhang, Weidong Qiu |
APWeb (1) | 5 |
| 2016 | Biclique cryptanalysis using balanced complete bipartite subgraphs
Shusheng Liu, Yamin Wen, Yiyuan Luo, Weidong Qiu |
Sci. China Inf. Sci. | 5 |
| 2016 | Optimizations for High Performance Network Virtualization
Fanfu Zhou, Ruhui Ma, Jian Li 0021, Li-Xia Chen, Weidong Qiu, Haibing Guan |
J. Comput. Sci. Technol. | 5 |
| 2015 | RCP Mining: Towards the Summarization of Spatial Co-location Patterns
Bozhong Liu, Ling Chen 0006, Chengqi Zhang, Weidong Qiu |
SSTD | 5 |
| 2014 | A New Approach to Multimedia Files CarvingabstractTraditional file recovery methods rely on file system information, which are ineffective when file system information isn't available. File carving is a file recovery method that recovers files according to their structure and content without file system information, which is widely used in digital forensics. As the important carriers of digital information, multimedia files are important digital evidence. In this paper, a new multimedia file carving approach is proposed to improve the recovery accuracy of high entropy file fragments. The fragmented files can be recovered by a hierarchical carving process, including file header identification via entropy, file fragment type classification, and file reassembly via parallel unique path approach. A new file type classification method is constructed based on support vector machine, by using the features of BFD (byte frequency distribution) and ROC (rate of change). Four different datasets, such as DFRWS 2006/2007 challenge datasets, dataset simulating actual disk, dataset with randomly disordered fragments, and dataset with biomedical images, are employed in our experiments. The results show that JPEG recovery accuracy is improved greatly compared with that of Photo Rec tool. Our method performs best in the situation where the order of fragments is completely confusing. Weidong Qiu, Run Zhu, Jie Guo 0011, Xiaoming Tang, Bozhong Liu |
BIBE | 1 |
| 2011 | On the Security of 4-Bit Involutive S-Boxes for Lightweight Designs
Bozhong Liu, Weidong Qiu, Dong Zheng 0001 |
ISPEC | 3 |
| 2011 | Restrictive partially blind signature for resource-constrained information systems
Weidong Qiu, Bozhong Liu, Yu Long 0001, Kefei Chen |
Knowl. Inf. Syst. | 1 |
| 2008 | Identity-Based Threshold Key-Insulated Encryption without Random Oracles
Jian Weng 0001, Shengli Liu 0001, Kefei Chen, Dong Zheng 0001, Weidong Qiu |
CT-RSA | 5 |
| 2008 | Bagging very weak learners with lazy local learningabstractBagging predictors have shown to be effective especially when the learners used to train the base classifiers are weak. In this paper, we argue that for very weak (VW) learners, such as DecisionStump, OneR, and SuperPipes, the base classifiers built from boostrap bags are strongly correlated with each other. As a result, a simple bagging (SB) predictor built on such VW learners has very little improvement compared to a single classifier trained from the same data. Alternatively, we propose a Local Lazy Learning based bagging approach (L3B), where base learners are trained from a small instance subset surrounding each test instance. More specifically, given a test instance x, L3B first discovers x¿s k nearest neighbours, and then applies progressive sampling to the selected neighbours to train a set of base classifiers, by using a given VW learner. At the last stage, x is labeled as the most frequently voted class of all base classifiers. Experimental results on 32 real-world datasets, including two high dimensional gene expression datasets, demonstrate that L3B significantly outperforms SB for building accurate classifier ensemble models for VW learners. Xingquan Zhu 0001, Chengyi Bao, Weidong Qiu |
ICPR | 3 |
| 2007 | Identity-Based Threshold Decryption Revisited
Shengli Liu 0001, Kefei Chen, Weidong Qiu |
ISPEC | 3 |
| 2005 | Fast on-line real-time scheduling algorithm for reconfigurable computingabstractPartially reconfigurable devices are able to execute several tasks in parallel and allow for on-line reconfiguration. To manage such devices at runtime, the scheduler in the operating system has two more modules for hardware tasks: placer and loader. Placer is to find appropriate places in reconfigurable devices and loader is to do the reconfiguration for hardware tasks. In order to satisfy the timing constraints in hard real-time systems, we propose a fast online real-time scheduling algorithm. Our algorithm will be based on correct empty resource management and will utilize the reconfiguration reuse. The experiments show that the developed scheduler leads to substantial performance gains. Weidong Qiu, Chenglian Peng |
CSCWD (2) | 1 |
| 2005 | SHUM-uCOS: A RTOS using multi-task model to reduce migration cost between SW/HW tasksabstractThe design of embedded systems has become more complex than ever, and the design qualities depend more on the cooperation of multidisciplinary design teams: hardware engineers and software engineers in general. However, due to the lack of uniform programming model and system components for these different teams, the migrations costs of a function model from software to hardware are high. But these actions are necessary in the hardware-software partitioning of embedded systems, especially in the prototype designs. To cope with this problem, we adopt a uniform multi-task model and implement a RTOS (real-time operating system), called SHUM-uCOS, which deals with hardware functions as same as software tasks. This RTOS uses uCOSII as its prototype, traces and manages the states of reconfigurable resources (FPGAs), which allows the execution of hardware tasks in a true multitasking manner. Moreover, SHUM-uCOS also defines a standard hardware-task interface, which supports share-bus protocol. It has been proved by experiments that SHUM-uCOS can shorten the migration time from software implementations to hardware implementations with the performance improvement. Weidong Qiu, Chenglian Peng |
CSCWD (2) | 2 |
| 2002 | How to Play Sherlock Holmes in the World of Mobile Agents
Biljana Cubaleska, Weidong Qiu, Markus Schneider 0002 |
ACISP | 2 |
| 2002 | Component Priority Assignment in tha Data Flow Dominated Embedded Systems with Timing ConstraintsabstractDataflow dominated embedded systems often use data flow graphs (DFG) as system models. To achieve the desired performance, these systems usually contain a lot of hardware/software components working in parallel. These concurrent and cooperative components result in the contentions for shared resources due to architecture and data dependencies. The approach to solve the contentions can be priority assignments. In this paper we introduce an algorithm which can find out a priority assignment for a given set of components working in parallel with a timing constraint. In addition, the algorithm also provides a fast way to calculate, whether a set of components working in parallel can guarantee a given timing constraint. Hence the algorithm can be applied both in designing phase and implementation phase of hardware/software co-design for embedded systems. Baifeng Wu, Chenglian Peng, Weidong Qiu, Xiaoguang Sun |
CSCWD | 3 |
| 2002 | A New Offline Privacy Protecting E-cash System with Revokable Anonymity
Weidong Qiu, Kefei Chen, Dawu Gu |
ISC | 1 |