Jinqiao Shi

dblp:53/10661 · DBLP profile ↗
← Back
72ranked-venue papers
4as first author
32since 2021 · last 2026
0000-0002-0111-1023ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 20 · 5 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 7 since 2021Security and privacy · 16 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 9 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Systems, architecture and hardware · 5 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Good Gradients Poison Your Model: Evading Defenses in Federated Learning via Boundary-adaptive Perturbation
abstract
Federated learning (FL) allows for collaborative model training while preserving data privacy, but its distributed nature makes it vulnerable to poisoning attacks. Existing defense methods typically rely on using gradients from multiple clients to define a trusted region, selecting only the trustworthy update (good gradients) within this region for aggregation. Mainstream defense boundaries are categorized as hard boundaries, soft boundaries, and semi-soft boundaries. However, we argue that even good gradients within these boundaries can still be exploited by attackers to poison the model. To tackle this challenge, we introduce a boundary-adaptive attack method that leverages the directional properties of optimization techniques to derive baseline poisoned gradients. Through iterative perturbation, it generates seemingly innocent gradients that subtly deviate from the global model. Our extensive study on benchmark datasets and mainstream defensive mechanisms confirms that the proposed attack raises a significantly threat to the integrity and security of FL practices, regardless of the flourishing of robust FL methods.
Jinqiao Shi, Junmin Huang, Chongru Fan
AAAI2
2026 Mitigating Cumulative Privacy Risk in Continual Information Sharing: A Dynamic Stackelberg Game Approach
abstract
Privacy leakage on Web-based platform has become a critical challenge as users continually share personal information through online social networks, health tracking platforms, and other Web services. Information recipients and third-party can progressively aggregate shared content across the Web, enabling increasingly accurate profiling of individuals. However, existing studies typically treat each disclosure as independent, overlooking the cumulative privacy risks that arise in continual information sharing. In addition, the subjective cognition of both users and adversaries in Web environments, where users and adversaries can dynamically adapt based on observable actions, remains underexplored. To address these challenges, we propose a dynamic Stackelberg game model for continual information sharing scenarios, where user's sequential privacy decisions are optimized to balance privacy protection and data utility. The model explicitly captures the cognitive behaviors of both the user and the adversary, allowing their subjective perceptions to shape the Stackelberg equilibrium. Building on this formulation, we develop a reinforcement learning-based algorithm to derive approximately optimal strategies for mitigating privacy leakage in the context of continual information sharing. Experiments on real-world datasets demonstrate that our method significantly reduces cumulative privacy risks while preserving the utility of shared content. The proposed model further provides actionable insights for the design of privacy-enhancing technologies and web platform policies.
Yuzi Yi, Yehong Luo, Jinqiao Shi, Jiwei Huang
WWW4
2026 Uncovering malicious sybil structures in Tor: A hybrid approach of constraint clustering and dynamic scoring
Cheng Huang 0003, Beining Zhang, Chang You, Jinqiao Shi
Comput. Networks6
2026 Vulnerabilities of label and model protection in vertical federated learning based on simulation models
Zhenhan Ke, Hongfan Chen, Hui Xiang, Jinqiao Shi
Neurocomputing9
2026 MUFFIN: A Meta-Knowledge Decoupling-Based Approach to Few-Shot IoT Traffic Classification
abstract
Traditional machine learning-based approaches to Internet of Things (IoT) traffic classification are hindered by the need for a large number of labeled flows in real-world network environments. To cope with this problem, some few-shot traffic classification approaches have been proposed to achieve accurate classification by comparing the similarity of flows to the given labeled flows. These approaches train neural networks to concurrently learn feature extraction and feature comparison meta-knowledge from sufficient labeled flows. However, they conflate these two types of meta-knowledge and ignore their differences. We propose MUFFIN, a meta-knowledge decoupling-based approach to few-shot IoT traffic classification. MUFFIN is based on our key insight that feature extraction and feature comparison meta-knowledge are two different types of meta-knowledge with distinct learning goals. Therefore, MUFFIN decouples their learning in the meta-knowledge phase and introduces a masked autoencoder as its feature extractor to enhance the learning of feature extraction meta-knowledge. Extensive experiments conducted on three publicly available datasets demonstrate the classification ability and robustness of MUFFIN. Furthermore, a comparison with five state-of-the-art few-shot traffic classification approaches reveals that MUFFIN outperforms them in classification accuracy and requires only 20 labeled flows.
Yipeng Wang 0001, Yingxu Lai, Shui Yu 0001, Jinqiao Shi
IEEE Trans. Netw.5
2025 Quadruplet Fingerprinting: Onion Website Fingerprinting Through Quadruplet Network
abstract
The onion service is designed to achieve anonymity in the Tor network. However, it is highly vulnerable to website fingerprinting (WF) attacks because websites have unique traffic patterns that can be identified. Previous server-side WF attacks have certain limitations. There is an accuracy gap of nearly 10% when comparing with the ideal modeling using server-side traces. To address this issue, we propose a remote fingerprinting attack based on quadruplet networks, named Quadruplet Fingerprinting (QF). We utilize the RP node as the remote node and construct a feature extractor with quadruplet loss. This approach enhances the generalization ability of the model. The experimental results demonstrate that higher accuracy has been achieved in both common and few-shot WF scenarios compared to the existing methods. And most importantly, the accuracy rate of the RP side has been improved by more than 10% at the client side.
Can Zhao 0005, Yefeng Qin, Qingyun Liu 0001, Jinqiao Shi
CSCWD6
2025 PromptAL: Sample-aware dynamic soft prompts for few-shot active learning
Hui Xiang, Jinqiao Shi, Yong Ma 0004
Knowl. Based Syst.2
2025 Proxied Traffic Fingerprinting for Hidden Service De-Anonymization With Burst Reshaping
abstract
Traffic fingerprinting attack is a promising approach for Tor hidden services (HS) de-anonymization. However, it is inherently difficult to acquire traffic of target HSs (HST) for fingerprinting model training, because the physical location of the services is hidden due to the design of Tor protocol. In order to solve this problem, some alternatives such as mirrored HST (MHST) and client-side HST (CHST) have been proposed for training fingerprinting model. These alternatives are easy to acquire and aim to closely match the characteristics of the target HST. However, they cannot perfectly replace the target HST for the aspects of consistency of both response and protocol. In this paper, we propose a proxied fingerprinting approach called PF. A Proxy HS is deployed to acquire proxied HS traffic (PHST) as an alternative to conduct traffic fingerprinting attack, which satisfies both response and protocol consistency and is easy to acquire. In order to mitigate the impact introduced by Proxy HS, PF also introduces Burst Reshaping which includes burst reconstruction and pseudo-label learning to enhance the similarities between PHST and target HST. Experiments show that, PHST is a superior alternative to target HST, fingerprinting model trained using PF achieved an accuracy of 92.2%, surpassing the models trained with MHST and CHST by 72% and 34%, respectively. Additionally, PF is an add-on approach capable of improving the HS deanonymization effectiveness of any fingerprinting model architecture. The source code and dataset are available at https://github.com/Lzreal/BurstReshapedPHST.
Yipeng Wang 0001, Haoting Liu, Jiapeng Zhao, Jinqiao Shi
IEEE Trans. Inf. Forensics Secur.6
2024 Seeing the Attack Paths: Improved Flow Correlation Scheme in Stepping-Stone Intrusion
abstract
Stepping-stones are widely used by attackers to conceal their identities and gain unauthorized access to restricted targets. Numerous strategies have been suggested to identify stepping-stones and counteract evasive behaviors, with flow correlation standing out as the key technique used in many deanonymization methods. Existing attempts to tackle the flow correlation issue rely on long-flow observations and feature-based methods. Yet, these approaches meet substantial limitations in their suitability across diverse scenarios, network noises influence and accuracy, particularly in the stepping-stone environment. In this paper, we introduce an improved flow correlation model, FlowLinker, which aims to resolve these issues. Specifically, we combined multidimensional statistical features and cumulative flow sequences to construct a robust traffic feature. Then, we leveraged the triplet network to produce an optimized representation, amplifying the difference between unrelated representations. Consequently, it significantly reduces the false correlation rate with flows that seem similar but are unrelated. The experiments on real-world datasets from different network environments we collected show that FlowLinker outperforms other state-of-the-art methods with shorter lengths of flow observations.
Chao Zheng 0001, Zhao Li 0010, Jinqiao Shi
CSCWD4
2024 AlterCell Attack: Exploiting a Logic Vulnerability in Tor Cell Integrity Validation
abstract
The hidden service is used to protect the anonymity of receivers in the Tor network, but it is often exploited for malicious purposes and becomes a breeding ground for crime. To protect the hidden service from abuse, we deeply analysed the Tor specifications and summarised a finite-state machine for Tor cell transmission protocol, deriving a new logical vulnerability of the Tor cell integrity validation. This vulnerability allowed the plaintext field of cell_command in the cell to be arbitrarily modified. Therefore, we proposed the AlterCell attack against potentially malicious hidden services, by introducing malformed cells whose fields of cell_command were modified. In the experiments, when our node acted as the guard for the target hidden service, the AlterCell attack was able to de-anonymise the hidden service if the responsible HSDir was under our control. Otherwise, our attack could block the descriptor publishing process or the client access process, to carry out a denial-of-service attack on the hidden service.
Can Zhao 0005, Baiwei Duan, Qingyun Liu 0001, Jinqiao Shi
HPCC6
2024 A Comprehensive Evaluation of the Impact on Tor Network Anonymity Caused by ShadowBridge
Baiwei Duan, Yujia Zhu, Can Zhao 0005, Jinqiao Shi
SecureComm (4)7
2024 Risk Assessment Based on Dataflow Dynamic Hypergraph for Cross-Border Data Transfer
abstract
In the era of big data, information security, particularly during data transmission, has become a top priority. However, existing risk assessments are often limited to local perspectives, missing a global view of data-related risks and hidden relationships. This paper introduces the Dataflow Dynamic Hypergraph (DDHG) framework, which abstracts data flow into a dynamic hypergraph, incorporating key features to better describe the data transfer process. To fully capture this information, we propose the Node-Edge-Hyperedge Autoencoder Graph Network (NEH-AGN), which frames risk assessment as a graph anomaly detection problem, identifying anomalous nodes for risk evaluation. The paper also explores future research directions in data flow scenarios, offering a reference for future work.
Yigang Diao, Jinqiao Shi
TrustCom6
2024 Knock-Knock: De-Anonymise Hidden Services by Exploiting Service Answer Vulnerability
abstract
The hidden service, devised by the Tor Project, serves to protect receiver anonymity. However, to address the potential abuse of hidden services, this paper introduces the “Knock-Knock attack,” a novel de-anonymization method that utilizes watermark to enable an attacker to conduct an attack with control over only the client and the guard relays. The key of our approach is manipulating the number of RELAY_COMMAND_BEGIN cells and RELAY_COMMAND_CONNECTED cells to construct flexible and robust watermarks during the Tor protocol handshake, thus facilitating parallel de-anonymization of hidden services while remaining insensitive to the network state. Empirical experiments demonstrate that our attack boasts a 100% true positive rate and 0% false positive rate. Additionally, we propose a theoretical framework to guide the optimal encoding form of watermark, leading to a notable 3.6 times improvement in speed compared to prior works. Lastly, we present a method to mitigate watermark attacks and report the design flaw to the Tor Project.
Muqian Chen, Can Zhao 0005, Qingyun Liu 0001, Jinqiao Shi
WCNC6
2024 HSDirSniper: A New Attack Exploiting Vulnerabilities in Tor's Hidden Service Directories
Zhiyang Teng, Yue Gao 0003, Qingyun Liu 0001, Jinqiao Shi
WWW6
2024 SPEFL: Efficient Security and Privacy-Enhanced Federated Learning Against Poisoning Attacks
abstract
Federated learning (FL) is a distributed machine learning paradigm in the Internet of Things (IoT), which allows multiple devices to collaboratively train models without leaking local data. In the open scenario of IoT, malicious devices can launch poisoning attacks to compromise the final model by submitting crafted gradients. Some previous studies defend against poisoning attacks by analyzing the statistical characteristics of plaintext gradients. However, plaintext gradients would expose private information to malicious FL devices or servers. To simultaneously resist poisoning attacks and preserve privacy, cryptography technology can be utilized to obfuscate the gradients in defense methods, but the private calculation of resisting poisoning attack methods will cause efficiency problems, especially imposing unaffordable overhead on resource-limited IoT devices. Therefore, resisting poisoning attacks efficiently while protecting privacy remains a challenge. This paper proposes a secure and privacy-enhanced FL (SPEFL) framework for efficient privacy-preserving and poisoning-resistant federated learning in IoT. We design an efficient secure computation protocol based on a three-server architecture to facilitate the cryptographic computation of large linear and complex nonlinear operators in the method against poisoning attacks. In SPEFL, most of the calculations are efficiently performed on the servers, which will not impose too much burden on resource-limited IoT devices. In addition, we design a security-enhanced verifiable protocol to detect the malicious behavior of the server and guarantee the correctness of FL aggregation results. Experimental and theoretical results demonstrate that SPEFL can efficiently complete FL training meanwhile guaranteeing the accuracy of the model.
Liyan Shen, Zhenhan Ke, Jinqiao Shi, Xi Zhang 0008, Jiapeng Zhao
IEEE Internet Things J.3
2024 Label-Aware Chinese Event Detection with Heterogeneous Graph Attention Network
Shiyao Cui, Xin Cong, Tingwen Liu, Qingfeng Tan, Jinqiao Shi
J. Comput. Sci. Technol.6
2024 Enhancing Multimodal Entity and Relation Extraction With Variational Information Bottleneck
abstract
This paper studies the multimodal named entity recognition (MNER) and multimodal relation extraction (MRE), which are important for content analysis and various applications. The core of MNER and MRE lies in incorporating evident visual information to enhance textual semantics, where two issues inherently demand investigations. The first issue is modality-noise, where the task-irrelevant information in each modality may be noises misleading the task prediction. The second issue is modality-gap, where representations from different modalities are inconsistent, preventing from building the semantic alignment between the text and image. To address these issues, we propose a novel method for MNER and MRE byMultiModal representation learning withInformationBottleneck (MMIB). For the first issue, a refinement-regularizer probes the information-bottleneck principle to balance the predictive evidence and noisy information, yielding expressive representations for prediction. For the second issue, an alignment-regularizer is proposed, where a mutual information-based item works in a contrastive manner to regularize the consistent text-image representations. To our best knowledge, we are the first to explore variational IB estimation for MNER and MRE. Experiments show that MMIB achieves the state-of-the-art performances on three public benchmarks.
Shiyao Cui, Jiangxia Cao, Xin Cong, Jiawei Sheng, Quangang Li, Tingwen Liu, Jinqiao Shi
IEEE ACM Trans. Audio Speech Lang. Process.7
2024 Event-Based Monocular Depth Estimation With Recurrent Transformers
abstract
Event cameras, offering high temporal resolutions and high dynamic ranges, have brought a new perspective to address common challenges in monocular depth estimation (e.g., motion blur and low light). However, existing CNN-based methods insufficiently exploit global spatial information from asynchronous events, while RNN-based methods show a limited capacity for effective temporal cues utilization for event-based monocular depth estimation. To this end, we propose a event-based monocular depth estimator with recurrent transformers, namely EReFormer. Technically, we first design a transformer-based encoder-decoder that utilizes multi-scale features to model global spatial information from events. Then, we propose a Gate Recurrent Vision Transformer (GRViT), introducing a recursive mechanism into transformers, to leverage rich temporal cues from events. Finally, we present a Cross Attention-guided Skip Connection (CASC), performing cross attention to fuse multi-scale features, to improve global spatial modeling capabilities. The experimental results show that our EReFormer outperforms state-of-the-art methods by a margin on both synthetic and real-world datasets. Our open-source code is available at https://github.com/liuxu0303/EReFormer.
Xu Liu 0006, Jianing Li 0001, Jinqiao Shi, Xiaopeng Fan 0001, Yonghong Tian 0001, Debin Zhao
IEEE Trans. Circuits Syst. Video Technol.3
2023 URM4DMU: An User Representation Model for Darknet Markets Users
abstract
Darknet markets provide a large platform for trading illicit goods and services due to their anonymity. Learning an invariant representation of each user based on their posts on different markets makes it easy to aggregate user information across different platforms, which helps identify anonymous users. Traditional user representation methods mainly rely on modeling the text information of posts and cannot capture the temporal content and the forum interaction of posts. While recent works mainly use CNN to model the text information of posts, failing to effectively model posts whose length changes frequently in an episode. To address the above problems, we propose a model named URM4DMU(User Representation Model for Darknet Markets Users) which mainly improves the post representation by augmenting convolutional operators and self-attention with an adaptive gate mechanism. It performs much better when combined with the temporal content and the forum interaction of posts. We demonstrate the effectiveness of URM4DMU on four darknet markets. The average improvements on MRR value and Recall@10 are 22.5% and 25.5% over the state-of-the-art method respectively.
Hongmeng Liu, Jiapeng Zhao, Yixuan Huo, Chun Liao, Liyan Shen, Shiyao Cui, Jinqiao Shi
ICASSP8
2023 A Comprehensive Evaluation of the Impact on Tor Network Anonymity Caused by ShadowRelay
abstract
As a distributed anonymous network run by volunteers, Tor relays are often manipulated by operators to achieve their goals. Our work reveals that some relays, named ShadowRelay, are bound to hidden nodes and actively forward user traffic to the next-hop relay or target without the user's knowledge. To detect ShadowRelays, we developed HiddenSniffer based on client and Tor relay collusion, and found 162 hidden nodes distributed across 22 countries, along with 85 Shadow Relays which account for 2.08% of the total relay bandwidth. Additionally, there exists a family relationship among the Shadow Relays, with the largest family containing 24 members. The experimental results indicate that ShadowRelays have increased the number of ASes capable of sniffing user traffic by 27.6%, and improved the ability of 14.7% of attackers to launch traffic confirmation attacks. Furthermore, ShadowRelays adversely impact the Tor network's availability by introducing increased transmission delay within the circuits.
Muqian Chen, Qingyun Liu 0001, Jinqiao Shi
ISCC6
2023 Reservoir Computing Transformer for Image-Text Retrieval
abstract
Although the attention mechanism in transformers has proven successful in image-text retrieval tasks, most transformer models suffer from a large number of parameters. Inspired by brain circuits that process information with recurrent connected neurons, we propose a novel Reservoir Computing Transformer Reasoning Network (RCTRN) for image-text retrieval. The proposed RCTRN employs a two-step strategy to focus on feature representation and data distribution of different modalities respectively. Specifically, we send visual and textual features through a unified meshed reasoning module, which encodes multi-level feature relationships with prior knowledge and aggregates the complementary outputs in a more effective way. The reservoir reasoning network is proposed to optimize memory connections between features at different stages and address the data distribution mismatch problem introduced by the unified scheme. To investigate the significance of the low power dissipation and low bandwidth characteristics of RRN in practical scenarios, we deployed the model in the wireless transmission system, demonstrating that RRN's optimization of data structures also has a certain robustness against channel noise. Extensive experiments on two benchmark datasets, Flickr30K and MS-COCO, demonstrate the superiority of RCTRN in terms of performance and low-power dissipation compared to state-of-the-art baselines.
Wenrui Li 0001, Zhengyu Ma, Liang-Jian Deng, Penghong Wang, Jinqiao Shi, Xiaopeng Fan 0001
ACM Multimedia5
2023 Deanonymize Tor Hidden Services Using Remote Website Fingerprinting
abstract
Website Fingerprinting (WF) is commonly used to deanonymize the Tor users. Meanwhile, WF can also be applied to the hidden service (HS) side, to destroy the anonymity of the HS operator. However, the attacker faces the problem of no available data for training. To solve this problem, previous researchers build mirror sites and use the data collected on them to train the classifier. However, this approach still has limitations: The attacker cannot tell whether the target HS is some HS that is not in the set of sites he mirrors, i.e, an HS in the wild. To address this, we propose a novel two-phase approach to deanonymize HSs that using the website fingerprint collected on the remote node, i.e., a node on the same RP circuit with HS. Our approach does not rely on mirror sites because we decouple feature extraction from classification, so that mirror sites are only involved in the training of the feature extractor and are not related to the classification process. The experimental results show that our attack is effective in both closed-world and open-world scenarios.
Muqian Chen, Jinqiao Shi, Binxing Fang
TrustCom5
2023 An efficient confidentiality protection solution for pub/sub system
abstract
Abstract Publish/subscribe(pub/sub) systems are widely used in large-scale messaging systems due to their asynchronous and decoupled nature. With the population of pub/sub cloud services, the privacy protection problem of pub/sub systems has started to emerge, and events and subscriptions are exposed when executing event matching on untrustworthy cloud brokers. However, as the number of subscriptions increases, the effectiveness of the previous confidentiality protection approaches declines drastically. In this paper, we propose SBM (scalable blind matching), an effective confidentiality protection scheme for pub/sub systems. To the best of our knowledge, SBM is the first scheme that applies order-preserving encryption algorithm to protect the system’s confidentiality and ensure its scalability. In this scheme, SBM-I is highly effective in subscription matching but is unable to achieve ideal security IND-OCPA, whereas SBM-II is suggested to ensure system security and SGX is used to reduce interaction and boost ciphertext matching performance. The experiment demonstrates that this method has better matching performance compared to others: the average matching time of SBM-I is 3–4 orders of magnitude faster than the matching algorithm MP and SGX-based algorithm SCBR when the number of subscriptions is 500,000, and the average matching time of SBM-II is 40 times faster than MP and 24 times than SCBR.
Jinglei Pei, Qingling Feng, Ruisheng Shi, Lina Lan, Shui Yu 0001, Jinqiao Shi, Zhaofeng Ma
Cybersecur.7
2023 Evicting and filling attack for linking multiple network addresses of Bitcoin nodes
abstract
Abstract Bitcoin is a decentralized P2P cryptocurrency. It supports users to use pseudonyms instead of network addresses to send and receive transactions at the data layer, hiding users’ real network identities. Traditional transaction tracing attack cuts through the network layer to directly associate each transaction with the network address that issued it, thus revealing the sender’s network identity. But this attack can be mitigated by Bitcoin’s network layer privacy protections. Since Bitcoin protects the unlinkability of Bitcoin addresses and there may be a many-to-one relationship between addresses and nodes, transactions sent from the same node via different addresses are seen as coming from different nodes because attackers can only use addresses as node identifiers. In this paper, we proposed the evicting and filling attack to expose the correlations between addresses and cluster transactions sent from different addresses of the same node. The attack exploited the unisolation of Bitcoin’s incoming connection processing mechanism. In particular, an attacker can utilize the shared connection pool and deterministic connection eviction strategy to infer the correlation between incoming and evicting connections, as well as the correlation between releasing and filling connections. Based on inferred results, different addresses of the same node with these connections can be linked together, whether they are of the same or different network types. We designed a multi-step attack procedure, and set reasonable attack parameters through analyzing the factors that affect the attack efficiency and accuracy. We mounted this attack on both our self-run nodes and multi-address nodes in real Bitcoin network, achieving an average accuracy of 96.9% and 82%, respectively. Furthermore, we found that the attack is also applicable to Zcash, Litecoin, Dogecoin, Bitcoin Cash, and Dash. We analyzed the cost of network-wide attacks, the application scenario, and proposed countermeasures of this attack.
Huashuang Yang, Jinqiao Shi, Yue Gao 0003, Ruisheng Shi, Dongbin Wang
Cybersecur.2
2023 The Style Transformer With Common Knowledge Optimization for Image-Text Retrieval
abstract
Image-text retrieval which associates different modalities has drawn broad attention due to its excellent research value and broad real-world application. However, most of the existing methods haven't taken the high-level semantic relationships (“style embedding”) and common knowledge from multi-modalities into full consideration. To this end, we introduce a novel style transformer network with common knowledge optimization (CKSTN) for image-text retrieval. The main module is the common knowledge adaptor (CKA) with both the style embedding extractor (SEE) and the common knowledge optimization (CKO) modules. Specifically, the SEE uses the sequential update strategy to effectively connect the features of different stages in SEE. The CKO module is introduced to dynamically capture the latent concepts of common knowledge from different modalities. Besides, to get generalized temporal common knowledge, we propose a sequential update strategy to effectively integrate the features of different layers in SEE with previous common feature units. CKSTN demonstrates the superiorities of the state-of-the-art methods in image-text retrieval on MSCOCO and Flickr30 K datasets. Moreover, CKSTN is constructed based on the lightweight transformer which is more convenient and practical for the application of real scenes, due to the better performance and lower parameters.
Wenrui Li 0001, Zhengyu Ma, Jinqiao Shi, Xiaopeng Fan 0001
IEEE Signal Process. Lett.3
2022 Event Causality Extraction with Event Argument Correlations
abstract
Event Causality Identification (ECI), which aims to detect whether a causality relation exists between two given textual events, is an important task for event causality understanding. However, the ECI task ignores crucial event structure and cause-effect causality component information, making it struggle for downstream applications. In this paper, we introduce a novel task, namely Event Causality Extraction (ECE), aiming to extract the cause-effect event causality pairs with their structured event information from plain texts. The ECE task is more challenging since each event can contain multiple event arguments, posing fine-grained correlations between events to decide the cause-effect event pair. Hence, we propose a method with a dual grid tagging scheme to capture the intra- and inter-event argument correlations for ECE. Further, we devise a event type-enhanced model architecture to realize the dual grid tagging scheme. Experiments demonstrate the effectiveness of our method, and extensive analyses point out several future directions for ECE.
Shiyao Cui, Jiawei Sheng, Xin Cong, Quangang Li, Tingwen Liu, Jinqiao Shi
COLING6
2022 ABNN2: secure two-party arbitrary-bitwidth quantized neural network predictions
abstract
Data privacy and security issues are preventing a lot of potential on-cloud machine learning as services from happening. In the recent past, secure multi-party computation (MPC) has been used to achieve the secure neural network predictions, guaranteeing the privacy of data. However, the cost of the existing two-party solutions is expensive and they are impractical in real-world setting.
Liyan Shen, Ye Dong, Binxing Fang, Jinqiao Shi, Shengli Pan 0001, Ruisheng Shi
DAC4
2022 Document-Level Event Extraction via Human-Like Reading Process
abstract
Document-level Event Extraction (DEE) is particularly tricky due to the two challenges it poses: scattering-arguments and multi-events. The first challenge means that arguments of one event record could reside in different sentences in the document, while the second one reflects that one document may simultaneously contain multiple such event records. Motivated by humans’ reading cognitive to extract information of interests, in this paper, we propose a method called HRE (Human Reading inspired Extractor for Document Events), where DEE is decomposed into these two iterative stages, rough reading and elaborate reading. Specifically, the first stage browses the document to detect the occurrence of events, and the second stage serves to extract specific event arguments. For each concrete event role, elaborate reading hierarchically works from sentences to characters to locate arguments across sentences, thus the scattering-arguments problem is tackled. Meanwhile, rough reading is explored in a multi-round manner to discover undetected events, thus the multi-events problem is handled. Experiment results show the superiority of HRE over prior competitive methods.
Shiyao Cui, Xin Cong, Bowen Yu 0002, Tingwen Liu, Jinqiao Shi
ICASSP6
2022 PrUE: Distilling Knowledge from Sparse Teacher Networks
Shaopu Wang, Xiaojun Chen 0004, Mengzhen Kou, Jinqiao Shi
ECML/PKDD (3)4
2022 Improved Network Pruning via Similarity-Based Regularization
Shaopu Wang, Jiaxin Zhang 0026, Xiaojun Chen 0004, Jinqiao Shi
PRICAI (2)5
2021 SEPC: Improving Joint Extraction of Entities and Relations by Strengthening Entity Pairs Connection
Jiapeng Zhao, Tingwen Liu, Jinqiao Shi
PAKDD (1)4
2021 A Differential Privacy Collaborative Deep Learning Algorithm in Pervasive Edge Computing Environment
abstract
With the development of 5G technology and intelligent terminals, the future direction of the Industrial Internet of Things (IIoT) evolution is Pervasive Edge Computing (PEC). In the pervasive edge computing environment, intelligent terminals can perform calculations and data processing. By migrating part of the original cloud computing model's calculations to intelligent terminals, the intelligent terminal can complete model training without uploading local data to a remote server. Pervasive edge computing solves the problem of data islands and is also successfully applied in scenarios such as vehicle interconnection and video surveillance. However, pervasive edge computing is facing great security problems. Suppose the remote server is honest but curious. In that case, it can still design algorithms for the intelligent terminal to execute and infer sensitive content such as their identity data and private pictures through the information returned by the intelligent terminal. In this paper, we research the problem of honest but curious remote servers infringing intelligent terminal privacy and propose a differential privacy collaborative deep learning algorithm in the pervasive edge computing environment. We use a Gaussian mechanism that meets the differential privacy guarantee to add noise on the first layer of the neural network to protect the data of the intelligent terminal and use analytical moments accountant technology to track the cumulative privacy loss. Experiments show that with the Gaussian mechanism, the training data of intelligent terminals can be protected reduction inaccuracy.
Dayin Zhang, Xiaojun Chen 0004, Jinqiao Shi, Dakui Wang
TrustCom3
2020 Scalable Blind Matching: An Efficient Ciphertext Matching Scheme for Content-Based Pub/Sub Cloud Services
abstract
Content-based publish/subscribe cloud services are prevailing recently. Confidentiality in publish/subscribe cloud services has become a major concern, especially for applications with sensitive data. Many methods have been proposed to achieve the confidentiality of events and subscriptions. Unfortunately, all these approaches suffer significant performance deterioration while matching on large-scale subscription set. In this paper, we propose an efficient ciphertext matching scheme called SBM. To the best of our knowledge, SBM is the first approach which can preserve confidentiality while at the same time provide scalable event matching. Furthermore, we have integrated SBM with an open source publish/subscribe middleware, PADRES and conducted extensive experiments to evaluate our scheme. The experimental results demonstrate that the matching speed of our solution is faster by two orders of magnitude than its counterparts.
Qingling Feng, Ruisheng Shi, Qifeng Luo, Lina Lan, Jinqiao Shi
IEEE BigData5
2020 An Efficient 3-Party Framework for Privacy-Preserving Neural Network Inference
Liyan Shen, Xiaojun Chen 0004, Jinqiao Shi, Ye Dong, Binxing Fang
ESORICS (1)3
2020 2ch-TCN: A Website Fingerprinting Attack over Tor Using 2-channel Temporal Convolutional Networks
abstract
In a website fingerprinting attack, an eavesdropper analyses the traffic between the Tor user and entry node of the Tor network to infer which websites the user has visited. Some recent work apply deep learning algorithms, however, most of them do not fully exploit the packet timing information. In this work, we propose a novel website fingerprinting attack based on a two-channel Temporal Convolutional Networks model that extracts features from both the packet sequences and packet timing information. Our attack is proved to perform better compared to the state-of-the-art attacks. Experiment results also show that the timing information is very useful for classification. Furthermore, we collect our own traffic traces between client and entry node, and transform them into three extraction layers: TCP, TLS and Tor cell layer, and meanwhile record Tor’s cell log at the entry node. The experimental results show that the data of the cell layer is the most divisible among the three layers. Based on the experimental results, we conclude that the adversary at the entry node has an advantage over the one who just listens to traffic between client and entry node.
Yanzeng Li, Tingwen Liu, Jinqiao Shi, Muqian Chen
ISCC5
2020 Napping Guard: Deanonymizing Tor Hidden Service in a Stealthy Way
abstract
In this paper, we propose the Napping Guard attack which can deanonymize hidden services in a stealthy way. The key insight of our method is utilizing a design flaw of hidden service's requests to build a simplex covert channel, which can send message from a malicious guard relay to the collusive Client-OP. With the help of this covert channel, the guard relay delivers the actual IP address of the hidden service to the collusive Client-OP. Considering the Client-OP knows the onion address of hidden service, the adversary is able to deanonymize the hidden service through correlating the actual IP address and onion address on Client-OP. In particular, compared with previous attacks, our covert channel utilizes latency signal instead of traffic signal, and eliminates the dependency of malicious Rend-Point, so as to achieve a better concealment and lower cost. Our experiment shows that the covert channel is reliable that has the precision and recall about 99.35% and 99.19%. In addition, we also propose a mitigation of Napping Guard attack, and report the design flaw to the Tor project.
Muqian Chen, Jinqiao Shi, Can Zhao 0005, Binxing Fang
TrustCom3
2019 A More Efficient Private Set Intersection Protocol Based on Random OT and Balance Hash
abstract
Private set intersection (PSI) is a specific application problem in the field of secure multi-party computation. It allows the participants to compute set intersection collaboratively without learning any additional information about the sets. It is a building block for many real-world privacy-related applications. The state-of-the-art PSI protocol is mainly based on Random Oblivious Transfer (ROT) and cuckoo hash. It is efficient enough in computation but needs high communication cost and high network bandwidth. In this paper, we describe a novel PSI protocol based on ROT and balance hash under semi-honest adversary model. Experiments demonstrate that our protocol has better performance than the cuckoo hash based PSI at the expense of leaking a part of the indexes of hash bucket. We have proved that the impact of information leakage on security is almost negligible. In addition, for the purpose of reducing the communication overhead, we utilize cuckoo filter to optimize our protocol.
Liyan Shen, Xiaojun Chen 0004, Jinqiao Shi, Binxing Fang
ICC3
2019 Towards Comprehensive Security Analysis of Hidden Services Using Binding Guard Relays
Muqian Chen, Jinqiao Shi, Yue Gao 0003, Can Zhao 0005, Wei Sun 0041
ICICS3
2019 ICNet: Incorporating Indicator Words and Contexts to Identify Functional Description Information
abstract
Functional description information refers to the texts that describe the functionality or performance characteristics of a certain object. This type of information is of great potential value for the field of intelligence discovery. Thus automatically and accurately identifying this information from large amounts of texts on the web is very important. In this paper we reduce the functional description problem to a binary classification task deciding whether the input sentence is a functional description sentence or not. However, there exist lots of comment texts in the web data, which are semantically very similar to description texts, making our task quite difficult. Also, existing methods only provide general sentence representation models, which can't lead to targeted ways to solve our problem. Therefore, to address the problem, we not only exploit contexts, like many other previous work did, but also introduce indicator word information to learn rich representations. And in order to incorporate them both, we propose two models, namely ICNet(multi-tasks) and ICNet(ensemble). ICNet(multitasks) exploits them jointly in a integrated process of learning representations, while ICNet(ensemble) exploits them by two respective but concatenated sub-models. Experimental results on the collected real-world dataset indicate that both ICNet(multitasks) and ICNet(ensemble) achieve higher F1 scores compared with FaxtText, CNN, RNN, LSTM and Bi-LSTM, QuickThought models on this task.
Qu Liu, Zhenyu Zhang 0006, Yanzeng Li, Tingwen Liu, Diying Li, Jinqiao Shi
IJCNN6
2019 NeuralAS: Deep Word-Based Spoofed URLs Detection Against Strong Similar Samples
abstract
Spoofed URLs are associated with various cyber crimes such as phishing and ransomware etc. Most existing detection approaches design a set of hand-crafted features and feed them to machine learning classifiers. However, designing such features is a time consuming and labor intensive process. This paper proposes an approach named NeuralAS (Neural Anti-Spoofing) by segmenting URLs into word sequences and detecting spoofed URLs with recurrent neural networks. As a result, NeuralAS can perform detection with high-abstract and poor-interpretable features learned automatically, and achieve accurate detection with contextual information in sequences. We also propose a novel method to construct indistinguishable data sets of strong similar samples, which can be used to evaluate the robustness of different approaches. Extensive experimental results show that NeuralAS works well on spoofed URLs detection, and has a significant effectiveness and robustness even on strong similar data sets.
Jing Ya, Tingwen Liu, Jinqiao Shi, Li Guo 0001, Zhaojun Gu
IJCNN4
2019 Topology Measurement and Analysis on Ethereum P2P Network
abstract
Ethereum, one of the most popular cryptocurrencies, has attracted increasing attention of people in various fields. As the backbone of Ethereum, its peer-to-peer network has an effect on almost every aspect of the ecosystem. Consequently, it's necessary to understand the topological properties of Ethereum P2P network. In this paper, we conducted a measurement of Ethereum P2P network. Our result shows that the graphs of Ethereum network have a small average shortest path length and a large clustering coefficient, and the degree distribution of nodes does not follow a pure power-law distribution. These indicate that Ethereum network is very close to a small world network. Though there are a large number of stale nodes and useless nodes, Ethereum is still resilient to both random failures and targeted attacks. What's more, we find that there are around one hundred abnormal nodes in the network. The IP addresses of nodes included in the neighbors messages they reply are replaced with their own IP addresses. Those nodes might have a bad influence on network routing.
Yue Gao 0003, Jinqiao Shi, Qingfeng Tan, Can Zhao 0005, Zelin Yin
ISCC2
2019 SignalCookie: Discovering Guard Relays of Hidden Services in Parallel
abstract
In this paper, we propose the SignalCookie attack which can reveal guard relays of multiple hidden services in parallel. The key insight of our method is utilizing Rendezvous Cookie and circuit watermark to deliver the hidden services' identifiers to our controlled relays. By conducting the attack in parallel, the speed of our method increases about 11.6 times compared with the previous work, so that we can continually monitor the guard relays of 13604 hidden services for 7 months. And we have an interesting finding by analyzing the distribution of hidden services binding on each guard relays. The distribution is quite uneven, that only 20% of guard relays serve about 89.32% hidden services. These guard relays would bind all hidden services at least one time in about 17 months. At last, we analyze several security problems aggravated by the uneven distribution, and find that just one corrupt guard relay may cause hundreds of hidden services being de-anonymized or eclipsed.
Muqian Chen, Tingwen Liu, Jinqiao Shi, Zelin Yin, Binxing Fang
ISCC4
2019 iMCircle: Automatic Mining of Indicators of Compromise from the Web
abstract
With the rapidly evolving landscape of cyber threats, Indicators of Compromise (IOCs) are aggressively exchanged as forensic artifacts to help security professionals quickly identify and response cyber threats. Previous related studies mostly focus on extracting and generating IOCs from some fixed-point monitoring data sources, which are passive and time-consuming. In this paper, we present iMCircle, an innovation system that automatically mines IOCs from the Web by checking suspicious indicators with the help of open-source threat information. Based on the initial input of several suspicious indicators, iMCircle first collects their relevant public threat information from the Web and generates IOCs by checking whether those indicators are threat indicators in the target threat field. Second, it actively extracts new indicators from the search results as new inputs and checks them as described above. In that way, the system works in a circle and generates IOCs continuously. Running this system for almost two months in the real world, it has the appreciable performances on the active checking of suspicious indicators and the automatic generation of IOCs.
Jing Ya, Tingwen Liu, Quangang Li, Jinqiao Shi, Zhaojun Gu
ISCC5
2019 Toward a Comprehensive Insight Into the Eclipse Attacks of Tor Hidden Services
abstract
Tor hidden services (HSs) are used to provide anonymity services to users on the Internet without disclosing the location of the servers so as to enable freedom of speech. However, existing Tor HSs use decentralized architecture that makes it easier for an adversary to launch DHT-based attacks. In this paper, we present practical Eclipse attacks on Tor HSs that allow an adversary with an extremely low cost to block arbitrary Tor HSs. We found that the dominant cost of this attack is IP address resources, the experimental results show that we can use only three IP addresses to eclipse an arbitrary HS with 100% success probability. To understand the severity of the Eclipse attack problems on Tor HSs, and its security implications, we present the first formal analysis to evaluate the extent of threat such vulnerabilities may cause and quantify the costs of Eclipse attacks involved in our attack via probabilistic analysis. Theoretical analysis suggests that adversaries with a modest number of IP address resources can block a large number of HSs at any time. Finally, we discuss countermeasures and future works.
Qingfeng Tan, Yue Gao 0003, Jinqiao Shi, Binxing Fang, Zhihong Tian 0001
IEEE Internet Things J.3
2019 A Novel Device Identification Method Based on Passive Measurement
abstract
Nowadays, with the continuous integration of production network and business network, more and more Industrial Internet of Things and Internal Office Network have been interconnected and evolved into a large-scale enterprise-level intraindustry network. Terminal devices are the basic units of internal network. Accurate identification of the type of device corresponding to the IP address and detailed description of the communication behavior of the device are of great significance for conducting network security risk assessment, hidden danger investigation, and threat warning. Traditional cyberspace surveying and mapping techniques take the form of active measurement, but they cannot be transplanted to large-scale intranet. Resources or specific targets in internal networks are often protected by firewalls, VPNs, gateways, and other technologies, so they are difficult to analyze and determine by active measurement. In this paper, a passive measurement method is proposed to identify and characterize devices in the network through real traffic data. Firstly, a new graph structure mining method is used to determine the server-like devices and host-like devices; then, the NAT-like devices are determined by quantitative analysis of traffic; finally, by qualitative analysis of the NAT-like device traffic, it is determined whether there are server-like devices behind the NAT-like device. This method will prove to be useful in identifying all kinds of devices in network data traffic, detecting unauthorized NAT-like devices and whether there are server-like devices behind the NAT-like devices.
Wei Sun 0041, Jinqiao Shi, Jian-Guo Jiang
Secur. Commun. Networks5
2018 Character-based BiLSTM-CRF Incorporating POS and Dictionaries for Chinese Opinion Target Extraction
abstract
Opinion target extraction (OTE) is a fundamental step for sentiment analysis and opinion summarization. We analyze the difference between Chinese and the Indo-European languages family, and reduce Chinese OTE to a character-based sequence tagging task. Then we introduce two novel features for each character by distributing POS differentially and using predefined templates over contexts and dictionaries. We further propose a character-based BiLSTM-CRF model incorporating the two feature sequences aligned with the character sequence. Experimental results on real-world consumer review datasets show that our work significantly outperforms the baseline methods for Chinese OTE.
Yanzeng Li, Tingwen Liu, Diying Li, Quangang Li, Jinqiao Shi, Yanqiu Wang
ACML5
2018 Improve Word Mover's Distance with Part-of-Speech Tagging
abstract
Word Mover's Distance (WMD) is a document distance metric with free parameter, intelligible interpretation and unprecedented accuracy on document classification. WMD is on the basis of word embedding and largely focuses on semantic relationships rather than syntactic relationships, which would bring some limitations on measuring document distance. To enhance the impact of syntactic information, we proposed a new method called WMD with Part-of-Speech (PWMD) that integrates part-of-speech (POS) into the original WMD model. POS is a kind of syntactic information, providing more valuable features combined with WMD in document distance metric. Two combination strategies of the POS tagging are provided in “WMD, “word level” and “document level”. The results of contrastive experiments have shown that the PWMD is able to get better document distance than WMD.
Xiaojun Chen 0004, Li Bai 0004, Dakui Wang, Jinqiao Shi
ICPR4
2018 FraudNE: a Joint Embedding Approach for Fraud Detection
abstract
Detecting fraudsters is a meaningful problem for both users and e-commerce platform. Existing graph-based approaches mainly adopt shallow models, which cannot capture the highly non-linear relationship between vertexes in a bipartite graph composed of users and items. To address this issue, in this paper we propose a joint deep structure embedding approach FraudNE for fraud detection that (a) can preserve the highly non-linear structural information of networks, (b) is robust to sparse networks, (c) embeds different types of vertexes jointly in the same latent space. It is worth mentioning that we can detect multiple fraudulent groups without the number of groups as a priori. Compared with baselines, our method achieved significant accuracy improvement.
Mengyu Zheng, Chuan Zhou 0001, Jia Wu 0001, Shirui Pan, Jinqiao Shi, Li Guo 0001
IJCNN5
2017 Efficient and Scalable Privacy-Preserving Similar Document Detection
abstract
Similar document detection has been well studied for many applications, such as file management systems, plagiarism and double submission detection. Traditional detection algorithms are challenged by the privacy-preserving problems. Recently, privacy-preserving similar document detection between two parties gains more attention. However, most of the existing works mainly focus on computing similarity between two documents, and they are inefficient with O(n2) computation complexity when processing secure comparison between two n-document sets. Focusing on this problem, this paper presents a new efficient and scalable privacy-preserving similar document detection protocol based on oblivious multi-garbled Bloom filter intersection and MinHash algorithm. Experimental evaluation shows that when processing large document sets, our protocol still remains linear computation complexity with the scale of document sets increasing and achieves overwhelming computational performance improvement against other major approaches.
Xiaojie Yu, Xiaojun Chen 0004, Jinqiao Shi, Liyan Shen, Dakui Wang
GLOBECOM3
2017 A closer look at Eclipse attacks against Tor hidden services
abstract
Tor hidden Services are used to provide anonymity service to users on the Internet without disclosing the location of the servers so as to enable freedom of speech. However, existing Tor hidden services use decentralized architecture making it easier for an adversary to launch DHT-based attacks. In this paper, we present practical Eclipse attacks on Tor hidden services that allow an adversary with an extremely low cost to block arbitrary Tor hidden services. We found that the dominant cost of this attack is IP address resources. The experimental results show that we can eclipse an arbitrary hidden service with 100% success probability with only 6 IP addresses. To understand the severity of the Eclipse attack problems on Tor's hidden services, and its security implications, we present the first formal analysis to evaluate the extent of threat such vulnerabilities may cause and quantify the costs of Eclipse attacks involved in our attack via probabilistic analysis. Theoretical analysis suggests that adversaries with a modest number of IP address resources can block a large number of hidden services at any time.
Qingfeng Tan, Yue Gao 0003, Jinqiao Shi, Binxing Fang
ICC3
2017 Flexible Expert Finding on the Web via Semantic Hypergraph Learning and Affinity Propagation Model
abstract
Expert finding (EF) task has received widespread attention as an important task of information retrieval.One key category of EF is expert finding on the web, which seeks to rank influential public figures from diverse webpage sources with respect to given query.Previous web expert finding approach relies on casting webpages to hypergraph structure and run heat diffusion process to find the top ranking person vertices according to their heat of popularity.Such approach suffer from two major drawbacks:First, previous web expert finding approach (CoDiffusion) suffers from unflexibility of selecting queries.This means that all corresponding queries must be stored as vertices in hypergraph index beforehand, otherwise CoDiffusion cannot run the expert finding process.Such defect make it ungeneric and incapable of handling the newly invented technical terms or phrases in real world scenarios.Second, the performance of previous approach is less satisfying.We incorporate semantic relatedness information with Hypergraph Learning Framework and Affinity Propagation ({HLFAP}) to handle the above drawbacks.In order to overcome the first disadvantage, we distribute initial heat to the related vertices according to their semantic similarity on given query. In order to solve the second disadvantage, we propose semantic labeled hypergraph learning framework and person influence affinity propagation model to make high quality candidates can receive more heat transition. Experimental results shows that our generic methodology achieves more satisfying results than the non-semantics state-of-the-art baseline method.
Tingwen Liu, Jinqiao Shi, Qiuyan Wang, Li Guo 0001
ICTAI3
2017 Large-scale discovery and empirical analysis for I2P eepSites
abstract
I2P is a widely used low-latency anonymous network that provides privacy to service providers, such as anonymous web services called eepSites. The large-scale discovery of eepSites allows us to grasp their size, content and popularity. In this paper, three approaches were proposed to discover eepSites: (1) running floodfill routers, (2) gathering hosts.txt files actively and (3) crawling popular portal eepSites. In our nineteen-day real-world experiments, the combination of the three methods in total discovered 1861 online eepSites covering over 80% of all eepSites in I2P network. And the coupon collector's problem was used for theoretical analysis, showing that eepSites discovery based on running floodfill routers is straightforward and efficient with low cost. Besides, the popularity and availability of eepSites were estimated and analyzed.
Yue Gao 0003, Qingfeng Tan, Jinqiao Shi, Muqian Chen
ISCC3
2017 Improving Password Guessing Using Byte Pair Encoding
Dakui Wang, Xiaojun Chen 0004, Jinqiao Shi, Li Guo 0001
ISC5
2016 Vision-based real-time 3D mapping for UAV with laser sensor
abstract
Real-time 3D mapping with MAV (Micro Aerial Vehicle) in GPS-denied environment is a challenging problem. In this paper, we present an effective vision-based 3D mapping system with 2D laser-scanner. All algorithms necessary for this system are on-board. In this system, two cameras work together with the laser-scanner for motion estimation. The distance of the points detected by laser-scanner are transformed and treated as the depth of image features, which improves the robustness and accuracy of the pose estimation. The output of visual odometry is used as an initial pose in the Iterative Closest Point (ICP) algorithm and the motion trajectory is optimized by the registration result. We finally get the MAV's state by fusing IMU with the pose estimation from mapping process. This method maximizes the utility of the point clouds information and overcomes the scale problem of lacking depth information in the monocular visual odometry. The results of the experiments prove that this method has good characteristics in real-time and accuracy.
Jinqiao Shi, Bingwei He, Jianwei Zhang 0001
IROS1
2016 An Unsupervised Framework Towards Sci-Tech Compound Entity Recognition
Tingwen Liu, Li Guo 0001, Jiapeng Zhao, Jinqiao Shi
KSEM5
2015 A Self-learning Rule-Based Approach for Sci-tech Compound Phrase Entity Recognition
Tingwen Liu, Jinqiao Shi, Li Guo 0001
APWeb4
2015 StegoP2P: Oblivious user-driven unobservable communications
abstract
With increasing concern for erosion of privacy, privacy preserving and censorship-resistance techniques are becoming more and more important. Anonymous communication techniques offer an important method defending against Internet surveillance, but these techniques don't conceal themselves when used. In this paper, we propose StegoP2P, an unobservable communication system with Internet users in overlay network that relies on Innocent users' oblivious data downloading, StegoP2P works by deploying a end-to-middle proxies, which inspect special steganography flows from StegoP2P users to innocent-looking destinations and mirror them to the true destination requested by oblivious P2P users. The hidden communication is indistinguishable from normal network communications to any adversaries without a private key, hence, making the StegoP2P clients unobservable. We have developed a proof-of-concept application based on Vuze and conducted evaluations through experiments.
Qingfeng Tan, Jinqiao Shi, Binxing Fang
ICC2
2015 Towards misdirected email detection based on multi-attributes
abstract
Email has become widely used in recent years bringing with it new problems. Although this event doesn't happen often, misdirected emails can bring out great information leakage. It is not easy to detect these misdirected emails from legitimate ones since they may be only distinguishable in the sender's perspective. Existing methods discover misdirected emails from user agent or gateway but are not appropriate for varied application environment. This paper proposes a misdirected mail detection method based on multi-attributes which can be deployed on server side. Three type of attributes including email content fingerprinting, social relationship and meta information are considered in this method. Based on SVM classification algorithm, experiments show that it can detect misdirected emails with up to 91.6% accuracy.
Yiguo Pu, Jinqiao Shi, Xiaojun Chen 0004, Li Guo 0001, Tingwen Liu
ISCC2
2014 A probabilistic approach towards modeling email network with realistic features
abstract
Email plays a very important role in our daily life. Much work have been put into practice on email network. Those studies mostly require real email network datasets and reliable models to analyze user information and understand the mechanisms of network evolution. However, much research work is constrained by the absence of real large-scale email datasets. Although email communication is ubiquitous, there are very few large-scale available email datasets satisfied different research purposes. Due to privacy policy and restricted permissions, it is arduous to collect a real large-scale email dataset in a short time. Various social network models are usually used to create synthetic email networks. However, these models focus on modeling several structural properties of network without considering user behaviour patterns. They are not appropriate to generate large-scale realistic synthetic email network datasets. Towards this end, we propose a probabilistic model by which we can construct large-scale synthetic email datasets with a small captured email log. What is more important is that the generated synthetic dataset matches real email network properties and individual communication patterns. Moreover, it has linear complexity, and can be paralleled easily. Experimental results on Enron dataset demonstrate the above benefits of our model.
Quangang Li, Jinqiao Shi, Tingwen Liu, Li Guo 0001, Zhiguang Qin
ICCCN2
2014 Towards misdirected email detection for preventing information leakage
abstract
With the widespread usage of emails, information leakage via misdirected emails becomes a practical and disastrous problem, which should be addressed at all costs. Prior methods have two limitations: privacy issue as relying on email contents to work, and high cost as building too many targeted models. In this paper, we reduce the detection of misdirected emails to a binary classification problem, and build only a universal model to detect misdirected emails. We introduce some representative features that can vividly describe the characteristics of misdirected emails while not infringe users' privacy. Then we design novel algorithms to get these features. The random forest classifier is chosen to perform the detecting task. Experimental results show that our work is able to detect misdirected emails with 89% precision rate and 82% recall rate in average.
Tingwen Liu, Yiguo Pu, Jinqiao Shi, Quangang Li, Xiaojun Chen 0004
ISCC3
2014 Winnowing Double Structure for Wildcard Query in Payload Attribution
Xiaojun Chen 0004, Yiguo Pu, Jinqiao Shi, Sihan Qing
ISC5
2014 A Moving Target Framework to Improve Network Service Accessibility
abstract
Nowadays, the problem of Internet services accessibility has become a hot topic with more and more cyber attacks and censorship. This has prompted the rapid development of jamming-resistance infrastructure consisting of multiple dynamic access points such as proxies, anonymous communication nodes and covert communication nodes. However, the channel between user and access point has become an emerging attacking target for the adversary. Once the channel is detected and identified by the adversary, the channel will be interrupted. Though users can require new access points from the infrastructure and resume the communication to their destination, the service quality will be downgraded dramatically due to the time-consuming bootstrapping process. In this paper, a moving target framework is proposed to improve the network service accessibility, which combines both time-consuming bootstrapping and frequent and low-cost channel refreshing operations. The constant channel refreshing operations is imported to limit the adversary's ability of detecting and blocking communications, thus to make the lifespan of effective communication longer. With theoretical and simulating analysis, the optimized refreshing strategy is proposed, which can help improve service accessibility.
Jinqiao Shi, Xiao Wang 0001, Binxing Fang, Li Guo 0001
NAS1
2014 Towards Improving Service Accessibility by Adaptive Resource Distribution Strategy
Jinqiao Shi, Xiao Wang 0001, Binxing Fang, Qingfeng Tan, Li Guo 0001
SecureComm (1)1
2014 Towards Fast and Optimal Grouping of Regular Expressions via DFA Size Estimation
abstract
Regular Expression (RegEx) matching, as a core operation in many network and security applications, is typically performed on Deterministic Finite Automata (DFA) to process packets at wire speed; however, DFA size is often exponential in the number of RegExes. RegEx grouping is the practical way to address DFA state explosion. Prior RegEx grouping algorithms are extremely slow and memory intensive. In this paper, we first propose DFAestimator, an algorithm that can quickly estimate DFA size for a given RegEx set without building the actual DFA. Second, we propose RegexGrouper, a RegEx grouping algorithm based on DFA size estimation. In terms of speed and memory consumption, our work is orders of magnitude more efficient than prior art because DFA size estimation is much faster and memory efficient than DFA construction. In terms of the resulting size sum of DFAs, our work is significantly more effective than prior art because we use a much finer grained quantification of the degree of interaction between two RegExes. For example, to divide the RegEx set of the L7-filter system into 7 groups, prior art uses 279.3 minutes and the resulting 7 DFAs have a total of 29047 states, whereas RegexGrouper uses 3.2 minutes and the resulting 7 DFAs have a total of 15578 states.
Tingwen Liu, Alex X. Liu, Jinqiao Shi, Li Guo 0001
IEEE J. Sel. Areas Commun.3
2013 Design and Evaluation of Access Control Model Based on Classification of Users' Network Behaviors
Peipeng Liu, Jinqiao Shi, Li Guo 0001
APWeb2
2013 An empirical analysis of family in the Tor network
abstract
As one of the most popular anonymous communication systems, Tor has become a research hotspot in this area. Recently, Tor nodes from Tor families (referred to as family nodes) have played an increasingly important role and caused significant influence on the Tor network. However, existing research about Tor mostly focuses on the entire Tor network without much consideration about the difference between family nodes and the others. In order to analyze family nodes' contribution to the entire Tor network as well as their influence, this paper distinguishes family nodes from the others, and gives an empirical analysis of family nodes based on the live Tor network data of 3 years. Results show that, family nodes compose a small but full functional subset of Tor nodes; and compared with the other Tor nodes, they can provide relatively stable and high-performance service to Tor users. Furthermore, family nodes naturally form a hot area in the Tor network, relaying increasingly high-density traffic through a small number of nodes. Compared with random node targets, selective attacks focusing on family nodes can cause serious availability downgrade of the Tor network with much lower cost.
Xiao Wang 0001, Jinqiao Shi, Binxing Fang, Li Guo 0001
ICC2
2013 IX-Level Adversaries on Entry- and Exit-Transmission Paths in Tor Network
abstract
Tor is a worldwide publicly deployed low-latency anonymity system. In order to prevent observers from telling where the data came from and where it's going, data packets on the Tor network take a pathway through several intermediate relays. However, nodes selection to build such a pathway is oblivious to Internet routing, so anonymity guarantees can break down in cases where an attacker can correlate traffic across the entry- and exit-segments of a Tor circuit. Although many works have been done to avoid this kind of collusion attack, recent researches [18] indicated that some Internet exchanges (IXes) locating at the entry- and exit-transmission paths in Tor network (that are the paths from the client to the chosen entry node and from the chosen exit node to the destination) are still possible to perform a correlation attack. However, few works have been done to suggest and verify modifications to Tor's path selection algorithm that would help clients avoid an IX-level observer. In this paper, we first, based on the entry-exit pairs chosen by Tor's path selection algorithm, demonstrated that the probability of a single IX observing both ends of an anonymous Tor connection is greater than previously thought. And then, we proposed and evaluated the effectiveness of a simple IX-awareness path selection algorithm that help to resist IX-level attackers.
Peipeng Liu, Jinqiao Shi, Xiao Wang 0001, Qingfeng Tan
NAS2
2013 Botnet Triple-Channel Model: Towards Resilient and Efficient Bidirectional Communication Botnets
Xiang Cui, Binxing Fang, Jinqiao Shi, Chaoge Liu
SecureComm3
2012 The Triple-Channel Model: Toward Robust and Efficient Advanced Botnets (Poster Abstract)
Xiang Cui, Jinqiao Shi, Chaoge Liu
RAID2
2011 A Covert Communication Method Based on User-Generated Content Sites
abstract
With the worldwide increasing of Internet censorship, censorship-resistance technology has attracted more and more attentions, some famous systems, such as Tor and JAP, have been deployed to provide public service for censorship-resistance. However, these systems all rely on dedicated infrastructure and entry points for service accessibility. The network infrastructure and entry points may become the target of censorship attack. In this paper, a UGC-based method is proposed (called user-generated content based covert communication, UGC3) for covert communication in a friends-to-friends (F2F) manner. It uses existing infrastructures (i.e., UGC sites ) to form a fully distributed overlay network. An efficient resource discovery algorithm is proposed to negotiate the rendezvous point. Analysis shows that this method is able to circumvent internet censorship with user repudiation and fault tolerance.
Qingfeng Tan, Peipeng Liu, Jinqiao Shi, Xiao Wang 0001, Li Guo 0001
ICTAI3
2008 Protecting Mobile Codes Using the Decentralized Label Model
abstract
For protection of the confidentiality and integrity of the mobile codes, this paper proposes a new decentralized label model and a implementation of this model in Linux system, MCGuard. Using MCGuard, the owners can flexibly define their security policies to control the dissemination of their mobile codes just by labelling them. By intercepting system calls, MCGuard inserts an interposition layer between the processes and system calls to control the data flows of mobile codes and guarantee them not to be transmitted to insecure channels and manipulated by malicious principals. In MCGuard, the labelling and control of the mobile codes and their transmitting channels is performed at the level of standard operating system abstractions, and the labels can migrate between hosts. This makes the MCGuard applicable in mobile code systems composed of the stock Linux OS and existing mobile codes.
Jian-Wei Ye, Binxing Fang, Jinqiao Shi, Zhi-Gang Wu
WAIM3
2004 Towards an Analysis of Source-Rewriting Anonymous Systems in a Lossy Environment
Jinqiao Shi, Binxing Fang
PDCAT1