VLDB 2026 Research / reviewers in the wild / expert
Yuexiang Yang
dblp:117/1744
· DBLP profile ↗
33ranked-venue papers
0as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 11 · 6 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Secure Optimization With Asynchronous Structured Skyline Predicates Under Vertical Data FederationabstractSkyline optimization is a powerful tool for filtering prominent data to support analysis and decision-making. However, traditional centralized skyline predicates are inadequate for contemporary data islands, and shallow data federation poses a threat to privacy with sensitive data. In existing distributed environments, achieving both efficiency and security in skyline computation remains a critical challenge. This paper addresses the challenge of performing secure skyline predicates on encrypted data federation while safeguarding both the dataset and skyline from unauthorized access. We propose a novel asynchronous structured skyline predicate based on vertical dominance and truth-value conversion, taking full advantage of distributed computing. Furthermore, we introduce a secure optimization that balances security and efficiency, thereby facilitating a distributed skyline predicate. We evaluate the efficiency and scalability across various parameters, demonstrating improvements in traversal overhead and expensive ciphertext operations. Yu Chen 0056, Rongmao Chen, Shaojing Fu, Mingwu Zhang, Yuexiang Yang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | ANODYNE: Mitigating backdoor attacks in federated learning
Zhipin Gu, Jiangyong Shi, Yuexiang Yang |
Expert Syst. Appl. | 3 |
| 2025 | Secure Optimization With Preferred Skyline Predicate on Incomplete DataabstractOutsourcing data storage and computations to cloud servers offers a cost-effective solution for remote data management and query processing. However, ensuring the privacy of sensitive information remains a critical concern, and existing secure algorithms rely on data completeness where all attribute values are valid to ignore the dominance issues under intransitivity and cyclicity. This paper addresses the challenge of executing secure skyline predicates on outsourced incomplete data, while keeping the dataset, queries, and results confidential from the cloud servers. We propose a novel secure dominance under incomplete data as a core component of various query types. To balance security and efficiency, we introduce two filtering methods around access patterns. Additionally, we present two secure skyline extensions concerning dimension and skyband to produce meaningful skylines. The proposed solutions are empirically evaluated for efficiency and scalability on diverse datasets, demonstrating the practical viability of our approach. Yu Chen 0056, Rongmao Chen, Shaojing Fu, Xinyi Huang 0001, Mingwu Zhang, Yuexiang Yang |
IEEE Trans. Serv. Comput. | 6 |
| 2024 | Honest-Majority Maliciously Secure Skyline Queries on Outsourced DataabstractThe application of skyline queries on outsourced databases significantly aids online analysis, yet efficiently handling encrypted queries remains a formidable obstacle. Moreover, query outcomes are vulnerable to potential malicious cloud services. To circumvent these limitations, this work presents the Honest-Majority and Maliciously Skyline Query scheme (HMMSQ), which facilitates efficient skyline queries while safeguarding the privacy of datasets, queries, and skylines, as well as detecting malevolent activities. The core of HMMSQ is an optimized skyline diagram constructed by a novel skyline region-splitting algorithm for accurate skyline queries. Furthermore, it mitigates the frequency of dataset accesses by leveraging a multi-path R-tree for secure skyline retrieval. Notably, the majority of malicious behavior detection is focused on the servers, thereby minimizing user authentication overhead. The complexity and security are thoroughly analyzed, and experimental evaluations on various datasets demonstrate its efficiency and practicality in terms of computational cost and communication overhead. Remarkably, HMMSQ outperforms existing methods in query latency, achieving up to an order of magnitude improvement. Yu Chen 0113, Lin Liu 0018, Rongmao Chen, Shaojing Fu, Yuexiang Yang |
CIKM | 5 |
| 2024 | Speedy Privacy-Preserving Skyline Queries on Outsourced Data
Yu Chen 0113, Lin Liu 0018, Rongmao Chen, Shaojing Fu, Yuexiang Yang, Jiangyong Shi, Liangzhong He |
ESORICS (2) | 5 |
| 2024 | Multiview Deep Anomaly Detection: A Systematic ExplorationabstractAnomaly detection (AD), which models a given normal class and distinguishes it from the rest of abnormal classes, has been a long-standing topic with ubiquitous applications. As modern scenarios often deal with massive high-dimensional complex data spawned by multiple sources, it is natural to consider AD from the perspective of multiview deep learning. However, it has not been formally discussed by the literature and remains underexplored. Motivated by this blank, this article makes fourfold contributions: First, to the best of our knowledge, this is the first work that formally identifies and formulates the multiview deep AD problem. Second, we take recent advances in relevant areas into account and systematically devise various baseline solutions, which lays the foundation for multiview deep AD research. Third, to remedy the problem that limited benchmark datasets are available for multiview deep AD, we extensively collect the existing public data and process them into more than 30 multiview benchmark datasets via multiple means, so as to provide a better evaluation platform for multiview deep AD. Finally, by comprehensively evaluating the devised solutions on different types of multiview deep AD benchmark datasets, we conduct a thorough analysis on the effectiveness of the designed baselines and hopefully provide other researchers with beneficial guidance and insight into the new multiview deep AD topic. Siqi Wang 0001, Jiyuan Liu 0003, Xinwang Liu 0002, Sihang Zhou 0001, En Zhu, Yuexiang Yang, Jianping Yin, Wenjing Yang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Defending against Poisoning Attacks in Federated Learning from a Spatial-temporal PerspectiveabstractIn federated learning, the central server aggregates local model updates from the participants in the network to generate a global model. For the purpose of protecting clients' privacy, the server is designed to have no visibility into how these updates are generated. The nature of federated learning makes detecting and defending against malicious model up-dates a challenging task. Unlike existing works that struggle to defend against poisoning attacks from a spatial perspective, the paper considers mitigating the impact of attacks from a spatial-temporal perspective. This paper proposes Fedmvae, a robust federated learning framework. Fedmvae uses multiple variational autoencoder models to detect and exclude malicious model updates from a spatial perspective. Moreover, to handle poisoning attacks with time-varying features, we propose generating a robust global model update according to momentum-based update speculation and historical global updates. Fedmvae is tested with extensive experiments on both IID and non-IID datasets, showing a competitive performance over existing aggregation methods under both Byzantine attacks and backdoor attacks. Zhipin Gu, Jiangyong Shi, Yuexiang Yang, Liangzhong He |
SRDS | 3 |
| 2023 | Defending against Adversarial Attacks in Federated Learning on Metric Learning ModelabstractThe industry has widely deployed federated learning (FL) due to its promise to protect clients’ privacy. However, FL is vulnerable to adversarial attacks when the participants are compromised. The defense against adversarial attacks is a challenging problem in FL. Moreover, existing defense methods optimize the dimensionality reduction and anomaly detection models separately, leading to a disappointing projection space and low detection accuracy. We propose a deep metric learning-based anomaly detection to project the model gradients into a metric space where the malicious gradients are separated from benign ones. Meanwhile, while existing methods require an auxiliary dataset to train the defense model, the auxiliary dataset is usually unavailable to the server in the FL setting. We propose a self-supervised method to distill the data between the training epochs of our defense model. To handle radical changes in malicious model gradients, we utilize a median-based aggregated gradient filter to discard improper aggregated gradients. We show experimentally that our algorithm has a competitive performance over existing methods under Byzantine attacks and backdoor attacks with various triggers. Zhipin Gu, Jiangyong Shi, Yuexiang Yang, Liangzhong He |
TrustCom | 3 |
| 2023 | Contrastive Multi-View Kernel LearningabstractKernel method is a proven technique in multi-view learning. It implicitly defines a Hilbert space where samples can be linearly separated. Most kernel-based multi-view learning algorithms compute a kernel function aggregating and compressing the views into a single kernel. However, existing approaches compute the kernels independently for each view. This ignores complementary information across views and thus may result in a bad kernel choice. In contrast, we propose the Contrastive Multi-view Kernel - a novel kernel function based on the emerging contrastive learning framework. The Contrastive Multi-view Kernel implicitly embeds the views into a joint semantic space where all of them resemble each other while promoting to learn diverse views. We validate the method's effectiveness in a large empirical study. It is worth noting that the proposed kernel functions share the types and parameters with traditional ones, making them fully compatible with existing kernel theory and application. On this basis, we also propose a contrastive multi-view clustering framework and instantiate it with multiple kernel k-means, achieving a promising performance. To the best of our knowledge, this is the first attempt to explore kernel generation in multi-view setting and the first approach to use contrastive learning for a multi-view kernel learning. Jiyuan Liu 0003, Xinwang Liu 0002, Yuexiang Yang, Qing Liao 0001, Yuanqing Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | MADDC: Multi-Scale Anomaly Detection, Diagnosis and Correction for Discrete Event LogsabstractAnomaly detection for discrete event logs can provide critical information for building secure and reliable systems in various application domains, such as large scale data centers, autonomous driving, and intrusion detection. However, the task is very challenging due to the lack of a clear understanding and definition of anomaly in the specific problem space, and the log data is often highly complex with temporal correlation. Existing deep learning based methods mostly suffer from such issues as overfitting, uncertainty or low interpretability; consequently, the detection results may be inaccurate, with little information to help security analysts diagnose the reported anomalies with high confidence. To tackle this challenge, in this research, we propose a general framework named MADDC, which aims to (1) accurately perform Multi-scale Anomaly Detection, Diagnosis and Correction for discrete event logs, and (2) help analysts further mitigate anomalies based on diagnosis results. Specifically, we first design a new anomaly critic for LSTM variational autoencoder based model to alleviate overfitting and reduce false negatives during anomaly detection. As one of our main contributions, we then introduce process mining technique to build process-centric workflow models in an unsupervised manner, which forms the ‘normal’ context of an event sequence and help perform accurate and consistent anomaly diagnosis through global sequence alignment. Experiments on publicly available datasets show that MADDC not only outperformed several representative methods in terms of detection accuracy, but also could improve the visibility to abnormal deviations from normal execution, hence helping security analysts understand anomalies and make further corrections. Xiaolei Wang 0003, Lin Yang 0031, Linru Ma, Junchao Xiao, Jiyuan Liu 0003, Yuexiang Yang |
ACSAC | 8 |
| 2022 | Optimal Neighborhood Multiple Kernel Clustering With Adaptive Local KernelsabstractMultiple kernel clustering (MKC) algorithm aims to group data into different categories by optimally integrating information from a group of pre-specified kernels. Though demonstrating superiorities in various applications, we observe that existing MKC algorithms usuallydo not sufficiently consider the local density around individual data samplesandexcessively limit the representation capacity of the learned optimal kernel, leading to unsatisfying performance. In this paper, we propose an algorithm, called optimal neighborhood MKC with adaptive local kernels (ON-ALK), to address the two issues. In specific, we construct adaptive local kernels to sufficiently consider the local density around individual data samples, where different numbers of neighbors are discriminatingly selected on each sample. Further, the proposed ON-ALK algorithm boosts the representation of the learned optimal kernel via relaxing it into the neighborhood area of weighted combination of the pre-specified kernels. To solve the resultant optimization problem, a three-step iterative algorithm is designed and theoretically proven to be convergent. After that, we also study the generalization bound of the proposed algorithm. Extensive experiments have been conducted to evaluate the clustering performance. As indicated, the algorithm significantly outperforms state-of-the-art methods in recent literatures on six challenging benchmark datasets, verifying its advantages and effectiveness. Jiyuan Liu 0003, Xinwang Liu 0002, Jian Xiong 0002, Qing Liao 0001, Sihang Zhou 0001, Siwei Wang 0001, Yuexiang Yang |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2022 | Multiview Subspace Clustering via Co-Training Robust Data RepresentationabstractTaking the assumption that data samples are able to be reconstructed with the dictionary formed by themselves, recent multiview subspace clustering (MSC) algorithms aim to find a consensus reconstruction matrix via exploring complementary information across multiple views. Most of them directly operate on the original data observations without preprocessing, while others operate on the corresponding kernel matrices. However, they both ignore that the collected features may be designed arbitrarily and hard guaranteed to be independent and nonoverlapping. As a result, original data observations and kernel matrices would contain a large number of redundant details. To address this issue, we propose an MSC algorithm that groups samples and removes data redundancy concurrently. In specific, eigendecomposition is employed to obtain the robust data representation of low redundancy for later clustering. By utilizing the two processes into a unified model, clustering results will guide eigendecomposition to generate more discriminative data representation, which, as feedback, helps obtain better clustering results. In addition, an alternate and convergent algorithm is designed to solve the optimization problem. Extensive experiments are conducted on eight benchmarks, and the proposed algorithm outperforms comparative ones in recent literature by a large margin, verifying its superiority. At the same time, its effectiveness, computational efficiency, and robustness to noise are validated experimentally. Jiyuan Liu 0003, Xinwang Liu 0002, Yuexiang Yang, Xifeng Guo 0001, Marius Kloft, Liangzhong He |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Hierarchical Multiple Kernel ClusteringabstractCurrent multiple kernel clustering algorithms compute a partition with the consensus kernel or graph learned from the pre-specified ones, while the emerging late fusion methods firstly construct multiple partitions from each kernel separately, and then obtain a consensus one with them. However, both of them directly distill the clustering information from kernels or graphs to partition matrices, where the sudden dimension drop would result in loss of advantageous details for clustering. In this paper, we provide a brief insight of the aforementioned issue and propose a hierarchical approach to perform clustering while preserving advantageous details maximumly. Specifically, we gradually group samples into fewer clusters, together with generating a sequence of intermediary matrices of descending sizes. The consensus partition with is simultaneously learned and conversely guides the construction of intermediary matrices. Nevertheless, this cyclic process is modeled into an unified objective and an alternative algorithm is designed to solve it. In addition, the proposed method is validated and compared with other representative multiple kernel clustering algorithms on benchmark datasets, demonstrating state-of-the-art performance by a large margin. Jiyuan Liu 0003, Xinwang Liu 0002, Siwei Wang 0001, Sihang Zhou 0001, Yuexiang Yang |
AAAI | 5 |
| 2021 | One-pass Multi-view Clustering for Large-scale DataabstractExisting non-negative matrix factorization based multi-view clustering algorithms compute multiple coefficient matrices respect to different data views, and learn a common consensus concurrently. The final partition is always obtained from the consensus with classical clustering techniques, such as k-means. However, the non-negativity constraint prevents from obtaining a more discriminative embedding. Meanwhile, this two-step procedure fails to unify multi-view matrix factorization with partition generation closely, resulting in unpromising performance. Therefore, we propose an one-pass multi-view clustering algorithm by removing the non-negativity constraint and jointly optimize the aforementioned two steps. In this way, the generated partition can guide multi-view matrix factorization to produce more purposive coefficient matrix which, as a feedback, improves the quality of partition. To solve the resultant optimization problem, we design an alternate strategy which is guaranteed to be convergent theoretically. Moreover, the proposed algorithm is free of parameter and of linear complexity, making it practical in applications. In addition, the proposed algorithm is compared with recent advances in literature on benchmarks, demonstrating its effectiveness, superiority and efficiency. Jiyuan Liu 0003, Xinwang Liu 0002, Yuexiang Yang, Li Liu 0002, Siqi Wang 0001, Weixuan Liang, Jiangyong Shi |
ICCV | 3 |
| 2021 | Detecting Malicious Model Updates from Federated Learning on Conditional Variational AutoencoderabstractIn federated learning, the central server combines local model updates from the clients in the network to create an aggregated model. To protect clients' privacy, the server is designed to have no visibility into how these updates are generated. The nature of federated learning makes detecting and defending against malicious model updates a challenging task. Unlike existing works that struggle to defend against Byzantine clients, the paper considers defending against targeted model poisoning attack in the federated learning setting. The adversary aims to reduce the model performance on targeted subtasks while maintaining the main task's performance. This paper proposes Fedcvae, a robust and unsupervised federated learning framework where the central server uses conditional variational autoencoder to detect and exclude malicious model updates. Since the reconstruction error of malicious updates is much larger than that of benign ones, it can be used as an anomaly score. We formulate a dynamic threshold of reconstruction error to differentiate malicious updates from normal ones based on this idea. Fedcvae is tested with extensive experiments on IID and non-IID federated benchmarks, showing a competitive performance over existing aggregation methods under Byzantine attack and targeted model poisoning attack. Zhipin Gu, Yuexiang Yang |
IPDPS | 2 |
| 2021 | Self-Representation Subspace Clustering for Incomplete Multi-view DataabstractIncomplete multi-view clustering is an important research topic in multimedia where partial data entries of one or more views are missing. Current subspace clustering approaches mostly employ matrix factorization on the observed feature matrices to address this issue. Meanwhile, self-representation technique is left unexplored, since it explicitly relies on full data entries to construct the coefficient matrix, which is contradictory to the incomplete data setting. However, it is widely observed that self-representation subspace method enjoys a better clustering performance over the factorization based one. Therefore, we adapt it to incomplete data by jointly performing data imputation and self-representation learning. To the best of our knowledge, this is the first attempt in incomplete multi-view clustering literature. Besides, the proposed method is carefully compared with current advances in experiment with respect to different missing ratios, verifying its effectiveness. Jiyuan Liu 0003, Xinwang Liu 0002, Yi Zhang 0104, Pei Zhang 0008, Wenxuan Tu, Siwei Wang 0001, Sihang Zhou 0001, Weixuan Liang, Siqi Wang 0001, Yuexiang Yang |
ACM Multimedia | 10 |
| 2021 | Detecting Malicious Gradients from Asynchronous SGD on Variational AutoencoderabstractIn asynchronous distributed learning, the parameter server updates the global model as soon as a new gradient is received from any device. The asynchronous systems are designed to address the existence of lagging devices which is inevitable due to device heterogeneity and network unreliability. However, the lack of synchrony incurs additional noise and makes detecting and defending against malicious model gradients a challenging task. Unlike existing works that struggle to design robust methods to tolerate untargeted model poisoning gradients, the paper considers detecting and removing targeted model poisoning gradients from the normal asynchronous training process. This paper proposes Asynvae, a robust distributed asynchronous learning framework where the parameter server uses variational autoencoder to detect and exclude malicious gradients. Since the reconstruction error of malicious updates is much larger than that of benign ones, it can be used as an anomaly score. We formulate a threshold of reconstruction error to differentiate malicious updates from normal ones based on this idea. Asynvae is tested with extensive experiments on distributed learning benchmarks, showing a competitive performance over existing distributed learning methods under untargeted model poisoning attack, targeted model poisoning attack and lagging attack. Zhipin Gu, Yuexiang Yang, Heyuan Shi |
SRDS | 2 |
| 2021 | Dual-branch combination network (DCN): Towards accurate diagnosis and lesion segmentation of COVID-19 using CT images
Kai Gao 0011, Jianpo Su, Zhongbiao Jiang, Zhichao Feng, Hui Shen 0004, Pengfei Rong, Xin Xu 0001, Yuexiang Yang, Wei Wang 0434, Dewen Hu |
Medical Image Anal. | 10 |
| 2019 | Optimizing partitioning strategies for faster inverted index compression
Xingshen Song, Yuexiang Yang, Kun Jiang 0001 |
Frontiers Comput. Sci. | 2 |
| 2019 | PrivacyContext: identifying malicious mobile privacy leak using program contextabstractSerious concerns have been raised about user's privacy leak in mobile apps, and many detection approaches are proposed. To evade detection, new mobile malware starts to mimic privacy-related behaviours of benign apps, and mix malicious privacy leak with benign ones to reduce the chance of being observed. Since prior proposed approaches primarily focus on the privacy leak discovery, these evasive techniques will make differentiating between malicious and benign privacy disclosures difficult during privacy leak analysis. In this paper, we propose PrivacyContext to identify malicious privacy leak using context. PrivacyContext can be used to purify privacy leak detection results for automatic and easy interpretation by filtering benign privacy disclosures. Experiments show PrivacyContext can perform an effective and efficient static privacy disclosure analysis enhancement and identify malicious privacy leak with 92.73% true positive rate. Evaluation also indicates that to keep the accuracy of privacy disclosure classification, our proposed contexts are all necessary. Xiaolei Wang 0003, Yuexiang Yang |
Int. J. Inf. Comput. Secur. | 2 |
| 2019 | Automated Hybrid Analysis of Android Malware through Augmenting Fuzzing with Forced ExecutionabstractAutomatically triggering malicious behaviors is an essential step to understand malware for developing effective solutions. Existing automated dynamic analysis approaches usually try to trigger the malicious behaviors by relying on simple fuzzing or complex input generation techniques (e.g., concolic execution). However, advanced malware often adopt various evasion techniques to hide malicious behaviors, e.g., by introducing complex condition checks which are very hard to trigger. In this paper, we propose a new approach named DirectDroid, which bypasses related checks through on-demand forced execution while adopting fuzzingto feed the necessary program input. In this way, many hidden malicious behaviors can be successfully triggered. To ensure the normal execution towards the malicious behaviors, DirectDroid also largely handles potential program crashes caused by forced execution. Finally, we implement a prototype of DirectDroid and evaluate it against 951 recent malware samples. Our experiment results show that DirectDroid can trigger many more malicious behaviors than several previous works, even when crashes happened. Our further analysis shows that DirectDroid has a low false positive rate even though it adopts forced execution. Xiaolei Wang 0003, Yuexiang Yang, Sencun Zhu |
IEEE Trans. Mob. Comput. | 2 |
| 2017 | Droid-AntiRM: Taming Control Flow Anti-analysis to Support Automated Dynamic Analysis of Android MalwareabstractWhile many test input generation techniques have been proposed to improve the code coverage of dynamic analysis, they are still inefficient in triggering hidden malicious behaviors protected by anti-analysis techniques. In this work, we design and implement Droid-AntiRM, a new approach seeking to tame anti-analysis automatically and improve automated dynamic analysis. Our approach leverages three key observations: 1) Logic-bomb based anti-analysis techniques control the execution of certain malicious behaviors; 2) Anti-analysis techniques are normally implemented through condition statements; 3) Anti-analysis techniques normally have no dependence on program inputs. Based on these observations, Droid-AntiRM uses various techniques to detect anti-analysis in malware samples, and rewrite the condition statements in anti-analysis cases through bytecode instrumentation, thus forcing the hidden behavior to be executed at runtime. Through a study of 3187 malware samples, we find that 32.50% of them employ various anti-analysis techniques. Our experiments demonstrate that Droid-AntiRM can identify anti-analysis instances from 30 malware samples with a true positive rate of 89.15% and zero false negative. By taming the identified anti-analysis, Droid-AntiRM can greatly improve the automated dynamic analysis, successfully triggering 44 additional hidden malicious behaviors from the 30 samples. Further performance evaluation shows that Droid-AntiRM has good efficiency to perform large-scale analysis. Xiaolei Wang 0003, Sencun Zhu, Dehua Zhou, Yuexiang Yang |
ACSAC | 4 |
| 2017 | Software Optimizations of Multiple Sets Intersection via SIMD InstructionsabstractConjunctive Boolean query is one fundamental operation for document retrieval in many information systems and databases. In its most basic and popular form, a conjunctive query can be seen as the intersection problem of multiple sets of sorted integers. Various algorithms have been put up in terms of maximizing the query efficiency. In recent years, researchers began to exploit the parallel advantage of single-instruction-multiple-data (SIMD) instructions to accelerate the intersection procedure and achieved substantial gains over previous scalar algorithms. However, these works only focus on intersecting two sets at a time and ignore the scenario of multiple sets intersection. Missing from the literature is a thorough study that explores the combination of traditional multiple sets intersection algorithms and SIMD instructions. This article discusses software optimizations for the intersection algorithms via AVX2 and AVX512 SIMD instructions of modern processor architectures. Through an experimental analysis we show that the proposed is able to reduce comparisons executed while improving instruction throughput, thus gaining performance enhancement over previous methods. Xingshen Song, Yuexiang Yang |
APSEC | 2 |
| 2017 | SIMD-Based Multiple Sets Intersection with Dual-Scale Search AlgorithmabstractConjunctive Boolean query is one fundamental operation for document retrieval in many information systems and databases. Various algorithms have been put up in terms of maximizing the query efficiency. In recent years, researchers began to exploit the parallel advantage of single-instruction-multiple-data (SIMD) instructions to accelerate the intersection procedure and achieved substantial gains over previous scalar algorithms. However, these works only focus on intersecting two sets at a time and ignore the scenario of multiple sets intersection. We present a flexible search algorithm which balances non-SIMD and SIMD comparisons in order to provide efficient and effective intersection. Xingshen Song, Yuexiang Yang, Xiaoyong Li 0002 |
CIKM | 2 |
| 2016 | On Optimizing Partitioning Strategies for Faster Inverted Index Compression
Xingshen Song, Kun Jiang 0001, Yuexiang Yang |
ICCSA (4) | 4 |
| 2016 | Architecture Support for Controllable VMI on Untrusted Cloud
Jiangyong Shi, Yuexiang Yang |
SecureComm | 2 |
| 2016 | IacCE: Extended Taint Path Guided Dynamic Analysis of Android Inter-App Data Leakage
Tianjun Wu, Yuexiang Yang |
SecureComm | 2 |
| 2016 | Efficient dynamic pruning on largest scores first (LSF) retrievalabstractInverted index traversal techniques have been studied in addressing the query processing performance challenges of web search engines, but still leave much room for improvement. In this paper, we focus on the inverted index traversal on document-sorted indexes and the optimization technique called dynamic pruning, which can efficiently reduce the hardware computational resources required. We propose another novel exhaustive index traversal scheme called largest scores first (LSF) retrieval, in which the candidates are first selected in the posting list of important query terms with the largest upper bound scores and then fully scored with the contribution of the remaining query terms. The scheme can effectively reduce the memory consumption of existing term-at-atime (TAAT) and the candidate selection cost of existing document-at-a-time (DAAT) retrieval at the expense of revisiting the posting lists of the remaining query terms. Preliminary analysis and implementation show comparable performance between LSF and the two well-known baselines. To further reduce the number of postings that need to be revisited, we present efficient rank safe dynamic pruning techniques based on LSF, including two important optimizations called list omitting (LSF_LO) and partial scoring (LSF_PS) that make full use of query term importance. Finally, experimental results with the TREC GOV2 collection show that our new index traversal approaches reduce the query latency by almost 27% over the WAND baseline and produce slightly better results compared with the MaxScore baseline, while returning the same results as exhaustive evaluation. Kun Jiang 0001, Yuexiang Yang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2015 | Fine-grained P2P traffic classification by simply counting flowsabstractThe continuous emerging of peer-to-peer (P2P) applications enriches resource sharing by networks, but it also brings about many challenges to network management. Therefore, P2P applications monitoring, in particular, P2P traffic classification, is becoming increasingly important. In this paper, we propose a novel approach for accurate P2P traffic classification at a fine-grained level. Our approach relies only on counting some special flows that are appearing frequently and steadily in the traffic generated by specific P2P applications. In contrast to existing methods, the main contribution of our approach can be summarized as the following two aspects. Firstly, it can achieve a high classification accuracy by exploiting only several generic properties of flows rather than complicated features and sophisticated techniques. Secondly, it can work well even if the classification target is running with other high bandwidth-consuming applications, outperforming most existing host-based approaches, which are incapable of dealing with this situation. We evaluated the performance of our approach on a real-world trace. Experimental results show that P2P applications can be classified with a true positive rate higher than 97.22% and a false positive rate lower than 2.78%. Jie He 0002, Yuexiang Yang, Yong Qiao, Wenping Deng |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2014 | Faster MaxScore Query Processing with Essential List Skipping
Kun Jiang 0001, Yuexiang Yang |
ADMA | 2 |
| 2014 | Identifying P2P Network Activities on Encrypted TrafficabstractPeer-to-Peer (P2P) traffic has always been a dominant portion of current Internet traffic and become more and more difficult to manage for Internet Service Producers (ISP) and network administrators. Although many methods have been proposed to classify different types of P2P applications and achieved satisfied performance, research on identifying network activities of a certain P2P application is still lacking to the best of our knowledge, which is urgently required in the context of forensic investigation for illegal P2P applications. In this paper, a novel approach based on Hidden Markov Model is proposed to identify network activities on the encrypted traffic, based on analysis of the time series characteristics and statistical properties of network traffic. After presenting a general model of network activities, Team Viewer is selected as a case study to verify the effectiveness of the approach to identify different activities. According to experiments using real network traces, our approach proves to be effective in identifying different activities of a P2P application with a high true positive 99.1% and low negligible false positive 3.6%. Xiaolei Wang 0003, Jie He 0002, Yuexiang Yang |
TrustCom | 3 |
| 2014 | Faster MaxScore Document Retrieval with Aggressive Processing
Kun Jiang 0001, Xingshen Song, Yuexiang Yang |
WAIM | 3 |
| 2013 | Detecting P2P bots by mining the regional periodicityabstractPeer-to-peer (P2P) botnets outperform the traditional Internet relay chat (IRC) botnets in evading detection and they have become a prevailing type of threat to the Internet nowadays. Current methods for detecting P2P botnets, such as similarity analysis of network behavior and machine-learning based classification, cannot handle the challenges brought about by different network scenarios and botnet variants. We noticed that one important but neglected characteristic of P2P bots is that they periodically send requests to update their peer lists or receive commands from botmasters in the command-and-control (C&C) phase. In this paper, we propose a novel detection model named detection by mining regional periodicity (DMRP), including capturing the event time series, mining the hidden periodicity of host behaviors, and evaluating the mined periodic patterns to identify P2P bot traffic. As our detection model is built based on the basic properties of P2P protocols, it is difficult for P2P bots to avoid being detected as long as P2P protocols are employed in their C&C. For hidden periodicity mining, we introduce the so-called regional periodic pattern mining in a time series and present our algorithms to solve the mining problem. The experimental evaluation on public datasets demonstrates that the algorithms are promising for efficient P2P bot detection in the C&C phase. Yong Qiao, Yuexiang Yang, Jie He 0002, Chuan Tang, Yingzhi Zeng |
J. Zhejiang Univ. Sci. C | 2 |