VLDB 2026 Research / reviewers in the wild / expert
Xuyang Jing
dblp:194/6281
· DBLP profile ↗
13ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0003-2274-0969ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GeminiSketch: An Accurate and Efficient Sketch for Summarizing Temporal Graph Streams with Rolling-Out Elimination
Xuyang Jing, Zheng Yan 0002, Qingze Jiang, Witold Pedrycz |
ICDE | 1 |
| 2026 | Octopus: A Robust and Privacy-Preserving Scheme for Compressed Gradients in Federated LearningabstractFederated learning is a distributed machine learning framework that allows multiple parties to collaboratively train a shared model without the need to share their original data with a central server. This approach minimizes the risk of data leakage and effectively addresses the challenge of isolated data silos. However, it necessitates multiple rounds of interactions to transmit models or gradients between the client and the server, leading to a significant communication cost and private data leakage. Consequently, some schemes compress the transmitted models or gradients through model pruning or knowledge distillation. However, they often overlook the privacy implications of compressed gradients or models, potentially resulting in privacy breaches. Additionally, they are susceptible to client disconnections, resulting in incomplete or delayed model updates. To address these challenges, this paper proposes Octopus, a robust and privacy-preserving scheme for compressed gradients in federated learning. Octopus employs Sketch to compress gradients and embeds masks for the compressed gradients, thereby safeguarding the gradients while concurrently reducing communication overhead. Moreover, we propose an anti-disconnection strategy to support model updates even in situations where some clients are disconnected. Lastly, we carry out comprehensive security and convergence analyses, along with extensive performance evaluations, demonstrating Octopus's robustness, stability, and efficiency over existing schemes. Wenxiu Ding, Yuxuan Xiao 0001, Zheng Yan 0002, Ciwei Chen, Xuyang Jing |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | MDSF-YOLO: Advancing Object Detection With a Multiscale Dilated Sequence Fusion NetworkabstractAccurate and fast detection of traffic signs is critical for autonomous driving, particularly in complex environments with diverse sign scales and varying detection distances. Existing approaches, incorporating attention modules or modifying detection heads, frequently encounter high rates of false positives and omissions due to the increased sampling depth. To address these limitations, we propose MDSF-you only look once (YOLO), a novel detection framework that integrates multiscale sequence fusion (MSF) for synergistic feature integration across granularities, enhancing the precision of both localization and semantic information fusion. Additionally, our dilated-wise residual (DWR) module leverages dilated convolutions and channel-wise reparameterization to improve fine-grained feature extraction. The architecture further introduces a $P_{2}$ detection head for shallow features and fully decouples all detection heads, optimizing target localization and category identification. Extensive experiments on the TT100K and CCTSDB2021 datasets demonstrate the superiority of MDSF-YOLO over benchmark models, including YOLOv11s, with significant improvements in mAP by 8.8% and 2.4% on respective datasets while substantially reducing false positives and leakage rate. Besides, the marked improvement of MDSF-YOLO on the VisDrone2019 dataset verifies its enhanced capability to address drone-based object detection. These advances underscore the efficiency and robustness of the proposed model, providing a promising solution for autonomous driving and similar object detection scenarios. Chong Zhang 0003, Xuyang Jing, Qing-Guo Wang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | TardySketch: A Framework for Cardinality Estimation Adaptable to Sliding WindowsabstractSliding cardinality estimation is crucial in many data analysis scenarios, e.g., detecting abnormal network behav-iors by monitoring unique connections in real time, detecting fraud in online transactions by monitoring unique user behavior patterns, and improving inventory management in supply chains by analyzing unique buyer behaviors. However, existing sliding cardinality estimation methods suffer from a cardinality barrel-down problem caused by unexpired item elimination in advance and item excessive removal, which remains unresolved so far. In this paper, we propose TardySketch, a sketch framework to make sliding cardinality estimation accurate and efficient by solving the above problem. The cornerstone of TardySketch is a Bidirectional Pointer-based Bitmap (BP-Bitmap), which stores the arrival sequence of items without timestamps. To prevent the premature elimination of unexpired items, we propose a Gap mechanism to enhance the accuracy of BP-Bitmap for identifying truly expired items through intermittent monitoring. To ensure an appropriate number of items are eliminated as the window moves, we design a Slow-Down mechanism to slacken the reset rate of bucket in BP- Bitmap to prevent over removal of items. Experimental results based on real-world datasets demonstrate that TardySketch significantly outperforms state-of-the-art methods, achieving a performance improvement of 5–40 times. The source code of TardySketch is available on GitHub. Xuyang Jing, Qinghua Cao, Zheng Yan 0002, Wenxiu Ding, Witold Pedrycz, Pu Wang 0003 |
ICDE | 1 |
| 2025 | LocalSketch: An Accurate and Efficient Sketch for Range Spread EstimationabstractSketch demonstrates good properties in spread estimation over network measurements, providing fast processing and accurate estimation under limited memory usage. However, most current methods remain limited to single-flow spread estimation, resulting in suboptimal performance when applied to range spread estimation that requires measuring the spread of a range of flows. In this paper, we propose LocalSketch, a novel sketch that achieves both high estimation accuracy and memory efficiency for range spread estimation with provable theoretical guarantees. LocalSketch has two key innovations: (1) local key aggregation within predefined ranges that eliminates duplicate spread information through locality correlation, and (2) adaptive counter sizing that dynamically allocates memory resources for large-spread ranges while maintaining compact representations for low-spread ranges. LocalSketch also features an efficient abnormal bucket detection mechanism by comparing identification sign, avoiding exhaustive bucket traversal during super range detection. Moreover, the main idea of LocalSketch can be adapted to existing plug-in spread counters, which has been experimentally proved. We provide a theoretical analysis of estimation accuracy and conduct comprehensive evaluations using real-world network traffic datasets. Experimental results demonstrate that LocalSketch outperforms state-of-the-art methods by achieving 76× higher estimation accuracy for range spread estimation, while showing 15× better accuracy and 39× faster detection speed for super range identification across all datasets. Xuyang Jing, Qinghua Cao, Zheng Yan 0002, Witold Pedrycz, Pu Wang 0003 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2023 | ALSketch: An adaptive learning-based sketch for accurate network measurement under dynamic traffic distribution
Xiaojun Cheng, Xuyang Jing, Zheng Yan 0002, Pu Wang 0003 |
J. Netw. Comput. Appl. | 2 |
| 2022 | Deceiving Learning-based Sketches to Cause Inaccurate Frequency EstimationabstractLearning-based sketches have been widely studied as an improvement of traditional sketches that achieves high efficiency in terms of both time and space. It uses a learning model to reveal and exploit underlying patterns of input data for helping traditional sketches obtain accurate frequency estimation with memory efficient. However, recent studies only focus on the performance improvement of learning-based sketches and pay little attention to security. The potential security problems can be easily exploited by an adversary to make learning-based sketches inaccurate. In this paper, we firstly explore the security issues of learning-based sketches with regard to estimation accuracy and memory overhead. Some adversarial scenarios of learning model and backup sketch are modeled according to the knowledge and capabilities of an adversary. Then, we propose four attacks to deceive learning-based sketch, namely counterfeit attack, targeted point attack, memory occupation attack, and blind increment attack. We conduct a series of experiments based on real-world datasets and verify that the proposed attacks highly degrade the performance of learning-based sketch even when the adversary knows nothing about it. Xuyang Jing, Xiaojun Cheng, Zheng Yan 0002 |
TrustCom | 1 |
| 2022 | A Group-Based Distance Learning Method for Semisupervised Fuzzy ClusteringabstractLearning a proper distance for clustering from prior knowledge falls into the realm of semisupervised fuzzy clustering. Although most existing learning methods take prior knowledge (e.g., pairwise constraints) into account, they pay little attention to local knowledge of data, which, however, can be utilized to optimize the distance. In this article, we propose a novel distance learning method, which learns from the Group-level information, for semisupervised fuzzing clustering. We first present a new format of constraint information, called Group-level constraints, by elevating the pairwise constraints (must-links and cannot-links) from point level to Group level. The Groups, generated around data points contained in the pairwise constraints, carry not only the local information of data (the relation between close data points) but also more background information under some given limited prior knowledge. Then, we propose a novel method to learn a distance by using the Group-level constraints, namely, Group-based distance learning, in order to optimize the performance of fuzzy clustering. The distance learning process aims to pull must-link Groups as close as possible while pushing cannot-link Groups as far as possible. We formulate the learning process with the weights of constraints by invoking some linear and nonlinear transformations. The linear Group-based distance learning method is realized by means of semidefinite programming, and the nonlinear learning method is realized by using the neural network, which can explicitly provide nonlinear mappings. Experimental results based on both synthetic and real-world datasets show that the proposed methods yield much better performance compared to other distance learning methods using pairwise constraints. Xuyang Jing, Zheng Yan 0002, Yinghua Shen, Witold Pedrycz |
IEEE Trans. Cybern. | 1 |
| 2022 | SuperSketch: A Multi-Dimensional Reversible Data Structure for Super Host IdentificationabstractFacing big network traffic data, effective data compression becomes crucially important and urgently needed for estimating host cardinalities and identifying super hosts. However, the current literature confronts several challenges: incapability of simultaneously measuring various types of host cardinalities and inability to efficiently reconstruct super host addresses. To address these challenges, in this article, we propose a novel sketch data structure, named SuperSketch, to simultaneously measure multiple types of host cardinalities with the purpose of efficiently identifying super hosts. SuperSketch has two significant characteristics: multi-dimensionality and reversibility. The multi-dimensionality makes SuperSketch capable of simultaneously measuring Source Cardinality, Destination Cardinality, and Destination Port Cardinality. The reversibility allows SuperSketch to accurately and quickly reconstruct the original addresses of super hosts once they are identified. We conduct both theoretical analysis and performance evaluation based on real-world network traffic. Experimental results show that SuperSketch achieves outstanding performance for multi-cardinality measurement, super host identification, and host address reconstruction. Xuyang Jing, Zheng Yan 0002, Witold Pedrycz |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | ExtendedSketch: Fusing Network Traffic for Super Host Identification With a Memory Efficient SketchabstractSuper host refers to the host that has a high cardinality or exhibits a big change in a network. Facing big-volume network traffic, sketches have been widely applied to identify super hosts in an efficient and accurate way. However, most sketches cannot flexibly balance memory usage and accuracy in host cardinality estimation. Setting an inappropriate counter size for a sketch could either lead to inaccurate host cardinality estimation or cause memory waste. In order to solve this issue, we propose a novel extensible and reversible sketch, named ExtendedSketch, to achieve accurate super host identification with high memory efficiency. The core idea of ExtendedSketch is to monitor low-cardinality hosts with small-sized counters while dynamically extending the size of counters when monitoring high-cardinality hosts by applying an adaptive extension strategy. Such the strategy can adaptively increase counter size according to network traffic status at runtime, which not only ensures the accuracy of high-cardinality host estimation but also avoids unnecessary memory consumption. We perform theoretical analysis and conduct a series of experimental evaluations on ExtendedSketch based on real world network traffic. Experimental results show that under same memory usage, compared to the state-of-the-art, ExtendedSketch achieves$1.4{ \sim }7.5$times smaller error rate in estimating host cardinality with$1.9{ \sim }26.7$times better accuracy on super host identification and$95 {\sim }2^{15}$times faster speed on abnormal address reconstruction. Its advance in accuracy and efficiency demonstrates the practical significance of ExtendedSketch for super host identification. Xuyang Jing, Zheng Yan 0002, Witold Pedrycz |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | Identification of Fuzzy Rule-Based Models With Output Space Knowledge GuidanceabstractIn this article, we advocate that a knowledge tidbit residing in the output space could be helpful in improving the performance (accuracy) of the fuzzy rule-based model. It states thatif two outputs are far apart from each other,it is advisable to place their corresponding inputs in different clusters when forming subspaces of the input space. Considering this knowledge guidance mechanism, we propose two different methods to partition the input space. In the first method, input data are first partitioned with the use of the standard clustering algorithm, say fuzzy C-means; here, a constructed partition matrix is reflective of the structure present in the input space. Then, the knowledge tidbit is used to adjust the entries of the original partition matrix in such a way that those input data whose corresponding output data are far apart from each other are assigned with low values of proximity. In the second method, we propose two strategies to modify the distance between input data and a prototype (cluster center) identified in the input space. The crux of this method is that if there are many input data (which, in virtue of the knowledge tidbit, are regarded as being far-apart from the input data of interest) around a certain prototype, the distance between the input data of interest and this prototype should be penalized. Thus, the membership of these input data to the prototype is reduced. The comprehensive experimental studies carried out on both synthetic and publicly available data are used to examine the usefulness of the proposed methods. Yinghua Shen, Witold Pedrycz, Xuyang Jing, Adam Gacek, Xianmin Wang, Bingsheng Liu |
IEEE Trans. Fuzzy Syst. | 3 |
| 2020 | An Adaptive Security Data Collection and Composition Recognition method for security measurement over LTE/LTE-A networks
Hanlu Chen, Zheng Yan 0002, Raimo Kantola, Xuyang Jing, Jin Cao 0001, Hui Li 0006 |
J. Netw. Comput. Appl. | 6 |
| 2019 | A reversible sketch-based method for detecting and mitigating amplification attacks
Xuyang Jing, Zheng Yan 0002, Witold Pedrycz |
J. Netw. Comput. Appl. | 1 |