VLDB 2026 Research / reviewers in the wild / expert
Xuying Meng
dblp:146/8088
· DBLP profile ↗
27ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-0095-4811ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpecDetect: Simple, Fast, and Training-Free Detection of LLM-Generated Text via Spectral AnalysisabstractThe proliferation of high-quality text from Large Language Models (LLMs) demands reliable and efficient detection methods. While existing training-free approaches show promise, they often rely on surface-level statistics and overlook fundamental signal properties of the text generation process. In this work, we reframe detection as a signal processing problem, introducing a novel paradigm that analyzes the sequence of token log-probabilities in the frequency domain. By systematically analyzing the signal's spectral properties using the global Discrete Fourier Transform (DFT) and the local Short-Time Fourier Transform (STFT), we find that human-written text consistently exhibits significantly higher spectral energy. This higher energy reflects the larger-amplitude fluctuations inherent in human writing compared to the suppressed dynamics of LLM-generated text. Based on this key insight, we construct SpecDetect, a detector built on a single, robust feature from the global DFT: DFT total energy. We also propose an enhanced version, SpecDetect++, which incorporates a sampling discrepancy mechanism to further boost robustness. Extensive experiments show that our approach outperforms the state-of-the-art model while running in nearly half the time. Our work introduces a new, efficient, and interpretable pathway for LLM-generated text detection, showing that classical signal processing techniques offer a surprisingly powerful solution to this modern challenge. Haitong Luo, Weiyao Zhang, Suhang Wang, Wenji Zou, Chungang Lin, Xuying Meng, Yujun Zhang 0001 |
AAAI | 6 |
| 2026 | Heterogeneity-Oblivious Robust Federated LearningabstractFederated Learning (FL) remains highly vulnerable to poisoning attacks, especially under real-world hyper-heterogeneity, where clients differ significantly in data distributions, communication capabilities, and model architectures. Such heterogeneity not only undermines the effectiveness of aggregation strategies but also makes attacks more difficult to detect. Furthermore, high-dimensional models expand the attack surface. To address these challenges, we propose Horus, a heterogeneity-oblivious robust FL framework centered on low-rank adaptations (LoRAs). Rather than aggregating full model parameters, Horus inserts LoRAs into empirically stable layers and aggregates only LoRAs to reduce the attack uncover a key empirical observation that the input projection (LoRA-A) is markedly more stable than the output projection (LoRA-B) under heterogeneity and poisoning. Leveraging this, we design a Heterogeneity-Oblivious Poisoning Score using the features from LoRA-A to filter poisoned clients. For the remaining benign clients, we propose projection-aware aggregation mechanism to preserve collaborative signals while suppressing drifts, which reweights client updates by consistency with the global directions. Extensive experiments across diverse datasets, model architectures, and attacks demonstrate that Horus consistently outperforms state-of-the-art baselines in both robustness and accuracy. Weiyao Zhang, Chungang Lin, Haitong Luo, Xuying Meng |
INFOCOM | 7 |
| 2026 | Enhance graph alignment for large language modelsabstractGraph-structured data is prevalent in the real world. Recently, due to the powerful emergent capabilities, Large Language Models (LLMs) have shown promising performance in modeling graphs. The key to effectively applying LLMs on graphs is converting graph data into a format LLMs can comprehend. Graph-to-token approaches are popular in enabling LLMs to process graph information. They transform graphs into sequences of tokens and align them with text tokens through instruction tuning, where self-supervised instruction tuning helps LLMs acquire general knowledge about graphs, and supervised fine-tuning specializes LLMs for the downstream tasks on graphs. Despite their initial success, we find that existing methods have a misalignment between self-supervised tasks and supervised downstream tasks, resulting in negative transfer from self-supervised fine-tuning to downstream tasks. To address these issues, we propose Graph Alignment Large Language Models (GALLM) to benefit from aligned task templates. In the self-supervised tuning stage, we introduce a novel text matching task using templates aligned with downstream tasks. In the task-specific tuning stage, we propose two category prompt methods that learn supervision information from additional explanation with further aligned templates. Experimental evaluations on four datasets demonstrate substantial improvements in supervised learning, multi-dataset generalizability, and particularly in zero-shot capability, highlighting the model's potential as a graph foundation model. Our code is available at the anonymous repository https://anonymous.4open.science/r/GALLM-AC54/. Haitong Luo, Xuying Meng, Suhang Wang, Tianxiang Zhao 0001, Fali Wang, Yujun Zhang 0001 |
Neural Networks | 2 |
| 2026 | Dynamic Federated Edge Anomaly Detection Based on Dual Fuzzing StrategiesabstractWith the rapid proliferation of edge computing, mobile edges interconnected via networks are increasingly ex posed to a wide range of attacks. Considering privacy concerns, federated edge anomaly detection, collaboratively training a global detection model across multiple edge clients without centralizing sensitive data, is a promising paradigm to detecting such attacks. However, mobile edge environment is inherently dynamic, with frequent evolving in both network traffic and participating clients. Most existing methods are pre-defined for specific attacks and fixed client settings, which significantly limits their adaptability in dynamic mobile edges. To address these issues, we define the federated adaptability from both data-level and client-level perspectives for the first time, and derive the decisive factors for improving adaptability, i.e., high smoothness and low gradient similarity. Driven by this theoretical foundation, we propose DIFF, a federated edge anomaly detection framework based on dual fuzzing strategies. DIFF incorporates two novel designs for improving adaptability: (i) a classless fuzzing strategy to improve data-level adaptability by fuzzing the boundary between normal and abnormal samples, thereby encouraging model smoothness and enhancing adaptability to unseen traffic and emerging attacks; and (ii) a direction fuzzing strategy to enhance client-level adaptability by perturbing the optimization directions of local models, enabling the aggregated global model to adapt effectively to new clients. Experiments on public datasets demonstrate the superior adapt ability of DIFF compared to the state-of-art methods. Weiyao Zhang, Jinyang Li 0009, Botao Peng, Xuying Meng, Yujun Zhang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | PGAE: A Perturbed Graph Autoencoder Integrating Explicit and Implicit Features for APT Detection
Chunmiao Xiang, Mengyao Han, Xuying Meng, Wenli Song |
ICIC (4) | 4 |
| 2025 | Not All Layers of LLMs Are Necessary During InferenceabstractDue to the large number of parameters, the inference phase of Large Language Models (LLMs) is resource-intensive. However, not all requests posed to LLMs are equally difficult to handle. Through analysis, we show that for some tasks, LLMs can achieve results comparable to the final output at some intermediate layers. That is, not all layers of LLMs are necessary during inference. If we can predict at which layer the inferred results match the final results (produced by evaluating all layers), we could significantly reduce the inference cost. To this end, we propose a simple yet effective algorithm named AdaInfer to adaptively terminate the inference process for an input instance. AdaInfer relies on easily obtainable statistical features and classic classifiers like SVM. Experiments on well-known LLMs like the Llama2 series and OPT, show that AdaInfer can achieve an average of 17.8% pruning ratio, and up to 43% on sentiment tasks, with nearly no performance drop (<1%). Because AdaInfer does not alter LLM parameters, the LLMs incorporated with AdaInfer maintain generalizability across tasks. Siqi Fan 0001, Xin Jiang 0005, Xiang Li 0001, Xuying Meng, Peng Han 0005, Shuo Shang, Aixin Sun, Yequan Wang |
IJCAI | 4 |
| 2025 | TraGe: A Generic Packet Representation for Traffic Classification Based on Header-Payload DifferencesabstractTraffic classification has a significant impact on maintaining the Quality of Service (QoS) of the network. Since traditional methods heavily rely on feature extraction and largescale labeled data, some recent pre-trained models manage to reduce the dependency by utilizing different pre-training tasks to train generic representations for network packets. However, existing pre-trained models typically adopt pre-training tasks developed for image or text data, which are not tailored to traffic data. As a result, the obtained traffic representations fail to fully reflect the information contained in the traffic, and may even disrupt the protocol information. To address this, we propose TraGe, a novel generic packet representation model for traffic classification. Based on the differences between the header and payload-the two fundamental components of a network packetwe perform differentiated pre-training according to the byte sequence variations (continuous in the header vs. discontinuous in the payload). A dynamic masking strategy is further introduced to prevent overfitting to fixed byte positions. Once the generic packet representation is obtained, TraGe can be finetuned for diverse traffic classification tasks using limited labeled data. Experimental results demonstrate that TraGe significantly outperforms state-of-the-art methods on two traffic classification tasks, with up to a 6.97% performance improvement. Moreover, TraGe exhibits superior robustness under parameter fluctuations and variations in sampling configurations. Chungang Lin, Yilong Jiang, Weiyao Zhang, Xuying Meng, Tianyu Zuo, Yujun Zhang 0001 |
IWQoS | 4 |
| 2025 | LogOW: A semi-supervised log anomaly detection model in open-world setting
Jingwei Ye, Zhaojun Gu, Xuying Meng, Weiyao Zhang, Yujun Zhang 0001 |
J. Syst. Softw. | 5 |
| 2025 | Burst-Sensitive Traffic Forecast via Multi-Property Personalized Fusion in Federated LearningabstractFor distributed network traffic prediction with data localization and privacy protection, Federated Learning (FL) enables collaborative training without raw data exchange across Base Stations (BSs). Nevertheless, traffic across BSs exhibit inherently heterogeneous trend burst and smooth fluctuation properties, but existing FL methods model single-scale series from only one view, which cannot simultaneously capture diverse trend and fluctuation properties, especially distinct burst distributions. In this paper, we proposePersonalized Federated Forecasting with Multi-property Self-fusion (P2FMS), which can represent multi-scale traffic properties from different views. With precise multi-property representations, a fusion-level prediction decision is learned for each client in a personalized manner to promptly sense traffic bursts and improve forecasting performance in non-IID settings. Specifically, P2FMS decomposes the traffic series into distinct time scales, based on which, we effectively extract closeness, period, and trend properties from different views. The closeness and period are embedded through global-view representations with spatial correlations, while non-stationary trends are individually fitted from the client-side view. Furthermore, a personalized combiner is designed to accurately quantify the proportion of general fluctuation raws (i.e., closeness and period) and specific trend property in predictions, which enables multi-property self-fusion for each client to accommodate heterogeneous traffic patterns and enhance prediction accuracy. Besides, an alternant training mechanism is introduced to optimize property representation and fusion modules with the convergence guarantee. Extensive experiments on real-world datasets show that P2FMS outperforms status quo methods in both prediction performance and convergence time. Min Liu 0001, Yuwei Wang 0003, Xuying Meng, Jingyuan Wang 0001, Junbo Zhang 0004, Ke Xu 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Spectral-Based Graph Neural Networks for Complementary Item RecommendationabstractModeling complementary relationships greatly helps recommender systems to accurately and promptly recommend the subsequent items when one item is purchased. Unlike traditional similar relationships, items with complementary relationships may be purchased successively (such as iPhone and Airpods Pro), and they not only share relevance but also exhibit dissimilarity. Since the two attributes are opposites, modeling complementary relationships is challenging. Previous attempts to exploit these relationships have either ignored or oversimplified the dissimilarity attribute, resulting in ineffective modeling and an inability to balance the two attributes. Since Graph Neural Networks (GNNs) can capture the relevance and dissimilarity between nodes in the spectral domain, we can leverage spectral-based GNNs to effectively understand and model complementary relationships. In this study, we present a novel approach called Spectral-based Complementary Graph Neural Networks (SComGNN) that utilizes the spectral properties of complementary item graphs. We make the first observation that complementary relationships consist of low-frequency and mid-frequency components, corresponding to the relevance and dissimilarity attributes, respectively. Based on this spectral observation, we design spectral graph convolutional networks with low-pass and mid-pass filters to capture the low-frequency and mid-frequency components. Additionally, we propose a two-stage attention mechanism to adaptively integrate and balance the two attributes. Experimental results on four e-commerce datasets demonstrate the effectiveness of our model, with SComGNN significantly outperforming existing baseline models. Haitong Luo, Xuying Meng, Suhang Wang, Hanyun Cao, Weiyao Zhang, Yequan Wang, Yujun Zhang 0001 |
AAAI | 2 |
| 2024 | Self-supervised multi-view clustering in computer vision: A surveyabstractAbstract In recent years, multi‐view clustering (MVC) has had significant implications in the fields of cross‐modal representation learning and data‐driven decision‐making. Its main objective is to cluster samples into distinct groups by leveraging consistency and complementary information among multiple views. However, the field of computer vision has witnessed the evolution of contrastive learning, and self‐supervised learning has made substantial research progress. Consequently, self‐supervised learning is progressively becoming dominant in MVC methods. It involves designing proxy tasks to extract supervisory information from image and video data, thereby guiding the clustering process. Despite the rapid development of self‐supervised MVC, there is currently no comprehensive survey analysing and summarising the current state of research progress. Hence, the authors aim to explore the emergence of self‐supervised MVC by discussing the reasons and advantages behind it. Additionally, the internal connections and classifications of common datasets, data issues, representation learning methods, and self‐supervised learning methods are investigated. The authors not only introduce the mechanisms for each category of methods, but also provide illustrative examples of their applications. Finally, some open problems are identified for further investigation and development. Jiatai Wang, Xuying Meng |
IET Comput. Vis. | 6 |
| 2024 | DMSTG: Dynamic Multiview Spatio-Temporal Networks for Traffic ForecastingabstractTraffic sensor networks are widely applied in smart cities to monitor traffic in real-time and record huge volumes of traffic data. Exploiting such data to forecast future traffic conditions have the potential to enhance the decision-making capabilities of intelligent transportation systems, which attracts widespread attention from both industries and academia. Among them, network-wide prediction based on graph convolutional neural networks(GCN) has become mainstream. It models the spatial dependencies of sensors in a graph with a pre-defined Laplacian matrix based on the distances among sensors. However, understanding spatio-temporal traffic patterns is quite challenging as there is a huge difference in terms of traffic patterns during different periods or in different regions. In addition, the actual data collected can be polluted due to unavoidable data loss from severe communication conditions or sensor failures. Considering these issues, we propose a novel dynamic multiview spatial-temporal prediction framework which takes into consideration various factors, including local/global, short/long term spatio-temporal dependencies and their dynamic changes. To comprehensively track the dynamic spatio-temporal dependencies among traffic data, we creatively design two different modules to perceive the changes in traffic patterns. We first propose a dynamic Laplacian matrix learning module based on our theoretical derivation to estimate the Laplacian matrix of the graph for GCN timely. We creatively incorporate tensor decomposition into this module, where real-time traffic data are decomposed into a global component that is stable and depends on long-term temporal-spatial traffic relationships and a local component that captures the traffic fluctuations. We also design a self-attention based module to dynamically assign a weight to each part in traffic data. The spatio-temporal features from multiple views are deeply fused by a feature fusion module. The forecasting performance is evaluated with 5 real-time traffic datasets. Experiment results demonstrate that our framework can consistently outperform the state-of-the-art baselines and be more robust under noisy environments. Zulong Diao, Xin Wang 0001, Da-Fang Zhang 0001, Gaogang Xie, Jianguo Chen 0001, Changhua Pei, Xuying Meng, Kun Xie 0001, Guangxing Zhang |
IEEE Trans. Mob. Comput. | 7 |
| 2023 | Multi-Layer Collaborative Bandit for Multivariate Time Series Anomaly DetectionabstractMultivariate Time Series Anomaly Detection (MTSAD) detects abnormal indicators from Multivariate Time Series (MTS), and provides the rank of the multiple abnormal indicators to meet the expert's detection interest in current environment, which underpin the security and stability of intelligent cyber-physical systems. However, popular integration-based methods, which are pre-defined, fall short in locating the exact abnormal indicator, nor can they perceive the environmental dynamic and evolve accordingly. Let alone meeting the expert's interest. As a result, the expert's workload is exaggerated. These issues motivate us to propose a novel multi-layer collaborative bandit framework MULA for MTSAD. MULA decomposes MTS and pairs individual time series with a bandit arm, which locates the abnormal indicator directly. Then, MULA sorts the indicators by abnormal scores computed based on the expert's feedback, which facilitates the experts. Besides, to address the adaptability issue, we devise a dual signal to comprehensively monitor environmental changes, and design a multi-layer collaborative mechanism for MULA to adapt to the dynamic environment. Theoretical analysis and experiments on public datasets demonstrate the superiority of MULA compared to the state-of-art. Weiyao Zhang, Xuying Meng, Jinyang Li 0009, Yequan Wang, Yujun Zhang 0001 |
IWQoS | 2 |
| 2023 | EC-GCN: A encrypted traffic classification framework based on multi-scale graph convolution networks
Zulong Diao, Gaogang Xie, Xin Wang 0001, Xuying Meng, Guangxing Zhang, Kun Xie 0001, Mingyu Qiao |
Comput. Networks | 5 |
| 2022 | CofeNet: Context and Former-Label Enhanced Net for Complicated Quotation ExtractionabstractQuotation extraction aims to extract quotations from written text. There are three components in a quotation: source refers to the holder of the quotation, cue is the trigger word(s), and content is the main body. Existing solutions for quotation extraction mainly utilize rule-based approaches and sequence labeling models. While rule-based approaches often lead to low recalls, sequence labeling models cannot well handle quotations with complicated structures. In this paper, we propose the Context and Former-Label Enhanced Net () for quotation extraction. is able to extract complicated quotations with components of variable lengths and complicated structures. On two public datasets (and ) and one proprietary dataset (), we show that our achieves state-of-the-art performance on complicated quotation extraction. Yequan Wang, Xiang Li 0001, Aixin Sun, Xuying Meng, Huaming Liao, Jiafeng Guo |
COLING | 4 |
| 2022 | A High-performance FPGA-based Accelerator for Gradient CompressionabstractGradient compression technology has attracted much attention in recent years, due to its high effectiveness in alleviating the communication bottleneck of distributed deep learning. However, except for the communication reduction, it also brings in a significant increase of computational overhead, which limits or even eliminates the communication-reduction benefit brought by gradient compression. To solve the high computational overhead problem, we propose an FPGA-based accelerator for gradient compression in this paper. A high-performance and programmable accelerator architecture is developed for accelerating various gradient compression algorithms by offloading compute-intensive compression operations to FPGA. Also, we design and implement the FPGA-based accelerator based on the popular gradient compression algorithm top-k sparsification. Experimental results show that the new accelerator achieves up to hundreds of times faster than the compression algorithm implemented on CPU and GPU. What's more, the stable and controllable performance under different datasets demonstrates that the proposed accelerator is insensitive to data distribution, which is essential for time-sensitive applications. Qingqing Ren, Shuyong Zhu, Xuying Meng, Yujun Zhang 0001 |
DCC | 3 |
| 2022 | Dual-track Protocol Reverse Analysis Based on Share LearningabstractPrivate protocols, whose specifications are agnostic, are widely used in the Industrial Internet. While providing customized service, they also raise essential security concerns as well, due to their agnostic nature. The Protocol Reverse Analysis (PRA) techniques are developed to infer the specifications of private protocols. However, the conventional PRA techniques are far from perfection for the following reasons: (i) Error propagation: Canonical solutions strictly follow the "from keyword extraction to message clustering" serial structure, which deteriorates the performance for ignoring the interplay between the sub-tasks, and the error will flow and accumulate through the sequential workflow. (ii) Increasing diversity: As the protocols’ diversities of characteristics increase, tailoring for specific types of protocols becomes infeasible. To address these issues, we design a novel dual-track framework SPRA, and propose Share Learning, a new concept of protocol reverse analysis. Particularly, based on the share layer for protocol learning, SPRA builds a parallel workflow to co-optimize both the generative model for keyword extraction and the probability-based model for message clustering, which delivers automatic and robust syntax inference across diverse protocols and greatly improves the performance. Experiments on five real-world datasets demonstrate that the proposed SPRA achieves better performance compared with the state-of-art PRA methods. Weiyao Zhang, Xuying Meng, Yujun Zhang 0001 |
INFOCOM | 2 |
| 2022 | Packet Representation Learning for Traffic ClassificationabstractWith the surging development of information technology, to provide a high quality of network services, there are increasing demands and challenges for network analysis. As all data on the Internet are encapsulated and transferred by network packets, packets are widely used for various network traffic analysis tasks, from application identification to intrusion detection. Considering the choice of features and how to represent them can greatly affect the performance of downstream tasks, it is critical to learn high-quality packet representations. In addition, existing packet-level works ignore packet representations but focus on trying to get good performance with independent analysis of different classification tasks. In the real world, although a packet may have different class labels for different tasks, the packet representation learned from one task can also help understand its complex packet patterns in other tasks, while existing works omit to leverage them. Xuying Meng, Yequan Wang, Runxin Ma, Haitong Luo, Yujun Zhang 0001 |
KDD | 1 |
| 2022 | A Linear Time Approach to Computing Time Series Similarity Based on Deep Metric LearningabstractTime series similarity computation is a fundamental primitive that underpins many time series data analysis tasks. However, many existing time series similarity measures have a high computation cost. While there has been much research effort for reducing the computational cost, such effort is usually specific to one similarity measure. We proposeNeuTS(Neural metric learning forTimeSeries) to accelerate time series similarity computation in a generic fashion.NeuTScomputes the similarity of a given time series pair in linear time and generic to handle any existing similarity measures.NeuTSsamples a number of seed time series from the given database, and then uses their pair-wise similarities as guidance to approximate the similarity function with a neural metric learning framework.NeuTSfeatures two novel modules to achieve accurate approximation of the similarity function: (1) a local attention memory module that augments existing recurrent neural networks for time series encoding; and (2) a distance-weighted ranking loss that effectively transcribes information from the seed-based guidance. With these two modules,NeuTScan yield high accuracies and fast convergence rates even if the training data is small. Our experiments with five real-life datasets and four similarity measures (Fréchet, Hausdorff, ERP and DTW) show thatNeuTSoutperforms baselines consistently and significantly. Specifically, it achieves over 80 percent accuracies in most settings, while obtaining 50x-1000x speedup over bruteforce methods and 3x-350x speedup over approximate algorithms for top-k similarity search. Di Yao 0001, Gao Cong, Chao Zhang 0014, Xuying Meng, Rongchang Duan, Jingping Bi |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | Semi-supervised anomaly detection in dynamic communication networksabstractTo ensure the security and stabilization of the communication networks, anomaly detection is the first line of defense. However, their learning process suffers two major issues: (1) inadequate labels : there are many different kinds of attacks but rare abnormal nodes in mt of these atstacks; and (2) inaccurate labels : considering the heavy network flows and new emerging attacks, providing accurate labels for all nodes is very expensive. The inadequate and inaccurate label problem challenges many existing methods because the majority normal nodes result in a biased classifier while the noisy labels will further degrade the performance of the classifier. To tackle these issues, we propose SemiADC, a Semi -supervised A nomaly D etection framework for dynamic C ommunication networks. SemiADC first approximately learns the feature distribution of normal nodes with regularization from abnormal ones. It then cleans the datasets and extracts the nodes sasainaccurate labels by the learned feature distribution and structure-based temporal correlations. These self-learning processes run iteratively with mutual promotion, and finally help increase the accuracy of anomaly detection. Experimental evaluations on real-world datasets demonstrate the effectiveness of our SemiADC, which performs substantially better than the state-of-art anomaly detection approaches without the demand of adequate and accurate supervision. Xuying Meng, Suhang Wang, Zhimin Liang, Di Yao 0001, Jihua Zhou, Yujun Zhang 0001 |
Inf. Sci. | 1 |
| 2021 | Interactive Anomaly Detection in Dynamic Communication NetworksabstractNetwork flows are the basic components of the Internet. Considering the serious consequences of abnormal flows, it is crucial to provide timely anomaly detection in dynamic communication networks. To obtain accurate anomaly detection results in dynamic networks, supervision from experts is highly demanded. However, to obtain high-quality ground truth of abnormal flows, we suffer from two major problems: (1)limited labor resources: experts with the latest domain knowledge are much fewer than the large number of flows; and (2)dynamic environment: considering the new abnormal patterns (i.e., new attacks) and continuously changing network structures, it requires timely supervision to adaptively update the parameters. To tackle these problems, we propose HADDN, a novel bandit framework for periodic-updated anomaly detection in dynamic communication networks. We formulate the task as a bandit problem, where by interactions, supervision is offered by human experts to provide the ground truth to a fraction of flows. We construct semi-parametric expected rewards to optimize the estimation of flows’ abnormality in limited interactions. Also, we utilize feature-based clusters and structural correlations to make connections between historical flows and new flows to improve both efficiency and accuracy of abnormality estimation. What’s more, we provide two implementations for the semi-parametric expected reward of the proposed HADDN with theoretical proof. Experimental evaluations on public datasets demonstrate the substantial improvement of our proposed approaches compared to state-of-art anomaly detection methods. Xuying Meng, Yequan Wang, Suhang Wang, Di Yao 0001, Yujun Zhang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2019 | A Practical Semi-Parametric Contextual BanditabstractClassic multi-armed bandit algorithms are inefficient for a large number of arms. On the other hand, contextual bandit algorithms are more efficient, but they suffer from a large regret due to the bias of reward estimation with finite dimensional features. Although recent studies proposed semi-parametric bandits to overcome these defects, they assume arms' features are constant over time. However, this assumption rarely holds in practice, since real-world problems often involve underlying processes that are dynamically evolving over time especially for the special promotions like Singles' Day sales. In this paper, we formulate a novel Semi-Parametric Contextual Bandit Problem to relax this assumption. For this problem, a novel Two-Steps Upper-Confidence Bound framework, called Semi-Parametric UCB (SPUCB), is presented. It can be flexibly applied to linear parametric function problem with a satisfied gap-free bound on the n-step regret. Moreover, to make our method more practical in online system, an optimization is proposed for dealing with high dimensional features of a linear function. Extensive experiments on synthetic data as well as a real dataset from one of the largest e-commercial platforms demonstrate the superior performance of our algorithm. Miao Xie, Xuying Meng, Nan Li 0019, Rong Jin 0001 |
IJCAI | 4 |
| 2019 | Towards privacy preserving social recommendation under personalized privacy settings
Xuying Meng, Suhang Wang, Kai Shu, Jundong Li, Bo Chen 0028, Huan Liu 0001, Yujun Zhang 0001 |
World Wide Web | 1 |
| 2018 | Exploiting Emotion on Reviews for Recommender SystemsabstractReview history is widely used by recommender systems to infer users' preferences and help find the potential interests from the huge volumes of data, whereas it also brings in great concerns on the sparsity and cold-start problems due to its inadequacy. Psychology and sociology research has shown that emotion information is a strong indicator for users' preferences. Meanwhile, with the fast development of online services, users are willing to express their emotion on others' reviews, which makes the emotion information pervasively available. Besides, recent research shows that the number of emotion on reviews is always much larger than the number of reviews. Therefore incorporating emotion on reviews may help to alleviate the data sparsity and cold-start problems for recommender systems. In this paper, we provide a principled and mathematical way to exploit both positive and negative emotion on reviews, and propose a novel framework MIRROR, exploiting eMotIon on Reviews for RecOmmendeR systems from both global and local perspectives. Empirical results on real-world datasets demonstrate the effectiveness of our proposed framework and further experiments are conducted to understand how emotion on reviews works for the proposed framework. Xuying Meng, Suhang Wang, Huan Liu 0001, Yujun Zhang 0001 |
AAAI | 1 |
| 2018 | Personalized Privacy-Preserving Social RecommendationabstractPrivacy leakage is an important issue for social recommendation. Existing privacy preserving social recommendation approaches usually allow the recommender to fully control users' information. This may be problematic since the recommender itself may be untrusted, leading to serious privacy leakage. Besides, building social relationships requires sharing interests as well as other private information, which may lead to more privacy leakage. Although sometimes users are allowed to hide their sensitive private data using privacy settings, the data being shared can still be abused by the adversaries to infer sensitive private information. Supporting social recommendation with least privacy leakage to untrusted recommender and other users (i.e., friends) is an important yet challenging problem. In this paper, we aim to address the problem of achieving privacy-preserving social recommendation under personalized privacy settings. We propose PrivSR, a novel framework for privacy-preserving social recommendation, in which users can model ratings and social relationships privately. Meanwhile, by allocating different noise magnitudes to personalized sensitive and non-sensitive ratings, we can protect users' privacy against the untrusted recommender and friends. Theoretical analysis and experimental evaluation on real-world datasets demonstrate that our framework can protect users' privacy while being able to retain effectiveness of the underlying recommender system. Xuying Meng, Suhang Wang, Kai Shu, Jundong Li, Bo Chen 0028, Huan Liu 0001, Yujun Zhang 0001 |
AAAI | 1 |
| 2016 | Adaptively modeling multi-feature preferences for personalized searchabstractThis paper is concerned with the adaptation to multi-feature preferences on personalized search. In the existing work, the personalized search mainly leverages semantic features extracted from user history and ignores other non-semantic latent features, lets alone adapt to preference distribution on non-semantic features. To tackle this problem, we propose an adaptive model for multi-feature preferences, in which we adapt latent non-semantic features extracted from visited pages to reflect diverse aspects of user preferences. We also utilize a novel algorithm in our model to improve the adaptation for diverse preference distribution on these features for different users. Our experimental results demonstrate that our model can improve personalized search performance by enhanced adaptation to diverse user preferences. Xuying Meng, Miao Wang 0007, Hanwen Zhang 0001, Yujun Zhang 0001 |
ISCC | 1 |
| 2013 | A reputation based incentive mechanism for selfish BitTorrent systemabstractThe current BitTorrent-like file sharing systems suffer from peer selfish behaviors. The uncooperative peers can freeload compliant users by free-riding and exploiting. To study the performance of BitTorrent's embedded incentive mechanism against selfishness, a fluid model with three different classes of peers, namely normal peers, exploiters and free-riders, is established. We point out that the current BitTorrent system can not provide an effectively differentiated service in accordance with contribution of peers. Therefore, a reputation based incentive (RBI) mechanism for selfish BitTorrent system is proposed. RBI defines a trust value for each peer associative to its historical performance to the whole system. With the trust value, the choking mechanism is modified to ensure the more trustworthy peers will have more chances to get served. Our simulation study indicates that RBI mechanism can remarkably prevent exploiting behaviors, severely penalize free-riders, and thus result in a fairer allocation of bandwidth among peers. Miao Wang 0007, Yujun Zhang 0001, Xuying Meng |
GLOBECOM | 3 |