Xiaoyu Kou

dblp:242/9273 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
4since 2021 · last 2026
0009-0005-0855-0693ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 CoDa: Privacy-preserving multi-dimensional dataset publishing based on consistent data masking
Xiaoyu Kou, Hui Zhu 0001, Jiezhen Tang, Jiaqi Zhao 0005, Fengwei Wang, Hui Li 0006
Inf. Sci.1
2026 Plog: An Efficient and Privacy-Preserving Collaborative Learning Framework on Vertically Partitioned Graph Data
abstract
With the rapid advancement and widespread ap plication of the graph neural network (GNN), the collaborative graph learning, in which multiple parties collaboratively construct a GNN model using their respective graph data, has attracted increasing attention. However, this paradigm also raises significant privacy concerns, as both nodes and edges may contain sensitive personal information, while existing privacy preserving schemes often come at the cost of degraded model performance or substantial system overhead. Therefore, this paper proposes an efficient and privacy-preserving collaborative, and Hui Li, Member, IEEE, Xiaoyu Kou Social Platform learning framework on vertically partitioned graph data, dubbed Plog. Specifically, we first design a decomposition algorithm to split the sparse adjacency matrix into the summation of multiple independent permutations, which are lightweight, parallelizable, and well-suited for secure multi-party computation. Building on this, a weighted oblivious batch permutation protocol is carefully customized based on correlated randomness to securely and efficiently compute adjacency matrix multiplications, addressing the core efficiency bottleneck in GNN inference and training. The selective security of Plog is formally verified under the ideal-real paradigm. Extensive experimental results on three real world datasets demonstrate that compared to the state-of-the art scheme, Plog can reduce online communication rounds by 46% and achieve a 1.73× speedup in the overall inference and training time.
Jiaqi Zhao 0005, Hui Zhu 0001, Xiaoyu Kou, Haonan Yan, Fengwei Wang, Hui Li 0006
IEEE Trans. Knowl. Data Eng.4
2024 Masked image: Visually protected image dataset privacy-preserving scheme for convolutional neural networks
Xiaoyu Kou, Fengwei Wang, Hui Zhu 0001, Yandong Zheng, Zhe Liu 0001
Peer Peer Netw. Appl.1
2022 Self-Supervised Augmentation and Generation for Multi-lingual Text Advertisements at Bing
abstract
Multi-lingual text advertisement generation is a critical task for international companies, such as Microsoft. Due to the lack of training data, scaling out text advertisements generation to low-resource languages is a grand challenge in the real industry setting. Although some methods transfer knowledge from rich-resource languages to low-resource languages through a pre-trained multi-lingual language model, they fail in balancing the transferability from the source language and the smooth expression in target languages. In this paper, we propose a unified Self-Supervised Augmentation and Generation (SAG) architecture to handle the multi-lingual text advertisements generation task in a real production scenario. To alleviate the problem of data scarcity, we employ multiple data augmentation strategies to synthesize training data in target languages. Moreover, a self-supervised adaptive filtering structure is developed to alleviate the impact of the noise in the augmented data. The new state-of-the-art results on a well-known benchmark verify the effectiveness and generalizability of our proposed framework, and deployment in Microsoft Bing demonstrates the superior performance of our method.
Xiaoyu Kou, Qi Zhang 0066
KDD1
2020 TextNAS: A Neural Architecture Search Space Tailored for Text Representation
abstract
Learning text representation is crucial for text classification and other language related tasks. There are a diverse set of text representation networks in the literature, and how to find the optimal one is a non-trivial problem. Recently, the emerging Neural Architecture Search (NAS) techniques have demonstrated good potential to solve the problem. Nevertheless, most of the existing works of NAS focus on the search algorithms and pay little attention to the search space. In this paper, we argue that the search space is also an important human prior to the success of NAS in different applications. Thus, we propose a novel search space tailored for text representation. Through automatic search, the discovered network architecture outperforms state-of-the-art models on various public datasets on text classification and natural language inference tasks. Furthermore, some of the design principles found in the automatic network agree well with human intuition.
Yujing Wang 0002, Yaming Yang 0001, Jing Bai 0010, Ce Zhang 0001, Guinan Su, Xiaoyu Kou, Yunhai Tong, Mao Yang 0004, Lidong Zhou
AAAI7
2020 NASE: : Learning Knowledge Graph Embedding for Link Prediction via Neural Architecture Search
abstract
Link prediction is the task of predicting missing connections between entities in the knowledge graph (KG). While various forms of models are proposed for the link prediction task, most of them are designed based on a few known relation patterns in several well-known datasets. Due to the diversity and complexity nature of the real-world KGs, it is inherently difficult to design a model that fits all datasets well. To address this issue, previous work has tried to use Automated Machine Learning (AutoML) to search for the best model for a given dataset. However, their search space is limited only to bilinear model families. In this paper, we propose a novel Neural Architecture Search (NAS) framework for the link prediction task. First, the embeddings of the input triplet are refined by the Representation Search Module. Then, the prediction score is searched within the Score Function Search Module. This framework entails a more general search space, which enables us to take advantage of several mainstream model families, and thus it can potentially achieve better performance. We relax the search space to be continuous so that the architecture can be optimized efficiently using gradient-based search strategies. Experimental results on several benchmark datasets demonstrate the effectiveness of our method compared with several state-of-the-art approaches.
Xiaoyu Kou, Bingfeng Luo, Huang Hu, Yan Zhang 0004
CIKM1
2020 Disentangle-based Continual Graph Representation Learning
abstract
Graph embedding (GE) methods embed nodes (and/or edges) in graph into a low-dimensional semantic space, and have shown its effectiveness in modeling multi-relational data.However, existing GE models are not practical in real-world applications since it overlooked the streaming nature of incoming data.To address this issue, we study the problem of continual graph representation learning which aims to continually train a graph embedding model on new data to learn incessantly emerging multi-relational data while avoiding catastrophically forgetting old learned knowledge.Moreover, we propose a disentangle-based continual graph representation learning (DiC-GRL) framework inspired by the human's ability to learn procedural knowledge.The experimental results show that DiCGRL could effectively alleviate the catastrophic forgetting problem and outperform state-of-the-art continual learning models.* This work is done when Xiaoyu Kou was interning at Pattern Recognition Center, WeChat AI, Tencent Inc, China !"#"$% &'"(" )*$+,--, &'"(" )"-*" .//&'"(" .//,01/+"( 2#,3*4,/5 6+, 7/*5,4 85"5,3 85"5, 9: ;"<"** ="<>,# ?#"3,# @.B9'*/39/ ?*#35 ="4> 9: 78 @+*$"C9
Xiaoyu Kou, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Yan Zhang 0004
EMNLP (1)1
2019 Time-Series Anomaly Detection Service at Microsoft
abstract
Large companies need to monitor various metrics (for example, Page Views and Revenue) of their applications and services in real time. At Microsoft, we develop a time-series anomaly detection service which helps customers to monitor the time-series continuously and alert for potential incidents on time. In this paper, we introduce the pipeline and algorithm of our anomaly detection service, which is designed to be accurate, efficient and general. The pipeline consists of three major modules, including data ingestion, experimentation platform and online compute. To tackle the problem of time-series anomaly detection, we propose a novel algorithm based on Spectral Residual (SR) and Convolutional Neural Network (CNN). Our work is the first attempt to borrow the SR model from visual saliency detection domain to time-series anomaly detection. Moreover, we innovatively combine SR and CNN together to improve the performance of SR model. Our approach achieves superior experimental results compared with state-of-the-art baselines on both public datasets and Microsoft production data.
Hansheng Ren, Bixiong Xu, Yujing Wang 0002, Chao Yi, Congrui Huang, Xiaoyu Kou, Tony Xing, Mao Yang 0004, Jie Tong, Qi Zhang 0066
KDD6