Siyue Wu

dblp:293/8140 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 GREAT: Generalized Reservoir Sampling based Triangle Counting Estimation over Streaming Graphs
abstract
The number of triangles of a streaming graph is a crucial metric with various applications, such as network evolution analysis, community detection, and anomaly detection. A practical solution for triangle counting in streaming graphs is the sampling-based approximation. Although a lot of research efforts have been devoted to the fixed-sized memory based algorithms, they suffer from the accuracy and the efficiency issues. To tackle these issues, we first propose the generalized reservoir sampling (GRS), which stores less edges for reducing the computational cost and can still generate uniformly random edge sample in the streaming graph. Then, we propose the GREAT algorithm based on GRS for efficient and accurate triangle counting estimation. To further improve the estimation accuracy, we propose the GREAT + algorithm for considering the dynamic timestamp interval distribution in real-world streaming graphs so that triangles with short and long timestamp intervals will be sampled following the ground-truth distribution. Extensive evaluations on real datasets demonstrate the efficiency and the accuracy of our algorithms. The relative error of our algorithm GREAT + is significantly (an order of magnitude) better than the competitors.
Siyue Wu, Dingming Wu 0001, Sinhong Cheuk, Tsz Nam Chan, Kezhong Lu
Proc. VLDB Endow.1
2024 CDFRS: A scalable sampling approach for efficient big data analysis
abstract
The sampling-based approximation method has demonstrated its potential in various domains such as machine learning, query processing, and data analysis. Most preceding sampling algorithms generate samples at the record level, making it impractical to apply them to very large datasets using a single machine. Even distributed solutions encounter efficiency issues when dealing with terabyte-scale datasets. In this paper, we introduce a scalable sampling approach named CDFRS, which can generate samples with a distribution-preserving guarantee from extensive datasets. CDFRS exhibits significantly improved speed compared to existing sampling algorithms when dealing with terabyte-scale datasets. We provide theoretical guarantees and empirical justifications, demonstrating that samples generated by the CDFRS approach maintain the distribution characteristics of the original dataset. Additionally, we propose a sample size determination algorithm, denoted as A2. Experiment results indicate that the running time of CDFRS shows at least an order of magnitude improvement over other distributed sampling methods. Notably, sampling a 10TB dataset using CDFRS only takes hundreds of seconds, while the compared method requires more than ten thousand seconds. In the context of big data analysis, including tasks such as classification and clustering, models trained with samples generated by CDFRS closely match those trained with the entire training set. Furthermore, the proposed A2 algorithm efficiently determines an appropriate sample size compared with traditional methods.
Yongda Cai, Dingming Wu 0001, Xudong Sun 0004, Siyue Wu, Jingsheng Xu, Joshua Zhexue Huang
Inf. Process. Manag.4
2024 Efficient and Accurate PageRank Approximation on Large Graphs
abstract
PageRank is a commonly used measurement in a wide range of applications, including search engines, recommendation systems, and social networks. However, this measurement suffers from huge computational overhead, which cannot be scaled to large graphs. Although many approximate algorithms have been proposed for computing PageRank values, these algorithms are either (i) not efficient or (ii) not accurate. Worse still, some of them cannot provide estimated PageRank values for all the vertices. In this paper, we first propose the CUR-Trans algorithm, which can reduce the time complexity for computing PageRank values and has lower error bound than existing matrix approximation-based PageRank algorithms. Then, we develop the T 2 -Approx algorithm to further reduce the time complexity for computing this measurement. Experiment results on three large-scale graphs show that both the CUR-Trans algorithm and the T 2 -Approx algorithm achieve the lowest response time for computing PageRank values with the best accuracy (for the CUR-Trans algorithm) or the competitive accuracy (for the T 2 -Approx algorithm). Besides, the two proposed algorithms are able to provide estimated PageRank values for all the vertices.
Siyue Wu, Dingming Wu 0001, Junyi Quan, Tsz Nam Chan, Kezhong Lu
Proc. ACM Manag. Data1
2024 Multi-Party Conversation Modeling for Emotion Recognition
abstract
Multi-party conversation modeling plays a vital role in emotion recognition in conversation (ERC). Aside from the intra- and inter-speaker dependencies between different speakers, the difficulty also lies in the fact that each conversation may contain several to many utterances that compose a long text sequence. In this article, we present two approaches to effective multi-party conversation modeling. First, to encode long sequences and capture long-range dependency between utterances, we introduce a dialog-oriented language model, DialogXL, with enhanced memory to store longer conversation sequences and dialog-aware self-attention to deal with multi-party dependencies. Second, we present a directed acyclic neural network, namely DAG-ERC, to encode the utterances with a directed acyclic graph (DAG) to better capture the intrinsic structure within a conversation. DAG-ERC combines the advantages of recurrent models and graph models and provides a more intuitive way to model information flow between sequential utterances. Extensive experiments are conducted on four ERC benchmarks with state-of-the-art models employed for comparison, and empirical results demonstrate the superiority of the two models in multi-party conversation modeling.
Xiaojun Quan, Siyue Wu, Weizhou Shen, Jianxing Yu
IEEE Trans. Affect. Comput.2
2023 AD-KD: Attribution-Driven Knowledge Distillation for Language Model Compression
abstract
Knowledge distillation has attracted a great deal of interest recently to compress pre-trained language models.However, existing knowledge distillation methods suffer from two limitations.First, the student model simply imitates the teacher's behavior while ignoring the underlying reasoning.Second, these methods usually focus on the transfer of sophisticated model-specific knowledge but overlook dataspecific knowledge.In this paper, we present a novel attribution-driven knowledge distillation approach, which explores the token-level rationale behind the teacher model based on Integrated Gradients (IG) and transfers attribution knowledge to the student model.To enhance the knowledge transfer of model reasoning and generalization, we further explore multi-view attribution distillation on all potential decisions of the teacher.Comprehensive experiments are conducted with BERT on the GLUE benchmark.The experimental results demonstrate the superior performance of our approach to several state-of-the-art methods.
Siyue Wu, Hongzhan Chen, Xiaojun Quan, Qifan Wang 0001, Rui Wang 0005
ACL (1)1
2021 Directed Acyclic Graph Network for Conversational Emotion Recognition
abstract
Weizhou Shen, Siyue Wu, Yunyi Yang, Xiaojun Quan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Weizhou Shen, Siyue Wu, Yunyi Yang, Xiaojun Quan
ACL/IJCNLP (1)2