Shaowen Peng

dblp:183/1435 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
8since 2021 · last 2027
0000-0003-4020-9100ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 RED: Retrieval-enhanced knowledge distillation for smaller language models
abstract
While large language models (LLMs) excel in open-domain question answering (QA) through in-context learning (ICL), their performance in domain-specific QA, where specialized knowledge is essential, remains suboptimal. Despite the technical feasibility of fine-tuning LLMs, the prohibitive computational cost during training and inference limits their practicality. Therefore, a more practical approach under limited computational resources is to enhance smaller language models for domain-specific QA through knowledge distillation (KD), leveraging both external domain knowledge and the knowledge embedded in LLMs. KD aims to improve student model performance by guiding it to mimic the prediction process of teacher models. With the rise of LLMs, student models can further benefit from chain of thought (CoT) rationales generated by LLMs. However, while LLMs contain extensive general knowledge, they often lack the specialized expertise required for domain-specific tasks that demand deep domain knowledge. To address this, we propose a novel KD scenario: R etrieval- E nhanced knowledge D istillation (RED), which expands the knowledge sources in KD by incorporating external knowledge. While external knowledge can enhance the model’s performance, retrieved text does not always perfectly align with the information needed to answer a given question due to the limitations in retrieval techniques and the content of external sources. Therefore, the key challenge in the retrieval-enhanced KD scenario is to effectively extract useful knowledge from external text while minimizing the impact of irrelevant information. In the proposed RED framework, LLMs are required not only to generate rationale for reasoning but also to evaluate the utility of retrieved text in facilitating reasoning. Additionally, we introduce compartment, filter, and traction mechanisms to mitigate the challenges associated with external text insertion. Experimental results demonstrate that our approach achieves consistent improvements for smaller models across four datasets in the scientific and biomedical domains. Particularly, under the perturbation setting where the retrieved text and the input text are mismatched, our method outperforms the competitive baseline of LLM inference by 9.14%, demonstrating its robustness.
Xinbai Li, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki
Expert Syst. Appl.2
2026 Estimating Shared Mental Models via Communication-Categorized Directed Graphs
abstract
Corporate organizations face increasingly complex tasks that demand effective team management. A key concept is the Shared Mental Model (SMM), which enables members to maintain performance despite limited communication. Traditional measurements rely on interviews or questionnaires, which are labor-intensive, context-specific, and unsuitable for continuous monitoring. Consequently, leaders lack practical tools to track shared cognition in real time. This paper’s empirical analysis shows that only specific categories (e.g., informative exchanges) correlate strongly with SMM, clarifying which forms of communication can influence shared cognition. This insight leads to our proposed approach, which estimates SMM from instant messaging systems like Slack. Our approach categorizes messages into communicative acts using large language models, constructs category-wise communication graphs, and applies a graph neural network for estimation. The model outperforms baselines, demonstrating the feasibility of continuous, scalable monitoring without intrusive surveys. While validated in corporate contexts, the approach extends to education, healthcare, and disaster response domains.
Wataru Yamada, Keiichi Ochiai, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki
CHI4
2025 GenKP: generative knowledge prompts for enhancing large language models
abstract
Large language models (LLMs) have demonstrated extensive capabilities across various natural language processing (NLP) tasks. Knowledge graphs (KGs) harbor vast amounts of facts, furnishing external knowledge for language models. The structured knowledge extracted from KGs must undergo conversion into sentences to align with the input format required by LLMs. Previous research has commonly utilized methods such as triple conversion and template-based conversion. However, sentences converted using existing methods frequently encounter issues such as semantic incoherence, ambiguity, and unnaturalness, which distort the original intent, and deviate the sentences from the facts. Meanwhile, despite the improvement that knowledge-enhanced pre-training and prompt-tuning methods have achieved in small-scale models, they are difficult to implement for LLMs in the absence of computational resources. The advanced comprehension of LLMs facilitates in-context learning (ICL), thereby enhancing their performance without the need for additional training. In this paper, we propose a knowledge prompts generation method, GenKP, which injects knowledge into LLMs by ICL. Compared to inserting triple-conversion or templated-conversion knowledge without selection, GenKP entails generating knowledge samples using LLMs in conjunction with KGs and makes a trade-off of knowledge samples through weighted verification and BM25 ranking, reducing knowledge noise. Experimental results illustrate that incorporating knowledge prompts enhances the performance of LLMs. Furthermore, LLMs augmented with GenKP exhibit superior improvements compared to the methods utilizing triple and template-based knowledge injection.
Xinbai Li, Shaowen Peng, Shuntaro Yada, Shoko Wakamiya, Eiji Aramaki
Appl. Intell.2
2025 Balancing Embedding Spectrum for Recommendation
abstract
Modern recommender systems heavily rely on high-quality representations learned from high-dimensional sparse data. While significant efforts have been invested in designing powerful algorithms for extracting user preferences, the factors contributing to good representations have remained relatively unexplored. In this work, we shed light on an issue in the existing pairwise learning paradigm (i.e., embedding collapse), that the representations tend to span a subspace of the whole embedding space, leading to a suboptimal solution and reducing the model capacity. Specifically, we show that alignment of positive pairs is equivalent to a low-pass filter causing users and items to collapse to a constant vector. While negative sampling can partially mitigate this issue by acting as a high-pass filter to balance the spectrum, leading to an incomplete collapse. To tackle this issue, we present a novel learning paradigm DirectSpec, which directly optimizes the spectrum distribution to ensure that users and items effectively span the entire embedding space. We demonstrate that many self-supervised learning algorithms without explicit negative sampling can be considered as special cases of DirectSpec. Furthermore, we show that optimizing the spectrum inappropriately could also be detrimental to data representation, where the key lies in a dynamic balance between alignment of positive pairs and spectrum balancing. Finally, we propose an enhanced and practical implementation DirectSpec + to balance the embedding spectrum more adaptively and effectively. We implement DirectSpec + on two popular recommender models: matrix factorization and LightGCN. Our experimental results demonstrate its effectiveness and efficiency over competitive baselines.
Shaowen Peng, Kazunari Sugiyama, Xin Liu 0020, Tsunenori Mine
Trans. Recomm. Syst.1
2024 How Powerful is Graph Filtering for Recommendation
abstract
It has been shown that the effectiveness of graph convolutional network (GCN) for recommendation is attributed to the spectral graph filtering. Most GCN-based methods consist of a graph filter or followed by a low-rank mapping optimized based on supervised training. However, we show two limitations suppressing the power of graph filtering: (1) Lack of generality. Due to the varied noise distribution, graph filters fail to denoise sparse data where noise is scattered across all frequencies, while supervised training results in worse performance on dense data where noise is concentrated in middle frequencies that can be removed by graph filters without training. (2) Lack of expressive power. We theoretically show that linear GCN (LGCN) that is effective on collaborative filtering (CF) cannot generate arbitrary embeddings, implying the possibility that optimal data representation might be unreachable.
Shaowen Peng, Xin Liu 0020, Kazunari Sugiyama, Tsunenori Mine
KDD1
2024 Less is More: Removing Redundancy of Graph Convolutional Networks for Recommendation
abstract
While Graph Convolutional Networks (GCNs) have shown great potential in recommender systems and collaborative filtering (CF), they suffer from expensive computational complexity and poor scalability. On top of that, recent works mostly combine GCNs with other advanced algorithms which further sacrifice model efficiency and scalability. In this work, we unveil the redundancy of existing GCN-based methods in three aspects: (1) Feature redundancy . By reviewing GCNs from a spectral perspective, we show that most spectral graph features are noisy for recommendation, while stacking graph convolution layers can suppress but cannot completely remove the noisy features, which we mostly summarize from our previous work; (2) Structure redundancy . By providing a deep insight into how user/item representations are generated, we show that what makes them distinctive lies in the spectral graph features, while the core idea of GCNs (i.e., neighborhood aggregation) is not the reason making GCNs effective; and (3) Distribution redundancy . Following observations from (1), we further show that the number of required spectral features is closely related to the spectral distribution, where important information tends to be concentrated in more (fewer) spectral features on a flatter (sharper) distribution. To make important information be concentrated in as few features as possible, we sharpen the spectral distribution by increasing the node similarity without changing the original data, thereby reducing the computational cost. To remove these three kinds of redundancies, we propose a Simplified Graph Denoising Encoder (SGDE) only exploiting the top- K singular vectors without explicitly aggregating neighborhood, which significantly reduces the complexity of GCN-based methods. We further propose a scalable contrastive learning framework to alleviate data sparsity and to boost model robustness and generalization, leading to significant improvement. Extensive experiments on three real-world datasets show that our proposed SGDE not only achieves state-of-the-art but also shows higher scalability and efficiency than our previously proposed GDE as well as traditional and GCN-based CF methods.
Shaowen Peng, Kazunari Sugiyama, Tsunenori Mine
ACM Trans. Inf. Syst.1
2022 SVD-GCN: A Simplified Graph Convolution Paradigm for Recommendation
abstract
With the tremendous success of Graph Convolutional Networks (GCNs), they have been widely applied to recommender systems and have shown promising performance. However, most GCN-based methods rigorously stick to a common GCN learning paradigm and suffer from two limitations: (1) the limited scalability due to the high computational cost and slow training convergence; (2) the notorious over-smoothing issue which reduces performance as stacking graph convolution layers. We argue that the above limitations are due to the lack of a deep understanding of GCN-based methods. To this end, we first investigate what design makes GCN effective for recommendation. By simplifying LightGCN, we show the close connection between GCN-based and low-rank methods such as Singular Value Decomposition (SVD) and Matrix Factorization (MF), where stacking graph convolution layers is to learn a low-rank representation by emphasizing (suppressing) components with larger (smaller) singular values. Based on this observation, we replace the core design of GCN-based methods with a flexible truncated SVD and propose a simplified GCN learning paradigm dubbed SVD-GCN, which only exploits K-largest singular vectors for recommendation. To alleviate the over-smoothing issue, we propose a renormalization trick to adjust the singular value gap, resulting in significant improvement. Extensive experiments on three real-world datasets show that our proposed SVD-GCN not only significantly outperforms state-of-the-arts but also achieves over 100x and 10x speedups over LightGCN and MF, respectively.
Shaowen Peng, Kazunari Sugiyama, Tsunenori Mine
CIKM1
2022 Less is More: Reweighting Important Spectral Graph Features for Recommendation
abstract
As much as Graph Convolutional Networks (GCNs) have shown tremendous success in recommender systems and collaborative filtering (CF), the mechanism of how they, especially the core components (\textiti.e., neighborhood aggregation) contribute to recommendation has not been well studied. To unveil the effectiveness of GCNs for recommendation, we first analyze them in a spectral perspective and discover two important findings: (1) only a small portion of spectral graph features that emphasize the neighborhood smoothness and difference contribute to the recommendation accuracy, whereas most graph information can be considered as noise that even reduces the performance, and (2) repetition of the neighborhood aggregation emphasizes smoothed features and filters out noise information in an ineffective way. Based on the two findings above, we propose a new GCN learning scheme for recommendation by replacing neihgborhood aggregation with a simple yet effective Graph Denoising Encoder (GDE), which acts as a band pass filter to capture important graph features. We show that our proposed method alleviates the over-smoothing and is comparable to an indefinite-layer GCN that can take any-hop neighborhood into consideration. Finally, we dynamically adjust the gradients over the negative samples to expedite model training without introducing additional complexity. Extensive experiments on five real-world datasets show that our proposed method not only outperforms state-of-the-arts but also achieves 12x speedup over LightGCN.
Shaowen Peng, Kazunari Sugiyama, Tsunenori Mine
SIGIR1
2018 Vector Representation Based Model Considering Randomness of User Mobility for Predicting Potential Users
Shaowen Peng, Xianzhong Xie, Tsunenori Mine, Chang Su 0003
PRIMA1