Zhizhuo Yin

dblp:275/7037 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-4700-7774ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 IMGNN: An Efficient, Effective and Generalizable Algorithm for Influence Maximization in Social Networks
abstract
Influence Maximization (IM) is a crucial problem in social network analysis and has been extensively studied. Traditional approaches rely on designing approximation algorithms using network sampling; however, these methods lack generalizability and depend on an explicit definition of the influence diffusion model as input. Recently, researchers have turned to deep learning methods to address the shortcomings of traditional IM algorithms, but current learning-based IM algorithms still suffer from severe deficiencies in scalability and generalizability. In this paper, we propose IMGNN, a simple, efficient, effective, and generalizable algorithm powered by graph neural networks. IMGNN is a learning-based IM algorithm with strong generalization capability that reduces the overhead of model retraining while also providing fast seed set inference speed. As a result, IMGNN achieves better performance in terms of both efficiency and effectiveness compared to existing IM algorithms. Its exceptional generalization capability also enables it to be trained on small-scale graphs and directly infer the seed node set for large-scale graphs. IMGNN achieves these advantages by adopting a novel design for feature construction and model training, utilizing features constructed from influence propagations over graphs with randomly skipped nodes. This approach enables IMGNN to avoid overfitting to specific network structures while employing a unique technique to improve time efficiency by training on smaller networks. We have conducted extensive experiments using real-world social networks with up to 40 million nodes, and the results strongly demonstrate the superiority of IMGNN in terms of influence spread, seed node set inference speed, and generalizability.
Haotian Zhang 0027, Kai Han 0003, Zhizhuo Yin, Jing Tang 0004, Pan Hui 0001
KDD (1)3
2025 PyraMotion: Attentional Pyramid-Structured Motion Integration for Co-Speech 3D Gesture Synthesis
abstract
Generating full-body human gestures encompassing face, body, hands, and global movements from audio is crucial yet challenging for virtual avatar creation. Existing systems tokenize gestures frame-wise, predicting tokens of each frame from the input audio. However, expressive human gestures consist of varied patterns with different frame lengths, and different body parts exhibit motion patterns of varying durations. Existing systems fail to capture motion patterns across body parts and temporal scales due to the fixed frame-count setting of their gesture tokens. Inspired by the success of the feature pyramid technique in the multi-scale visual information extraction, we propose a novel framework named PyraMotion and an adaptive multi-scale feature capturing model called Attentive Pyramidal VQ-VAE (APVQ-VAE). Objective and subjective experiments demonstrate that the PyraMotion outperforms state-of-the-art methods in terms of generating natural and expressive full-body human gestures. Extensive ablation experiments highlight that the self-adaptiveness integration through attention maps contributes to performance.
Zhizhuo Yin, Yuk Hang Tsui, Pan Hui 0001
NeurIPS1
2024 Text2VRScene: Exploring the Framework of Automated Text-driven Generation System for VR Experience
abstract
With the recent development of the Virtual Reality (VR) industry, the increasing number of VR users pushes the demand for the massive production of immersive and expressive VR scenes in related industries. However, creating expressive VR scenes involves the reasonable organization of various digital content to express a coherent and logical theme, which is time-consuming and labor-intensive. In recent years, Large Language Models (LLMs) such as ChatGPT 3.5 and generative models such as stable diffusion have emerged as powerful tools for comprehending natural language and generating digital contents such as text, code, images, and 3D objects. In this paper, we have explored how we can generate VR scenes from text by incorporating LLMs and various generative models into an automated system. To achieve this, we first identify the possible limitations of LLMs for an automated system and propose a systematic framework to mitigate them. Subsequently, we developed Text2VRScene, a VR scene generation system, based on our proposed framework with well-designed prompts. To validate the effectiveness of our proposed framework and the designed prompts, we carry out a series of test cases. The results show that the proposed framework contributes to improving the reliability of the system and the quality of the generated VR scenes. The results also illustrate the promising performance of the Text2VRScene in generating satisfying VR scenes with a clear theme regularized by our well-designed prompts. This paper ends with a discussion about the limitations of the current system and the potential of developing similar generation systems based on our framework.
Zhizhuo Yin, Yuyang Wang 0002, Theodoros Papatheodorou, Pan Hui 0001
VR1
2024 APT-Pipe: A Prompt-Tuning Tool for Social Data Annotation using ChatGPT
abstract
Recent research has highlighted the potential of LLMs, like ChatGPT, for performing label annotation on social computing data. However, it is already well known that performance hinges on the quality of the input prompts. To address this, there has been a flurry of research into prompt tuning --- techniques and guidelines that attempt to improve the quality of prompts. Yet these largely rely on manual effort and prior knowledge of the dataset being annotated. To address this limitation, we propose APT-Pipe, an automated prompt-tuning pipeline. APT-Pipe aims to automatically tune prompts to enhance ChatGPT's text classification performance on any given dataset. We implement APT-Pipe and test it across twelve distinct text classification datasets. We find that prompts tuned by APT-Pipe help ChatGPT achieve higher weighted F1-score on nine out of twelve experimented datasets, with an improvement of 7.01% on average. We further highlight APT-Pipe's flexibility as a framework by showing how it can be extended to support additional tuning mechanisms.
Zhizhuo Yin, Gareth Tyson, Ehsan ul Haq, Lik-Hang Lee, Pan Hui 0001
WWW2
2024 Spatial and Temporal User Interest Representations for Sequential Recommendation
abstract
In recent years, recommendation systems have become increasingly prevalent in various fields, facilitating quick access to the information users need. As a result, many models have been proposed to model user interests, leading to more accurate recommendation lists, superior user experience, and business value. However, characterizing the dynamically changing interests of users is a challenging task. User interests shift over time while maintaining some long-term interests, and at each time, users’ interests are diverse. To investigate the benefits of multidimensional interests for users, this article proposes to characterize user preferences based on their spatiotemporal interests. Utilizing temporal and spatial information is critical for improving recommendation accuracy. To achieve this, we present a novel approach called multilong short-term interest (MLSI) user representation for recommendation. This method extracts long-term and short-term interests of users from their behavioural sequences using decoupled self-supervised learning with different optimizers. Self-attention is then employed to capture the diverse interests of users through their behavioral sequences. Final, long-term and short-term interests, as well as diversified interests, are aggregated to represent user interests. Extensive experiments on real-world datasets show that MLSI not only outperforms state-of-the-art methods but also more effectively characterizes user interests, reflecting an improvement ranging from 5% to 20% across various metrics on multiple datasets.
Haibing Hu, Kai Han 0003, Zhizhuo Yin, Defu Lian
IEEE Trans. Comput. Soc. Syst.3
2024 H3GNN: Hybrid Hierarchical HyperGraph Neural Network for Personalized Session-based Recommendation
abstract
Personalized Session-based recommendation (PSBR) is a general and challenging task in the real world, aiming to recommend a session’s next clicked item based on the session’s item transition information and the corresponding user’s historical sessions. A session is defined as a sequence of interacted items during a short period. The PSBR problem has a natural hierarchical architecture in which each session consists of a series of items, and each user owns a series of sessions. However, the existing PSBR methods can merely capture the pairwise relation information within items and users. To effectively capture the hierarchical information, we propose a novel hierarchical hypergraph neural network to model the hierarchical architecture. Moreover, considering that the items in sessions are sequentially ordered, while the hypergraph can only model the set relation, we propose a directed graph aggregator (DGA) to aggregate the sequential information from the directed global item graph. By attentively combining the embeddings of the above two modules, we propose a framework dubbed H3GNN (Hybrid Hierarchical HyperGraph Neural Network). Extensive experiments on three benchmark datasets demonstrate the superiority of our proposed model compared to the state-of-the-art methods, and ablation experiment results validate the effectiveness of all the proposed components.
Zhizhuo Yin, Kai Han 0003, Pengzi Wang, Xi Zhu 0004
ACM Trans. Inf. Syst.1
2023 Multi Global Information Assisted Streaming Session-Based Recommendation System
abstract
Streaming Session-Based Recommendation (SSBR) is a challenging problem as user preferences in sessions are continually drifting with sessions generated chronologically. In recent years, some SSBR models have been proposed to address this problem by reservoir technique and Graph Neural Networks (GNN) which help to preserve a representative sketch of the historical data and extract item transition information in sessions. However, there are two critical problems in existing methods: (1) most existing methods only focus on the local session information without exploiting the information of other sessions and users; (2) GNN models in existing SSBR methods are unable to capture the importance of different user features. To address the problems mentioned above, we propose a novel architecture namedGlobalItem andUser embeddingAssistedGraphNeuralNetwork (GIUA-GNN) for combining the global user and item information in an attentional manner with local session information for the recommendation. We also propose a novel architecture of graph neural network which utilizes the attention mechanism for better extracting the importance of different features of user embeddings namedBi-directedAttentionalGraphConvolutionalNetwork (BA-GCN). Extensive experiments on three different sizes of real-world datasets have been conducted to demonstrate the superiority of our model on metrics MRR and Recall.
Zhizhuo Yin, Kai Han 0003, Pengzi Wang, Haibing Hu
IEEE Trans. Knowl. Data Eng.1
2022 Bilinear Multi-Head Attention Graph Neural Network for Traffic Prediction
Haibing Hu, Kai Han 0003, Zhizhuo Yin
ICAART (2)3
2020 Character-Oriented Video Summarization With Visual and Textual Cues
abstract
With the booming of content “re-creation” in social media platforms, character-orientedvideo summary has become a crucial form of user-generated video content. However, artificial extraction could be time-consuming with high missing rate, while traditional techniques on person search may incur heavy burden of computing resources. At the same time, in social media platforms, videos are usually accompanied with rich textual information, e.g., subtitles or bullet-screen comments which provide the multi-view description of videos. Thus, there exists a potential to leverage textual information to enhance the character-oriented video summarization. To that end, in this paper, we propose a novel framework for jointly modeling visual and textual information. Specifically, we first locate characters indiscriminately through detection methods, and then identify these characters via re-identification to extract potential key-frames, in which appropriate source of textual information will be automatically selected and integrated based on the features of specific frame. Finally, key-frames will be aggregated as the character-oriented summarization. Experiments on real-world data sets validate that our solution outperforms several state-of-the-art baselines on both person search and summarization tasks, which prove the effectiveness of our solution on the character-oriented video summarization problem.
Peilun Zhou, Tong Xu 0001, Zhizhuo Yin, Dong Liu 0002, Enhong Chen, Guangyi Lv, Changliang Li
IEEE Trans. Multim.3