Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hongyun Cai 0001

dblp:247/9265-1 · DBLP profile ↗
← Back
19ranked-venue papers
10as first author
4since 2021 · last 2026
0000-0001-5853-2841ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Web and social media mining · 38% Information retrieval · 16% Graph data management · 13%
Artificial intelligence
2 papers
Graph learning · 73% Representation and self-supervised learning · 27%
Computer graphics and multimedia
3 papers
Multimedia analysis and retrieval · 60% Visualization and visual analytics · 40%

Topics — the 27 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Web and social media mining
social network analysis
0.922023
Friend Ranking in Online Games via Pre-training Edge Transformers · SIGIR 2023
SocialLens: Searching and Browsing Communities by Content and Interaction · ICDE 2017
Web and social media mining › event detection
social event detection
0.742016
Indexing Evolving Events from Tweet Streams · IEEE Trans. Knowl. Data Eng. 2015
What are Popular: Exploring Twitter Features for Event Detection, Tracking and Visualization · ACM Multimedia 2015
EventEye: Monitoring Evolving Events from Tweet Streams · ACM Multimedia 2014
Recommender systems › social recommendation
friend recommendation
0.712023
Friend Ranking in Online Games via Pre-training Edge Transformers · SIGIR 2023
Knowledge graphs
link prediction
0.712023
Friend Ranking in Online Games via Pre-training Edge Transformers · SIGIR 2023
Web and social media mining › online community analysis
community profiling
0.622017
From Community Detection to Community Profiling · Proc. VLDB Endow. 2017
SocialLens: Searching and Browsing Communities by Content and Interaction · ICDE 2017
Information retrieval › indexing
inverted index
0.522016
Indexing evolving events from tweet streams · ICDE 2016
Indexing Evolving Events from Tweet Streams · IEEE Trans. Knowl. Data Eng. 2015
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
embedding initialization
0.412020
Initialization for Network Embedding: A Graph Partition Approach · WSDM 2020
Machine learning › Graph learning
network embedding
0.412020
Initialization for Network Embedding: A Graph Partition Approach · WSDM 2020
Machine learning › Graph learning › graph neural network
node classification
0.412020
Initialization for Network Embedding: A Graph Partition Approach · WSDM 2020
Graph data management
graph analytics
0.312018
A Comprehensive Survey of Graph Embedding: Problems, Techniques, and Applications · IEEE Trans. Knowl. Data Eng. 2018
Graph data management
graph embedding
0.312018
A Comprehensive Survey of Graph Embedding: Problems, Techniques, and Applications · IEEE Trans. Knowl. Data Eng. 2018
Data mining › structured data mining › graph mining
community detection
0.312017
From Community Detection to Community Profiling · Proc. VLDB Endow. 2017
Graph data management
community search
0.312017
SocialLens: Searching and Browsing Communities by Content and Interaction · ICDE 2017
Information retrieval › interactive information retrieval
search and browsing
0.312017
SocialLens: Searching and Browsing Communities by Content and Interaction · ICDE 2017
Web and social media mining
event detection
0.322015
Indexing Evolving Events from Tweet Streams · IEEE Trans. Knowl. Data Eng. 2015
What are Popular: Exploring Twitter Features for Event Detection, Tracking and Visualization · ACM Multimedia 2015
Data stream processing
complex event processing
0.212016
Indexing evolving events from tweet streams · ICDE 2016
Information retrieval
indexing
0.212015
Indexing Evolving Events from Tweet Streams · IEEE Trans. Knowl. Data Eng. 2015
Multimedia analysis and retrieval › multimodal learning
multimodal topic modeling
0.212015
What are Popular: Exploring Twitter Features for Event Detection, Tracking and Visualization · ACM Multimedia 2015
Machine learning › Graph learning
graph neural network
0.212023
Friend Ranking in Online Games via Pre-training Edge Transformers · SIGIR 2023
Data mining › clustering › online clustering
data stream clustering
0.212014
EventEye: Monitoring Evolving Events from Tweet Streams · ACM Multimedia 2014
Web and social media mining › event analysis
event evolution tracking
0.212014
EventEye: Monitoring Evolving Events from Tweet Streams · ACM Multimedia 2014
Information retrieval › document retrieval › temporal information retrieval
event retrieval
0.212014
EventEye: Monitoring Evolving Events from Tweet Streams · ACM Multimedia 2014
Machine learning › Graph learning
link prediction
0.112020
Initialization for Network Embedding: A Graph Partition Approach · WSDM 2020
Data mining
network analysis
0.112017
From Community Detection to Community Profiling · Proc. VLDB Endow. 2017
Visualization and visual analytics › graph visualization
interactive graph visualization
0.112017
SocialLens: Searching and Browsing Communities by Content and Interaction · ICDE 2017
Data mining › text mining
topic model
0.112015
What are Popular: Exploring Twitter Features for Event Detection, Tracking and Visualization · ACM Multimedia 2015
Visualization and visual analytics › spatiotemporal visualization
event visualization
0.112014
EventEye: Monitoring Evolving Events from Tweet Streams · ACM Multimedia 2014

Methods — techniques the papers use, named apart from their topics

pre-training · 1.3masked autoencoder · 1.3profile-aware search · 0.6community ranking · 0.6upper bound pruning · 0.5convolutional neural network features · 0.4graph partition · 0.4embedding propagation · 0.4taxonomy · 0.3scalable inference · 0.3joint profiling and detection model · 0.3nearest neighbour search · 0.2topic model · 0.2stream clustering · 0.2multi-layer indexing · 0.2
YearPublicationVenuePosition
2026 LPS-GNN: Deploying Graph Neural Networks on Graphs with 100-Billion Edges
abstract
Graph Neural Networks (GNNs) have emerged as powerful tools for various graph mining tasks, yet existing scalable solutions often struggle to balance execution efficiency with prediction accuracy. These difficulties stem from iterative message-passing techniques, which place significant computational demands and require extensive GPU memory, particularly when dealing with the neighbor explosion issue inherent in large-scale graphs. This paper introduces a scalable, low-cost, flexible, and efficient GNN framework called LPS-GNN, which can perform representation learning on 100 billion graphs with a single GPU in 10 hours and shows a 13.8% improvement in User Acquisition scenarios. We examine existing graph partitioning methods and design a superior graph partition algorithm named LPMetis. In particular, LPMetis outperforms current state-of-the-art (SOTA) approaches on various evaluation metrics. In addition, our paper proposes a subgraph augmentation strategy to enhance the model's predictive performance. It exhibits excellent compatibility, allowing the entire framework to accommodate various GNN algorithms. Successfully deployed on the Tencent platform, LPS-GNN has been tested on public and real-world datasets, achieving performance lifts of 8. 24% to 13. 89% over SOTA models in online applications.
Yukuo Cen, Wenzheng Feng, Hongyun Cai 0001, Jie Tang 0001
ACM Trans. Knowl. Discov. Data8
2024 A Robust Sequential Recommendation Model Based on Multiple Feedback Behavior Denoising and Trusted Neighbors
abstract
Abstract At present, most of the personalized sequential recommendations utilize users’ implicit positive feedback (such as clicks) to predict user behavior, ignoring the impact of implicit negative feedback and explicit feedback on the accuracy of recommendation results prediction. In this paper, we propose a robust sequence recommendation model based on multi feedback behavior denoising and trusted neighbors, which utilizes multiple feedback behavior data for feature denoising and considers trusted nearest neighbor information to improve model performance. Firstly, by learning the feature representations and interactions of various types of feedback, explicit feedback is used to map and purify implicit feedback with the same and different attributes, resulting in unbiased user performance. Then, we design a filter attention network to identify highly trusted neighbor information. Finally, we integrate pure user interest representations and trusted nearest neighbor representations to improve the accuracy and robustness of the model. The experimental results on two publicly available datasets show that the proposed sequential recommendation model can achieve superior results to baseline methods in both AUC and RelaImpr.
Hongyun Cai 0001, Shilin Yuan, Jichao Ren
Neural Process. Lett.1
2023 Friend Ranking in Online Games via Pre-training Edge Transformers
abstract
Friend recall is an important way to improve Daily Active Users (DAU) in online games. The problem is to generate a proper inactive (lost) friend ranking list essentially. Traditional friend recall methods focus on rules like friend intimacy or training a classifier for predicting lost players' return probability, but ignore feature information of (active) players and historical friend recall events. In this work, we treat friend recall as a link prediction problem and explore several link prediction methods which can use features of both active and lost players, as well as historical events. Furthermore, we propose a novel Edge Transformer model and pre-train the model via masked auto-encoders. Our method achieves state-of-the-art results in the offline experiments and online A/B Tests of three Tencent games.
Jiazhen Peng, Shenggong Ji, Qiang Liu 0005, Hongyun Cai 0001
SIGIR5
2022 Learning high-order structural and attribute information by knowledge graph attention networks for enhancing knowledge graph embedding
Hongyun Cai 0001, Sifa Xie, Yipeng Yu, dukehyzhang
Knowl. Based Syst.2
2020 Initialization for Network Embedding: A Graph Partition Approach
abstract
Network embedding has been intensively studied in the literature and widely used in various applications, such as link prediction and node classification. While previous work focus on the design of new algorithms or are tailored for various problem settings, the discussion of initialization strategies in the learning process is often missed. In this work, we address this important issue of initialization for network embedding that could dramatically improve the performance of the algorithms on both effectiveness and efficiency. Specifically, we first exploit the graph partition technique that divides the graph into several disjoint subsets, and then construct an abstract graph based on the partitions. We obtain the initialization of the embedding for each node in the graph by computing the network embedding on the abstract graph, which is much smaller than the input graph, and then propagating the embedding among the nodes in the input graph. With extensive experiments on various datasets, we demonstrate that our initialization technique significantly improves the performance of the state-of-the-art algorithms on the evaluations of link prediction and node classification by up to 7.76% and 8.74% respectively. Besides, we show that the technique of initialization reduces the running time of the state-of-the-arts by at least 20%.
Wenqing Lin, Faqiang Zhang, Hongyun Cai 0001
WSDM5
2019 An Attention-Based Model for Learning Dynamic Interaction Networks
abstract
In the physical world, complex systems are generally created as the composition of multiple primitive components that interact with each other rather than a single monolithic structure. Recently, spatio-temporal graphs received a reasonable amount of attention from the research community since they emerged as a natural representational tool able to capture the interactive and interrelated structure of a complex problem. To better understand the nature of complex systems, there is the need to define models that can easily explain the learned causal relationship. To this end, we propose an attentive model able to learn and project the relational structure into a fixed-size embedding. Such representation naturally captures the dynamic influence that each neighbors exert over a given vertex providing a valuable description of the problem setting. The proposed architecture has been extensively evaluated against strong baselines on toy as well as real-world tasks, such as prediction of household energy load and traffic congestion.
Sandro Cavallari, Soujanya Poria, Erik Cambria, Vincent Wenchen Zheng, Hongyun Cai 0001
IJCNN5
2018 A Comprehensive Survey of Graph Embedding: Problems, Techniques, and Applications
abstract
Graph is an important data representation which appears in a wide diversity of real-world scenarios. Effective graph analytics provides users a deeper understanding of what is behind the data, and thus can benefit a lot of useful applications such as node classification, node recommendation, link prediction, etc. However, most graph analytics methods suffer the high computation and space cost. Graph embedding is an effective yet efficient way to solve the graph analytics problem. It converts the graph data into a low dimensional space in which the graph structural information and graph properties are maximumly preserved. In this survey, we conduct a comprehensive review of the literature in graph embedding. We first introduce the formal definition of graph embedding as well as the related concepts. After that, we propose two taxonomies of graph embedding which correspond to what challenges exist in different graph embedding problem settings and how the existing work addresses these challenges in their solutions. Finally, we summarize the applications that graph embedding enables and suggest four promising future research directions in terms of computation efficiency, problem settings, techniques, and application scenarios.
Hongyun Cai 0001, Vincent Wenchen Zheng, Kevin Chen-Chuan Chang
IEEE Trans. Knowl. Data Eng.1
2017 Learning Community Embedding with Community Detection and Node Embedding on Graphs
abstract
In this paper, we study an important yet largely under-explored setting of graph embedding, i.e., embedding communities instead of each individual nodes. We find that community embedding is not only useful for community-level applications such as graph visualization, but also beneficial to both community detection and node classification. To learn such embedding, our insight hinges upon a closed loop among community embedding, community detection and node embedding. On the one hand, node embedding can help improve community detection, which outputs good communities for fitting better community embedding. On the other hand, community embedding can be used to optimize the node embedding by introducing a community-aware high-order proximity. Guided by this insight, we propose a novel community embedding framework that jointly solves the three tasks together. We evaluate such a framework on multiple real-world datasets, and show that it improves graph visualization and outperforms state-of-the-art baselines in various application tasks, e.g., community detection and node classification.
Sandro Cavallari, Vincent Wenchen Zheng, Hongyun Cai 0001, Kevin Chen-Chuan Chang, Erik Cambria
CIKM3
2017 SocialLens: Searching and Browsing Communities by Content and Interaction
abstract
Community analysis is an important task in graph mining. Most of the existing community studies are community detection, which aim to find the community membership for each user based on the user friendship links. However, membership alone, without a complete profile of what a community is and how it interacts with other communities, has limited applications. This motivates us to consider systematically profiling the communities and thereby developing useful community-level applications. In this paper, we introduce a novel concept of community profiling, upon which we build a SocialLens system1 to enable searching and browsing communities by content and interaction. We deploy SocialLens on two social graphs: Twitter and DBLP. We demonstrate two useful applications of SocialLens, including interactive community visualization and profile-aware community ranking.
Hongyun Cai 0001, Vincent Wenchen Zheng, Penghe Chen, Fanwei Zhu, Kevin Chen-Chuan Chang, Zi Huang
ICDE1
2017 From Community Detection to Community Profiling
abstract
Most existing community-related studies focus on detection, which aim to find the community membership for each user from user friendship links. However, membership alone, without a complete profile of what a community is and how it interacts with other communities, has limited applications. This motivates us to consider systematically profiling the communities and thereby developing useful community-level applications. In this paper, we for the first time formalize the concept of community profiling. With rich user information on the network, such as user published content and user diffusion links, we characterize a community in terms of both its internal content profile and external diffusion profile. The difficulty of community profiling is often underestimated. We novelly identify three unique challenges and propose a joint Community Profiling and Detection (CPD) model to address them accordingly. We also contribute a scalable inference algorithm, which scales linearly with the data size and it is easily parallelizable. We evaluate CPD on large-scale real-world data sets, and show that it is significantly better than the state-of-the-art baselines in various tasks.
Hongyun Cai 0001, Vincent Wenchen Zheng, Fanwei Zhu, Kevin Chen-Chuan Chang, Zi Huang
Proc. VLDB Endow.1
2016 Indexing evolving events from tweet streams
abstract
Tweet streams provide a variety of real-time information on dynamic social events. Although event detection has been actively studied, most of the existing approaches do not address the issue of efficient event monitoring in the presence of a large number of events detected from continuous tweet streams. In this paper, we capture the dynamics of events using four event operations: creation, absorption, split and merge.We also propose a novel event indexing structure, named Multi-layer Inverted List (MIL), for the acceleration of large-scale event search and update. We thoroughly study the problem of nearest neighbour search using MIL based on upper bound pruning. Extensive experiments have been conducted on a large-scale tweet dataset. The results demonstrate the promising performance of our method in terms of both efficiency and effectiveness.
Hongyun Cai 0001, Zi Huang, Divesh Srivastava, Qing Zhang 0001
ICDE1
2015 What are Popular: Exploring Twitter Features for Event Detection, Tracking and Visualization
abstract
As one of the most representative social media platforms, Twitter provides various real-life information on social events in real time. Despite that social event detection has been actively studied, tweet images, which appear in around 36 percent of the total tweets, have not been well utilized for this research problem. Most existing event detection methods tend to represent an image as a bag-of-visual-words and then process these visual words in the same way as textual words. This may not fully exploit the visual properties of images. State-of-the-art visual features like convolutional neural network (CNN) features have shown significant performance gains over the traditional bag-of-visual-words in unveiling the image's semantics. Unfortunately, they have not been employed in detecting events from social websites. Hence, how to make the most of tweet images to improve the performance of social event detection and visualization remains open. In this paper, we thoroughly study the impact of tweet images on social event detection for different event categories using various visual features. A novel topic model which jointly models five Twitter features (text, image, location, timestamp and hashtag) is designed to discover events from the sheer amount of tweets. Moreover, the evolutions of events are tracked by linking the events detected on adjacent days and each event is visualized by representative images selected on three predefined criteria. Extensive experiments have been conducted on a real-life tweet dataset to verify the effectiveness of our method.
Hongyun Cai 0001, Yang Yang 0002, Zi Huang
ACM Multimedia1
2015 Indexing Evolving Events from Tweet Streams
abstract
Tweet streams provide a variety of real-life and real-time information on social events that dynamically change over time. Although social event detection has been actively studied, how to efficiently monitor evolving events from continuous tweet streams remains open and challenging. One common approach for event detection from text streams is to use single-pass incremental clustering. However, this approach does not track the evolution of events, nor does it address the issue of efficient monitoring in the presence of a large number of events. In this paper, we capture the dynamics of events using four event operations (create, absorb, split, and merge), which can be effectively used to monitor evolving events. Moreover, we propose a novel event indexing structure, called Multi-layer Inverted List (MIL), to manage dynamic event databases for the acceleration of large-scale event search and update. We thoroughly study the problem of nearest neighbour search using MIL based on upper bound pruning, along with incremental index maintenance. Extensive experiments have been conducted on a large-scale real-life tweet dataset. The results demonstrate the promising performance of our event indexing and monitoring methods on both efficiency and effectiveness.
Hongyun Cai 0001, Zi Huang, Divesh Srivastava, Qing Zhang 0001
IEEE Trans. Knowl. Data Eng.1
2015 Social event identification and ranking on flickr
Hongyun Cai 0001, Zi Huang, Yang Yang 0002, Xiaofang Zhou 0001
World Wide Web2
2014 Multi-Output Regression with Tag Correlation Analysis for Effective Image Tagging
Hongyun Cai 0001, Zi Huang, Xiaofeng Zhu 0001, Qing Zhang 0001
DASFAA (2)1
2014 EventEye: Monitoring Evolving Events from Tweet Streams
abstract
With the rapid growth in popularity of social websites, social event detection has become one of the hottest research topics. However, continuously monitoring social events has not been well studied. In this demo, we present a novel system called EventEye to effectively monitor evolving events and visualize their evolving paths, which are discovered from tweet streams. In particular, four event operations are defined for our proposed stream clustering algorithm to capture the evolutions over time and a multi-layer indexing structure is designed to support efficient event search from large-scale event databases. In our system, events are visualized in different views, including evolution graph, timeline, map view, etc.
Hongyun Cai 0001, Zhongxian Tang, Yang Yang 0002, Zi Huang
ACM Multimedia1
2013 Spatio-temporal Event Modeling and Ranking
Hongyun Cai 0001, Zi Huang, Yang Yang 0002, Xiaofang Zhou 0001
WISE (2)2
2013 Imagilar: A Real-Time Image Similarity Search System on Mobile Platform
Bicheng Luo, Zi Huang, Hongyun Cai 0001, Yang Yang 0002
WISE (2)3
2012 Context Sensitive Tag Expansion with Information Inference
Hongyun Cai 0001, Zi Huang, Jie Shao 0001, Xue Li 0001
DASFAA (1)1