EDBT 2026 Demo / reviewers in the wild / expert
Jun Pang 0001
dblp:p/JunPang
· DBLP profile ↗
31ranked-venue papers in the field
2as first author
14since 2021 · last 2026
0000-0002-4521-4112ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 17 (2 first)Data Mining & Knowledge Discovery · 9Other / Interdisciplinary · 3Database Systems & Data Management · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large Language Models-Enhanced Semantic Diffusion for User-Centric RecommendationabstractRecently, knowledge graphs have been utilised in recommendation systems to improve accuracy by integrating item-side auxiliary information. However, structural user-side knowledge is difficult to construct and integrate due to inherent scarcity and improper granularity. This paper introduces a graph contrastive learning with Semantic transitions-Enhanced DIffusion architecture based on Large Language Models (LLMs) for user-side knowledge-aware Recommendation (SEDIRec). Specifically, our SEDIRec first leverages LLMs to infer user interests from historical behaviors, integrating this user-side information with item-side and collaborative data to construct main views. Then, two contrastive views are generated using diffusion models with semantic transitions: one at the user-side level and the other at the item-side level. For both contrastive views, we integrate user-side or item-side information with collaborative data to generate a user-item graph. Subsequently, each user-item graph is transformed into collaborative data spaces via diffusion models for generating contrastive views. This procedure not only enhances the alignment between user/item-side information and the semantic spaces of collaborative data but also effectively eliminates noise. Extensive experiments on three datasets reveal the superiority of SEDIRec, especially for users with sparse interactions. Xian Mo, Jun Pang 0001 |
WWW | 3 |
| 2025 | "Double vaccinated, 5G boosted!": Learning Attitudes towards COVID-19 Vaccination from Social MediaabstractThe sudden onset of the recently concluded COVID-19 pandemic has driven substantial progress in various scientific fields. One notable example is the comprehension of public vaccination attitudes and the timely monitoring of their fluctuations through social media platforms. This approach can serve as a cost-effective means to supplement surveys in gathering public vaccine hesitancy levels. In this article, we propose a deep learning framework leveraging textual posts on social media to extract and track users’ vaccination stances in near real time. Compared to previous works, we integrate into the framework the recent posts of a user’s social network friends to collaboratively detect the user’s genuine attitude towards vaccination. Based on our annotated dataset from X (formerly known as Twitter), the models instantiated from our framework can increase the performance of attitude extraction by up to 23% compared to the state-of-the-art text-only models. Using this framework, we successfully confirm the feasibility of using social media to track the evolution of vaccination attitudes in real life. In addition, we illustrate the generality of our framework in extracting other public opinions such as political ideology. We further show one practical use of our framework by validating the possibility of forecasting a user’s vaccine hesitancy changes with information perceived from social media. Ninghan Chen, Xihui Chen, Zhiqiang Zhong 0001, Jun Pang 0001 |
ACM Trans. Web | 4 |
| 2024 | Multi-Grained Semantics-Aware Graph Neural Networks (Extended abstract)abstractGraph Neural Networks (GNNs) are powerful techniques in representation learning for graphs and have been increasingly deployed in a multitude of different applications that involve node- and graph-wise tasks. Most existing studies solve either the node-wise task or the graph-wise task independently while they are inherently correlated. This work proposes a unified model, AdamGNN, to interactively learn node and graph representations in a mutual-optimisation manner. Compared with existing GNN models and graph pooling methods, AdamGNN enhances the node representation with the learned multi-grained semantics and avoids losing node features and graph structure information during pooling. Experiments on 14 real-world graph datasets show that AdamGNN can significantly outperform 17 competing models on both node- and graph-wise tasks. The ablation studies confirm the effectiveness of AdamGNN's components, and the last empirical analysis further reveals the ingenious ability of AdamGNN in capturing long-range interactions. This work was published at IEEE TKDE11Full paper is available at https://ieeexplore.ieee.org/document/9844866/. Zhiqiang Zhong 0001, Cheng-Te Li, Jun Pang 0001 |
ICDE | 3 |
| 2024 | Uplift Modeling Under Limited Supervision
George Panagopoulos, Daniele Malitesta, Fragkiskos D. Malliaros, Jun Pang 0001 |
ECML/PKDD (6) | 4 |
| 2024 | Exploring Unconfirmed Transactions for Effective Bitcoin Address ClusteringabstractThe advancement of clustering heuristics has demonstrated that the addresses of Bitcoin, which are protected by their anonymous mechanisms, can be de-anonymized. While the state-of-the-art (SOTA) clustering heuristics focus on confirmed transactions stored in the blockchain, they ignore unconfirmed transactions in the mempool. These unconfirmed transactions contain information about transactions before being stored in the blockchain, covering additional address associations that can improve Bitcoin address clustering. Kai Wang 0062, Yakun Cheng, Michael Wen Tong, Zhenghao Niu, Jun Pang 0001, Weili Han |
WWW | 5 |
| 2024 | A tale of two roles: exploring topic-specific susceptibility and influence in cascade predictionabstractAbstract We propose a new deep learning cascade prediction model CasSIM that can simultaneously achieve two most demanded objectives: popularity prediction and final adopter prediction. Compared to existing methods based on cascade representation, CasSIM simulates information diffusion processes by exploring users’ dual roles in information propagation with three basic factors: users’ susceptibilities, influences and message contents. With effective user profiling, we are the first to capture the topic-specific property of susceptibilities and influences. In addition, the use of graph neural networks allows CasSIM to capture the dynamics of susceptibilities and influences during information diffusion. We evaluate the effectiveness of CasSIM on three real-life datasets and the results show that CasSIM outperforms the state-of-the-art methods in popularity and final adopter prediction. Ninghan Chen, Xihui Chen, Zhiqiang Zhong 0001, Jun Pang 0001 |
Data Min. Knowl. Discov. | 4 |
| 2024 | Bridging Performance of X (formerly known as Twitter) Users: A Predictor of Subjective Well-Being During the PandemicabstractThe outbreak of the COVID-19 pandemic triggered the perils of misinformation over social media. By amplifying the spreading speed and popularity of trustworthy information, influential social media users have been helping overcome the negative impacts of such flooding misinformation. In this article, we use the COVID-19 pandemic as a representative global health crisisand examine the impact of the COVID-19 pandemic on these influential users’ subjective well-being (SWB), one of the most important indicators of mental health. We leverage X (formerly known as Twitter) as a representative social media platform and conduct the analysis with our collection of 37,281,824 tweets spanning almost two years. To identify influential X users, we propose a new measurement called user bridging performance (UBM) to evaluate the speed and wideness gain of information transmission due to their sharing. With our tweet collection, we manage to reveal the more significant mental sufferings of influential users during the COVID-19 pandemic. According to this observation, through comprehensive hierarchical multiple regression analysis , we are the first to discover the strong relationship between individual social users’ subjective well-being and their bridging performance. We proceed to extend bridging performance from individuals to user subgroups. The new measurement allows us to conduct a subgroup analysis according to users’ multilingualism and confirm the bridging role of multilingual users in the COVID-19 information propagation. We also find that multilingual users not only suffer from a much lower SWB in the pandemic, but also experienced a more significant SWB drop. Ninghan Chen, Xihui Chen, Zhiqiang Zhong 0001, Jun Pang 0001 |
ACM Trans. Web | 4 |
| 2024 | XRAD: Ransomware Address Detection Method based on Bitcoin Transaction RelationshipsabstractRecently, there is a surge in ransomware activities that encrypt users’ sensitive data and demand bitcoins for ransom payments to conceal the criminal’s identity. It is crucial for regulatory agencies to identify as many ransomware addresses as possible to accurately estimate the impact of these ransomware activities. However, existing methods for detecting ransomware addresses rely primarily on time-consuming data collection and clustering heuristics, and they face two major issues: (1) The features of an address itself are insufficient to accurately represent its activity characteristics, and (2) the number of disclosed ransomware addresses is extremely less than the number of unlabeled addresses. These issues lead to a significant number of ransomware addresses being undetected, resulting in a substantial underestimation of the impact of ransomware activities. To solve the above two issues, we propose an optimized ransomware address detection method based on Bitcoin transaction relationships, named XRAD , to detect more ransomware addresses with high performance. To address the first one, we present a cascade feature extraction method for Bitcoin transactions to aggregate features of related addresses after exploring transaction relationships. To address the second one, we build a classification model based on Positive-unlabeled learning to detect ransomware addresses with high performance. Extensive experiments demonstrate that XRAD significantly improves average accuracy, recall, and F1 score by 15.07%, 19.71%, and 34.83%, respectively, compared to state-of-the-art methods. In total, XRAD detects 120,335 ransomware activities from 2009 to 2023, revealing a development trend and average ransom payment per year that aligns with three reports by FinCEN, Chainalysis, and Coveware. Kai Wang 0062, Michael Wen Tong, Jun Pang 0001, Jitao Wang, Weili Han |
ACM Trans. Web | 3 |
| 2023 | Hierarchical message-passing graph neural networksabstractAbstract Graph Neural Networks (GNNs) have become a prominent approach to machine learning with graphs and have been increasingly applied in a multitude of domains. Nevertheless, since most existing GNN models are based on flat message-passing mechanisms, two limitations need to be tackled: (i) they are costly in encoding long-range information spanning the graph structure; (ii) they are failing to encode features in the high-order neighbourhood in the graphs as they only perform information aggregation across the observed edges in the original graph. To deal with these two issues, we propose a novel Hierarchical Message-passing Graph Neural Networks framework. The key idea is generating a hierarchical structure that re-organises all nodes in a flat graph into multi-level super graphs, along with innovative intra- and inter-level propagation manners. The derived hierarchy creates shortcuts connecting far-away nodes so that informative long-range interactions can be efficiently accessed via message passing and incorporates meso- and macro-level semantics into the learned node representations. We present the first model to implement this framework, termed Hierarchical Community-aware Graph Neural Network (HC-GNN), with the assistance of a hierarchical community detection algorithm. The theoretical analysis illustrates HC-GNN’s remarkable capacity in capturing long-range information without introducing heavy additional computation complexity. Empirical experiments conducted on 9 datasets under transductive, inductive, and few-shot settings exhibit that HC-GNN can outperform state-of-the-art GNN models in network analysis tasks, including node classification, link prediction, and community detection. Moreover, the model analysis further demonstrates HC-GNN’s robustness facing graph sparsity and the flexibility in incorporating different GNN encoders. Zhiqiang Zhong 0001, Cheng-Te Li, Jun Pang 0001 |
Data Min. Knowl. Discov. | 3 |
| 2023 | Multi-Grained Semantics-Aware Graph Neural NetworksabstractGraph Neural Networks (GNNs) are powerful techniques in representation learning for graphs and have been increasingly deployed in a multitude of different applications that involve node- and graph-wise tasks. Most existing studies solve either the node-wise task or the graph-wise task independently while they are inherently correlated. This work proposes a unified model, AdamGNN, to interactively learn node and graph representations in a mutual-optimisation manner. Compared with existing GNN models and graph pooling methods, AdamGNN enhances the node representation with the learned multi-grained semantics and avoids losing node features and graph structure information during pooling. Specifically, a differentiable pooling operator is proposed to adaptively generate a multi-grained structure that involves meso- and macro-level semantic information in the graph. We also devise the unpooling operator and theflybackaggregator in AdamGNN to better leverage the multi-grained semantics to enhance node representations. The updated node representations can further adjust the graph representation in the next iteration. Experiments on 14 real-world graph datasets show that AdamGNN can significantly outperform 17 competing models on both node- and graph-wise tasks. The ablation studies confirm the effectiveness of AdamGNN's components, and the last empirical analysis further reveals the ingenious ability of AdamGNN in capturing long-range interactions. Zhiqiang Zhong 0001, Cheng-Te Li, Jun Pang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | The Burden of Being a Bridge: Analysing Subjective Well-Being of Twitter Users During the COVID-19 Pandemic
Ninghan Chen, Xihui Chen, Zhiqiang Zhong 0001, Jun Pang 0001 |
ECML/PKDD (2) | 4 |
| 2022 | Personalised meta-path generation for heterogeneous graph neural networksabstractAbstract Recently, increasing attention has been paid to heterogeneous graph representation learning (HGRL), which aims to embed rich structural and semantic information in heterogeneous information networks (HINs) into low-dimensional node representations. To date, most HGRL models rely on hand-crafted meta-paths. However, the dependency on manually-defined meta-paths requires domain knowledge, which is difficult to obtain for complex HINs. More importantly, the pre-defined or generated meta-paths of all existing HGRL methods attached to each node type or node pair cannot be personalised to each individual node. To fully unleash the power of HGRL, we present a novel framework, Personalised Meta-path based Heterogeneous Graph Neural Networks (PM-HGNN), to jointly generate meta-paths that are personalised for each individual node in a HIN and learn node representations for the target downstream task like node classification. Precisely, PM-HGNN treats the meta-path generation as a Markov Decision Process and utilises a policy network to adaptively generate a meta-path for each individual node and simultaneously learn effective node representations. The policy network is trained with deep reinforcement learning by exploiting the performance improvement on a downstream task. We further propose an extension, PM-HGNN++, to better encode relational structure and accelerate the training during the meta-path generation. Experimental results reveal that both PM-HGNN and PM-HGNN++ can significantly and consistently outperform 16 competing baselines and state-of-the-art methods in various settings of node classification. Qualitative analysis also shows that PM-HGNN++ can identify meaningful meta-paths overlooked by human knowledge. Zhiqiang Zhong 0001, Cheng-Te Li, Jun Pang 0001 |
Data Min. Knowl. Discov. | 3 |
| 2022 | A Large-scale Empirical Analysis of Ransomware Activities in BitcoinabstractExploiting the anonymous mechanism of Bitcoin, ransomware activities demanding ransom in bitcoins have become rampant in recent years. Several existing studies quantify the impact of ransomware activities, mostly focusing on the amount of ransom. However, victims’ reactions in Bitcoin that can well reflect the impact of ransomware activities are somehow largely neglected. Besides, existing studies track ransom transfers at the Bitcoin address level, making it difficult for them to uncover the patterns of ransom transfers from a macro perspective beyond Bitcoin addresses. In this article, we conduct a large-scale analysis of ransom payments, ransom transfers, and victim migrations in Bitcoin from 2012 to 2021. First, we develop a fine-grained address clustering method to cluster Bitcoin addresses into users, which enables us to identify more addresses controlled by ransomware criminals. Second, motivated by the fact that Bitcoin activities and their participants already formed stable industries, such as Darknet and Miner , we train a multi-label classification model to identify the industry identifiers of users. Third, we identify ransom payment transactions and then quantify the amount of ransom and the number of victims in 63 ransomware activities. Finally, after we analyze the trajectories of ransom transferred across different industries and track victims’ migrations across industries, we find out that to obscure the purposes of their transfer trajectories, most ransomware criminals (e.g., operators of Locky and Wannacry) prefer to spread ransom into multiple industries instead of utilizing the services of Bitcoin mixers. Compared with other industries, Investment is highly resilient to ransomware activities in the sense that the number of users in Investment remains relatively stable. Moreover, we also observe that a few victims become active in the Darknet after paying ransom. Our findings in this work can help authorities deeply understand ransomware activities in Bitcoin. While our study focuses on ransomware, our methods are potentially applicable to other cybercriminal activities that have similarly adopted bitcoins as their payments. Kai Wang 0062, Jun Pang 0001, Dingjie Chen, Dapeng Huang, Chen Chen 0112, Weili Han |
ACM Trans. Web | 2 |
| 2021 | From #jobsearch to #mask: improving COVID-19 cascade prediction with spillover effectsabstractAn information outbreak occurs on social media along with the COVID-19 pandemic and leads to infodemic. Predicting the popularity of online content, known as cascade prediction, allows for not only catching in advance hot information that deserves attention, but also identifying false information that will widely spread and require quick response to mitigate its impact. Among the various information diffusion patterns leveraged in previous works, the spillover effect of the information exposed to users on their decision to participate in diffusing certain information is still not studied. In this paper, we focus on the diffusion of information related to COVID-19 preventive measures. Through our collected Twitter dataset, we validated the existence of this spillover effect. Building on the finding, we proposed extensions to three cascade prediction methods based on Graph Neural Networks (GNNs). Experiments conducted on our dataset demonstrated that the use of the identified spillover effect significantly improves the state-of-the-art GNNs methods in predicting the popularity of not only preventive measure messages, but also other COVID-19 related messages. Ninghan Chen, Xihui Chen, Zhiqiang Zhong 0001, Jun Pang 0001 |
ASONAM | 4 |
| 2020 | Higher-Order Graph Convolutional Embedding for Temporal Networks
Xian Mo, Jun Pang 0001, Zhiming Liu 0001 |
WISE (1) | 2 |
| 2020 | NeuLP: An End-to-End Deep-Learning Model for Link Prediction
Zhiqiang Zhong 0001, Yang Zhang 0016, Jun Pang 0001 |
WISE (1) | 3 |
| 2019 | A Graph-Based Approach to Explore Relationship Between Hashtags and Images
Zhiqiang Zhong 0001, Yang Zhang 0016, Jun Pang 0001 |
WISE | 3 |
| 2019 | An active learning-based approach for location-aware acquaintance inference
Bo-Heng Chen, Cheng-Te Li, Kun-Ta Chuang, Jun Pang 0001, Yang Zhang 0016 |
Knowl. Inf. Syst. | 4 |
| 2018 | Tagvisor: A Privacy Advisor for Sharing HashtagsabstractHashtag has emerged as a widely used concept of popular culture and campaigns, but its implications on people»s privacy have not been investigated so far. In this paper, we present the first systematic analysis of privacy issues induced by hashtags. We concentrate in particular on location, which is recognized as one of the key privacy concerns in the Internet era. By relying on a random forest model, we show that we can infer a user»s precise location from hashtags with accuracy of 70% to 76%, depending on the city. To remedy this situation, we introduce a system called Tagvisor that systematically suggests alternative hashtags if the user-selected ones constitute a threat to location privacy. Tagvisor realizes this by means of three conceptually different obfuscation techniques and a semantics-based metric for measuring the consequent utility loss. Our findings show that obfuscating as little as two hashtags already provides a near-optimal trade-off between privacy and utility in our dataset. This in particular renders Tagvisor highly time-efficient, and thus, practical in real-world settings. Yang Zhang 0016, Mathias Humbert, Tahleen A. Rahman, Cheng-Te Li, Jun Pang 0001, Michael Backes 0001 |
WWW | 5 |
| 2017 | Semantic Annotation for Places in LBSN through Graph EmbeddingabstractWith the prevalence of location-based social networks (LBSNs), automated semantic annotation for places plays a critical role in many LBSN-related applications. Although a line of research continues to enhance labeling accuracy, there is still a lot of room for improvement. The crucial problem is to find a high-quality representation for each place. In previous works, the representation is usually derived directly from observed patterns of places or indirectly from calculated proximity amongst places or their combination. In this paper, we also exploit the combination to represent places but present a novel semi-supervised learning framework based on graph embedding, called Predictive Place Embedding (PPE). For place proximity, PPE first learns user embeddings from a user-tag bipartite graph by minimizing supervised loss in order to preserve the similarity of users visiting analogous places. User similarity is then transformed into place proximity by optimizing each place embedding as the centroid of the vectors of its check-in users. Our underlying idea is that a place can be considered as a representative of all its visitors. For observed patterns, a place-temporal bipartite graph is used to further adjust place embeddings by reducing unsupervised loss. Extensive experiments on real large LBSNs show that PPE outperforms state-of-the-art methods significantly. Yan Wang 0014, Zongxu Qin, Jun Pang 0001, Yang Zhang 0016, Jin Xin |
CIKM | 3 |
| 2017 | DeepCity: A Feature Learning Framework for Mining Location Check-Ins
Jun Pang 0001, Yang Zhang 0016 |
ICWSM | 1 |
| 2017 | Does #like4like indeed provoke more likes?abstractHashtags, created by social network users, have gained a huge popularity in recent years. As a kind of metatag for organizing information, hashtags in online social networks, especially in Instagram, have greatly facilitated users' interactions. In recent years, academia starts to use hashtags to reshape our understandings on how users interact with each other. #like4like is one of the most popular hashtags in Instagram with more than 290 million photos appended with it, when a publisher uses #like4like in one photo, it means that he will like back photos of those who like this photo. Different from other hashtags, #like4like implies an interaction between a photo's publisher and a user who likes this photo, and both of them aim to attract likes in Instagram. In this paper, we study whether #like4like indeed serves the purpose it is created for, i.e., will #like4like provoke more likes? We first perform a general analysis of #like4like with 1.8 million photos collected from Instagram, and discover that its quantity has dramatically increased by 1,300 times from 2012 to 2016. Then, we study whether #like4like will attract likes for photo publishers; results show that it is not #like4like but actually photo contents attract more likes, and the lifespan of a #like4like photo is quite limited. In the end, we study whether users who like #like4like photos will receive likes from #like4like publishers. However, results show that more than 90% of the publishers do not keep their promises, i.e., they will not like back others who like their #like4like photos; and for those who keep their promises, the photos which they like back are often randomly selected. Yang Zhang 0016, Minyue Ni, Weili Han, Jun Pang 0001 |
WI | 4 |
| 2016 | On Impact of Weather on Human Mobility in Cities
Jun Pang 0001, Polina Zablotskaia, Yang Zhang 0016 |
WISE (2) | 1 |
| 2015 | Distance and Friendship: A Distance-Based Model for Link Prediction in Social Networks
Yang Zhang 0016, Jun Pang 0001 |
APWeb | 2 |
| 2015 | Inferring Friendship from Check-in Data of Location-Based Social NetworksabstractWith the ubiquity of GPS-enabled devices and location-based social network services, research on human mobility becomes quantitatively achievable. Understanding it could lead to appealing applications such as city planning and epidemiology. In this paper, we focus on predicting whether two individuals are friends based on their mobility information. Intuitively, friends tend to visit similar places, thus the number of their co-occurrences should be a strong indicator of their friendship. Besides, the visiting time interval between two users also has an effect on friendship prediction. By exploiting machine learning techniques, we construct two friendship prediction models based on mobility information. The first model focuses on predicting friendship of two individuals with only one of their co-occurred places' information. The second model proposes a solution for predicting friendship of two individuals based on all their co-occurred places. Experimental results show that both of our models outperform the state-of-the-art solutions. Jun Pang 0001, Yang Zhang 0016 |
ASONAM | 2 |
| 2015 | Community-Driven Social Influence Analysis and Applications
Yang Zhang 0016, Jun Pang 0001 |
ICWE | 2 |
| 2014 | Measuring User Similarity with Trajectory Patterns: Principles and New Metrics
Xihui Chen, Ruipeng Lu, Xiaoxing Ma, Jun Pang 0001 |
APWeb | 4 |
| 2014 | MinUS: Mining User Similarity with Trajectory Patterns
Xihui Chen, Piotr Kordy, Ruipeng Lu, Jun Pang 0001 |
ECML/PKDD (3) | 4 |
| 2014 | Protecting query privacy in location-based services
Xihui Chen, Jun Pang 0001 |
GeoInformatica | 2 |
| 2014 | Constructing and Comparing User Mobility ProfilesabstractNowadays, the accumulation of people's whereabouts due to location-based applications has made it possible to construct their mobility profiles. This access to users' mobility profiles subsequently brings benefits back to location-based applications. For instance, in on-line social networks, friends can be recommended not only based on the similarity between their registered information, for instance, hobbies and professions but also referring to the similarity between their mobility profiles. In this article, we propose a new approach to construct and compare users' mobility profiles. First, we improve and apply frequent sequential pattern mining technologies to extract the sequences of places that a user frequently visits and use them to model his mobility profile. Second, we present a new method to calculate the similarity between two users using their mobility profiles. More specifically, we identify the weaknesses of a similarity metric in the literature, and propose a new one which not only fixes the weaknesses but also provides more precise and effective similarity estimation. Third, we consider the semantics of spatio-temporal information contained in user mobility profiles and add them into the calculation of user similarity. It enables us to measure users' similarity from different perspectives. Two specific types of semantics are explored in this article: location semantics and temporal semantics . Last, we validate our approach by applying it to two real-life datasets collected by Microsoft Research Asia and Yonsei University, respectively. The results show that our approach outperforms the existing works from several aspects. Xihui Chen, Jun Pang 0001, Ran Xue |
ACM Trans. Web | 2 |
| 2011 | Fast leader election in anonymous rings with bounded expected delay
Rena Bakhshi, Jörg Endrullis, Wan J. Fokkink, Jun Pang 0001 |
Inf. Process. Lett. | 4 |