VLDB 2026 Research / reviewers in the wild / expert
Heli Sun
dblp:87/6001
· DBLP profile ↗
64ranked-venue papers
20as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 28 · 11 first-author · 12 since 2021Artificial intelligence and machine learning · 25 · 10 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Software engineering, systems software and programming languages · 2Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MgSAN: Multi-Graph Semantic-Aware Adaptive Graph Convolutional Network for Fake News DetectionabstractThe widespread dissemination and misleading impact of fake news on the web have become a significant concern for the public and the government. Discovering fake news is crucial for ensuring that users receive authentic information and maintaining social harmony. However, most existing entity-based fake news detection methods have two issues: i) methods for acquiring additional information through entities lack flexibility and real-time capabilities. ii) approaches using entities to capture news semantics have not adequately revealed the interactions between words in the text. To address these issues, we propose aMulti-graphSemantic-awareAdaptive Graph ConvolutionalNetwork (MgSAN), which comprehensively captures the semantic information of news texts by constructing multiple semantic graphs and learns the features from these graph structures using an adaptive graph convolutional network (SwiGCN). Specifically, we design a global semantic interaction graph to capture the complex interactions between words, generating a comprehensive textual semantic representation. We also employ an entity-noun relationship graph to mine deep semantic associations, enhancing the model's understanding of fine-grained textual deep meanings. Additionally, we develop an adaptive graph convolutional network to effectively extract and aggregate feature information from different graph structures. Finally, we introduce a fusion module to integrate both global and local fine-grained semantic information, forming a rich composite semantic representation, thereby improving the effectiveness of fake news detection. Extensive experimental results on three public benchmark datasets verify the effectiveness and superior performance of MgSAN, outperforming state-of-the-art detection models. Liang He 0006, Heli Sun |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2026 | Self-supervised Graph Neural Sequential Recommendation with Disentangling Long and Short-Term InterestabstractIn real-world scenarios, a large amount of noise in user historical behaviors obstructs the reflection of their genuine interests. The long-tail distribution of user-item interactions also makes it difficult to capture interest evolution patterns from historical sequences. Moreover, as user behavior sequences continue to grow, solely relying on conventional sequence models is insufficient to extract user interest information and learn accurate sequence representations, thus limiting recommendation accuracy. To address these issues, we propose a self-supervised graph neural sequential recommendation model called LS4SRec, which disentangles users’ long- and short-term interests. Specifically, LS4SRec constructs two independent interest encoders to extract users’ long- and short-term interests. By utilizing the global user behavior sequence graph WITG to provide additional collaborative signals for each interaction sequence, we alleviate the issue of data sparsity. Subsequently, contrastive learning is applied to WITG to remove noise information and enhance the sequence representation. Further, interest allocation matrices and sequence models are utilized to model users’ interest evolution patterns. Finally, we introduce sequence graph data augmentation methods and long- and short-term interest pseudo-label construction methods to generate unsupervised signals that assist in model training. Extensive experiments conducted on real-world data validate the effectiveness of our proposed model. Our model implementation codes are available at the link https://github.com/jiubaoyibao/LS4SRec . Liang He 0006, Wujie Yan, Tingzhou Yi, Heli Sun |
Trans. Recomm. Syst. | 4 |
| 2025 | Aspect Enhancement and Text Simplification in Multimodal Aspect-Based Sentiment Analysis for Multi-Aspect and Multi-Sentiment ScenariosabstractMultimodal Aspect-Based Sentiment Analysis (MABSA) plays a pivotal role in the advancement of sentiment analysis technology. Although current methods strive to integrate multimodal information to enhance the performance of sentiment analysis, they still face two critical challenges when dealing with multi-aspect and multi-sentiment data: i) the importance of aspect terms within multimodal data is often overlooked, and ii) models fail to accurately associate specific aspect terms with corresponding sentiment words in multi-aspect and multi-sentiment sentences. To tackle these problems, we propose a novel multimodal aspect-based sentiment analysis method that combines Aspect Enhancement and Text Simplification (AETS). Specifically, we develop an aspect enhancement module that boosts the ability of model to discern relevant aspect terms. Concurrently, we employ text simplification module to simplify and restructure multi-aspect and multi-sentiment texts, accurately capturing aspects and their corresponding sentiments while reducing irrelevant information. Leveraging this method, we perform three tasks including multimodal aspect term extraction, multimodal aspect sentiment classification, and joint multimodal aspect-based sentiment analysis. Experimental results indicate that our proposed AETS model achieved state-of-the-art performance on two benchmark datasets. Heli Sun, Qunshu Gao, Liang He 0006 |
AAAI | 2 |
| 2025 | AVF-MAE++: Scaling Affective Video Facial Masked Autoencoders via Efficient Audio-Visual Self-Supervised LearningabstractAffective Video Facial Analysis (AVFA) is important for advancing emotion-aware AI, yet the persistent data scarcity in AVFA presents challenges. Recently, the self-supervised learning (SSL) technique of Masked Autoencoders (MAE) has gained significant attention, particularly in its audio-visual adaptation. Insights from general domains suggest that scaling is vital for unlocking impressive improvements, though its effects on AVFA remain largely unexplored. Additionally, capturing both intra- and inter-modal correlations through scalable representations is a crucial challenge in this field. To tackle these gaps, we introduce AVF-MAE++, a series audio-visual MAE designed to explore the impact of scaling on AVFA with a focus on advanced correlation modeling. Our method incorporates a novel audio-visual dual masking strategy and an improved modality encoder with a holistic view to better support scalable pre-training. Furthermore, we propose the Iteratively Audio-Visual Correlations Learning Module to improve correlations capture within the SSL framework, bridging the limitations of prior methods. To support smooth adaptation and mitigate overfitting, we also introduce a progressive semantics injection strategy, which structures training in three stages. Extensive experiments across 17 datasets, spanning three key AVFA tasks, demonstrate the superior performance of AVFMAE++, establishing new state-of-the-art outcomes. Ablation studies provide further insights into the critical design choices driving these gains. Code is released at this URL. Heli Sun, Jiayu Nie, Junxiao Xue, Liang He 0006 |
CVPR | 2 |
| 2025 | Sentiment-enhanced Multi-hop Connected Graph Attention Network for Multimodal Aspect-Based Sentiment AnalysisabstractMultimodal aspect-based sentiment analysis aims to extract aspects from different data sources and recognize the corresponding sentiments. While current research has broadly focused on syntax relation-driven semantic comprehension, the impact of the importance of different syntactic relations on semantic understanding has not been adequately investigated. To address this issue, we propose a Sentiment-enhanced Multi-hop Connected Graph Attention Network (MCG), aiming to enhance the discriminative capability of model for sentiments and to delve into the syntactic relationships within the text. Firstly, we design a contrastive sentiment-enhanced pre-training task that expands the diversity and complexity of training samples to improve the recognition of multiple sentiments. Secondly, we construct a multi-hop connected syntactic dependency graph to deeply explore the rich syntactic dependencies in the text and to reveal the differences among syntactic relations. Moreover, we develop a multi-hop connected graph attention mechanism that enables the model to focus on the key syntactic relations within the syntactic structure, thereby enhancing the comprehension and predictive capabilities of model in multimodal sentiment analysis. Experimental results on two benchmark datasets demonstrate that our method outperforms state-of-the-art methods. The source code is provided in the supplementary materials. Heli Sun, Xiaoyong Huang, Ruichen Cao |
IJCAI | 2 |
| 2025 | Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and BaselineabstractNowadays, short-form videos (SVs) are essential to web information acquisition and sharing in our daily life. The prevailing use of SVs to spread emotions leads to the necessity of conducting video emotion analysis (VEA) towards SVs. Considering the lack of SVs emotion data, we introduce a large-scale dataset named eMotions, comprising 27,996 videos. Meanwhile, we alleviate the impact of subjectivities on labeling quality by emphasizing better personnel allocations and multi-stage annotations. In addition, we provide the category-balanced and test-oriented variants through targeted data sampling. Some commonly used videos, such as facial expressions, have been well studied. However, it is still challenging to analysis the emotions in SVs. Since the broader content diversity brings more distinct semantic gaps and difficulties in learning emotion-related features, and there exists local biases and collective information gaps caused by the emotion inconsistence under the prevalently audio-visual co-expressions. To tackle these challenges, we present an end-to-end audio-visual baseline AV-CANet which employs the video transformer to better learn semantically relevant representations. We further design the Local-Global Fusion Module to progressively capture the correlations of audio-visual features. The EP-CE Loss is then introduced to guide model optimization. Extensive experimental results across seven datasets demonstrate the effectiveness of AV-CANet, while providing broad insights for future works. Besides, we explore the key components of AV-CANet by ablation studies. Datasets and code are released at https://github.com/XuecWu/eMotions. Heli Sun, Junxiao Xue, Jiayu Nie, Xiangyan Kong, Ruofan Zhai, Danlei Huang, Liang He 0006 |
ICMR | 2 |
| 2025 | HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-trainingabstractAdvances in Generative AI have made video-level deepfake detection increasingly challenging, exposing the limitations of current detection techniques. In this paper, we present HOLA, our solution to the Video-Level Deepfake Detection track of 2025 1M-Deepfakes Detection Challenge. Inspired by the success of large-scale pre-training in the general domain, we first scale audio-visual self-supervised pre-training in the multimodal video-level deepfake detection, which leverages our self-built dataset of 1.81M samples, thereby leading to a unified two-stage framework. To be specific, HOLA features an iterative-aware cross-modal learning module for selective audio-visual interactions, hierarchical contextual modeling with gated aggregations under the local-global perspective, and a pyramid-like refiner for scale-aware cross-grained semantic enhancements. Moreover, we propose the pseudo supervised singal injection strategy to further boost model performance. Extensive experiments across expert models and MLLMs impressivly demonstrate the effectiveness of our proposed HOLA. We also conduct a series of ablation studies to explore the crucial design factors of our introduced components. Remarkably, our HOLA ranks 1st, outperforming the second by 0.0476 AUC on the TestA set. Heli Sun, Danlei Huang, Xinyi Yin, Hao Wang 0182, Jia Zhang 0016, Fei Wang 0128, Peihao Guo, Suyu Xing, Junxiao Xue, Liang He 0006 |
ACM Multimedia | 2 |
| 2025 | Contrastive deep graph clustering with hard boundary sample awareness
Heli Sun, Xiaoyong Huang, Pan Lou, Liang He 0006 |
Inf. Process. Manag. | 2 |
| 2024 | Joint Multimodal Aspect Sentiment Analysis with Aspect Enhancement and Syntactic Adaptive Learning
Heli Sun, Qunshu Gao, Tingzhou Yi, Liang He 0006 |
IJCAI | 2 |
| 2024 | Semantic community query in a large-scale attributed graph based on an attribute cohesiveness optimization strategyabstractAbstract The task of a semantic community query is to obtain a subgraph based on a given query vertex (or vertex set) and other query parameters in an attributed graph such that belongs to , contains and satisfies a predefined community cohesiveness model. In most cases, existing community query models based on the network structure for traditional attributed networks usually lack community semantics. However, the features of vertex attributes, especially the attributes of the query vertices, which are closely related to the community semantics, are rarely considered in an attributed graph. Existing community query algorithms based on both structure cohesiveness and attribute cohesiveness usually do not take the attributes of the query vertex as an important factor of the community cohesiveness model, which leads to weak semantics of the communities. This paper proposes a semantic community query method named in a large‐scale attributed graph. First, the k‐core structure model is adopted as the structure cohesiveness of our community query model to obtain a subgraph of the original graph. Second, we define attribute cohesiveness based on the average distance between the query vertices and other vertices in terms of attributes in the community to prune the subgraph and obtain the semantic community. In order to improve the community query efficiency in large‐scale attributed graphs, applies two heuristic pruning strategies. The experimental results show that our method outperforms the existing community query methods in multiple evaluation metrics and is ideal for querying semantic communities in large‐scale attributed graphs. Jinhuan Ge, Heli Sun, Yezhi Lin, Liang He 0006 |
Expert Syst. J. Knowl. Eng. | 2 |
| 2024 | Novel behavior-enhanced long- and short-term interest model for sequential recommendation
Heli Sun, Liang He 0006 |
Inf. Sci. | 2 |
| 2024 | Public Bike Scheduling Strategy Based on Demand Prediction for Unbalanced Life-Value DistributionabstractPublic bikes have emerged as a significant mode of transportation for commuters. However, public bikes can be damaged to varying degrees as use frequency increases. The stations occupied by damaged bikes are of low usability, which may result in fewer users and a waste of resources. In this paper, we investigate a global scheduling solution for a bike system with unbalanced life-value distribution and design a scheduling approach. Firstly, we employ the Weibull model to estimate bike life to quantify use load. Moreover, we design a clustering algorithm to partition station regions. We also design a demand prediction model named ST-SAGCN to capture dynamic spatial correlations and spatio-temporal correlations simultaneously. We also propose a scheduling method that meets station demand as much as possible to balance station use and bike-life distribution. We conduct experiments on two public bike datasets covering different time spans in New York and Washington, and compare the bike-life distribution after scheduling with that of real-world actual conditions. The experimental results attest to the effectiveness of our approach to balancing bike use-load. The code is publicly available athttps://github.com/Zayn-Tang/BikeSchedulingStrategy. Heli Sun, Zunye Tang, Mengting Cao, Yu Wang 0330, Zhou Yang 0004, Haokun Xue, Ruirui Xue, Liang He 0006, Hui Xiong 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | KTPGN: Novel event-based group recommendation method considering implicit social trust and knowledge propagation
Heli Sun, Liang He 0006 |
Inf. Sci. | 2 |
| 2023 | What Your Next Check-in Might Look Like: Next Check-in Behavior PredictionabstractIn recent years, the next-POI recommendation has become a trending research topic in the field of trajectory data mining. For protection of user privacy, users’ complete GPS trajectories are difficult to obtain. The check-in information posted by users on social networks has become an important data source for Spatio-temporal Trajectory research. However, state-of-the-art methods neglect the social meaning and the information dissemination function of check-in behavior. The social meaning is an important reason why users are willing to post check-in on social networks, and the information dissemination function means, users can affect each other’s behavior by check-ins. The above characteristics of the check-in behavior make it different from the visiting behavior. We consider a new problem of predicting the next check-in behavior including the check-in time, the POI (point-of-interest) where the check-in is located, functional semantics of the POI, and so on. To solve the proposed problem, we build a multi-task learning model called DPMTM, and a pre-training module is designed to extract dynamic social semantics of check-in behaviors. Our results show that the DPMTM model works well in the check-in behavior problem. Heli Sun, Xuguang Chu, Junzhi Lu, Liang He 0006, Zhi Wang 0002, Hui Xiong 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2023 | Graph Neural Networks with Motisf-aware for Tenuous Subgraph FindingabstractTenuous subgraph finding aims to detect a subgraph with few social interactions and weak relationships among nodes. Despite significant efforts made on this task, they are mostly carried out in view of graph-structured data. These methods depend on calculating the shortest path and need to enumerate all the paths between nodes, which suffer the combinatorial explosion. Moreover, they all lack the integration of neighborhood information. To this end, we propose a novel model named Graph Neural Network with Motif-aware for tenuous subgraph finding (GNNM), a neighborhood aggregation-based GNN framework that can capture the latent relationship between nodes. We design a GNN module to project nodes into a low-dimensional vector combining the higher-order correlation within nodes based on a motif-aware module. Then we design greedy algorithms in vector space to obtain a tenuous subgraph whose size is greater than a specified constraint. Particularly, considering that existing evaluation indicators cannot capture the latent friendship between nodes, we introduce a novel Potential Friend concept to measure the tenuity of a graph from a new perspective. Experimental results on the real-world and synthetic datasets demonstrate that our proposed method GNNM outperforms existing algorithms in efficiency and subgraph quality. Heli Sun, Miaomiao Sun, Xuechun Liu, Liang He 0006, Xiaolin Jia |
ACM Trans. Knowl. Discov. Data | 1 |
| 2023 | Platform-Oriented Event Time AllocationabstractOnline Event-based social networks (EBSNs), such as Meetup and Whova, which provide platforms for users to publish, arrange and participate in events, have become increasingly popular. A major challenge for managing EBSNs is to generate the most satisfactory event arrangement, i.e. events are scheduled at the reasonable time to attract maximum number of participants. Existing approaches usually focus on assigning a set of events organized by the same group to time intervals, but ignore the competitive relationships among different event organizers, which will lead to event time allocations unacceptable to organizers. Thus, a more intelligent EBSNs platform that allocates social events properly in a global view (i.e. the perspective of platform) is desired. In this paper, we first formally define the problem of Platform-oriented Event Time Allocation (PETA), which contains two parts: the prediction of event feasible time period and the event time allocation. Unfortunately, we find that the PETA problem is NP-hard due to the global conflict constraints on events. Thus, we propose design a greedy algorithm and two approximation algorithms to solve the PETA problem. Finally, we conduct extensive experiments on both real and synthetic datasets to test the effectiveness and efficiency of the proposed algorithms. Heli Sun, Jingyu Jia, Hui Xiong 0001, Liang He 0006, Xinwang Liu 0002, Shaojie Qiao, Jizhong Zhao |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Platform-Oriented Event Time Allocation(Extended Abstract)abstractOnline Event-based social networks (EBSNs), such as Meetup and Whova, which provide platforms for users to publish, arrange and participate in events, have become increasingly popular. A major challenge for managing EBSNs is to generate the most satisfactory event arrangement. Existing approaches usually focus on assigning a set of events organized to time intervals, but ignore the competitive relationships among different event organizers, which will lead to event time allocations unacceptable to organizers. Thus, a more intelligent EBSNs platform that allocates social events properly in a global view (i.e. the perspective of platform) is desired. In this work, we first formally define the problem of Platform-oriented Event Time Allocation (PETA), which contains two parts: the prediction of event feasible time period and the event time allocation. We propose a method to calculate event feasible time period based on event time prediction, and design a greedy algorithm and two approximation algorithms to solve the PETA problem. Extensive experiments on both real and synthetic datasets demonstrate that the proposed algorithms have high effectiveness and efficiency. Heli Sun, Jingyu Jia, Hui Xiong 0001, Liang He 0006, Xinwang Liu 0002, Shaojie Qiao, Jizhong Zhao |
ICDE | 1 |
| 2022 | Central Station-Based Demand Prediction for Determining Target Inventory in a Bike-Sharing SystemabstractAbstract Predicting the bike demand can help rebalance the bikes and improve the service quality of a bike-sharing system. A lot of works focus on predicting the bike demand for all the stations, which is unnecessary as the travel cost of rebalance operations increases sharply as the number of stations increases. In this paper, we propose a framework for predicting the hourly bike demand based on the central stations we define. Firstly, we propose Two-Stage Station Clustering Algorithm to assign central stations and common stations into each cluster. Secondly, we propose a hierarchical prediction model to predict the hourly bike demand for every cluster and each central station progressively. Thirdly, we use a well-studied queuing model to determine the target initial inventory for each central station. The most innovative contribution of this paper is proposing the concept of central station, the use of a novel algorithm to cluster the central stations and present a hierarchical model, containing the Time and Weather Similarity Weighted K-Nearest Neighbor Algorithm and a linear model to predict the bike demand for central stations. The experimental results on the New York citi bike system demonstrate that our proposed method is more accurate than other methods in solving existing problems. Heli Sun, He Li 0006, Longji Huang |
Comput. J. | 2 |
| 2022 | Querying Tenuous Group in Attributed NetworksabstractAbstract Finding groups in networks is very common in many practical applications, and most work mainly focus on dense groups. However, in scenarios like reviewer selection or weak social friends recommendation, we need to emphasize the privacy of individuals or minimize the possibility of information dissemination. So the internal relationship between individuals should be as tenuous as possible, but existing works cannot suit well to the requirement. Some works have focused on finding tenuous groups. However, these works only aim to find the most tenuous group and do not consider containing certain vertices. In this paper, we study the problem of finding tenuous groups in attributed networks that contain specific vertices. We first propose a new problem called Tenuous Attributed Group Query, and a new indicator, k-tenuity, to measure the structural tenuity of a group. Then we propose a method TAG-Basic to find proper groups by gradually selecting the vertices with optimal influence. We further design an advanced method TAG-ADV to improve the efficiency by forming a candidate set before selecting the optimal vertex. Experiment results show that k-tenuity is more effective than other state-of-the-art measurements, and our methods obtain the best result on group quality compared with other benchmark methods. Heli Sun, Liang He 0006, Jiyin Chen, Xiaolin Jia |
Comput. J. | 2 |
| 2022 | Regional-based multi-module spatial-temporal networks predicting city-wide taxi pickup/dropoff demand from origin to destinationabstractAbstract Taxi demand forecasting from origin to destination (OD) is an important component in managing public transportation needs on a city‐wide scale. Accurate taxi demand forecasting may provide several benefits, including economic and traffic flow optimization. However, due to complicated spatial–temporal connections and irregular distant locations, predicting taxi demand becomes difficult. To address these issues, we proposed a novel architecture of multi‐module spatial–temporal networks to collectively predict city‐wide OD taxi demand. In our work to deal with adjacent areas in a city, we employed the 3D convolutional neural networks to extract the spatial–temporal dependencies and learn the OD taxi demand pattern. To handle remote areas, we created an attention‐based auto encoder‐decoder, in which the input of the set of convolutional layers creates a feature matrix. The feature matrix embeds the spatial–temporal correlation jointly and passing through the encoder layer. We encode the spatial–temporal features with the help of pooling layer, then flatten layer used the back‐propagation method to decode the weight matrix. We apply the normalization function to determine the demand pattern influences. The influence vector we compute with the Euclidean distance formula to determine the similarity of all distant regions. Finally, we use the attention mechanism to calculate the attention weight score for each region that its neighbour impacted. Then we used long short‐term memory, we captured the significant relationship of spatial–temporal dependencies with external factors. We train our model simultaneously to forecast city‐wide taxi services OD. We have performed comprehensive experiments and large‐scale dataset comparisons that reveal the taxi demand prediction problem. Zain Ul Abideen 0004, Heli Sun, Zhou Yang 0004, Hamza Fahim |
Expert Syst. J. Knowl. Eng. | 2 |
| 2022 | A novel meta-graph-based attention model for event recommendation
Heli Sun, Liang He 0006, Xiaolin Jia |
Neural Comput. Appl. | 2 |
| 2022 | Predicting Future Locations with Semantic TrajectoriesabstractLocation prediction has attracted much attention due to its important role in many location-based services, including taxi services, route navigation, traffic planning, and location-based advertisements. Traditional methods only use spatial-temporal trajectory data to predict where a user will go next. The divorce of semantic knowledge from the spatial-temporal one inhibits our better understanding of users’ activities. Inspired by the architecture of Long Short Term Memory (LSTM), we design ST-LSTM, which draws on semantic trajectories to predict future locations. Semantic data add a new dimension to our study, increasing the accuracy of prediction. Since semantic trajectories are sparser than the spatial-temporal ones, we propose a strategic filling algorithm to solve this problem. In addition, as the prediction is based on the historical trajectories of users, the cold-start problem arises. We build a new virtual social network for users to resolve the issue. Experiments on two real-world datasets show that the performance of our method is superior to those of the baselines. Heli Sun, Xianglan Guo, Zhou Yang 0004, Xuguang Chu, Xinwang Liu 0002, Liang He 0006 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2022 | Robust Traffic Speed Inference With Ensemble LearningabstractTraffic speed inference enables many applications that are essential for everyday life. Most traffic-prediction approaches assume that a constant number of sensors are deployed on the roads, whether they are either stationary loop detectors or vehicles equipped with Global Positioning System (GPS) tracking devices. The static nature of those fixtures limits their ability to adapt to scenarios that are more dynamic. Rather than relying on several fixed sensors to detect changes and infer traffic, we use crowdsourcing to judiciously select individuals and then make predictions. Our solution consists of three core components: dynamic seed selection, regional cluster building and ensemble traffic prediction. In the first phase, we employ Efficient Transition Probability (ETP) to evaluate candidate seed sets. Road clusters are then formed using hierarchical clustering that is tweaked by a dynamic programming technique. This method assesses the eccentricity of every cluster to bond every road within each cluster more closely. Subsequently, we develop an ensemble-learning strategy in conjunction with Lasso regression to forecast traffic. The strength of our ensemble approach is its ability to manage absent seeds, a condition that has never been investigated, to our knowledge. Substantial experimental evaluation indicates that our claim of dynamic updates is valid and effective. Our solution outperforms the state-of-the-art techniques by a wide margin, in terms of prediction accuracy. Zhou Yang 0004, Heli Sun, Liang He 0006, Xiaolin Jia, Jizhong Zhao, Shaojie Qiao |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Graph Community InfomaxabstractGraph representation learning aims at learning low-dimension representations for nodes in graphs, and has been proven very useful in several downstream tasks. In this article, we propose a new model, Graph Community Infomax (GCI), that can adversarial learn representations for nodes in attributed networks. Different from other adversarial network embedding models, which would assume that the data follow some prior distributions and generate fake examples, GCI utilizes the community information of networks, using nodes as positive(or real) examples and negative(or fake) examples at the same time. An autoencoder is applied to learn the embedding vectors for nodes and reconstruct the adjacency matrix, and a discriminator is used to maximize the mutual information between nodes and communities. Experiments on several real-world and synthetic networks have shown that GCI outperforms various network embedding methods on community detection tasks. Heli Sun, Bing Lv, Wujie Yan, Liang He 0006, Shaojie Qiao |
ACM Trans. Knowl. Discov. Data | 1 |
| 2021 | LPX: Overlapping community detection based on X-means and label propagation algorithm in attributed networksabstractAbstract Traditional community detection methods in attributed networks (eg, social network) usually disregard abundant node attribute information and only focus on structural information of a graph. Existing community detection methods in attributed networks are mostly applied in the detection of nonoverlapping communities and cannot be directly used to detect the overlapping structures. This article proposes an overlapping community detection algorithm in attributed networks. First, we employ the modified X‐means algorithm to cluster attributes to form different themes. Second, we employ the label propagation algorithm (LPA), which is based on neighborhood network conductance for priority and the rule of theme weight, to detect communities in each theme. Finally, we perform redundant processing to form the final community division. The proposed algorithm improves the X‐means algorithm to avoid the effects of outliers. Problems of LPA such as instability of division and adjacent communities being easily merged can be corrected by prioritizing the node neighborhood network conductance. As the community is detected in the attribute subspace, the algorithm can find overlapping communities. Experimental results on real‐attributed and synthetic‐attributed networks show that the performance of the proposed algorithm is excellent with multiple evaluation metrics. Jinhuan Ge, Heli Sun, Chenhao Xue, Liang He 0006, Xiaolin Jia, Jiyin Chen |
Comput. Intell. | 2 |
| 2021 | Distance dynamics based overlapping semantic community detection for node-attributed networksabstractAbstract In recent years, due to the rise of social, biological, and other rich content graphs, several novel community detection methods using structure and node attributes have been proposed. Moreover, nodes in a network are naturally characterized by multiple community memberships and there is growing interest in overlapping community detection algorithms. In this paper, we design a weighted vertex interaction model based on distance dynamics to divide the network, furthermore, we propose a distance Dynamics‐based Overlapping Semantic Community detection algorithm(DOSC) for node‐attribute networks. The method is divided into three phases: Firstly, we detect local single‐attribute subcommunities in each attribute‐induced graph based on the weighted vertex interaction model. Then, a hypergraph is constructed by using the subcommunities obtained in the previous step. Finally, the weighted vertex interaction model is used in the hypergraph to get global semantic communities. Experimental results in real‐world networks demonstrate that DOSC is a more effective semantic community detection method compared with state‐of‐the‐art methods. Heli Sun, Xiaolin Jia, Ruodan Huang |
Comput. Intell. | 1 |
| 2021 | A truss-based approach for densest homogeneous subgraph mining in node-attributed graphsabstractAbstract In a wide range of graph analysis tasks such as community detection and event detection, densest subgraph mining is important and primitive. With the development of social network, densest subgraph mining not only need to consider the structural data but also the attributes information, which descripts the features of nodes or edges. However, there are few researches on densest subgraph mining with attribute description. In this article, we only focus on the node‐attributed graph. According to the properties of structure and attribute in node‐attributed graphs, we define a novel dense subgraph pattern, called hybridized k‐truss in attribute‐augmented graph. A hybridized k‐truss is a subgraph that consists of structural nodes and attribute nodes, of which there are at least (k − 2) common neighbors between any two connected nodes. We introduce the densest hybridized truss problem, and the densest hybridized truss mapping to a densely connected subgraph with homogenous attributes in the original graph. We propose a densest hybridized truss extraction (DHTE) algorithm for node‐attributed graphs, to automatically find the densest subgraph with high density and homogenous attributes at the same time. Extensive experimental results of 21 real world datasets demonstrate the effectiveness and efficiency of DHTE over state‐of‐the‐art methods, through comparison about structural cohesiveness and attributive homogeneity. Heli Sun, Xiaolin Jia, Ruodan Huang, Liang He 0006, Zhongbin Sun |
Comput. Intell. | 1 |
| 2021 | CMG2Vec: A composite meta-graph based heterogeneous information network embedding approach
Qinglin Tan, Heli Sun, Yu Zhou 0019 |
Knowl. Based Syst. | 4 |
| 2021 | Dynamic Community Evolution Analysis Framework for Large-Scale Complex Networks Based on Strong and Weak EventsabstractCommunity evolution remains a heavily researched and challenging area in the analysis of dynamic complex network structures. Currently, the primary limitation of traditional event-based approaches for community evolution analysis is the lack of strict constraint conditions for distinguishing evolutionary events, which entails that as the cardinality of discovered events increases, so does the number of redundant events. Another limitation of existing approaches is the lack of consideration for weak events. Weak events can be generated by small changes in communities, which are empirically prevalent, and are typically not captured by traditional events. To manage these two aforementioned limitations, this research aims to formalize a weak and strong events-based framework, which includes the following newly discovered events: “weak shrink,” “weak expand,” “weak merge,” and “weak splity” predicated on the community overlapping degree and community degree membership, this article refines these traditional strong events, as well as new constraints for weak events. In addition, a community evolution mining framework, which is based on both strong and weak events, is proposed and denoted by a weak-event-based community evolution method (WECEM). The framework can be summarized by the following: 1) communities in complex networks with adjacent time-stamps are compared to determine the community overlapping degree and community membership degree; 2) the values of the community overlapping degree and membership degree meet the definition of events; and 3) weak events are effectively identified. Extensive experimental results, on real and synthetic data sets consisting of dynamic complex networks and online social networks, demonstrate that WECEM is able to identify weak events more effectively than traditional frameworks. Specifically, WECEM outperforms traditional frameworks by 22.9% in the number of discovered strong events. The detection accuracy of evolutionary events is approximately 12.2% higher than that of traditional event-based frameworks. It is also worth noting that, as the cardinality of the data grows, the proposed framework, when compared with traditional frameworks, can more effectively, and efficiently, mine large-scale complex networks. Shaojie Qiao, Nan Han, Yunjun Gao, Rong-Hua Li 0001, Heli Sun, Xindong Wu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2020 | A personalized event-participant arrangement framework based on user interests in social network
Heli Sun, Liang He 0006, Zhangtian Duan |
Comput. Networks | 1 |
| 2020 | Leader-aware community detection in complex networks
Heli Sun, Hongxia Du, Zhongbin Sun, Liang He 0006, Xiaolin Jia, Zhongmeng Zhao |
Knowl. Inf. Syst. | 1 |
| 2020 | Community search for multiple nodes on attribute graphs
Heli Sun, Ruodan Huang, Xiaolin Jia, Liang He 0006, Miaomiao Sun, Zhongbin Sun |
Knowl. Based Syst. | 1 |
| 2020 | Network Embedding for Community Detection in Attributed NetworksabstractCommunity detection aims to partition network nodes into a set of clusters, such that nodes are more densely connected to each other within the same cluster than other clusters. For attributed networks, apart from the denseness requirement of topology structure, the attributes of nodes in the same community should also be homogeneous. Network embedding has been proved extremely useful in a variety of tasks, such as node classification, link prediction, and graph visualization, but few works dedicated to unsupervised embedding of node features specified for clustering task, which is vital for community detection and graph clustering. By post-processing with clustering algorithms like k -means, most existing network embedding methods can be applied to clustering tasks. However, the learned embeddings are not designed for clustering task, they only learn topological and attributed information of networks, and no clustering-oriented information is explored. In this article, we propose an algorithm named Network Embedding for node Clustering (NEC) to learn network embedding for node clustering in attributed graphs. Specifically, the presented work introduces a framework that simultaneously learns graph structure-based representations and clustering-oriented representations together. The framework consists of the following three modules: graph convolutional autoencoder module, soft modularity maximization module, and self-clustering module. Graph convolutional autoencoder module learns node embeddings based on topological structure and node attributes. We introduce soft modularity, which can be easily optimized using gradient descent algorithms, to exploit the community structure of networks. By integrating clustering loss and embedding loss, NEC can jointly optimize node cluster labels assignment and learn representations that keep local structure of network. This model can be effectively optimized using stochastic gradient algorithm. Empirical experiments on real-world networks and synthetic networks validate the feasibility and effectiveness of our algorithm on community detection task compared with network embedding based methods and traditional community detection methods. Heli Sun, Yizhou Sun, Liang He 0006, Zhongbin Sun, Xiaolin Jia |
ACM Trans. Knowl. Discov. Data | 1 |
| 2020 | An Efficient Destination Prediction Approach Based on Future Trajectory Prediction and Transition Matrix OptimizationabstractDestination prediction is an essential task in various mobile applications and up to now many methods have been proposed. However, existing methods usually suffer from the problems of heavy computational burden, data sparsity, and low coverage. Therefore, a novel approach named DestPD is proposed to tackle the aforementioned problems. Differing from an earlier approach that only considers the starting and current location of a partial trip, DestPD first determines the most likely future location and then predicts the destination. It comprises two phases, the offline training and the online prediction. During the offline training, transition probabilities between two locations are obtained via Markov transition matrix multiplication. In order to improve the efficiency of matrix multiplication, we propose two data constructs, Efficient Transition Probability (ETP) and Transition Probabilities with Detours (TPD). They are capable of pinpointing the minimum amount of needed computation. During the online prediction, we design Obligatory Update Point (OUP) and Transition Affected Area (TAA) to accelerate the frequent update of ETP and TPD for recomputing the transition probabilities. Moreover, a new future trajectory prediction approach is devised. It captures the most recent movement based on a query trajectory. It consists of two components: similarity finding through Best Path Notation (BPN) and best node selection. Our novel BPN similarity finding scheme keeps track of the nodes that induces inefficiency and then finds similarity fast based on these nodes. It is particularly suitable for trajectories with overlapping segments. Finally, the destination is predicted by combining transition probabilities and the most probable future location through Bayesian reasoning. The DestPD method is proved to achieve one order of cut in both time and space complexity. Furthermore, the experimental results on real-world and synthetic datasets have shown that DestPD consistently surpasses the state-of-the-art methods in terms of both efficiency (approximately over 100 times faster) and accuracy. Zhou Yang 0004, Heli Sun, Zhongbin Sun, Hui Xiong 0001, Shaojie Qiao, Ziyu Guan, Xiaolin Jia |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2019 | Central Station Based Demand Prediction in a Bike Sharing SystemabstractPredicting the bike demand can help rebalance the bikes and improve the service quality of a bike sharing system. A lot of work focuses on predicting the bike demand for all the stations. It is not necessary because the travel cost of rebalance operations increases sharply as the number of stations increases. In this paper, we take more attention to those stations with higher bike demand, which are called "central stations" in the following narrative. We propose a framework to predict the hourly bike demand based on the central stations we define. Firstly, we propose a novel clustering algorithm to assign different types of stations into each cluster. Secondly, we propose a hierarchical prediction model to predict the hourly bike demand for every cluster and each central station progressively. The experimental results on the NYC Citi Bike system show the advantages of our approach to these problems. Heli Sun |
MDM | 3 |
| 2019 | Parameter-free Community Detection through Distance Dynamic SynchronizationabstractAbstract Community detection plays a significant role in understanding the essence of a network. A recently proposed algorithm Attractor, which is based on distance dynamics, can spot communities effectively, but it depends on a cohesion parameter. Moreover, no efficient way is provided to find an optimal cohesion parameter setting. In this paper, we propose a parameter-free community detection algorithm by synchronizing distances iteratively. In each iteration, the distance of each edge will change dynamically according to the effect generated by its related neighbours. Several iterations later, distances between vertices belonging to the same community will synchronize to 0, while distances between vertices not in the same community will synchronize to 1. Besides, merging and division strategies are built up in the process of community detection. Experiments on both real-world and synthetic networks demonstrate benefits of our method compared to the baseline methods. Qingquan Bian, Heli Sun, Yaming Yang 0002, Yu Zhou 0019 |
Comput. J. | 3 |
| 2019 | Recurrent Meta-Structure for Robust Similarity Measure in Heterogeneous Information NetworksabstractSimilarity measure is one of the fundamental task in heterogeneous information network (HIN) analysis. It has been applied to many areas, such as product recommendation, clustering, and Web search. Most of the existing metrics can provide personalized services for users by taking a meta-path or meta-structure as input. However, these metrics may highly depend on the user-specified meta-path or meta-structure. In addition, users must know how to select an appropriate meta-path or meta-structure. In this article, we propose a novel similarity measure in HINs, called Recurrent Meta-Structure (RecurMS)-based Similarity (RMSS). The RecurMS as a schematic structure in HINs provides a unified framework for integrating all of the meta-paths and meta-structures, and can be constructed automatically by means of repetitively traversing the network schema. In order to formalize the semantics, the RecurMS is decomposed into several recurrent meta-paths and recurrent meta-trees, and we then define the commuting matrices of the recurrent meta-paths and meta-trees. All of these commuting matrices are combined together according to different weights. We propose two kinds of weighting strategies to determine the weights. The first is called the local weighting strategy that depends on the sparsity of the commuting matrices, and the second is called the global weighting strategy that depends on the strength of the commuting matrices. As a result, RMSS is defined by means of the weighted summation of the commuting matrices. Note that RMSS can also provide personalized services for users by means of the weights of the recurrent meta-paths and meta-trees. Experimental evaluations show that the proposed RMSS is robust and outperforms the existing metrics in terms of ranking and clustering task. Yu Zhou 0019, Heli Sun, Yizhou Sun, Shaojie Qiao, Stephen Manko Wambura |
ACM Trans. Knowl. Discov. Data | 3 |
| 2018 | Detecting semantic-based communities in node-attributed graphsabstractAbstract In social network analysis, community detection on plain graphs has been widely studied. With the proliferation of available data, each user in the network is usually associated with additional attributes for elaborate description. However, many existing methods only concentrate on the topological structure and fail to deal with node‐attributed networks. These approaches are incapable of extracting clear semantic meanings for communities detected. In this paper, we combine the topological structure and attribute information into a unified process and propose a novel algorithm to detect overlapping semantic communities. Moreover, a new metric is designed to measure the density of semantic communities. The proposed algorithm is divided into 3 phases. First, we detect local semantic subcommunities from each node's perspective using a greedy strategy on the metric. Then, a supergraph, which consists of all these subcommunities is created. Finally, we find global semantic communities on the supergraph. The experimental results on real‐world data sets show the efficiency and effectiveness of our approach against other state‐of‐the‐art methods. Heli Sun, Hongxia Du, Zhongbin Sun, Liang He 0006, Xiaolin Jia, Zhongmeng Zhao |
Comput. Intell. | 1 |
| 2018 | Corrigendum: A Team Formation Model with Personnel Work Hours and Project Workload Quantifiedabstractdoi:10.1093/comjnl/bxx009 An acknowledgements section was missing from the Advance Access section of this paper. It should read as follows: We would like to thank anonymous reviewers greatly for their valuable comments. The work was supported in part by the National Science Foundation of China grants 61472299, 61672417 and 61602354, the Fundamental Research Funds for the Central Universities of China grants BDY10, Shaanxi Postdoctoral Science Foundation, Natural Science Basic Research Plan in Shaanxi Province of China grants 2014JQ8359. Any opinions, findings and conclusions expressed here are those of the authors and do not necessarily reflect the views of the funding agencies. This has been corrected online and in print. The author apologises for this error. Xiaojing Sun, Yu Zhou 0019, Heli Sun |
Comput. J. | 4 |
| 2018 | A semantic-rich similarity measure in heterogeneous information networks
Yu Zhou 0019, He Li 0006, Heli Sun, Yueshen Xu |
Knowl. Based Syst. | 4 |
| 2017 | A Balanced Assignment Mechanism for Online Taxi RecommendationabstractMajority of taxi recommender systems mainly focused on satisfaction of passengers without considering fairness in assignment of taxi drivers. In this paper we propose a balanced assignment mechanism for online taxi recommendation (BAMOTR). BAMOTR provides a mechanism for fair assignment of drivers at some locations with specific routes to pick up passengers and ensures a short waiting time for passengers. Fair assignment is intended to minimize the differences in income among the taxi drivers. Analysis shows out that fair assignment of drivers and shortening the time the passenger wait before pick up is a trade-off problem. In this paper, we set a regulatory factor that can adjust the trade-off between fair assignment of drivers and shortening of waiting time of passengers. We also propose an efficient range refinement algorithm to solve online taxi recommendation problem in BAMOTR. It is theoretically and experimentally proved that range refinement algorithm ensures the same recommendation result as brute-force algorithm, however it greatly reduces the time overhead. We validate the performances of BAMOTR with extensive evaluations. Experimental results show that BAMOTR achieve better recommendation fairness than compared approaches and guarantee a short waiting time for passengers to be picked up. Guang Dai, Stephen Manko Wambura, Heli Sun |
MDM | 4 |
| 2017 | A Centrality-Based Local-First Approach for Analyzing Overlapping Communities in Dynamic Networks
Ximan Chen, Heli Sun, Hongxia Du |
PAKDD (2) | 2 |
| 2017 | Mining Cohesive Clusters with Interpretations in Labeled Graphs
Hongxia Du, Heli Sun, Zhongbin Sun, Liang He 0006, Hong Cheng 0001 |
PAKDD (2) | 2 |
| 2017 | Predicting Bugs in Software Code Changes Using Isolation ForestabstractIdentifying bug immediately when it is introduced can help improve the validity and effectiveness of bug fixing. Predicting bugs in software code changes makes such identification possible. Buggy changes, changes that introduce bugs into source code, can be viewed as anomalies relative to clean changes for that they are rare and irregular. Thus, anomaly detection techniques can be applied to buggy change prediction. Isolation Forest, which detects anomalies based on the hypothesis that the anomalies have the shortest average path length on the constructed random forest, has exhibited its good performance on anomaly detection compared to other anomaly detection methods. In this paper, we adopt it in predicting bugs in software code changes. Empirical study with eight practical open source projects are conducted to validate the effective of Isolation Forest in bug prediction in software code changes. Results of the empirical study show that compared to traditional classification methods used in literature, Isolation Forest can achieve better clean precision, buggy recall, buggy F-measure, AUC and Gmean. Yueyang He, Guangtao Wang, Heli Sun, Yong Wang 0076 |
QRS | 4 |
| 2017 | LinkLPA: A Link-Based Label Propagation Algorithm for Overlapping Community Detection in NetworksabstractCommunity detection is an important methodology for understanding the intrinsic structure and function of complex networks. Because overlapping community is one of the characteristics of real‐world networks and should be considered for community detection, in this article, we propose an algorithm, called link‐based label propagation algorithm (LinkLPA), to detect overlapping communities. Because the link partition is conceptually natural for the problem of overlapping community detection, LinkLPA first transforms node partition problem into link partition problem and employs a new label propagation algorithm with preference on links instead of nodes to detect communities due to the simplicity and efficiency of label propagation algorithm. Then the proposed LinkLPA performs a postprocessing to refine the detected overlapping communities by avoiding over‐overlapping and incorrect partition of weak ties. Experimental results on a large number of real‐world and synthetic networks show that the proposed method achieves high accuracy on detecting overlapping communities in networks. Heli Sun, Guangtao Wang, Xiaolin Jia, Qinbao Song |
Comput. Intell. | 1 |
| 2017 | Forming Grouped Teams with Efficient Collaboration in Social NetworksabstractNot only the expertise of people but also the collaboration among people are of great importance for a team. Given a set of experts with different skills, a social network that reflects the collaboration among people and a task, the team formation problem in social networks aims at forming a team to complete the task. The team is required to satisfy the skill requirements of the task and collaborates efficiently. Different communication cost functions have been proposed to have a good measure on the collaboration strength of a team in the existing work. However, the grouped organization structure inside team is never considered, which is very common in real life scenarios. In a grouped team, we are more concerned with the collaboration among people in same group and among leaders. In this paper, a novel communication cost function for a grouped team is proposed, and we define the Grouped Team Formation problem. To solve the problem, an exact algorithm is proposed. We further modify the exact algorithm to propose two heuristic algorithms with higher efficiency. Extensive experiments evaluate the effectiveness and efficiency of the proposed methods, and validate the reasonability of our problem definition in practical settings. Ze Lv, Yu Zhou 0019, He Li 0006, Heli Sun, Xiaolin Jia |
Comput. J. | 5 |
| 2017 | A Team Formation Model with Personnel Work Hours and Project Workload QuantifiedabstractTeam formation is a problem of gathering a group of experts with complementary skills to complete a given task in a cooperative way. This mode secures a flexible team for a project, and in the meantime, it allows individual experts to seek a position to give full play in their expertise. This work considers a setting where each expert has an available work time and a set of skills, each of which is associated with a skill level indicating his competence in this skill. We are further presented with a specific project requiring a set of skills and respective demands on skill level and work amount. Then, we study a problem of building a team for the project from a given candidate set so that the project is accomplished in both quality and quantity and the overall effect team score is maximized. We refer to this as Quality Team Formation problem. The problem is proven to be NP-hard, and two approximation algorithms are provided with demonstration of their effectiveness and practicability through experiments on real-world data sets. Xiaojing Sun, Yu Zhou 0019, Heli Sun |
Comput. J. | 4 |
| 2017 | A Novel Social Event Organization Approach for Diverse User ChoicesabstractThe boom of social networking services makes it convenient to organize or participate in social events. Recent studies consider the social event organization problem, but they cannot provide diverse choices for users when there are multiple events. In this paper, we seek to devise an event organization scheme that provides diverse choices for users. For this goal, we explicitly distinguish the subjective preferences of users to events and the objective preferences of users to events. The former is considered to be the generative probabilities of users appearing in events, and the latter is viewed as the posterior probabilities of users participating in events. The Expectation–Maximization algorithm is employed to connect them together. After extracting these features, we extract an edge-weighted bipartite graph as the scheme from all the objective preferences. It captures the maximum sum of the objective preferences, and simultaneously satisfies soft user capacity constraints and hard event capacity constraints. We prove this problem is in P via transforming it into linear program (LP). In consideration that LP may have no feasible solutions, we alternatively devise an efficient polynomial time algorithm which can yield an approximately feasible solution. Experimental evaluation shows the effectiveness and efficiency of our proposed method. Yu Zhou 0019, Xiaolin Jia, Heli Sun |
Comput. J. | 4 |
| 2017 | On Participation Constrained Team Formation
Yu Zhou 0019, Xiaolin Jia, Heli Sun |
J. Comput. Sci. Technol. | 4 |
| 2016 | Grouped Team Formation in Social Networks
Ze Lv, Yu Zhou 0019, Heli Sun, Xiaolin Jia |
APWeb (2) | 4 |
| 2016 | Profit Maximizing Route Recommendation for Vehicle Sharing Requests
Hua Gao, Heli Sun, Xiaolin Jia |
APWeb (2) | 4 |
| 2016 | A machine learning based software process model recommendation method
Qinbao Song, Xiaoyan Zhu 0003, Guangtao Wang, Heli Sun, Chenhao Xue, Baowen Xu |
J. Syst. Softw. | 4 |
| 2016 | Efficient k-edge connected component detection through an early merging and splitting strategy
Heli Sun, Zhongmeng Zhao, Xiaolin Jia |
Knowl. Based Syst. | 1 |
| 2015 | Label propagation based evolutionary clustering for detecting overlapping and non-overlapping communities in dynamic networks
Heli Sun, Mengjie Wan, Yutao Qi, He Li 0006 |
Knowl. Based Syst. | 3 |
| 2015 | A novel ensemble method for classifying imbalanced data
Zhongbin Sun, Qinbao Song, Xiaoyan Zhu 0003, Heli Sun, Baowen Xu, Yuming Zhou |
Pattern Recognit. | 4 |
| 2015 | Backward Path Growth for Efficient Mobile Sequential RecommendationabstractThe problem of mobile sequential recommendation is to suggest a route connecting a set of pick-up points for a taxi driver so that he/she is more likely to get passengers with less travel cost. Essentially, a key challenge of this problem is its high computational complexity. In this paper, we propose a novel dynamic programming based method to solve the mobile sequential recommendation problem consisting of two separate stages: an offline pre-processing stage and an online search stage. The offline stage pre-computes potential candidate sequences from a set of pick-up points. A backward incremental sequence generation algorithm is proposed based on the identified iterative property of the cost function. Simultaneously, an incremental pruning policy is adopted in the process of sequence generation to reduce the search space of the potential sequences effectively. In addition, a batch pruning algorithm is further applied to the generated potential sequences to remove some non-optimal sequences of a given length. Since the pruning effectiveness keeps growing with the increase of the sequence length, at the online stage, our method can efficiently find the optimal driving route for an unloaded taxi in the remaining candidate sequences. Moreover, our method can handle the problem of optimal route search with a maximum cruising distance or a destination constraint. Experimental results on real and synthetic data sets show that both the pruning ability and the efficiency of our method surpass the state-of-the-art methods. Our techniques can therefore be effectively employed to address the problem of mobile sequential recommendation with many pick-up points in real-world applications. Xuejun Huangfu, Heli Sun, Hui Li 0005, Peixiang Zhao 0001, Hong Cheng 0001, Qinbao Song |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2014 | IncOrder: Incremental density-based community detection in dynamic networks
Heli Sun, Huailiang Liu, Jianhua Zou, Qinbao Song |
Knowl. Based Syst. | 1 |
| 2013 | A Feature Subset Selection Algorithm Automatic Recommendation MethodabstractMany feature subset selection (FSS) algorithms have been proposed, but not all of them are appropriate for a given feature selection problem. At the same time, so far there is rarely a good way to choose appropriate FSS algorithms for the problem at hand. Thus, FSS algorithm automatic recommendation is very important and practically useful. In this paper, a meta learning based FSS algorithm automatic recommendation method is presented. The proposed method first identifies the data sets that are most similar to the one at hand by the k-nearest neighbor classification algorithm, and the distances among these data sets are calculated based on the commonly-used data set characteristics. Then, it ranks all the candidate FSS algorithms according to their performance on these similar data sets, and chooses the algorithms with best performance as the appropriate ones. The performance of the candidate FSS algorithms is evaluated by a multi-criteria metric that takes into account not only the classification accuracy over the selected features, but also the runtime of feature selection and the number of selected features. The proposed recommendation method is extensively tested on 115 real world data sets with 22 well-known and frequently-used different FSS algorithms for five representative classifiers. The results show the effectiveness of our proposed FSS algorithm recommendation method. Guangtao Wang, Qinbao Song, Heli Sun, Baowen Xu, Yuming Zhou |
J. Artif. Intell. Res. | 3 |
| 2013 | ESC: An efficient synchronization-based clustering algorithm
Heli Sun, Jianmei Kang, Junjie Qi, Hongbo Deng, Qinbao Song |
Knowl. Based Syst. | 2 |
| 2013 | Revealing Density-Based Clustering Structure from the Core-Connected Tree of a NetworkabstractClustering is an important technique for mining the intrinsic community structures in networks. The density-based network clustering method is able to not only detect communities of arbitrary size and shape, but also identify hubs and outliers. However, it requires manual parameter specification to define clusters, and is sensitive to the parameter of density threshold which is difficult to determine. Furthermore, many real-world networks exhibit a hierarchical structure with communities embedded within other communities. Therefore, the clustering result of a global parameter setting cannot always describe the intrinsic clustering structure accurately. In this paper, we introduce a novel density-based network clustering method, called graph-skeleton-based clustering (gSkeletonClu). By projecting an undirected network to its core-connected maximal spanning tree, the clustering problem can be converted to detect core connectivity components on the tree. The density-based clustering of a specific parameter setting and the hierarchical clustering structure both can be efficiently extracted from the tree. Moreover, it provides a convenient way to automatically select the parameter and to achieve the meaningful cluster tree in a network. Extensive experiments on both real-world and synthetic networks demonstrate the superior performance of gSkeletonClu for effective and efficient density-based clustering. Heli Sun, Qinbao Song, Hongbo Deng, Jiawei Han 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | QoRank: A query-dependent ranking model using LSE-based weighted multiple hyperplanes aggregation for information retrievalabstractRanking is a core problem for information retrieval since the performance of the search system is directly impacted by the accuracy of ranking results. Ranking model construction has been the focus of both the fields of information retrieval and machine learning, and learning to rank in particular has attracted much interest. Many ranking models have been proposed, for example, RankSVM is a state-of-the-art method for learning to rank and has been empirically demonstrated to be effective. However, most of the proposed methods do not consider about the significant differences between queries, only resort to a single function in ranking. In this paper, we present a novel ranking model named QoRank, which performs the learning task dependent on queries. We also propose a LSE (least-squares estimation) -based weighted method to aggregate the ranking lists produced by base decision functions as the final ranking. Comparison of QoRank with other ranking techniques is conducted, and several evaluation criteria are employed to evaluate its performance. Experimental results on the LETOR OHSUMED data set show that QoRank strikes a good balance of accuracy and complexity, and outperforms the baseline methods. © 2010 Wiley Periodicals, Inc. Heli Sun, Boqin Feng |
Int. J. Intell. Syst. | 1 |
| 2010 | SHRINK: a structural clustering algorithm for detecting hierarchical communities in networksabstractCommunity detection is an important task for mining the structure and function of complex networks. Generally, there are several different kinds of nodes in a network which are cluster nodes densely connected within communities, as well as some special nodes like hubs bridging multiple communities and outliers marginally connected with a community. In addition, it has been shown that there is a hierarchical structure in complex networks with communities embedded within other communities. Therefore, a good algorithm is desirable to be able to not only detect hierarchical communities, but also identify hubs and outliers. In this paper, we propose a parameter-free hierarchical network clustering algorithm SHRINK by combining the advantages of density-based clustering and modularity optimization methods. Based on the structural connectivity information, the proposed algorithm can effectively reveal the embedded hierarchical community structure with multiresolution in large-scale weighted undirected networks, and identify hubs and outliers as well. Moreover, it overcomes the sensitive threshold problem of density-based clustering algorithms and the resolution limit possessed by other modularity-based methods. To illustrate our methodology, we conduct experiments with both real-world and synthetic datasets for community detection, and compare with many other baseline methods. Experimental results demonstrate that SHRINK achieves the best performance with consistent improvements. Heli Sun, Jiawei Han 0001, Hongbo Deng, Yizhou Sun, Yaguang Liu |
CIKM | 2 |
| 2010 | gSkeletonClu: Density-Based Network Clustering via Structure-Connected Tree Division or AgglomerationabstractCommunity detection is an important task for mining the structure and function of complex networks. Many pervious approaches are difficult to detect communities with arbitrary size and shape, and are unable to identify hubs and outliers. A recently proposed network clustering algorithm, SCAN, is effective and can overcome this difficulty. However, it depends on a sensitive parameter: minimum similarity threshold ε, but provides no automated way to find it. In this paper, we propose a novel density-based network clustering algorithm, called gSkeletonClu (graph-skeleton based clustering). By projecting a network to its Core-Connected Maximal Spanning Tree (CCMST), the network clustering problem is converted to finding core-connected components in the CCMST. We discover that all possible values of the parameter ε lie in the edge weights of the corresponding CCMST. By means of tree divisive or agglomerative clustering, our algorithm can find the optimal parameter ε and detect communities, hubs and outliers in large-scale undirected networks automatically without any user interaction. Extensive experiments on both real-world and synthetic networks demonstrate the superior performance of gSkeletonClu over the baseline methods. Heli Sun, Jiawei Han 0001, Hongbo Deng, Peixiang Zhao 0001, Boqin Feng |
ICDM | 1 |
| 2009 | OrdRank: Learning to Rank with Ordered Multiple HyperplanesabstractRanking is a central problem for information retrieval systems, because the performance of an information retrieval system is mainly evaluated by the effectiveness of its ranking results. Learning to rank has received much attention in recent years due to its importance in information retrieval. This paper focuses on learning to rank in document retrieval and presents a ranking model named OrdRank that ranks documents with ordered multiple hyperplanes. Comparison of OrdRank with other state-of-the-art ranking techniques is conducted and several evaluation criteria are employed to evaluate its performance. Experimental results on the OHSUMED dataset show that OrdRank outperforms other methods, both in terms of quality of ranking results and efficiency. Heli Sun, Boqin Feng, Yingliang Zhao, Jun Liu 0002 |
Web Intelligence | 1 |