Hong Shen 0001

dblp:74/3247-1 · DBLP profile ↗
← Back
36ranked-venue papers in the field
1as first author
3since 2021 · last 2025
0000-0002-3663-6591ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 19Other / Interdisciplinary · 7 (1 first)Information Retrieval & Web Search · 6Database Systems & Data Management · 2Knowledge Engineering, Semantic Web & Information Systems · 2
YearPublicationVenuePosition
2025 Cross-Modal Sequential Point-of-Interest Recommendation with Lightweight Hybrid Fusion Strategy
Tianxing Wang 0004, Can Wang 0004, Hui Tian 0001, Hong Shen 0001
DaWaK4
2025 Local-Aware Convolutional Modulation for Short-Term Sequential Recommendation
Tianxing Wang 0004, Can Wang 0004, Hui Tian 0001, Hong Shen 0001
DaWaK4
2023 GeoMixer: The MLP-Based Sequential POI Recommender with Travel Routing Modelling
abstract
Nowadays, with the rise of location-based services, the personalized sequential POI recommendation has become a pivotal element for enhancing customer experiences. Although many previous POI recommendation models have shown promising results and improvements in this area, several challenges still exist in this field. Firstly, the previous sequential recommenders do not well-utilize the geographical features that are highly affecting the user’s future choices of visits. Furthermore, the self-attention mechanism, which is a popular method used in sequential POI recommendation, has a limitation in treating the input user sequence as an unordered set. Using positional embedding is a typical way to overcome this limitation. However, the use of such embeddings may potentially restrict the model’s ability to learn meaningful patterns in user preferences among POIs. To address these challenges, we propose GeoMixer, a novel MLP-based sequential POI recommender that incorporates travel routing distance to capture geographical features and leverages Multi-layer Perceptron (MLP) architecture to model the spatial and sequential patterns in the sequential POI recommendations. By adopting MLP mixing layers, GeoMixer has the capability of memorizing the chronological order of the input POIs without the positional embedding and can emphasize the important latent features of each POI. The use of the travel routing information improves the model’s ability of capturing spatial patterns during the model learning process. Extensive experiments on real-world datasets show that GeoMixer outperforms state-of-theart methods in various metrics, highlighting the significance of incorporating travel routing distance and leveraging MLP architecture in sequential POI recommendation systems.
Tianxing Wang 0004, Can Wang 0004, Hui Tian 0001, Hong Shen 0001
ICDM4
2019 A Convergent Differentially Private k-Means Clustering Algorithm
Zhigang Lu 0001, Hong Shen 0001
PAKDD (1)2
2019 Neural Variational Matrix Factorization with Side Information for Collaborative Filtering
Teng Xiao, Hong Shen 0001
PAKDD (1)2
2019 Variational Deep Collaborative Matrix Factorization for Social Recommendation
Teng Xiao, Hui Tian 0001, Hong Shen 0001
PAKDD (1)3
2019 Item diversified recommendation based on influence diffusion
Huimin Huang 0001, Hong Shen 0001, Zaiqiao Meng
Inf. Process. Manag.2
2019 Fast top-k similarity search in large dynamic attributed networks
Zaiqiao Meng, Hong Shen 0001
Inf. Process. Manag.2
2018 Search result diversification on attributed networks via nonnegative matrix factorization
Zaiqiao Meng, Hong Shen 0001, Huimin Huang 0001, Wei Liu 0061, Jing Wang 0030, Arun Kumar Sangaiah
Inf. Process. Manag.2
2017 Secured Privacy Preserving Data Aggregation with Semi-honest Servers
Zhigang Lu 0001, Hong Shen 0001
PAKDD (2)2
2017 Weighted Ensemble Classification of Multi-label Data Streams
Lulu Wang 0008, Hong Shen 0001, Hui Tian 0001
PAKDD (2)2
2017 User clustering in a dynamic social network topic model for short text streams
Zhangcheng Qiu, Hong Shen 0001
Inf. Sci.2
2017 Search Result Diversification in Short Text Streams
abstract
We consider the problem of search result diversification for streams of short texts. Diversifying search results in short text streams is more challenging than in the case of long documents, as it is difficult to capture the latent topics of short documents. To capture the changes of topics and the probabilities of documents for a given query at a specific time in a short text stream, we propose a dynamic Dirichlet multinomial mixture topic model, called D2M3, as well as a Gibbs sampling algorithm for the inference. We also propose a streaming diversification algorithm, SDA, that integrates the information captured by D2M3 with our proposed modified version of the PM-2 (Proportionality-based diversification Method -- second version) diversification algorithm. We conduct experiments on a Twitter dataset and find that SDA statistically significantly outperforms state-of-the-art non-streaming retrieval methods, plain streaming retrieval methods, as well as streaming diversification methods that use other dynamic topic models.
Shangsong Liang, Emine Yilmaz, Hong Shen 0001, Maarten de Rijke, W. Bruce Croft
ACM Trans. Inf. Syst.3
2016 Community Detection in Networks with Less Significant Community Structure
Ba Dung Le, Hung X. Nguyen, Hong Shen 0001
ADMA3
2015 A Security-assured Accuracy-maximised Privacy Preserving Collaborative Filtering Recommendation Algorithm
abstract
The neighbourhood-based Collaborative Filtering is a widely used method in recommender systems. However, the risks of revealing customers' privacy during the process of filtering have attracted noticeable public concern recently. Specifically, kNN attack discloses the target user's sensitive information by creating k fake nearest neighbours by non-sensitive information. Among the current solutions against kNN attack, the probabilistic methods showed a powerful privacy preserving effect. However, the existing probabilistic methods neither guarantee enough prediction accuracy due to the global randomness, nor provide assured security enforcement against kNN attack. To overcome the problems of current probabilistic methods, we propose a novel approach, Probabilistic Partitioned Neighbour Selection, to ensure a required security guarantee while achieving the optimal prediction accuracy against kNN attack. In this paper, we define the sum of k neighbours' similarity as the accuracy metric α, the number of user partitions, across which we select the k neighbours, as the security metric β. Differing from the present methods that globally selected neighbours, our method selects neighbours from each group with exponential differential privacy to decrease the magnitude of noise. Theoretical and experimental analysis show that to achieve the same security guarantee against kNN attack, our approach ensures the optimal prediction accuracy.
Zhigang Lu 0001, Hong Shen 0001
IDEAS2
2014 Two-Phase Layered Learning Recommendation via Category Structure
Ke Ji, Hong Shen 0001, Hui Tian 0001, Yanbo Wu, Jun Wu 0007
PAKDD (2)2
2014 A New Evaluation Function for Entropy-Based Feature Selection from Incomplete Data
Wenhao Shu, Hong Shen 0001, Yingpeng Sang, Yidong Li, Jun Wu 0007
PAKDD (2)2
2014 A Selectively Re-train Approach Based on Clustering to Classify Concept-Drifting Data Streams with Skewed Distribution
Hong Shen 0001, Hui Tian 0001, Yidong Li, Jun Wu 0007, Yingpeng Sang
PAKDD (2)2
2013 A Self-immunizing Manifold Ranking for Image Retrieval
Jun Wu 0007, Yidong Li, Songhe Feng, Hong Shen 0001
PAKDD (2)4
2012 Towards Identity Disclosure Control in Private Hypergraph Publishing
Yidong Li, Hong Shen 0001
PAKDD (2)2
2010 Multivariate Equi-width Data Swapping for Private Data Publication
Yidong Li, Hong Shen 0001
PAKDD (1)2
2009 Reconstructing Data Perturbed by Random Projections When the Mixing Matrix Is Known
Yingpeng Sang, Hong Shen 0001, Hui Tian 0001
ECML/PKDD (2)2
2009 Privacy-Preserving Tuple Matching in Distributed Databases
abstract
We address the problems of privacy-preserving duplicate tuple matching (PPDTM) and privacy-preserving threshold attributes matching (PPTAM) in the scenario of a horizontally partitioned database among N parties, where each party holds a private share of the database's tuples and all tuples have the same set of attributes. In PPDTM, each party determines whether its tuples have any duplicate on other parties' private databases. In PPTAM, each party determines whether all attribute values of each tuple appear at least a threshold number of times in the attribute unions. We propose protocols for the two problems using additive homomorphic cryptosystem based on the subgroup membership assumption, e.g., Paillier's and ElGamal's schemes. By analysis on the total numbers of modular exponentiations, modular multiplications and communication bits, with a reduced computation cost which dominates the total cost, by trading off communication cost, our PPDTM protocol for the semihonest model is superior to the solution derivable from existing techniques in total cost. Our PPTAM protocol is superior in both computation and communication costs. The efficiency improvements are achieved mainly by using random numbers instead of random polynomials as existing techniques for perturbation, without causing successful attacks by polynomial interpolations. We also give detailed constructions on the required zero-knowledge proofs and extend our two protocols to the malicious model, which were previously unknown.
Yingpeng Sang, Hong Shen 0001, Hui Tian 0001
IEEE Trans. Knowl. Data Eng.2
2006 The Probability of Success of Mobile Agents When Routing in Faulty Networks
Wenyu Qu, Hong Shen 0001
APWeb2
2005 Cache Replacement for Transcoding Proxy Caching
abstract
In this paper, we address the problem of cache replacement for transcoding proxy caching. First, an efficient cache replacement algorithm is proposed. Our algorithm considers both the aggregate effect of caching multiple versions of the same multimedia object and cache consistency. Second, a complexity analysis is presented to show the efficiency of our algorithm. Finally, some preliminary simulation experiments are conducted to compare the performance of our algorithm with some existing algorithms. The results show that our algorithm outperforms others in terms of the various performance metrics.
Keqiu Li, Keishi Tajima, Hong Shen 0001
Web Intelligence3
2004 Coordinated En-Route Web Caching in Transcoding Proxies
Keqiu Li, Hong Shen 0001
APWeb2
2004 An Improved GreedyDual Cache Document Replacement Algorithm
abstract
Web caching is an important technique for reducing web traffic, user access latency, and server load and cache replacement plays an important role in the functionality of web caching. In this paper we propose an improved GreedyDual (GD) cache document replacement algorithm, which considers update frequency as a factor in its utility function. We use both trace data and statistical data to simulate our proposed algorithm. The experimental results show that our improved GD algorithm can outperform the existing GD algorithm over the performance metrics considered.
Keqiu Li, Hong Shen 0001
Web Intelligence2
2004 Mining Informative Rule Set for Prediction
Jiuyong Li, Hong Shen 0001, Rodney W. Topor
J. Intell. Inf. Syst.2
2003 Transversal of disjoint convex polygons
Francis Y. L. Chin, Hong Shen 0001, Fu Lee Wang
Inf. Process. Lett.2
2002 Construct robust rule sets for classification
abstract
We study the problem of computing classification rule sets from relational databases so that accurate predictions can be made on test data with missing attribute values. Traditional classifiers perform badly when test data are not as complete as the training data because they tailor a training database too much. We introduce the concept of one rule set being more robust than another, that is, able to make more accurate predictions on test data with missing attribute values. We show that the optimal class association rule set is as robust as the complete class association rule set. We then introduce the k-optimal rule set, which provides predictions exactly the same as the optimal class association rule set on test data with up to k missing attribute values. This leads to a hierarchy of k-optimal rule sets in which decreasing size corresponds to decreasing robustness, and they all more robust than a traditional classification rule set. We introduce two methods to find k-optimal rule sets, i.e. an optimal association rule mining approach and a heuristic approximate approach. We show experimentally that a k-optimal rule set generated by the optimal association rule mining approach performs better than that by the heuristic approximate approach and both rule sets perform significantly better than a typical classification rule set (C4.5Rules) on incomplete test data.
Jiuyong Li, Rodney W. Topor, Hong Shen 0001
KDD3
2001 Mining the Smallest Association Rule Set for Predictions
abstract
Mining transaction databases for association rules usually generates a large number of rules, most of which are unnecessary when used for subsequent prediction. In this paper we define a rule set for a given transaction database that is much smaller than the association rule set but makes the same predictions as the association rule set by the confidence priority. We call this subset the informative rule set. The informative rule set is not constrained to particular target items; and it is smaller than the non-redundant association rule set. We present an algorithm to directly generate the informative rule set, i.e., without generating all frequent itemsets first, and that accesses the database less often than other unconstrained direct methods. We show experimentally that the informative rule set is much smaller than both the association rule set and the non-redundant association rule set, and that it can be generated more efficiently.
Jiuyong Li, Hong Shen 0001, Rodney W. Topor
ICDM2
2001 Mining Optimal Class Association Rule Set
Jiuyong Li, Hong Shen 0001, Rodney W. Topor
PAKDD2
1996 Optimal Parallel Selection in Sorted Matrices
Hong Shen 0001, Sarnath Ramnath
Inf. Process. Lett.1
1996 NC Algorithms for the Single Most Vital Edge Problem with Respect to Shortest Path
Sven Venema, Hong Shen 0001, Francis Suraweera
Inf. Process. Lett.2
1994 An Efficient Permutation-Based Parallel Range-Join Algorithm on N-Dimensional Torus Computers
Shao Dong Chen, Hong Shen 0001, Rodney W. Topor
Inf. Process. Lett.2
1990 Improved Nonconservative Sequential and Parallel Integer Sorting
Torben Hagerup, Hong Shen 0001
Inf. Process. Lett.2