VLDB 2026 Research / reviewers in the wild / expert
Junheng Huang
dblp:93/8349
· DBLP profile ↗
11ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A novel model on improving Chinese dialogue summarization with multi-perspective information enhancement
Kaikun Dong, Zongwei Du, Junheng Huang, Hongri Liu, Bailing Wang |
Neural Networks | 4 |
| 2024 | Co-Training-Teaching: A Robust Semi-Supervised Framework for Review-Aware Rating RegressionabstractReview-aware Rating Regression (RaRR) suffers the severe challenge of extreme data sparsity as the multi-modality interactions of ratings accompanied by reviews are costly to obtain. Although some studies of semi-supervised rating regression are proposed to mitigate the impact of sparse data, they bear the risk of learning from noisy pseudo-labeled data. In this article, we propose a simple yet effective paradigm, called co-training-teaching ( CoT 2 ), for integrating the merits of both co-training and co-teaching toward robust semi-supervised RaRR. CoT 2 employs two predictors trained with different feature sets of textual reviews, each of which functions as both “labeler” and “validator.” Specifically, one predictor (labeler) first labels unlabeled data for its peer predictor (validator); after that, the validator samples reliable instances from the noisy pseudo-labeled data it received and sends them back to the labeler for updating. By exchanging and validating pseudo-labeled instances, the two predictors are reinforced by each other in an iterative learning process. The final prediction is made by averaging the outputs of both the refined predictors. Extensive experiments show that our CoT 2 considerably outperforms the state-of-the-art recommendation techniques in the RaRR task, especially when the training data is severely insufficient. Xiangkui Lu, Jun Wu 0007, Junheng Huang, Fangyuan Luo |
ACM Trans. Knowl. Discov. Data | 3 |
| 2022 | Mining trading patterns of pyramid schemes from financial time series data
Linxuan Han, Yulong Pei, Junheng Huang, Bailing Wang, Mykola Pechenizkiy |
Future Gener. Comput. Syst. | 6 |
| 2021 | Semi-supervised Factorization Machines for Review-Aware Recommendation
Junheng Huang, Fangyuan Luo, Jun Wu 0007 |
DASFAA (3) | 1 |
| 2021 | A network representation learning method based on topology
Dongyang Ma, Guodong Xin, Yunpeng Han, Junheng Huang, Bailing Wang |
Inf. Sci. | 5 |
| 2018 | Detecting bursts in sentiment-aware topics from social media
Kang Xu 0001, Guilin Qi, Junheng Huang, Tianxing Wu 0001, Xuefeng Fu |
Knowl. Based Syst. | 3 |
| 2017 | Incorporating Wikipedia concepts and categories as prior knowledge into topic modelsabstractTopic models have been widely applied in discovering topics that underly a collection of documents. Incorporating human knowledge can guide conventional topic models to produce topics which are easily interpreted and semantically coherent. Several knowledge-based topic models have been proposed, bu t these models just leverage lexical knowledge of words that are often not in accordance with topics. To solve the problem, we recognize entity mentions, besides words, in the documents and incorporate entity knowledge from external knowledge bases. In this paper, we study to utilize entity knowledge, concepts and categories in Wikipedia, as prior knowledge into topic models to discover more coherent topics. A novel knowledge-based topic model, WCM-LDA (Wikipedia-Category-concept-Mention Latent Dirichlet Allocation), is proposed, which not only models the relationship between words and topics, but also utilizes concept and category knowledge of entities to model the semantic relation of entities and topics. We compare WCM-LDA with the state-of-the-art knowledge-based topic models, on three datasets. Experimental results show that our approach outperforms the existing baseline methods on all three datasets. Moreover, our model can visualize topics with top words, concepts and categories such that topics are made easily to be interpreted and classified. Kang Xu 0001, Guilin Qi, Junheng Huang, Tianxing Wu 0001 |
Intell. Data Anal. | 3 |
| 2016 | A Joint Model for Sentiment-Aware Topic Detection on Social MediaabstractJoint sentiment/topic models are widely applied in detecting sentiment-aware topics on the lengthy review data and they are achieved with Latent Dirichlet Allocation (LDA) based model. Nowadays plenty of user-generated posts, e.g., tweets and E-commerce short reviews, are published on the social media and the posts imply the public's sentiments (i.e., positive and negative) towards various topics. However, the existing sentiment/topic models are not applicable to detect sentiment-aware topics on the posts, i.e., short texts, because applying the models to the short texts directly will suffer from the context sparsity problem. In this paper, we propose a Time-User Sentiment/Topic Latent Dirichlet Allocation (TUS-LDA) which aggregates posts in the same timeslice or user as a pseudo-document to alleviate the context sparsity problem. Moreover, we design approaches for parameter inference and incorporating prior knowledge into TUS-LDA. Experiments on the Sentiment140 and tweets of electronic products from Twitter7 show that TUS-LDA outperforms previous models in the tasks of sentiment classification and sentiment-aware topic extraction. Finally, we visualize the sentiment-aware topics discovered by TUS-LDA. Kang Xu 0001, Guilin Qi, Junheng Huang, Tianxing Wu 0001 |
ECAI | 3 |
| 2015 | A proactive discovery and filtering solution on phishing websitesabstractPhishing website is becoming a major threat to the information security in Social Network. The attacks not only lessen the users' trust but also influence the benefit of the third party who develops the platform. In order to solve the time lag in phishing website passive detection, this paper proposes a solution to discover phishing website initiatively based on blacklist, in which the anomalies of its URL and WHOIS information are analyzed, and based on this, the heuristic rules that aim to generate suspicious URLs are made. In order to filter out noise sites in the suspicious set, a website filtering solution based on webpages image-layout is presented. We firstly propose a Ray Scan Method to generate the location feature of webpage images quickly, and then, we proposed a method of calculating the webpage layout similarity, which will be compared against the preset threshold to decide whether it will be filtered. The experimental results show that the solution successfully detects some phishing websites out before they are widely spread, and further, the webpage filtering method guarantees both high filtration ratio and high phishing website retention ratio. Bailing Wang, Junheng Huang, Yushan Sun, Yuliang Wei |
IEEE BigData | 3 |
| 2015 | A collaborative filtering algorithm fusing user-based, item-based and social networksabstractThe traditional collaborative filtering recommendation algorithm can be divided into the user-based and the item-based two methods, which only uses the information in the rating matrix. Because of the limitation of the information capacity they used, it is difficult to further improve the accuracy of the recommendation, and cold start problem also affects the normal operation of the recommendation system. This paper presented a collaborative filtering recommendation algorithm (UISA) fusing user-based, item-based and social networks data. The algorithm uses the data of the neighbor relations in social networks, calculating the users' friends not reflected in the rating matrix. At the same time, we can calculate the similarity between items by using the data of item text in social networks, mining similar items not reflected in the rating matrix. In this way, it can fundamentally expand available information capacity of the traditional filtering collaboration recommendation algorithms, improve the recommendation accuracy, alleviate cold start problem. Experimental results based on KDD CUP 2012 real data show that compared with the traditional collaborative filtering system, this system has obvious advantages in the recommendation accuracy and ease of cold start. Bailing Wang, Junheng Huang, Libing Ou |
IEEE BigData | 2 |
| 2015 | A collaborative filtering algorithm based on social network informationabstractIn traditional collaborative filtering recommendation, the matrix sparsity and cold start restricted the accuracy of system. In this paper, we develop a way to enhance the recommendation effectiveness by merging neighborhood relationship and users keyword of social network information into collaborative filtering. We extend the calculation method of the TOP N neighbors which is the most important from two aspects. Our method expands the information capacity which can be used by collaborative filtering, improves the accuracy of recommendation and eases the cold start problem in recommendation system. We conducts experiment based on KDD 2012 real data set. The result indicates that our algorithm performs more superior than traditional collaborative filtering algorithm. Bailing Wang, Junheng Huang |
IEEE BigData | 3 |