VLDB 2026 Research / reviewers in the wild / expert
Hao Lin 0002
dblp:89/3472-2
· DBLP profile ↗
16ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0002-1921-3036ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 1 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Catching fraudulent "invisible hands": An approach to finding associative dense blocks in bipartite graphs
Yili Ren, Hao Lin 0002, Jiazhou Yu, Zhou Feng, Jixian Zhou, Ling Man, Guannan Liu 0004 |
Inf. Process. Manag. | 2 |
| 2024 | Deterring the Gray Market: Product Diversion Detection via Learning Disentangled Representations of Multivariate Time SeriesabstractA gray market emerges when some distributors divert products to unauthorized distributors/retailers to make sneaky profits from the manufacturers’ differential channel incentives, such as quantity discounts. Traditionally, manufacturers rely heavily on internal audits to periodically investigate the flows of products and funds so as to deter the gray market; however, this is too costly given the large number of distributors and their huge volumes of orders. Owing to the advances in data analytics techniques, the ordering quantities of a distributor over time, which form multivariate time series, can help reveal suspicious product diversion behaviors and narrow the audit scope drastically. To that end, in this paper, we build on the recent advancement of representation learning for time series and adopt a sequence autoencoder to automatically characterize the overall demand patterns. To cope with the underlying entangled factors and interfering information in the multivariate time series of ordering quantities, we develop a disentangled learning scheme to construct more effective sequence representations. An interdistributor correlation regularization is also proposed to ensure more reliable representations. Finally, given the highly scarce anomaly labels for the detection task, an unsupervised deep generative model based on the learned representations of the distributors is developed to estimate the densities of distributions, which enables the anomaly scores generated through end-to-end learning. Extensive experiments on a real-world distribution channel data set and a larger simulated data set empirically validate our model’s superior and robust performances compared with several state-of-the-art baselines. Additionally, our illustrative economic analysis demonstrates that the manufacturers can launch more targeted and cost-effective audits toward the suspected distributors recommended by our model so as to deter the gray market. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Funding: This work was supported by the National Natural Science Foundation of China [Grants 72031001, 72301017, 72371011, and 72242101]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.0155 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2022.0155 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Hao Lin 0002, Guannan Liu 0004, Junjie Wu 0002, J. Leon Zhao |
INFORMS J. Comput. | 1 |
| 2024 | BoostXML: Gradient Boosting for Extreme Multilabel Text Classification With Tail LabelsabstractMultilabel learning involving hundreds of thousands or even millions of labels is referred to as extreme multilabel learning (XML), in which the labels often follow a power-law distribution with the majority occurring in very few data points as tail labels. Recent years have witnessed the intensive use of deep-learning methods for high-performance XML, but they are typically optimized for the head labels with abundant training instances and less consider the performance on tail labels, which, however, like the needles in haystacks, are often the focus of attention in real-life applications. In light of this, we present BoostXML, a deep learning-based XML method for extreme multilabel text classification, enhanced greatly by gradient boosting. In BoostXML, we pay more attention to tail labels in each Boosting Step by optimizing the residual mostly from unfitted training instances with tail labels. A Corrective Step is further proposed to avoid the mismatching between the text encoder and weak learners during optimization, which reduces the risk of falling into local optima and improves model performance. A Pretraining Step is also introduced in the initial stage of BoostXML to avoid exorbitant bias to tail labels. Extensive experiments on five benchmark datasets with state-of-the-art baselines demonstrate the advantage of BoostXML in tail-label prediction. Fengzhi Li, Yuan Zuo, Hao Lin 0002, Junjie Wu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Telecom Fraud Detection via Hawkes-Enhanced Sequence ModelabstractDetecting frauds from a massive amount of user behavioral data is often regarded as finding a needle in a haystack. While tremendous efforts have been devoted to fraud detection from behavioral sequences, existing studies rarely consider behavioral targets and companions and their interactions simultaneously in a sequence model. In this paper, we suggest extracting source and target neighbor sequences from the temporal bipartite network of user behaviors, and disclose the interesting correlation mode and repetition mode hidden inside the two types of sequences as important clues for fraudsters distinguishment. We then propose a novel Hawkes-enhanced sequence model (HESM) by integrating the Hawkes process into LSTM for historical influence learning. A historical attention mechanism is also proposed to enhance the strength of the long-term historical influence in response to the repetition mode. Moreover, in order to collectively model both types of neighbor sequences for capturing the correlation mode, we propose a correlation gate to control the information flow in sequences. We conduct extensive experiments on real-world datasets and demonstrate that HESM outperforms competitive baseline methods consistently in telecom fraud detection. Particularly, the abilities of HESM in historical influence leaning and sequence correlation learning have been explored visually and intensively. Guannan Liu 0004, Junjie Wu 0002, Hao Lin 0002 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Topic Modeling of Short Texts: A Pseudo-Document View With Word Embedding EnhancementabstractRecent years have witnessed the unprecedented growth of online social media, resulting in short texts being the prevalent format of information on the Internet. Given the sparsity of data, however, short-text topic modeling remains a critical yet much-watched challenge in both academia and industry. Research has been devoted to building different types of probabilistic topic models for short texts, among which self-aggregation methods emerged recently to provide informative cross-text word co-occurrences. However, models along this line are still in their infancy and typically yield overfit results and exhibit high computational costs. In this paper, we propose a novel model called Pseudo-document-based Topic Model (PTM), which introduces the concept of pseudo-document to implicitly aggregate short texts against data sparsity. By modeling the topic distributions of latent pseudo-documents rather than short texts, PTM yields excellent performance in accuracy and efficiency. A word embedding-enhanced PTM (WE-PTM) is also proposed to leverage pre-trained word embeddings, which is essential to further alleviating data sparsity. Extensive experiments with self-aggregation or word embedding-based baselines on four real-world datasets including two online media short texts, demonstrate the high-quality topics learned by our models. Robustness to limited training samples and the explainable semantics of topics are also investigated. Yuan Zuo, Congrui Li, Hao Lin 0002, Junjie Wu 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Algorithm 1038: KCC: A MATLAB Package for k-Means-based Consensus ClusteringabstractConsensus clustering is gaining increasing attention for its high quality and robustness. In particular, k -means-based Consensus Clustering (KCC) converts the usual computationally expensive problem to a classic k -means clustering with generalized utility functions, bringing potentials for large-scale data clustering on different types of data. Despite KCC’s applicability and generalizability, implementing this method such as representing the binary dataset in the k -means heuristic is challenging and has seldom been discussed in prior work. To fill this gap, we present a MATLAB package, KCC, that completely implements the KCC framework and utilizes a sparse representation technique to achieve a low space complexity. Compared to alternative consensus clustering packages, the KCC package is of high flexibility, efficiency, and effectiveness. Extensive numerical experiments are also included to show its usability on real-world datasets. Hao Lin 0002, Hongfu Liu 0001, Junjie Wu 0002, Stephan Günnemann |
ACM Trans. Math. Softw. | 1 |
| 2021 | Where to go? Predicting next location in IoT environment
Hao Lin 0002, Guannan Liu 0004, Fengzhi Li, Yuan Zuo |
Frontiers Comput. Sci. | 1 |
| 2020 | Fraud Detection in Dynamic Interaction NetworkabstractFraud detection from massive user behaviors is often regarded as trying to find a needle in a haystack. In this paper, we suggest abnormal behavioral patterns can be better revealed if both sequential and interaction behaviors of users can be modeled simultaneously, which however has rarely been addressed in prior work. Along this line, we propose a COllective Sequence and INteraction (COSIN) model, in which the behavioral sequences and interactions between source and target users in a dynamic interaction network are modeled uniformly in a probabilistic graphical model. More specifically, the sequential schema is modeled with a hierarchical Hidden Markov Model, and meanwhile it is shifted to the interaction schema to generate the interaction counts through Poisson factorization. A hybrid Gibbs-Variational algorithm is then proposed for efficient parameter estimation of the COSIN model. We conduct extensive experiments on both synthetic and real-world telecom datasets in different scales, and the results show that the proposed model outperforms some competitive baseline methods and is scalable. A case is further presented to show the precious explainability of the model. Hao Lin 0002, Guannan Liu 0004, Junjie Wu 0002, Yuan Zuo |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Enhancing Employer Brand Evaluation with Collaborative Topic Regression ModelsabstractEmployer Brand Evaluation (EBE) is to understand an employer’s unique characteristics to identify competitive edges. Traditional approaches rely heavily on employers’ financial information, including financial reports and filings submitted to the Securities and Exchange Commission (SEC), which may not be readily available for private companies. Fortunately, online recruitment services provide a variety of employers’ information from their employees’ online ratings and comments, which enables EBE from an employee’s perspective. To this end, in this article, we propose a method named Company Profiling–based Collaborative Topic Regression (CPCTR) to collaboratively model both textual (i.e., reviews) and numerical information (i.e., salaries and ratings) for learning latent structural patterns of employer brands. With identified patterns, we can effectively conduct both qualitative opinion analysis and quantitative salary benchmarking. Moreover, a Gaussian processes--based extension, GPCTR, is proposed to capture the complex correlation among heterogeneous information. Extensive experiments are conducted on three real-world datasets to validate the effectiveness and generalizability of our methods in real-life applications. The results clearly show that our methods outperform state-of-the-art baselines and enable a comprehensive understanding of EBE. Hao Lin 0002, Hengshu Zhu, Junjie Wu 0002, Yuan Zuo, Chen Zhu 0003, Hui Xiong 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2020 | A Pseudo-document-based Topical N-grams model for short texts
Hao Lin 0002, Yuan Zuo, Guannan Liu 0004, Junjie Wu 0002, Zhiang Wu 0001 |
World Wide Web | 1 |
| 2018 | Embedding Temporal Network via Neighborhood FormationabstractGiven the rich real-life applications of network mining as well as the surge of representation learning in recent years, network embedding has become the focal point of increasing research interests in both academic and industrial domains. Nevertheless, the complete temporal formation process of networks characterized by sequential interactive events between nodes has yet seldom been modeled in the existing studies, which calls for further research on the so-called temporal network embedding problem. In light of this, in this paper, we introduce the concept of neighborhood formation sequence to describe the evolution of a node, where temporal excitation effects exist between neighbors in the sequence, and thus we propose a Hawkes process based Temporal Network Embedding (HTNE) method. HTNE well integrates the Hawkes process into network embedding so as to capture the influence of historical neighbors on the current neighbors. In particular, the interactions of low-dimensional vectors are fed into the Hawkes process as base rate and temporal influence, respectively. In addition, attention mechanism is also integrated into HTNE to better determine the influence of historical neighbors on current neighbors of a node. Experiments on three large-scale real-life networks demonstrate that the embeddings learned from the proposed HTNE model achieve better performance than state-of-the-art methods in various tasks including node classification, link prediction, and embedding visualization. In particular, temporal recommendation based on arrival rate inferred from node embeddings shows excellent predictive power of the proposed model. Yuan Zuo, Guannan Liu 0004, Hao Lin 0002, Xiaoqian Hu, Junjie Wu 0002 |
KDD | 3 |
| 2017 | Collaborative Company Profiling: Insights from an Employee's PerspectiveabstractCompany profiling is an analytical process to build an in-depth understanding of company's fundamental characteristics. It serves as an effective way to gain vital information of the target company and acquire business intelligence. Traditional approaches for company profiling rely heavily on the availability of rich finance information about the company, such as finance reports and SEC filings, which may not be readily available for many private companies. However, the rapid prevalence of online employment services enables a new paradigm — to obtain the variety of company's information from their employees' online ratings and comments. This, in turn, raises the challenge to develop company profiles from an employee's perspective. To this end, in this paper, we propose a method named Company Profiling based Collaborative Topic Regression (CPCTR), for learning the latent structural patterns of companies. By formulating a joint optimization framework, CPCTR has the ability in collaboratively modeling both textual (e.g., reviews) and numerical information (e.g., salaries and ratings). Indeed, with the identified patterns, including the positive/negative opinions and the latent variable that influences salary, we can effectively carry out opinion analysis and salary prediction. Extensive experiments were conducted on a real-world data set to validate the effectiveness of CPCTR. The results show that our method provides a comprehensive understanding of company characteristics and delivers a more effective prediction of salaries than other baselines. Hao Lin 0002, Hengshu Zhu, Yuan Zuo, Chen Zhu 0003, Junjie Wu 0002, Hui Xiong 0001 |
AAAI | 1 |
| 2016 | Topic Modeling of Short Texts: A Pseudo-Document ViewabstractRecent years have witnessed the unprecedented growth of online social media, which empower short texts as the prevalent format for information of Internet. Given the nature of sparsity, however, short text topic modeling remains a critical yet much-watched challenge in both academy and industry. Rich research efforts have been put on building different types of probabilistic topic models for short texts, among which the self aggregation methods without using auxiliary information become an emerging solution for providing informative cross-text word co-occurrences. However, models along this line are still rarely seen, and the representative one Self-Aggregation Topic Model (SATM) is prone to overfitting and computationally expensive. In light of this, in this paper, we propose a novel probabilistic model called Pseudo-document-based Topic Model (PTM) for short text topic modeling. PTM introduces the concept of pseudo document to implicitly aggregate short texts against data sparsity. By modeling the topic distributions of latent pseudo documents rather than short texts, PTM is expected to gain excellent performance in both accuracy and efficiency. A Sparsity-enhanced PTM (SPTM for short) is also proposed by applying Spike and Slab prior, with the purpose of eliminating undesired correlations between pseudo documents and latent topics. Extensive experiments on various real-world data sets with state-of-the-art baselines demonstrate the high quality of topics learned by PTM and its robustness with reduced training samples. It is also interesting to show that i) SPTM gains a clear edge over PTM when the number of pseudo documents is relatively small, and ii) the constraint that a short text belongs to only one pseudo document is critically important for the success of PTM. We finally take an in-depth semantic analysis to unveil directly the fabulous function of pseudo documents in finding cross-text word co-occurrences for topic modeling. Yuan Zuo, Junjie Wu 0002, Hui Zhang 0028, Hao Lin 0002, Fei Wang 0148, Ke Xu 0001, Hui Xiong 0001 |
KDD | 4 |
| 2015 | Complementary Aspect-Based Opinion Mining Across Asymmetric CollectionsabstractAspect-based opinion mining is to find elaborate opinions towards an underlying theme, perspective or viewpoint as to a subject such as a product or an event. Nowadays, with rapid growing of opinionated text on the Web, mining aspect-level opinions has become a promising means for online public opinion analysis. In particular, the booming of various types of online media provide diverse yet complementary information, bringing unprecedented opportunities for public opinion analysis across different populations. Along this line, in this paper, we propose CAMEL, a novel topic model for complementary aspect-based opinion mining across asymmetric collections. CAMEL gains complementarity by modeling both common and specific aspects across different collections, and keeping all the corresponding opinions for contrastive study. To further boost CAMEL, we propose AME, an automatic labeling scheme for maximum entropy model, to help discriminate aspect and opinion words without heavy human labeling. Extensive experiments on synthetic multicollection data sets demonstrate the superiority of CAMEL to baseline methods, in leveraging cross-collection complementarity to find higher-quality aspects and more coherent opinions as well as aspect-opinion relationships. This is particularly true when the collections get seriously imbalanced. Experimental results also show that the AME model indeed outperforms manual labeling in suggesting true opinion words. Finally, case study on two public events further demonstrates the practical value of CAMEL for real-world public opinion analysis. Yuan Zuo, Junjie Wu 0002, Hui Zhang 0028, Deqing Wang 0001, Hao Lin 0002, Fei Wang 0148, Ke Xu 0001 |
ICDM | 5 |
| 2013 | How Many Zombies Around You?abstractRecent years have witnessed the explosive growth of online social media. Weibo, a famous "Chinese Twitter", has attracted over half billion users in less than four years. Among them are zombie users or bogus users, who are seemingly active common users but actually marionettes manipulated by intelligent software for economic interests. To probe such users thus becomes critically important for a healthy Weibo, but the existing studies along this line are still in initial stage due to the serious lack of labeled zombies and the limited attributes for user profiling. In light of this, in this paper, we figure out a commercial way for training set labeling, and propose a two-stage cascading model called ProZombie for zombie user recognition. ProZombie decomposes the training/predicting process into fast and refined phases in cascade, which greatly improves the modeling efficiency without sacrificing the accuracy. Moreover, 35 attributes including 16 newly proposed ones are employed for a panoramic description of Weibo users. Experiments on real-world labeled Weibo users demonstrate the effectiveness and efficiency of ProZombie. More interestingly, two case studies based on ProZombie successfully unveil the zombies hidden around common users, and their impact to information propagation on Weibo. To our best knowledge, this study is among the first to quantify these interesting observations on Weibo. Hongfu Liu 0001, Hao Lin 0002, Junjie Wu 0002, Zhiang Wu 0001 |
ICDM | 3 |
| 2013 | SEA: a system for event analysis on chinese tweetsabstractRecent years have witnessed the explosive growth of online social media. Weibo, a famous "Chinese Twitter", has attracted over 0.5 billion users in less than four years, with more than 1000 tweets generated in every second. These tweets are informative but very fragmented, and thus would be better archived from an event perspective, as done by Weibo itself in the "Micro-Topic" program. This effort, however, is yet far from satisfaction for not providing enough analytical power to events. In light of this, in this demo paper, we propose SEA, a System for Event Analysis on Chinese tweets. In general, SEA is an event-centric, multi-functional platform that conducts panoramic analysis on Weibo events from various aspects, including the semantic information of the events, the temporal and spatial trends, the public sentiments, the hidden sub-events, the key users in the event diffusion and their preferences, etc. These functions are enabled by the integration of various analytical models and by the noSQL techniques adopted purposefully for massive tweets management. Finally, a case study on the "Spring Festival" event demonstrates the effectiveness of SEA. To our best knowledge, SEA is the first third-party system that provides panoramic analysis to Weibo events. Yaqiong Wang, Hongfu Liu 0001, Hao Lin 0002, Junjie Wu 0002, Zhiang Wu 0001, Jie Cao 0001 |
KDD | 3 |