Linglong Kong

dblp:35/8525 · DBLP profile ↗
← Back
14ranked-venue papers in the field
0as first author
9since 2021 · last 2025
0000-0003-3011-9216ORCID · corroborated

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 10Information Retrieval & Web Search · 4
YearPublicationVenuePosition
2025 Oblivious Johnson-Lindenstrauss embeddings for compressed Tucker decompositions
Matthew Pietrosanu, Bei Jiang, Linglong Kong
CIKM3
2025 Adaptive Conformal Prediction Intervals for Invariant Learning
abstract
Distribution shifts between training and testing datasets can significantly affect the effectiveness of machine learning models, as the typical assumption of data homogeneity often does not hold in practical applications. To address this challenge, previous literature explores invariant representation learning methods, including but not limited to domain adversarial learning, feature alignment, and invariant risk minimization. In this paper, we introduce a novel method for generating distribution-free prediction regions, which effectively quantify the uncertainty of predictions in different environments. Our methodology incorporates a weighted conformity score, tailored to the environment of each test sample, to construct adaptive conformal intervals. The proposed method is computationally efficient and can be implemented by subsampling. We demonstrate how these intervals maintain a consistent conditional coverage under some conditions. The numerical experiments, including simulations and real data analysis, show that our method is applicable and robust in dealing with distribution shifts.
Shuxin Liang, Yihan Xiao, Linglong Kong, Wenlu Tang
KDD (2)3
2025 A Bayesian hierarchical model for orthogonal Tucker decomposition with oblivious tensor compression
Matthew Pietrosanu, Bei Jiang, Linglong Kong
Knowl. Inf. Syst.3
2024 A Bayesian Hierarchical Model for Orthogonal Tucker Decomposition with Oblivious Tensor Compression
abstract
Low-rank representations such as the Tucker decomposition underlie many frequentist methods for tensor analysis. Bayesian analogues, in contrast, have received less attention. Notably missing in the literature is a Bayesian Tucker decomposition with orthogonal factor matrices-a standard interpretability restriction in frequentist settings. We propose a Bayesian hierarchical model for the orthogonal Tucker decomposition, which we implement via conditionally conjugate Gibbs sampler. To reduce the complexity of tensor operations in MCMC estimation, we incorporate a mechanism that uses Johnson-Lindenstrauss embeddings to compress data. Our theoretical analysis bounds change in the full-conditional posterior distributions of tensor components due to compression (with respect to Hellinger distance). We further establish posterior consistency for the decomposition's factor matrices in settings where these parameters are shared across tensor observations. Empirical results show that, for large tensor datasets, moderate compression can significantly reduce draw time (by about 50%) with only a moderate increase (up to about 15%) in median reconstruction error. Compression in the proposed model additionally enables analyses of tensor data that are too large to be held wholly in memory, thus making large-scale analyses tractable on even moderate computing resources.
Matthew Pietrosanu, Bei Jiang, Linglong Kong
ICDM3
2023 Exploring the Training Robustness of Distributional Reinforcement Learning Against Noisy State Observations
Ke Sun 0013, Shangling Jui, Linglong Kong
ECML/PKDD (5)4
2023 Optimal Smooth Approximation for Quantile Matrix Factorization
abstract
Matrix Factorization (MF) is essential to many estimation tasks. Most existing matrix factorization methods focus on least squares matrix factorization (LSMF), which aims to minimize a smooth L2 loss between observations and their dependent matrix measurement variables. In reality, however, L1 loss and check loss are widely used in regression to deal with outliers or observations contaminated by skewed or heavy-tailed noise. Although under certain conditions, linear convergence to the global optimality can be established for matrix factorization under the L2 loss, there is a lack of provably efficient algorithms for solving matrix factorization under non-smooth losses. In this paper, we investigate Quantile Matrix Factorization (QMF), the counterpart of Quantile Regression in matrix estimation, that adopts a tunable check loss and introduces robustness to matrix estimation for skewed and heavy- tailed observations, which are prevalent in reality. To deal with the non-smooth loss, we propose Nesterov- smoothed QMF (NsQMF), extending Nesterov's optimal smooth approximation technique to the matrix factorization setting. We then present an alternating minimization algorithm to solve the smooth NsQMF efficiently. We mathematically prove that solving the smoothed NsQMF is equivalent to solving the original non-smooth QMF problem and that our proposed algorithm achieves linear convergence to the global optimality of QMF. Numerical evaluations verify our theoretical findings and demonstrate that NsQMF significantly outperforms the commonly used LSMF and prior approximate smoothing heuristics for QMF under various noise distributions.
Peng Liu 0048, Yi Liu 0062, Rui Zhu 0007, Linglong Kong, Bei Jiang, Di Niu 0002
SDM4
2022 TAG: Toward Accurate Social Media Content Tagging with a Concept Graph
abstract
Although conceptualization has been widely studied in semantics and knowledge representation, it is still challenging to find the most accurate concept terms to tag fast-growing social media content. This is partly attributed to the fact that most traditional knowledge bases contain general terms of the world, such as trees and cars, which are not interesting to users, and do not have the defining power for social media content. Another reason is that the intricate use of tense, negation and grammar in social media content may change the logic or emphasis of the content, thus focusing on different main ideas. In this paper, we present TAG, a high-quality concept matching dataset consisting of 10,000 labeled pairs of fine-grained concepts and web-styled natural language sentences, mined from open-domain social media content. The concepts we provide are the trending terms on social media and have the right granularity to define user interests, e.g., highly educated actors instead of just actors. In the meantime, TAG offers a concept graph which interconnects these fine-grained concepts and entities to provide contextual information. We evaluate a wide range of neural text matching models as well as pre-trained language models for the concept matching task on TAG, and point out their insufficiency to tag social media content to characterize its main idea. We further propose a novel graph-graph matching framework that demonstrates superior abstraction and generalization performance by better utilizing both the structural information in the concept graph and logic interactions between semantic units in the natural language sentence via syntactic dependency parsing.
Jiuding Yang, Weidong Guo, Bang Liu 0003, Yakun Yu, Jinwen Luo, Linglong Kong, Di Niu 0002
KDD7
2022 Factorizing Historical User Actions for Next-Day Purchase Prediction
abstract
It is common practice for many large e-commerce operators to analyze daily logged transaction data to predict customer purchase behavior, which may potentially lead to more effective recommendations and increased sales. Traditional recommendation techniques based on collaborative filtering, although having gained success in video and music recommendation, are not sufficient to fully leverage the diverse information contained in the implicit user behavior on e-commerce platforms. In this article, we analyze user action records in the Alibaba Mobile Recommendation dataset from the Alibaba Tianchi Data Lab, as well as the Retailrocket recommender system dataset from the Retail Rocket website. To estimate the probability that a user will purchase a certain item tomorrow, we propose a new model called Time-decayed Multifaceted Factorizing Personalized Markov Chains (Time-decayed Multifaceted-FPMC), taking into account multiple types of user historical actions not only limited to past purchases but also including various behaviors such as clicks, collects and add-to-carts. Our model also considers the time-decay effect of the influence of past actions. To learn the parameters in the proposed model, we further propose a unified framework named Bayesian Sparse Factorization Machines. It generalizes the theory of traditional Factorization Machines to a more flexible learning structure and trains the Time-decayed Multifaceted-FPMC with the Markov Chain Monte Carlo method. Extensive evaluations based on multiple real-world datasets demonstrate that our proposed approaches significantly outperform various existing purchase recommendation algorithms.
Bang Liu 0003, Linglong Kong, Di Niu 0002
ACM Trans. Web3
2021 L2NAS: Learning to Optimize Neural Architectures via Continuous-Action Reinforcement Learning
abstract
Neural architecture search (NAS) has achieved remarkable results in deep neural network design. Differentiable architecture search converts the search over discrete architectures into a hyperparameter optimization problem which can be solved by gradient descent. However, questions have been raised regarding the effectiveness and generalizability of gradient methods for solving non-convex architecture hyperparameter optimization problems. In this paper, we propose L2NAS, which learns to intelligently optimize and update architecture hyperparameters via an actor neural network based on the distribution of high-performing architectures in the search history. We introduce a quantile-driven training procedure which efficiently trains L2NAS in an actor-critic framework via continuous-action reinforcement learning. Experiments show that L2NAS achieves state-of-the-art results on NAS-Bench-201 benchmark as well as DARTS search space and Once-for-All MobileNetV3 search space. We also show that search policies generated by L2NAS are generalizable and transferable across different training datasets with minimal fine-tuning.
Keith G. Mills, Fred X. Han, Mohammad Salameh, Seyed Saeed Changiz Rezaei, Linglong Kong, Wei Lu 0023, Shuo Lian, Shangling Jui, Di Niu 0002
CIKM5
2020 Story Forest: Extracting Events and Telling Stories from Breaking News
abstract
Extracting events accurately from vast news corpora and organize events logically is critical for news apps and search engines, which aim to organize news information collected from the Internet and present it to users in the most sensible forms. Intuitively speaking, an event is a group of news documents that report the same news incident possibly in different ways. In this article, we describe our experience of implementing a news content organization system at Tencent to discover events from vast streams of breaking news and to evolve news story structures in an online fashion. Our real-world system faces unique challenges in contrast to previous studies on topic detection and tracking (TDT) and event timeline or graph generation, in that we (1) need to accurately and quickly extract distinguishable events from massive streams of long text documents, and (2) must develop the structures of event stories in an online manner, in order to guarantee a consistent user viewing experience. In solving these challenges, we propose Story Forest , a set of online schemes that automatically clusters streaming documents into events, while connecting related events in growing trees to tell evolving stories. A core novelty of our Story Forest system is EventX , a semi-supervised scheme to extract events from massive Internet news corpora. EventX relies on a two-layered, graph-based clustering procedure to group documents into fine-grained events. We conducted extensive evaluations based on (1) 60 GB of real-world Chinese news data, (2) a large Chinese Internet news dataset that contains 11,748 news articles with truth event labels, and (3) the 20 News Groups English dataset, through detailed pilot user experience studies. The results demonstrate the superior capabilities of Story Forest to accurately identify events and organize news text into a logical structure that is appealing to human readers.
Bang Liu 0003, Fred X. Han, Di Niu 0002, Linglong Kong, Kunfeng Lai
ACM Trans. Knowl. Discov. Data4
2019 M-estimation in Low-Rank Matrix Factorization: A General Framework
abstract
Many problems in science and engineering can be reduced to the recovery of an unknown large matrix from a small number of random linear measurements. Matrix factorization arguably is the most popular approach for low-rank matrix recovery. Many methods have been proposed using different loss functions, for example the most widely used L2loss, more robust choices such as L1and Huber loss, quantile and expectile loss for skewed data. All of them can be unified into the framework of M-estimation. In this paper, we present a general framework of low-rank matrix factorization based on M-estimation in statistics. The framework mainly involves two steps: firstly we apply Nesterov's smoothing technique to obtain an optimal smooth approximation for non-smooth loss function, such as L1and quantile loss; secondly we exploit an alternative updating scheme along with Nesterov's momentum method at each step to minimize the smoothed loss function. Strong theoretical convergence guarantee has been developed for the general framework, and extensive numerical experiments have been conducted to illustrate the performance of proposed algorithm.
Peng Liu 0048, Jingyu Zhao 0001, Yi Liu 0062, Linglong Kong, Bei Jiang, Guangjian Tian, Hengshuai Yao
ICDM5
2017 Growing Story Forest Online from Massive Breaking News
abstract
We describe our experience of implementing a news content organization system at Tencent that discovers events from vast streams of breaking news and evolves news story structures in an online fashion. Our real-world system has distinct requirements in contrast to previous studies on topic detection and tracking (TDT) and event timeline or graph generation, in that we 1) need to accurately and quickly extract distinguishable events from massive streams of long text documents that cover diverse topics and contain highly redundant information, and 2) must develop the structures of event stories in an online manner, without repeatedly restructuring previously formed stories, in order to guarantee a consistent user viewing experience. In solving these challenges, we propose Story Forest, a set of online schemes that automatically clusters streaming documents into events, while connecting related events in growing trees to tell evolving stories. We conducted extensive evaluation based on 60 GB of real-world Chinese news data, although our ideas are not language-dependent and can easily be extended to other languages, through detailed pilot user experience studies. The results demonstrate the superior capability of Story Forest to accurately identify events and organize news text into a logical structure that is appealing to human readers, compared to multiple existing algorithm frameworks.
Bang Liu 0003, Di Niu 0002, Kunfeng Lai, Linglong Kong
CIKM4
2017 Recover Fine-Grained Spatial Data from Coarse Aggregation
abstract
In this paper, we study a new type of spatial sparse recovery problem, that is to infer the fine-grained spatial distribution of certain density data in a region only based on the aggregate observations recorded for each of its subregions. One typical example of this spatial sparse recovery problem is to infer spatial distribution of cellphone activities based on aggregate mobile traffic volumes observed at sparsely scattered base stations. We propose a novel Constrained Spatial Smoothing (CSS) approach, which exploits the local continuity that exists in many types of spatial data to perform sparse recovery via finite-element methods, while enforcing the aggregated observation constraints through an innovative use of the ADMM algorithm. We also improve the approach to further utilize additional geographical attributes. Extensive evaluations based on a large dataset of phone call records and a demographical dataset from the city of Milan show that our approach significantly outperforms various state-of-the-art approaches, including Spatial Spline Regression (SSR).
Bang Liu 0003, Borislav Mavrin, Linglong Kong, Di Niu 0002
ICDM3
2016 House Price Modeling over Heterogeneous Regions with Hierarchical Spatial Functional Analysis
abstract
Online real-estate information systems such as Zillow and Trulia have gained increasing popularity in recent years. One important feature offered by these systems is the online home price estimate through automated data-intensive computation based on housing information and comparative market value analysis. State-of-the-art approaches model house prices as a combination of a latent land desirability surface and a regression from house features. However, by using uniformly damping kernels, they are unable to handle irregularly shaped regions or capture land value discontinuities within the same region due to the existence of implicit sub-communities, which are common in real-world scenarios. In this paper, we explore the novel application of recent advances in spatial functional analysis to house price modeling and propose the Hierarchical Spatial Functional Model (HSFM), which decomposes house values into land desirability at both the global scale and hidden local scales as well as the feature regression component. We propose statistical learning algorithms based on finite-element spatial functional analysis and spatial constrained clustering to train our model. Extensive evaluations based on housing data in a major Canadian city show that our proposed approach can reduce the mean relative house price estimation error down to 6.60%.
Bang Liu 0003, Borislav Mavrin, Di Niu 0002, Linglong Kong
ICDM4