Linhong Zhu

dblp:23/6504 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
2since 2021 · last 2026
0000-0001-8350-392XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 20 · 9 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Data mining · 45% Recommender systems · 19% Graph data management · 17%
Artificial intelligence
3 papers
Vision and language · 50% Graph learning · 38% Information extraction and text analysis · 12%
Theoretical computer science
2 papers
Graph algorithms and graph theory · 100%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Smart cities and intelligent transportation · 100%

Topics — the 26 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
A General Framework for Multimodal LLM-Based Multimedia Understanding in Large-Scale Recommendation Systems · SIGIR 2026
Recommender systems
multimodal recommendation
1.012026
A General Framework for Multimodal LLM-Based Multimedia Understanding in Large-Scale Recommendation Systems · SIGIR 2026
Data mining › representation learning
latent space models
0.522017
Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks (Extended Abstract) · ICDE 2017
Latent Space Model for Road Networks to Predict Time-Varying Traffic · KDD 2016
Data mining › structured data mining
graph mining
0.512021
Label Propagation on K-Partite Graphs with Heterophily · IEEE Trans. Knowl. Data Eng. 2021
Data mining › semi-supervised learning
label propagation
0.512021
Label Propagation on K-Partite Graphs with Heterophily · IEEE Trans. Knowl. Data Eng. 2021
Smart cities and intelligent transportation
traffic prediction
0.322017
Latent Space Model for Road Networks to Predict Time-Varying Traffic · KDD 2016
Situation Aware Multi-task Learning for Traffic Prediction · ICDM 2017
Graph data management
dynamic graph processing
0.312017
Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks (Extended Abstract) · ICDE 2017
Knowledge graphs
link prediction
0.312017
Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks (Extended Abstract) · ICDE 2017
Knowledge graphs › link prediction
temporal link prediction
0.312017
Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks (Extended Abstract) · ICDE 2017
Machine learning › Graph learning › dynamic graph learning
dynamic link prediction
0.212016
Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks · IEEE Trans. Knowl. Data Eng. 2016
Machine learning › Graph learning
link prediction
0.212016
Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks · IEEE Trans. Knowl. Data Eng. 2016
Machine learning › Graph learning
network embedding
0.212016
Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks · IEEE Trans. Knowl. Data Eng. 2016
Data mining
spatiotemporal data mining
0.212016
Latent Space Model for Road Networks to Predict Time-Varying Traffic · KDD 2016
Graph data management › cohesive subgraph mining
maximal clique enumeration
0.222011
Finding maximal cliques in massive networks · ACM Trans. Database Syst. 2011
Finding maximal cliques in massive networks by H*-graph · SIGMOD Conference 2010
Data mining
clustering
0.212014
Tripartite graph clustering for dynamic sentiment analysis on social media · SIGMOD Conference 2014
Data mining › clustering
graph clustering
0.212014
Tripartite graph clustering for dynamic sentiment analysis on social media · SIGMOD Conference 2014
Data mining › text mining
sentiment analysis
0.212014
Tripartite graph clustering for dynamic sentiment analysis on social media · SIGMOD Conference 2014
Web and social media mining
social network analysis
0.112021
Label Propagation on K-Partite Graphs with Heterophily · IEEE Trans. Knowl. Data Eng. 2021
Parallel and multicore computing
parallel graph algorithms
0.112012
Fast algorithms for maximal clique enumeration with limited memory · KDD 2012
Graph algorithms and graph theory › graph algorithms › subgraph enumeration
clique enumeration
0.112012
Fast algorithms for maximal clique enumeration with limited memory · KDD 2012
Graph algorithms and graph theory › graph algorithms › subgraph enumeration › clique enumeration
maximal clique enumeration
0.112012
Fast algorithms for maximal clique enumeration with limited memory · KDD 2012
Graph data management
graph algorithms
0.112011
Finding maximal cliques in massive networks · ACM Trans. Database Syst. 2011
Indexing and storage engines
external memory algorithms
0.112010
Finding maximal cliques in massive networks by H*-graph · SIGMOD Conference 2010
Graph data management
graph indexing
0.112010
Finding maximal cliques in massive networks by H*-graph · SIGMOD Conference 2010
Knowledge graphs › knowledge graph construction
concept graph construction
0.112016
Modeling Concept Dependencies in a Scientific Corpus · ACL (1) 2016
Web and social media mining
social media analysis
0.112014
Tripartite graph clustering for dynamic sentiment analysis on social media · SIGMOD Conference 2014

Methods — techniques the papers use, named apart from their topics

tokenized categorical features · 2.0caption generation · 2.0latent space model · 1.0incremental online algorithm · 0.8global optimization · 0.5k-partite graph model · 0.5information-theoretic modeling · 0.5incremental algorithm · 0.5human evaluation · 0.5partition-based algorithm · 0.3nested partitioning · 0.3cost model · 0.3group lasso · 0.3clustering · 0.3block coordinate gradient descent · 0.3FISTA · 0.3incremental update · 0.2
YearPublicationVenuePosition
2026 A General Framework for Multimodal LLM-Based Multimedia Understanding in Large-Scale Recommendation Systems
abstract
Conventional recommendation systems frequently fail to fully exploit the high-dimensional semantic signals inherent in multimedia content, thereby limiting the fidelity of user preference modeling. While Multimodal Large Language Models (MM-LLMs) offer robust mechanisms for interpreting such complex data, their integration into latency-constrained, industrial-scale architectures remains a significant challenge. To address this, we propose a generalized framework for MM-LLM-driven multimedia understanding. Our methodology employs a tripartite architecture encompassing content interpretation, representation extraction, and systematic pipeline integration, instantiated via a LLaMA2-based model that generates descriptive captions subsequently ingested as tokenized categorical features. Empirical evaluation demonstrates the efficacy of this approach, yielding a 0.35% increase in offline AUC and a 0.02% improvement in online metrics at scale, substantiating the practical viability of leveraging MM-LLMs to enhance large-scale recommendation performance.
Ziyun Xu, Joena Zhang, Sirius Chen, Chenheli Hua, Silvester Yao, Qichao Que, Wentao Shi 0002, Junfeng Pan, Linhong Zhu
SIGIR12
2021 Label Propagation on K-Partite Graphs with Heterophily
abstract
In this paper, for the first time, we study label propagation in heterogeneous graphs under heterophily assumption. Homophily label propagation (i.e., two connected nodes share similar labels) in homogeneous graph (with same types of vertices and relations) has been extensively studied before. Unfortunately, real-life networks (e.g., social networks) are heterogeneous, they contain different types of vertices (e.g., users, images, and texts) and relations (e.g., friendships and co-tagging) and allow for each node to propagate both the same and opposite copy of labels to its neighbors. We propose a IC-partite label propagation model to handle the mystifying combination of heterogeneous nodes/relations and heterophily propagation. With this model, we develop a novel label inference algorithm framework with update rules in near-linear time complexity. Since real networks change overtime, we devise an incremental approach, which supports fast updates for both new data and evidence (e.g., ground truth labels) with guaranteed efficiency. We further provide a utility function to automatically determine whether an incremental or a re-modeling approach is favored. Extensive experiments on real datasets have verified the effectiveness and efficiency of our approach, and its superiority over the state-of-the-art label propagation methods.
Dingxiong Deng, Fan Bai 0001, Yiqi Tang, Shuigeng Zhou, Cyrus Shahabi, Linhong Zhu
IEEE Trans. Knowl. Data Eng.6
2019 Coupled Clustering of Time-Series and Networks
abstract
Motivated by the problem of human-trafficking, where it is often observed that criminal organizations are linked and behave similarly over time, we introduce the problem of Coupled Clustering of Time-series and their underlying Network. The goal is to find tightly connected subgroups of nodes that also have similar node-specific time series (temporal—not necessarily structural—behavior). We formulate the problem as a coupled matrix factorization for the time series, combined with regularization for network smoothness. We propose CCTN, and an incrementally-updated counterpart, CCTN-inc, which efficiently handles network updates. Extensive experiments show that CCTN is up to 4x more accurate than baselines that consider graph structure or time series alone, and CCTN-inc is up to 55x faster than CCTN. As an application, we explore an exclusive database with millions of online ads on human trafficking, and successfully deploy our technique to detect criminal organizations.
Linhong Zhu, Pedro A. Szekely, Aram Galstyan, Danai Koutra
SDM2
2017 Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks (Extended Abstract)
abstract
We propose to model dependence within a network view using the temporal latent space model, which uses a time-dependent low-dimensional geometric projections to represent the high-dimensional dependence structure in time-varying networks. Once we obtain the lowdimensional temporal latent space representation for graphs from time 1 to t, we can accurately predict future links in time t + 1 (i.e., Gt+1). We present a global optimization algorithm to effectively infer the temporal latent space using block coordinate gradient descent (BCGD). We further introduce two new variants of BCGD: a local BCGD algorithm and an incremental BCGD algorithm, to scale the inference algorithm to massive networks.
Linhong Zhu, Junming Yin, Greg Ver Steeg, Aram Galstyan
ICDE1
2017 Situation Aware Multi-task Learning for Traffic Prediction
abstract
Due to the recent vast availability of transportation traffic data, major research efforts have been devoted to traffic prediction, which is useful in many applications such as urban planning, traffic management and navigations systems. Current prediction methods that independently train a model per traffic sensor cannot accurately predict traffic in every situation (e.g., rush hours, constructions and accidents) because there may not exist sufficient training samples per sensor for all situations. To address this shortcoming, our core idea is to explore the commonalities of prediction tasks across multiple sensors who behave similarly in a specific traffic situation. Instead of building a model independently per sensor, we propose a Multi-Task Learning (MTL) framework that aims to first automatically identify the traffic situations and then simultaneously build one forecasting model for similar-behaving sensors per traffic situation. The key innovation here is that instead of the straightforward application of MTL where each "task" corresponds to a sensor, we relate each MTL's "task" to a traffic situation. Specifically, we first identify these traffic situations by running clustering algorithms on all sensors' data. Subsequently, to enforce the commonalities under each identified situation, we use the group Lasso regularization in MTL to select a common set of features for the prediction tasks, and we adapt efficient FISTA algorithm with guaranteed convergence rate. We evaluated our methods with a large volume of real-world traffic sensor data; our results show that by incorporating traffic situations, our proposed MTL framework performs consistently better than naively applying MTL per sensor. Moreover, our holistic approach, under different traffic situations, outperforms all the best traffic prediction approaches for a given situation by up to 18% and 30% in short and long term predictions, respectively.
Dingxiong Deng, Cyrus Shahabi, Ugur Demiryurek, Linhong Zhu
ICDM4
2016 Modeling Concept Dependencies in a Scientific Corpus
abstract
Our goal is to generate reading lists for students that help them optimally learn technical material.Existing retrieval algorithms return items directly relevant to a query but do not return results to help users read about the concepts supporting their query.This is because the dependency structure of concepts that must be understood before reading material pertaining to a given query is never considered.Here we formulate an information-theoretic view of concept dependency and present methods to construct a "concept graph" automatically from a text corpus.We perform the first human evaluation of concept dependency edges (to be published as open data), and the results verify the feasibility of automatic approaches for inferring concepts and their dependency relations.This result can support search capabilities that may be tuned to help users learn a subject rather than retrieve documents based on a single query.
Jonathan Gordon 0001, Linhong Zhu, Aram Galstyan, Premkumar Natarajan, Gully A. P. C. Burns
ACL (1)2
2016 Latent Space Model for Road Networks to Predict Time-Varying Traffic
abstract
Real-time traffic prediction from high-fidelity spatiotemporal traffic sensor datasets is an important problem for intelligent transportation systems and sustainability. However, it is challenging due to the complex topological dependencies and high dynamism associated with changing road conditions. In this paper, we propose a Latent Space Model for Road Networks (LSM-RN) to address these challenges holistically. In particular, given a series of road network snapshots, we learn the attributes of vertices in latent spaces which capture both topological and temporal properties. As these latent attributes are time-dependent, they can estimate how traffic patterns form and evolve. In addition, we present an incremental online algorithm which sequentially and adaptively learns the latent attributes from the temporal graph changes. Our framework enables real-time traffic prediction by 1) exploiting real-time sensor readings to adjust/update the existing latent spaces, and 2) training as data arrives and making predictions on-the-fly. By conducting extensive experiments with a large volume of real-world traffic sensor data, we demonstrate the superiority of our framework for real-time traffic prediction on large road networks over competitors as well as baseline graph-based LSM's.
Dingxiong Deng, Cyrus Shahabi, Ugur Demiryurek, Linhong Zhu, Rose Yu, Yan Liu 0002
KDD4
2016 Unsupervised Entity Resolution on Multi-type Graphs
Linhong Zhu, Majid Ghasemi-Gol, Pedro A. Szekely, Aram Galstyan, Craig A. Knoblock
ISWC (1)1
2016 Task selection in spatial crowdsourcing from worker's perspective
Dingxiong Deng, Cyrus Shahabi, Ugur Demiryurek, Linhong Zhu
GeoInformatica4
2016 Partitioning Networks with Node Attributes by Compressing Information Flow
abstract
Real-world networks are often organized as modules or communities of similar nodes that serve as functional units. These networks are also rich in content, with nodes having distinguished features or attributes. In order to discover a network’s modular structure, it is necessary to take into account not only its links but also node attributes. We describe an information-theoretic method that identifies modules by compressing descriptions of information flow on a network. Our formulation introduces node content into the description of information flow, which we then minimize to discover groups of nodes with similar attributes that also tend to trap the flow of information. The method is conceptually simple and does not require ad-hoc parameters to specify the number of modules or to control the relative contribution of links and node attributes to network structure. We apply the proposed method to partition real-world networks with known community structure. We demonstrate that adding node attributes helps recover the underlying community structure in content-rich networks more effectively than using links alone. In addition, we show that our method is faster and more accurate than alternative state-of-the-art algorithms.
Laura M. Smith, Linhong Zhu, Kristina Lerman, Allon G. Percus
ACM Trans. Knowl. Discov. Data2
2016 Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks
abstract
We propose a temporal latent space model for link prediction in dynamic social networks, where the goal is to predict links over time based on a sequence of previous graph snapshots. The model assumes that each user lies in an unobserved latent space, and interactions are more likely to occur between similar users in the latent space representation. In addition, the model allows each user to gradually move its position in the latent space as the network structure evolves over time. We present a global optimization algorithm to effectively infer the temporal latent space. Two alternative optimization algorithms with local and incremental updates are also proposed, allowing the model to scale to larger networks without compromising prediction accuracy. Empirically, we demonstrate that our model, when evaluated on a number of real-world dynamic networks, significantly outperforms existing approaches for temporal link prediction in terms of both scalability and predictive power.
Linhong Zhu, Junming Yin, Greg Ver Steeg, Aram Galstyan
IEEE Trans. Knowl. Data Eng.1
2015 Task matching and scheduling for multiple workers in spatial crowdsourcing
abstract
A new platform, termed spatial crowdsourcing, is emerging which enables a requester to commission workers to physically travel to some specified locations to perform a set of spatial tasks (i.e., tasks related to a geographical location and time). The current approach is to formulate spatial crowdsourcing as a matching problem between tasks and workers; hence the primary objective of the existing solutions is to maximize the number of matched tasks. Our goal is to solve the spatial crowdsourcing problem in the presence of multiple workers where we optimize for both travel cost and the number of completed tasks, while taking the tasks' expiration times into consideration. The challenge is that the solution should be a mixture of task-matching and task-scheduling, which are fundamentally different. In this paper, we show that a baseline approach that performs a task-matching first, and subsequently schedules the tasks assigned per worker in a following phase, does not perform well. Hence, we add a third phase in which we iterate back to the matching phase to improve the assignment per the output of the scheduling phase, and thus further improves the quality of matching and scheduling. Even though this 3-phase approach generates high quality results, it is very slow and does not scale. Hence, to scale our algorithm to large number of workers and tasks, we propose a Bisection-based framework which recursively divides all the workers and tasks into different partitions such that assignment and scheduling can be performed locally in a much smaller and promising space. Our experiments show that this approach is three orders of magnitude faster than the 3-phase approach while it only sacrifices 4% of the results' quality.
Dingxiong Deng, Cyrus Shahabi, Linhong Zhu
SIGSPATIAL/GIS3
2014 Tripartite graph clustering for dynamic sentiment analysis on social media
abstract
The growing popularity of social media (e.g., Twitter) allows users to easily share information with each other and influence others by expressing their own sentiments on various subjects. In this work, we propose an unsupervised tri-clustering framework, which analyzes both user-level and tweet-level sentiments through co-clustering of a tripartite graph. A compelling feature of the proposed framework is that the quality of sentiment clustering of tweets, users, and features can be mutually improved by joint clustering. We further investigate the evolution of user-level sentiments and latent feature vectors in an online framework and devise an efficient online algorithm to sequentially update the clustering of tweets, users and features with newly arrived data. The online framework not only provides better quality of both dynamic user-level and tweet-level sentiment analysis, but also improves the computational and storage efficiency. We verified the effectiveness and efficiency of the proposed approaches on the November 2012 California ballot Twitter data.
Linhong Zhu, Aram Galstyan, James Cheng, Kristina Lerman
SIGMOD Conference1
2013 Graph-based informative-sentence selection for opinion summarization
abstract
In this paper, we propose a new framework for opinion summarization based on sentence selection. Our goal is to assist users to get helpful opinion suggestions from reviews by only reading a short summary with few informative sentences, where the quality of summary is evaluated in terms of both aspect coverage and viewpoints preservation. More specifically, we formulate the informative-sentence selection problem in opinion summarization as a community-leader detection problem, where a community consists of a cluster of sentences towards the same aspect of an entity. The detected leaders of the communities can be considered as the most informative sentences of the corresponding aspect, while informativeness of a sentence is defined by its informativeness within both its community and the document it belongs to. Review data from six product domains from Amazon.com are used to verify the effectiveness of our method for opinion summarization.
Linhong Zhu, Sinno Jialin Pan, Haizhou Li 0001, Dingxiong Deng, Cyrus Shahabi
ASONAM1
2012 Fast algorithms for maximal clique enumeration with limited memory
abstract
Maximal clique enumeration (MCE) is a long-standing problem in graph theory and has numerous important applications. Though extensively studied, most existing algorithms become impractical when the input graph is too large and is disk-resident. We first propose an efficient partition-based algorithm for MCE that addresses the problem of processing large graphs with limited memory. We then further reduce the high cost of CPU computation of MCE by a careful nested partition based on a cost model. Finally, we parallelize our algorithm to further reduce the overall running time. We verified the efficiency of our algorithms by experiments in large real-world graphs.
James Cheng, Linhong Zhu, Yiping Ke, Shumo Chu
KDD2
2011 Detecting spam blogs from blog search results
Linhong Zhu, Aixin Sun, Byron Choi
Inf. Process. Manag.1
2011 Structure and attribute index for approximate graph matching in large graphs
Linhong Zhu, Wee Keong Ng, James Cheng
Inf. Syst.1
2011 Finding maximal cliques in massive networks
abstract
Maximal clique enumeration is a fundamental problem in graph theory and has important applications in many areas such as social network analysis and bioinformatics. The problem is extensively studied; however, the best existing algorithms require memory space linear in the size of the input graph. This has become a serious concern in view of the massive volume of today's fast-growing networks. We propose a general framework for designing external-memory algorithms for maximal clique enumeration in large graphs. The general framework enables maximal clique enumeration to be processed recursively in small subgraphs of the input graph, thus allowing in-memory computation of maximal cliques without the costly random disk access. We prove that the set of cliques obtained by the recursive local computation is both correct (i.e., globally maximal) and complete. The subgraph to be processed each time is defined based on a set of base vertices that can be flexibly chosen to achieve different purposes. We discuss the selection of the base vertices to fully utilize the available memory in order to minimize I/O cost in static graphs, and for update maintenance in dynamic graphs. We also apply our framework to design an external-memory algorithm for maximum clique computation in a large graph.
James Cheng, Yiping Ke, Ada Wai-Chee Fu, Jeffrey Xu Yu, Linhong Zhu
ACM Trans. Database Syst.5
2010 Finding maximal cliques in massive networks by H*-graph
abstract
Maximal clique enumeration (MCE) is a fundamental problem in graph theory and has important applications in many areas such as social network analysis and bioinformatics. The problem is extensively studied; however, the best existing algorithms require memory space linear in the size of the input graph. This has become a serious concern in view of the massive volume of today's fast-growing network graphs. Since MCE requires random access to different parts of a large graph, it is difficult to divide the graph into smaller parts and process one part at a time, because either the result may be incorrect and incomplete, or it incurs huge cost on merging the results from different parts. We propose a novel notion, H*-graph, which defines the core of a network and extends to encompass the neighborhood of the core for MCE computation. We propose the first external-memory algorithm for MCE (ExtMCE) that uses the H*-graph to bound the memory usage. We prove both the correctness and completeness of the result computed by ExtMCE. Extensive experiments verify that ExtMCE efficiently processes large networks that cannot be fit in the memory. We also show that the H*-graph captures important properties of the network; thus, updating the maximal cliques in the H*-graph retains the most essential information, with a low update cost, when it is infeasible to perform update on the entire network.
James Cheng, Yiping Ke, Ada Wai-Chee Fu, Jeffrey Xu Yu, Linhong Zhu
SIGMOD Conference5
2009 A Uniform Framework for Ad-Hoc Indexes to Answer Reachability Queries on Large Graphs
Linhong Zhu, Byron Choi, Bingsheng He, Jeffrey Xu Yu, Wee Keong Ng
DASFAA1
2008 Online spam-blog detection through blog search
abstract
In this work, we propose a novel post-indexing spam-blog (or splog) detection method, which capitalizes on the re-sults returned by blog search engines. More specifically, we analyze the search results of a sequence of temporally-ordered queries returned by a blog search engine, and build and maintain blog profiles for those blogs whose posts fre-quently appear in the top-ranked search results. With the blog profiles, 4 splog scoring functions were evaluated us-ing real data collected from a popular blog search engine. Our experiments show that the proposed method could ef-fectively detect splogs with a high accuracy.
Linhong Zhu, Aixin Sun, Byron Choi
CIKM1