Jerry Jiang

dblp:185/4257 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 42% Recommender systems · 30% Graph data management · 24%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › structured data mining › graph mining
community detection
0.412020
SimClusters: Community-Based Representations for Heterogeneous Recommendations at Twitter · KDD 2020
Data mining › structured data mining › graph mining › community detection
overlapping community detection
0.412020
SimClusters: Community-Based Representations for Heterogeneous Recommendations at Twitter · KDD 2020
Recommender systems
content recommendation
0.212016
GraphJet: Real-Time Content Recommendations at Twitter · Proc. VLDB Endow. 2016
Recommender systems
graph-based recommendation
0.212016
GraphJet: Real-Time Content Recommendations at Twitter · Proc. VLDB Endow. 2016
Graph data management › graph processing
graph processing systems
0.212016
GraphJet: Real-Time Content Recommendations at Twitter · Proc. VLDB Endow. 2016
Graph data management › graph processing › graph processing systems
in-memory graph processing
0.212016
GraphJet: Real-Time Content Recommendations at Twitter · Proc. VLDB Endow. 2016
Recommender systems
heterogeneous recommendation
0.112020
SimClusters: Community-Based Representations for Heterogeneous Recommendations at Twitter · KDD 2020

Methods — techniques the papers use, named apart from their topics

sparse vector representation · 0.4metropolis-hastings sampling · 0.4random walk · 0.2bipartite graph · 0.2
YearPublicationVenuePosition
2020 SimClusters: Community-Based Representations for Heterogeneous Recommendations at Twitter
abstract
Personalized recommendation products at Twitter target a multitude of heterogeneous items: Tweets, Events, Topics, Hashtags, and users. Each of these targets varies in their cardinality (which affects the scale of the problem) and their "shelf life'' (which constrains the latency of generating the recommendations). Although Twitter has built a variety of recommendation systems before dating back a decade, solutions to the broader problem were mostly tackled piecemeal. In this paper, we present SimClusters, a general-purpose representation layer based on overlapping communities into which users as well as heterogeneous content can be captured as sparse, interpretable vectors to support a multitude of recommendation tasks. We propose a novel algorithm for community discovery based on Metropolis-Hastings sampling, which is both more accurate and significantly faster than off-the-shelf alternatives. SimClusters scales to networks with billions of users and has been effective across a variety of deployed applications at Twitter.
Venu Satuluri, Xun Zheng, Yilei Qian, Brian Wichers, Qieyun Dai, Gui Ming Tang, Jerry Jiang, Jimmy Lin
KDD8
2016 GraphJet: Real-Time Content Recommendations at Twitter
abstract
This paper presents GraphJet, a new graph-based system for generating content recommendations at Twitter. As motivation, we trace the evolution of our formulation and approach to the graph recommendation problem, embodied in successive generations of systems. Two trends can be identified: supplementing batch with real-time processing and a broadening of the scope of recommendations from users to content. Both of these trends come together in Graph-Jet, an in-memory graph processing engine that maintains a real-time bipartite interaction graph between users and tweets. The storage engine implements a simple API, but one that is sufficiently expressive to support a range of recommendation algorithms based on random walks that we have refined over the years. Similar to Cassovary, a previous graph recommendation engine developed at Twitter, GraphJet assumes that the entire graph can be held in memory on a single server. The system organizes the interaction graph into temporally-partitioned index segments that hold adjacency lists. GraphJet is able to support rapid ingestion of edges while concurrently serving lookup queries through a combination of compact edge encoding and a dynamic memory allocation scheme that exploits power-law characteristics of the graph. Each GraphJet server ingests up to one million graph edges per second, and in steady state, computes up to 500 recommendations per second, which translates into several million edge read operations per second.
Aneesh Sharma, Jerry Jiang, Praveen Bommannavar, Brian Larson, Jimmy Lin
Proc. VLDB Endow.2