Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Victor E. Lee

dblp:86/7123 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 2 since 2021Artificial intelligence and machine learning · 6 · 1 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Data models and query languages · 31% Query processing and optimization · 19% Transaction processing and concurrency control · 16%
Artificial intelligence
2 papers
Graph learning · 92% Probabilistic and Bayesian machine learning · 8%
Theoretical computer science
3 papers
Graph algorithms and graph theory · 96% Mathematical optimization · 4%
Network and information security
1 paper
Privacy and data protection · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 17 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
0.612022
Efficient Machine Learning on Large-Scale Graphs · KDD 2022
Machine learning › Graph learning › efficient graph learning
large-scale graph learning
0.612022
Efficient Machine Learning on Large-Scale Graphs · KDD 2022
Graph algorithms and graph theory
network analysis
0.612022
Efficient Machine Learning on Large-Scale Graphs · KDD 2022
Transaction processing and concurrency control
data integrity
0.512021
PG-Keys: Keys for Property Graphs · SIGMOD Conference 2021
Query processing and optimization
aggregation
0.412020
Aggregation Support for Modern Graph Analytics in TigerGraph · SIGMOD Conference 2020
Graph data management
graph analytics
0.412020
Aggregation Support for Modern Graph Analytics in TigerGraph · SIGMOD Conference 2020
Data models and query languages
graph query language
0.412020
Aggregation Support for Modern Graph Analytics in TigerGraph · SIGMOD Conference 2020
Recommender systems
collaborative filtering
0.212014
k-CoRating: Filling Up Data to Obtain Privacy and Utility · AAAI 2014
Privacy and data protection
anonymization
0.212014
k-CoRating: Filling Up Data to Obtain Privacy and Utility · AAAI 2014
Privacy and data protection › anonymization
k-anonymity
0.212014
k-CoRating: Filling Up Data to Obtain Privacy and Utility · AAAI 2014
Query processing and optimization
similarity query processing
0.112012
A highway-centric labeling approach for answering distance queries on large sparse graphs · SIGMOD Conference 2012
Web and social media mining
social network analysis
0.112011
Axiomatic ranking of network role similarity · KDD 2011
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
time series causal discovery
0.112009
A Sparsification Approach for Temporal Graphical Model Decomposition · ICDM 2009
Data mining
clustering
0.112009
A Sparsification Approach for Temporal Graphical Model Decomposition · ICDM 2009
Data mining › spatiotemporal data mining
spatio-temporal pattern mining
0.112009
Migration motif: a spatial - temporal pattern mining approach for financial markets · KDD 2009
Graph data management
graph indexing
0.012012
A highway-centric labeling approach for answering distance queries on large sparse graphs · SIGMOD Conference 2012
Mathematical optimization
continuous optimization
0.012009
A Sparsification Approach for Temporal Graphical Model Decomposition · ICDM 2009

Methods — techniques the papers use, named apart from their topics

pagerank · 1.7graph neural network · 1.7embedding · 1.7constraint specification · 0.5rating prediction · 0.4quasi-newton method · 0.3maximum weight independent set · 0.3generalized ridge regression · 0.3iterative computation · 0.2axiomatic analysis · 0.2bipartite set cover · 0.1spatial-temporal pattern mining · 0.1
YearPublicationVenuePosition
2022 Efficient Machine Learning on Large-Scale Graphs
abstract
Machine learning on graph data has become a common area of interest across academia and industry. However, due to the size of real-world industry graphs (hundreds of millions of vertices and billions of edges) and the special architecture of graph neural net- works, it is still a challenge for practitioners and researchers to perform machine learning tasks on large-scale graph data. It typi- cally takes a powerful and expensive GPU machine to train a graph neural network on a million-vertex scale graph, let alone doing deep learning on real enterprise graphs. In this tutorial, we will cover how to develop and run performant graph algorithms and graph neural network models with TigerGraph [3], a massively parallel platform for graph analytics, and its Machine Learning Workbench with PyTorch Geometric [4] and DGL [8] support. Using an NFT transaction dataset [6], we will first investigate transactions using graph algorithms by themselves as methods of graph traversing, clustering, classification, and determining similarities between data. Secondly, we will show how to use those graph-derived features such as PageRank and embeddings to empower traditional machine learning models. Finally, we will demonstrate how to train common graph neural networks with TigerGraph and how to implement novel graph neural network models. Participants will use the Tiger- Graph ML Workbench Cloud to perform graph feature engineering and train their machine learning algorithms during the session.
Parker Erickson, Victor E. Lee, Jiliang Tang
KDD2
2021 PG-Keys: Keys for Property Graphs
abstract
We report on a community effort between industry and academia to shape the future of property graph constraints. The standardization for a property graph query language is currently underway through the ISO Graph Query Language (GQL) project. Our position is that this project should pay close attention to schemas and constraints, and should focus next on key constraints. The main purposes of keys are enforcing data integrity and allowing the referencing and identifying of objects. Motivated by use cases from our industry partners, we argue that key constraints should be able to have different modes, which are combinations of basic restriction that require the key to be exclusive, mandatory, and singleton. Moreover, keys should be applicable to nodes, edges, and properties since these all can represent valid real-life entities. Our result is PG-Keys, a flexible and powerful framework for defining key constraints, which fulfills the above goals. PG-Keys is a design by the Linked Data Benchmark Council's Property Graph Schema Working Group, consisting of members from industry, academia, and ISO GQL standards group, intending to bring the best of all worlds to property graph practitioners. PG-Keys aims to guide the evolution of the standardization efforts towards making systems more useful, powerful, and expressive.
Renzo Angles, Angela Bonifati, Stefania Dumbrava, George Fletcher 0001, Keith W. Hare, Jan Hidders, Victor E. Lee, Leonid Libkin, Wim Martens, Filip Murlak, Josh Perryman, Ognjen Savkovic, Michael Schmidt 0002, Juan F. Sequeda, Slawomir Staworko, Dominik Tomaszuk
SIGMOD Conference7
2020 Aggregation Support for Modern Graph Analytics in TigerGraph
abstract
We describe how GSQL, TigerGraph's graph query language, supports the specification of aggregation in graph analytics. GSQL makes several unique design decisions with respect to both the expressive power and the evaluation complexity of the specified aggregation. We detail our design showing how our ideas transcend GSQL and are eminently portable to the upcoming graph query language standards as well as the existing pattern-based declarative query languages.
Alin Deutsch, Mingxi Wu, Victor E. Lee
SIGMOD Conference4
2019 Privacy-aware smart city: A case study in collaborative filtering recommender systems
Feng Zhang 0012, Victor E. Lee, Ruoming Jin, Saurabh Kumar Garg 0001, Kim-Kwang Raymond Choo, Michele Maasberg, Lijun Dong, Chi Cheng 0003
J. Parallel Distributed Comput.2
2018 Jo-DPMF: Differentially private matrix factorization learning through joint optimization
Feng Zhang 0012, Victor E. Lee, Kim-Kwang Raymond Choo
Inf. Sci.2
2016 Fast algorithms to evaluate collaborative filtering recommender systems
Feng Zhang 0012, Ti Gong, Victor E. Lee, Gansen Zhao, Chunming Rong, Guangzhi Qu
Knowl. Based Syst.3
2015 Simple is Beautiful: An Online Collaborative Filtering Recommendation Solution with Higher Accuracy
Feng Zhang 0012, Ti Gong, Victor E. Lee, Gansen Zhao, Guangzhi Qu
APWeb3
2015 Fast and Accurate Computation of Role Similarity via Vertex Centrality
Longjie Li 0001, Lvjian Qian, Victor E. Lee, Mingwei Leng
WAIM3
2014 k-CoRating: Filling Up Data to Obtain Privacy and Utility
abstract
For datasets in Collaborative Filtering (CF) recommendations, even if the identifier is deleted and some trivial perturbation operations are applied to ratings before they are released, there are research results claiming that the adversary could discriminate the individual's identity with a little bit of information. In this paper, we propose $k$-coRating, a novel privacy-preserving model, to retain data privacy by replacing some null ratings with "well-predicted" scores. They do not only mask the original ratings such that a $k$-anonymity-like data privacy is preserved, but also enhance the data utility (measured by prediction accuracy in this paper), which shows that the traditional assumption that accuracy and privacy are two goals in conflict is not necessarily correct. We show that the optimal $k$-coRated mapping is an NP-hard problem and design a naive but efficient algorithm to achieve $k$-coRating. All claims are verified by experimental results.
Feng Zhang 0012, Victor E. Lee, Ruoming Jin
AAAI2
2014 Scalable and axiomatic ranking of network role similarity
abstract
A key task in analyzing social networks and other complex networks is role analysis: describing and categorizing nodes according to how they interact with other nodes. Two nodes have the same role if they interact with equivalent sets of neighbors. The most fundamental role equivalence is automorphic equivalence. Unfortunately, the fastest algorithms known for graph automorphism are nonpolynomial. Moreover, since exact equivalence is rare, a more meaningful task is measuring the role similarity between any two nodes. This task is closely related to the structural or link-based similarity problem that SimRank addresses. However, SimRank and other existing similarity measures are not sufficient because they do not guarantee to recognize automorphically or structurally equivalent nodes. This article makes two contributions. First, we present and justify several axiomatic properties necessary for a role similarity measure or metric. Second, we present RoleSim, a new similarity metric that satisfies these axioms and can be computed with a simple iterative algorithm. We rigorously prove that RoleSim satisfies all of these axiomatic properties. We also introduce Iceberg RoleSim, a scalable algorithm that discovers all pairs with RoleSim scores above a user-defined threshold θ. We demonstrate the interpretative power of RoleSim on both both synthetic and real datasets.
Ruoming Jin, Victor E. Lee, Longjie Li 0001
ACM Trans. Knowl. Discov. Data2
2012 A highway-centric labeling approach for answering distance queries on large sparse graphs
abstract
The distance query, which asks the length of the shortest path from a vertex $u$ to another vertex v, has applications ranging from link analysis, semantic web and other ontology processing, to social network operations. Here, we propose a novel labeling scheme, referred to as Highway-Centric Labeling, for answering distance queries in a large sparse graph. It empowers the distance labeling with a highway structure and leverages a novel bipartite set cover framework/algorithm. Highway-centric labeling provides better labeling size than the state-of-the-art $2$-hop labeling, theoretically and empirically. It also offers both exact distance and approximate distance with bounded accuracy. A detailed experimental evaluation on both synthetic and real datasets demonstrates that highway-centric labeling can outperform the state-of-the-art distance computation approaches in terms of both index size and query time.
Ruoming Jin, Ning Ruan, Yang Xiang 0007, Victor E. Lee
SIGMOD Conference4
2011 Axiomatic ranking of network role similarity
abstract
A key task in analyzing social networks and other complex networks is role analysis: describing and categorizing nodes by how they interact with other nodes. Two nodes have the same role if they interact with equivalent sets of neighbors. The most fundamental role equivalence is automorphic equivalence. Unfortunately, the fastest algorithm known for graph automorphism is nonpolynomial. Moreover, since exact equivalence is rare, a more meaningful task is measuring the role similarity between any two nodes. This task is closely related to the link-based similarity problem that SimRank addresses. However, SimRank and other existing simliarity measures are not sufficient because they do not guarantee to recognize automorphically or structurally equivalent nodes. This paper makes two contributions. First, we present and justify several axiomatic properties necessary for a role similarity measure or metric. Second, we present RoleSim, a role similarity metric which satisfies these axioms and which can be computed with a simple iterative algorithm. We rigorously prove that RoleSim satisfies all the axiomatic properties and demonstrate its superior interpretative power on both synthetic and real datasets.
Ruoming Jin, Victor E. Lee, Hui Hong
KDD2
2009 A Sparsification Approach for Temporal Graphical Model Decomposition
abstract
Temporal causal modeling can be used to recover the causal structure among a group of relevant time series variables. Several methods have been developed to explicitly construct temporal causal graphical models. However, how to best understand and conceptualize these complicated causal relationships is still an open problem. In this paper, we propose a decomposition approach to simplify the temporal graphical model. Our method clusters time series variables into groups such that strong interactions appear among the variables within each group and weak (or no) interactions exist for cross-group variable pairs. Specifically, we formulate the clustering problem for temporal graphical models as a regression-coefficient sparsification problem and define an interesting objective function which balances the model prediction power and its cluster structure. We introduce an iterative optimization approach utilizing the Quasi-Newton method and generalized ridge regression to minimize the objective function and to produce a clustered temporal graphical model. We also present a novel optimization procedure utilizing a graph theoretical tool based on the maximum weight independent set problem to speed up the Quasi-Newton method for a large number of variables. Finally, our detailed experimental study on both synthetic and real datasets demonstrates the effectiveness of our methods.
Ning Ruan, Ruoming Jin, Victor E. Lee
ICDM3
2009 Migration motif: a spatial - temporal pattern mining approach for financial markets
abstract
A recent study by two prominent finance researchers, Fama and French, introduces a new framework for studying risk vs. return: the migration of stocks across size-value portfolio space. Given the financial events of 2008, this first attempt to disentangle the relationships between migration behavior and stock returns is especially timely. Their work, however, derives results only for market segments, not individual companies, and only for one-year moves. Thus, we see a new challenge for financial data mining: how to capture and categorize the migration of individual companies, and how such behavior affects their returns.
Xiaoxi Du, Ruoming Jin, Victor E. Lee, John H. Thornton Jr.
KDD4