Yuetian Sun

dblp:304/5069 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2024
0009-0000-4094-1380ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Graph data management · 100%
Theoretical computer science
1 paper
Algorithms and data structures · 50% Graph algorithms and graph theory · 50%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Graph data management
graph similarity
0.712023
Efficient and Accurate SimRank-based Similarity Joins: Experiments, Analysis, and Improvement · Proc. VLDB Endow. 2023
Graph data management › graph similarity
simrank
0.712023
Efficient and Accurate SimRank-based Similarity Joins: Experiments, Analysis, and Improvement · Proc. VLDB Endow. 2023
Graph algorithms and graph theory › centrality › pagerank
personalized pagerank
0.712023
Efficient and Accurate SimRank-based Similarity Joins: Experiments, Analysis, and Improvement · Proc. VLDB Endow. 2023
Algorithms and data structures › similarity search
similarity join
0.712023
Efficient and Accurate SimRank-based Similarity Joins: Experiments, Analysis, and Improvement · Proc. VLDB Endow. 2023

Methods — techniques the papers use, named apart from their topics

randomized local push · 1.3approximation algorithm · 1.3
YearPublicationVenuePosition
2024 Multilevel Stochastic Optimization for Imputation in Massive Medical Data Records
abstract
It has long been a recognized problem that many datasets contain significant levels of missing numerical data. A potentially critical predicate for application of machine learning methods to datasets involves addressing this problem. However, this is a challenging task. In this paper, we apply a recently developed multi-level stochastic optimization approach to the problem of imputation in massive medical records. The approach is based on computational applied mathematics techniques and is highly accurate. In particular, for the Best Linear Unbiased Predictor (BLUP) this multi-level formulation is exact, and is significantly faster and more numerically stable. This permits practical application of Kriging methods to data imputation problems for massive datasets. We test this approach on data from the National Inpatient Sample (NIS) data records, Healthcare Cost and Utilization Project (HCUP), Agency for Healthcare Research and Quality. Numerical results show that the multi-level method significantly outperforms current approaches and is numerically robust. It has superior accuracy as compared with methods recommended in the recent report from HCUP. Benchmark tests show up to 75% reductions in error. Furthermore, the results are also superior to recent state of the art methods such as discriminative deep learning
Yuetian Sun, Snezana Milanovic, Mark Kon, Julio Enrique Castrillón-Candás
IEEE Trans. Big Data3
2023 Efficient and Accurate SimRank-based Similarity Joins: Experiments, Analysis, and Improvement
abstract
SimRank-based similarity joins, which mainly include threshold-based and top- k similarity joins, are important types of all-pair SimRank queries. Although a line of related algorithms have been proposed recently, they still fall short of providing approximation guarantee and suffer from scalability issues on medium and large graphs. Meanwhile, we also lack an extensive analysis of existing techniques in terms of accuracy and efficiency. Motivated by these challenges, we first conduct detailed analysis of state-of-the-art algorithms and provide additional theoretical results. Second, to address the limitations of existing techniques, we propose simple yet effective algorithm frameworks for both queries to theoretically guarantee the approximation bound, and present a more efficient all-pair algorithm inspired by randomized local push of Personalized PageRank. Next, we analyze the algorithmic complexity of threshold-based and top- k similarity joins by leveraging a reasonable assumption of SimRank distribution. Through extensive experiments, we find that our proposed methods far exceed existing ones with respect to query efficiency, approximation guarantee and practical accuracy, while our theoretical analysis nicely matches the empirical study.
Yu Liu 0070, Yuetian Sun, Lei Zou 0001, Yuxing Chen 0003, Anqun Pan
Proc. VLDB Endow.4