VLDB 2026 Research / reviewers in the wild / expert
Jiashun Jin
dblp:56/2406
· DBLP profile ↗
11ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0002-7442-1962ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
5 papers |
Computational complexity · 25% Graph algorithms and graph theory · 22% Mathematical optimization · 19% | |
| Databases, data mining, and information retrieval
3 papers |
Data mining · 65% Data models and query languages · 21% Recommender systems · 14% | |
| Artificial intelligence
3 papers |
Learning theory · 90% Optimization for machine learning · 10% |
Topics — the 25 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › structured data mining › graph mining
community detection |
1.2 | 2 | 2025 | Fitting Networks with a Cancellation Trick · ICLR 2025 Network Global Testing by Counting Graphlets · ICML 2018 |
Graph algorithms and graph theory › graph clustering
community detection |
1.2 | 2 | 2023 | Phase transition for detecting a small community in a large network · ICLR 2023 Sharp Impossibility Results for Hyper-graph Testing · NeurIPS 2021 |
Data models and query languages › data modeling
network data model |
0.9 | 1 | 2025 | Fitting Networks with a Cancellation Trick · ICLR 2025 |
Mathematical optimization › convergence analysis
non-asymptotic bounds |
0.8 | 1 | 2024 | Improved algorithm and bounds for successive projection · ICLR 2024 |
Computational complexity
phase transition |
0.7 | 1 | 2023 | Phase transition for detecting a small community in a large network · ICLR 2023 |
Computational complexity
statistical-computational gaps |
0.7 | 1 | 2023 | Phase transition for detecting a small community in a large network · ICLR 2023 |
Recommender systems › collaborative filtering
matrix factorization |
0.6 | 1 | 2022 | A sharp NMF result with applications in network modeling · NeurIPS 2022 |
Data mining
network modeling |
0.6 | 1 | 2022 | A sharp NMF result with applications in network modeling · NeurIPS 2022 |
Data mining › dimensionality reduction
nonnegative matrix factorization |
0.6 | 1 | 2022 | A sharp NMF result with applications in network modeling · NeurIPS 2022 |
Machine learning › Learning theory › computational learning theory
impossibility result |
0.5 | 1 | 2021 | Sharp Impossibility Results for Hyper-graph Testing · NeurIPS 2021 |
Combinatorics and discrete mathematics
hypergraph |
0.5 | 1 | 2021 | Sharp Impossibility Results for Hyper-graph Testing · NeurIPS 2021 |
Information theory
hypothesis testing |
0.5 | 1 | 2021 | Sharp Impossibility Results for Hyper-graph Testing · NeurIPS 2021 |
Data mining › structured data mining
graph mining |
0.3 | 1 | 2018 | Network Global Testing by Counting Graphlets · ICML 2018 |
Mathematical optimization
statistical estimation |
0.3 | 1 | 2025 | Fitting Networks with a Cancellation Trick · ICLR 2025 |
Bioinformatics and computational biology
statistical genetics |
0.2 | 1 | 2016 | Component-wise gradient boosting and false discovery control in survival analysis with high-dimensional covariates · Bioinform. 2016 |
Machine learning › Learning theory
high-dimensional statistics |
0.2 | 1 | 2014 | Optimality of graphlet screening in high dimensional variable selection · J. Mach. Learn. Res. 2014 |
Machine learning › Learning theory
minimax optimality |
0.2 | 1 | 2014 | Optimality of graphlet screening in high dimensional variable selection · J. Mach. Learn. Res. 2014 |
Machine learning › Learning theory › model selection
variable selection |
0.2 | 1 | 2014 | Optimality of graphlet screening in high dimensional variable selection · J. Mach. Learn. Res. 2014 |
Machine learning › Learning theory
high-dimensional regression |
0.1 | 1 | 2012 | A Comparison of the Lasso and Marginal Regression · J. Mach. Learn. Res. 2012 |
Machine learning › Optimization for machine learning › regularized risk minimization › regularized regression
lasso |
0.1 | 1 | 2012 | A Comparison of the Lasso and Marginal Regression · J. Mach. Learn. Res. 2012 |
Machine learning › Learning theory
statistical estimation |
0.1 | 1 | 2012 | A Comparison of the Lasso and Marginal Regression · J. Mach. Learn. Res. 2012 |
Bioinformatics and computational biology › statistical genetics
genetic association study |
0.1 | 1 | 2016 | Component-wise gradient boosting and false discovery control in survival analysis with high-dimensional covariates · Bioinform. 2016 |
Privacy and data protection
privacy metrics |
0.0 | 1 | 2004 | When do data mining results violate privacy? · KDD 2004 |
Privacy and data protection
privacy-preserving data analysis |
0.0 | 1 | 2004 | When do data mining results violate privacy? · KDD 2004 |
Cryptographic protocols and secure computation
secure multiparty computation |
0.0 | 1 | 2004 | When do data mining results violate privacy? · KDD 2004 |
Methods — techniques the papers use, named apart from their topics
spectral methods · 1.7recursive algorithm · 1.7cancellation trick · 1.7spectral analysis · 1.1degree matching · 1.0successive projection algorithm · 0.8projection · 0.8denoising · 0.8random graph analysis · 0.7nonconvex optimization · 0.6non-convex optimization · 0.6tensor scaling · 0.5test statistic · 0.3short path and cycle counting · 0.3degree heterogeneity correction · 0.3stability selection · 0.2random permutation · 0.2component-wise gradient boosting · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STAR-Net: an interpretable model-aided network for remote sensing image denoising
Jingjing Liu 0004, Jiashun Jin, Xianchao Xiu, Wanquan Liu |
Pattern Recognit. | 2 |
| 2025 | Fitting Networks with a Cancellation TrickabstractThe degree-corrected block model (DCBM), latent space model (LSM), and $\beta$-model are all popular network models. We combine their modeling ideas and propose the logit-DCBM as a new model. Similar as the $\beta$-model and LSM, the logit-DCBM contains nonlinear factors, where fitting the parameters is a challenging open problem. We resolve this problem by introducing a cancellation trick. We also propose R-SCORE as a recursive community detection algorithm, where in each iteration, we first use the idea above to update our parameter estimation, and then use the results to remove the nonlinear factors in the logit-DCBM so the renormalized model approximately satisfies a low-rank model, just like the DCBM. Our numerical study suggests that R-SCORE significantly improves over existing spectral approaches in many cases. Also, theoretically, we show that the Hamming error rate of R-SCORE is faster than that of SCORE in a specific sparse region, and is at least as fast outside this region. Jiashun Jin, Jingming Wang |
ICLR | 1 |
| 2024 | Improved algorithm and bounds for successive projectionabstractConsider a $K$-vertex simplex in a $d$-dimensional space. We measure $n$ points on the simplex, but due to the measurement noise,
some of the observed points fall outside the simplex. The interest is vertex hunting (i.e., estimating the vertices of the simplex). The successive projection algorithm (SPA) is one of the most popular approaches to vertex hunting, but it is vulnerable to noise and outliers, and may perform unsatisfactorily. We propose pseudo-point SPA (pp-SPA) as a new approach to vertex hunting. The approach contains
two novel ideas (a projection step and a denoise step) and generates roughly $n$ pseudo-points, which can be fed in to SPA for vertex hunting. For theory, we first derive an improved non-asymptotic bound for the orthodox SPA, and then use the result to derive the bounds for pp-SPA. Compared with the orthodox SPA, pp-SPA has a faster rate and more satisfactory numerical performance in a broad setting. The analysis is quite delicate: the non-asymptotic bound is hard to derive, and we need precise results on the extreme values of (possibly) high-dimensional random vectors. Jiashun Jin, Zheng Tracy Ke, Gabriel Moryoussef, Jingming Wang |
ICLR | 1 |
| 2023 | Phase transition for detecting a small community in a large network
Jiashun Jin, Zheng Tracy Ke, Paxton Turner, Anru Zhang |
ICLR | 1 |
| 2022 | A sharp NMF result with applications in network modelingabstractGiven an $n \times n$ non-negative rank-$K$ matrix $\Omega$ where $m$ eigenvalues are negative, when can we write $\Omega = Z P Z'$ for non-negative matrices $Z \in \mathbb{R}^{n, K}$ and $P \in \mathbb{R}^{K, K}$? While most existing works focused on the case of $m = 0$, our primary interest is on the case of general $m$. With new proof ideas we develop, we present sharp results on when the NMF problem is solvable, which significantly extend existing results on this topic. The NMF problem is partially motivated by applications in network modeling. For a network with $K$ communities, rank-$K$ models are popular, with many proposals. The DCMM model is a recent rank-$K$ model which is especially useful and interpretable in practice. To enjoy such properties, it is of interest to study when a rank-$K$ model can be rewritten as a DCMM model. Using our NMF results, we show that for a rank-$K$ model with parameters in the most interesting range, we can always rewrite it as a DCMM model. Jiashun Jin |
NeurIPS | 1 |
| 2021 | Sharp Impossibility Results for Hyper-graph TestingabstractIn a broad Degree-Corrected Mixed-Membership (DCMM) setting, we test whether a non-uniform hypergraph has only one community or has multiple communities. Since both the null and alternative hypotheses have many unknown parameters, the challenge is, given an alternative, how to identify the null that is hardest to separate from the alternative. We approach this by proposing a degree matching strategy where the main idea is leveraging the theory for tensor scaling to create a least favorable pair of hypotheses. We present a result on standard minimax lower bound theory and a result on Region of Impossibility (which is more informative than the minimax lower bound). We show that our lower bounds are tight by introducing a new test that attains the lower bound up to a logarithmic factor. We also discuss the case where the hypergraphs may have mixed-memberships. Jiashun Jin, Zheng Tracy Ke, Jiajun Liang |
NeurIPS | 1 |
| 2018 | Network Global Testing by Counting GraphletsabstractConsider a large social network with possibly severe degree heterogeneity and mixed-memberships. We are interested in testing whether the network has only one community or there are more than one communities. The problem is known to be non-trivial, partially due to the presence of severe degree heterogeneity. We construct a class of test statistics using the numbers of short paths and short cycles, and the key to our approach is a general framework for canceling the effects of degree heterogeneity. The tests compare favorably with existing methods. We support our methods with careful analysis and numerical study with simulated data and a real data example. Jiashun Jin, Zheng Tracy Ke, Shengming Luo |
ICML | 1 |
| 2016 | Component-wise gradient boosting and false discovery control in survival analysis with high-dimensional covariatesabstractMOTIVATION: Technological advances that allow routine identification of high-dimensional risk factors have led to high demand for statistical techniques that enable full utilization of these rich sources of information for genetics studies. Variable selection for censored outcome data as well as control of false discoveries (i.e. inclusion of irrelevant variables) in the presence of high-dimensional predictors present serious challenges. This article develops a computationally feasible method based on boosting and stability selection. Specifically, we modified the component-wise gradient boosting to improve the computational feasibility and introduced random permutation in stability selection for controlling false discoveries. RESULTS: We have proposed a high-dimensional variable selection method by incorporating stability selection to control false discovery. Comparisons between the proposed method and the commonly used univariate and Lasso approaches for variable selection reveal that the proposed method yields fewer false discoveries. The proposed method is applied to study the associations of 2339 common single-nucleotide polymorphisms (SNPs) with overall survival among cutaneous melanoma (CM) patients. The results have confirmed that BRCA2 pathway SNPs are likely to be associated with overall survival, as reported by previous literature. Moreover, we have identified several new Fanconi anemia (FA) pathway SNPs that are likely to modulate survival of CM patients. AVAILABILITY AND IMPLEMENTATION: The related source code and documents are freely available at https://sites.google.com/site/bestumich/issues. CONTACT: [email protected]. Kevin He, Jeffrey E. Lee, Christopher I. Amos, Terry Hyslop, Jiashun Jin, Huazhen Lin, Qinyi Wei, Yi Li 0019 |
Bioinform. | 8 |
| 2014 | Optimality of graphlet screening in high dimensional variable selection
Jiashun Jin, Cun-Hui Zhang |
J. Mach. Learn. Res. | 1 |
| 2012 | A Comparison of the Lasso and Marginal Regression
Christopher R. Genovese, Jiashun Jin, Larry A. Wasserman, Zhigang Yao |
J. Mach. Learn. Res. | 2 |
| 2004 | When do data mining results violate privacy?abstractPrivacy-preserving data mining has concentrated on obtaining valid results when the input data is private. An extreme example is Secure Multiparty Computation-based methods, where only the results are revealed. However, this still leaves a potential privacy breach: Do the results themselves violate privacy? This paper explores this issue, developing a framework under which this question can be addressed. Metrics are proposed, along with analysis that those metrics are consistent in the face of apparent problems. Murat Kantarcioglu, Jiashun Jin, Chris Clifton |
KDD | 2 |