EDBT 2026 Demo / reviewers in the wild / expert
Junming Yin
dblp:53/7065
· DBLP profile ↗
15ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0001-6018-7813ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 30% Probabilistic and Bayesian machine learning · 29% Graph learning · 26% | |
| Databases, data mining, and information retrieval
4 papers |
Data mining · 73% Knowledge graphs · 18% Graph data management · 9% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 30 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-armed bandit |
0.6 | 1 | 2022 | Thompson Sampling for Bandit Learning in Matching Markets · IJCAI 2022 |
Machine learning › Reinforcement learning
thompson sampling |
0.6 | 1 | 2022 | Thompson Sampling for Bandit Learning in Matching Markets · IJCAI 2022 |
Algorithmic game theory and mechanism design › market design › matching markets
bandit learning in matching markets |
0.6 | 1 | 2022 | Thompson Sampling for Bandit Learning in Matching Markets · IJCAI 2022 |
Algorithmic game theory and mechanism design › market design
matching markets |
0.6 | 1 | 2022 | Thompson Sampling for Bandit Learning in Matching Markets · IJCAI 2022 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
stochastic variational inference |
0.4 | 2 | 2016 | Latent Space Inference of Internet-Scale Networks · J. Mach. Learn. Res. 2016 A Scalable Approach to Probabilistic Latent Space Inference of Large-Scale Networks · NIPS 2013 |
Data mining › text mining › topic modeling
hierarchical dirichlet process |
0.3 | 1 | 2018 | Supervised Topic Modeling Using Hierarchical Dirichlet Process-Based Inverse Regression: Experiments on E-Commerce Applications · IEEE Trans. Knowl. Data Eng. 2018 |
Data mining › text mining › topic modeling
supervised topic modeling |
0.3 | 1 | 2018 | Supervised Topic Modeling Using Hierarchical Dirichlet Process-Based Inverse Regression: Experiments on E-Commerce Applications · IEEE Trans. Knowl. Data Eng. 2018 |
Data mining
text mining |
0.3 | 1 | 2018 | Supervised Topic Modeling Using Hierarchical Dirichlet Process-Based Inverse Regression: Experiments on E-Commerce Applications · IEEE Trans. Knowl. Data Eng. 2018 |
Data mining › text mining
topic modeling |
0.3 | 1 | 2018 | Supervised Topic Modeling Using Hierarchical Dirichlet Process-Based Inverse Regression: Experiments on E-Commerce Applications · IEEE Trans. Knowl. Data Eng. 2018 |
Graph data management
dynamic graph processing |
0.3 | 1 | 2017 | Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks (Extended Abstract) · ICDE 2017 |
Data mining › representation learning
latent space models |
0.3 | 1 | 2017 | Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks (Extended Abstract) · ICDE 2017 |
Knowledge graphs
link prediction |
0.3 | 1 | 2017 | Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks (Extended Abstract) · ICDE 2017 |
Knowledge graphs › link prediction
temporal link prediction |
0.3 | 1 | 2017 | Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks (Extended Abstract) · ICDE 2017 |
Machine learning › Graph learning › dynamic graph learning
dynamic link prediction |
0.2 | 1 | 2016 | Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks · IEEE Trans. Knowl. Data Eng. 2016 |
Machine learning › Generative modeling
latent space model |
0.2 | 1 | 2016 | Latent Space Inference of Internet-Scale Networks · J. Mach. Learn. Res. 2016 |
Machine learning › Graph learning
link prediction |
0.2 | 1 | 2016 | Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks · IEEE Trans. Knowl. Data Eng. 2016 |
Machine learning › Graph learning
network embedding |
0.2 | 1 | 2016 | Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks · IEEE Trans. Knowl. Data Eng. 2016 |
Machine learning › Graph learning › graph clustering
overlapping community detection |
0.2 | 1 | 2016 | Latent Space Inference of Internet-Scale Networks · J. Mach. Learn. Res. 2016 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
scalable inference |
0.2 | 1 | 2016 | Latent Space Inference of Internet-Scale Networks · J. Mach. Learn. Res. 2016 |
Bioinformatics and computational biology › genomics
genome-wide association study |
0.2 | 1 | 2016 | A time-varying group sparse additive model for genome-wide association studies of dynamic complex traits · Bioinform. 2016 |
Bioinformatics and computational biology
statistical genetics |
0.2 | 1 | 2016 | A time-varying group sparse additive model for genome-wide association studies of dynamic complex traits · Bioinform. 2016 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.2 | 1 | 2013 | A Scalable Approach to Probabilistic Latent Space Inference of Large-Scale Networks · NIPS 2013 |
Data mining
network analysis |
0.2 | 1 | 2013 | A Scalable Approach to Probabilistic Latent Space Inference of Large-Scale Networks · NIPS 2013 |
Machine learning › Trustworthy machine learning › interpretability › explainable AI
additive models |
0.1 | 1 | 2012 | Group Sparse Additive Models · ICML 2012 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.1 | 1 | 2012 | On Triangular versus Edge Representations --- Towards Scalable Modeling of Networks · NIPS 2012 |
Machine learning › Optimization for machine learning › sparse learning
group sparsity |
0.1 | 1 | 2012 | Group Sparse Additive Models · ICML 2012 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
sparse additive model |
0.1 | 1 | 2012 | Group Sparse Additive Models · ICML 2012 |
Data mining
clustering |
0.1 | 1 | 2012 | On Triangular versus Edge Representations --- Towards Scalable Modeling of Networks · NIPS 2012 |
Data mining › structured data mining › graph mining
community detection |
0.1 | 1 | 2012 | On Triangular versus Edge Representations --- Towards Scalable Modeling of Networks · NIPS 2012 |
Data mining › structured data mining › graph mining
network motif |
0.1 | 1 | 2012 | On Triangular versus Edge Representations --- Towards Scalable Modeling of Networks · NIPS 2012 |
Methods — techniques the papers use, named apart from their topics
upper confidence bound · 1.1thompson sampling · 1.1regret analysis · 1.1triangular motif representation · 0.6stochastic variational inference · 0.6global optimization · 0.5inverse regression · 0.3bayesian inference · 0.3block coordinate gradient descent · 0.3parameter server · 0.2latent space model · 0.2incremental update · 0.2group sparse additive model · 0.2functional regression · 0.2subsampling · 0.1group lasso · 0.1approximate inference · 0.1likelihood-based method · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation ModelsabstractSince the seminal work of TabPFN, research on tabular foundation models (TFMs) based on in-context learning (ICL) has challenged long-standing paradigms in machine learning. Without seeing any real-world data, models pretrained on purely synthetic datasets generalize remarkably well across diverse datasets, often using only a moderate number of in-context examples. This shifts the focus in tabular machine learning from model architecture design to the design of synthetic datasets, or, more precisely, to the prior distributions that generate them. Yet the guiding principles for prior design remain poorly understood. This work marks the first attempt to address the gap. We systematically investigate and identify key properties of synthetic priors that allow pretrained TFMs to generalize well. Based on these insights, we introduce Mitra, a TFM trained on a curated mixture of synthetic priors selected for their diversity, distinctiveness, and performance on real-world tabular data. Mitra consistently outperforms state-of-the-art TFMs, such as TabPFNv2 and TabICL, across both classification and regression benchmarks, with better sample efficiency. Danielle Maddix Robinson, Junming Yin, Nick Erickson, Abdul Fatir Ansari, Boran Han, Shuai Zhang 0007, Leman Akoglu, Christos Faloutsos, Michael W. Mahoney, Tony Hu, Huzefa Rangwala, George Karypis, Yuyang Wang 0001 |
NeurIPS | 3 |
| 2022 | Bandit Learning in Many-to-One Matching MarketsabstractThe problem of two-sided matching markets is well-studied in social science and economics. Some recent works study how to match while learning the unknown preferences of agents in one-to-one matching markets. However, in many cases like the online recruitment platform for short-term workers, a company can select more than one agent while an agent can only select one company at a time. These short-term workers try many times in different companies to find the most suitable jobs for them. Thus we consider a more general bandit learning problem in many-to-one matching markets where each arm has a fixed capacity and agents make choices with multiple rounds of iterations. We develop algorithms in both centralized and decentralized settings and prove regret bounds of order O(log T) and Olog2 T) respectively. Extensive experiments show the convergence and effectiveness of our algorithms. Zilong Wang 0010, Liya Guo, Junming Yin, Shuai Li 0010 |
CIKM | 3 |
| 2022 | Thompson Sampling for Bandit Learning in Matching MarketsabstractThe problem of two-sided matching markets has a wide range of real-world applications and has been extensively studied in the literature. A line of recent works have focused on the problem setting where the preferences of one-side market participants are unknown a priori and are learned by iteratively interacting with the other side of participants. All these works are based on explore-then-commit (ETC) and upper confidence bound (UCB) algorithms, two common strategies in multi-armed bandits (MAB). Thompson sampling (TS) is another popular approach, which attracts lots of attention due to its easier implementation and better empirical performances. In many problems, even when UCB and ETC-type algorithms have already been analyzed, researchers are still trying to study TS for its benefits. However, the convergence analysis of TS is much more challenging and remains open in many problem settings. In this paper, we provide the first regret analysis for TS in the new setting of iterative matching markets. Extensive experiments demonstrate the practical advantages of the TS-type algorithm over the ETC and UCB-type baselines. Fang Kong 0002, Junming Yin, Shuai Li 0010 |
IJCAI | 2 |
| 2020 | Relaxed Multivariate Bernoulli Distribution and Its Applications to Deep Generative ModelsabstractRecent advances in variational auto-encoder (VAE) have demonstrated the possibility of approximating the intractable posterior distribution with a variational distribution parameterized by a neural network. To optimize the variational objective of VAE, the reparameterization trick is commonly applied to obtain a low-variance estimator of the gradient. The main idea of the trick is to express the variational distribution as a differentiable function of parameters and a random variable with a fixed distribution. To extend the reparameterization trick to inference involving discrete latent variables, a common approach is to use a continuous relaxation of the categorical distribution as the approximate posterior. However, when applying continuous relaxation to the multivariate cases, multiple variables are typically assumed to be independent, making it suboptimal in applications where modeling dependency is crucial to the overall performance. In this work, we propose a multivariate generalization of the Relaxed Bernoulli distribution, which can be reparameterized and can capture the correlation between variables via a Gaussian copula. We demonstrate its effectiveness in two tasks: density estimation with Bernoulli VAE and semi-supervised multi-label classification. Junming Yin |
UAI | 2 |
| 2018 | Supervised Topic Modeling Using Hierarchical Dirichlet Process-Based Inverse Regression: Experiments on E-Commerce ApplicationsabstractThe proliferation of e-commerce calls for mining consumer preferences and opinions from user-generated text. To this end, topic models have been widely adopted to discover the underlying semantic themes (i.e., topics). Supervised topic models have emerged to leverage discovered topics for predicting the response of interest (e.g., product quality and sales). However, supervised topic modeling remains a challenging problem because of the need to prespecify the number of topics, the lack of predictive information in topics, and limited scalability. In this paper, we propose a novel supervised topic model, Hierarchical Dirichlet Process-based Inverse Regression (HDP-IR). HDP-IR characterizes the corpus with a flexible number of topics, which prove to retain as much predictive information as the original corpus. Moreover, we develop an efficient inference algorithm capable of examining large-scale corpora (millions of documents or more). Three experiments were conducted to evaluate the predictive performance over major e-commerce benchmark testbeds of online reviews. Overall, HDP-IR outperformed existing state-of-the-art supervised topic models. Particularly, retaining sufficient predictive information improved predictive R-squared by over 17.6 percent; having topic structure flexibility contributed to predictive R-squared by at least 4.1 percent. HDP-IR provides an important step for future study on user-generated texts from a topic perspective. Weifeng Li 0002, Junming Yin, Hsinchun Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks (Extended Abstract)abstractWe propose to model dependence within a network view using the temporal latent space model, which uses a time-dependent low-dimensional geometric projections to represent the high-dimensional dependence structure in time-varying networks. Once we obtain the lowdimensional temporal latent space representation for graphs from time 1 to t, we can accurately predict future links in time t + 1 (i.e., Gt+1). We present a global optimization algorithm to effectively infer the temporal latent space using block coordinate gradient descent (BCGD). We further introduce two new variants of BCGD: a local BCGD algorithm and an incremental BCGD algorithm, to scale the inference algorithm to massive networks. Linhong Zhu, Junming Yin, Greg Ver Steeg, Aram Galstyan |
ICDE | 3 |
| 2017 | Convex-constrained Sparse Additive Modeling and Its Extensions
Junming Yin, Yaoliang Yu |
UAI | 1 |
| 2016 | Targeting key data breach services in underground supply chainabstractOver the past decade, a growing body of cybercriminals have founded an underground supply chain to facilitate data breaches, leading to the leak of personal information for hundreds of millions of individuals. As many service providers in the supply chain are rippers, cybercriminals tend to rely on a few key services. Identifying key services is of great interest to both cybersecurity researchers and practitioners. This study presents a text-mining framework for identifying key data breach services based on analysis of service reviews. The framework includes crawlers with counter anti-crawling measures, text preprocessing, and supervised topic models. In our experiment, more than 70% of the key services were identified by our framework. Weifeng Li 0002, Junming Yin, Hsinchun Chen |
ISI | 2 |
| 2016 | A time-varying group sparse additive model for genome-wide association studies of dynamic complex traitsabstractMOTIVATION: Despite the widespread popularity of genome-wide association studies (GWAS) for genetic mapping of complex traits, most existing GWAS methodologies are still limited to the use of static phenotypes measured at a single time point. In this work, we propose a new method for association mapping that considers dynamic phenotypes measured at a sequence of time points. Our approach relies on the use of Time-Varying Group Sparse Additive Models (TV-GroupSpAM) for high-dimensional, functional regression. RESULTS: This new model detects a sparse set of genomic loci that are associated with trait dynamics, and demonstrates increased statistical power over existing methods. We evaluate our method via experiments on synthetic data and perform a proof-of-concept analysis for detecting single nucleotide polymorphisms associated with two phenotypes used to assess asthma severity: forced vital capacity, a sensitive measure of airway obstruction and bronchodilator response, which measures lung response to bronchodilator drugs. AVAILABILITY AND IMPLEMENTATION: Source code for TV-GroupSpAM freely available for download at http://www.cs.cmu.edu/~mmarchet/projects/tv_group_spam, implemented in MATLAB. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Micol Marchetti-Bowick, Junming Yin, Judie A. Howrylak, Eric P. Xing |
Bioinform. | 2 |
| 2016 | Latent Space Inference of Internet-Scale NetworksabstractThe rise of Internet-scale networks, such as web graphs and social media with hundreds of millions to billions of nodes, presents new scientific opportunities, such as overlapping community detection to discover the structure of the Internet, or to analyze trends in online social behavior. However, many existing probabilistic network models are difficult or impossible to deploy at these massive scales. We propose a scalable approach for modeling and inferring latent spaces in Internet-scale networks, with an eye towards overlapping community detection as a key application. By applying a succinct representation of networks as a bag of triangular motifs, developing a parsimonious statistical model, deriving an efficient stochastic variational inference algorithm, and implementing it as a distributed cluster program via the Petuum parameter server system, we demonstrate overlapping community detection on real networks with up to 100 million nodes and 1000 communities on 5 machines in under 40 hours. Compared to other state-of-the-art probabilistic network approaches, our method is several orders of magnitude faster, with competitive or improved accuracy at overlapping community detection. Qirong Ho, Junming Yin, Eric P. Xing |
J. Mach. Learn. Res. | 2 |
| 2016 | Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social NetworksabstractWe propose a temporal latent space model for link prediction in dynamic social networks, where the goal is to predict links over time based on a sequence of previous graph snapshots. The model assumes that each user lies in an unobserved latent space, and interactions are more likely to occur between similar users in the latent space representation. In addition, the model allows each user to gradually move its position in the latent space as the network structure evolves over time. We present a global optimization algorithm to effectively infer the temporal latent space. Two alternative optimization algorithms with local and incremental updates are also proposed, allowing the model to scale to larger networks without compromising prediction accuracy. Empirically, we demonstrate that our model, when evaluated on a number of real-world dynamic networks, significantly outperforms existing approaches for temporal link prediction in terms of both scalability and predictive power. Linhong Zhu, Junming Yin, Greg Ver Steeg, Aram Galstyan |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | A Scalable Approach to Probabilistic Latent Space Inference of Large-Scale NetworksabstractWe propose a scalable approach for making inference about latent spaces of large networks. With a succinct representation of networks as a bag of triangular motifs, a parsimonious statistical model, and an efficient stochastic variational inference algorithm, we are able to analyze real networks with over a million vertices and hundreds of latent roles on a single machine in a matter of hours, a setting that is out of reach for many existing methods. When compared to the state-of-the-art probabilistic approaches, our method is several orders of magnitude faster, with competitive or improved accuracy for latent space recovery and link prediction. Junming Yin, Qirong Ho, Eric P. Xing |
NIPS | 1 |
| 2012 | Group Sparse Additive Models
Junming Yin, Xi Chen 0010, Eric P. Xing |
ICML | 1 |
| 2012 | On Triangular versus Edge Representations --- Towards Scalable Modeling of NetworksabstractIn this paper, we argue for representing networks as a bag of {\it triangular motifs}, particularly for important network problems that current model-based approaches handle poorly due to computational bottlenecks incurred by using edge representations. Such approaches require both 1-edges and 0-edges (missing edges) to be provided as input, and as a consequence, approximate inference algorithms for these models usually require $\Omega(N^2)$ time per iteration, precluding their application to larger real-world networks. In contrast, triangular modeling requires less computation, while providing equivalent or better inference quality. A triangular motif is a vertex triple containing 2 or 3 edges, and the number of such motifs is $\Theta(\sum_{i}D_{i}^{2})$ (where $D_i$ is the degree of vertex $i$), which is much smaller than $N^2$ for low-maximum-degree networks. Using this representation, we develop a novel mixed-membership network model and approximate inference algorithm suitable for large networks with low max-degree. For networks with high maximum degree, the triangular motifs can be naturally subsampled in a {\it node-centric} fashion, allowing for much faster inference at a small cost in accuracy. Empirically, we demonstrate that our approach, when compared to that of an edge-based model, has faster runtime and improved accuracy for mixed-membership community detection. We conclude with a large-scale demonstration on an $N\approx 280,000$-node network, which is infeasible for network models with $\Omega(N^2)$ inference cost. Qirong Ho, Junming Yin, Eric P. Xing |
NIPS | 2 |
| 2009 | Joint estimation of gene conversion rates and mean conversion tract lengths from population SNP dataabstractMOTIVATION: Two known types of meiotic recombination are crossovers and gene conversions. Although they leave behind different footprints in the genome, it is a challenging task to tease apart their relative contributions to the observed genetic variation. In particular, for a given population SNP dataset, the joint estimation of the crossover rate, the gene conversion rate and the mean conversion tract length is widely viewed as a very difficult problem. RESULTS: In this article, we devise a likelihood-based method using an interleaved hidden Markov model (HMM) that can jointly estimate the aforementioned three parameters fundamental to recombination. Our method significantly improves upon a recently proposed method based on a factorial HMM. We show that modeling overlapping gene conversions is crucial for improving the joint estimation of the gene conversion rate and the mean conversion tract length. We test the performance of our method on simulated data. We then apply our method to analyze real biological data from the telomere of the X chromosome of Drosophila melanogaster, and show that the ratio of the gene conversion rate to the crossover rate for the region may not be nearly as high as previously claimed. AVAILABILITY: A software implementation of the algorithms discussed in this article is available at http://www.cs.berkeley.edu/ approximately yss/software.html. Junming Yin, Michael I. Jordan, Yun S. Song |
Bioinform. | 1 |