EDBT 2026 Demo / reviewers in the wild / expert
Arthur U. Asuncion
dblp:78/2352
· DBLP profile ↗
13ranked-venue papers
4as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Probabilistic and Bayesian machine learning · 42% Information extraction and text analysis · 36% Graph learning · 16% | |
| Network and information security
1 paper |
Privacy and data protection · 67% Security and privacy of machine learning · 33% | |
| Databases, data mining, and information retrieval
3 papers |
Recommender systems · 39% Graph data management · 31% Data mining · 31% | |
| Software engineering, system software, and programming languages
1 paper |
Software maintenance and evolution · 67% Empirical software engineering · 33% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Distributed systems · 100% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
topic model |
0.3 | 4 | 2009 | Distributed Algorithms for Topic Models · J. Mach. Learn. Res. 2009 Asynchronous Distributed Learning of Topic Models · NIPS 2008 Fast collapsed gibbs sampling for latent dirichlet allocation · KDD 2008 |
Natural language and speech › Information extraction and text analysis › topic model
latent dirichlet allocation |
0.2 | 3 | 2008 | Asynchronous Distributed Learning of Topic Models · NIPS 2008 Fast collapsed gibbs sampling for latent dirichlet allocation · KDD 2008 Distributed Inference for Latent Dirichlet Allocation · NIPS 2007 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.2 | 2 | 2010 | Particle Filtered MCMC-MLE with Connections to Contrastive Divergence · ICML 2010 Fast collapsed gibbs sampling for latent dirichlet allocation · KDD 2008 |
Privacy and data protection › statistical database privacy
database privacy |
0.2 | 1 | 2013 | Nonadaptive Mastermind Algorithms for String and Vector Databases, with Case Studies · IEEE Trans. Knowl. Data Eng. 2013 |
Security and privacy of machine learning › privacy attack
data reconstruction attack |
0.2 | 1 | 2013 | Nonadaptive Mastermind Algorithms for String and Vector Databases, with Case Studies · IEEE Trans. Knowl. Data Eng. 2013 |
Privacy and data protection
privacy-preserving data analysis |
0.2 | 1 | 2013 | Nonadaptive Mastermind Algorithms for String and Vector Databases, with Case Studies · IEEE Trans. Knowl. Data Eng. 2013 |
Recommender systems
collaborative filtering |
0.2 | 2 | 2013 | Bayesian Matrix Factorization with Side Information and Dirichlet Process Mixtures · AAAI 2010 Nonadaptive Mastermind Algorithms for String and Vector Databases, with Case Studies · IEEE Trans. Knowl. Data Eng. 2013 |
Machine learning › Graph learning › dynamic graph learning
dynamic graph modeling |
0.1 | 1 | 2011 | Continuous-Time Regression Models for Longitudinal Networks · NIPS 2011 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › relational model
statistical network models |
0.1 | 1 | 2011 | Continuous-Time Regression Models for Longitudinal Networks · NIPS 2011 |
Graph data management
citation network |
0.1 | 1 | 2011 | Dynamic Egocentric Models for Citation Networks · ICML 2011 |
Data mining › structured data mining › graph mining › dynamic network analysis
dynamic network model |
0.1 | 1 | 2011 | Dynamic Egocentric Models for Citation Networks · ICML 2011 |
Machine learning › Representation and self-supervised learning › matrix factorization
bayesian matrix factorization |
0.1 | 1 | 2010 | Bayesian Matrix Factorization with Side Information and Dirichlet Process Mixtures · AAAI 2010 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation |
0.1 | 1 | 2010 | Particle Filtered MCMC-MLE with Connections to Contrastive Divergence · ICML 2010 |
Software maintenance and evolution › traceability
automated traceability |
0.1 | 1 | 2010 | Software traceability with topic modeling · ICSE (1) 2010 |
Empirical software engineering
mining software repositories |
0.1 | 1 | 2010 | Software traceability with topic modeling · ICSE (1) 2010 |
Software maintenance and evolution
traceability |
0.1 | 1 | 2010 | Software traceability with topic modeling · ICSE (1) 2010 |
Distributed systems
distributed algorithms |
0.1 | 1 | 2009 | Distributed Algorithms for Topic Models · J. Mach. Learn. Res. 2009 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference |
0.1 | 1 | 2008 | Fast collapsed gibbs sampling for latent dirichlet allocation · KDD 2008 |
Machine learning › Probabilistic and Bayesian machine learning › sampling
collapsed gibbs sampling |
0.1 | 1 | 2008 | Fast collapsed gibbs sampling for latent dirichlet allocation · KDD 2008 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
gibbs sampling |
0.1 | 1 | 2008 | Fast collapsed gibbs sampling for latent dirichlet allocation · KDD 2008 |
Distributed systems › distributed machine learning
distributed inference |
0.1 | 1 | 2007 | Distributed Inference for Latent Dirichlet Allocation · NIPS 2007 |
Computational social science and digital humanities
social network analysis |
0.0 | 1 | 2011 | Continuous-Time Regression Models for Longitudinal Networks · NIPS 2011 |
Methods — techniques the papers use, named apart from their topics
gibbs sampling · 0.6sparsity exploitation · 0.3nonadaptive group testing · 0.3survival analysis · 0.2maximum likelihood inference · 0.2event history analysis · 0.2dynamic egocentric models · 0.2dirichlet process mixture · 0.2variational inference · 0.2topic modeling · 0.1probabilistic topic model · 0.1particle filtering · 0.1contrastive divergence · 0.1collapsed gibbs sampling · 0.1hierarchical bayesian modeling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Nonadaptive Mastermind Algorithms for String and Vector Databases, with Case StudiesabstractIn this paper, we study sparsity-exploiting Mastermind algorithms for attacking the privacy of an entire database of character strings or vectors, such as DNA strings, movie ratings, or social network friendship data. Based on reductions to nonadaptive group testing, our methods are able to take advantage of minimal amounts of privacy leakage, such as contained in a single bit that indicates if two people in a medical database have any common genetic mutations, or if two people have any common friends in an online social network. We analyze our Mastermind attack algorithms using theoretical characterizations that provide sublinear bounds on the number of queries needed to clone the database, as well as experimental tests on genomic information, collaborative filtering data, and online social networks. By taking advantage of the generally sparse nature of these real-world databases and modulating a parameter that controls query sparsity, we demonstrate that relatively few nonadaptive queries are needed to recover a large majority of each database. Arthur U. Asuncion, Michael T. Goodrich |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | TopicNets: Visual Analysis of Large Text Corpora with Topic ModelingabstractWe present TopicNets , a Web-based system for visual and interactive analysis of large sets of documents using statistical topic models. A range of visualization types and control mechanisms to support knowledge discovery are presented. These include corpus- and document-specific views, iterative topic modeling, search, and visual filtering. Drill-down functionality is provided to allow analysts to visualize individual document sections and their relations within the global topic space. Analysts can search across a dataset through a set of expansion techniques on selected document and topic nodes. Furthermore, analysts can select relevant subsets of documents and perform real-time topic modeling on these subsets to interactively visualize topics at various levels of granularity, allowing for a better understanding of the documents. A discussion of the design and implementation choices for each visual analysis technique is presented. This is followed by a discussion of three diverse use cases in which TopicNets enables fast discovery of information that is otherwise hard to find. These include a corpus of 50,000 successful NSF grant proposals, 10,000 publications from a large research center, and single documents including a grant proposal and a PhD thesis. Brynjar Gretarsson, John O'Donovan, Svetlin Bostandjiev, Tobias Höllerer, Arthur U. Asuncion, David Newman 0001, Padhraic Smyth |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2011 | Dynamic Egocentric Models for Citation Networks
Duy Quang Vu, Arthur U. Asuncion, David R. Hunter, Padhraic Smyth |
ICML | 2 |
| 2011 | Continuous-Time Regression Models for Longitudinal NetworksabstractThe development of statistical models for continuous-time longitudinal network data is of increasing interest in machine learning and social science. Leveraging ideas from survival and event history analysis, we introduce a continuous-time regression modeling framework for network event data that can incorporate both time-dependent network statistics and time-varying regression coefficients. We also develop an efficient inference scheme that allows our approach to scale to large networks. On synthetic and real-world data, empirical results demonstrate that the proposed inference approach can accurately estimate the coefficients of the regression model, which is useful for interpreting the evolution of the network; furthermore, the learned model has systematically better predictive performance compared to standard baseline methods. Duy Quang Vu, Arthur U. Asuncion, David R. Hunter, Padhraic Smyth |
NIPS | 2 |
| 2010 | Bayesian Matrix Factorization with Side Information and Dirichlet Process MixturesabstractMatrix factorization is a fundamental technique in machine learning that is applicable to collaborative filtering, information retrieval and many other areas. In collaborative filtering and many other tasks, the objective is to fill in missing elements of a sparse data matrix. One of the biggest challenges in this case is filling in a column or row of the matrix with very few observations. In this paper we introduce a Bayesian matrix factorization model that performs regression against side information known about the data in addition to the observations. The side information helps by adding observed entries to the factored matrices. We also introduce a nonparametric mixture model for the prior of the rows and columns of the factored matrices that gives a different regularization for each latent class. Besides providing a richer prior, the posterior distribution of mixture assignments reveals the latent classes. Using Gibbs sampling for inference, we apply our model to the Netflix Prize problem of predicting movie ratings given an incomplete user-movie ratings matrix. Incorporating rating information with gathered metadata information, our Bayesian approach outperforms other matrix factorization techniques even when using fewer dimensions. Ian Porteous, Arthur U. Asuncion, Max Welling |
AAAI | 2 |
| 2010 | Particle Filtered MCMC-MLE with Connections to Contrastive Divergence
Arthur U. Asuncion, Qiang Liu 0001, Alexander Ihler, Padhraic Smyth |
ICML | 1 |
| 2010 | Software traceability with topic modelingabstractSoftware traceability is a fundamentally important task in software engineering. The need for automated traceability increases as projects become more complex and as the number of artifacts increases. We propose an automated technique that combines traceability with a machine learning technique known as topic modeling. Our approach automatically records traceability links during the software development process and learns a probabilistic topic model over artifacts. The learned model allows for the semantic categorization of artifacts and the topical visualization of the software system. To test our approach, we have implemented several tools: an artifact search tool combining keyword-based search and topic modeling, a recording tool that performs prospective traceability, and a visualization tool that allows one to navigate the software architecture and view semantic topics associated with relevant artifacts and architectural components. We apply our approach to several data sets and discuss how topic modeling enhances software traceability, and vice versa. Categories and Subject Descriptors Hazeline U. Asuncion, Arthur U. Asuncion, Richard N. Taylor |
ICSE (1) | 2 |
| 2009 | On Smoothing and Inference for Topic Models
Arthur U. Asuncion, Max Welling, Padhraic Smyth, Yee Whye Teh |
UAI | 1 |
| 2009 | Distributed Algorithms for Topic Models
David Newman 0001, Arthur U. Asuncion, Padhraic Smyth, Max Welling |
J. Mach. Learn. Res. | 2 |
| 2008 | Fast collapsed gibbs sampling for latent dirichlet allocationabstractIn this paper we introduce a novel collapsed Gibbs sampling method for the widely used latent Dirichlet allocation (LDA) model. Our new method results in significant speedups on real world text corpora. Conventional Gibbs sampling schemes for LDA require O(K) operations per sample where K is the number of topics in the model. Our proposed method draws equivalent samples but requires on average significantly less then K operations per sample. On real-word corpora FastLDA can be as much as 8 times faster than the standard collapsed Gibbs sampler for LDA. No approximations are necessary, and we show that our fast sampling scheme produces exactly the same results as the standard (but slower) sampling scheme. Experiments on four real world data sets demonstrate speedups for a wide range of collection sizes. For the PubMed collection of over 8 million documents with a required computation time of 6 CPU months for LDA, our speedup of 5.7 can save 5 CPU months of computation. Ian Porteous, David Newman 0001, Alexander Ihler, Arthur U. Asuncion, Padhraic Smyth, Max Welling |
KDD | 4 |
| 2008 | Asynchronous Distributed Learning of Topic ModelsabstractDistributed learning is a problem of fundamental interest in machine learning and cognitive science. In this paper, we present asynchronous distributed learning algorithms for two well-known unsupervised learning frameworks: Latent Dirichlet Allocation (LDA) and Hierarchical Dirichlet Processes (HDP). In the proposed approach, the data are distributed across P processors, and processors independently perform Gibbs sampling on their local data and communicate their information in a local asynchronous manner with other processors. We demonstrate that our asynchronous algorithms are able to learn global topic models that are statistically as accurate as those learned by the standard LDA and HDP samplers, but with significant improvements in computation time and memory. We show speedup results on a 730-million-word text corpus using 32 processors, and we provide perplexity results for up to 1500 virtual processors. As a stepping stone in the development of asynchronous HDP, a parallel HDP sampler is also introduced. Arthur U. Asuncion, Padhraic Smyth, Max Welling |
NIPS | 1 |
| 2007 | Distributed Inference for Latent Dirichlet Allocationabstract processors only sees We investigate the problem of learning a widely-used latent-variable model – the Latent Dirichlet Allocation (LDA) or “topic” model – using distributed compu- of the total data set. We pro- tation, where each of pose two distributed inference schemes that are motivated from different perspec- tives. The first scheme uses local Gibbs sampling on each processor with periodic updates—it is simple to implement and can be viewed as an approximation to a single processor implementation of Gibbs sampling. The second scheme re- lies on a hierarchical Bayesian extension of the standard LDA model to directly processors—it has a theo- account for the fact that data are distributed across retical guarantee of convergence but is more complex to implement than the ap- proximate method. Using five real-world text corpora we show that distributed learning works very well for LDA models, i.e., perplexity and precision-recall scores for distributed learning are indistinguishable from those obtained with single-processor learning. Our extensive experimental results include large-scale distributed computation on 1000 virtual processors; and speedup experiments of learning topics in a 100-million word corpus using 16 processors. David Newman 0001, Arthur U. Asuncion, Padhraic Smyth, Max Welling |
NIPS | 2 |
| 2005 | Incremental Parallelization Using Navigational Programming: A Case StudyabstractWe show how a series of transformations can be applied to incrementally parallelize sequential programs. Our navigational programming (NavP) methodology is based on the principle of self-migrating computations and is truly incremental, in that each step represents a functioning program and every intermediate program is an improvement over its predecessor. The transformations are mechanical and straightforward to apply. We illustrate our methodology in the context of matrix multiplication. Our final stage is similar to the classical Gentleman's algorithm. The NavP methodology is conducive to new ways of thinking that lead to ease of programming and high performance. Lei Pan 0001, Wenhui Zhang 0001, Arthur U. Asuncion, Ming Kin Lai, Michael B. Dillencourt, Lubomir F. Bic |
ICPP | 3 |