VLDB 2026 Research / reviewers in the wild / expert
Boya Ma
dblp:330/1105
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2024
0009-0001-4917-5050ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Low Rank Multi-Dictionary Selection at ScaleabstractThe sparse dictionary coding framework represents signals as a linear combination of a few predefined dictionary atoms. It has been employed for images, time series, graph signals and recently for 2-way (or 2D) spatio-temporal data employing jointly temporal and spatial dictionaries. Large and over-complete dictionaries enable high-quality models, but also pose scalability challenges which are exacerbated in multi-dictionary settings. Hence, an important problem that we address in this paper is: How to scale multi-dictionary coding for large dictionaries and datasets?We propose a multi-dictionary atom selection technique for low-rank sparse coding named LRMDS. To enable scalability to large dictionaries and datasets, it progressively selects groups of row-column atom pairs based on their alignment with the data and performs convex relaxation coding via the corresponding sub-dictionaries. We demonstrate both theoretically and experimentally that when the data has a low-rank encoding with a sparse subset of the atoms, LRMDS is able to select them with strong guarantees under mild assumptions. Furthermore, we demonstrate the scalability and quality of LRMDS in both synthetic and real-world datasets and for a range of coding dictionaries. It achieves 3 times to 10 times speed-up compared to baselines, while obtaining up to two orders of magnitude improvement in representation quality on some of the real world datasets given a fixed target number of atoms. Boya Ma, Maxwell McNeil, Abram Magner, Petko Bogdanov |
KDD | 1 |
| 2023 | GIST: Graph Inference for Structured Time SeriesabstractMachine learning and data analytics tasks on graphs enjoy a lot of attention from both researchers and practitioners due to the utility that a graph structure among data entities adds for downstream tasks. In many cases, however, a graph structure is not known a priori, and instead has to be inferred from data. Specifically, learning a graph associating time series may elucidate hidden dependencies and also enable improved performance in tasks like classification, forecasting and clustering. While approaches based on pairwise correlation and precision matrix estimation have been employed widely, recent approaches that model observations as signals on graphs have been shown to be more advantageous. Boya Ma, Maxwell McNeil, Petko Bogdanov |
SDM | 1 |
| 2022 | SAGA: Signal-Aware Graph AggregationabstractGraphs are widely employed models for structural dependencies in complex systems such as social, infrastructure, and information networks. Graph datasets are often massive, computationally challenging to mine, and non-trivial to understand by humans. Hence, there is a large body of literature on graph summarization aiming to improve algorithmic efficiency, quality of analytics tasks, and visualization. Most existing graph summarization methods focus solely on the structure of the graph and assume that node properties are static. In this paper we focus on graphs with temporal measurements on their nodes, which we call temporal graph signals. Our goal is to learn a graph aggregation (summary) that best reflects both the structure and the temporal node measurements. We propose a signal-aware graph aggregation framework called SAGA. The key idea is to group well-connected nodes whose behavior exhibits consistent temporal patterns. SAGA learns simultaneously how to (i) aggregate the graph into supernode groups and (ii) represent the groups' collective temporal behavior succinctly via a sparse dictionary encoding. The obtained aggregations offer insights into the functional organization of the graph and the learned model enables improved performance for state-of-the-art approaches for downstream tasks like temporal graph signal decomposition, forecasting and link prediction. We demonstrate, in both synthetic and real-world data sets, that SAGA's learned aggregations improve (i) reconstruction quality for temporal graph signals by up to 75%, (ii) link prediction accuracy by up to 40% and (iii) the accuracy of forecasting by up to 63% while also offering 50% scalability improvements. Maxwell McNeil, Boya Ma, Petko Bogdanov |
SDM | 2 |