VLDB 2026 Research / reviewers in the wild / expert
Rico Angell
dblp:184/9716
· DBLP profile ↗
9ranked-venue papers
6as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
2 papers |
Mathematical optimization · 48% Algorithms and data structures · 36% Graph algorithms and graph theory · 16% | |
| Artificial intelligence
4 papers |
Trustworthy machine learning · 69% Representation and self-supervised learning · 22% Probabilistic and Bayesian machine learning · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Software testing · 77% Debugging and program repair · 23% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Environmental and earth informatics · 100% |
Topics — the 13 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning |
0.8 | 1 | 2024 | Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder |
0.8 | 1 | 2024 | Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models · NeurIPS 2024 |
Mathematical optimization
semidefinite programming |
0.8 | 1 | 2024 | Fast, Scalable, Warm-Start Semidefinite Programming with Spectral Bundling and Sketching · ICML 2024 |
Information retrieval › similarity search
nearest neighbor search |
0.6 | 1 | 2022 | Efficient Nearest Neighbor Search for Cross-Encoder Models using Matrix Factorization · EMNLP 2022 |
Algorithms and data structures
clustering |
0.6 | 1 | 2022 | Interactive Correlation Clustering with Existential Cluster Constraints · ICML 2022 |
Graph algorithms and graph theory › graph clustering
correlation clustering |
0.6 | 1 | 2022 | Interactive Correlation Clustering with Existential Cluster Constraints · ICML 2022 |
Algorithms and data structures › clustering › clustering with queries
interactive clustering |
0.6 | 1 | 2022 | Interactive Correlation Clustering with Existential Cluster Constraints · ICML 2022 |
Machine learning › Trustworthy machine learning
fairness |
0.3 | 1 | 2018 | Themis: automatically testing software for discrimination · ESEC/SIGSOFT FSE 2018 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.3 | 1 | 2018 | Inferring Latent Velocities from Weather Radar Data using Gaussian Processes · NeurIPS 2018 |
Software testing
test generation |
0.3 | 1 | 2018 | Themis: automatically testing software for discrimination · ESEC/SIGSOFT FSE 2018 |
Mathematical optimization › optimization
warm-start optimization |
0.2 | 1 | 2024 | Fast, Scalable, Warm-Start Semidefinite Programming with Spectral Bundling and Sketching · ICML 2024 |
Algorithms and data structures › clustering
constrained clustering |
0.2 | 1 | 2022 | Interactive Correlation Clustering with Existential Cluster Constraints · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
inference algorithm · 1.1existential cluster constraints · 1.1warm-start initialization · 0.8spectral bundling · 0.8sparse autoencoder · 0.8p-annealing · 0.8matrix sketching · 0.8laplace's method · 0.7fast fourier transform · 0.7causal analysis · 0.7automated test suite generation · 0.7matrix factorization · 0.6CUR decomposition · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Fast, Scalable, Warm-Start Semidefinite Programming with Spectral Bundling and SketchingabstractWhile semidefinite programming (SDP) has traditionally been limited to moderate-sized problems, recent algorithms augmented with matrix sketching techniques have enabled solving larger SDPs. However, these methods achieve scalability at the cost of an increase in the number of necessary iterations, resulting in slower convergence as the problem size grows. Furthermore, they require iteration-dependent parameter schedules that prohibit effective utilization of warm-start initializations important in practical applications with incrementally-arriving data or mixed-integer programming. We present Unified Spectral Bundling with Sketching (USBS), a provably correct, fast and scalable algorithm for solving massive SDPs that can leverage a warm-start initialization to further accelerate convergence. Our proposed algorithm is a spectral bundle method for solving general SDPs containing both equality and inequality constraints. Moveover, when augmented with an optional matrix sketching technique, our algorithm achieves the dramatically improved scalability of previous work while sustaining convergence speed. We empirically demonstrate the effectiveness of our method across multiple applications, with and without warm-starting. For example, USBS provides a 500x speed-up over the state-of-the-art scalable SDP solver on an instance with over 2 billion decision variables. We make our implementation in pure JAX publicly available. Rico Angell, Andrew McCallum |
ICML | 1 |
| 2024 | Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game ModelsabstractWhat latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-truth collection of interpretable features which we expect good SAEs to identify. We thus propose to measure progress in interpretable dictionary learning by working in the setting of LMs trained on Chess and Othello transcripts. These settings carry natural collections of interpretable features—for example, “there is a knight on F3”—which we leverage into metrics for SAE quality. To guide progress in interpretable dictionary learning, we introduce a new SAE training technique, $p$-annealing, which demonstrates improved performance on our metric. Adam Karvonen, Benjamin Wright, Can Rager, Rico Angell, Jannik Brinkmann, Logan Smith, Claudio Mayrink Verdun, David Bau, Samuel Marks |
NeurIPS | 4 |
| 2022 | Efficient Nearest Neighbor Search for Cross-Encoder Models using Matrix FactorizationabstractEfficient k-nearest neighbor search is a fundamental task, foundational for many problems in NLP.When the similarity is measured by dot-product between dual-encoder vectors or ℓ 2 -distance, there already exist many scalable and efficient search methods.But not so when similarity is measured by more accurate and expensive black-box neural similarity models, such as cross-encoders, which jointly encode the query and candidate neighbor.The cross-encoders' high computational cost typically limits their use to reranking candidates retrieved by a cheaper model, such as dual encoder or TF-IDF.However, the accuracy of such a two-stage approach is upper-bounded by the recall of the initial candidate set, and potentially requires additional training to align the auxiliary retrieval model with the cross-encoder model.In this paper, we present an approach that avoids the use of a dual-encoder for retrieval, relying solely on the cross-encoder.Retrieval is made efficient with CUR decomposition, a matrix decomposition approach that approximates all pairwise cross-encoder distances from a small subset of rows and columns of the distance matrix.Indexing items using our approach is computationally cheaper than training an auxiliary dual-encoder model through distillation.Empirically, for k > 10, our approach provides test-time recall-vs-computational cost trade-offs superior to the current widely-used methods that re-rank items retrieved using a dual-encoder or TF-IDF. Nishant Yadav, Nicholas Monath, Rico Angell, Manzil Zaheer, Andrew McCallum |
EMNLP | 3 |
| 2022 | Interactive Correlation Clustering with Existential Cluster ConstraintsabstractWe consider the problem of clustering with user feedback. Existing methods express constraints about the input data points, most commonly through must-link and cannot-link constraints on data point pairs. In this paper, we introduce existential cluster constraints: a new form of feedback where users indicate the features of desired clusters. Specifically, users make statements about the existence of a cluster having (and not having) particular features. Our approach has multiple advantages: (1) constraints on clusters can express user intent more efficiently than point pairs; (2) in cases where the users’ mental model is of the desired clusters, it is more natural for users to express cluster-wise preferences; (3) it functions even when privacy restrictions prohibit users from seeing raw data. In addition to introducing existential cluster constraints, we provide an inference algorithm for incorporating our constraints into the output clustering. Finally, we demonstrate empirically that our proposed framework facilitates more accurate clustering with dramatically fewer user feedback inputs. Rico Angell, Nicholas Monath, Nishant Yadav, Andrew McCallum |
ICML | 1 |
| 2022 | Entity Linking via Explicit Mention-Mention Coreference ModelingabstractDhruv Agarwal, Rico Angell, Nicholas Monath, Andrew McCallum. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Dhruv Agarwal 0003, Rico Angell, Nicholas Monath, Andrew McCallum |
NAACL-HLT | 2 |
| 2021 | Clustering-based Inference for Biomedical Entity LinkingabstractRico Angell, Nicholas Monath, Sunil Mohan, Nishant Yadav, Andrew McCallum. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Rico Angell, Nicholas Monath, Sunil Mohan, Nishant Yadav, Andrew McCallum |
NAACL-HLT | 1 |
| 2018 | Inferring Latent Velocities from Weather Radar Data using Gaussian ProcessesabstractArchived data from the US network of weather radars hold detailed information about bird migration over the last 25 years, including very high-resolution partial measurements of velocity. Historically, most of this spatial resolution is discarded and velocities are summarized at a very small number of locations due to modeling and algorithmic limitations. This paper presents a Gaussian process (GP) model to reconstruct high-resolution full velocity fields across the entire US. The GP faithfully models all aspects of the problem in a single joint framework, including spatially random velocities, partial velocity measurements, station-specific geometries, measurement noise, and an ambiguity known as aliasing. We develop fast inference algorithms based on the FFT; to do so, we employ a creative use of Laplace's method to sidestep the fact that the kernel of the joint process is non-stationary. Rico Angell, Daniel Sheldon |
NeurIPS | 1 |
| 2018 | Themis: automatically testing software for discriminationabstractBias in decisions made by modern software is becoming a common and serious problem. We present Themis, an automated test suite generator to measure two types of discrimination, including causal relationships between sensitive inputs and program behavior. We explain how Themis can measure discrimination and aid its debugging, describe a set of optimizations Themis uses to reduce test suite size, and demonstrate Themis' effectiveness on open-source software. Themis is open-source and all our evaluation data are available at http://fairness.cs.umass.edu/. See a video of Themis in action: https://youtu.be/brB8wkaUesY Rico Angell, Brittany Johnson, Yuriy Brun, Alexandra Meliou |
ESEC/SIGSOFT FSE | 1 |
| 2017 | Don't Be Greedy: Leveraging Community Structure to Find High Quality Seed Sets for Influence Maximization
Rico Angell, Grant Schoenebeck |
WINE | 1 |