VLDB 2026 Research / reviewers in the wild / expert
Gregory Polyakov
dblp:351/9847
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0003-7536-9670ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 100% | |
| Artificial intelligence
1 paper |
Trustworthy machine learning · 67% Language models and text generation · 33% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › retrieval models › sparse retrieval
learned sparse retrieval |
1.0 | 1 | 2026 | Understanding Wacky Weights: A Dissection of SPLADE's Learned Term Importance · SIGIR 2026 |
Information retrieval › retrieval models › sparse retrieval › learned sparse retrieval
SPLADE |
1.0 | 1 | 2026 | Understanding Wacky Weights: A Dissection of SPLADE's Learned Term Importance · SIGIR 2026 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | Interpretability Analysis of Arithmetic In-Context Learning in Large Language Models · EMNLP 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Interpretability Analysis of Arithmetic In-Context Learning in Large Language Models · EMNLP 2025 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
0.9 | 1 | 2025 | Interpretability Analysis of Arithmetic In-Context Learning in Large Language Models · EMNLP 2025 |
Information retrieval › retrieval models › neural retrieval
neural ranking model |
0.9 | 1 | 2025 | Towards Best Practices of Axiomatic Activation Patching in Information Retrieval · SIGIR 2025 |
Information retrieval
retrieval evaluation |
0.9 | 1 | 2025 | Towards Best Practices of Axiomatic Activation Patching in Information Retrieval · SIGIR 2025 |
Information retrieval › cross-modal retrieval
image-text retrieval |
0.7 | 1 | 2023 | Sinkhorn Transformations for Single-Query Postprocessing in Text-Video Retrieval · SIGIR 2023 |
Information retrieval
multimodal retrieval |
0.7 | 1 | 2023 | Sinkhorn Transformations for Single-Query Postprocessing in Text-Video Retrieval · SIGIR 2023 |
Methods — techniques the papers use, named apart from their topics
activation patching · 1.7transformer backbones · 1.0sparse regularization · 1.0loss function variation · 1.0mechanistic interpretability · 0.9logit lens · 0.9information flow analysis · 0.9circuit discovery · 0.9sinkhorn transformations · 0.7dual-softmax loss · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Understanding Wacky Weights: A Dissection of SPLADE's Learned Term ImportanceabstractLearned sparse retrieval models such as SPLADE combine the effectiveness of neural architectures with the efficiency of inverted indices. As these models assign weights to terms from a fixed vocabulary, interpretability is often touted as a major benefit of these models. However, the emergence of wacky weights, i.e., expansion terms that appear semantically unrelated to the input, limits interpretability. While prior research has anecdotally observed this phenomenon, there is a lack of systematic understanding regarding their origins, prevalence, and contribution to retrieval effectiveness. In this paper, we reproduce SPLADE-v2 to systematically investigate wacky weights across the SPLADE family of models. We present a comprehensive dissection of wacky weights, providing a formal definition of wackiness based on the lexical utility of expansion terms. Furthermore, we introduce a novel measure to compare the prevalence of these tokens across models with varying vocabularies and sparsity levels. Beyond reproducing the original SPLADE-v2, we train it with various loss functions, datasets, and backbone transformers to isolate the factors contributing to wackiness. Our results show that larger vocabularies are associated with a higher prevalence of wacky tokens, while stricter sparsity regularizers are associated with lower prevalence. Finally, we find that wacky weights are used primarily for in-domain effectiveness rather than out-of-domain generalization. Gregory Polyakov, Harrisen Scells, Carsten Eickhoff |
SIGIR | 1 |
| 2025 | Interpretability Analysis of Arithmetic In-Context Learning in Large Language ModelsabstractLarge language models (LLMs) exhibit sophisticated behavior, notably solving arithmetic with only a few in-context examples (ICEs).Yet the computations that connect those examples to the answer remain opaque.We probe four open-weight LLMs, Pythia-12B, Llama-3.1-8B,MPT-7B, and OPT-6.7B, on basic arithmetic to illustrate how they process ICEs.Our study integrates activation patching, information-flow analysis, automatic circuit discovery, and the logit-lens perspective into a unified pipeline.Within this framework we isolate partial-sum representations in three-operand tasks, investigate their influence on final logits, and derive linear function vectors that characterize tasks and align with ICEinduced activations.Controlled ablations show that strict pattern consistency in the formatting of ICEs guides the models more strongly than the symbols chosen or even the factual correctness of the examples.By unifying four complementary interpretability tools, this work delivers one of the most comprehensive interpretability studies of LLM arithmetic to date, and the first on three-operand tasks.Our code is publicly available 1 . Gregory Polyakov, Christian Hepting, Carsten Eickhoff, Seyed Ali Bahrainian |
EMNLP | 1 |
| 2025 | Towards Best Practices of Axiomatic Activation Patching in Information RetrievalabstractMechanistic interpretability research, which aims to uncover the internal processes of machine learning models, has gained significant attention. One state-of-the-art technique, activation patching, has been applied to analyzing neural ranker behavior in relation to information retrieval (IR) axioms. To date, however, this remains a rapidly evolving topic in IR, with no established methodology for measuring results or constructing datasets to ensure pronounced, robust, and consistent patching effects. In this study, based on experimental results, we provide recommendations on measuring patching effects and designing diagnostic datasets for investigating term frequency. We identify the rareness and informativeness of injected terms as a key factor influencing the magnitude of patching effects. Additionally, we find that low score differences between baseline and perturbed documents introduce significant noise, which can be mitigated by filtering or applying penalty scores to the metric. More generally, we provide practical recommendations for the reliable application of activation patching in IR, advancing future interpretability research of neural ranking models. Our code is available at https://github.com/polgrisha/best-practices-ir-patching. Gregory Polyakov, Catherine Chen 0001, Carsten Eickhoff |
SIGIR | 1 |
| 2023 | Sinkhorn Transformations for Single-Query Postprocessing in Text-Video RetrievalabstractA recent trend in multimodal retrieval is related to postprocessing test set results via the dual-softmax loss (DSL). While this approach can bring significant improvements, it usually presumes that an entire matrix of test samples is available as DSL input. This work introduces a new postprocessing approach based on Sinkhorn transformations that outperforms DSL. Further, we propose a new postprocessing setting that does not require access to multiple test queries. We show that our approach can significantly improve the results of state of the art models such as CLIP4Clip, BLIP, X-CLIP, and DRL, thus achieving a new state-of-the-art on several standard text-video retrieval datasets both with access to the entire test set and in the single-query setting. Konstantin Yakovlev, Gregory Polyakov, Ilseyar Alimova, Alexander Podolskiy, Andrey Bout, Sergey I. Nikolenko, Irina Piontkovskaya |
SIGIR | 2 |