Till Speicher

dblp:144/7849 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0000-1172-2525ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 31% Trustworthy machine learning · 22% Representation and self-supervised learning · 15%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
in-context learning
1.322026
Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective · ACL (1) 2026
Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction · WSDM 2025
Machine learning › Transfer learning and domain adaptation
fine-tuning
1.012026
Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective · ACL (1) 2026
Machine learning › Learning theory › inductive inference
formal language learning
1.012026
Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective · ACL (1) 2026
Machine learning › Efficient and distributed learning
model compression
0.712023
Diffused Redundancy in Pre-trained Representations · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation analysis
representation similarity
0.612022
Measuring Representational Robustness of Neural Networks Through Shared Invariances · ICML 2022
Machine learning › Trustworthy machine learning
robustness
0.612022
Measuring Representational Robustness of Neural Networks Through Shared Invariances · ICML 2022
Machine learning › Trustworthy machine learning › robustness
robust representations
0.612022
Measuring Representational Robustness of Neural Networks Through Shared Invariances · ICML 2022
Machine learning › Trustworthy machine learning
fairness
0.312018
A Unified Approach to Quantifying Algorithmic Unfairness: Measuring Individual &Group Unfairness via Inequality Indices · KDD 2018
Machine learning › Trustworthy machine learning › fairness › fairness criteria
individual and group fairness
0.312018
A Unified Approach to Quantifying Algorithmic Unfairness: Measuring Individual &Group Unfairness via Inequality Indices · KDD 2018
Natural language and speech › Language models and text generation › language modeling
n-gram language model
0.212014
A Generalized Language Model as the Combination of Skipped n-grams and Modified Kneser Ney Smoothing · ACL (1) 2014
Natural language and speech › Language models and text generation › language modeling
smoothing
0.212014
A Generalized Language Model as the Combination of Skipped n-grams and Modified Kneser Ney Smoothing · ACL (1) 2014
Computational social science and digital humanities
algorithmic decision-making
0.112018
A Unified Approach to Quantifying Algorithmic Unfairness: Measuring Individual &Group Unfairness via Inequality Indices · KDD 2018

Methods — techniques the papers use, named apart from their topics

formal language theory · 1.0zero-prompt probing · 0.9in-context learning · 0.9transformer · 0.7resnet-50 · 0.7linear probe · 0.7inequality indices · 0.7representation similarity measures · 0.6
YearPublicationVenuePosition
2026 Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective
abstract
Bishwamittra Ghosh, Soumi Das, Till Speicher, Qinyuan Wu, Mohammad Aflah Khan, Deepak Garg, Krishna P. Gummadi, Evimaria Terzi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Bishwamittra Ghosh, Soumi Das, Till Speicher, Qinyuan Wu, Mohammad Aflah Khan, Krishna P. Gummadi, Evimaria Terzi
ACL (1)3
2025 Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction
abstract
In this paper, we focus on the challenging task of reliably estimating factual knowledge that is embedded inside large language models (LLMs). To avoid reliability concerns with prior approaches, we propose to eliminate prompt engineering when probing LLMs for factual knowledge. Our approach, called Zero-Prompt Latent Knowledge Estimator (ZP-LKE), leverages the in-context learning ability of LLMs to communicate both the factual knowledge question as well as the expected answer format. Our knowledge estimator is both conceptually simpler (i.e., doesn't depend on meta-linguistic judgments of LLMs) and easier to apply (i.e., is not LLM-specific), and we demonstrate that it can surface more of the latent knowledge embedded in LLMs. We also investigate how different design choices affect the performance of ZP-LKE. Using the proposed estimator, we perform a large-scale evaluation of the factual knowledge of a variety of open-source LLMs, like OPT, Pythia, Llama(2), Mistral, Gemma, etc. over a large set of relations and facts from the Wikidata knowledge base. We observe differences in the factual knowledge between different model families and models of different sizes, that some relations are consistently better known than others but that models differ in the precise facts they know, and differences in the knowledge of base models and their finetuned counterparts. Code available at: https://github.com/QinyuanWu0710/ZeroPrompt_LKE
Qinyuan Wu, Mohammad Aflah Khan, Soumi Das, Vedant Nanda, Bishwamittra Ghosh, Camila Kolling, Till Speicher, Laurent Bindschaedler, Krishna P. Gummadi, Evimaria Terzi
WSDM7
2023 Diffused Redundancy in Pre-trained Representations
abstract
Representations learned by pre-training a neural network on a large dataset are increasingly used successfully to perform a variety of downstream tasks. In this work, we take a closer look at how features are encoded in such pre-trained representations. We find that learned representations in a given layer exhibit a degree of diffuse redundancy, ie, any randomly chosen subset of neurons in the layer that is larger than a threshold size shares a large degree of similarity with the full layer and is able to perform similarly as the whole layer on a variety of downstream tasks. For example, a linear probe trained on $20\%$ of randomly picked neurons from the penultimate layer of a ResNet50 pre-trained on ImageNet1k achieves an accuracy within $5\%$ of a linear probe trained on the full layer of neurons for downstream CIFAR10 classification. We conduct experiments on different neural architectures (including CNNs and Transformers) pre-trained on both ImageNet1k and ImageNet21k and evaluate a variety of downstream tasks taken from the VTAB benchmark. We find that the loss \& dataset used during pre-training largely govern the degree of diffuse redundancy and the "critical mass" of neurons needed often depends on the downstream task, suggesting that there is a task-inherent redundancy-performance Pareto frontier. Our findings shed light on the nature of representations learned by pre-trained deep neural networks and suggest that entire layers might not be necessary to perform many downstream tasks. We investigate the potential for exploiting this redundancy to achieve efficient generalization for downstream tasks and also draw caution to certain possible unintended consequences. Our code is available at \url{https://github.com/nvedant07/diffused-redundancy}.
Vedant Nanda, Till Speicher, John Dickerson 0001, Krishna P. Gummadi, Soheil Feizi, Adrian Weller
NeurIPS2
2022 Measuring Representational Robustness of Neural Networks Through Shared Invariances
abstract
A major challenge in studying robustness in deep learning is defining the set of “meaningless” perturbations to which a given Neural Network (NN) should be invariant. Most work on robustness implicitly uses a human as the reference model to define such perturbations. Our work offers a new view on robustness by using another reference NN to define the set of perturbations a given NN should be invariant to, thus generalizing the reliance on a reference “human NN” to any NN. This makes measuring robustness equivalent to measuring the extent to which two NNs share invariances. We propose a measure called \stir, which faithfully captures the extent to which two NNs share invariances. \stir re-purposes existing representation similarity measures to make them suitable for measuring shared invariances. Using our measure, we are able to gain insights about how shared invariances vary with changes in weight initialization, architecture, loss functions, and training dataset. Our implementation is available at: \url{https://github.com/nvedant07/STIR}.
Vedant Nanda, Till Speicher, Camila Kolling, John Dickerson 0001, Krishna P. Gummadi, Adrian Weller
ICML2
2018 A Unified Approach to Quantifying Algorithmic Unfairness: Measuring Individual &Group Unfairness via Inequality Indices
abstract
Discrimination via algorithmic decision making has received considerable attention. Prior work largely focuses on defining conditions for fairness, but does not define satisfactory measures of algorithmic unfairness. In this paper, we focus on the following question: Given two unfair algorithms, how should we determine which of the two is more unfair? Our core idea is to use existing inequality indices from economics to measure how unequally the outcomes of an algorithm benefit different individuals or groups in a population. Our work offers a justified and general framework to compare and contrast the (un)fairness of algorithmic predictors. This unifying approach enables us to quantify unfairness both at the individual and the group level. Further, our work reveals overlooked tradeoffs between different fairness notions: using our proposed measures, the overall individual-level unfairness of an algorithm can be decomposed into a between-group and a within-group component. Earlier methods are typically designed to tackle only between-group un- fairness, which may be justified for legal or other reasons. However, we demonstrate that minimizing exclusively the between-group component may, in fact, increase the within-group, and hence the overall unfairness. We characterize and illustrate the tradeoffs between our measures of (un)fairness and the prediction accuracy.
Till Speicher, Hoda Heidari, Nina Grgic-Hlaca, Krishna P. Gummadi, Adish Singla, Adrian Weller, Muhammad Bilal Zafar
KDD1
2014 A Generalized Language Model as the Combination of Skipped n-grams and Modified Kneser Ney Smoothing
abstract
Rene Pickhardt, Thomas Gottron, Martin Körner, Paul Georg Wagner, Till Speicher, Steffen Staab. Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2014.
Rene Pickhardt, Thomas Gottron, Martin Körner, Paul Georg Wagner, Till Speicher, Steffen Staab
ACL (1)5