Wojciech Gajewski

dblp:200/8068 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 34% Question answering and dialogue systems · 31% Efficient and distributed learning · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
question answering evaluation
0.612022
Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation · EMNLP 2022
Information retrieval
evaluation
0.612022
Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation · EMNLP 2022
Information retrieval
retrieval models
0.612022
Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation · EMNLP 2022
Machine learning › Efficient and distributed learning
model compression
0.512021
Sparse is Enough in Scaling Transformers · NeurIPS 2021
Machine learning › Deep learning architectures and training › transformer › efficient transformer
sparse transformer
0.512021
Sparse is Enough in Scaling Transformers · NeurIPS 2021
Machine learning › Deep learning architectures and training
transformer
0.512021
Sparse is Enough in Scaling Transformers · NeurIPS 2021
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering
0.312018
Ask the Right Questions: Active Question Reformulation with Reinforcement Learning · ICLR 2018
Machine learning › Reinforcement learning
policy learning
0.312018
Ask the Right Questions: Active Question Reformulation with Reinforcement Learning · ICLR 2018
Natural language and speech › Language models and text generation
large language model evaluation
0.212022
Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

human annotation · 1.1BERT matching · 1.1sparsity · 0.5attention · 0.5reinforcement learning · 0.3
YearPublicationVenuePosition
2022 Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation
abstract
The predictions of question answering (QA) systems are typically evaluated against manually annotated finite sets of one or more answers.This leads to a coverage limitation that results in underestimating the true performance of systems, and is typically addressed by extending over exact match (EM) with predefined rules or with the token-level F 1 measure.In this paper, we present the first systematic conceptual and data-driven analysis to examine the shortcomings of token-level equivalence measures.To this end, we define the asymmetric notion of answer equivalence (AE), accepting answers that are equivalent to or improve over the reference, and publish over 23k human judgments for candidates produced by multiple QA systems on SQuAD. 1 Through a careful analysis of this data, we reveal and quantify several concrete limitations of the F 1 measure, such as a false impression of graduality, or missing dependence on the question.Since collecting AE annotations for each evaluated model is expensive, we learn a BERT matching (BEM) measure to approximate this task.Being a simpler task than QA, we find BEM to provide significantly better AE approximations than F 1 , and to more accurately reflect the performance of systems.Finally, we demonstrate the practical utility of AE and BEM on the concrete application of minimal accurate prediction sets, reducing the number of required answers by up to ×2.6.
Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, Tal Schuster
EMNLP3
2021 Sparse is Enough in Scaling Transformers
abstract
Large Transformer models yield impressive results on many tasks, but are expensive to train, or even fine-tune, and so slow at decoding that their use and study becomes out of reach. We address this problem by leveraging sparsity. We study sparse variants for all layers in the Transformer and propose Scaling Transformers, a family of next generation Transformer models that use sparse layers to scale efficiently and perform unbatched decoding much faster than the standard Transformer as we scale up the model size. Surprisingly, the sparse layers are enough to obtain the same perplexity as the standard Transformer with the same number of parameters. We also integrate with prior sparsity approaches to attention and enable fast inference on long sequences even with limited memory. This results in performance competitive to the state-of-the-art on long text summarization.
Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, Jonni Kanerva
NeurIPS5
2018 Ask the Right Questions: Active Question Reformulation with Reinforcement Learning
Christian Buck, Jannis Bulian, Massimiliano Ciaramita, Wojciech Gajewski, Andrea Gesmundo, Neil Houlsby, Wei Wang 0236
ICLR4