Sebastian Goodman

dblp:192/1217 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Vision and language · 34% Language models and text generation · 22% Efficient and distributed learning · 11%

Topics — the 23 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language pretraining
0.922023
PreSTU: Pre-Training for Scene-Text Understanding · ICCV 2023
PaLI: A Jointly-Scaled Multilingual Language-Image Model · ICLR 2023
Machine learning › Optimization for machine learning › convergence analysis
convergence analysis of transformers
0.812024
CausalLM is not optimal for in-context learning · ICLR 2024
Natural language and speech › Language models and text generation
in-context learning
0.812024
CausalLM is not optimal for in-context learning · ICLR 2024
Computer vision › Vision and language › vision-language model › vision-language foundation model
multilingual vision-language models
0.812024
On Scaling Up a Multilingual Vision and Language Model · CVPR 2024
Machine learning › Efficient and distributed learning › distributed training
distributed training systems
0.712023
Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training
0.712023
Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023
Natural language and speech › Language models and text generation
multilingual language models
0.712023
PaLI: A Jointly-Scaled Multilingual Language-Image Model · ICLR 2023
Computer vision › Vision and language › multimodal understanding
scene text understanding
0.712023
PreSTU: Pre-Training for Scene-Text Understanding · ICCV 2023
Computer vision › Vision and language
image captioning
0.622024
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning · ACL (1) 2018
On Scaling Up a Multilingual Vision and Language Model · CVPR 2024
Machine learning › Transfer learning and domain adaptation › meta-learning
few-shot meta-learning
0.512021
Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning · NeurIPS 2021
Machine learning › Transfer learning and domain adaptation
meta-learning
0.512021
Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning · NeurIPS 2021
Machine learning › Learning theory
PAC-Bayesian analysis
0.512021
Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning · NeurIPS 2021
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
exposure bias mitigation
0.412020
TeaForN: Teacher-Forcing with N-grams · EMNLP (1) 2020
Natural language and speech › Language models and text generation
pre-trained language model
0.412020
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations · ICLR 2020
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation
0.412020
TeaForN: Teacher-Forcing with N-grams · EMNLP (1) 2020
Natural language and speech › Language models and text generation
teacher forcing
0.412020
TeaForN: Teacher-Forcing with N-grams · EMNLP (1) 2020
Computer vision › Vision and language
caption generation
0.312018
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning · ACL (1) 2018
Computer vision › Image recognition and object detection
object detection
0.212024
On Scaling Up a Multilingual Vision and Language Model · CVPR 2024
Computer vision › Vision and language
visual question answering
0.212024
On Scaling Up a Multilingual Vision and Language Model · CVPR 2024
Natural language and speech › Language models and text generation
large language model
0.212023
Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models
bayesian deep learning
0.112021
Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning · NeurIPS 2021
Natural language and speech › Machine translation
neural machine translation
0.112020
TeaForN: Teacher-Forcing with N-grams · EMNLP (1) 2020
Natural language and speech › Language models and text generation
text summarization
0.112020
TeaForN: Teacher-Forcing with N-grams · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

scaling · 0.8online gradient descent analysis · 0.8multimodal pretraining · 0.8linear regression analysis · 0.8transformer encoder-decoder · 0.7transfer learning · 0.7joint scaling · 0.7OCR · 0.7PAC-Bayes bounds · 0.5MAML · 0.5
YearPublicationVenuePosition
2024 On Scaling Up a Multilingual Vision and Language Model
abstract
We explore the boundaries of scaling up a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture. Our model achieves new levels of performance on a wide-range of varied and complex tasks, including multiple image-based captioning and question-answering tasks, image-based document understanding and few-shot (in-context) learning, as well as object detection, video question answering, and video captioning. Our model advances the state-of-the-art on most vision-and-language benchmarks considered (20+ of them). Finally, we observe emerging capabilities, such as complex counting and multilingual object detection, tasks that are not explicitly in the training mix.
Xi Chen 0071, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Carlos Riquelme, Sebastian Goodman, Xiao Wang 0038, Yi Tay, Siamak Shakeri, Mostafa Dehghani 0001, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang 0001, Ceslee Montgomery, Paulina Pietrzyk, Marvin Ritter, A. J. Piergiovanni, Matthias Minderer, Filip Pavetic, Austin Waters, Gang Li 0021, Ibrahim Alabdulmohsin, Lucas Beyer, Julien Amelot, Kenton Lee, Andreas Steiner 0001, Yang Li 0058, Daniel Keysers, Anurag Arnab, Yuanzhong Xu, Keran Rong, Alexander Kolesnikov 0003, Mojtaba Seyedhosseini, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, Radu Soricut
CVPR8
2024 CausalLM is not optimal for in-context learning
abstract
Recent empirical evidence indicates that transformer based in-context learning performs better when using a prefix language model (prefixLM), in which in-context samples can all attend to each other, compared to causal language models (causalLM), which use auto-regressive attention that prohibits in-context samples to attend to future samples. While this result is intuitive, it is not understood from a theoretical perspective. In this paper we take a theoretical approach and analyze the convergence behavior of prefixLM and causalLM under a certain parameter construction. Our analysis shows that both LM types converge to their stationary points at a linear rate, but that while prefixLM converges to the optimal solution of linear regression, causalLM convergence dynamics follows that of an online gradient descent algorithm, which is not guaranteed to be optimal even as the number of samples grows infinitely. We supplement our theoretical claims with empirical experiments over synthetic and real tasks and using various types of transformers. Our experiments verify that causalLM consistently underperforms prefixLM in all settings.
Nan Ding 0002, Tomer Levinboim, Sebastian Goodman, Radu Soricut
ICLR4
2023 PreSTU: Pre-Training for Scene-Text Understanding
abstract
The ability to recognize and reason about text embedded in visual inputs is often lacking in vision-and-language (V&L) models, perhaps because V&L pre-training methods have often failed to include such an ability in their training objective. In this paper, we propose PreSTU, a novel pre-training recipe dedicated to scene-text understanding (STU). PreSTU introduces OCR-aware pre-training objectives that encourage the model to recognize text from an image and connect it to the rest of the image content. We implement PreSTU using a simple transformer-based encoder-decoder architecture, combined with large-scale image-text datasets with scene text obtained from an off-the-shelf OCR system. We empirically demonstrate the effectiveness of this pre-training approach on eight visual question answering and four image captioning benchmarks.
Jihyung Kil, Soravit Changpinyo, Xi Chen 0071, Hexiang Hu, Sebastian Goodman, Wei-Lun Chao, Radu Soricut
ICCV5
2023 PaLI: A Jointly-Scaled Multilingual Language-Image Model
Xi Chen 0071, Xiao Wang 0038, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov 0003, Joan Puigcerver, Nan Ding 0002, Keran Rong, Hassan Akbari, Linting Xue, Ashish V. Thapliyal, Weicheng Kuo
ICLR7
2023 Scaling Up Models and Data with t5x and seqio
abstract
Scaling up training datasets and model parameters have benefited neural network-based language models, but also present challenges like distributed compute, input data bottlenecks and reproducibility of results. We introduce two simple and scalable software libraries that simplify these issues: t5x enables training large language models at scale, while seqio enables reproducible input and evaluation pipelines. These open-source libraries have been used to train models with hundreds of billions of parameters on multi-terabyte datasets. Configurations and instructions for T5-like and GPT-like models are also provided. The libraries can be found at https://github.com/google-research/t5x and https://github.com/google/seqio.
Adam Roberts, Hyung Won Chung, Anselm Levskaya, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio B. Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Kathleen Kenealy, Kehang Han, Michelle Casbon, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Tachard Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, Andrea Gesmundo
J. Mach. Learn. Res.16
2021 Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning
abstract
Despite recent advances in its theoretical understanding, there still remains a significant gap in the ability of existing PAC-Bayesian theories on meta-learning to explain performance improvements in the few-shot learning setting, where the number of training examples in the target tasks is severely limited. This gap originates from an assumption in the existing theories which supposes that the number of training examples in the observed tasks and the number of training examples in the target tasks follow the same distribution, an assumption that rarely holds in practice. By relaxing this assumption, we develop two PAC-Bayesian bounds tailored for the few-shot learning setting and show that two existing meta-learning algorithms (MAML and Reptile) can be derived from our bounds, thereby bridging the gap between practice and PAC-Bayesian theories. Furthermore, we derive a new computationally-efficient PACMAML algorithm, and show it outperforms existing meta-learning algorithms on several few-shot benchmark datasets.
Nan Ding 0002, Xi Chen 0071, Tomer Levinboim, Sebastian Goodman, Radu Soricut
NeurIPS4
2020 TeaForN: Teacher-Forcing with N-grams
abstract
Sequence generation models trained with teacher-forcing suffer from issues related to exposure bias and lack of differentiability across timesteps.Our proposed method, Teacher-Forcing with N-grams (TeaForN), addresses both these problems directly, through the use of a stack of N decoders trained to decode along a secondary time axis that allows modelparameter updates based on N prediction steps.TeaForN can be used with a wide class of decoder architectures and requires minimal modifications from a standard teacher-forcing setup.Empirically, we show that TeaForN boosts generation quality on one Machine Translation benchmark, WMT 2014 English-French, and two News Summarization benchmarks, CNN/Dailymail and Gigaword.
Sebastian Goodman, Nan Ding 0002, Radu Soricut
EMNLP (1)1
2020 ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhen-Zhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut
ICLR3
2018 Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
abstract
We present a new dataset of image caption annotations, Conceptual Captions, which contains an order of magnitude more images than the MS-COCO dataset (Lin et al., 2014) and represents a wider variety of both images and image caption styles.We achieve this by extracting and filtering image caption annotations from billions of webpages.We also present quantitative evaluations of a number of image captioning models and show that a model architecture based on Inception-ResNet-v2 (Szegedy et al., 2016) for image-feature extraction and Transformer (Vaswani et al., 2017) for sequence modeling achieves the best performance when trained on the Conceptual Captions dataset.
Piyush Sharma, Nan Ding 0002, Sebastian Goodman, Radu Soricut
ACL (1)3