EDBT 2026 Demo / reviewers in the wild / expert
Sebastian Goodman
dblp:192/1217
· DBLP profile ↗
9ranked-venue papers
1as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Vision and language · 34% Language models and text generation · 22% Efficient and distributed learning · 11% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
vision-language pretraining |
0.9 | 2 | 2023 | PreSTU: Pre-Training for Scene-Text Understanding · ICCV 2023 PaLI: A Jointly-Scaled Multilingual Language-Image Model · ICLR 2023 |
Machine learning › Optimization for machine learning › convergence analysis
convergence analysis of transformers |
0.8 | 1 | 2024 | CausalLM is not optimal for in-context learning · ICLR 2024 |
Natural language and speech › Language models and text generation
in-context learning |
0.8 | 1 | 2024 | CausalLM is not optimal for in-context learning · ICLR 2024 |
Computer vision › Vision and language › vision-language model › vision-language foundation model
multilingual vision-language models |
0.8 | 1 | 2024 | On Scaling Up a Multilingual Vision and Language Model · CVPR 2024 |
Machine learning › Efficient and distributed learning › distributed training
distributed training systems |
0.7 | 1 | 2023 | Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023 |
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training |
0.7 | 1 | 2023 | Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023 |
Natural language and speech › Language models and text generation
multilingual language models |
0.7 | 1 | 2023 | PaLI: A Jointly-Scaled Multilingual Language-Image Model · ICLR 2023 |
Computer vision › Vision and language › multimodal understanding
scene text understanding |
0.7 | 1 | 2023 | PreSTU: Pre-Training for Scene-Text Understanding · ICCV 2023 |
Computer vision › Vision and language
image captioning |
0.6 | 2 | 2024 | Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning · ACL (1) 2018 On Scaling Up a Multilingual Vision and Language Model · CVPR 2024 |
Machine learning › Transfer learning and domain adaptation › meta-learning
few-shot meta-learning |
0.5 | 1 | 2021 | Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning · NeurIPS 2021 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.5 | 1 | 2021 | Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning · NeurIPS 2021 |
Machine learning › Learning theory
PAC-Bayesian analysis |
0.5 | 1 | 2021 | Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
exposure bias mitigation |
0.4 | 1 | 2020 | TeaForN: Teacher-Forcing with N-grams · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.4 | 1 | 2020 | ALBERT: A Lite BERT for Self-supervised Learning of Language Representations · ICLR 2020 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation |
0.4 | 1 | 2020 | TeaForN: Teacher-Forcing with N-grams · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation
teacher forcing |
0.4 | 1 | 2020 | TeaForN: Teacher-Forcing with N-grams · EMNLP (1) 2020 |
Computer vision › Vision and language
caption generation |
0.3 | 1 | 2018 | Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning · ACL (1) 2018 |
Computer vision › Image recognition and object detection
object detection |
0.2 | 1 | 2024 | On Scaling Up a Multilingual Vision and Language Model · CVPR 2024 |
Computer vision › Vision and language
visual question answering |
0.2 | 1 | 2024 | On Scaling Up a Multilingual Vision and Language Model · CVPR 2024 |
Natural language and speech › Language models and text generation
large language model |
0.2 | 1 | 2023 | Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023 |
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models
bayesian deep learning |
0.1 | 1 | 2021 | Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning · NeurIPS 2021 |
Natural language and speech › Machine translation
neural machine translation |
0.1 | 1 | 2020 | TeaForN: Teacher-Forcing with N-grams · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation
text summarization |
0.1 | 1 | 2020 | TeaForN: Teacher-Forcing with N-grams · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
scaling · 0.8online gradient descent analysis · 0.8multimodal pretraining · 0.8linear regression analysis · 0.8transformer encoder-decoder · 0.7transfer learning · 0.7joint scaling · 0.7OCR · 0.7PAC-Bayes bounds · 0.5MAML · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | On Scaling Up a Multilingual Vision and Language ModelabstractWe explore the boundaries of scaling up a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture. Our model achieves new levels of performance on a wide-range of varied and complex tasks, including multiple image-based captioning and question-answering tasks, image-based document understanding and few-shot (in-context) learning, as well as object detection, video question answering, and video captioning. Our model advances the state-of-the-art on most vision-and-language benchmarks considered (20+ of them). Finally, we observe emerging capabilities, such as complex counting and multilingual object detection, tasks that are not explicitly in the training mix. Xi Chen 0071, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Carlos Riquelme, Sebastian Goodman, Xiao Wang 0038, Yi Tay, Siamak Shakeri, Mostafa Dehghani 0001, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang 0001, Ceslee Montgomery, Paulina Pietrzyk, Marvin Ritter, A. J. Piergiovanni, Matthias Minderer, Filip Pavetic, Austin Waters, Gang Li 0021, Ibrahim Alabdulmohsin, Lucas Beyer, Julien Amelot, Kenton Lee, Andreas Steiner 0001, Yang Li 0058, Daniel Keysers, Anurag Arnab, Yuanzhong Xu, Keran Rong, Alexander Kolesnikov 0003, Mojtaba Seyedhosseini, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, Radu Soricut |
CVPR | 8 |
| 2024 | CausalLM is not optimal for in-context learningabstractRecent empirical evidence indicates that transformer based in-context learning performs better when using a prefix language model (prefixLM), in which in-context samples can all attend to each other, compared to causal language models (causalLM), which use auto-regressive attention that prohibits in-context samples to attend to future samples. While this result is intuitive, it is not understood from a theoretical perspective. In this paper we take a theoretical approach and analyze the convergence behavior of prefixLM and causalLM under a certain parameter construction. Our analysis shows that both LM types converge to their stationary points at a linear rate, but that while prefixLM converges to the optimal solution of linear regression, causalLM convergence dynamics follows that of an online gradient descent algorithm, which is not guaranteed to be optimal even as the number of samples grows infinitely. We supplement our theoretical claims with empirical experiments over synthetic and real tasks and using various types of transformers. Our experiments verify that causalLM consistently underperforms prefixLM in all settings. Nan Ding 0002, Tomer Levinboim, Sebastian Goodman, Radu Soricut |
ICLR | 4 |
| 2023 | PreSTU: Pre-Training for Scene-Text UnderstandingabstractThe ability to recognize and reason about text embedded in visual inputs is often lacking in vision-and-language (V&L) models, perhaps because V&L pre-training methods have often failed to include such an ability in their training objective. In this paper, we propose PreSTU, a novel pre-training recipe dedicated to scene-text understanding (STU). PreSTU introduces OCR-aware pre-training objectives that encourage the model to recognize text from an image and connect it to the rest of the image content. We implement PreSTU using a simple transformer-based encoder-decoder architecture, combined with large-scale image-text datasets with scene text obtained from an off-the-shelf OCR system. We empirically demonstrate the effectiveness of this pre-training approach on eight visual question answering and four image captioning benchmarks. Jihyung Kil, Soravit Changpinyo, Xi Chen 0071, Hexiang Hu, Sebastian Goodman, Wei-Lun Chao, Radu Soricut |
ICCV | 5 |
| 2023 | PaLI: A Jointly-Scaled Multilingual Language-Image Model
Xi Chen 0071, Xiao Wang 0038, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov 0003, Joan Puigcerver, Nan Ding 0002, Keran Rong, Hassan Akbari, Linting Xue, Ashish V. Thapliyal, Weicheng Kuo |
ICLR | 7 |
| 2023 | Scaling Up Models and Data with t5x and seqioabstractScaling up training datasets and model parameters have benefited neural network-based language models, but also present challenges like distributed compute, input data bottlenecks and reproducibility of results. We introduce two simple and scalable software libraries that simplify these issues: t5x enables training large language models at scale, while seqio enables reproducible input and evaluation pipelines. These open-source libraries have been used to train models with hundreds of billions of parameters on multi-terabyte datasets. Configurations and instructions for T5-like and GPT-like models are also provided. The libraries can be found at https://github.com/google-research/t5x and https://github.com/google/seqio. Adam Roberts, Hyung Won Chung, Anselm Levskaya, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio B. Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Kathleen Kenealy, Kehang Han, Michelle Casbon, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Tachard Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, Andrea Gesmundo |
J. Mach. Learn. Res. | 16 |
| 2021 | Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-LearningabstractDespite recent advances in its theoretical understanding, there still remains a significant gap in the ability of existing PAC-Bayesian theories on meta-learning to explain performance improvements in the few-shot learning setting, where the number of training examples in the target tasks is severely limited. This gap originates from an assumption in the existing theories which supposes that the number of training examples in the observed tasks and the number of training examples in the target tasks follow the same distribution, an assumption that rarely holds in practice. By relaxing this assumption, we develop two PAC-Bayesian bounds tailored for the few-shot learning setting and show that two existing meta-learning algorithms (MAML and Reptile) can be derived from our bounds, thereby bridging the gap between practice and PAC-Bayesian theories. Furthermore, we derive a new computationally-efficient PACMAML algorithm, and show it outperforms existing meta-learning algorithms on several few-shot benchmark datasets. Nan Ding 0002, Xi Chen 0071, Tomer Levinboim, Sebastian Goodman, Radu Soricut |
NeurIPS | 4 |
| 2020 | TeaForN: Teacher-Forcing with N-gramsabstractSequence generation models trained with teacher-forcing suffer from issues related to exposure bias and lack of differentiability across timesteps.Our proposed method, Teacher-Forcing with N-grams (TeaForN), addresses both these problems directly, through the use of a stack of N decoders trained to decode along a secondary time axis that allows modelparameter updates based on N prediction steps.TeaForN can be used with a wide class of decoder architectures and requires minimal modifications from a standard teacher-forcing setup.Empirically, we show that TeaForN boosts generation quality on one Machine Translation benchmark, WMT 2014 English-French, and two News Summarization benchmarks, CNN/Dailymail and Gigaword. Sebastian Goodman, Nan Ding 0002, Radu Soricut |
EMNLP (1) | 1 |
| 2020 | ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhen-Zhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut |
ICLR | 3 |
| 2018 | Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image CaptioningabstractWe present a new dataset of image caption annotations, Conceptual Captions, which contains an order of magnitude more images than the MS-COCO dataset (Lin et al., 2014) and represents a wider variety of both images and image caption styles.We achieve this by extracting and filtering image caption annotations from billions of webpages.We also present quantitative evaluations of a number of image captioning models and show that a model architecture based on Inception-ResNet-v2 (Szegedy et al., 2016) for image-feature extraction and Transformer (Vaswani et al., 2017) for sequence modeling achieves the best performance when trained on the Conceptual Captions dataset. Piyush Sharma, Nan Ding 0002, Sebastian Goodman, Radu Soricut |
ACL (1) | 3 |