VLDB 2026 Research / reviewers in the wild / expert
Xi Chen 0071
dblp:16/3283-71
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2024
0000-0002-1581-4627ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | On Scaling Up a Multilingual Vision and Language ModelabstractWe explore the boundaries of scaling up a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture. Our model achieves new levels of performance on a wide-range of varied and complex tasks, including multiple image-based captioning and question-answering tasks, image-based document understanding and few-shot (in-context) learning, as well as object detection, video question answering, and video captioning. Our model advances the state-of-the-art on most vision-and-language benchmarks considered (20+ of them). Finally, we observe emerging capabilities, such as complex counting and multilingual object detection, tasks that are not explicitly in the training mix. Xi Chen 0071, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Carlos Riquelme, Sebastian Goodman, Xiao Wang 0038, Yi Tay, Siamak Shakeri, Mostafa Dehghani 0001, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang 0001, Ceslee Montgomery, Paulina Pietrzyk, Marvin Ritter, A. J. Piergiovanni, Matthias Minderer, Filip Pavetic, Austin Waters, Gang Li 0021, Ibrahim Alabdulmohsin, Lucas Beyer, Julien Amelot, Kenton Lee, Andreas Steiner 0001, Yang Li 0058, Daniel Keysers, Anurag Arnab, Yuanzhong Xu, Keran Rong, Alexander Kolesnikov 0003, Mojtaba Seyedhosseini, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, Radu Soricut |
CVPR | 1 |
| 2024 | Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
Brian Gordon, Yonatan Bitton, Yonatan Shafir, Roopal Garg, Xi Chen 0071, Dani Lischinski, Daniel Cohen-Or, Idan Szpektor |
ECCV (57) | 5 |
| 2023 | Improving Robust Generalization by Direct PAC-Bayesian Bound MinimizationabstractRecent research in robust optimization has shown an overfitting-like phenomenon in which models trained against adversarial attacks exhibit higher robustness on the training set compared to the test set. Although previous work provided theoretical explanations for this phenomenon using a robust PAC-Bayesian bound over the adversarial test error, related algorithmic derivations are at best only loosely connected to this bound, which implies that there is still a gap between their empirical success and our understanding of adversarial robustness theory. To close this gap, in this paper we consider a different form of the robust PAC-Bayesian bound and directly minimize it with respect to the model posterior. The derivation of the optimal solution connects PAC-Bayesian learning to the geometry of the robust loss surface through a Trace of Hessian (TrH) regularizer that measures the surface flatness. In practice, we restrict the TrH regularizer to the top layer only, which results in an analytical solution to the bound whose computational cost does not depend on the network depth. Finally, we evaluate our TrH regularization approach over CIFAR-10/100 and ImageNet using Vision Transformers (ViT) and compare against baseline adversarial robustness algorithms. Experimental results show that TrH regularization leads to improved ViT robustness that either matches or surpasses previous state-of-the-art approaches while at the same time requires less memory and computational cost. Zifan Wang 0001, Nan Ding 0002, Tomer Levinboim, Xi Chen 0071, Radu Soricut |
CVPR | 4 |
| 2023 | PreSTU: Pre-Training for Scene-Text UnderstandingabstractThe ability to recognize and reason about text embedded in visual inputs is often lacking in vision-and-language (V&L) models, perhaps because V&L pre-training methods have often failed to include such an ability in their training objective. In this paper, we propose PreSTU, a novel pre-training recipe dedicated to scene-text understanding (STU). PreSTU introduces OCR-aware pre-training objectives that encourage the model to recognize text from an image and connect it to the rest of the image content. We implement PreSTU using a simple transformer-based encoder-decoder architecture, combined with large-scale image-text datasets with scene text obtained from an off-the-shelf OCR system. We empirically demonstrate the effectiveness of this pre-training approach on eight visual question answering and four image captioning benchmarks. Jihyung Kil, Soravit Changpinyo, Xi Chen 0071, Hexiang Hu, Sebastian Goodman, Wei-Lun Chao, Radu Soricut |
ICCV | 3 |
| 2023 | PaLI: A Jointly-Scaled Multilingual Language-Image Model
Xi Chen 0071, Xiao Wang 0038, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov 0003, Joan Puigcerver, Nan Ding 0002, Keran Rong, Hassan Akbari, Linting Xue, Ashish V. Thapliyal, Weicheng Kuo |
ICLR | 1 |
| 2022 | PACTran: PAC-Bayesian Metrics for Estimating the Transferability of Pretrained Models to Classification Tasks
Nan Ding 0002, Xi Chen 0071, Tomer Levinboim, Soravit Changpinyo, Radu Soricut |
ECCV (34) | 2 |
| 2022 | MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and TextabstractWhile language Models store a massive amount of world knowledge implicitly in their parameters, even very large models often fail to encode information about rare entities and events, while incurring huge computational costs.Recently, retrieval-augmented models, such as REALM, RAG, and RETRO, have incorporated world knowledge into language generation by leveraging an external nonparametric index and have demonstrated impressive performance with constrained model sizes.However, these methods are restricted to retrieving only textual knowledge, neglecting the ubiquitous amount of knowledge in other modalities like images -much of which contains information not covered by any text.To address this limitation, we propose the first Multimodal Retrieval-Augmented Transformer (MuRAG), which accesses an external non-parametric multimodal memory to augment language generation.MuRAG is pretrained with a mixture of large-scale imagetext and text-only corpora using a joint contrastive and generative loss.We perform experiments on two different datasets that require retrieving and reasoning over both images and text to answer a given query: We-bQA, and MultimodalQA.Our results show that MuRAG achieves state-of-the-art accuracy, outperforming existing models by 10-20% absolute on both datasets and under both distractor and full-wiki settings. Wenhu Chen, Hexiang Hu, Xi Chen 0071, Patrick Verga, William W. Cohen |
EMNLP | 3 |
| 2022 | Crossmodal-3600: A Massively Multilingual Multimodal Evaluation DatasetabstractResearch in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.In this paper we present the Crossmodal-3600 dataset (XM3600 in short), a geographically-diverse set of 3600 images annotated with humangenerated reference captions in 36 languages.The images were selected from across the world, covering regions where the 36 languages are spoken, and annotated with captions that achieve consistency in terms of style across all languages, while avoiding annotation artifacts due to direct translation.We apply this benchmark to model selection for massively multilingual image captioning models, and show strong correlation results with human evaluations when using XM3600 as golden references for automatic metrics. Ashish V. Thapliyal, Jordi Pont-Tuset, Xi Chen 0071, Radu Soricut |
EMNLP | 3 |
| 2022 | All You May Need for VQA are Image CaptionsabstractSoravit Changpinyo, Doron Kukliansy, Idan Szpektor, Xi Chen, Nan Ding, Radu Soricut. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Soravit Changpinyo, Doron Kukliansky, Idan Szpektor, Xi Chen 0071, Nan Ding 0002, Radu Soricut |
NAACL-HLT | 4 |
| 2021 | Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-LearningabstractDespite recent advances in its theoretical understanding, there still remains a significant gap in the ability of existing PAC-Bayesian theories on meta-learning to explain performance improvements in the few-shot learning setting, where the number of training examples in the target tasks is severely limited. This gap originates from an assumption in the existing theories which supposes that the number of training examples in the observed tasks and the number of training examples in the target tasks follow the same distribution, an assumption that rarely holds in practice. By relaxing this assumption, we develop two PAC-Bayesian bounds tailored for the few-shot learning setting and show that two existing meta-learning algorithms (MAML and Reptile) can be derived from our bounds, thereby bridging the gap between practice and PAC-Bayesian theories. Furthermore, we derive a new computationally-efficient PACMAML algorithm, and show it outperforms existing meta-learning algorithms on several few-shot benchmark datasets. Nan Ding 0002, Xi Chen 0071, Tomer Levinboim, Sebastian Goodman, Radu Soricut |
NeurIPS | 2 |