EDBT 2026 Demo / reviewers in the wild / expert
Radu Soricut
dblp:83/3497
· DBLP profile ↗
48ranked-venue papers
11as first author
21since 2021 · last 2024
0000-0003-1565-3365ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 9 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | On Scaling Up a Multilingual Vision and Language ModelabstractWe explore the boundaries of scaling up a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture. Our model achieves new levels of performance on a wide-range of varied and complex tasks, including multiple image-based captioning and question-answering tasks, image-based document understanding and few-shot (in-context) learning, as well as object detection, video question answering, and video captioning. Our model advances the state-of-the-art on most vision-and-language benchmarks considered (20+ of them). Finally, we observe emerging capabilities, such as complex counting and multilingual object detection, tasks that are not explicitly in the training mix. Xi Chen 0071, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Carlos Riquelme, Sebastian Goodman, Xiao Wang 0038, Yi Tay, Siamak Shakeri, Mostafa Dehghani 0001, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang 0001, Ceslee Montgomery, Paulina Pietrzyk, Marvin Ritter, A. J. Piergiovanni, Matthias Minderer, Filip Pavetic, Austin Waters, Gang Li 0021, Ibrahim Alabdulmohsin, Lucas Beyer, Julien Amelot, Kenton Lee, Andreas Steiner 0001, Yang Li 0058, Daniel Keysers, Anurag Arnab, Yuanzhong Xu, Keran Rong, Alexander Kolesnikov 0003, Mojtaba Seyedhosseini, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, Radu Soricut |
CVPR | 43 |
| 2024 | Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-Rank ExpertsabstractLarge multi-modal models (LMMs) exhibit remarkable performance across numerous tasks. However, generalist LMMs often suffer from performance degradation when tuned over a large collection of tasks. Recent research suggests that Mixture of Experts (MoE) architectures are useful for instruction tuning, but for LMMs of parameter size around O(50-100B), the prohibitive cost of replicating and storing the expert models severely limits the number of experts we can use. We propose Omni-SMoLA, an architecture that uses the Soft MoE approach to (softly) mix many multimodal low rank experts, and avoids introducing a significant number of new parameters compared to conventional MoE models. The core intuition here is that the large model provides a foundational backbone, while different lightweight experts residually learn specialized knowledge, either per-modality or multimodally. Extensive experiments demonstrate that the SMoLA approach helps improve the generalist performance across a broad range of generative vision-and-language tasks, achieving new SoTA generalist performance that often matches or outperforms single specialized LMM baselines, as well as new SoTA specialist performance. Yaqing Wang 0007, Bo Pang 0001, Radu Soricut |
CVPR | 5 |
| 2024 | ImageInWords: Unlocking Hyper-Detailed Image DescriptionsabstractRoopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton, Ceslee Montgomery, Yasumasa Onoe, Andrew Bunner, Ranjay Krishna, Jason Michael Baldridge, Radu Soricut. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Roopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton, Ceslee Montgomery, Yasumasa Onoe, Andrew Bunner, Ranjay Krishna, Jason Baldridge, Radu Soricut |
EMNLP | 10 |
| 2024 | CausalLM is not optimal for in-context learningabstractRecent empirical evidence indicates that transformer based in-context learning performs better when using a prefix language model (prefixLM), in which in-context samples can all attend to each other, compared to causal language models (causalLM), which use auto-regressive attention that prohibits in-context samples to attend to future samples. While this result is intuitive, it is not understood from a theoretical perspective. In this paper we take a theoretical approach and analyze the convergence behavior of prefixLM and causalLM under a certain parameter construction. Our analysis shows that both LM types converge to their stationary points at a linear rate, but that while prefixLM converges to the optimal solution of linear regression, causalLM convergence dynamics follows that of an online gradient descent algorithm, which is not guaranteed to be optimal even as the number of samples grows infinitely. We supplement our theoretical claims with empirical experiments over synthetic and real tasks and using various types of transformers. Our experiments verify that causalLM consistently underperforms prefixLM in all settings. Nan Ding 0002, Tomer Levinboim, Sebastian Goodman, Radu Soricut |
ICLR | 5 |
| 2023 | Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image InpaintingabstractText-guided image editing can have a transformative impact in supporting creative applications. A key challenge is to generate edits that are faithful to input text prompts, while consistent with input images. We present Imagen Editor, a cascaded diffusion model built, by fine-tuning Imagen [36] on text-guided image inpainting. Imagen Editor's edits are faithful to the text prompts, which is accomplished by using object detectors to propose inpainting masks during training. In addition, Imagen Editor captures fine details in the input image by conditioning the cascaded pipeline on the original high resolution image. To improve qualitative and quantitative evaluation, we introduce EditBench, a systematic benchmark for text-guided image inpainting. EditBench evaluates inpainting edits on natural and generated images exploring objects, attributes, and scenes. Through extensive human evaluation on EditBench, we find that object-masking during training leads to across-the-board improvements in text-image alignment – such that Imagen Editor is preferred over DALL-E 2 [31] and Stable Diffusion [33] – and, as a cohort, these models are better at object-rendering than text-rendering, and handle material/color/size attributes better than count/shape attributes. Su Wang 0001, Chitwan Saharia, Ceslee Montgomery, Jordi Pont-Tuset, Shai Noy, Stefano Pellegrini, Yasumasa Onoe, Sarah Laszlo, David J. Fleet, Radu Soricut, Jason Baldridge, Mohammad Norouzi 0002 |
CVPR | 10 |
| 2023 | Connecting Vision and Language with Video Localized NarrativesabstractWe propose Video Localized Narratives, a new form of multimodal video annotations connecting vision and language. In the original Localized Narratives [36], annotators speak and move their mouse simultaneously on an image, thus grounding each word with a mouse trace segment. However, this is challenging on a video. Our new protocol empowers annotators to tell the story of a video with Localized Narratives, capturing even complex events involving multiple actors interacting with each other and with several passive objects. We annotated 20k videos of the OVIS, UVO, and Oops datasets, totalling 1.7M words. Based on this data, we also construct new benchmarks for the video narrative grounding and video question answering tasks, and provide reference results from strong baseline models. Our annotations are available at https://google.github.io/video-localized-narratives/ Paul Voigtlaender, Soravit Changpinyo, Jordi Pont-Tuset, Radu Soricut, Vittorio Ferrari |
CVPR | 4 |
| 2023 | Improving Robust Generalization by Direct PAC-Bayesian Bound MinimizationabstractRecent research in robust optimization has shown an overfitting-like phenomenon in which models trained against adversarial attacks exhibit higher robustness on the training set compared to the test set. Although previous work provided theoretical explanations for this phenomenon using a robust PAC-Bayesian bound over the adversarial test error, related algorithmic derivations are at best only loosely connected to this bound, which implies that there is still a gap between their empirical success and our understanding of adversarial robustness theory. To close this gap, in this paper we consider a different form of the robust PAC-Bayesian bound and directly minimize it with respect to the model posterior. The derivation of the optimal solution connects PAC-Bayesian learning to the geometry of the robust loss surface through a Trace of Hessian (TrH) regularizer that measures the surface flatness. In practice, we restrict the TrH regularizer to the top layer only, which results in an analytical solution to the bound whose computational cost does not depend on the network depth. Finally, we evaluate our TrH regularization approach over CIFAR-10/100 and ImageNet using Vision Transformers (ViT) and compare against baseline adversarial robustness algorithms. Experimental results show that TrH regularization leads to improved ViT robustness that either matches or surpasses previous state-of-the-art approaches while at the same time requires less memory and computational cost. Zifan Wang 0001, Nan Ding 0002, Tomer Levinboim, Xi Chen 0071, Radu Soricut |
CVPR | 5 |
| 2023 | PreSTU: Pre-Training for Scene-Text UnderstandingabstractThe ability to recognize and reason about text embedded in visual inputs is often lacking in vision-and-language (V&L) models, perhaps because V&L pre-training methods have often failed to include such an ability in their training objective. In this paper, we propose PreSTU, a novel pre-training recipe dedicated to scene-text understanding (STU). PreSTU introduces OCR-aware pre-training objectives that encourage the model to recognize text from an image and connect it to the rest of the image content. We implement PreSTU using a simple transformer-based encoder-decoder architecture, combined with large-scale image-text datasets with scene text obtained from an off-the-shelf OCR system. We empirically demonstrate the effectiveness of this pre-training approach on eight visual question answering and four image captioning benchmarks. Jihyung Kil, Soravit Changpinyo, Xi Chen 0071, Hexiang Hu, Sebastian Goodman, Wei-Lun Chao, Radu Soricut |
ICCV | 7 |
| 2022 | Denoising Large-Scale Image Captioning from Alt-text Data Using Content Selection ModelsabstractTraining large-scale image captioning (IC) models demands access to a rich and diverse set of training examples that are expensive to curate both in terms of time and man-power. Instead, alt-text based captions gathered from the web is a far cheaper alternative to scale with the downside of being noisy. Recent modeling approaches to IC often fall short in terms of performance in leveraging these noisy datasets in favor of clean annotations. We address this problem with a simple yet effective technique of breaking down the task into two smaller, more controllable tasks – skeleton prediction and skeleton-based caption generation. Specifically, we show that sub-selecting content words as skeletons helps in generating improved and denoised captions when leveraging rich yet noisy alt-text–based uncurated datasets. We also show that the predicted English skeletons can further cross-lingually be leveraged to generate non-English captions, and present experimental results covering caption generation in French, Italian, German, Spanish and Hindi. We also show that skeleton-based prediction allows for better control of certain caption properties, such as length, content, and gender expression, providing a handle to perform human-in-the-loop interpretable semi-automatic corrections. Khyathi Raghavi Chandu, Piyush Sharma, Soravit Changpinyo, Ashish V. Thapliyal, Radu Soricut |
COLING | 5 |
| 2022 | End-to-end Dense Video Captioning as Sequence GenerationabstractDense video captioning aims to identify the events of interest in an input video, and generate descriptive captions for each event. Previous approaches usually follow a two-stage generative process, which first proposes a segment for each event, then renders a caption for each identified segment. Recent advances in large-scale sequence generation pretraining have seen great success in unifying task formulation for a great variety of tasks, but so far, more complex tasks such as dense video captioning are not able to fully utilize this powerful paradigm. In this work, we show how to model the two subtasks of dense video captioning jointly as one sequence generation task, and simultaneously predict the events and the corresponding descriptions. Experiments on YouCook2 and ViTT show encouraging results and indicate the feasibility of training complex tasks such as end-to-end dense video captioning integrated into large-scale pretrained models. Wanrong Zhu, Bo Pang 0001, Ashish V. Thapliyal, William Yang Wang, Radu Soricut |
COLING | 5 |
| 2022 | PACTran: PAC-Bayesian Metrics for Estimating the Transferability of Pretrained Models to Classification Tasks
Nan Ding 0002, Xi Chen 0071, Tomer Levinboim, Soravit Changpinyo, Radu Soricut |
ECCV (34) | 5 |
| 2022 | Crossmodal-3600: A Massively Multilingual Multimodal Evaluation DatasetabstractResearch in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.In this paper we present the Crossmodal-3600 dataset (XM3600 in short), a geographically-diverse set of 3600 images annotated with humangenerated reference captions in 36 languages.The images were selected from across the world, covering regions where the 36 languages are spoken, and annotated with captions that achieve consistency in terms of style across all languages, while avoiding annotation artifacts due to direct translation.We apply this benchmark to model selection for massively multilingual image captioning models, and show strong correlation results with human evaluations when using XM3600 as golden references for automatic metrics. Ashish V. Thapliyal, Jordi Pont-Tuset, Xi Chen 0071, Radu Soricut |
EMNLP | 4 |
| 2022 | All You May Need for VQA are Image CaptionsabstractSoravit Changpinyo, Doron Kukliansy, Idan Szpektor, Xi Chen, Nan Ding, Radu Soricut. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Soravit Changpinyo, Doron Kukliansky, Idan Szpektor, Xi Chen 0071, Nan Ding 0002, Radu Soricut |
NAACL-HLT | 6 |
| 2022 | 2.5D visual relationship detection
Yu-Chuan Su, Soravit Changpinyo, Xiangning Chen, Sathish Thoppay, Cho-Jui Hsieh, Lior Shapira, Radu Soricut, Hartwig Adam, Matthew Brown 0001, Ming-Hsuan Yang 0001, Boqing Gong |
Comput. Vis. Image Underst. | 7 |
| 2021 | H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for SequencesabstractZhenhai Zhu, Radu Soricut. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhenhai Zhu, Radu Soricut |
ACL/IJCNLP (1) | 2 |
| 2021 | Understanding Guided Image Captioning Performance across DomainsabstractImage captioning models generally lack the capability to take into account user interest, and usually default to global descriptions that try to balance readability, informativeness, and information overload.We present a Transformerbased model with the ability to produce captions focused on specific objects, concepts or actions in an image by providing them as guiding text to the model.Further, we evaluate the quality of these guided captions when trained on Conceptual Captions which contain 3.3M image-level captions compared to Visual Genome which contain 3.6M object-level captions.Counter-intuitively, we find that guided captions produced by the model trained on Conceptual Captions generalize better on outof-domain data.Our human-evaluation results indicate that attempting in-the-wild guided image captioning requires access to large, unrestricted-domain training datasets, and that increased style diversity (even without increasing the number of unique tokens) is a key factor for improved performance. Edwin G. Ng, Bo Pang 0001, Piyush Sharma, Radu Soricut |
CoNLL | 4 |
| 2021 | Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual ConceptsabstractThe availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pretraining. However, these datasets are often collected with overrestrictive requirements inherited from their original target tasks (e.g., image caption generation), which limit the resulting dataset scale and diversity. We take a step further in pushing the limits of vision-and-language pretraining data by relaxing the data collection pipeline used in Conceptual Captions 3M (CC3M) [54] and introduce the Conceptual 12M (CC12M), a dataset with 12 million image-text pairs specifically meant to be used for visionand-language pre-training. We perform an analysis of this dataset and benchmark its effectiveness against CC3M on multiple downstream tasks with an emphasis on long-tail visual recognition. Our results clearly illustrate the benefit of scaling up pre-training data for vision-and-language tasks, as indicated by the new state-of-the-art results on both the nocaps and Conceptual Captions benchmarks.1 Soravit Changpinyo, Piyush Sharma, Nan Ding 0002, Radu Soricut |
CVPR | 4 |
| 2021 | CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA GeneralizationabstractOne challenge in evaluating visual question answering (VQA) models in the cross-dataset adaptation setting is that the distribution shifts are multi-modal, making it difficult to identify if it is the shifts in visual or language features that play a key role.In this paper, we propose a semi-automatic framework for generating disentangled shifts by introducing a controllable visual question-answer generation (VQAG) module that is capable of generating highly-relevant and diverse questionanswer pairs with the desired dataset style.We use it to create CrossVQA, a collection of test splits for assessing VQA generalization based on the VQA2, VizWiz, and Open Images datasets.We provide an analysis of our generated datasets and demonstrate its utility by using them to evaluate several state-of-theart VQA systems.One important finding is that the visual shifts in cross-dataset VQA matter more than the language shifts.More broadly, we present a scalable framework for systematically evaluating the machine with little human intervention. Arjun R. Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma, Song-Chun Zhu, Radu Soricut |
EMNLP (1) | 6 |
| 2021 | Telling the What while Pointing to the Where: Multimodal Queries for Image RetrievalabstractMost existing image retrieval systems use text queries as a way for the user to express what they are looking for. However, fine-grained image retrieval often requires the ability to also express where in the image the content they are looking for is. The text modality can only cumbersomely express such localization preferences, whereas pointing is a more natural fit. In this paper, we propose an image retrieval setup with a new form of multimodal queries, where the user simultaneously uses both spoken natural language (the what) and mouse traces over an empty canvas (the where) to express the characteristics of the desired target image. We then describe simple modifications to an existing image retrieval model, enabling it to operate in this setup. Qualitative and quantitative experiments show that our model effectively takes this spatial guidance into account, and provides significantly more accurate retrieval results compared to text-only equivalent systems. Soravit Changpinyo, Jordi Pont-Tuset, Vittorio Ferrari, Radu Soricut |
ICCV | 4 |
| 2021 | Quality Estimation for Image Captions Based on Large-scale Human EvaluationsabstractTomer Levinboim, Ashish V. Thapliyal, Piyush Sharma, Radu Soricut. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tomer Levinboim, Ashish V. Thapliyal, Piyush Sharma, Radu Soricut |
NAACL-HLT | 4 |
| 2021 | Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-LearningabstractDespite recent advances in its theoretical understanding, there still remains a significant gap in the ability of existing PAC-Bayesian theories on meta-learning to explain performance improvements in the few-shot learning setting, where the number of training examples in the target tasks is severely limited. This gap originates from an assumption in the existing theories which supposes that the number of training examples in the observed tasks and the number of training examples in the target tasks follow the same distribution, an assumption that rarely holds in practice. By relaxing this assumption, we develop two PAC-Bayesian bounds tailored for the few-shot learning setting and show that two existing meta-learning algorithms (MAML and Reptile) can be derived from our bounds, thereby bridging the gap between practice and PAC-Bayesian theories. Furthermore, we derive a new computationally-efficient PACMAML algorithm, and show it outperforms existing meta-learning algorithms on several few-shot benchmark datasets. Nan Ding 0002, Xi Chen 0071, Tomer Levinboim, Sebastian Goodman, Radu Soricut |
NeurIPS | 5 |
| 2020 | Reinforcing an Image Caption Generator Using Off-Line Human FeedbackabstractHuman ratings are currently the most accurate way to assess the quality of an image captioning model, yet most often the only used outcome of an expensive human rating evaluation is a few overall statistics over the evaluation dataset. In this paper, we show that the signal from instance-level human caption ratings can be leveraged to improve captioning models, even when the amount of caption ratings is several orders of magnitude less than the caption training data. We employ a policy gradient method to maximize the human ratings as rewards in an off-policy reinforcement learning setting, where policy gradients are estimated by samples from a distribution that focuses on the captions in a caption ratings dataset. Our empirical evidence indicates that the proposed method learns to generalize the human raters' judgments to a previously unseen set of images, as judged by a different set of human judges, and additionally on a different, multi-dimensional side-by-side human evaluation procedure. Hongsuck Seo, Piyush Sharma, Tomer Levinboim, Bohyung Han, Radu Soricut |
AAAI | 5 |
| 2020 | Cross-modal Coherence Modeling for Caption GenerationabstractWe use coherence relations inspired by computational models of discourse to study the information needs and goals of image captioning.Using an annotation protocol specifically devised for capturing image-caption coherence relations, we annotate 10,000 instances from publicly-available image-caption pairs.We introduce a new task for learning inferences in imagery and text, coherence relation prediction, and show that these coherence annotations can be exploited to learn relation classifiers as an intermediary step, and also train coherence-aware, controllable image captioning models.The results show a dramatic improvement in the consistency and quality of the generated captions with respect to information needs specified via coherence relations. Malihe Alikhani, Piyush Sharma, Shengjie Li 0002, Radu Soricut, Matthew Stone |
ACL | 4 |
| 2020 | Cross-modal Language Generation using Pivot Stabilization for Web-scale Language CoverageabstractCross-modal language generation tasks such as image captioning are directly hurt in their ability to support non-English languages by the trend of data-hungry models combined with the lack of non-English annotations.We investigate potential solutions for combining existing language-generation annotations in English with translation capabilities in order to create solutions at web-scale in both domain and language coverage.We describe an approach called Pivot-Language Generation Stabilization (PLuGS), which leverages directly at training time both existing English annotations (gold data) as well as their machinetranslated versions (silver data); at run-time, it generates first an English caption and then a corresponding target-language caption.We show that PLuGS models outperform other candidate solutions in evaluations performed over 5 different target languages, under a largedomain testset using images from the Open Images dataset.Furthermore, we find an interesting effect where the English captions generated by the PLuGS models are better than the captions generated by the original, monolingual English model. Ashish V. Thapliyal, Radu Soricut |
ACL | 2 |
| 2020 | Connecting Vision and Language with Localized Narratives
Jordi Pont-Tuset, Jasper R. R. Uijlings, Soravit Changpinyo, Radu Soricut, Vittorio Ferrari |
ECCV (5) | 4 |
| 2020 | TeaForN: Teacher-Forcing with N-gramsabstractSequence generation models trained with teacher-forcing suffer from issues related to exposure bias and lack of differentiability across timesteps.Our proposed method, Teacher-Forcing with N-grams (TeaForN), addresses both these problems directly, through the use of a stack of N decoders trained to decode along a secondary time axis that allows modelparameter updates based on N prediction steps.TeaForN can be used with a wide class of decoder architectures and requires minimal modifications from a standard teacher-forcing setup.Empirically, we show that TeaForN boosts generation quality on one Machine Translation benchmark, WMT 2014 English-French, and two News Summarization benchmarks, CNN/Dailymail and Gigaword. Sebastian Goodman, Nan Ding 0002, Radu Soricut |
EMNLP (1) | 3 |
| 2020 | Beyond Instructional Videos: Probing for More Diverse Visual-Textual Grounding on YouTubeabstractPretraining from unlabelled web videos has quickly become the de-facto means of achieving high performance on many video understanding tasks.Features are learned via prediction of grounded relationships between visual content and automatic speech recognition (ASR) tokens.However, prior pretraining work has been limited to only instructional videos; a priori, we expect this domain to be relatively "easy:" speakers in instructional videos will often reference the literal objects/actions being depicted.We ask: can similar models be trained on more diverse video corpora?And, if so, what types of videos are "grounded" and what types are not?We fit a representative pretraining model to the diverse YouTube8M dataset, and study its success and failure cases.We find that visualtextual grounding is indeed possible across previously unexplored video categories, and that pretraining on a more diverse set results in representations that generalize to both noninstructional and instructional domains. Jack Hessel, Zhenhai Zhu, Bo Pang 0001, Radu Soricut |
EMNLP (1) | 4 |
| 2020 | ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhen-Zhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut |
ICLR | 6 |
| 2019 | Informative Image Captioning with External Sources of InformationabstractAn image caption should fluently present the essential information in a given image, including informative, fine-grained entity mentions and the manner in which these entities interact. However, current captioning models are usually trained to generate captions that only contain common object names, thus falling short on an important “informativeness” dimension. We present a mechanism for integrating image information together with fine-grained labels (assumed to be generated by some upstream models) into a caption that describes the image in a fluent and informative manner. We introduce a multimodal, multi-encoder model based on Transformer that ingests both image features and multiple sources of entity labels. We demonstrate that we can learn to control the appearance of these entity labels in the output, resulting in captions that are both fluent and informative. Sanqiang Zhao, Piyush Sharma, Tomer Levinboim, Radu Soricut |
ACL (1) | 4 |
| 2019 | A Case Study on Combining ASR and Visual Features for Generating Instructional Video CaptionsabstractInstructional videos get high-traffic on video sharing platforms, and prior work suggests that providing time-stamped, subtask annotations (e.g., "heat the oil in the pan") improves user experiences.However, current automatic annotation methods based on visual features alone perform only slightly better than constant prediction.Taking cues from prior work, we show that we can improve performance significantly by considering automatic speech recognition (ASR) tokens as input.Furthermore, jointly modeling ASR tokens and visual features results in higher performance compared to training individually on either modality.We find that unstated background information is better explained by visual features, whereas fine-grained distinctions (e.g., "add oil" vs. "add olive oil") are disambiguated more easily via ASR tokens. Jack Hessel, Bo Pang 0001, Zhenhai Zhu, Radu Soricut |
CoNLL | 4 |
| 2019 | Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels Improve Image Captioning and Visual Question AnsweringabstractSoravit Changpinyo, Bo Pang, Piyush Sharma, Radu Soricut. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Soravit Changpinyo, Bo Pang 0001, Piyush Sharma, Radu Soricut |
EMNLP/IJCNLP (1) | 4 |
| 2018 | Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image CaptioningabstractWe present a new dataset of image caption annotations, Conceptual Captions, which contains an order of magnitude more images than the MS-COCO dataset (Lin et al., 2014) and represents a wider variety of both images and image caption styles.We achieve this by extracting and filtering image caption annotations from billions of webpages.We also present quantitative evaluations of a number of image captioning models and show that a model architecture based on Inception-ResNet-v2 (Szegedy et al., 2016) for image-feature extraction and Transformer (Vaswani et al., 2017) for sequence modeling achieves the best performance when trained on the Conceptual Captions dataset. Piyush Sharma, Nan Ding 0002, Sebastian Goodman, Radu Soricut |
ACL (1) | 4 |
| 2018 | SHAPED: Shared-Private Encoder-Decoder for Text Style AdaptationabstractYe Zhang, Nan Ding, Radu Soricut. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Nan Ding 0002, Radu Soricut |
NAACL-HLT | 3 |
| 2017 | Cold-Start Reinforcement Learning with Softmax Policy GradientabstractPolicy-gradient approaches to reinforcement learning have two common and undesirable overhead procedures, namely warm-start training and sample variance reduction. In this paper, we describe a reinforcement learning method based on a softmax value function that requires neither of these procedures. Our method combines the advantages of policy-gradient methods with the efficiency and simplicity of maximum-likelihood approaches. We apply this new cold-start reinforcement learning method in training sequence generation models for structured output prediction problems. Empirical evidence validates this method on automatic summarization and image captioning tasks. Nan Ding 0002, Radu Soricut |
NIPS | 2 |
| 2016 | Morpho-syntactic Lexicon Generation Using Graph-based Semi-supervised LearningabstractMorpho-syntactic lexicons provide information about the morphological and syntactic roles of words in a language. Such lexicons are not available for all languages and even when available, their coverage can be limited. We present a graph-based semi-supervised learning method that uses the morphological, syntactic and semantic relations between words to automatically construct wide coverage lexicons from small seed sets. Our method is language-independent, and we show that we can expand a 1000 word seed lexicon to more than 100 times its size with high quality for 11 languages. In addition, the automatically created lexicons provide features that improve performance in two downstream tasks: morphological tagging and dependency parsing. Manaal Faruqui, Ryan T. McDonald, Radu Soricut |
Trans. Assoc. Comput. Linguistics | 3 |
| 2015 | Unsupervised Morphology Induction Using Word EmbeddingsabstractWe present a language agnostic, unsupervised method for inducing morphological transformations between words.The method relies on certain regularities manifest in highdimensional vector spaces.We show that this method is capable of discovering a wide range of morphological rules, which in turn are used to build morphological analyzers.We evaluate this method across six different languages and nine datasets, and show significant improvements across all languages. Radu Soricut, Franz Josef Och |
HLT-NAACL | 1 |
| 2013 | Quality estimation for machine translation: preface
Lucia Specia, Radu Soricut |
Mach. Transl. | 2 |
| 2010 | TrustRank: Inducing Trust in Automatic Translations via Ranking
Radu Soricut, Abdessamad Echihabi |
ACL | 1 |
| 2008 | Automatic Prediction of Parser Accuracy
Sujith Ravi, Kevin Knight, Radu Soricut |
EMNLP | 3 |
| 2007 | Abstractive headline generation using WIDL-expressions
Radu Soricut, Daniel Marcu |
Inf. Process. Manag. | 1 |
| 2006 | Stochastic Language Generation Using WIDL-Expressions and its Application in Machine Translation and SummarizationabstractWe propose WIDL-expressions as a flexible formalism that facilitates the integration of a generic sentence realization system within end-to-end language processing applications. WIDL-expressions represent compactly probability distributions over finite sets of candidate realizations, and have optimal algorithms for realization via interpolation with language model probability distributions. We show the effectiveness of a WIDL-based NLG system in two sentence realization tasks: automatic translation and headline generation. Radu Soricut, Daniel Marcu |
ACL | 1 |
| 2006 | Discourse Generation Using Utility-Trained Coherence Models
Radu Soricut, Daniel Marcu |
ACL | 1 |
| 2006 | Automatic question answering using the web: Beyond the Factoid
Radu Soricut, Eric Brill |
Inf. Retr. | 1 |
| 2005 | Natural Language Generation for Text-to-Text Applications Using an Information-Slim Representation
Radu Soricut |
AAAI | 1 |
| 2005 | Towards Developing Generation Algorithms for Text-to-Text ApplicationsabstractWe describe a new sentence realization framework for text-to-text applications. This framework uses IDL-expressions as a representation formalism, and a generation mechanism based on algorithms for intersecting IDL-expressions with probabilistic language models. We present both theoretical and empirical results concerning the correctness and efficiency of these algorithms. Radu Soricut, Daniel Marcu |
ACL | 1 |
| 2004 | A Unified Framework For Automatic Evaluation Using 4-Gram Co-occurrence StatisticsabstractIn this paper we propose a unified framework for automatic evaluation of NLP applications using N-gram co-occurrence statistics. The automatic evaluation metrics proposed to date for Machine Translation and Automatic Summarization are particular instances from the family of metrics we propose. We show that different members of the same family of metrics explain best the variations obtained with human evaluations, according to the application being evaluated (Machine Translation, Automatic Summarization, and Automatic Question Answering) and the evaluation guidelines used by humans for evaluating such applications. Radu Soricut, Eric Brill |
ACL | 1 |
| 2004 | Automatic Question Answering: Beyond the Factoid
Radu Soricut, Eric Brill |
HLT-NAACL | 1 |
| 2003 | Sentence Level Discourse Parsing using Syntactic and Lexical Information
Radu Soricut, Daniel Marcu |
HLT-NAACL | 1 |