Ashish V. Thapliyal

dblp:42/4147 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
6since 2021 · last 2023
0000-0002-7219-0515ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Theory of computation · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 44% Language models and text generation · 32% Representation and self-supervised learning · 12%
Theoretical computer science
1 paper
Information theory · 50% Quantum computing and quantum information · 33% Coding theory · 17%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
image captioning
1.022022
Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset · EMNLP 2022
Cross-modal Language Generation using Pivot Stabilization for Web-scale Language Coverage · ACL 2020
Natural language and speech › Language models and text generation
multilingual language models
0.712023
PaLI: A Jointly-Scaled Multilingual Language-Image Model · ICLR 2023
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation
0.712023
Emergence of Abstract State Representations in Embodied Sequence Modeling · EMNLP 2023
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning
0.712023
Emergence of Abstract State Representations in Embodied Sequence Modeling · EMNLP 2023
Computer vision › Vision and language › image captioning
cross-lingual image captioning
0.612022
Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset · EMNLP 2022
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation
0.612022
Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset · EMNLP 2022
Natural language and speech › Language models and text generation › text generation
multilingual text generation
0.412020
Cross-modal Language Generation using Pivot Stabilization for Web-scale Language Coverage · ACL 2020
Computer vision › Vision and language
vision-language pretraining
0.212023
PaLI: A Jointly-Scaled Multilingual Language-Image Model · ICLR 2023
Information theory
channel capacity
0.012002
Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002
Coding theory › channel coding
channel simulation
0.012002
Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002
Information theory › communication channels › channel models
discrete memoryless channel
0.012002
Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002
Quantum computing and quantum information › quantum channel capacity
entanglement-assisted capacity
0.012002
Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002
Quantum computing and quantum information
quantum channel capacity
0.012002
Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002
Information theory › channel capacity
reverse shannon theorem
0.012002
Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002

Methods — techniques the papers use, named apart from their topics

transformer sequence modeling · 0.7transfer learning · 0.7probing · 0.7joint scaling · 0.7human evaluation · 0.6automatic caption metrics · 0.6training with silver data · 0.4pivot stabilization · 0.4machine translation · 0.4entropy-based capacity formula · 0.0
YearPublicationVenuePosition
2023 Emergence of Abstract State Representations in Embodied Sequence Modeling
abstract
Decision making via sequence modeling aims to mimic the success of language models, where actions taken by an embodied agent are modeled as tokens to predict.Despite their promising performance, it remains unclear if embodied sequence modeling leads to the emergence of internal representations that represent the environmental state information.A model that lacks abstract state representations would be liable to make decisions based on surface statistics which fail to generalize.We take the BabyAI environment, a grid world in which language-conditioned navigation tasks are performed, and build a sequence modeling Transformer, which takes a language instruction, a sequence of actions, and environmental observations as its inputs.In order to investigate the emergence of abstract state representations, we design a "blindfolded" navigation task, where only the initial environmental layout, the language instruction, and the action sequence to complete the task are available for training.Our probing results show that intermediate environmental layouts can be reasonably reconstructed from the internal activations of a trained model, and that language instructions play a role in the reconstruction accuracy.Our results suggest that many key features of state representations can emerge via embodied sequence modeling, supporting an optimistic outlook for applications of sequence modeling objectives to more complex embodied decision-making domains.1
Tian Yun 0001, Zilai Zeng, Kunal Handa, Ashish V. Thapliyal, Bo Pang 0001, Ellie Pavlick, Chen Sun 0002
EMNLP4
2023 PaLI: A Jointly-Scaled Multilingual Language-Image Model
Xi Chen 0071, Xiao Wang 0038, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov 0003, Joan Puigcerver, Nan Ding 0002, Keran Rong, Hassan Akbari, Linting Xue, Ashish V. Thapliyal, Weicheng Kuo
ICLR18
2022 Denoising Large-Scale Image Captioning from Alt-text Data Using Content Selection Models
abstract
Training large-scale image captioning (IC) models demands access to a rich and diverse set of training examples that are expensive to curate both in terms of time and man-power. Instead, alt-text based captions gathered from the web is a far cheaper alternative to scale with the downside of being noisy. Recent modeling approaches to IC often fall short in terms of performance in leveraging these noisy datasets in favor of clean annotations. We address this problem with a simple yet effective technique of breaking down the task into two smaller, more controllable tasks – skeleton prediction and skeleton-based caption generation. Specifically, we show that sub-selecting content words as skeletons helps in generating improved and denoised captions when leveraging rich yet noisy alt-text–based uncurated datasets. We also show that the predicted English skeletons can further cross-lingually be leveraged to generate non-English captions, and present experimental results covering caption generation in French, Italian, German, Spanish and Hindi. We also show that skeleton-based prediction allows for better control of certain caption properties, such as length, content, and gender expression, providing a handle to perform human-in-the-loop interpretable semi-automatic corrections.
Khyathi Raghavi Chandu, Piyush Sharma, Soravit Changpinyo, Ashish V. Thapliyal, Radu Soricut
COLING4
2022 End-to-end Dense Video Captioning as Sequence Generation
abstract
Dense video captioning aims to identify the events of interest in an input video, and generate descriptive captions for each event. Previous approaches usually follow a two-stage generative process, which first proposes a segment for each event, then renders a caption for each identified segment. Recent advances in large-scale sequence generation pretraining have seen great success in unifying task formulation for a great variety of tasks, but so far, more complex tasks such as dense video captioning are not able to fully utilize this powerful paradigm. In this work, we show how to model the two subtasks of dense video captioning jointly as one sequence generation task, and simultaneously predict the events and the corresponding descriptions. Experiments on YouCook2 and ViTT show encouraging results and indicate the feasibility of training complex tasks such as end-to-end dense video captioning integrated into large-scale pretrained models.
Wanrong Zhu, Bo Pang 0001, Ashish V. Thapliyal, William Yang Wang, Radu Soricut
COLING3
2022 Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset
abstract
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.In this paper we present the Crossmodal-3600 dataset (XM3600 in short), a geographically-diverse set of 3600 images annotated with humangenerated reference captions in 36 languages.The images were selected from across the world, covering regions where the 36 languages are spoken, and annotated with captions that achieve consistency in terms of style across all languages, while avoiding annotation artifacts due to direct translation.We apply this benchmark to model selection for massively multilingual image captioning models, and show strong correlation results with human evaluations when using XM3600 as golden references for automatic metrics.
Ashish V. Thapliyal, Jordi Pont-Tuset, Xi Chen 0071, Radu Soricut
EMNLP1
2021 Quality Estimation for Image Captions Based on Large-scale Human Evaluations
abstract
Tomer Levinboim, Ashish V. Thapliyal, Piyush Sharma, Radu Soricut. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Tomer Levinboim, Ashish V. Thapliyal, Piyush Sharma, Radu Soricut
NAACL-HLT2
2020 Cross-modal Language Generation using Pivot Stabilization for Web-scale Language Coverage
abstract
Cross-modal language generation tasks such as image captioning are directly hurt in their ability to support non-English languages by the trend of data-hungry models combined with the lack of non-English annotations.We investigate potential solutions for combining existing language-generation annotations in English with translation capabilities in order to create solutions at web-scale in both domain and language coverage.We describe an approach called Pivot-Language Generation Stabilization (PLuGS), which leverages directly at training time both existing English annotations (gold data) as well as their machinetranslated versions (silver data); at run-time, it generates first an English caption and then a corresponding target-language caption.We show that PLuGS models outperform other candidate solutions in evaluations performed over 5 different target languages, under a largedomain testset using images from the Open Images dataset.Furthermore, we find an interesting effect where the English captions generated by the PLuGS models are better than the captions generated by the original, monolingual English model.
Ashish V. Thapliyal, Radu Soricut
ACL1
2003 Rank two bipartite bound entangled states do not exist
Pawel Horodecki, John A. Smolin, Barbara M. Terhal, Ashish V. Thapliyal
Theor. Comput. Sci.4
2002 Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem
abstract
The entanglement-assisted classical capacity of a noisy quantum channel (C/sub E/) is the amount of information per channel use that can be sent over the channel in the limit of many uses of the channel, assuming that the sender and receiver have access to the resource of shared quantum entanglement, which may be used up by the communication protocol. We show that the capacity C/sub E/ is given by an expression parallel to that for the capacity of a purely classical channel: i.e., the maximum, over channel inputs /spl rho/, of the entropy of the channel input plus the entropy of the channel output minus their joint entropy, the latter being defined as the entropy of an entangled purification of /spl rho/ after half of it has passed through the channel. We calculate entanglement-assisted capacities for two interesting quantum channels, the qubit amplitude damping channel and the bosonic channel with amplification/attenuation and Gaussian noise. We discuss how many independent parameters are required to completely characterize the asymptotic behavior of a general quantum channel, alone or in the presence of ancillary resources such as prior entanglement. In the classical analog of entanglement-assisted communication - communication over a discrete memoryless channel (DMC) between parties who share prior random information - we show that one parameter is sufficient, i.e., that in the presence of prior shared random information, all DMCs of equal capacity can simulate one another with unit asymptotic efficiency.
Charles H. Bennett, Peter W. Shor, John A. Smolin, Ashish V. Thapliyal
IEEE Trans. Inf. Theory4