EDBT 2026 Demo / reviewers in the wild / expert
Ashish V. Thapliyal
dblp:42/4147
· DBLP profile ↗
9ranked-venue papers
2as first author
6since 2021 · last 2023
0000-0002-7219-0515ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Theory of computation · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Vision and language · 44% Language models and text generation · 32% Representation and self-supervised learning · 12% | |
| Theoretical computer science
1 paper |
Information theory · 50% Quantum computing and quantum information · 33% Coding theory · 17% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
image captioning |
1.0 | 2 | 2022 | Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset · EMNLP 2022 Cross-modal Language Generation using Pivot Stabilization for Web-scale Language Coverage · ACL 2020 |
Natural language and speech › Language models and text generation
multilingual language models |
0.7 | 1 | 2023 | PaLI: A Jointly-Scaled Multilingual Language-Image Model · ICLR 2023 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state representation |
0.7 | 1 | 2023 | Emergence of Abstract State Representations in Embodied Sequence Modeling · EMNLP 2023 |
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning |
0.7 | 1 | 2023 | Emergence of Abstract State Representations in Embodied Sequence Modeling · EMNLP 2023 |
Computer vision › Vision and language › image captioning
cross-lingual image captioning |
0.6 | 1 | 2022 | Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset · EMNLP 2022 |
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation |
0.6 | 1 | 2022 | Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset · EMNLP 2022 |
Natural language and speech › Language models and text generation › text generation
multilingual text generation |
0.4 | 1 | 2020 | Cross-modal Language Generation using Pivot Stabilization for Web-scale Language Coverage · ACL 2020 |
Computer vision › Vision and language
vision-language pretraining |
0.2 | 1 | 2023 | PaLI: A Jointly-Scaled Multilingual Language-Image Model · ICLR 2023 |
Information theory
channel capacity |
0.0 | 1 | 2002 | Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002 |
Coding theory › channel coding
channel simulation |
0.0 | 1 | 2002 | Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002 |
Information theory › communication channels › channel models
discrete memoryless channel |
0.0 | 1 | 2002 | Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002 |
Quantum computing and quantum information › quantum channel capacity
entanglement-assisted capacity |
0.0 | 1 | 2002 | Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002 |
Quantum computing and quantum information
quantum channel capacity |
0.0 | 1 | 2002 | Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002 |
Information theory › channel capacity
reverse shannon theorem |
0.0 | 1 | 2002 | Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem · IEEE Trans. Inf. Theory 2002 |
Methods — techniques the papers use, named apart from their topics
transformer sequence modeling · 0.7transfer learning · 0.7probing · 0.7joint scaling · 0.7human evaluation · 0.6automatic caption metrics · 0.6training with silver data · 0.4pivot stabilization · 0.4machine translation · 0.4entropy-based capacity formula · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Emergence of Abstract State Representations in Embodied Sequence ModelingabstractDecision making via sequence modeling aims to mimic the success of language models, where actions taken by an embodied agent are modeled as tokens to predict.Despite their promising performance, it remains unclear if embodied sequence modeling leads to the emergence of internal representations that represent the environmental state information.A model that lacks abstract state representations would be liable to make decisions based on surface statistics which fail to generalize.We take the BabyAI environment, a grid world in which language-conditioned navigation tasks are performed, and build a sequence modeling Transformer, which takes a language instruction, a sequence of actions, and environmental observations as its inputs.In order to investigate the emergence of abstract state representations, we design a "blindfolded" navigation task, where only the initial environmental layout, the language instruction, and the action sequence to complete the task are available for training.Our probing results show that intermediate environmental layouts can be reasonably reconstructed from the internal activations of a trained model, and that language instructions play a role in the reconstruction accuracy.Our results suggest that many key features of state representations can emerge via embodied sequence modeling, supporting an optimistic outlook for applications of sequence modeling objectives to more complex embodied decision-making domains.1 Tian Yun 0001, Zilai Zeng, Kunal Handa, Ashish V. Thapliyal, Bo Pang 0001, Ellie Pavlick, Chen Sun 0002 |
EMNLP | 4 |
| 2023 | PaLI: A Jointly-Scaled Multilingual Language-Image Model
Xi Chen 0071, Xiao Wang 0038, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov 0003, Joan Puigcerver, Nan Ding 0002, Keran Rong, Hassan Akbari, Linting Xue, Ashish V. Thapliyal, Weicheng Kuo |
ICLR | 18 |
| 2022 | Denoising Large-Scale Image Captioning from Alt-text Data Using Content Selection ModelsabstractTraining large-scale image captioning (IC) models demands access to a rich and diverse set of training examples that are expensive to curate both in terms of time and man-power. Instead, alt-text based captions gathered from the web is a far cheaper alternative to scale with the downside of being noisy. Recent modeling approaches to IC often fall short in terms of performance in leveraging these noisy datasets in favor of clean annotations. We address this problem with a simple yet effective technique of breaking down the task into two smaller, more controllable tasks – skeleton prediction and skeleton-based caption generation. Specifically, we show that sub-selecting content words as skeletons helps in generating improved and denoised captions when leveraging rich yet noisy alt-text–based uncurated datasets. We also show that the predicted English skeletons can further cross-lingually be leveraged to generate non-English captions, and present experimental results covering caption generation in French, Italian, German, Spanish and Hindi. We also show that skeleton-based prediction allows for better control of certain caption properties, such as length, content, and gender expression, providing a handle to perform human-in-the-loop interpretable semi-automatic corrections. Khyathi Raghavi Chandu, Piyush Sharma, Soravit Changpinyo, Ashish V. Thapliyal, Radu Soricut |
COLING | 4 |
| 2022 | End-to-end Dense Video Captioning as Sequence GenerationabstractDense video captioning aims to identify the events of interest in an input video, and generate descriptive captions for each event. Previous approaches usually follow a two-stage generative process, which first proposes a segment for each event, then renders a caption for each identified segment. Recent advances in large-scale sequence generation pretraining have seen great success in unifying task formulation for a great variety of tasks, but so far, more complex tasks such as dense video captioning are not able to fully utilize this powerful paradigm. In this work, we show how to model the two subtasks of dense video captioning jointly as one sequence generation task, and simultaneously predict the events and the corresponding descriptions. Experiments on YouCook2 and ViTT show encouraging results and indicate the feasibility of training complex tasks such as end-to-end dense video captioning integrated into large-scale pretrained models. Wanrong Zhu, Bo Pang 0001, Ashish V. Thapliyal, William Yang Wang, Radu Soricut |
COLING | 3 |
| 2022 | Crossmodal-3600: A Massively Multilingual Multimodal Evaluation DatasetabstractResearch in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.In this paper we present the Crossmodal-3600 dataset (XM3600 in short), a geographically-diverse set of 3600 images annotated with humangenerated reference captions in 36 languages.The images were selected from across the world, covering regions where the 36 languages are spoken, and annotated with captions that achieve consistency in terms of style across all languages, while avoiding annotation artifacts due to direct translation.We apply this benchmark to model selection for massively multilingual image captioning models, and show strong correlation results with human evaluations when using XM3600 as golden references for automatic metrics. Ashish V. Thapliyal, Jordi Pont-Tuset, Xi Chen 0071, Radu Soricut |
EMNLP | 1 |
| 2021 | Quality Estimation for Image Captions Based on Large-scale Human EvaluationsabstractTomer Levinboim, Ashish V. Thapliyal, Piyush Sharma, Radu Soricut. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tomer Levinboim, Ashish V. Thapliyal, Piyush Sharma, Radu Soricut |
NAACL-HLT | 2 |
| 2020 | Cross-modal Language Generation using Pivot Stabilization for Web-scale Language CoverageabstractCross-modal language generation tasks such as image captioning are directly hurt in their ability to support non-English languages by the trend of data-hungry models combined with the lack of non-English annotations.We investigate potential solutions for combining existing language-generation annotations in English with translation capabilities in order to create solutions at web-scale in both domain and language coverage.We describe an approach called Pivot-Language Generation Stabilization (PLuGS), which leverages directly at training time both existing English annotations (gold data) as well as their machinetranslated versions (silver data); at run-time, it generates first an English caption and then a corresponding target-language caption.We show that PLuGS models outperform other candidate solutions in evaluations performed over 5 different target languages, under a largedomain testset using images from the Open Images dataset.Furthermore, we find an interesting effect where the English captions generated by the PLuGS models are better than the captions generated by the original, monolingual English model. Ashish V. Thapliyal, Radu Soricut |
ACL | 1 |
| 2003 | Rank two bipartite bound entangled states do not exist
Pawel Horodecki, John A. Smolin, Barbara M. Terhal, Ashish V. Thapliyal |
Theor. Comput. Sci. | 4 |
| 2002 | Entanglement-assisted capacity of a quantum channel and the reverse Shannon theoremabstractThe entanglement-assisted classical capacity of a noisy quantum channel (C/sub E/) is the amount of information per channel use that can be sent over the channel in the limit of many uses of the channel, assuming that the sender and receiver have access to the resource of shared quantum entanglement, which may be used up by the communication protocol. We show that the capacity C/sub E/ is given by an expression parallel to that for the capacity of a purely classical channel: i.e., the maximum, over channel inputs /spl rho/, of the entropy of the channel input plus the entropy of the channel output minus their joint entropy, the latter being defined as the entropy of an entangled purification of /spl rho/ after half of it has passed through the channel. We calculate entanglement-assisted capacities for two interesting quantum channels, the qubit amplitude damping channel and the bosonic channel with amplification/attenuation and Gaussian noise. We discuss how many independent parameters are required to completely characterize the asymptotic behavior of a general quantum channel, alone or in the presence of ancillary resources such as prior entanglement. In the classical analog of entanglement-assisted communication - communication over a discrete memoryless channel (DMC) between parties who share prior random information - we show that one parameter is sufficient, i.e., that in the presence of prior shared random information, all DMCs of equal capacity can simulate one another with unit asymptotic efficiency. Charles H. Bennett, Peter W. Shor, John A. Smolin, Ashish V. Thapliyal |
IEEE Trans. Inf. Theory | 4 |