Josiah Wang

dblp:34/9514 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
2since 2021 · last 2022
0000-0003-0048-3893ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 54% Transfer learning and domain adaptation · 23% Image recognition and object detection · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
image captioning
0.622019
VIFIDEL: Evaluating the Visual Fidelity of Image Descriptions · ACL (1) 2019
Combining Geometric, Textual and Visual Features for Predicting Prepositions in Image Descriptions · EMNLP 2015
Machine learning › Transfer learning and domain adaptation
knowledge transfer
0.622018
Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Large Scale Semi-Supervised Object Detection Using Visual and Semantic Knowledge Transfer · CVPR 2016
Computer vision › Image recognition and object detection › object detection
semi-supervised object detection
0.622018
Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Large Scale Semi-Supervised Object Detection Using Visual and Semantic Knowledge Transfer · CVPR 2016
Computer vision › Vision and language
cross-modal retrieval
0.412019
Phrase Localization Without Paired Training Examples · ICCV 2019
Computer vision › Vision and language › image captioning
image caption evaluation
0.412019
VIFIDEL: Evaluating the Visual Fidelity of Image Descriptions · ACL (1) 2019
Computer vision › Vision and language › visual grounding
phrase grounding
0.412019
Phrase Localization Without Paired Training Examples · ICCV 2019
Information retrieval
evaluation
0.412019
VIFIDEL: Evaluating the Visual Fidelity of Image Descriptions · ACL (1) 2019
Machine learning › Transfer learning and domain adaptation
cross-domain transfer
0.312018
Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › Vision and language › cross-modal alignment
visual-semantic similarity
0.312018
Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › 3D vision › 3d scene understanding
spatial relation understanding
0.212015
Combining Geometric, Textual and Visual Features for Predicting Prepositions in Image Descriptions · EMNLP 2015
Computer vision › Vision and language
visual entity recognition
0.112015
Combining Geometric, Textual and Visual Features for Predicting Prepositions in Image Descriptions · EMNLP 2015

Methods — techniques the papers use, named apart from their topics

semantic similarity · 1.1CNN · 0.6object detection · 0.4knowledge transfer · 0.3visual similarity · 0.2semantic relatedness · 0.2visual features · 0.2textual features · 0.2geometric features · 0.2
YearPublicationVenuePosition
2022 MultiSubs: A Large-scale Multimodal and Multilingual Dataset
abstract
This paper introduces a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. The dataset consists of images selected to unambiguously illustrate concepts expressed in sentences from movie subtitles. The dataset is a valuable resource as (i) the images are aligned to text fragments rather than whole sentences; (ii) multiple images are possible for a text fragment and a sentence; (iii) the sentences are free-form and real-world like; (iv) the parallel texts are multilingual. We also set up a fill-in-the-blank game for humans to evaluate the quality of the automatic image selection process of our dataset. Finally, we propose a fill-in-the-blank task to demonstrate the utility of the dataset, and present some baseline prediction models. The dataset will benefit research on visual grounding of words especially in the context of free-form sentences, and can be obtained from https://doi.org/10.5281/zenodo.5034604 under a Creative Commons licence.
Josiah Wang, Josiel Figueiredo, Lucia Specia
LREC1
2021 Read, spot and translate
abstract
Abstract We propose multimodal machine translation (MMT) approaches that exploit the correspondences between words and image regions. In contrast to existing work, our referential grounding method considers objects as the visual unit for grounding, rather than whole images or abstract image regions, and performs visual grounding in the source language, rather than at the decoding stage via attention. We explore two referential grounding approaches: (i) implicit grounding, where the model jointly learns how to ground the source language in the visual representation and to translate; and (ii) explicit grounding, where grounding is performed independent of the translation model, and is subsequently used to guide machine translation. We performed experiments on the Multi30K dataset for three language pairs: English–German, English–French and English–Czech. Our referential grounding models outperform existing MMT models according to automatic and human evaluation metrics.
Lucia Specia, Josiah Wang, Sun Jae Lee, Alissa Ostapenko, Pranava Swaroop Madhyastha
Mach. Transl.2
2019 VIFIDEL: Evaluating the Visual Fidelity of Image Descriptions
abstract
We address the task of evaluating image description generation systems.We propose a novel image-aware metric for this task: VIFIDEL.It estimates the faithfulness of a generated caption with respect to the content of the actual image, based on the semantic similarity between labels of objects depicted in images and words in the description.The metric is also able to take into account the relative importance of objects mentioned in human reference descriptions during evaluation.Even if these human reference descriptions are not available, VIFIDEL can still reliably evaluate system descriptions.The metric achieves high correlation with human judgments on two well-known datasets and is competitive with metrics that depend on and rely exclusively on human references.
Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia
ACL (1)2
2019 Phrase Localization Without Paired Training Examples
abstract
Localizing phrases in images is an important part of image understanding and can be useful in many applications that require mappings between textual and visual information. Existing work attempts to learn these mappings from examples of phrase-image region correspondences (strong supervision) or from phrase-image pairs (weak supervision). We postulate that such paired annotations are unnecessary, and propose the first method for the phrase localization problem where neither training procedure nor paired, task-specific data is required. Our method is simple but effective: we use off-the-shelf approaches to detect objects, scenes and colours in images, and explore different approaches to measure semantic similarity between the categories of detected visual elements and words in phrases. Experiments on two well-known phrase localization datasets show that this approach surpasses all weakly supervised methods by a large margin and performs very competitively to strongly supervised methods, and can thus be considered a strong baseline to the task. The non-paired nature of our method makes it applicable to any domain and where no paired phrase localization annotation is available.
Josiah Wang, Lucia Specia
ICCV1
2018 End-to-end Image Captioning Exploits Distributional Similarity in Multimodal Space
Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia
BMVC2
2018 Object Counts! Bringing Explicit Detections Back into Image Captioning
abstract
Josiah Wang, Pranava Swaroop Madhyastha, Lucia Specia. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Josiah Wang, Pranava Swaroop Madhyastha, Lucia Specia
NAACL-HLT1
2018 The role of image representations in vision to language tasks
abstract
Abstract Tasks that require modeling of both language and visual information, such as image captioning, have become very popular in recent years. Most state-of-the-art approaches make use of image representations obtained from a deep neural network, which are used to generate language information in a variety of ways with end-to-end neural-network-based models. However, it is not clear how different image representations contribute to language generation tasks. In this paper, we probe the representational contribution of the image features in an end-to-end neural modeling framework and study the properties of different types of image representations. We focus on two popular vision to language problems: The task of image captioning and the task of multimodal machine translation. Our analysis provides interesting insights into the representational properties and suggests that end-to-end approaches implicitly learn a visual-semantic subspace and exploit the subspace to generate captions.
Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia
Nat. Lang. Eng.2
2018 Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection
abstract
Deep CNN-based object detection systems have achieved remarkable success on several large-scale object detection benchmarks. However, training such detectors requires a large number of labeled bounding boxes, which are more difficult to obtain than image-level annotations. Previous work addresses this issue by transforming image-level classifiers into object detectors. This is done by modeling the differences between the two on categories with both image-level and bounding box annotations, and transferring this information to convert classifiers to detectors for categories without bounding box annotations. We improve this previous work by incorporating knowledge about object similarities from visual and semantic domains during the transfer process. The intuition behind our proposed method is that visually and semantically similar categories should exhibit more common transferable properties than dissimilar categories, e.g. a better detector would result by transforming the differences between a dog classifier and a dog detector onto the cat class, than would by transforming from the violin class. Experimental results on the challenging ILSVRC2013 detection dataset demonstrate that each of our proposed object similarity based knowledge transfer methods outperforms the baseline methods. We found strong evidence that visual similarity and semantic relatedness are complementary for the task, and when combined notably improve detection, achieving state-of-the-art detection performance in a semi-supervised setting.
Yuxing Tang, Josiah Wang, Boyang Gao, Emmanuel Dellandréa, Robert J. Gaizauskas, Liming Chen 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Large Scale Semi-Supervised Object Detection Using Visual and Semantic Knowledge Transfer
abstract
Deep CNN-based object detection systems have achieved remarkable success on several large-scale object detection benchmarks. However, training such detectors requires a large number of labeled bounding boxes, which are more difficult to obtain than image-level annotations. Previous work addresses this issue by transforming image-level classifiers into object detectors. This is done by modeling the differences between the two on categories with both imagelevel and bounding box annotations, and transferring this information to convert classifiers to detectors for categories without bounding box annotations. We improve this previous work by incorporating knowledge about object similarities from visual and semantic domains during the transfer process. The intuition behind our proposed method is that visually and semantically similar categories should exhibit more common transferable properties than dissimilar categories, e.g. a better detector would result by transforming the differences between a dog classifier and a dog detector onto the cat class, than would by transforming from the violin class. Experimental results on the challenging ILSVRC2013 detection dataset demonstrate that each of our proposed object similarity based knowledge transfer methods outperforms the baseline methods. We found strong evidence that visual similarity and semantic relatedness are complementary for the task, and when combined notably improve detection, achieving state-of-the-art detection performance in a semi-supervised setting.
Yuxing Tang, Josiah Wang, Boyang Gao, Emmanuel Dellandréa, Robert J. Gaizauskas, Liming Chen 0002
CVPR2
2016 Harvesting Training Images for Fine-Grained Object Categories Using Visual Descriptions
Josiah Wang, Katja Markert, Mark Everingham
ECIR1
2016 Don't Mention the Shoe! A Learning to Rank Approach to Content Selection for Image Description Generation
abstract
We tackle the sub-task of content selection as part of the broader challenge of automatically generating image descriptions.More specifically, we explore how decisions can be made to select what object instances should be mentioned in an image description, given an image and labelled bounding boxes.We propose casting the content selection problem as a learning to rank problem, where object instances that are most likely to be mentioned by humans when describing an image are ranked higher than those that are less likely to be mentioned.Several features are explored: those derived from bounding box localisations, from concept labels, and from image regions.Object instances are then selected based on the ranked list, where we investigate several methods for choosing a stopping criterion as the 'cut-off' point for objects in the ranked list.Our best-performing method achieves state-of-the-art performance on the ImageCLEF2015 sentence generation challenge.
Josiah Wang, Robert J. Gaizauskas
INLG1
2016 Cross-validating Image Description Datasets and Evaluation Metrics
Josiah Wang, Robert J. Gaizauskas
LREC1
2015 Combining Geometric, Textual and Visual Features for Predicting Prepositions in Image Descriptions
abstract
We investigate the role that geometric, textual and visual features play in the task of predicting a preposition that links two visual entities depicted in an image.The task is an important part of the subsequent process of generating image descriptions.We explore the prediction of prepositions for a pair of entities, both in the case when the labels of such entities are known and unknown.In all situations we found clear evidence that all three features contribute to the prediction task.
Arnau Ramisa, Josiah Wang, Ying Lu 0007, Emmanuel Dellandréa, Francesc Moreno-Noguer, Robert J. Gaizauskas
EMNLP2
2009 Learning Models for Object Recognition from Natural Language Descriptions
abstract
We investigate the task of learning models for visual object recognition from natural language descriptions alone. The approach contributes to the recognition of fine-grain object categories, such as animal and plant species, where it may be difficult to collect many images for training, but where textual descriptions of visual attributes are readily available. As an example we tackle recognition of butterfly species, learning models from descriptions in an online nature guide. We propose natural language processing methods for extracting salient visual attributes from these descriptions to use as ‘templates ’ for the object categories, and apply vision methods to extract corresponding attributes from test images. A generative model is used to connect textual terms in the learnt templates to visual attributes. We report experiments comparing the performance of humans and the proposed method on a dataset of ten butterfly categories. 1
Josiah Wang, Katja Markert, Mark Everingham
BMVC1