Alexander Ku

dblp:215/4289 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Vision and language · 31% Generative modeling · 18% Trustworthy machine learning · 15%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 24 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-and-language navigation
1.942023
A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning · CVPR 2023
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding · EMNLP (1) 2020
Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019
Machine learning › Generative modeling
autoregressive model
0.922022
Vector-quantized Image Modeling with Improved VQGAN · ICLR 2022
Image Transformer · ICML 2018
Natural language and speech › Language models and text generation
instruction following
0.822023
A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning · CVPR 2023
Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019
Computer vision › Vision and language › compositionality
binding problem
0.812024
Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › model debugging
failure mode analysis
0.812024
Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem · NeurIPS 2024
Computer vision › Vision and language
image captioning
0.812024
DOCCI: Descriptions of Connected and Contrasting Images · ECCV (60) 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem · NeurIPS 2024
Computer vision › Vision and language
multimodal reasoning
0.812024
Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
Prompt Expansion for Adaptive Text-to-Image Generation · ACL (1) 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.712023
Gaussian Process Probes (GPP) for Uncertainty-Aware Probing · NeurIPS 2023
Natural language and speech › Language models and text generation › instruction tuning
instruction data generation
0.712023
A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning · CVPR 2023
Machine learning › Representation and self-supervised learning
probing
0.712023
Gaussian Process Probes (GPP) for Uncertainty-Aware Probing · NeurIPS 2023
Machine learning › Trustworthy machine learning
uncertainty estimation
0.712023
Gaussian Process Probes (GPP) for Uncertainty-Aware Probing · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling
0.612022
Vector-quantized Image Modeling with Improved VQGAN · ICLR 2022
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › discrete latent variable model
vector-quantized image modeling
0.612022
Vector-quantized Image Modeling with Improved VQGAN · ICLR 2022
Machine learning › Generative modeling › generative adversarial network › GAN architecture
VQGAN
0.612022
Vector-quantized Image Modeling with Improved VQGAN · ICLR 2022
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation learning
0.412019
Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019
Machine learning › Representation and self-supervised learning › transferable representation
transferable representation learning
0.412019
Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019
Machine learning › Generative modeling
image generation
0.312018
Image Transformer · ICML 2018
Machine learning › Deep learning architectures and training
transformer
0.312018
Image Transformer · ICML 2018
Image and video processing
super-resolution
0.312018
Image Transformer · ICML 2018
Human-AI interaction
prompt engineering
0.212024
Prompt Expansion for Adaptive Text-to-Image Generation · ACL (1) 2024
Computer vision › 3D vision
photorealistic simulation
0.112020
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding · EMNLP (1) 2020
Robotics › Robot navigation and mapping
embodied navigation
0.112019
Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019

Methods — techniques the papers use, named apart from their topics

transformer · 1.2feedforward processing analysis · 0.8cognitive science theory · 0.8imitation learning · 0.7image-to-image GAN · 0.7gaussian process probes · 0.7bayesian linear probing · 0.7vector quantization · 0.6GAN · 0.6multi-task learning · 0.4self-attention · 0.3autoregressive modeling · 0.3
YearPublicationVenuePosition
2024 Prompt Expansion for Adaptive Text-to-Image Generation
abstract
Text-to-image generation models are powerful but difficult to use. Users craft specific prompts to get better images, though the images can be repetitive. This paper proposes the Prompt Expansion framework that helps users generate high-quality, diverse images with less effort. The Prompt Expansion model takes a text query as input and outputs a set of expanded text prompts that are optimized such that when passed to a text-to-image model, they generate a wider variety of appealing images. We conduct a human evaluation study that shows that images generated through Prompt Expansion are more aesthetically pleasing and diverse than those generated by baseline methods. Overall, this paper presents a novel and effective approach to improving the text-to-image generation experience.
Siddhartha Datta, Alexander Ku, Deepak Ramachandran
ACL (1)2
2024 Can Generative Multimodal Models Count to Ten?
Sunayana Rane, Alexander Ku, Jason Baldridge, Ian Tenney, Thomas L. Griffiths 0001, Been Kim
CogSci2
2024 DOCCI: Descriptions of Connected and Contrasting Images
Yasumasa Onoe, Sunayana Rane, Zachary Berger, Yonatan Bitton, Jaemin Cho 0001, Roopal Garg, Alexander Ku, Zarana Parekh, Jordi Pont-Tuset, Garrett Tanzer, Su Wang 0001, Jason Baldridge
ECCV (60)7
2024 Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem
abstract
Recent work has documented striking heterogeneity in the performance of state-of-the-art vision language models (VLMs), including both multimodal language models and text-to-image models. These models are able to describe and generate a diverse array of complex, naturalistic images, yet they exhibit surprising failures on basic multi-object reasoning tasks -- such as counting, localization, and simple forms of visual analogy -- that humans perform with near perfect accuracy. To better understand this puzzling pattern of successes and failures, we turn to theoretical accounts of the binding problem in cognitive science and neuroscience, a fundamental problem that arises when a shared set of representational resources must be used to represent distinct entities (e.g., to represent multiple objects in an image), necessitating the use of serial processing to avoid interference. We find that many of the puzzling failures of state-of-the-art VLMs can be explained as arising due to the binding problem, and that these failure modes are strikingly similar to the limitations exhibited by rapid, feedforward processing in the human brain.
Declan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata, Kia Ghods, Amogh Joshi 0004, Alexander Ku, Steven Frankland, Thomas L. Griffiths 0001, Jonathan D. Cohen 0003, Taylor W. Webb
NeurIPS7
2023 A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning
abstract
Recent studies in Vision-and-Language Navigation (VLN) train RL agents to execute natural-language navigation instructions in photorealistic environments, as a step towards robots that can follow human instructions. However, given the scarcity of human instruction data and limited diversity in the training environments, these agents still struggle with complex language grounding and spatial language understanding. Pretraining on large text and image-text datasets from the web has been extensively explored but the improvements are limited. We investigate large-scale augmentation with synthetic instructions. We take 500+ indoor environments captured in densely-sampled 360 ° panoramas, construct navigation trajectories through these panoramas, and generate a visually-grounded instruction for each trajectory using Marky [63], a high-quality multilingual navigation instruction generator. We also synthesize image observations from novel viewpoints using an image-to-image GAN [27]. The resulting dataset of 4.2M instruction-trajectory pairs is two orders of magnitude larger than existing human-annotated datasets, and contains a wider variety of environments and viewpoints. To efficiently leverage data at this scale, we train a simple transformer agent with imitation learning. On the challenging RxR dataset, our approach outperforms all existing RL agents, improving the state-of-the-art NDTW from 71.1 to 79.1 in seen environments, and from 64.6 to 66.8 in unseen test environments. Our work points to a new path to improving instruction-following agents, emphasizing large-scale training on near-human quality synthetic instructions.
Aishwarya Kamath, Su Wang 0001, Jing Yu Koh, Alexander Ku, Austin Waters, Yinfei Yang, Jason Baldridge, Zarana Parekh
CVPR5
2023 Gaussian Process Probes (GPP) for Uncertainty-Aware Probing
abstract
Understanding which concepts models can and cannot represent has been fundamental to many tasks: from effective and responsible use of models to detecting out of distribution data. We introduce Gaussian process probes (GPP), a unified and simple framework for probing and measuring uncertainty about concepts represented by models. As a Bayesian extension of linear probing methods, GPP asks what kind of distribution over classifiers (of concepts) is induced by the model. This distribution can be used to measure both what the model represents and how confident the probe is about what the model represents. GPP can be applied to any pre-trained model with vector representations of inputs (e.g., activations). It does not require access to training data, gradients, or the architecture. We validate GPP on datasets containing both synthetic and real images. Our experiments show it can (1) probe a model's representations of concepts even with a very small number of examples, (2) accurately measure both epistemic uncertainty (how confident the probe is) and aleatory uncertainty (how fuzzy the concepts are to the model), and (3) detect out of distribution data using those uncertainty measures as well as classic methods do. By using Gaussian processes to expand what probing can offer, GPP provides a data-efficient, versatile and uncertainty-aware tool for understanding and evaluating the capabilities of machine learning models.
Alexander Ku, Jason Baldridge, Thomas L. Griffiths 0001, Been Kim
NeurIPS2
2022 Vector-quantized Image Modeling with Improved VQGAN
Jing Yu Koh, Han Zhang 0010, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge
ICLR7
2021 On the Evaluation of Vision-and-Language Navigation Instructions
abstract
Ming Zhao, Peter Anderson, Vihan Jain, Su Wang, Alexander Ku, Jason Baldridge, Eugene Ie. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Vihan Jain, Su Wang 0001, Alexander Ku, Jason Baldridge, Eugene Ie
EACL5
2020 Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
abstract
We introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset.RxR is multilingual (English, Hindi, and Telugu) and larger (more paths and instructions) than other VLN datasets.It emphasizes the role of language in VLN by addressing known biases in paths and eliciting more references to visible entities.Furthermore, each word in an instruction is time-aligned to the virtual poses of instruction creators and validators.We establish baseline scores for monolingual and multilingual settings and multitask learning when including Room-to-Room annotations (Anderson et al., 2018b).We also provide results for a model that learns from synchronized pose traces by focusing only on portions of the panorama attended to in human demonstrations.The size, scope and detail of RxR dramatically expands the frontier for research on embodied language agents in simulated, photo-realistic environments.
Alexander Ku, Roma Patel, Eugene Ie, Jason Baldridge
EMNLP (1)1
2019 Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation
abstract
Advances in learning and representations have reinvigorated work that connects language to other modalities.A particularly exciting direction is Vision-and-Language Navigation (VLN), in which agents interpret natural language instructions and visual scenes to move through environments and reach goals.Despite recent progress, current research leaves unclear how much of a role language understanding plays in this task, especially because dominant evaluation metrics have focused on goal completion rather than the sequence of actions corresponding to the instructions.Here, we highlight shortcomings of current metrics for the Room-to-Room dataset (Anderson et al., 2018b) and propose a new metric, Coverage weighted by Length Score (CLS).We also show that the existing paths in the dataset are not ideal for evaluating instruction following because they are direct-to-goal shortest paths.We join existing short paths to form more challenging extended paths to create a new data set, Room-for-Room (R4R).Using R4R and CLS, we show that agents that receive rewards for instruction fidelity outperform agents that focus on goal completion.
Vihan Jain, Gabriel Ilharco, Alexander Ku, Ashish Vaswani, Eugene Ie, Jason Baldridge
ACL (1)3
2019 Transferable Representation Learning in Vision-and-Language Navigation
abstract
Vision-and-Language Navigation (VLN) tasks such as Room-to-Room (R2R) require machine agents to interpret natural language instructions and learn to act in visually realistic environments to achieve navigation goals. The overall task requires competence in several perception problems: successful agents combine spatio-temporal, vision and language understanding to produce appropriate action sequences. Our approach adapts pre-trained vision and language representations to relevant in-domain tasks making them more effective for VLN. Specifically, the representations are adapted to solve both a cross-modal sequence alignment and sequence coherence task. In the sequence alignment task, the model determines whether an instruction corresponds to a sequence of visual frames. In the sequence coherence task, the model determines whether the perceptual sequences are predictive sequentially in the instruction-conditioned latent space. By transferring the domain-adapted representations, we improve competitive agents in R2R as measured by the success rate weighted by path length (SPL) metric.
Haoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku, Gabriel Ilharco, Jason Baldridge, Eugene Ie
ICCV4
2018 Image Transformer
abstract
Image generation has been successfully cast as an autoregressive sequence generation or transformation problem. Recent work has shown that self-attention is an effective way of modeling textual sequences. In this work, we generalize a recently proposed model architecture based on self-attention, the Transformer, to a sequence modeling formulation of image generation with a tractable likelihood. By restricting the self-attention mechanism to attend to local neighborhoods we significantly increase the size of images the model can process in practice, despite maintaining significantly larger receptive fields per layer than typical convolutional neural networks. While conceptually simple, our generative models significantly outperform the current state of the art in image generation on ImageNet, improving the best published negative log-likelihood on ImageNet from 3.83 to 3.77. We also present results on image super-resolution with a large magnification ratio, applying an encoder-decoder configuration of our architecture. In a human evaluation study, we find that images generated by our super-resolution model fool human observers three times more often than the previous state of the art.
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, Dustin Tran
ICML6