EDBT 2026 Demo / reviewers in the wild / expert
Alexander Ku
dblp:215/4289
· DBLP profile ↗
12ranked-venue papers
1as first author
8since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Vision and language · 31% Generative modeling · 18% Trustworthy machine learning · 15% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 24 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
vision-and-language navigation |
1.9 | 4 | 2023 | A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning · CVPR 2023 Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding · EMNLP (1) 2020 Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019 |
Machine learning › Generative modeling
autoregressive model |
0.9 | 2 | 2022 | Vector-quantized Image Modeling with Improved VQGAN · ICLR 2022 Image Transformer · ICML 2018 |
Natural language and speech › Language models and text generation
instruction following |
0.8 | 2 | 2023 | A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning · CVPR 2023 Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019 |
Computer vision › Vision and language › compositionality
binding problem |
0.8 | 1 | 2024 | Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › interpretability › model debugging
failure mode analysis |
0.8 | 1 | 2024 | Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem · NeurIPS 2024 |
Computer vision › Vision and language
image captioning |
0.8 | 1 | 2024 | DOCCI: Descriptions of Connected and Contrasting Images · ECCV (60) 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem · NeurIPS 2024 |
Computer vision › Vision and language
multimodal reasoning |
0.8 | 1 | 2024 | Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | Prompt Expansion for Adaptive Text-to-Image Generation · ACL (1) 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.7 | 1 | 2023 | Gaussian Process Probes (GPP) for Uncertainty-Aware Probing · NeurIPS 2023 |
Natural language and speech › Language models and text generation › instruction tuning
instruction data generation |
0.7 | 1 | 2023 | A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning · CVPR 2023 |
Machine learning › Representation and self-supervised learning
probing |
0.7 | 1 | 2023 | Gaussian Process Probes (GPP) for Uncertainty-Aware Probing · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.7 | 1 | 2023 | Gaussian Process Probes (GPP) for Uncertainty-Aware Probing · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling |
0.6 | 1 | 2022 | Vector-quantized Image Modeling with Improved VQGAN · ICLR 2022 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › discrete latent variable model
vector-quantized image modeling |
0.6 | 1 | 2022 | Vector-quantized Image Modeling with Improved VQGAN · ICLR 2022 |
Machine learning › Generative modeling › generative adversarial network › GAN architecture
VQGAN |
0.6 | 1 | 2022 | Vector-quantized Image Modeling with Improved VQGAN · ICLR 2022 |
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation learning |
0.4 | 1 | 2019 | Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019 |
Machine learning › Representation and self-supervised learning › transferable representation
transferable representation learning |
0.4 | 1 | 2019 | Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019 |
Machine learning › Generative modeling
image generation |
0.3 | 1 | 2018 | Image Transformer · ICML 2018 |
Machine learning › Deep learning architectures and training
transformer |
0.3 | 1 | 2018 | Image Transformer · ICML 2018 |
Image and video processing
super-resolution |
0.3 | 1 | 2018 | Image Transformer · ICML 2018 |
Human-AI interaction
prompt engineering |
0.2 | 1 | 2024 | Prompt Expansion for Adaptive Text-to-Image Generation · ACL (1) 2024 |
Computer vision › 3D vision
photorealistic simulation |
0.1 | 1 | 2020 | Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding · EMNLP (1) 2020 |
Robotics › Robot navigation and mapping
embodied navigation |
0.1 | 1 | 2019 | Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.2feedforward processing analysis · 0.8cognitive science theory · 0.8imitation learning · 0.7image-to-image GAN · 0.7gaussian process probes · 0.7bayesian linear probing · 0.7vector quantization · 0.6GAN · 0.6multi-task learning · 0.4self-attention · 0.3autoregressive modeling · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Prompt Expansion for Adaptive Text-to-Image GenerationabstractText-to-image generation models are powerful but difficult to use. Users craft specific prompts to get better images, though the images can be repetitive. This paper proposes the Prompt Expansion framework that helps users generate high-quality, diverse images with less effort. The Prompt Expansion model takes a text query as input and outputs a set of expanded text prompts that are optimized such that when passed to a text-to-image model, they generate a wider variety of appealing images. We conduct a human evaluation study that shows that images generated through Prompt Expansion are more aesthetically pleasing and diverse than those generated by baseline methods. Overall, this paper presents a novel and effective approach to improving the text-to-image generation experience. Siddhartha Datta, Alexander Ku, Deepak Ramachandran |
ACL (1) | 2 |
| 2024 | Can Generative Multimodal Models Count to Ten?
Sunayana Rane, Alexander Ku, Jason Baldridge, Ian Tenney, Thomas L. Griffiths 0001, Been Kim |
CogSci | 2 |
| 2024 | DOCCI: Descriptions of Connected and Contrasting Images
Yasumasa Onoe, Sunayana Rane, Zachary Berger, Yonatan Bitton, Jaemin Cho 0001, Roopal Garg, Alexander Ku, Zarana Parekh, Jordi Pont-Tuset, Garrett Tanzer, Su Wang 0001, Jason Baldridge |
ECCV (60) | 7 |
| 2024 | Understanding the Limits of Vision Language Models Through the Lens of the Binding ProblemabstractRecent work has documented striking heterogeneity in the performance of state-of-the-art vision language models (VLMs), including both multimodal language models and text-to-image models. These models are able to describe and generate a diverse array of complex, naturalistic images, yet they exhibit surprising failures on basic multi-object reasoning tasks -- such as counting, localization, and simple forms of visual analogy -- that humans perform with near perfect accuracy. To better understand this puzzling pattern of successes and failures, we turn to theoretical accounts of the binding problem in cognitive science and neuroscience, a fundamental problem that arises when a shared set of representational resources must be used to represent distinct entities (e.g., to represent multiple objects in an image), necessitating the use of serial processing to avoid interference. We find that many of the puzzling failures of state-of-the-art VLMs can be explained as arising due to the binding problem, and that these failure modes are strikingly similar to the limitations exhibited by rapid, feedforward processing in the human brain. Declan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata, Kia Ghods, Amogh Joshi 0004, Alexander Ku, Steven Frankland, Thomas L. Griffiths 0001, Jonathan D. Cohen 0003, Taylor W. Webb |
NeurIPS | 7 |
| 2023 | A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation LearningabstractRecent studies in Vision-and-Language Navigation (VLN) train RL agents to execute natural-language navigation instructions in photorealistic environments, as a step towards robots that can follow human instructions. However, given the scarcity of human instruction data and limited diversity in the training environments, these agents still struggle with complex language grounding and spatial language understanding. Pretraining on large text and image-text datasets from the web has been extensively explored but the improvements are limited. We investigate large-scale augmentation with synthetic instructions. We take 500+ indoor environments captured in densely-sampled 360 ° panoramas, construct navigation trajectories through these panoramas, and generate a visually-grounded instruction for each trajectory using Marky [63], a high-quality multilingual navigation instruction generator. We also synthesize image observations from novel viewpoints using an image-to-image GAN [27]. The resulting dataset of 4.2M instruction-trajectory pairs is two orders of magnitude larger than existing human-annotated datasets, and contains a wider variety of environments and viewpoints. To efficiently leverage data at this scale, we train a simple transformer agent with imitation learning. On the challenging RxR dataset, our approach outperforms all existing RL agents, improving the state-of-the-art NDTW from 71.1 to 79.1 in seen environments, and from 64.6 to 66.8 in unseen test environments. Our work points to a new path to improving instruction-following agents, emphasizing large-scale training on near-human quality synthetic instructions. Aishwarya Kamath, Su Wang 0001, Jing Yu Koh, Alexander Ku, Austin Waters, Yinfei Yang, Jason Baldridge, Zarana Parekh |
CVPR | 5 |
| 2023 | Gaussian Process Probes (GPP) for Uncertainty-Aware ProbingabstractUnderstanding which concepts models can and cannot represent has been fundamental to many tasks: from effective and responsible use of models to detecting out of distribution data. We introduce Gaussian process probes (GPP), a unified and simple framework for probing and measuring uncertainty about concepts represented by models. As a Bayesian extension of linear probing methods, GPP asks what kind of distribution over classifiers (of concepts) is induced by the model. This distribution can be used to measure both what the model represents and how confident the probe is about what the model represents. GPP can be applied to any pre-trained model with vector representations of inputs (e.g., activations). It does not require access to training data, gradients, or the architecture. We validate GPP on datasets containing both synthetic and real images. Our experiments show it can (1) probe a model's representations of concepts even with a very small number of examples, (2) accurately measure both epistemic uncertainty (how confident the probe is) and aleatory uncertainty (how fuzzy the concepts are to the model), and (3) detect out of distribution data using those uncertainty measures as well as classic methods do. By using Gaussian processes to expand what probing can offer, GPP provides a data-efficient, versatile and uncertainty-aware tool for understanding and evaluating the capabilities of machine learning models. Alexander Ku, Jason Baldridge, Thomas L. Griffiths 0001, Been Kim |
NeurIPS | 2 |
| 2022 | Vector-quantized Image Modeling with Improved VQGAN
Jing Yu Koh, Han Zhang 0010, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge |
ICLR | 7 |
| 2021 | On the Evaluation of Vision-and-Language Navigation InstructionsabstractMing Zhao, Peter Anderson, Vihan Jain, Su Wang, Alexander Ku, Jason Baldridge, Eugene Ie. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Vihan Jain, Su Wang 0001, Alexander Ku, Jason Baldridge, Eugene Ie |
EACL | 5 |
| 2020 | Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingabstractWe introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset.RxR is multilingual (English, Hindi, and Telugu) and larger (more paths and instructions) than other VLN datasets.It emphasizes the role of language in VLN by addressing known biases in paths and eliciting more references to visible entities.Furthermore, each word in an instruction is time-aligned to the virtual poses of instruction creators and validators.We establish baseline scores for monolingual and multilingual settings and multitask learning when including Room-to-Room annotations (Anderson et al., 2018b).We also provide results for a model that learns from synchronized pose traces by focusing only on portions of the panorama attended to in human demonstrations.The size, scope and detail of RxR dramatically expands the frontier for research on embodied language agents in simulated, photo-realistic environments. Alexander Ku, Roma Patel, Eugene Ie, Jason Baldridge |
EMNLP (1) | 1 |
| 2019 | Stay on the Path: Instruction Fidelity in Vision-and-Language NavigationabstractAdvances in learning and representations have reinvigorated work that connects language to other modalities.A particularly exciting direction is Vision-and-Language Navigation (VLN), in which agents interpret natural language instructions and visual scenes to move through environments and reach goals.Despite recent progress, current research leaves unclear how much of a role language understanding plays in this task, especially because dominant evaluation metrics have focused on goal completion rather than the sequence of actions corresponding to the instructions.Here, we highlight shortcomings of current metrics for the Room-to-Room dataset (Anderson et al., 2018b) and propose a new metric, Coverage weighted by Length Score (CLS).We also show that the existing paths in the dataset are not ideal for evaluating instruction following because they are direct-to-goal shortest paths.We join existing short paths to form more challenging extended paths to create a new data set, Room-for-Room (R4R).Using R4R and CLS, we show that agents that receive rewards for instruction fidelity outperform agents that focus on goal completion. Vihan Jain, Gabriel Ilharco, Alexander Ku, Ashish Vaswani, Eugene Ie, Jason Baldridge |
ACL (1) | 3 |
| 2019 | Transferable Representation Learning in Vision-and-Language NavigationabstractVision-and-Language Navigation (VLN) tasks such as Room-to-Room (R2R) require machine agents to interpret natural language instructions and learn to act in visually realistic environments to achieve navigation goals. The overall task requires competence in several perception problems: successful agents combine spatio-temporal, vision and language understanding to produce appropriate action sequences. Our approach adapts pre-trained vision and language representations to relevant in-domain tasks making them more effective for VLN. Specifically, the representations are adapted to solve both a cross-modal sequence alignment and sequence coherence task. In the sequence alignment task, the model determines whether an instruction corresponds to a sequence of visual frames. In the sequence coherence task, the model determines whether the perceptual sequences are predictive sequentially in the instruction-conditioned latent space. By transferring the domain-adapted representations, we improve competitive agents in R2R as measured by the success rate weighted by path length (SPL) metric. Haoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku, Gabriel Ilharco, Jason Baldridge, Eugene Ie |
ICCV | 4 |
| 2018 | Image TransformerabstractImage generation has been successfully cast as an autoregressive sequence generation or transformation problem. Recent work has shown that self-attention is an effective way of modeling textual sequences. In this work, we generalize a recently proposed model architecture based on self-attention, the Transformer, to a sequence modeling formulation of image generation with a tractable likelihood. By restricting the self-attention mechanism to attend to local neighborhoods we significantly increase the size of images the model can process in practice, despite maintaining significantly larger receptive fields per layer than typical convolutional neural networks. While conceptually simple, our generative models significantly outperform the current state of the art in image generation on ImageNet, improving the best published negative log-likelihood on ImageNet from 3.83 to 3.77. We also present results on image super-resolution with a large magnification ratio, applying an encoder-decoder configuration of our architecture. In a human evaluation study, we find that images generated by our super-resolution model fool human observers three times more often than the previous state of the art. Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, Dustin Tran |
ICML | 6 |