VLDB 2026 Research / reviewers in the wild / expert
Yuge Shi
dblp:227/4684
· DBLP profile ↗
11ranked-venue papers
6as first author
9since 2021 · last 2024
0000-0003-1905-9320ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Position: Open-Endedness is Essential for Artificial Superhuman IntelligenceabstractIn recent years there has been a tremendous surge in the general capabilities of AI systems, mainly fuelled by training foundation models on internet-scale data. Nevertheless, the creation of open-ended, ever self-improving AI remains elusive. **In this position paper, we argue that the ingredients are now in place to achieve *open-endedness* in AI systems with respect to a human observer. Furthermore, we claim that such open-endedness is an essential property of any artificial superhuman intelligence (ASI).** We begin by providing a concrete formal definition of open-endedness through the lens of novelty and learnability. We then illustrate a path towards ASI via open-ended systems built on top of foundation models, capable of making novel, human-relevant discoveries. We conclude by examining the safety implications of generally-capable open-ended AI. We expect that open-ended foundation models will prove to be an increasingly fertile and safety-critical area of research in the near future. Edward Hughes 0001, Michael Dennis 0001, Jack Parker-Holder, Feryal M. P. Behbahani, Aditi Mavalankar, Yuge Shi, Tom Schaul, Tim Rocktäschel |
ICML | 6 |
| 2024 | Memory Consolidation Enables Long-Context Video UnderstandingabstractMost transformer-based video encoders are limited to short temporal contexts due to their quadratic complexity. While various attempts have been made to extend this context, this has often come at the cost of both conceptual and computational complexity. We propose to instead re-purpose existing pre-trained video transformers by simply fine-tuning them to attend to memories derived non-parametrically from past activations. By leveraging redundancy reduction, our memory-consolidated vision transformer (MC-ViT) effortlessly extends its context far into the past and exhibits excellent scaling behavior when learning from longer videos. In doing so, MC-ViT sets a new state-of-the-art in long-context video understanding on EgoSchema, Perception Test, and Diving48, outperforming methods that benefit from orders of magnitude more parameters. Ivana Balazevic, Yuge Shi, Pinelopi Papalampidi, Rahma Chaabouni 0001, Skanda Koppula, Olivier J. Hénaff |
ICML | 2 |
| 2024 | Genie: Generative Interactive EnvironmentsabstractWe introduce Genie, the first *generative interactive environment* trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual worlds described through text, synthetic images, photographs, and even sketches. At 11B parameters, Genie can be considered a *foundation world model*. It is comprised of a spatiotemporal video tokenizer, an autoregressive dynamics model, and a simple and scalable latent action model. Genie enables users to act in the generated environments on a frame-by-frame basis *despite training without any ground-truth action labels* or other domain specific requirements typically found in the world model literature. Further the resulting learned latent action space facilitates training agents to imitate behaviors from unseen videos, opening the path for training generalist agents of the future. Jake Bruce, Michael Dennis 0001, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes 0001, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal M. P. Behbahani, Stephanie C. Y. Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott E. Reed, Jingwei Zhang 0001, Konrad Zolna, Jeff Clune, Nando de Freitas, Satinder Singh 0001, Tim Rocktäschel |
ICML | 5 |
| 2023 | How robust is unsupervised representation learning to distribution shift?
Yuge Shi, Imant Daunhawer, Julia E. Vogt, Philip Torr 0001, Amartya Sanyal |
ICLR | 1 |
| 2023 | Tuning Computer Vision Models With Task RewardsabstractMisalignment between model predictions and intended usage can be detrimental for the deployment of computer vision models. The issue is exacerbated when the task involves complex structured outputs, as it becomes harder to design procedures which address this misalignment. In natural language processing, this is often addressed using reinforcement learning techniques that align models with a task reward. We adopt this approach and show its surprising effectiveness to improve generic models pretrained to imitate example outputs across multiple computer vision tasks, such as object detection, panoptic segmentation, colorization and image captioning. We believe this approach has the potential to be widely useful for better aligning models with a diverse range of computer vision tasks. André Susano Pinto, Alexander Kolesnikov 0003, Yuge Shi, Lucas Beyer, Xiaohua Zhai |
ICML | 3 |
| 2022 | Learning Multimodal VAEs through Mutual Supervision
Tom Joy, Yuge Shi, Philip Torr 0001, Tom Rainforth, Sebastian M. Schmon, N. Siddharth 0001 |
ICLR | 2 |
| 2022 | Gradient Matching for Domain Generalization
Yuge Shi, Jeffrey Seely, Philip Torr 0001, N. Siddharth 0001, Awni Y. Hannun, Nicolas Usunier, Gabriel Synnaeve |
ICLR | 1 |
| 2022 | Adversarial Masking for Self-Supervised LearningabstractWe propose ADIOS, a masked image model (MIM) framework for self-supervised learning, which simultaneously learns a masking function and an image encoder using an adversarial objective. The image encoder is trained to minimise the distance between representations of the original and that of a masked image. The masking function, conversely, aims at maximising this distance. ADIOS consistently improves on state-of-the-art self-supervised learning (SSL) methods on a variety of tasks and datasets—including classification on ImageNet100 and STL10, transfer learning on CIFAR10/100, Flowers102 and iNaturalist, as well as robustness evaluated on the backgrounds challenge (Xiao et al., 2021)—while generating semantically meaningful masks. Unlike modern MIM models such as MAE, BEiT and iBOT, ADIOS does not rely on the image-patch tokenisation construction of Vision Transformers, and can be implemented with convolutional backbones. We further demonstrate that the masks learned by ADIOS are more effective in improving representation learning of SSL methods than masking schemes used in popular MIM models. Yuge Shi, N. Siddharth 0001, Philip Torr 0001, Adam R. Kosiorek |
ICML | 1 |
| 2021 | Relating by Contrasting: A Data-efficient Framework for Multimodal Generative Models
Yuge Shi, Brooks Paige, Philip Torr 0001, N. Siddharth 0001 |
ICLR | 1 |
| 2019 | Variational Mixture-of-Experts Autoencoders for Multi-Modal Deep Generative ModelsabstractLearning generative models that span multiple data modalities, such as vision and language, is often motivated by the desire to learn more useful, generalisable representations that faithfully capture common underlying factors between the modalities. In this work, we characterise successful learning of such models as the fulfilment of four criteria: i) implicit latent decomposition into shared and private subspaces, ii) coherent joint generation over all modalities, iii) coherent cross-generation across individual modalities, and iv) improved model learning for individual modalities through multi-modal integration. Here, we propose a mixture-of-experts multi-modal variational autoencoder (MMVAE) for learning of generative models on different sets of modalities, including a challenging image <-> language dataset, and demonstrate its ability to satisfy all four criteria, both qualitatively and quantitatively. Yuge Shi, N. Siddharth 0001, Brooks Paige, Philip Torr 0001 |
NeurIPS | 1 |
| 2018 | Action Anticipation with RBF Kernelized Feature Mapping RNN
Yuge Shi, Basura Fernando, Richard I. Hartley |
ECCV (10) | 1 |