EDBT 2026 Demo / reviewers in the wild / expert
Roy Miles
dblp:256/1672
· DBLP profile ↗
10ranked-venue papers
7as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMsabstractLarge Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to contextual noise (e.g., background clutter). To address this issue, we introduce a visual comprehension stage, which we call ViCToR (Visual Comprehension via Token Reconstruction), a novel pretraining framework for LMMs. ViCToR employs a learnable visual token pool and utilizes the Hungarian matching algorithm to select semantically relevant tokens from this pool for visual token replacement. Furthermore, by integrating a visual token reconstruction loss with dense semantic supervision, ViCToR can learn tokens which retain high visual detail, thereby enhancing the large language model's (LLM's) understanding of visual information. After pretraining on 3 million publicly accessible images and captions, ViCToR achieves state-of-the-art results, improving over LLaVA-NeXT-8B by 10.4%, 3.2%, and 7.2% on the MMStar, SEEDI, and RealWorldQA benchmarks, respectively. Yin Xie, Kaicheng Yang 0002, Peirou Liang, Xiang An, Yongle Zhao, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng |
AAAI | 8 |
| 2025 | Region-based Cluster Discrimination for Visual Representation LearningabstractLearning visual representations is foundational for a broad spectrum of downstream tasks. Although recent vision-language contrastive models, such as CLIP and SigLIP, have achieved impressive zero-shot performance via large-scale vision-language alignment, their reliance on global representations constrains their effectiveness for dense prediction tasks, such as grounding, OCR, and segmentation. To address this gap, we introduce Region-Aware Cluster Discrimination (RICE), a novel method that enhances region-level visual and OCR capabilities. We first construct a billion-scale candidate region dataset and propose a Region Transformer layer to extract rich regional semantics. We further design a unified region cluster discrimination loss that jointly supports object and OCR learning within a single classification framework, enabling efficient and scalable distributed training on large-scale data. Extensive experiments show that RICE consistently outperforms previous methods on tasks, including segmentation, dense detection, and visual perception for Multimodal Large Language Models (MLLMs). The pre-trained models have been released at https://github.com/deepglint/MVT. Yin Xie, Kaicheng Yang 0002, Xiang An, Yongle Zhao, Weimo Deng, Zimin Ran, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng |
ICCV | 10 |
| 2024 | Understanding the Role of the Projector in Knowledge DistillationabstractIn this paper we revisit the efficacy of knowledge distillation as a function matching and metric learning problem. In doing so we verify three important design decisions, namely the normalisation, soft maximum function, and projection layers as key ingredients. We theoretically show that the projector implicitly encodes information on past examples, enabling relational gradients for the student. We then show that the normalisation of representations is tightly coupled with the training dynamics of this projector, which can have a large impact on the students performance. Finally, we show that a simple soft maximum function can be used to address any significant capacity gap problems. Experimental results on various benchmark datasets demonstrate that using these insights can lead to superior or comparable performance to state-of-the-art knowledge distillation techniques, despite being much more computationally efficient. In particular, we obtain these results across image classification (CIFAR100 and ImageNet), object detection (COCO2017), and on more difficult distillation objectives, such as training data efficient transformers, whereby we attain a 77.2% top-1 accuracy with DeiT-Ti on ImageNet. Code and models are publicly available. Roy Miles, Krystian Mikolajczyk |
AAAI | 1 |
| 2024 | Learning to Project for Cross-Task Knowledge Distillation
Dylan Auty, Roy Miles, Benedikt Kolbeinsson, Krystian Mikolajczyk |
BMVC | 2 |
| 2024 | $V_{k}D$: Improving Knowledge Distillation Using Orthogonal ProjectionsabstractKnowledge distillation is an effective method for training small and efficient deep learning models. However, the efficacy of a single method can degenerate when transferring to other tasks, modalities, or even other architectures. To address this limitation, we propose a novel constrained feature distillation method. This method is derived from a small set of core principles, which results in two emerging components: an orthogonal projection and a task-specific normalisation. Equipped with both of these components, our transformer models can outperform all previous methods on ImageNet and reach up to a 4.4% relative improvement over the previous state-of-the-art methods. To further demonstrate the generality of our method, we apply it to object detection and image generation, whereby we obtain consistent and substantial performance improvements over state-of-the-art. Code and models are publicly available11https://github.com/roymiles/vkd. Roy Miles, Ismail Elezi, Jiankang Deng |
CVPR | 1 |
| 2024 | VeLoRA: Memory Efficient Training using Rank-1 Sub-Token ProjectionsabstractLarge language models (LLMs) have recently emerged as powerful tools for tackling many language-processing tasks. Despite their success, training and fine-tuning these models is still far too computationally and memory intensive. In this paper, we identify and characterise the important components needed for effective model convergence using gradient descent. In doing so we find that the intermediate activations used to implement backpropagation can be excessively compressed without incurring any degradation in performance. This result leads us to a cheap and memory-efficient algorithm for both fine-tuning and pre-training LLMs. The proposed algorithm simply divides the tokens up into smaller sub-tokens before projecting them onto a fixed 1-dimensional subspace during the forward pass. These features are then coarsely reconstructed during the backward pass to implement the update rules. We confirm the effectiveness of our algorithm as being complimentary to many state-of-the-art PEFT methods on the VTAB-1k fine-tuning benchmark. Furthermore, we outperform QLoRA for fine-tuning LLaMA and show competitive performance against other memory-efficient pre-training methods on the large-scale C4 dataset. Roy Miles, Pradyumna Reddy, Ismail Elezi, Jiankang Deng |
NeurIPS | 1 |
| 2023 | MobileVOS: Real-Time Video Object Segmentation Contrastive Learning meets Knowledge DistillationabstractThis paper tackles the problem of semi-supervised video object segmentation on resource-constrained devices, such as mobile phones. We formulate this problem as a distillation task, whereby we demonstrate that small space-time-memory networks with finite memory can achieve competitive results with state of the art, but at a fraction of the computational cost (32 milliseconds per frame on a Samsung Galaxy S22). Specifically, we provide a theoretically grounded framework that unifies knowledge distillation with supervised contrastive representation learning. These models are able to jointly benefit from both pixel-wise contrastive learning and distillation from a pre-trained teacher. We validate this loss by achieving competitive J &F to state of the art on both the standard DAVIS and YouTube benchmarks, despite running up to × 5 faster, and with × 32 fewer parameters. Roy Miles, Mehmet Kerim Yucel, Bruno Manganelli, Albert Saà-Garriga |
CVPR | 1 |
| 2022 | Information Theoretic Representation Distillation
Roy Miles, Adrián López Rodríguez, Krystian Mikolajczyk |
BMVC | 1 |
| 2021 | Compressing Local Descriptor Models for Mobile ApplicationsabstractFeature-based image matching has been significantly improved through the use of deep learning and new large datasets. However, there has been little work addressing the computational cost, model size, and matching accuracy tradeoffs for the state of the art models. In this paper, we consider these practical aspects and improve the state-of-the-art HardNet model through the use of depthwise separable layers and an efficient tensor decomposition. We propose the Convolution-Depthwise-Pointwise (CDP) layer, which partitions the weights into a low and full rank decomposition to exploit the naturally emergent structure in the convolutional weights. We can achieve an 8× reduction in the number of parameters on the HardNet model, 13× reduction in the computational complexity, while sacrificing less than 1% on the overall accuracy across the HPatches benchmarks. To further demonstrate the generalisation of this approach, we apply it to other state-of-the-art descriptor models, where we are able to a significant performance improvement. Roy Miles, Krystian Mikolajczyk |
ICASSP | 1 |
| 2020 | Cascaded channel pruning using hierarchical self-distillation
Roy Miles, Krystian Mikolajczyk |
BMVC | 1 |