EDBT 2026 Demo / reviewers in the wild / expert
Nicholas Apostoloff
dblp:92/3793
· DBLP profile ↗
16ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Trustworthy machine learning · 52% Language models and text generation · 31% Generative modeling · 14% | |
| Computer graphics and multimedia
2 papers |
Image and video processing · 100% |
Topics — the 24 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering |
0.9 | 1 | 2025 | Controlling Language and Diffusion Models by Transporting Activations · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Controlling Language and Diffusion Models by Transporting Activations · ICLR 2025 |
Machine learning › Trustworthy machine learning
fairness |
0.9 | 1 | 2025 | Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs · ICML 2025 |
Machine learning › Trustworthy machine learning › fairness
fairness evaluation |
0.9 | 1 | 2025 | Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs · ICML 2025 |
Natural language and speech › Language models and text generation › model steering
language model steering |
0.9 | 1 | 2025 | Controlling Language and Diffusion Models by Transporting Activations · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs · ICML 2025 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.9 | 1 | 2025 | Controlling Language and Diffusion Models by Transporting Activations · ICLR 2025 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.9 | 1 | 2025 | Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs · ICML 2025 |
Machine learning › Trustworthy machine learning › adversarial machine learning › adversarial natural language processing
adversarial prompt defense |
0.8 | 1 | 2024 | Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models · ICML 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.8 | 1 | 2024 | Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models · ICML 2024 |
Machine learning › Trustworthy machine learning
toxicity reduction |
0.8 | 1 | 2024 | Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models · ICML 2024 |
Natural language and speech › Language models and text generation
controllable text generation |
0.6 | 1 | 2022 | Self-conditioning Pre-Trained Language Models · ICML 2022 |
Machine learning › Trustworthy machine learning
interpretability |
0.6 | 1 | 2022 | Self-conditioning Pre-Trained Language Models · ICML 2022 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.6 | 1 | 2022 | Self-conditioning Pre-Trained Language Models · ICML 2022 |
Computer vision › 3D vision
occlusion detection |
0.1 | 1 | 2005 | Learning Spatiotemporal T-Junctions for Occlusion Detection · CVPR (2) 2005 |
Image and video processing › video segmentation
motion segmentation |
0.1 | 1 | 2005 | Learning Spatiotemporal T-Junctions for Occlusion Detection · CVPR (2) 2005 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.0 | 1 | 2004 | Bayesian Video Matting Using Learnt Image Priors · CVPR (1) 2004 |
Image and video processing
image matting |
0.0 | 1 | 2004 | Bayesian Video Matting Using Learnt Image Priors · CVPR (1) 2004 |
Image and video processing › image matting
video matting |
0.0 | 1 | 2004 | Bayesian Video Matting Using Learnt Image Priors · CVPR (1) 2004 |
Robotics › Autonomous driving
driver assistance |
0.0 | 1 | 2003 | Driver assistance: an integration of vehicle monitoring and control · ICRA 2003 |
Robotics › Autonomous driving › driver assistance
driver monitoring |
0.0 | 1 | 2003 | Driver assistance: an integration of vehicle monitoring and control · ICRA 2003 |
Computer vision › Face, body and person analysis
gaze tracking |
0.0 | 1 | 2003 | Driver assistance: an integration of vehicle monitoring and control · ICRA 2003 |
Robotics › Autonomous driving › driver assistance
lane keeping |
0.0 | 1 | 2003 | Driver assistance: an integration of vehicle monitoring and control · ICRA 2003 |
Robotics › Autonomous driving
perception |
0.0 | 1 | 2003 | Driver assistance: an integration of vehicle monitoring and control · ICRA 2003 |
Methods — techniques the papers use, named apart from their topics
activation steering · 1.4uncertainty calibration · 0.9optimal transport · 0.9equalized odds · 0.9activation editing · 0.8AUROC-based neuron selection · 0.8product of experts · 0.6relevance vector machine · 0.1canny edge detection · 0.1SIFT · 0.1image priors · 0.0bayesian inference · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Controlling Language and Diffusion Models by Transporting ActivationsabstractThe increasing capabilities of large generative models and their ever more widespread deployment have raised concerns about their reliability, safety, and potential misuse. To address these issues, recent works have proposed to control model generation by steering model activations in order to effectively induce or prevent the emergence of concepts or behaviors in the generated output.
In this paper we introduce Activation Transport (AcT), a general framework to steer activations guided by optimal transport theory that generalizes many previous activation-steering works. AcT is modality-agnostic and provides fine-grained control over the model behavior with negligible computational overhead, while minimally impacting model abilities. We experimentally show the effectiveness and versatility of our approach by addressing key challenges in large language models (LLMs) and text-to-image diffusion models (T2Is). For LLMs, we show that AcT can effectively mitigate toxicity, induce arbitrary concepts, and increase their truthfulness. In T2Is, we show how AcT enables fine-grained style control and concept negation. Pau Rodríguez, Arno Blaas, Michal Klein, Luca Zappella, Nicholas Apostoloff, Marco Cuturi, Xavier Suau |
ICLR | 5 |
| 2025 | Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMsabstractThe recent rapid adoption of large language models (LLMs) highlights the critical need for benchmarking their fairness. Conventional fairness metrics, which focus on discrete accuracy-based evaluations (i.e., prediction correctness), fail to capture the implicit impact of model uncertainty (e.g., higher model confidence about one group over another despite similar accuracy). To address this limitation, we propose an uncertainty-aware fairness metric, UCerf, to enable a fine-grained evaluation of model fairness that is more reflective of the internal bias in model decisions. Furthermore, observing data size, diversity, and clarity issues in current datasets, we introduce a new gender-occupation fairness evaluation dataset with 31,756 samples for co-reference resolution, offering a more diverse and suitable benchmark for modern LLMs. Combining our metric and dataset, we provide insightful comparisons of eight open-source LLMs. For example, Mistral-8B exhibits suboptimal fairness due to high confidence in incorrect predictions, a detail overlooked by Equalized Odds but captured by UCerF. Overall, this work provides a holistic framework for LLM evaluation by jointly assessing fairness and uncertainty, enabling the development of more transparent and accountable AI systems. Yinong Wang 0001, Nivedha Sivakumar, Falaah Arif Khan, Katherine Metcalf, Adam Golinski, Natalie Mackraz, Barry-John Theobald, Luca Zappella, Nicholas Apostoloff |
ICML | 9 |
| 2024 | Whispering Experts: Neural Interventions for Toxicity Mitigation in Language ModelsabstractAn important issue with Large Language Models (LLMs) is their undesired ability to generate toxic language. In this work, we show that the neurons responsible for toxicity can be determined by their power to discriminate toxic sentences, and that toxic language can be mitigated by reducing their activation levels proportionally to this power. We propose AUROC adaptation (AurA), an intervention that can be applied to any pre-trained LLM to mitigate toxicity. As the intervention is proportional to the ability of each neuron to discriminate toxic content, it is free of any model-dependent hyperparameters. We show that AurA can achieve up to $2.2\times$ reduction in toxicity with only a $0.72$ perplexity increase. We also show that AurA is effective with models of different scale (from 1.5B to 40B parameters), and its effectiveness in mitigating toxic language, while preserving common-sense zero-shot abilities, holds across all scales. AurA can be combined with pre-prompting strategies, boosting its average mitigation potential from $1.28\times$ to $2.35\times$. Moreover, AurA can counteract adversarial pre-prompts that maliciously elicit toxic content, making it an effective method for deploying safer and less toxic models. Xavier Suau, Pieter Delobelle, Katherine Metcalf, Armand Joulin, Nicholas Apostoloff, Luca Zappella, Pau Rodríguez |
ICML | 5 |
| 2023 | On the Role of LIP Articulation in Visual Speech PerceptionabstractGenerating realistic lip motion from audio to simulate speech production is critical for driving natural character animation. Previous research has shown that traditional metrics used to optimize and assess models for generating lip motion from speech are not a good indicator of subjective opinion of animation quality. Devising metrics that align with subjective opinion first requires understanding what impacts human perception of quality. In this work, we focus on the degree of articulation and run a series of experiments to study how articulation strength impacts human perception of lip motion accompanying speech. Specifically, we study how increasing under-articulated (dampened) and over-articulated (exaggerated) lip motion affects human perception of quality. We examine the impact of articulation strength on human perception when considering only lip motion, where viewers are presented with talking faces represented by landmarks, and in the context of embodied characters, where viewers are presented with photo-realistic videos. Our results show that viewers prefer over-articulated lip motion consistently more than under-articulated lip motion and that this preference generalizes across different speakers and embodiments. Zakaria Aldeneh, Masha Fedzechkina, Skyler Seto, Katherine Metcalf, Miguel Sarabia, Nicholas Apostoloff, Barry-John Theobald |
ICASSP | 6 |
| 2023 | Spatial LibriSpeech: An Augmented Dataset for Spatial Audio LearningabstractWe present Spatial LibriSpeech, a spatial audio dataset with over 650 hours of 19-channel audio, first-order ambisonics, and optional distractor noise. Spatial LibriSpeech is designed for machine learning model training, and it includes labels for source position, speaking direction, room acoustics and geometry. Spatial LibriSpeech is generated by augmenting LibriSpeech samples with 200k+ simulated acoustic conditions across 8k+ synthetic rooms. To demonstrate the utility of our dataset, we train models on four spatial audio tasks, resulting in a median absolute error of 6.60{\deg} on 3D source localization, 0.43m on distance, 90.66ms on T30, and 2.74dB on DRR estimation. We show that the same models generalize well to widely-used evaluation datasets, e.g., obtaining a median absolute error of 12.43{\deg} on 3D source localization on TUT Sound Events 2018, and 157.32ms on T30 estimation on ACE Challenge. Miguel Sarabia, Elena Menyaylenko, Alessandro Toso, Skyler Seto, Zakaria Aldeneh, Shadi Pirhosseinloo, Luca Zappella, Barry-John Theobald, Nicholas Apostoloff, Jonathan Sheaffer |
INTERSPEECH | 9 |
| 2022 | Self-conditioning Pre-Trained Language ModelsabstractIn this paper we aim to investigate the mechanisms that guide text generation with pre-trained Transformer-based Language Models (TLMs). Grounded on the Product of Experts formulation by Hinton (1999), we describe a generative mechanism that exploits expert units which naturally exist in TLMs. Such units are responsible for detecting concepts in the input and conditioning text generation on such concepts. We describe how to identify expert units and how to activate them during inference in order to induce any desired concept in the generated output. We find that the activation of a surprisingly small amount of units is sufficient to steer text generation (as little as 3 units in a model with 345M parameters). While the objective of this work is to learn more about how TLMs work, we show that our method is effective for conditioning without fine-tuning or using extra parameters, even on fine-grained homograph concepts. Additionally, we show that our method can be used to correct gender bias present in the output of TLMs and achieves gender parity for all evaluated contexts. We compare our method with FUDGE and PPLM-BoW, and show that our approach is able to achieve gender parity at a lower perplexity and better Self-BLEU score. The proposed method is accessible to a wide audience thanks to its simplicity and minimal compute needs. The findings in this paper are a step forward in understanding the generative mechanisms of TLMs. Xavier Suau, Luca Zappella, Nicholas Apostoloff |
ICML | 3 |
| 2021 | MorphGAN: One-Shot Face Synthesis GAN for Detecting Recognition Bias
Nataniel Ruiz, Barry-John Theobald, Anurag Ranjan, Ahmed Hussen Abdelaziz, Nicholas Apostoloff |
BMVC | 5 |
| 2021 | Multimodal Punctuation Prediction with Contextual DropoutabstractAutomatic speech recognition (ASR) is widely used in consumer electronics. ASR greatly improves the utility and accessibility of technology, but usually the output is only word sequences without punctuation. This can result in ambiguity in inferring user-intent. We first present a transformer-based approach for punctuation prediction that achieves 8% improvement on the IWSLT 2012 TED Task, beating the previous state of the art [1]. We next describe our multimodal model that learns from both text and audio, which achieves 8% improvement over the text-only algorithm on an internal dataset for which we have both the audio and transcriptions. Finally, we present an approach to learning a model using contextual dropout that allows us to handle variable amounts of future context at test time. Andrew Silva, Barry-John Theobald, Nicholas Apostoloff |
ICASSP | 3 |
| 2020 | Modality Dropout for Improved Performance-driven Talking FacesabstractWe describe our novel deep learning approach for driving animated faces using both acoustic and visual information. In particular, speech-related facial movements are generated using audiovisual information, and non-verbal facial movements are generated using only visual information. To ensure that our model exploits both modalities during training, batches are generated that contain audio-only, video-only, and audiovisual input features. The probability of dropping a modality allows control over the degree to which the model exploits audio and visual information during training. Our trained model runs in real-time on resource limited hardware (e.g. a smart phone), it is user agnostic, and it is not dependent on a potentially error-prone transcription of the speech. We use subjective testing to demonstrate: 1) the improvement of audiovisual-driven animation over the equivalent video-only approach, and 2) the improvement in the animation of speech-related facial movements after introducing modality dropout. Without modality dropout, viewers prefer audiovisual-driven animation in 51% of the test sequences compared with only 18% for video-driven. After introducing dropout viewer preference for audiovisual-driven animation increases to 74%, but decreases to 8% for video-only. Ahmed Hussen Abdelaziz, Barry-John Theobald, Paul Dixon, Reinhard Knothe, Nicholas Apostoloff, Sachin Kajareker |
ICMI | 5 |
| 2020 | Filter Distillation for Network CompressionabstractIn this paper we introduce Principal Filter Analysis (PFA), an easy to use and effective method for neural network compression. PFA exploits the correlation between filter responses within network layers to recommend a smaller network that maintain as much as possible the accuracy of the full model. We propose two algorithms: the first allows users to target compression to specific network property, such as number of trainable variable (footprint), and produces a compressed model that satisfies the requested property while preserving the maximum amount of spectral energy in the responses of each layer, while the second is a parameter-free heuristic that selects the compression used at each layer by trying to mimic an ideal set of uncorrelated responses. Since PFA compresses networks based on the correlation of their responses we show in our experiments that it gains the additional flexibility of adapting each architecture to a specific domain while compressing. PFA is evaluated against several architectures and datasets, and shows considerable compression rates without compromising accuracy, e.g., for VGG-16 on CIFAR-10, CIFAR-100 and ImageNet, PFA achieves a compression rate of 8×, 3×, and 1.4× with an accuracy gain of 0.4%, 1.4% points, and 2.4% respectively. Our tests show that PFA is competitive with state-of-the-art approaches while removing adoption barriers thanks to its practical implementation, intuitive philosophy and ease of use. Xavier Suau, Luca Zappella, Nicholas Apostoloff |
WACV | 3 |
| 2019 | Speaker-Independent Speech-Driven Visual Speech Synthesis using Domain-Adapted Acoustic ModelsabstractSpeech-driven visual speech synthesis involves mapping acoustic speech features to the corresponding lip animation controls for a face model. This mapping can take many forms, but a powerful approach is to use deep neural networks (DNNs). The lack of synchronized audio, video, and depth data is a limitation to reliably train DNNs, especially for speaker-independent models. In this paper, we investigate adapting an automatic speech recognition (ASR) acoustic model (AM) for the visual speech synthesis problem. We train the ASR-AM on ten thousand hours of audio-only transcribed speech. The ASR-AM is then adapted to the visual speech synthesis domain using ninety hours of synchronized audio-visual speech. Using a subjective assessment test, we compared the performance of the AM-initialized DNN to a randomly initialized model. The results show that viewers significantly prefer animations generated from the AM-initialized DNN than the ones generated using the randomly initialized model. We conclude that visual speech synthesis can significantly benefit from the powerful representation of speech in the ASR acoustic models. Ahmed Hussen Abdelaziz, Barry-John Theobald, Justin Binder, Gabriele Fanelli, Paul Dixon, Nicholas Apostoloff, Thibaut Weise, Sachin Kajareker |
ICMI | 6 |
| 2019 | Mirroring to Build Trust in Digital AssistantsabstractWe describe experiments towards building a conversational digital assistant that considers the preferred conversational style of the user. In particular, these experiments are designed to measure whether users prefer and trust an assistant whose conversational style matches their own. To this end we conducted a user study where subjects interacted with a digital assistant that responded in a way that either matched their conversational style, or did not. Using self-reported personality attributes and subjects' feedback on the interactions, we built models that can reliably predict a user's preferred conversational style. Katherine Metcalf, Barry-John Theobald, Garrett Weinberg, Robert Lee, Ing-Marie Jonsson, Russell Webb, Nicholas Apostoloff |
INTERSPEECH | 7 |
| 2006 | Automatic Video Segmentation using Spatiotemporal T-JunctionsabstractThe problem of figure–ground segmentation is of great importance in both video editing and visual perception tasks. Classical video segmentation algorithms approach the problem from one of two perspectives. At one extreme, global approaches constrain the camera motion to simplify the image structure. At the other extreme, local approaches estimate motion in small image regions over a small number of frames and tend to produce noisy signals that are difficult to segment. With recent advances in image segmentation showing that sparse information is often sufficient for figure– ground segmentation it seems surprising then that with the extra temporal information of video, an unconstrained automatic figure–ground segmentation algorithm still eludes the research community. In this paper we present an automatic video segmentation algorithm that is intermediate between these two extremes and uses spatiotemporal features to regularize the segmentation. Detecting spatiotemporal T-junctions that indicate occlusion edges, we learn an occlusion edge model that is used within a colour contrast sensitive MRF to segment individual frames of a video sequence. T-junctions are learnt and classified using a support vector machine and a Gaussian mixture model is fitted to the (foreground, background) pixel pairs sampled from the detected T-junctions. Graph cut is then used to segment each frame of the video showing that sparse occlusion edge information can automatically initialize the video segmentation problem. 1 Nicholas Apostoloff, Andrew W. Fitzgibbon |
BMVC | 1 |
| 2005 | Learning Spatiotemporal T-Junctions for Occlusion DetectionabstractThe goal of motion segmentation and layer extraction can be viewed as the detection and localization of occluding surfaces. A feature that has been shown to be a particularly strong indicator of occlusion, in both computer vision and neuroscience, is the T-junction; however, little progress has been made in T-junction detection. One reason for this is the difficulty in distinguishing false T-junctions (i.e. those not on an occluding edge) and real T-junctions in cluttered images. In addition to this, their photometric profile alone is not enough for reliable detection. This paper overcomes the first problem by searching for T-junctions not in space, but in space-time. This removes many false T-junctions and creates a simpler image structure to explore. The second problem is mitigated by learning the appearance of T-junctions in these spatiotemporal images. An RVM T-junction classifier is learnt from hand-labelled data using SIFT to capture its redundancy. This detector is then demonstrated in a novel occlusion detector that fuses Canny edges and T-junctions in the spatiotemporal domain to detect occluding edges in the spatial domain. Nicholas Apostoloff, Andrew W. Fitzgibbon |
CVPR (2) | 1 |
| 2004 | Bayesian Video Matting Using Learnt Image Priors
Nicholas Apostoloff, Andrew W. Fitzgibbon |
CVPR (1) | 1 |
| 2003 | Driver assistance: an integration of vehicle monitoring and controlabstractAbout 1.17 million people die in road crashes around the world each year. It is estimated that up to 30% of these fatalities are caused by fatigue and inattention. There are systems able to detect what is happening outside of the car, e.g., lane tracking, obstacle detection, pedestrian detection etc. Further on, there are also means for monitoring the actions of the driver. A natural step is to fuse the available data from within and outside of the car, and suggest a suitable response. This paper discusses driver assistance systems, lists a set of necessary core competencies of such a system and in particular presents a system for force-feedback in the steering wheel when crossing lanes. The presented system utilises a robust lane tracker which is experimentally evaluated for the purpose of driver assistance. In addition, preliminary results from simultaneous driver monitoring and lane tracking are presented that indicates a good correlation between the two, i.e. the driver's gaze direction and the structure of the road. These data can in turn be used for more advanced driver assistance systems in the future. Lars Petersson, Nicholas Apostoloff, Alexander Zelinsky |
ICRA | 2 |