Chris Russell 0001

dblp:57/9988-1 · also Christopher Russell 0001 · DBLP profile ↗
← Back
62ranked-venue papers
8as first author
25since 2021 · last 2025
0000-0003-1665-1759ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 8 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
abstract
To collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated counterfactual explanations (SCEs), where a model explains its prediction by modifying the input such that it would have predicted a different outcome. We evaluate whether LLMs can produce SCEs that are valid, achieving the intended outcome, and minimal, modifying the input no more than necessary. When asked to generate counterfactuals, we find that LLMs typically produce SCEs that are valid, but far from minimal, offering little insight into their decision-making behaviour. Worryingly, when asked to generate minimal counterfactuals, LLMs typically make excessively small edits that fail to change predictions. The observed validity-minimality trade-off is consistent across several LLMs, datasets, and evaluation settings. Our findings suggest that SCEs are, at best, an ineffective explainability tool and, at worst, can provide misleading insights into model behaviour. Proposals to deploy LLMs in high-stakes settings must consider the impact of unreliable self-explanations on downstream decision-making. Our code is available at https://github.com/HarryMayne/SCEs.
Harry Mayne, Ryan Othniel Kearns, Andrew M. Bean 0001, Eoin Delaney, Chris Russell 0001, Adam Mahdi
EMNLP6
2025 FairImagen: Post-Processing for Bias Mitigation in Text-to-Image Models
abstract
Text-to-image diffusion models, such as Stable Diffusion, have demonstrated remarkable capabilities in generating high-quality and diverse images from natural language prompts. However, recent studies reveal that these models often replicate and amplify societal biases, particularly along demographic attributes like gender and race. In this paper, we introduce FairImagen (https://github.com/fuzihaofzh/FairImagen), a post-hoc debiasing framework that operates on prompt embeddings to mitigate such biases without retraining or modifying the underlying diffusion model. Our method integrates Fair Principal Component Analysis to project CLIP-based input embeddings into a subspace that minimizes group-specific information while preserving semantic content. We further enhance debiasing effectiveness through empirical noise injection and propose a unified cross-demographic projection method that enables simultaneous debiasing across multiple demographic attributes. Extensive experiments across gender, race, and intersectional settings demonstrate that FairImagen significantly improves fairness with a moderate trade-off in image quality and prompt fidelity. Our framework outperforms existing post-hoc methods and offers a simple, scalable, and model-agnostic solution for equitable text-to-image generation.
Ryan Brown, Shun Shao, Kai Rawal, Eoin Delaney, Chris Russell 0001
NeurIPS6
2024 OxonFair: A Flexible Toolkit for Algorithmic Fairness
abstract
We present OxonFair, a new open source toolkit for enforcing fairness in binary classification. Compared to existing toolkits: (i) We support NLP and Computer Vision classification as well as standard tabular problems. (ii) We support enforcing fairness on validation data, making us robust to a wide range of overfitting challenges. (iii) Our approach can optimize any measure based on True Positives, False Positive, False Negatives, and True Negatives. This makes it easily extensible and much more expressive than existing toolkits. It supports all 9 and all 10 of the decision-based group metrics of two popular review articles. (iv) We jointly optimize a performance objective alongside fairness constraints. This minimizes degradation while enforcing fairness, and even improves the performance of inadequately tuned unfair baselines. OxonFair is compatible with standard ML toolkits, including sklearn, Autogluon, and PyTorch and is available at https://github.com/oxfordinternetinstitute/oxonfair.
Eoin Delaney, Sandra Wachter, Brent D. Mittelstadt 0002, Chris Russell 0001
NeurIPS5
2023 Evaluating the Fairness of Discriminative Foundation Models in Computer Vision
abstract
We propose a novel taxonomy for bias evaluation of discriminative foundation models, such as Contrastive Language-Pretraining (CLIP), that are used for labeling tasks. We then systematically evaluate existing methods for mitigating bias in these models with respect to our taxonomy. Specifically, we evaluate OpenAI’s CLIP and OpenCLIP models for key applications, such as zero-shot classification, image retrieval and image captioning. We categorize desired behaviors based around three axes: (i) if the task concerns humans; (ii) how subjective the task is (i.e., how likely it is that people from a diverse range of backgrounds would agree on a labeling); and (iii) the intended purpose of the task and if fairness is better served by impartiality (i.e., making decisions independent of the protected attributes) or representation (i.e., making decisions to maximize diversity). Finally, we provide quantitative fairness evaluations for both binary-valued and multi-valued protected attributes over ten diverse datasets. We find that fair PCA, a post-processing method for fair representations, works very well for debiasing in most of the aforementioned tasks while incurring only minor loss of performance. However, different debiasing approaches vary in their effectiveness depending on the task. Hence, one should choose the debiasing approach depending on the specific use case.
Junaid Ali 0001, Matthäus Kleindessner, Florian Wenzel, Kailash Budhathoki, Volkan Cevher, Chris Russell 0001
AIES6
2023 Efficient fair PCA for fair representation learning
abstract
We revisit the problem of fair principal component analysis (PCA), where the goal is to learn the best low-rank linear approximation of the data that obfuscates demographic information. We propose a conceptually simple approach that allows for an analytic solution similar to standard PCA and can be kernelized. Our methods have the same complexity as standard PCA, or kernel PCA, and run much faster than existing methods for fair PCA based on semidefinite programming or manifold optimization, while achieving similar results.
Matthäus Kleindessner, Michele Donini, Chris Russell 0001, Muhammad Bilal Zafar
AISTATS3
2023 Learning Adaptive Neighborhoods for Graph Neural Networks
abstract
Graph convolutional networks (GCNs) enable end-to-end learning on graph structured data. However, many works assume a given graph structure. When the input graph is noisy or unavailable, one approach is to construct or learn a latent graph structure. These methods typically fix the choice of node degree for the entire graph, which is suboptimal. Instead, we propose a novel end-to-end differentiable graph generator which builds graph topologies where each node selects both its neighborhood and its size. Our module can be readily integrated into existing pipelines involving graph convolution operations, replacing the predetermined or existing adjacency matrix with one that is learned, and optimized, as part of the general objective. As such it is applicable to any GCN. We integrate our module into trajectory prediction, point cloud classification and node classification pipelines resulting in improved accuracy over other structure-learning methods across a wide range of datasets and GCN backbones. We will release the code.
Avishkar Saha, Oscar Mendez 0001, Chris Russell 0001, Richard Bowden
ICCV3
2023 Kick Back & Relax: Learning to Reconstruct the World by Watching SlowTV
abstract
Self-supervised monocular depth estimation (SS-MDE) has the potential to scale to vast quantities of data. Unfortunately, existing approaches limit themselves to the automotive domain, resulting in models incapable of generalizing to complex environments such as natural or indoor settings.To address this, we propose a large-scale SlowTV dataset curated from YouTube, containing an order of magnitude more data than existing automotive datasets. SlowTV contains 1.7M images from a rich diversity of environments, such as worldwide seasonal hiking, scenic driving and scuba diving. Using this dataset, we train an SS-MDE model that provides zero-shot generalization to a large collection of indoor/outdoor datasets. The resulting model outperforms all existing SSL approaches and closes the gap on supervised SoTA, despite using a more efficient architecture.We additionally introduce a collection of best-practices to further maximize performance and zero-shot generalization. This includes 1) aspect ratio augmentation, 2) camera intrinsic estimation, 3) support frame randomization and 4) flexible motion estimation. Code is available at https://github.com/jspenmar/slowtv_monodepth.
Jaime Spencer, Simon Hadfield, Chris Russell 0001, Richard Bowden
ICCV3
2023 When do Minimax-fair Learning and Empirical Risk Minimization Coincide?
abstract
Minimax-fair machine learning minimizes the error for the worst-off group. However, empirical evidence suggests that when sophisticated models are trained with standard empirical risk minimization (ERM), they often have the same performance on the worst-off group as a minimax-trained model. Our work makes this counter-intuitive observation concrete. We prove that if the hypothesis class is sufficiently expressive and the group information is recoverable from the features, ERM and minimax-fairness learning formulations indeed have the same performance on the worst-off group. We provide additional empirical evidence of how this observation holds on a wide range of datasets and hypothesis classes. Since ERM is fundamentally easier than minimax optimization, our findings have implications on the practice of fair machine learning.
Harvineet Singh, Matthäus Kleindessner, Volkan Cevher, Rumi Chunara, Chris Russell 0001
ICML5
2023 Translating Images into Maps (Extended Abstract)
abstract
We approach instantaneous mapping, converting images to a top-down view of the world, as a translation problem. We show how a novel form of transformer network can be used to map from images and video directly to an overhead map or bird's-eye-view (BEV) of the world, in a single end-to-end network. We assume a 1-1 correspondence between a vertical scanline in the image, and rays passing through the camera location in an overhead map. This lets us formulate map generation from an image as a set of sequence-to-sequence translations. This constrained formulation, based upon a strong physical grounding of the problem, leads to a restricted transformer network that is convolutional in the horizontal direction only. The structure allows us to make efficient use of data when training, and obtains state-of-the-art results for instantaneous mapping of three large-scale datasets, including a 15\% and 30\% relative gain against existing best performing methods on the nuScenes and Argoverse datasets, respectively.
Avishkar Saha, Oscar Mendez 0001, Chris Russell 0001, Richard Bowden
IJCAI3
2023 Explaining holistic image regressors and classifiers in urban analytics with plausible counterfactuals
abstract
We propose a new form of plausible counterfactual explanation designed to explain the behaviour of computer vision systems used in urban analytics that make predictions based on properties across the entire image, rather than specific regions of it. We illustrate the merits of our approach by explaining computer vision models used to analyse street imagery, which are now widely used in GeoAI and urban analytics. Such explanations are important in urban analytics as researchers and practioners are increasingly reliant on it for decision making. Finally, we perform a user study that demonstrate our approach can be used by non-expert users, who might not be machine learning experts, to be more confident and to better understand the behaviour of image-based classifiers/regressors for street view analysis. Furthermore, the method can potentially be used as an engagement tool to visualise how public spaces can plausibly look like. The limited realism of the counterfactuals is a concern which we hope to improve in the future.
Stephen Law, Rikuo Hasegawa, Brooks Paige, Chris Russell 0001, Andrew Elliott
Int. J. Geogr. Inf. Sci.4
2022 Pairwise Fairness for Ordinal Regression
abstract
We initiate the study of fairness for ordinal regression. We adapt two fairness notions previously considered in fair ranking and propose a strategy for training a predictor that is approximately fair according to either notion. Our predictor has the form of a threshold model, composed of a scoring function and a set of thresholds, and our strategy is based on a reduction to fair binary classification for learning the scoring function and local search for choosing the thresholds. We provide generalization guarantees on the error and fairness violation of our predictor, and we illustrate the effectiveness of our approach in extensive experiments.
Matthäus Kleindessner, Samira Samadi, Muhammad Bilal Zafar, Krishnaram Kenthapadi, Chris Russell 0001
AISTATS5
2022 'The Pedestrian next to the Lamppost" Adaptive Object Graphs for Better Instantaneous Mapping
abstract
Estimating a semantically segmented bird's-eye-view (BEV) map from a single image has become a popular technique for autonomous control and navigation. However, they show an increase in localization error with distance from the camera. While such an increase in error is entirely expected -localization is harder at distance - much of the drop in performance can be attributed to the cues used by current texture-based models, in particular, they make heavy use of object-ground intersections (such as shadows) [10], which become increasingly sparse and uncertain for distant objects. In this work, we address these shortcomings in BEV-mapping by learning the spatial relationship between objects in a scene. We propose a graph neural network which predicts BEV objects from a monocular image by spatially reasoning about an object within the context of other objects. Our approach sets a new state-of-the-art in BEV estimation from monocular images across three large-scale datasets, including a 50% relative improvement for objects on nuScenes.
Avishkar Saha, Oscar Mendez 0001, Chris Russell 0001, Richard Bowden
CVPR3
2022 Leveling Down in Computer Vision: Pareto Inefficiencies in Fair Deep Classifiers
abstract
Algorithmic fairness is frequently motivated in terms of a trade-off in which overall performance is decreased so as to improve performance on disadvantaged groups where the algorithm would otherwise be less accurate. Contrary to this, we find that applying existing fairness approaches to computer vision improve fairness by degrading the performance of classifiers across all groups (with increased degradation on the best performing groups). Extending the bias-variance decomposition for classification to fairness, we theoretically explain why the majority of fairness methods designed for low capacity models should not be used in settings involving high-capacity models, a scenario common to computer vision. We corroborate this analysis with extensive experimental support that shows that many of the fairness heuristics used in computer vision also degrade performance on the most disadvantaged groups. Building on these insights, we propose an adaptive augmentation strategy that, uniquely, of all methods tested, improves performance for the disadvantaged groups.
Dominik Zietlow, Michael Lohaus, Guha Balakrishnan, Matthäus Kleindessner, Francesco Locatello, Bernhard Schölkopf, Chris Russell 0001
CVPR7
2022 Visual Representation Learning Does Not Generalize Strongly Within the Same Domain
Lukas Schott, Julius von Kügelgen, Frederik Träuble, Peter V. Gehler, Chris Russell 0001, Matthias Bethge, Bernhard Schölkopf, Francesco Locatello, Wieland Brendel
ICLR5
2022 Active Sampling for Min-Max Fairness
abstract
We propose simple active sampling and reweighting strategies for optimizing min-max fairness that can be applied to any classification or regression model learned via loss minimization. The key intuition behind our approach is to use at each timestep a datapoint from the group that is worst off under the current model for updating the model. The ease of implementation and the generality of our robust formulation make it an attractive option for improving model performance on disadvantaged groups. For convex learning problems, such as linear or logistic regression, we provide a fine-grained analysis, proving the rate of convergence to a min-max fair solution.
Jacob D. Abernethy, Pranjal Awasthi, Matthäus Kleindessner, Jamie Morgenstern, Chris Russell 0001, Jie Zhang 0096
ICML5
2022 Score Matching Enables Causal Discovery of Nonlinear Additive Noise Models
abstract
This paper demonstrates how to recover causal graphs from the score of the data distribution in non-linear additive (Gaussian) noise models. Using score matching algorithms as a building block, we show how to design a new generation of scalable causal discovery methods. To showcase our approach, we also propose a new efficient method for approximating the score’s Jacobian, enabling to recover the causal graph. Empirically, we find that the new algorithm, called SCORE, is competitive with state-of-the-art causal discovery methods while being significantly faster.
Paul Rolland, Volkan Cevher, Matthäus Kleindessner, Chris Russell 0001, Dominik Janzing, Bernhard Schölkopf, Francesco Locatello
ICML4
2022 Translating Images into Maps
abstract
We approach instantaneous mapping, converting images to a top-down view of the world, as a translation problem. We show how a novel form of transformer network can be used to map from images and video directly to an overhead map or bird's-eye-view (BEV) of the world, in a single end-to-end network. We assume a 1–1 correspondence between a vertical scanline in the image, and rays passing through the camera location in an overhead map. This lets us formulate map generation from an image as a set of sequence-to-sequence translations. Posing the problem as translation allows the network to use the context of the image when interpreting the role of each pixel. This constrained formulation, based upon a strong physical grounding of the problem, leads to a restricted transformer network that is convolutional in the horizontal direction only. The structure allows us to make efficient use of data when training, and obtains state-of-the-art results for instantaneous mapping of three large-scale datasets, including a 15% and 30% relative gain against existing best performing methods on the nuScenes and Argoverse datasets, respectively.
Avishkar Saha, Oscar Mendez 0001, Chris Russell 0001, Richard Bowden
ICRA3
2022 Are Two Heads the Same as One? Identifying Disparate Treatment in Fair Neural Networks
abstract
We show that deep networks trained to satisfy demographic parity often do so through a form of race or gender awareness, and that the more we force a network to be fair, the more accurately we can recover race or gender from the internal state of the network. Based on this observation, we investigate an alternative fairness approach: we add a second classification head to the network to explicitly predict the protected attribute (such as race or gender) alongside the original task. After training the two-headed network, we enforce demographic parity by merging the two heads, creating a network with the same architecture as the original network. We establish a close relationship between existing approaches and our approach by showing (1) that the decisions of a fair classifier are well-approximated by our approach, and (2) that an unfair and optimally accurate classifier can be recovered from a fair classifier and our second head predicting the protected attribute. We use our explicit formulation to argue that the existing fairness approaches, just as ours, demonstrate disparate treatment and that they are likely to be unlawful in a wide range of scenarios under US law.
Michael Lohaus, Matthäus Kleindessner, Krishnaram Kenthapadi, Francesco Locatello, Chris Russell 0001
NeurIPS5
2022 Assaying Out-Of-Distribution Generalization in Transfer Learning
abstract
Since out-of-distribution generalization is a generally ill-posed problem, various proxy targets (e.g., calibration, adversarial robustness, algorithmic corruptions, invariance across shifts) were studied across different research programs resulting in different recommendations. While sharing the same aspirational goal, these approaches have never been tested under the same experimental conditions on real data. In this paper, we take a unified view of previous work, highlighting message discrepancies that we address empirically, and providing recommendations on how to measure the robustness of a model and how to improve it. To this end, we collect 172 publicly available dataset pairs for training and out-of-distribution evaluation of accuracy, calibration error, adversarial attacks, environment invariance, and synthetic corruptions. We fine-tune over 31k networks, from nine different architectures in the many- and few-shot setting. Our findings confirm that in- and out-of-distribution accuracies tend to increase jointly, but show that their relation is largely dataset-dependent, and in general more nuanced and more complex than posited by previous, smaller scale studies.
Florian Wenzel, Andrea Dittadi, Peter V. Gehler, Carl-Johann Simon-Gabriel, Max Horn, Dominik Zietlow, David Kernert, Chris Russell 0001, Thomas Brox, Bernt Schiele, Bernhard Schölkopf, Francesco Locatello
NeurIPS8
2022 4D Temporally Coherent Multi-Person Semantic Reconstruction and Segmentation
abstract
Abstract We introduce the first approach to solve the challenging problem of automatic 4D visual scene understanding for complex dynamic scenes with multiple interacting people from multi-view video. Our approach simultaneously estimates a detailed model that includes a per-pixel semantically and temporally coherent reconstruction, together with instance-level segmentation exploiting photo-consistency, semantic and motion information. We further leverage recent advances in 3D pose estimation to constrain the joint semantic instance segmentation and 4D temporally coherent reconstruction. This enables per person semantic instance segmentation of multiple interacting people in complex dynamic scenes. Extensive evaluation of the joint visual scene understanding framework against state-of-the-art methods on challenging indoor and outdoor sequences demonstrates a significant ( $$\approx 40\%$$ ≈ 40 % ) improvement in semantic segmentation, reconstruction and scene flow accuracy. In addition to the evaluation on several indoor and outdoor scenes, the proposed joint 4D scene understanding framework is applied to challenging outdoor sports scenes in the wild captured with manually operated wide-baseline broadcast cameras.
Armin Mustafa, Chris Russell 0001, Adrian Hilton 0001
Int. J. Comput. Vis.2
2021 Fair and Adequate Explanations
Nicholas Asher, Soumya Paul, Chris Russell 0001
CD-MAKE3
2021 Explaining Classifiers Using Adversarial Perturbations on the Perceptual Ball
abstract
We present a simple regularization of adversarial perturbations based upon the perceptual loss. While the resulting perturbations remain imperceptible to the human eye, they differ from existing adversarial perturbations in that they are semi-sparse alterations that highlight objects and regions of interest while leaving the background unaltered. As a semantically meaningful adverse perturbations, it forms a bridge between counterfactual explanations and adversarial perturbations in the space of images.We evaluate our approach on several standard explainability benchmarks, namely, weak localization, insertion-deletion, and the pointing game demonstrating that perceptually regularized counterfactuals are an effective explanation for image-based classifiers.
Andrew Elliott, Stephen Law, Chris Russell 0001
CVPR3
2021 Human Pose Manipulation and Novel View Synthesis using Differentiable Rendering
abstract
We present a new approach for synthesizing novel views of people in new poses. Our novel differentiable renderer enables the synthesis of highly realistic images from any viewpoint. Rather than operating over mesh-based structures, our renderer makes use of diffuse Gaussian primitives that directly represent the underlying skeletal structure of a human. Rendering these primitives gives results in a high-dimensional latent image, which is then transformed into an RGB image by a decoder network. The formulation gives rise to a fully differentiable framework that can be trained end-to-end. We demonstrate the effectiveness of our approach to image reconstruction on both the Human3.6M and Panoptic Studio datasets. We show how our approach can be used for motion transfer between individuals; novel view synthesis of individuals captured from just a single camera; to synthesize individuals from any virtual viewpoint; and to re-render people in novel poses. Code and video results are available at https://github.com/GuillaumeRochette/HumanViewSynthesis.
Guillaume Rochette, Chris Russell 0001, Richard Bowden
FG2
2021 Enabling spatio-temporal aggregation in Birds-Eye-View Vehicle Estimation
abstract
Constructing Birds-Eye-View (BEV) maps from monocular images is typically a complex multi-stage process involving the separate vision tasks of ground plane estimation, road segmentation and 3D object detection. However, recent approaches have adopted end-to-end solutions which warp image-based features from the image-plane to BEV while implicitly taking account of camera geometry. In this work, we show how such instantaneous BEV estimation of a scene can be learnt, and a better state estimation of the world can be achieved by incorporating temporal information. Our model learns a representation from monocular video through factorised 3D convolutions and uses this to estimate a BEV occupancy grid of the final frame. We achieve state-of-the-art results for BEV estimation from monocular images, and establish a new benchmark for single-scene BEV estimation from monocular video.
Avishkar Saha, Oscar Mendez 0001, Chris Russell 0001, Richard Bowden
ICRA3
2021 Why fairness cannot be automated: Bridging the gap between EU non-discrimination law and AI
Sandra Wachter, Brent D. Mittelstadt 0002, Chris Russell 0001
Comput. Law Secur. Rev.3
2020 What Did You Think Would Happen? Explaining Agent Behaviour through Intended Outcomes
abstract
We present a novel form of explanation for Reinforcement Learning, based around the notion of intended outcome. These explanations describe the outcome an agent is trying to achieve by its actions. We provide a simple proof that general methods for post-hoc explanations of this nature are impossible in traditional reinforcement learning. Rather, the information needed for the explanations must be collected in conjunction with training the agent. We derive approaches designed to extract local explanations based on intention for several variants of Q-function approximation and prove consistency between the explanations and the Q-values learned. We demonstrate our method on multiple reinforcement learning problems, and provide code to help researchers introspecting their RL environments and algorithms.
Herman Yau, Chris Russell 0001, Simon Hadfield
NeurIPS2
2019 Weakly-Supervised 3D Pose Estimation from a Single Image using Multi-View Consistency
Guillaume Rochette, Chris Russell 0001, Richard Bowden
BMVC2
2019 U4D: Unsupervised 4D Dynamic Scene Understanding
abstract
We introduce the first approach to solve the challenging problem of unsupervised 4D visual scene understanding for complex dynamic scenes with multiple interacting people from multi-view video. Our approach simultaneously estimates a detailed model that includes a per-pixel semantically and temporally coherent reconstruction, together with instance-level segmentation exploiting photo-consistency, semantic and motion information. We further leverage recent advances in 3D pose estimation to constrain the joint semantic instance segmentation and 4D temporally coherent reconstruction. This enables per person semantic instance segmentation of multiple interacting people in complex dynamic scenes. Extensive evaluation of the joint visual scene understanding framework against state-of-the-art methods on challenging indoor and outdoor sequences demonstrates a significant (approx 40%) improvement in semantic segmentation, reconstruction and scene flow accuracy.
Armin Mustafa, Chris Russell 0001, Adrian Hilton 0001
ICCV2
2019 Making Decisions that Reduce Discriminatory Impacts
abstract
As machine learning algorithms move into real-world settings, it is crucial to ensure they are aligned with societal values. There has been much work on one aspect of this, namely the discriminatory prediction problem: How can we reduce discrimination in the predictions themselves? While an important question, solutions to this problem only apply in a restricted setting, as we have full control over the predictions. Often we care about the non-discrimination of quantities we do not have full control over. Thus, we describe another key aspect of this challenge, the discriminatory impact problem: How can we reduce discrimination arising from the real-world impact of decisions? To address this, we describe causal methods that model the relevant parts of the real-world system in which the decisions are made. Unlike previous approaches, these models not only allow us to map the causal pathway of a single decision, but also to model the effect of interference–how the impact on an individual depends on decisions made about other people. Often, the goal of decision policies is to maximize a beneficial impact overall. To reduce the discrimination of these benefits, we devise a constraint inspired by recent work in counterfactual fairness, and give an efficient procedure to solve the constrained optimization problem. We demonstrate our approach with an example: how to increase students taking college entrance exams in New York City public schools.
Matt J. Kusner, Chris Russell 0001, Joshua R. Loftus, Ricardo Silva 0001
ICML2
2019 Fixing Implicit Derivatives: Trust-Region Based Learning of Continuous Energy Functions
abstract
We present a new technique for the learning of continuous energy functions that we refer to as Wibergian Learning. One common approach to inverse problems is to cast them as an energy minimisation problem, where the minimum cost solution found is used as an estimator of hidden parameters. Our new approach formally characterises the dependency between weights that control the shape of the energy function, and the location of minima, by describing minima as fixed points of optimisation methods. This allows for the use of gradient-based end-to- end training to integrate deep-learning and the classical inverse problem methods. We show how our approach can be applied to obtain state-of-the-art results in the diverse applications of tracker fusion and multiview 3D reconstruction.
Chris Russell 0001, Matteo Toso, Neill D. F. Campbell
NeurIPS1
2019 Linear programming-based submodular extensions for marginal estimation
Pankaj Pansari, Chris Russell 0001, M. Pawan Kumar
Comput. Vis. Image Underst.2
2019 Take a Look Around: Using Street View and Satellite Images to Estimate House Prices
abstract
When an individual purchases a home, they simultaneously purchase its structural features, its accessibility to work, and the neighborhood amenities. Some amenities, such as air quality, are measurable while others, such as the prestige or the visual impression of a neighborhood, are difficult to quantify. Despite the well-known impacts intangible housing features have on house prices, limited attention has been given to systematically quantifying these difficult to measure amenities. Two issues have led to this neglect. Not only do few quantitative methods exist that can measure the urban environment, but that the collection of such data is both costly and subjective. We show that street image and satellite image data can capture these urban qualities and improve the estimation of house prices. We propose a pipeline that uses a deep neural network model to automatically extract visual features from images to estimate house prices in London, UK. We make use of traditional housing features such as age, size, and accessibility as well as visual features from Google Street View images and Bing aerial images in estimating the house price model. We find encouraging results where learning to characterize the urban quality of a neighborhood improves house price prediction, even when generalizing to previously unseen London boroughs. We explore the use of non-linear vs. linear methods to fuse these cues with conventional models of house pricing, and show how the interpretability of linear models allows us to directly extract proxy variables for visual desirability of neighborhoods that are both of interest in their own right, and could be used as inputs to other econometric methods. This is particularly valuable as once the network has been trained with the training data, it can be applied elsewhere, allowing us to generate vivid dense maps of the visual appeal of London streets.
Stephen Law, Brooks Paige, Chris Russell 0001
ACM Trans. Intell. Syst. Technol.3
2018 Rethinking Pose in 3D: Multi-stage Refinement and Recovery for Markerless Motion Capture
abstract
We propose a CNN-based approach for multi-camera markerless motion capture of the human body. Unlike existing methods that first perform pose estimation on individual cameras and generate 3D models as post-processing, our approach makes use of 3D reasoning throughout a multi-stage approach. This novelty allows us to use provisional 3D models of human pose to rethink where the joints should be located in the image and to recover from past mistakes. Our principled refinement of 3D human poses lets us make use of image cues, even from images where we previously misdetected joints, to refine our estimates as part of an end-to-end approach. Finally, we demonstrate how the high-quality output of our multi-camera setup can be used as an additional training source to improve the accuracy of existing single camera models.
Denis Tomè, Matteo Toso, Lourdes Agapito, Chris Russell 0001
3DV4
2018 Optimal Submodular Extensions for Marginal Estimation
abstract
Submodular extensions of an energy function can be used to efficiently compute approximate marginals via variational inference. The accuracy of the marginals depends crucially on the quality of the submodular extension. To identify the best possible extension, we show an equivalence between the submodular extensions of the energy and the objective functions of linear programming (LP) relaxations for the corresponding MAP estimation problem. This allows us to (i) establish the optimality of the submodular extension for Potts model used in the literature; (ii) identify the optimal submodular extension for the more general class of metric labeling; and (iii) efficiently compute the marginals for the widely used dense CRF model using a recently proposed Gaussian filtering method. Using both synthetic and real data, we show that our approach provides comparable upper bounds on the log-partition function to those obtained using tree-reweighted message passing (TRW) in cases where the latter is computationally feasible. Importantly, unlike TRW, our approach provides the first practical algorithm to compute an upper bound on the dense CRF model.
Pankaj Pansari, Chris Russell 0001, M. Pawan Kumar
AISTATS2
2017 Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image
abstract
We propose a unified formulation for the problem of 3D human pose estimation from a single raw RGB image that reasons jointly about 2D joint estimation and 3D pose reconstruction to improve both tasks. We take an integrated approach that fuses probabilistic knowledge of 3D human pose with a multi-stage CNN architecture and uses the knowledge of plausible 3D landmark locations to refine the search for better 2D locations. The entire process is trained end-to-end, is extremely efficient and obtains state-of-the-art results on Human3.6M outperforming previous approaches both on 2D and 3D errors.
Denis Tomè, Chris Russell 0001, Lourdes Agapito
CVPR2
2017 Counterfactual Fairness
abstract
Machine learning can impact people with legal or ethical consequences when it is used to automate decisions in areas such as insurance, lending, hiring, and predictive policing. In many of these scenarios, previous decisions have been made that are unfairly biased against certain subpopulations, for example those of a particular race, gender, or sexual orientation. Since this past data may be biased, machine learning predictors must account for this to avoid perpetuating or creating discriminatory practices. In this paper, we develop a framework for modeling fairness using tools from causal inference. Our definition of counterfactual fairness captures the intuition that a decision is fair towards an individual if it the same in (a) the actual world and (b) a counterfactual world where the individual belonged to a different demographic group. We demonstrate our framework on a real-world problem of fair prediction of success in law school.
Matt J. Kusner, Joshua R. Loftus, Chris Russell 0001, Ricardo Silva 0001
NIPS3
2017 When Worlds Collide: Integrating Different Counterfactual Assumptions in Fairness
abstract
Machine learning is now being used to make crucial decisions about people's lives. For nearly all of these decisions there is a risk that individuals of a certain race, gender, sexual orientation, or any other subpopulation are unfairly discriminated against. Our recent method has demonstrated how to use techniques from counterfactual inference to make predictions fair across different subpopulations. This method requires that one provides the causal model that generated the data at hand. In general, validating all causal implications of the model is not possible without further assumptions. Hence, it is desirable to integrate competing causal models to provide counterfactually fair decisions, regardless of which causal "world" is the correct one. In this paper, we show how it is possible to make predictions that are approximately fair with respect to multiple possible causal models at once, thus mitigating the problem of exact causal specification. We frame the goal of learning a fair classifier as an optimization problem with fairness constraints entailed by competing causal explanations. We show how this optimization problem can be efficiently solved using gradient-based methods. We demonstrate the flexibility of our model on two real-world fair classification problems. We show that our model can seamlessly balance fairness in multiple worlds with prediction accuracy.
Chris Russell 0001, Matt J. Kusner, Joshua R. Loftus, Ricardo Silva 0001
NIPS1
2017 VEEGAN: Reducing Mode Collapse in GANs using Implicit Variational Learning
abstract
Deep generative models provide powerful tools for distributions over complicated manifolds, such as those of natural images. But many of these methods, including generative adversarial networks (GANs), can be difficult to train, in part because they are prone to mode collapse, which means that they characterize only a few modes of the true distribution. To address this, we introduce VEEGAN, which features a reconstructor network, reversing the action of the generator by mapping from data to noise. Our training objective retains the original asymptotic consistency guarantee of GANs, and can be interpreted as a novel autoencoder loss over the noise. In sharp contrast to a traditional autoencoder over data points, VEEGAN does not require specifying a loss function over the data, but rather only over the representations, which are standard normal by assumption. On an extensive set of synthetic and real world image datasets, VEEGAN indeed resists mode collapsing to a far greater extent than other recent GAN variants, and produces more realistic samples.
Akash Srivastava, Lazar Valkov, Chris Russell 0001, Michael U. Gutmann, Charles Sutton
NIPS3
2017 Efficient minimization of higher order submodular functions using monotonic Boolean functions
Srikumar Ramalingam, Chris Russell 0001, Lubor Ladicky, Philip Torr 0001
Discret. Appl. Math.2
2016 Better Together: Joint Reasoning for Non-rigid 3D Reconstruction with Specularities and Shading
Chris Russell 0001, Lourdes Agapito, Andrew W. Fitzgibbon, Liu-Yin Qi
BMVC1
2016 Solving Jigsaw Puzzles with Linear Programming
Chris Russell 0001, Lourdes Agapito
BMVC2
2015 Part-based modelling of compound scenes from images
abstract
We propose a method to recover the structure of a compound scene from multiple silhouettes. Structure is expressed as a collection of 3D primitives chosen from a predefined library, each with an associated pose. This has several advantages over a volume or mesh representation both for estimation and the utility of the recovered model. The main challenge in recovering such a model is the combinatorial number of possible arrangements of parts. We address this issue by exploiting the intrinsic structure and sparsity of the problem, and show that our method scales to scenes constructed from large libraries of parts.
Anton van den Hengel, Chris Russell 0001, Anthony R. Dick, John W. Bastian, Daniel Pooley, Lachlan Fleming, Lourdes Agapito
CVPR2
2015 Direct, Dense, and Deformable: Template-Based Non-rigid 3D Reconstruction from RGB Video
abstract
In this paper we tackle the problem of capturing the dense, detailed 3D geometry of generic, complex non-rigid meshes using a single RGB-only commodity video camera and a direct approach. While robust and even real-time solutions exist to this problem if the observed scene is static, for non-rigid dense shape capture current systems are typically restricted to the use of complex multi-camera rigs, take advantage of the additional depth channel available in RGB-D cameras, or deal with specific shapes such as faces or planar surfaces. In contrast, our method makes use of a single RGB video as input, it can capture the deformations of generic shapes, and the depth estimation is dense, per-pixel and direct. We first compute a dense 3D template of the shape of the object, using a short rigid sequence, and subsequently perform online reconstruction of the non-rigid mesh as it evolves over time. Our energy optimization approach minimizes a robust photometric cost that simultaneously estimates the temporal correspondences and 3D deformations with respect to the template mesh. In our experimental evaluation we show a range of qualitative results on novel datasets, we compare against an existing method that requires multi-frame optical flow, and perform a quantitative evaluation against other template-based approaches on a ground truth dataset.
Chris Russell 0001, Neill D. F. Campbell, Lourdes Agapito
ICCV2
2014 Video Pop-up: Monocular 3D Reconstruction of Dynamic Scenes
Chris Russell 0001, Lourdes Agapito
ECCV (7)1
2014 Semi-supervised Learning Using an Unsupervised Atlas
Nikolaos Pitelis, Chris Russell 0001, Lourdes Agapito
ECML/PKDD (2)2
2014 Associative Hierarchical Random Fields
abstract
This paper makes two contributions: the first is the proposal of a new model-The associative hierarchical random field (AHRF), and a novel algorithm for its optimization; the second is the application of this model to the problem of semantic segmentation. Most methods for semantic segmentation are formulated as a labeling problem for variables that might correspond to either pixels or segments such as super-pixels. It is well known that the generation of super pixel segmentations is not unique. This has motivated many researchers to use multiple super pixel segmentations for problems such as semantic segmentation or single view reconstruction. These super-pixels have not yet been combined in a principled manner, this is a difficult problem, as they may overlap, or be nested in such a way that the segmentations form a segmentation tree. Our new hierarchical random field model allows information from all of the multiple segmentations to contribute to a global energy. MAP inference in this model can be performed efficiently using powerful graph cut based move making algorithms. Our framework generalizes much of the previous work based on pixels or segments, and the resulting labelings can be viewed both as a detailed segmentation at the pixel level, or at the other extreme, as a segment selector that pieces together a solution like a jigsaw, selecting the best segments from different segmentations as pieces. We evaluate its performance on some of the most challenging data sets for object class segmentation, and show that this ability to perform inference using multiple overlapping segmentations leads to state-of-the-art results.
Lubor Ladicky, Chris Russell 0001, Pushmeet Kohli, Philip Torr 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Learning a Manifold as an Atlas
abstract
In this work, we return to the underlying mathematical definition of a manifold and directly characterise learning a manifold as finding an atlas, or a set of overlapping charts, that accurately describe local structure. We formulate the problem of learning the manifold as an optimisation that simultaneously refines the continuous parameters defining the charts, and the discrete assignment of points to charts. In contrast to existing methods, this direct formulation of a manifold does not require "unwrapping" the manifold into a lower dimensional space and allows us to learn closed manifolds of interest to vision, such as those corresponding to gait cycles or camera pose. We report state-of-the-art results for manifold based nearest neighbour classification on vision datasets, and show how the same techniques can be applied to the 3D reconstruction of human motion from a single image.
Nikolaos Pitelis, Chris Russell 0001, Lourdes Agapito
CVPR2
2013 Looking Beyond the Image: Unsupervised Learning for Object Saliency and Detection
abstract
We propose a principled probabilistic formulation of object saliency as a sampling problem. This novel formulation allows us to learn, from a large corpus of unlabelled images, which patches of an image are of the greatest interest and most likely to correspond to an object. We then sample the object saliency map to propose object locations. We show that using only a single object location proposal per image, we are able to correctly select an object in over 42% of the images in the Pascal VOC 2007 dataset, substantially outperforming existing approaches. Furthermore, we show that our object proposal can be used as a simple unsupervised approach to the weakly supervised annotation problem. Our simple unsupervised approach to annotating objects of interest in images achieves a higher annotation accuracy than most weakly supervised approaches.
Parthipan Siva, Chris Russell 0001, Tao Xiang 0002, Lourdes Agapito
CVPR2
2013 Inference Methods for CRFs with Co-occurrence Statistics
Lubor Ladicky, Chris Russell 0001, Pushmeet Kohli, Philip Torr 0001
Int. J. Comput. Vis.2
2012 In Defence of Negative Mining for Annotating Weakly Labelled Data
Parthipan Siva, Chris Russell 0001, Tao Xiang 0002
ECCV (3)2
2012 Dense multibody motion estimation and reconstruction from a handheld camera
abstract
Existing approaches to camera tracking and reconstruction from a single handheld camera for Augmented Reality (AR) focus on the reconstruction of static scenes. However, most real world scenarios are dynamic and contain multiple independently moving rigid objects. This paper addresses the problem of simultaneous segmentation, motion estimation and dense 3D reconstruction of dynamic scenes. We propose a dense solution to all three elements of this problem: depth estimation, motion label assignment and rigid transformation estimation directly from the raw video by optimizing a single cost function using a hill-climbing approach. We do not require prior knowledge of the number of objects present in the scene - the number of independent motion models and their parameters are automatically estimated. The resulting inference method combines the best techniques in discrete and continuous optimization: a state of the art variational approach is used to estimate the dense depth maps while the motion segmentation is achieved using discrete graph-cut based optimization. For the rigid motion estimation of the independently moving objects we propose a novel tracking approach designed to cope with the small fields of view they induce and agile motion. Our experimental results on real sequences show how accurate segmentations and dense depth maps can be obtained in a completely automated way and used in marker-free AR applications.
Anastasios Roussos, Chris Russell 0001, Ravi Garg, Lourdes Agapito
ISMAR2
2012 Joint Optimization for Object Class Segmentation and Dense Stereo Reconstruction
Lubor Ladicky, Paul Sturgess, Chris Russell 0001, Sunando Sengupta, Yalin Bastanlar, W. F. Clocksin, Philip Torr 0001
Int. J. Comput. Vis.3
2011 Efficient Second Order Multi-Target Tracking with Exclusion Constraints
Chris Russell 0001, Lourdes Agapito, Francesco Setti
BMVC1
2011 Energy based multiple model fitting for non-rigid structure from motion
abstract
In this paper we reformulate the 3D reconstruction of deformable surfaces from monocular video sequences as a labeling problem. We solve simultaneously for the assignment of feature points to multiple local deformation models and the fitting of models to points to minimize a geometric cost, subject to a spatial constraint that neighboring points should also belong to the same model. Piecewise reconstruction methods rely on features shared between models to enforce global consistency on the 3D surface. To account for this overlap between regions, we consider a super-set of the classic labeling problem in which a set of labels, instead of a single one, is assigned to each variable. We propose a mathematical formulation of this new model and show how it can be efficiently optimized with a variant of α-expansion. We demonstrate how this framework can be applied to Non-Rigid Structure from Motion and leads to simpler explanations of the same data. Compared to existing methods run on the same data, our approach has up to half the reconstruction error, and is more robust to over-fitting and outliers.
Chris Russell 0001, João Fayad, Lourdes Agapito
CVPR1
2011 Automated articulated structure and 3D shape recovery from point correspondences
abstract
In this paper we propose a new method for the simultaneous segmentation and 3D reconstruction of interest point based articulated motion. We decompose a set of point tracks into rigid-bodied overlapping regions which are associated with skeletal links, while joint centres can be derived from the regions of overlap. This allows us to formulate the problem of 3D reconstruction as one of model assignment, where each model corresponds to the motion and shape parameters of an articulated body part. We show how this labelling can be optimised using a combination of pre-existing graph-cut based inference, and robust structure from motion factorization techniques. The strength of our approach comes from viewing both the decomposition into parts, and the 3D reconstruction as the optimisation of a single cost function, namely the image re-projection error. We show results of full 3D shape recovery on challenging real-world sequences with one or more articulated bodies, in the presence of outliers and missing data.
João Fayad, Chris Russell 0001, Lourdes Agapito
ICCV2
2010 Joint Optimisation for Object Class Segmentation and Dense Stereo Reconstruction
abstract
This work is supported by EPSRC research grants, HMGCC, TUBITAK researcher exchange grant, the IST Programme of the European Community, under the PASCAL2 Network of Excellence, IST-2007-216886.
Lubor Ladicky, Paul Sturgess, Chris Russell 0001, Sunando Sengupta, Yalin Bastanlar, W. F. Clocksin, Philip Torr 0001
BMVC3
2010 Efficient piecewise learning for conditional random fields
abstract
Conditional Random Field models have proved effective for several low-level computer vision problems. Inference in these models involves solving a combinatorial optimization problem, with methods such as graph cuts, belief propagation. Although several methods have been proposed to learn the model parameters from training data, they suffer from various drawbacks. Learning these parameters involves computing the partition function, which is intractable. To overcome this, state-of-the-art structured learning methods frame the problem as one of large margin estimation. Iterative solutions have been proposed to solve the resulting convex optimization problem. Each iteration involves solving an inference problem over all the labels, which limits the efficiency of these structured methods. In this paper we present an efficient large margin piece-wise learning method which is widely applicable. We show how the resulting optimization problem can be reduced to an equivalent convex problem with a small number of constraints, and solve it using an efficient scheme. Our method is both memory and computationally efficient. We show results on publicly available standard datasets.
Karteek Alahari, Chris Russell 0001, Philip Torr 0001
CVPR2
2010 Graph Cut Based Inference with Co-occurrence Statistics
Lubor Ladicky, Chris Russell 0001, Pushmeet Kohli, Philip Torr 0001
ECCV (5)2
2010 What, Where and How Many? Combining Object Detectors and CRFs
Lubor Ladicky, Paul Sturgess, Karteek Alahari, Chris Russell 0001, Philip Torr 0001
ECCV (4)4
2010 Exact and Approximate Inference in Associative Hierarchical Networks using Graph Cuts
Chris Russell 0001, Lubor Ladicky, Pushmeet Kohli, Philip Torr 0001
UAI1
2009 Associative hierarchical CRFs for object class image segmentation
abstract
Most methods for object class segmentation are formulated as a labelling problem over a single choice of quantisation of an image space - pixels, segments or group of segments. It is well known that each quantisation has its fair share of pros and cons; and the existence of a common optimal quantisation level suitable for all object categories is highly unlikely. Motivated by this observation, we propose a hierarchical random field model, that allows integration of features computed at different levels of the quantisation hierarchy. MAP inference in this model can be performed efficiently using powerful graph cut based move making algorithms. Our framework generalises much of the previous work based on pixels or segments. We evaluate its efficiency on some of the most challenging data-sets for object class segmentation, and show it obtains state-of-the-art results.
Lubor Ladicky, Chris Russell 0001, Pushmeet Kohli, Philip Torr 0001
ICCV2
2007 Using the Pn Potts model with learning methods to segment live cell images
abstract
We present a segmentation method for live cell images, using graph cuts and learning methods. The images used here are particularly challenging because of the shared grey-level distributions of cells and background, which only differ by their textures, and the local imprecision around cell borders. We use the PnPotts model recently presented by Kohli et al. [9]: functions on higher-order cliques of pixels are included into the traditional Potts model, allowing us to account for local texture features, and to find the optimal solution efficiently. We use learning methods to define the potential functions used in the PnPotts model. We present the model and the learning methods we used, and compare our segmentation results with similar work in cytometry. While our method performs similarly, it requires little manual tuning and thus is straightforward to adapt to other images.
Chris Russell 0001, Dimitris N. Metaxas, Christophe Restif, Philip Torr 0001
ICCV1