Ian Davies

dblp:37/1217 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorComputer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 79% Probabilistic and Bayesian machine learning · 7% Kernel, tree and ensemble methods · 7%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
actor-critic methods
0.912025
Wasserstein Policy Optimization · ICML 2025
Machine learning › Reinforcement learning
continuous control
0.912025
Wasserstein Policy Optimization · ICML 2025
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.922020
Learning to Communicate Implicitly by Actions · AAAI 2020
Learning to Model Opponent Learning (Student Abstract) · AAAI 2020
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.912025
Wasserstein Policy Optimization · ICML 2025
Machine learning › Reinforcement learning
policy optimization
0.912025
Wasserstein Policy Optimization · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.512021
Kernel Identification Through Transformers · NeurIPS 2021
Machine learning › Kernel, tree and ensemble methods
kernel selection
0.512021
Kernel Identification Through Transformers · NeurIPS 2021
Machine learning › Deep learning architectures and training
transformer
0.512021
Kernel Identification Through Transformers · NeurIPS 2021
Machine learning › Reinforcement learning › multi-agent reinforcement learning › multi-agent communication
implicit communication
0.412020
Learning to Communicate Implicitly by Actions · AAAI 2020
Machine learning › Reinforcement learning › multi-agent reinforcement learning
non-stationarity
0.412020
Learning to Model Opponent Learning (Student Abstract) · AAAI 2020
Machine learning › Reinforcement learning › multi-agent reinforcement learning
opponent modeling
0.412020
Learning to Model Opponent Learning (Student Abstract) · AAAI 2020
Machine learning › Reinforcement learning › reward design › reward shaping
auxiliary reward
0.112020
Learning to Communicate Implicitly by Actions · AAAI 2020

Methods — techniques the papers use, named apart from their topics

wasserstein gradient flow · 0.9actor-critic · 0.9transformer · 0.5self-attention · 0.5policy search · 0.4policy belief learning · 0.4opponent modeling · 0.4belief modeling · 0.4behavior cloning · 0.4auxiliary reward · 0.4
YearPublicationVenuePosition
2025 Wasserstein Policy Optimization
abstract
We introduce Wasserstein Policy Optimization (WPO), an actor-critic algorithm for reinforcement learning in continuous action spaces. WPO can be derived as an approximation to Wasserstein gradient flow over the space of all policies projected into a finite-dimensional parameter space (e.g., the weights of a neural network), leading to a simple and completely general closed-form update. The resulting algorithm combines many properties of deterministic and classic policy gradient methods. Like deterministic policy gradients, it exploits knowledge of the *gradient* of the action-value function with respect to the action. Like classic policy gradients, it can be applied to stochastic policies with arbitrary distributions over actions -- without using the reparameterization trick. We show results on the DeepMind Control Suite and a magnetic confinement fusion task which compare favorably with state-of-the-art continuous control methods.
David Pfau, Ian Davies, Diana Borsa, João G. M. Araújo, Brendan D. Tracey, Hado van Hasselt
ICML2
2021 Kernel Identification Through Transformers
abstract
Kernel selection plays a central role in determining the performance of Gaussian Process (GP) models, as the chosen kernel determines both the inductive biases and prior support of functions under the GP prior. This work addresses the challenge of constructing custom kernel functions for high-dimensional GP regression models. Drawing inspiration from recent progress in deep learning, we introduce a novel approach named KITT: Kernel Identification Through Transformers. KITT exploits a transformer-based architecture to generate kernel recommendations in under 0.1 seconds, which is several orders of magnitude faster than conventional kernel search algorithms. We train our model using synthetic data generated from priors over a vocabulary of known kernels. By exploiting the nature of the self-attention mechanism, KITT is able to process datasets with inputs of arbitrary dimension. We demonstrate that kernels chosen by KITT yield strong performance over a diverse collection of regression benchmarks.
Fergus Simpson, Ian Davies, Vidhi Lalchand, Alessandro Vullo, Nicolas Durrande, Carl E. Rasmussen
NeurIPS2
2020 Learning to Model Opponent Learning (Student Abstract)
abstract
Multi-Agent Reinforcement Learning (MARL) considers settings in which a set of coexisting agents interact with one another and their environment. The adaptation and learning of other agents induces non-stationarity in the environment dynamics. This poses a great challenge for value function-based algorithms whose convergence usually relies on the assumption of a stationary environment. Policy search algorithms also struggle in multi-agent settings as the partial observability resulting from an opponent's actions not being known introduces high variance to policy training. Modelling an agent's opponent(s) is often pursued as a means of resolving the issues arising from the coexistence of learning opponents. An opponent model provides an agent with some ability to reason about other agents to aid its own decision making. Most prior works learn an opponent model by assuming the opponent is employing a stationary policy or switching between a set of stationary policies. Such an approach can reduce the variance of training signals for policy search algorithms. However, in the multi-agent setting, agents have an incentive to continually adapt and learn. This means that the assumptions concerning opponent stationarity are unrealistic. In this work, we develop a novel approach to modelling an opponent's learning dynamics which we term Learning to Model Opponent Learning (LeMOL). We show our structured opponent model is more accurate and stable than naive behaviour cloning baselines. We further show that opponent modelling can improve the performance of algorithmic agents in multi-agent settings.
Ian Davies, Zheng Tian 0002, Jun Wang 0012
AAAI1
2020 Learning to Communicate Implicitly by Actions
abstract
In situations where explicit communication is limited, human collaborators act by learning to: (i) infer meaning behind their partner's actions, and (ii) convey private information about the state to their partner implicitly through actions. The first component of this learning process has been well-studied in multi-agent systems, whereas the second — which is equally crucial for successful collaboration — has not. To mimic both components mentioned above, thereby completing the learning process, we introduce a novel algorithm: Policy Belief Learning (PBL). PBL uses a belief module to model the other agent's private information and a policy module to form a distribution over actions informed by the belief module. Furthermore, to encourage communication by actions, we propose a novel auxiliary reward which incentivizes one agent to help its partner to make correct inferences about its private information. The auxiliary reward for communication is integrated into the learning of the policy module. We evaluate our approach on a set of environments including a matrix game, particle environment and the non-competitive bidding problem from contract bridge. We show empirically that this auxiliary reward is effective and easy to generalize. These results demonstrate that our PBL algorithm can produce strong pairs of agents in collaborative games where explicit communication is disabled.
Zheng Tian 0002, Shihao Zou, Ian Davies, Tim Warr, Lisheng Wu, Haitham Bou-Ammar, Jun Wang 0012
AAAI3
2019 The ASC-Inclusion Perceptual Serious Gaming Platform for Autistic Children
abstract
“Serious games” are becoming extremely relevant to individuals who have specific needs, such as children with an autism spectrum condition (ASC). Often, individuals with an ASC have difficulties in interpreting verbal and nonverbal communication cues during social interactions. The ASC-Inclusion EU-FP7 funded project aims to provide children who have an ASC with a platform to learn emotion expression and recognition, through play in the virtual world. In particular, the ASC-Inclusion platform focuses on the expression of emotion via facial, vocal, and bodily gestures. The platform combines multiple analysis tools, using onboard microphone and webcam capabilities. The platform utilizes these capabilities via training games, text-based communication, animations, video, and audio clips. This paper introduces current findings and evaluations of the ASC-Inclusion platform and provides detailed description for the different modalities.
Erik Marchi, Tadas Baltrusaitis, Andra Adams, Marwa Mahmoud, Ofer Golan, Shimrit Fridenson-Hayo, Shahar Tal, Shai Newman, Noga Meir-Goren, Antonio Camurri, Stefano Piana, Björn W. Schuller, Sven Bölte, Tevfik Metin Sezgin, Nese Alyüz, Agnieszka Rynkiewicz, Aurelie Baranger, Alice Baird, Simon Baron-Cohen, Amandine Lassalle, Helen O'Reilly, Delia Pigat, Peter Robinson 0001, Ian Davies
IEEE Trans. Games24
2016 Behavior Identification in Two-Stage Games for Incentivizing Citizen Science Exploration
Yexiang Xue, Ian Davies, Daniel Fink 0002, Christopher Wood, Carla P. Gomes
CP2
2016 Supporting Scalable Data Sharing in Online Education
abstract
Online educational tools often generate learning data, and sharing such data between tutors and students can often improve learning outcomes. Unfortunately the process of sharing learning data today is not always transparent to students. Our aim is to improve the transparency and user control aspects of sharing data whilst maintaining the educational utility of data sharing between tutors and students. To do so, we start by surveying the possible methods of sharing data, and we use this to design a token-based scheme for facilitating data sharing. We implemented our scheme and observed it in use by 7,798 students over the course of one year. We find that our proposed scheme provides a good balance between transparency, user control, educational utility and scalability.
Stephen Cummins, Alastair R. Beresford, Ian Davies, Andrew C. Rice
L@S3
2016 Investigating the Use of Hints in Online Problem Solving
abstract
We investigate the use of hints as a form of scaffolding for 4,652 eligible users on a large-scale online learning environment called Isaac, which allows users to answer physics questions with up to five hints. We investigate user behaviour when using hints, users' engagement with fading (the process of gradually becoming less reliant on the hints provided), and hint strategies including Decomposition, Correction, Verification, or Comparison. Finally, we present recommendations for the design and development of online teaching tools that provide open access to hints, including a mechanism that may improve the speed at which users begin fading.
Stephen Cummins, Alistair Stead, Lisa Jardine-Wright, Ian Davies, Alastair R. Beresford, Andrew C. Rice
L@S4
2015 Equality: A Tool for Free-form Equation Editing
abstract
We describe a new tool, Equality, for equation entry using free-form layout of components drawn from a palette of symbols. Our approach is designed to enable learners to easily manipulate the structure of their equations, to be functional in both desktop and mobile environments, and to minimize the amount of learning required to use the tool. We present the results of a study comparing a prototype of our approach with Microsoft Equation Editor using a desktop machine. The initial results are promising with participants reporting that the mechanism is easy to learn and an easy way to manipulate their equations. We report the results of the study and the views of the participants and identify how these will inform the future development of Equality.
Stephen Cummins, Ian Davies, Andrew C. Rice, Alastair R. Beresford
ICALT2
2011 Emotional Investment in Naturalistic Data Collection
Ian Davies, Peter Robinson 0001
ACII (1)1
2009 Multimodal inference for driver-vehicle interaction
abstract
In this paper we present a novel system for driver-vehicle interaction which combines speech recognition with facial-expression recognition to increase intention recognition accuracy in the presence of engine- and road-noise. Our system would allow drivers to interact with in-car devices such as satellite navigation and other telematic or control systems. We describe a pilot study and experiment in which we tested the system, and show that multimodal fusion of speech and facial expression recognition provides higher accuracy than either would do alone.
Tevfik Metin Sezgin, Ian Davies, Peter Robinson 0001
ICMI2
1988 ISPBXs and terminals
Ian Davies, Alistair McBain
Comput. Commun.1