Shaobo Hou

dblp:78/6677 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 61% Trustworthy machine learning · 21% Generative modeling · 10%
Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 100%

Topics — the 24 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › computational creativity
creative generation
0.912025
Generating Creative Chess Puzzles · NeurIPS 2025
Machine learning › Reinforcement learning
reward design
0.912025
Generating Creative Chess Puzzles · NeurIPS 2025
Machine learning › Reinforcement learning
diversity optimization
0.712023
Discovering Policies with DOMiNO: Diversity Optimization Maintaining Near Optimality · ICLR 2023
Machine learning › Reinforcement learning › reinforcement learning theory
near-optimal policy identification
0.712023
Discovering Policies with DOMiNO: Diversity Optimization Maintaining Near Optimality · ICLR 2023
Machine learning › Reinforcement learning
policy diversity
0.712023
Discovering Policies with DOMiNO: Diversity Optimization Maintaining Near Optimality · ICLR 2023
Machine learning › Trustworthy machine learning
robustness
0.612022
Underspecification Presents Challenges for Credibility in Modern Machine Learning · J. Mach. Learn. Res. 2022
Machine learning › Trustworthy machine learning › robustness
underspecification
0.612022
Underspecification Presents Challenges for Credibility in Modern Machine Learning · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning
robust reinforcement learning
0.512021
Discovering a set of policies for the worst case reward · ICLR 2021
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.412019
The Option Keyboard: Combining Skills in Reinforcement Learning · NeurIPS 2019
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery
0.412019
The Option Keyboard: Combining Skills in Reinforcement Learning · NeurIPS 2019
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill composition
0.412019
The Option Keyboard: Combining Skills in Reinforcement Learning · NeurIPS 2019
Games and playful interaction › board games
chess
0.312025
Generating Creative Chess Puzzles · NeurIPS 2025
Natural language and speech › Speech recognition and synthesis
speech-driven animation
0.212013
Visual Speech Synthesis Using a Variable-Order Switching Shared Gaussian Process Dynamical Model · IEEE Trans. Multim. 2013
Computer animation and physical simulation › audio-driven animation
speech-driven animation
0.212013
Visual Speech Synthesis Using a Variable-Order Switching Shared Gaussian Process Dynamical Model · IEEE Trans. Multim. 2013
Computer animation and physical simulation › facial animation
speech-driven facial animation
0.212013
Visual Speech Synthesis Using a Variable-Order Switching Shared Gaussian Process Dynamical Model · IEEE Trans. Multim. 2013
Machine learning › Reinforcement learning
transfer learning in reinforcement learning
0.112019
The Option Keyboard: Combining Skills in Reinforcement Learning · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
mixture model learning
0.112008
Robust estimation of gaussian mixtures from noisy input data · CVPR 2008
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.112008
Robust estimation of gaussian mixtures from noisy input data · CVPR 2008
Computer vision › 3D vision › motion capture
articulated body tracking
0.112007
Real-time Body Tracking Using a Gaussian Process Latent Variable Model · ICCV 2007
Computer vision › Face, body and person analysis › human pose estimation
human pose tracking
0.112007
Real-time Body Tracking Using a Gaussian Process Latent Variable Model · ICCV 2007
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.112007
Real-time Body Tracking Using a Gaussian Process Latent Variable Model · ICCV 2007
Robotics › Motion planning and robot control
robot learning
0.112007
Real-time Body Tracking Using a Gaussian Process Latent Variable Model · ICCV 2007
Data mining
clustering
0.012008
Robust estimation of gaussian mixtures from noisy input data · CVPR 2008
Data mining › clustering › robust clustering
noisy clustering
0.012008
Robust estimation of gaussian mixtures from noisy input data · CVPR 2008

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.7chess engine search statistics · 1.7diversity optimization · 0.7constrained optimization · 0.7policy optimization · 0.5pseudo-rewards · 0.4options framework · 0.4cumulants · 0.4forced phonetic alignment · 0.3variable length markov model · 0.2variable-length markov model · 0.2switching state space model · 0.2shared gaussian process dynamical model · 0.2variational bayes · 0.1uncertainty modeling · 0.1
YearPublicationVenuePosition
2025 Generating Creative Chess Puzzles
abstract
While Generative AI rapidly advances in various domains, generating truly creative, aesthetic, and counter-intuitive outputs remains a challenge. This paper presents an approach to tackle these difficulties in the domain of chess puzzles. We start by benchmarking Generative AI architectures, and then introduce an RL framework with novel rewards based on chess engine search statistics to overcome some of those shortcomings. The rewards are designed to enhance a puzzle's uniqueness, counter-intuitiveness, diversity, and realism. Our RL approach dramatically increases counter-intuitive puzzle generation by 10x, from 0.22\% (supervised) to 2.5\%, surpassing existing dataset rates (2.1\%) and the best Lichess-trained model (0.4\%). Our puzzles meet novelty and diversity benchmarks, retain aesthetic themes, and are rated by human experts as more creative, enjoyable, and counter-intuitive than composed book puzzles, even approaching classic compositions. Our final outcome is a curated booklet of these novel AI-generated puzzles, which is acknowledged for creativity by three world-renowned experts.
Xidong Feng, Vivek Veeriah, Marcus Chiam, Michael Dennis 0001, Federico Barbero, Johan S. Obando-Ceron, Jiaxin Shi, Satinder Singh 0001, Shaobo Hou, Nenad Tomasev, Tom Zahavy
NeurIPS9
2023 Discovering Policies with DOMiNO: Diversity Optimization Maintaining Near Optimality
Tom Zahavy, Yannick Schroecker, Feryal M. P. Behbahani, Kate Baumli, Sebastian Flennerhag, Shaobo Hou, Satinder Singh 0001
ICLR6
2022 Underspecification Presents Challenges for Credibility in Modern Machine Learning
abstract
Machine learning (ML) systems often exhibit unexpectedly poor behavior when they are deployed in real-world domains. We identify underspecification in ML pipelines as a key reason for these failures. An ML pipeline is the full procedure followed to train and validate a predictor. Such a pipeline is underspecified when it can return many distinct predictors with equivalently strong test performance. Underspecification is common in modern ML pipelines that primarily validate predictors on held-out data that follow the same distribution as the training data. Predictors returned by underspecified pipelines are often treated as equivalent based on their training domain performance, but we show here that such predictors can behave very differently in deployment domains. This ambiguity can lead to instability and poor model behavior in practice, and is a distinct failure mode from previously identified issues arising from structural mismatch between training and deployment domains. We provide evidence that underspecfication has substantive implications for practical ML pipelines, using examples from computer vision, medical imaging, natural language processing, clinical risk prediction based on electronic health records, and medical genomics. Our results show the need to explicitly account for underspecification in modeling pipelines that are intended for real-world deployment in any domain.
Alexander D'Amour, Katherine A. Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew Hoffman 0001, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yi-An Ma, Cory Y. McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman 0003, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin G. Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang 0002, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai, D. Sculley
J. Mach. Learn. Res.13
2021 Discovering a set of policies for the worst case reward
Tom Zahavy, André Barreto 0001, Daniel J. Mankowitz, Shaobo Hou, Brendan O'Donoghue, Iurii Kemaev, Satinder Singh 0001
ICLR4
2019 The Option Keyboard: Combining Skills in Reinforcement Learning
abstract
The ability to combine known skills to create new ones may be crucial in the solution of complex reinforcement learning problems that unfold over extended periods. We argue that a robust way of combining skills is to define and manipulate them in the space of pseudo-rewards (or "cumulants"). Based on this premise, we propose a framework for combining skills using the formalism of options. We show that every deterministic option can be unambiguously represented as a cumulant defined in an extended domain. Building on this insight and on previous results on transfer learning, we show how to approximate options whose cumulants are linear combinations of the cumulants of known options. This means that, once we have learned options associated with a set of cumulants, we can instantaneously synthesise options induced by any linear combination of them, without any learning involved. We describe how this framework provides a hierarchical interface to the environment whose abstract actions correspond to combinations of basic skills. We demonstrate the practical benefits of our approach in a resource management problem and a navigation task involving a quadrupedal simulated robot.
André Barreto 0001, Diana Borsa, Shaobo Hou, Gheorghe Comanici, Eser Aygün, Philippe Hamel, Daniel Toyama, Jonathan J. Hunt, Shibl Mourad, David Silver 0001, Doina Precup
NeurIPS3
2013 Visual Speech Synthesis Using a Variable-Order Switching Shared Gaussian Process Dynamical Model
abstract
In this paper, we present a novel approach to speech- driven facial animation using a non-parametric switching state space model based on Gaussian processes. The model is an extension of the shared Gaussian process dynamical model, augmented with switching states. Two talking head corpora are processed by extracting visual and audio data from the sequences followed by a parameterization of both data streams. Phonetic labels are obtained by performing forced phonetic alignment on the audio. The switching states are found using a variable length Markov model trained on the labelled phonetic data. The audio and visual data corresponding to phonemes matching each switching state are extracted and modelled together using a shared Gaussian process dynamical model. We propose a synthesis method that takes into account both previous and future phonetic context, thus accounting for forward and backward coarticulation in speech. Both objective and subjective evaluation results are presented. The quantitative results demonstrate that the proposed method outperforms other state-of-the-art methods in visual speech synthesis and the qualitative results reveal that the synthetic videos are comparable to ground truth in terms of visual perception and intelligibility.
Salil Deena, Shaobo Hou, Aphrodite Galata
IEEE Trans. Multim.2
2008 Robust estimation of gaussian mixtures from noisy input data
abstract
We propose a variational bayes approach to the problem of robust estimation of gaussian mixtures from noisy input data. The proposed algorithm explicitly takes into account the uncertainty associated with each data point, makes no assumptions about the structure of the covariance matrices and is able to automatically determine the number of the gaussian mixture components. Through the use of both synthetic and real world data examples, we show that by incorporating uncertainty information into the clustering algorithm, we get better results at recovering the true distribution of the training data compared to other variational bayesian clustering algorithms.
Shaobo Hou, Aphrodite Galata
CVPR1
2007 Real-time Body Tracking Using a Gaussian Process Latent Variable Model
abstract
In this paper, we present a tracking framework for capturing articulated human motions in real-time, without the need for attaching markers onto the subject's body. This is achieved by first obtaining a low dimensional representation of the training motion data, using a nonlinear dimensionality reduction technique called back-constrained GPLVM. A prior dynamics model is then learnt from this low dimensional representation by partitioning the motion sequences into elementary movements using an unsupervised EM clustering algorithm. The temporal dependencies between these elementary movements are efficiently captured by a Variable Length Markov Model. The learnt dynamics model is used to bias the propagation of candidate pose feature vectors in the low dimensional space. By combining this with an efficient volumetric reconstruction algorithm, our framework can quickly evaluate each candidate pose against image evidence captured from multiple views. We present results that show our system can accurately track complex structured activities such as ballet dancing in real-time.
Shaobo Hou, Aphrodite Galata, Fabrice Caillette, Neil A. Thacker, Paul A. Bromiley
ICCV1