Wenzhe Shi

dblp:46/8487 · DBLP profile ↗
← Back
35ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-9750-6379ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Pixels to Play: A Foundation Model for 3D Gameplay
abstract
We introduce Pixels2Play-0.1 (P2P0.1), a foundation model that learns to play a wide range of 3 D video games with recognizable human-like behavior. Motivated by emerging consumer and developer use cases-AI teammates, controllable NPCs, personalized live-streamers, assistive testers-we argue that an agent must rely on the same pixel stream available to players and generalize to new titles with minimal gamespecific engineering. P2P0.1 is trained end-to-end with behavior cloning: labeled demonstrations collected from instrumented human game-play are complemented by unlabeled public videos, to which we impute actions via an inverse-dynamics model. A decoder-only transformer with auto-regressive action output handles the large action space while remaining latency-friendly on a single consumer GPU. We report qualitative results showing competent play across simple Roblox and classic MS-DOS titles, ablations on unlabeled data, and outline the scaling and evaluation steps required to reach expert-level, text-conditioned control.
Yuguang Yue, Chris Green, Samuel Hunt, Irakli Salia, Wenzhe Shi, Jonathan J. Hunt
CoG5
2025 Two grids are better than one: Hybrid indoor scene reconstruction framework with adaptive priors
abstract
Indoor scene reconstruction from multi-view images is a pivotal technology within the field of robotics and augmented reality . Previous researches have predominantly focused on neural radiance fields aided by geometric monocular priors. However, due to the inductive smoothness bias introduced by deep Multi-Layer Perceptron (MLP) networks, these methods struggle to recover the scene surface with complex and fine geometry details. Additionally, when used as additional supervision signals during optimization, priors in different regions make different contributions. Simply incorporating them in all regions may lead to a decrease in the accuracy. To tackle these issues, we present a generic end-to-end framework named AdaptSurf, which combines Signed Distance Field (SDF) voxel grids and feature voxel grids to enhance the capability of reconstructing accurate geometry details, respectively. Furthermore, we design a policy network to adaptively enable the estimated depth or normal priors to supervise the learning process, which improves the reconstruction accuracy and accelerates neural surface reconstruction. Qualitative and quantitative experiments show that AdaptSurf yields high-quality surfaces, especially for fine-grained details and smooth regions. Furthermore, the policy network exhibits an interpretable behavior that depends on the voxel features, which helps to improve the quality of surface reconstruction.
Boyuan Bai, Xiuquan Qiao, Hongru Zhao, Wenzhe Shi, Hengjia Zhang, Yakun Huang
Neurocomputing5
2025 FRPGS: Fast, Robust, and Photorealistic Monocular Dynamic Scene Reconstruction With Deformable 3D Gaussians
abstract
Dynamic reconstruction technology presents significant promise for applications in visual and interactive fields. Current techniques utilizing 3D Gaussian Splatting show favorable results and fast reconstruction speed. However, as scene expanding, using individual Gaussian structure (i) leads to instability in large-scale dynamic reconstruction, marked by abrupt deformation, and (ii) the heuristic densification of individuals suffers significant redundancy. Tackling these issues, we propose a jointed Gaussian representation method named FRPGS, which learns the global information and the deformation using center Gaussians and generates the neural Gaussians around them for local detail. Specifically, FRPGS employs center Gaussians initialized from point clouds, which are learned with a deformation field for representing global relationships and dynamic motion over time. Then, for each center Gaussian, attribute networks generate neural Gaussians that move under the linked center Gaussian driving, thereby ensuring structural integrity during movement within this joint-based representation. Finally, to reduce Gaussian redundancy, a densification strategy is developed based on the average cumulative gradient of the associated neural Gaussians, imposing strict limits on the growing of center Gaussians without compromising accuracy. Additionally, we established a large-scale dynamic indoor dataset at the MuLong Laboratory of ZTE Corporation. Evaluations demonstrate that FRPGS significantly outperforms state-of-the-art methods in both training efficiency and reconstruction quality, achieving over a 50% (up to 74%) improvement in efficiency on an RTX 4090. FRPGS also supports the 4K resolution reconstruction of 60 frames simultaneously.
Xiao Pan 0001, Daquan Feng, Wenzhe Shi
IEEE Trans. Circuits Syst. Video Technol.6
2025 Orhlr-net: one-stage residual learning network for joint single-image specular highlight detection and removal
Wenzhe Shi, Ziqi Hu, Hengjia Zhang, Jiale Yang
Vis. Comput.1
2024 Multi-Objective Recommendation via Multivariate Policy Learning
abstract
Real-world recommender systems often need to balance multiple objectives when deciding which recommendations to present to users. These include behavioural signals (e.g. clicks, shares, dwell time), as well as broader objectives (e.g. diversity, fairness). Scalarisation methods are commonly used to handle this balancing task, where a weighted average of per-objective reward signals determines the final score used for ranking. Naturally, how these weights are computed exactly, is key to success for any online platform.
Olivier Jeunen, Jatin Mandav, Ivan Potapov, Nakul Agarwal, Sourabh Vaid, Wenzhe Shi, Aleksei Ustimenko
RecSys6
2023 Investigating the effects of incremental training on neural ranking models
abstract
Recommender systems are an essential component of online platforms providing users with personalized experiences. Some recommendation scenarios such as social networks and news are extremely dynamic in nature with user interests changing over time and new items being continuously added due to breaking news and trending events.
Benedikt Schifferer, Wenzhe Shi, Gabriel de Souza Pereira Moreira, Even Oldridge, Chris Deotte, Gilberto Titericz, Kazuki Onodera, Praveen Dhinwa, Vishal Agrawal, Chris Green
RecSys2
2021 RecSys 2021 Challenge Workshop: Fairness-aware engagement prediction at scale on Twitter's Home Timeline
abstract
The workshop features presentations of accepted contributions to the RecSys Challenge 2021, organized by Politecnico di Bari, ETH Zürich, Jönköping University, and the data set is provided by Twitter. The challenge focuses on a real-world task of tweet engagement prediction in a dynamic environment. For 2021, the challenge considers four different engagement types: Likes, Retweet, Quote, and replies. This year’s challenge brings the problem even closer to Twitter’s real recommender systems by introducing latency constraints. We also increases the data size to encourage novel methods. Also, the data density is increased in terms of the graph where users are considered to be nodes and interactions as edges. The goal is twofold: to predict the probability of different engagement types of a target user for a set of Tweets based on heterogeneous input data while providing fair recommendations. In fact, multi-goal optimization considering accuracy and fairness is particularly challenging. However, we believed that the recommendation community was nowadays mature enough to face the challenge of providing accurate and, at the same time, fair recommendations. To this end, Twitter has released a public dataset of close to 1 billion data points, > 40 million each day over 28 days. Week 1 − 3 will be used for training and week 4 for evaluation and testing. Each datapoint contains the tweet along with engagement features, user features, and tweet features. A peculiarity of this challenge is related to keeping the dataset updated with the platform: if a user deletes a Tweet, or their data from Twitter, the dataset is promptly updated. Moreover, each change in the dataset implied new evaluations of all submissions and the update of the leaderboard metrics. The challenge was well received with 578 registered users, and 386 submissions.
Vito Walter Anelli, Saikishore Kalloori, Bruce Ferwerda, Luca Belli, Alykhan Tejani, Frank Portman, Alexandre Lung-Yut-Fong, Benjamin Paul Chamberlain, Yuanpu Xie, Jonathan J. Hunt, Michael M. Bronstein, Wenzhe Shi
RecSys12
2021 3-D FEM Azimuth Forward Modeling of Hydraulic Fractures Based on Electromagnetic Theory
abstract
Accurate monitoring of hydraulic fracturing fracture morphology is of great significance to the follow-up exploration and development work. Electromagnetic monitoring of fractures can effectively identify the effective propped volume (EPV) of fracturing fractures, making up for the disadvantages of other monitoring methods. This letter introduces a modeling method and a physical model based on very-low-frequency (VLF) electromagnetic scattering theory for monitoring the development of asymmetric fracturing fractures. The 3-D finite-element method (FEM) of equivalent fracturing fracture on transition boundary condition (TBC) surface is used to realize the fast forward modeling of 3-D space receiving response in large formations. A sector-shaped receiver is designed and the relationship between electromagnetic receiving signals and orientation parameters of asymmetric fracturing fractures, such as direction of fracture development and tilt angle, is discussed. By analyzing 3-D signals obtained from the rotating receiver sector, the spatial state of multiple asymmetric fracturing fractures can be determined. The problem of how to identify the 3-D direction of crack growth with electromagnetic monitoring method is solved, which provides a theoretical reference for the development and inversion of detection instruments.
Yang Li 0095, Dejun Liu, Ying Zhai, Wenzhe Shi, Yixuan Xian
IEEE Geosci. Remote. Sens. Lett.4
2020 RecSys 2020 Challenge Workshop: Engagement Prediction on Twitter's Home Timeline
abstract
The workshop features presentations of accepted contributions to the RecSys Challenge 2020, organized by Politecnico di Bari, Free University of Bozen-Bolzano, TU Wien, University of Colorado, Boulder, and Universidade Federal de Campina Grande, and sponsored by Twitter. The challenge focuses on a real-world task of Tweet engagement prediction in a dynamic environment. The goal is to predict the probability for different types of engagement (Like, Reply, Retweet, and Retweet with comment) of a target user for a set of Tweets, based on heterogeneous input data. To this end, Twitter has released a large public dataset of ~160M public Tweets, obtained by subsampling within ~2 weeks, that contains engagement features, user features, and Tweet features. A peculiarity of this challenge is related to the recent regulations on data protection and privacy. The challenge data set was compliant: if a user deleted a Tweet, or their data from Twitter, the dataset was promptly updated. Moreover, each change in the dataset implied new evaluations of all submissions and the update of the leaderboard metrics.
Vito Walter Anelli, Amra Delic, Gabriele Sottocornola, Jessie Smith, Nazareno Andrade, Luca Belli, Michael M. Bronstein, Sofia Ira Ktena, Alexandre Lung-Yut-Fong, Frank Portman, Alykhan Tejani, Yuanpu Xie, Wenzhe Shi
RecSys15
2020 Deep Bayesian Bandits: Exploring in Online Personalized Recommendations
abstract
Recommender systems trained in a continuous learning fashion are plagued by the feedback loop problem, also known as algorithmic bias. This causes a newly trained model to act greedily and favor items that have already been engaged by users. This behavior is particularly harmful in personalised ads recommendations, as it can also cause new campaigns to remain unexplored. Exploration aims to address this limitation by providing new information about the environment, which encompasses user preference, and can lead to higher long-term reward. In this work, we formulate a display advertising recommender as a contextual bandit and implement exploration techniques that require sampling from the posterior distribution of click-through-rates in a computationally tractable manner. Traditional large-scale deep learning models do not provide uncertainty estimates by default. We approximate these uncertainty measurements of the predictions by employing a bootstrapped model with multiple heads and dropout units. We benchmark a number of different models in an offline simulation environment using a publicly available dataset of user-ads engagements. We test our proposed deep Bayesian bandits algorithm in the offline simulation and online AB setting with large-scale production traffic, where we demonstrate a positive gain of our exploration model.
Dalin Guo, Sofia Ira Ktena, Pranay Kumar Myana, Ferenc Huszar, Wenzhe Shi, Alykhan Tejani, Michael Kneier
RecSys5
2020 Model Size Reduction Using Frequency Based Double Hashing for Recommender Systems
abstract
Deep Neural Networks (DNNs) with sparse input features have been widely used in recommender systems in industry. These models have large memory requirements and need a huge amount of training data. The large model size usually entails a cost, in the range of millions of dollars, for storage and communication with the inference services. In this paper, we propose a hybrid hashing method to combine frequency hashing and double hashing techniques for model size reduction, without compromising performance. We evaluate the proposed models on two product surfaces. In both cases, experiment results demonstrated that we can reduce the model size by around 90 while keeping the performance on par with the original baselines.
Caojin Zhang, Yicun Liu, Yuanpu Xie, Sofia Ira Ktena, Alykhan Tejani, Pranay Kumar Myana, Deepak Dilipkumar, Suvadip Paul, Ikuhiro Ihara, Prasang Upadhyaya, Ferenc Huszar, Wenzhe Shi
RecSys13
2019 Channel Estimation and Equalization Based on Deep BLSTM for FBMC-OQAM Systems
abstract
Channel estimation and equalization is one of the challenges of the filter bank multicarrier (FBMC) systems because of the existence of intrinsic imaginary interference. In this paper, we try to solve this challenge from a learning-based perspective. Based on the study of the intrinsic relationship between bidirectional long short-term memory (BLSTM) recurrent neural networks (RNN) and the FBMC radio signals, we propose a novel deep BLSTM network based channel estimation and equalization scheme, abbreviated as BLSTM-CE scheme. In the BLSTM-CE scheme, the input and output of some traditional independent modules are seen as an unknown nonlinear mapping and use a deep BLSTM network to approximate it. Numerical simulation shows that our proposed BLSTM-CE scheme can obtain perfect channel estimation performance in the FBMC systems.
Dejun Liu, Wenzhe Shi, Yang Zhao 0012
ICC4
2019 Addressing delayed feedback for continuous training with neural networks in CTR prediction
abstract
One of the challenges in display advertising is that the distribution of features and click through rate (CTR) can exhibit large shifts over time due to seasonality, changes to ad campaigns and other factors. The predominant strategy to keep up with these shifts is to train predictive models continuously, on fresh data, in order to prevent them from becoming stale. However, in many ad systems positive labels are only observed after a possibly long and random delay. These delayed labels pose a challenge to data freshness in continuous training: fresh data may not have complete label information at the time they are ingested by the training algorithm. Naive strategies which consider any data point a negative example until a positive label becomes available tend to underestimate CTR, resulting in inferior user experience and suboptimal performance for advertisers. The focus of this paper is to identify the best combination of loss functions and models that enable large-scale learning from a continuous stream of data in the presence of delayed labels. In this work, we compare 5 different loss functions, 3 of them applied to this problem for the first time. We benchmark their performance in offline settings on both public and proprietary datasets in conjunction with shallow and deep model architectures. We also discuss the engineering cost associated with implementing each loss function in a production environment. Finally, we carried out online experiments with the top performing methods, in order to validate their performance in a continuous training scheme. While training on 668 million in-house data points offline, our proposed methods outperform previous state-of-the-art by 3% relative cross entropy (RCE). During online experiments, we observed 55% gain in revenue per thousand requests (RPMq) against naive log loss.
Sofia Ira Ktena, Alykhan Tejani, Lucas Theis, Pranay Kumar Myana, Deepak Dilipkumar, Ferenc Huszar, Steven Yoo, Wenzhe Shi
RecSys8
2018 Three-dimensional cardiovascular imaging-genetics: a mass univariate framework
abstract
Motivation: Left ventricular (LV) hypertrophy is a strong predictor of cardiovascular outcomes, but its genetic regulation remains largely unexplained. Conventional phenotyping relies on manual calculation of LV mass and wall thickness, but advanced cardiac image analysis presents an opportunity for high-throughput mapping of genotype-phenotype associations in three dimensions (3D). Results: High-resolution cardiac magnetic resonance images were automatically segmented in 1124 healthy volunteers to create a 3D shape model of the heart. Mass univariate regression was used to plot a 3D effect-size map for the association between wall thickness and a set of predictors at each vertex in the mesh. The vertices where a significant effect exists were determined by applying threshold-free cluster enhancement to boost areas of signal with spatial contiguity. Experiments on simulated phenotypic signals and SNP replication show that this approach offers a substantial gain in statistical power for cardiac genotype-phenotype associations while providing good control of the false discovery rate. This framework models the effects of genetic variation throughout the heart and can be automatically applied to large population cohorts. Availability and implementation: The proposed approach has been coded in an R package freely available at https://doi.org/10.5281/zenodo.834610 together with the clinical data used in this work. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Carlo Biffi, Antonio M. Simoes Monteiro de Marvao, Mark I Attard, Timothy Dawes, Nicola Whiffin, Wenjia Bai, Wenzhe Shi, Catherine Francis, Hannah Meyer, Rachel J. Buchan, Stuart A. Cook, Daniel Rueckert, Declan P. O'Regan
Bioinform.7
2017 Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation
abstract
Convolutional neural networks have enabled accurate image super-resolution in real-time. However, recent attempts to benefit from temporal correlations in video super-resolution have been limited to naive or inefficient architectures. In this paper, we introduce spatio-temporal sub-pixel convolution networks that effectively exploit temporal redundancies and improve reconstruction accuracy while maintaining real-time speed. Specifically, we discuss the use of early fusion, slow fusion and 3D convolutions for the joint processing of multiple consecutive video frames. We also propose a novel joint motion compensation and video super-resolution algorithm that is orders of magnitude more efficient than competing methods, relying on a fast multi-resolution spatial transformer module that is end-to-end trainable. These contributions provide both higher accuracy and temporally more consistent videos, which we confirm qualitatively and quantitatively. Relative to single-frame models, spatio-temporal networks can either reduce the computational cost by 30% whilst maintaining the same quality or provide a 0.2dB gain for a similar computational cost. Results on publicly available datasets demonstrate that the proposed algorithms surpass current state-of-the-art performance in both accuracy and efficiency.
Jose Caballero, Christian Ledig, Andrew P. Aitken, Alejandro Acosta, Johannes Totz, Wenzhe Shi
CVPR7
2017 Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network
abstract
Despite the breakthroughs in accuracy and speed of single image super-resolution using faster and deeper convolutional neural networks, one central problem remains largely unsolved: how do we recover the finer texture details when we super-resolve at large upscaling factors? The behavior of optimization-based super-resolution methods is principally driven by the choice of the objective function. Recent work has largely focused on minimizing the mean squared reconstruction error. The resulting estimates have high peak signal-to-noise ratios, but they are often lacking high-frequency details and are perceptually unsatisfying in the sense that they fail to match the fidelity expected at the higher resolution. In this paper, we present SRGAN, a generative adversarial network (GAN) for image super-resolution (SR). To our knowledge, it is the first framework capable of inferring photo-realistic natural images for 4x upscaling factors. To achieve this, we propose a perceptual loss function which consists of an adversarial loss and a content loss. The adversarial loss pushes our solution to the natural image manifold using a discriminator network that is trained to differentiate between the super-resolved images and original photo-realistic images. In addition, we use a content loss motivated by perceptual similarity instead of similarity in pixel space. Our deep residual network is able to recover photo-realistic textures from heavily downsampled images on public benchmarks. An extensive mean-opinion-score (MOS) test shows hugely significant gains in perceptual quality using SRGAN. The MOS scores obtained with SRGAN are closer to those of the original high-resolution images than to those obtained with any state-of-the-art method.
Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew P. Aitken, Alykhan Tejani, Johannes Totz, Wenzhe Shi
CVPR11
2017 Fast Face-Swap Using Convolutional Neural Networks
abstract
We consider the problem of face swapping in images, where an input identity is transformed into a target identity while preserving pose, facial expression and lighting. To perform this mapping, we use convolutional neural networks trained to capture the appearance of the target identity from an unstructured collection of his/her photographs. This approach is enabled by framing the face swapping problem in terms of style transfer, where the goal is to render an image in the style of another one. Building on recent advances in this area, we devise a new loss function that enables the network to produce highly photorealistic results. By combining neural networks with simple pre- and post-processing steps, we aim at making face swap work in real-time with no input from the user.
Iryna Korshunova, Wenzhe Shi, Joni Dambre, Lucas Theis
ICCV2
2017 Amortised MAP Inference for Image Super-resolution
Casper Kaae Sønderby, Jose Caballero, Lucas Theis, Wenzhe Shi, Ferenc Huszar
ICLR4
2017 Lossy Image Compression with Compressive Autoencoders
Lucas Theis, Wenzhe Shi, Andrew Cunningham, Ferenc Huszar
ICLR (Poster)2
2016 Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network
abstract
Recently, several models based on deep neural networks have achieved great success in terms of both reconstruction accuracy and computational performance for single image super-resolution. In these methods, the low resolution (LR) input image is upscaled to the high resolution (HR) space using a single filter, commonly bicubic interpolation, before reconstruction. This means that the super-resolution (SR) operation is performed in HR space. We demonstrate that this is sub-optimal and adds computational complexity. In this paper, we present the first convolutional neural network (CNN) capable of real-time SR of 1080p videos on a single K2 GPU. To achieve this, we propose a novel CNN architecture where the feature maps are extracted in the LR space. In addition, we introduce an efficient sub-pixel convolution layer which learns an array of upscaling filters to upscale the final LR feature maps into the HR output. By doing so, we effectively replace the handcrafted bicubic filter in the SR pipeline with more complex upscaling filters specifically trained for each feature map, whilst also reducing the computational complexity of the overall SR operation. We evaluate the proposed approach using images and videos from publicly available datasets and show that it performs significantly better (+0.15dB on Images and +0.39dB on Videos) and is an order of magnitude faster than previous CNN-based methods.
Wenzhe Shi, Jose Caballero, Ferenc Huszar, Johannes Totz, Andrew P. Aitken, Rob Bishop, Daniel Rueckert
CVPR1
2015 Multi-atlas segmentation with augmented features for cardiac MR images
Wenjia Bai, Wenzhe Shi, Christian Ledig, Daniel Rueckert
Medical Image Anal.2
2015 A bi-ventricular cardiac atlas built from 1000+ high resolution MR images of healthy subjects and an analysis of shape and motion
Wenjia Bai, Wenzhe Shi, Antonio M. Simoes Monteiro de Marvao, Timothy Dawes, Declan P. O'Regan, Stuart A. Cook, Daniel Rueckert
Medical Image Anal.2
2015 Right ventricle segmentation from cardiac MRI: A collation study
Caroline Petitjean, Maria A. Zuluaga, Wenjia Bai, Jean-Nicolas Dacher, Damien Grosgeorge, Jérôme Caudron, Su Ruan, Ismail Ben Ayed, Manuel Jorge Cardoso, Hsiang-Chou Chen, Daniel Jimenez-Carretero, María J. Ledesma-Carbayo, Christos Davatzikos, Jimit Doshi, Güray Erus, Oskar M. O. Maier, Cyrus M. S. Nambakhsh, Yangming Ou, Sébastien Ourselin, Chun-Wei Peng, Nicholas S. Peters, Terry M. Peters, Martin Rajchl, Daniel Rueckert, Wenzhe Shi, Ching-Wei Wang, Haiyan Wang 0018, Jing Yuan 0001
Medical Image Anal.26
2015 4D Blood Flow Reconstruction Over the Entire Ventricle From Wall Motion and Blood Velocity Derived From Ultrasound Data
abstract
We demonstrate a new method to recover 4D blood flow over the entire ventricle from partial blood velocity measurements using multiple 3D+t colour Doppler images and ventricular wall motion estimated using 3D+t BMode images. We apply our approach to realistic simulated data to ascertain the ability of the method to deal with incomplete data, as typically happens in clinical practice. Experiments using synthetic data show that the use of wall motion improves velocity reconstruction, shows more accurate flow patterns and improves mean accuracy particularly when coverage of the ventricle is poor. The method was applied to patient data from 6 congenital cases, producing results consistent with the simulations. The use of wall motion produced more plausible flow patterns and reduced the reconstruction error in all patients.
Alberto Gómez 0002, Adelaide de Vecchi, Martin Jantsch, Wenzhe Shi, Kuberan Pushparajah, John M. Simpson, Nicolas Smith, Daniel Rueckert, Tobias Schaeffter, Graeme P. Penney
IEEE Trans. Medical Imaging4
2014 Patch-Based Evaluation of Image Segmentation
abstract
The quantification of similarity between image segmentations is a complex yet important task. The ideal similarity measure should be unbiased to segmentations of different volume and complexity, and be able to quantify and visualise segmentation bias. Similarity measures based on overlap, e.g. Dice score, or surface distances, e.g. Hausdorff distance, clearly do not satisfy all of these properties. To address this problem, we introduce Patch-based Evaluation of Image Segmentation (PEIS), a general method to assess segmentation quality. Our method is based on finding patch correspondences and the associated patch displacements, which allow the estimation of segmentation bias. We quantify both the agreement of the segmentation boundary and the conservation of the segmentation shape. We further assess the segmentation complexity within patches to weight the contribution of local segmentation similarity to the global score. We evaluate PEIS on both synthetic data and two medical imaging datasets. On synthetic segmentations of different shapes, we provide evidence that PEIS, in comparison to the Dice score, produces more comparable scores, has increased sensitivity and estimates segmentation bias accurately. On cardiac magnetic resonance (MR) images, we demonstrate that PEIS can evaluate the performance of a segmentation method independent of the size or complexity of the segmentation under consideration. On brain MR images, we compare five different automatic hippocampus segmentation techniques using PEIS. Finally, we visualise the segmentation bias on a selection of the cases.
Christian Ledig, Wenzhe Shi, Wenjia Bai, Daniel Rueckert
CVPR2
2014 Multi-atlas Spectral PatchMatch: Application to Cardiac Image Segmentation
Wenzhe Shi, Hervé Lombaert, Wenjia Bai, Christian Ledig, Xiahai Zhuang, Antonio M. Simoes Monteiro de Marvao, Timothy Dawes, Declan P. O'Regan, Daniel Rueckert
MICCAI (1)1
2013 Model-Guided Directional Minimal Path for Fully Automatic Extraction of Coronary Centerlines from Cardiac CTA
Wenzhe Shi, Daniel Rueckert, Mingxing Hu, Sébastien Ourselin, Xiahai Zhuang
MICCAI (1)2
2013 Cardiac Image Super-Resolution with Global Correspondence Using Multi-Atlas PatchMatch
Wenzhe Shi, Jose Caballero, Christian Ledig, Xiahai Zhuang, Wenjia Bai, Kanwal K. Bhatia, Antonio M. Simoes Monteiro de Marvao, Timothy Dawes, Declan P. O'Regan, Daniel Rueckert
MICCAI (3)1
2013 Temporal sparse free-form deformations
Wenzhe Shi, Martin Jantsch, Paul Aljabar, Luis Pizarro, Wenjia Bai, Haiyan Wang 0018, Declan P. O'Regan, Xiahai Zhuang, Daniel Rueckert
Medical Image Anal.1
2013 Benchmarking framework for myocardial tracking and deformation algorithms: An open access database
Catalina Tobon-Gomez, Mathieu De Craene, Kristin McLeod, Lennart Tautz, Wenzhe Shi, Anja Hennemuth, Adityo Prakosa, Gerry Carr-White, Stam Kapetanakis, Anja Lutz, Volker Rasche, Tobias Schaeffter, Constantine Butakoff, Ola Friman, Tommaso Mansi, Maxime Sermesant, Xiahai Zhuang, Sébastien Ourselin, Heinz-Otto Peitgen, Xavier Pennec, Reza Razavi, Daniel Rueckert, Alejandro F. Frangi, Kawal S. Rhode
Medical Image Anal.5
2013 The estimation of patient-specific cardiac diastolic functions from clinical measurements
abstract
An unresolved issue in patients with diastolic dysfunction is that the estimation of myocardial stiffness cannot be decoupled from diastolic residual active tension (AT) because of the impaired ventricular relaxation during diastole. To address this problem, this paper presents a method for estimating diastolic mechanical parameters of the left ventricle (LV) from cine and tagged MRI measurements and LV cavity pressure recordings, separating the passive myocardial constitutive properties and diastolic residual AT. Dynamic C1-continuous meshes are automatically built from the anatomy and deformation captured from dynamic MRI sequences. Diastolic deformation is simulated using a mechanical model that combines passive and active material properties. The problem of non-uniqueness of constitutive parameter estimation using the well known Guccione law is characterized by reformulation of this law. Using this reformulated form, and by constraining the constitutive parameters to be constant across time points during diastole, we separate the effects of passive constitutive properties and the residual AT during diastolic relaxation. Finally, the method is applied to two clinical cases and one control, demonstrating that increased residual AT during diastole provides a potential novel index for delineating healthy and pathological cases.
Jiahe Xi, Pablo Lamata, Steven A. Niederer, Sander Land, Wenzhe Shi, Xiahai Zhuang, Sébastien Ourselin, Simon G. Duckett, Anoop Shetty, C. Aldo Rinaldi, Daniel Rueckert, Reza Razavi, Nicolas Smith
Medical Image Anal.5
2013 A Probabilistic Patch-Based Label Fusion Model for Multi-Atlas Segmentation With Registration Refinement: Application to Cardiac MR Images
abstract
The evaluation of ventricular function is important for the diagnosis of cardiovascular diseases. It typically involves measurement of the left ventricular (LV) mass and LV cavity volume. Manual delineation of the myocardial contours is time-consuming and dependent on the subjective experience of the expert observer. In this paper, a multi-atlas method is proposed for cardiac magnetic resonance (MR) image segmentation. The proposed method is novel in two aspects. First, it formulates a patch-based label fusion model in a Bayesian framework. Second, it improves image registration accuracy by utilizing label information, which leads to improvement of segmentation accuracy. The proposed method was evaluated on a cardiac MR image set of 28 subjects. The average Dice overlap metric of our segmentation is 0.92 for the LV cavity, 0.89 for the right ventricular cavity and 0.82 for the myocardium. The results show that the proposed method is able to provide accurate information for clinical diagnosis.
Wenjia Bai, Wenzhe Shi, Declan P. O'Regan, Tong Tong 0001, Haiyan Wang 0018, Shahnaz Jamil-Copley, Nicholas S. Peters, Daniel Rueckert
IEEE Trans. Medical Imaging2
2012 Registration Using Sparse Free-Form Deformations
Wenzhe Shi, Xiahai Zhuang, Luis Pizarro, Wenjia Bai, Haiyan Wang 0018, Kai-Pin Tung, Philip J. Edwards, Daniel Rueckert
MICCAI (2)1
2012 A Comprehensive Cardiac Motion Estimation Framework Using Both Untagged and 3-D Tagged MR Images Based on Nonrigid Registration
abstract
In this paper, we present a novel technique based on nonrigid image registration for myocardial motion estimation using both untagged and 3-D tagged MR images. The novel aspect of our technique is its simultaneous usage of complementary information from both untagged and 3-D tagged MR images. To estimate the motion within the myocardium, we register a sequence of tagged and untagged MR images during the cardiac cycle to a set of reference tagged and untagged MR images at end-diastole. The similarity measure is spatially weighted to maximize the utility of information from both images. In addition, the proposed approach integrates a valve plane tracker and adaptive incompressibility into the framework. We have evaluated the proposed approach on 12 subjects. Our results show a clear improvement in terms of accuracy compared to approaches that use either 3-D tagged or untagged MR image information alone. The relative error compared to manually tracked landmarks is less than 15% throughout the cardiac cycle. Finally, we demonstrate the automatic analysis of cardiac function from the myocardial deformation fields.
Wenzhe Shi, Xiahai Zhuang, Haiyan Wang 0018, Simon G. Duckett, Duy V. N. Luong, Catalina Tobon-Gomez, Kai-Pin Tung, Philip J. Edwards, Kawal S. Rhode, Reza Razavi, Sébastien Ourselin, Daniel Rueckert
IEEE Trans. Medical Imaging1
2011 HCI⁁2 Workbench: A development tool for multimodal human-computer interaction systems
abstract
In this paper, we present a novel software tool designed and implemented to simplify the development process of Multimodal Human-Computer Interaction (MHCI) systems. This tool, which is called the HCI^2 Workbench, exploits a Publish / Subscribe (P/S) architecture to facilitate efficient and reliable inter-module data communication and runtime system management. In addition, through a combination of SDK, software tools, and standardized description / configuration file semantics, the HCI^2 Workbench provides an easy-to-follow procedure for developing highly flexible and reusable modules. Moreover, the HCI^2 Workbench features system persistence and portability by using standardized module packaging method and system configuration files. Last but not least, usability was another major concern. Unlike other similar tool, including Psyclone and ActiveMQ, the HCI^2 Workbench provides a complete graphical environment to support every step in a typical MHCI system development process, including module program development and debugging, module packaging, module management, system configuration, module and system testing, in a convenient and intuitive manner. To help demonstrating the HCI^2 Workbench, we also present a readily-applicable system developed using our tool. This open-source demo system, which is called the CamGame, consists of an interactive system allowing users to play a computer game using hand-held marker(s) and low-cost camera(s) instead of keyboard and mouse.
Jie Shen 0008, Wenzhe Shi, Maja Pantic
FG2