Björn Stenger

dblp:s/BStenger · also Bjoern Stenger · DBLP profile ↗
← Back
69ranked-venue papers
10as first author
13since 2021 · last 2025
0009-0008-0465-5545ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 10 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 50 · 7 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 LLDiffusion: Learning degradation representations in diffusion models for low-light image enhancement
Tao Wang 0052, Kaihao Zhang, Yong Zhang 0034, Wenhan Luo, Björn Stenger, Tong Lu 0002, Tae-Kyun Kim 0001, Wei Liu 0005
Pattern Recognit.5
2024 Text Removal In E-Commerce Images: A Comparison Of Inpainting Methods
Hiya Roy, Björn Stenger
BMVC2
2024 Linearly Controllable GAN: Unsupervised Feature Categorization and Decomposition for Image Generation and Manipulation
Sehyung Lee, Mijung Kim, Yeongnam Chae, Björn Stenger
ECCV (4)4
2024 GridFormer: Residual Dense Transformer with Grid Structure for Image Restoration in Adverse Weather Conditions
Tao Wang 0052, Kaihao Zhang, Ziqian Shao, Wenhan Luo, Björn Stenger, Tong Lu 0002, Tae-Kyun Kim 0001, Wei Liu 0005, Hongdong Li
Int. J. Comput. Vis.5
2024 MC-Blur: A Comprehensive Benchmark for Image Deblurring
abstract
Blur artifacts can seriously degrade the visual quality of images, and numerous deblurring methods have been proposed for specific scenarios. However, in most real-world images, blur is caused by different factors, e.g., motion, and defocus. In this paper, we address how other deblurring methods perform in the case of multiple types of blur. For in-depth performance evaluation, we construct a new large-scale multi-cause image deblurring dataset (MC-Blur), including real-world and synthesized blurry images with different blur factors. The images in the proposed MC-Blur dataset are collected using other techniques: averaging sharp images captured by a 1000-fps high-speed camera, convolving Ultra-High-Definition (UHD) sharp images with large-size kernels, adding defocus to images, and real-world blurry images captured by various camera models. Based on the MC-Blur dataset, we conduct extensive benchmarking studies to compare SOTA methods in different scenarios, analyze their efficiency, and investigate the buildataset’s capacity. These benchmarking results provide a comprehensive overview of the advantages and limitations of current deblurring methods, revealing our dataset’s advances. The dataset is available to the public athttps://github.com/HDCVLab/MC-Blur-Dataset.
Kaihao Zhang, Tao Wang 0052, Wenhan Luo, Wenqi Ren, Björn Stenger, Wei Liu 0005, Hongdong Li, Ming-Hsuan Yang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Ultra-High-Definition Low-Light Image Enhancement: A Benchmark and Transformer-Based Method
abstract
As the quality of optical sensors improves, there is a need for processing large-scale images. In particular, the ability of devices to capture ultra-high definition (UHD) images and video places new demands on the image processing pipeline. In this paper, we consider the task of low-light image enhancement (LLIE) and introduce a large-scale database consisting of images at 4K and 8K resolution. We conduct systematic benchmarking studies and provide a comparison of current LLIE algorithms. As a second contribution, we introduce LLFormer, a transformer-based low-light enhancement method. The core components of LLFormer are the axis-based multi-head self-attention and cross-layer attention fusion block, which significantly reduces the linear complexity. Extensive experiments on the new dataset and existing public datasets show that LLFormer outperforms state-of-the-art methods. We also show that employing existing LLIE methods trained on our benchmark as a pre-processing step significantly improves the performance of downstream tasks, e.g., face detection in low-light conditions. The source code and pre-trained models are available at https://github.com/TaoWangzj/LLFormer.
Tao Wang 0052, Kaihao Zhang, Tianrun Shen, Wenhan Luo, Björn Stenger, Tong Lu 0002
AAAI5
2023 PCT-Net: Full Resolution Image Harmonization Using Pixel-Wise Color Transformations
abstract
In this paper, we present PCT-Net, a simple and general image harmonization method that can be easily applied to images at full-resolution. The key idea is to learn a parameter network that uses downsampled input images to predict the parameters for pixel-wise color transforms (PCTs) which are applied to each pixel in the full-resolution image. We show that affine color transforms are both efficient and effective, resulting in state-of-the-art harmonization results. Moreover, we explore both CNNs and Transformers as the parameter network, and show that Transformers lead to better results. We evaluate the proposed method on the public full-resolution iHarmony4 dataset, which is comprised of four datasets, and show a reduction of the foreground MSE (fMSE) and MSE values by more than 20% and an increase of the PSNR value by 1.4dB, while keeping the architecture light-weight. In a user study with 20 people, we show that the method achieves a higher B-T score than two other recent methods.
Julian Jorge Andrade Guerreiro, Mitsuru Nakazawa, Björn Stenger
CVPR3
2023 Online Knowledge Distillation for Multi-task Learning
abstract
Multi-task learning (MTL) has found wide application in computer vision tasks. We train a backbone network to learn a shared representation for different tasks such as semantic segmentation, depth- and normal estimation. In many cases negative transfer, i.e. impaired performance in the target domain, causes the MTL accuracy to be lower than training the corresponding single-task networks. To mitigate this issue, we propose an online knowledge distillation method, where single-task networks are trained simultaneously with the MTL network to guide the optimization process. We propose selectively training layers for each task using an adaptive feature distillation (AFD) loss with an online task weighting (OTW) scheme. This task-wise feature distillation enables the MTL network to be trained in a similar way to the single-task networks. On the NYUv2 and Cityscapes datasets we show improvements over a baseline MTL model by 6.22% and 9.19%, respectively, outperforming recent MTL methods. We validate the design choices in ablative experiments, including the use of online task weighting and the adaptive feature distillation loss.
Geethu Miriam Jacob, Vishal Agarwal, Björn Stenger
WACV3
2022 Action Spotting in Soccer Videos Using Multiple Scene Encoders
abstract
Action spotting, which temporally localizes specific actions in a video, is an important task for understanding high-level semantic information. In this paper, we formulate the action spotting task to one of scene sequence recognition and propose a model with multiple scene encoders to capture scene changes around the timestamp where an action occurs. We divide the input into multiple subsets to reduce the influence of scene context that is temporally distant, and feed every subset into a scene encoder to learn scene context in every subset. Because the optimal temporal length for time windows (chunks) is different for each action, we analyze the influence of chunk sizes for action spotting. The experimental results on the public SoccerNet-v2 dataset demonstrate state-of-the-art accuracy. By using embedding features, our method obtains an Average-mAP of 75.3%. In addition, we confirm that the performance can be improved by using optimal chunk sizes for different actions.
Yuzhi Shi, Hiroaki Minoura, Takayoshi Yamashita, Tsubasa Hirakawa, Hironobu Fujiyoshi, Mitsuru Nakazawa, Yeongnam Chae, Björn Stenger
ICPR8
2022 Parsing Line Chart Images Using Linear Programming
abstract
This paper proposes a method for automatically recovering data from chart images. In particular we focus on the task of estimating line charts, as the most common chart type, in a fully automatic way that handles line occlusions, as well as lines of different styles, e.g. dashed or dotted. For this, we first train a single semantic segmentation network to predict probability maps for each different line styles. We then construct a graph based on this output and formulate the line tracing task as a minimum-cost-flow problem, optimizing a cost function using linear programming. From the traced lines, the axes, and text labels, we recover the numerical values used to generate the chart. In experiments on six datasets, containing both synthesized and crawled images, we show significant improvements over prior work.
Hajime Kato, Mitsuru Nakazawa, Hsuan-Kung Yang, Björn Stenger
WACV5
2022 Deep Image Deblurring: A Survey
Kaihao Zhang, Wenqi Ren, Wenhan Luo, Wei-Sheng Lai, Björn Stenger, Ming-Hsuan Yang 0001, Hongdong Li
Int. J. Comput. Vis.5
2021 Facial Action Unit Detection With Transformers
abstract
The Facial Action Coding System is a taxonomy for fine-grained facial expression analysis. This paper proposes a method for detecting Facial Action Units (FAU), which de-fine particular face muscle activity, from an input image. FAU detection is formulated as a multi-task learning problem, where image features and attention maps are input to a branch for each action unit to extract discriminative feature embeddings, using a new loss function, the center contrastive (CC) loss. We employ a new FAU correlation net-work, based on a transformer encoder architecture, to capture the relationships between different action units for the wide range of expressions in the training data. The resulting features are shown to yield high classification performance. We validate our design choices, including the use of CC-loss and Tversky loss functions, in ablative experiments. We show that the proposed method outperforms state-of-the-art techniques on two public datasets, BP4D and DISFA, with an absolute improvement of the F1-score of over 2% on each.
Geethu Miriam Jacob, Björn Stenger
CVPR2
2021 Benchmarking Ultra-High-Definition Image Super-resolution
abstract
Increasingly, modern mobile devices allow capturing images at Ultra-High-Definition (UHD) resolution, which includes 4K and 8K images. However, current single image super-resolution (SISR) methods focus on super-resolving images to ones with resolution up to high definition (HD) and ignore higher-resolution UHD images. To explore their performance on UHD images, in this paper, we first introduce two large-scale image datasets, UHDSR4K and UHDSR8K, to benchmark existing SISR methods. With 70,000 V100 GPU hours of training, we benchmark these methods on 4K and 8K resolution images under seven different settings to provide a set of baseline models. Moreover, we propose a baseline model, called Mesh Attention Network (MANet) for SISR. The MANet applies the attention mechanism in both different depths (horizontal) and different levels of receptive field (vertical). In this way, correlations among feature maps are learned, enabling the network to focus on more important features.
Kaihao Zhang, Dongxu Li 0003, Wenhan Luo, Wenqi Ren, Björn Stenger, Wei Liu 0005, Hongdong Li, Ming-Hsuan Yang 0001
ICCV5
2020 Deblurring by Realistic Blurring
abstract
Existing deep learning methods for image deblurring typically train models using pairs of sharp images and their blurred counterparts. However, synthetically blurring images does not necessarily model the blurring process in real-world scenarios with sufficient accuracy. To address this problem, we propose a new method which combines two GAN models, i.e., a learning-to-Blur GAN (BGAN) and learning-to-DeBlur GAN (DBGAN), in order to learn a better model for image deblurring by primarily learning how to blur images. The first model, BGAN, learns how to blur sharp images with unpaired sharp and blurry image sets, and then guides the second model, DBGAN, to learn how to correctly deblur such images. In order to reduce the discrepancy between real blur and synthesized blur, a relativistic blur loss is leveraged. As an additional contribution, this paper also introduces a Real-World Blurred Image (RWBI) dataset including diverse blurry images. Our experiments show that the proposed method achieves consistently superior quantitative performance as well as higher perceptual quality on both the newly proposed dataset and the public GOPRO dataset.
Kaihao Zhang, Wenhan Luo, Yiran Zhong, Lin Ma 0002, Björn Stenger, Wei Liu 0005, Hongdong Li
CVPR5
2020 Every Moment Matters: Detail-Aware Networks to Bring a Blurry Image Alive
abstract
Motion-blurred images are the result of light accumulation over the period of camera exposure time, during which the camera and objects in the scene are in relative motion to each other. The inverse process of extracting an image sequence from a single motion-blurred image is an ill-posed vision problem. One key challenge is that the motions across frames are subtle, which makes the generating networks difficult to capture them and thus the recovery sequences lack motion details. In order to alleviate this problem, we propose a detail-aware network with three consecutive stages to improve the reconstruction quality by addressing specific aspects in the recovery process. The detail-aware network firstly models the dynamics using a cycle flow loss, resolving the temporal ambiguity of the reconstruction in the first stage. Then, a GramNet is proposed in the second stage to refine subtle motion between continuous frames using Gram matrices as motion representation. Finally, we introduce a HeptaGAN in the third stage to bridge the continuous and discrete nature of exposure time and recovered frames, respectively, in order to maintain rich detail. Experiments show that the proposed detail-aware networks produce sharp image sequences with rich details and subtle motion, outperforming the state-of-the-art methods.
Kaihao Zhang, Wenhan Luo, Björn Stenger, Wenqi Ren, Lin Ma 0002, Hongdong Li
ACM Multimedia3
2019 Learning Classifiers on Positive and Unlabeled Data with Policy Gradient
abstract
Existing algorithms aiming to learn a binary classifier from positive (P) and unlabeled (U) data generally require estimating the class prior or label noises ahead of building a classification model. However, the estimation and classifier learning are normally conducted in a pipeline instead of being jointly optimized. In this paper, we propose to alternatively train the two steps using reinforcement learning. Our proposal adopts a policy network to adaptively make assumptions on the labels of unlabeled data, while a classifier is built upon the output of the policy network and provides rewards to learn a better strategy. The dynamic and interactive training between the policy maker and the classifier can exploit the unlabeled data in a more effective manner and yield a significant improvement on the classification performance. Furthermore, we present two different approaches to represent the actions sampled from the policy. The first approach considers continuous actions as soft labels, while the other uses discrete actions as hard assignment of labels for unlabeled examples. We validate the effectiveness of the proposed method on two benchmark datasets as well as one e-commerce dataset. The result shows the proposed method is able to consistently outperform state-of-the-art methods in various settings.
Tianyu Li 0007, Chien-Chih Wang, Patricia Ortal, Qifang Zhao, Björn Stenger, Yu Hirate
ICDM6
2019 Trajectories as Topics: Multi-Object Tracking by Topic Discovery
abstract
This paper proposes a new approach to multi-object tracking by semantic topic discovery. We dynamically cluster frame-by-frame detections and treat objects as topics, allowing the application of the Dirichlet process mixture model. The tracking problem is cast as a topic-discovery task, where the video sequence is treated analogously to a document. It addresses tracking issues such as object exclusivity constraints as well as tracking management without the need for heuristic thresholds. Variation of object appearance is modeled as the dynamics of word co-occurrence and handled by updating the cluster parameters across the sequence in the dynamical clustering procedure. We develop two kinds of visual representation based on super-pixel and deformable part model and integrate them into the model of automatic topic discovery for tracking rigid and non-rigid objects, respectively. In experiments on public data sets, we demonstrate the effectiveness of the proposed algorithm.
Wenhan Luo, Björn Stenger, Tae-Kyun Kim 0001
IEEE Trans. Image Process.2
2018 Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals
abstract
In this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods on three tasks: single frame 3D pose estimation, 3D hand tracking, and hand pose estimation during object interaction. We analyze the performance of different CNN structures with regard to hand shape, joint visibility, view point and articulation distributions. Our findings include: (1) isolated 3D hand pose estimation achieves low mean errors (10 mm) in the view point range of [70, 120] degrees, but it is far from being solved for extreme view points; (2) 3D volumetric representations outperform 2D CNNs, better capturing the spatial structure of the depth data; (3) Discriminative methods still generalize poorly to unseen hand shapes; (4) While joint occlusions pose a challenge for most methods, explicit modeling of structure constraints can significantly narrow the gap between errors on visible and occluded joints.
Shanxin Yuan, Guillermo Garcia-Hernando, Björn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov 0001, Jan Kautz, Sina Honari, Liuhao Ge, Junsong Yuan 0001, Xinghao Chen 0001, Guijin Wang, Fan Yang 0032, Kai Akiyama, Yang Wu 0001, Qingfu Wan, Meysam Madadi, Sergio Escalera, Shile Li, Dongheui Lee, Iasonas Oikonomidis, Antonis A. Argyros, Tae-Kyun Kim 0001
CVPR3
2018 Deep Heterogeneous Autoencoders for Collaborative Filtering
abstract
This paper leverages heterogeneous auxiliary information to address the data sparsity problem of recommender systems. We propose a model that learns a shared feature space from heterogeneous data, such as item descriptions, product tags and online purchase history, to obtain better predictions. Our model consists of autoencoders, not only for numerical and categorical data, but also for sequential data, which enables capturing user tastes, item characteristics and the recent dynamics of user preference. We learn the autoencoder architecture for each data source independently in order to better model their statistical properties. Our evaluation on two MovieLens datasets and an e-commerce dataset shows that mean average precision and recall improve over state-of-the-art methods.
Jiu Xu, Björn Stenger, Yu Hirate
ICDM3
2018 Enhancing Product Images for Click-Through Rate Improvement
abstract
This paper proposes a statistical method to enhance image quality in order to increase the click-through rate (CTR) of product images. We build a joint probability model of global image features for photos of different product categories. The images are modified in terms of brightness, contrast, and sharpness in order to increase the expected CTR. The effectiveness of the method is evaluated using a perceptual user study, comparing it to histogram equalization methods, and by conducting an A/B test over a one-week period on the e-commerce site Rakuten Ichiba.
Yeongnam Chae, Mitsuru Nakazawa, Björn Stenger
ICIP3
2018 Dense Bynet: Residual Dense Network for Image Super Resolution
abstract
This paper proposes a method, Dense ByNet, for single image super-resolution based on a convolutional neural network (CNN). The main innovation is a new architecture that combines several CNN design choices. Using a residual network as a basis, it introduces dense connections inside residual blocks, significantly reducing the number of parameters. Second, we apply dilation convolutions to increase the spatial context. Lastly, we propose modifications to the activation and cost functions. We evaluate the method on benchmark datasets and show that it achieves state-of-the-art results over multiple upscaling factors in terms of peak SNR and structural similarity (SSIM).
Jiu Xu, Yeongnam Chae, Björn Stenger, Ankur Datta
ICIP3
2017 BigHand2.2M Benchmark: Hand Pose Dataset and State of the Art Analysis
abstract
In this paper we introduce a large-scale hand pose dataset, collected using a novel capture method. Existing datasets are either generated synthetically or captured using depth sensors: synthetic datasets exhibit a certain level of appearance difference from real depth images, and real datasets are limited in quantity and coverage, mainly due to the difficulty to annotate them. We propose a tracking system with six 6D magnetic sensors and inverse kinematics to automatically obtain 21-joints hand pose annotations of depth maps captured with minimal restriction on the range of motion. The capture protocol aims to fully cover the natural hand pose space. As shown in embedding plots, the new dataset exhibits a significantly wider and denser range of hand poses compared to existing benchmarks. Current state-of-the-art methods are evaluated on the dataset, and we demonstrate significant improvements in cross-benchmark performance. We also show significant improvements in egocentric hand pose estimation with a CNN trained on the new dataset.
Shanxin Yuan, Qi Ye 0001, Björn Stenger, Siddhant Jain, Tae-Kyun Kim 0001
CVPR3
2017 BYNET-SR: Image super resolution with a bypass connection network
abstract
This paper proposes a deep residual network, ByNet, for the single image super resolution task. The main innovation is the introduction of two effective components, bypass connections and a feature scaling layer. Bypass connections are formed either by skip connections that jump multiple layers or by adding a convolution layer in such a jump. The final feature scaling layer enables more robust convergence. Experiments on standard benchmarks show that the proposed method achieves state of the art results over multiple scales in terms of PSNR and structural similarity (SSIM).
Jiu Xu, Yeongnam Chae, Björn Stenger
ICIP3
2017 Pano2CAD: Room Layout from a Single Panorama Image
abstract
This paper presents a method of estimating the geometry of a room and the 3D pose of objects from a single 360 panorama image. Assuming ManhattanWorld geometry, we formulate the task as an inference problem in which we estimate positions and orientations of walls and objects. The method combines surface normal estimation, 2D object detection and 3D object pose estimation. Quantitative results are presented on a dataset of synthetically generated 3D rooms containing objects, as well as on a subset of handlabeled images from the public SUN360 dataset.
Jiu Xu, Björn Stenger, Tommi Kerola, Tony Tung
WACV2
2016 How many bits do I need for matching local binary descriptors?
abstract
In this paper we provide novel insights about the performance and design of popular pairwise tests-based local binary descriptors with the aim of answering the question: How many bits are needed for matching local binary descriptors? We use the interpretation of binary descriptors as a Locality Sensitive Hashing (LSH) scheme for approximating Kendall's tau rank distance between image patches. Based on this understanding we compare local binary descriptors in terms of the number of bits that are required to achieve a certain performance in feature-based matching problems. Furthermore, we introduce a calibration method to automatically determine a suitable number of bits required in an image matching scenario. We provide a performance analysis in image matching and structure from motion benchmarks, showing calibration results in visual odometry and object recognition problems. Our results show that excellent performance can be achieved using a small fraction of the total number of bits from the whole descriptor, speeding-up matching and reducing storage requirements.
Pablo Fernández Alcantarilla, Björn Stenger
ICRA2
2016 Precise deterministic change detection for smooth surfaces
abstract
We introduce a precise deterministic approach for pixel-wise change detection in images taken of a scene of interest over time. Our motivation is for applications such as artefact condition monitoring and structural inspection, where a common problem is the need to efficiently and accurately identify subtle signs of damage and deterioration. The approach we describe is designed to compensate for the three most common sources of nuisance variation encountered when tackling the problem of change detection, namely: viewpoint variation due to camera motion between images, photometric variation due to lighting differences, and changes in image resolution/focal settings. To tackle viewpoint variation, particularly in areas of low texture, we propose the use of the generalised PatchMatch (PM) correspondence algorithm to compute a dense flow field. The flow field is regularized using a Thin Plate Spline (TPS) model which assumes a smooth underlying geometry and allows registration to be interpolated precisely through areas of low texture or uncertain flow. To compensate for low-frequency lighting variation, we fit a second TPS model to the photometric differences between registered images. Finally, to account for changes in focal settings, we estimate and apply a blurring kernel via optimisation over image differences. We provide a thorough evaluation of the performance of our method on an illustrative toy dataset and on two recent, real-world inspection datasets. Our approach performs favourably versus state-of-the-art baselines in both cases, while remaining relatively transparent to understand and simple to compute.
Simon Stent, Riccardo Gherardi, Björn Stenger, Roberto Cipolla
WACV3
2016 Expressive visual text-to-speech as an assistive technology for individuals with autism spectrum conditions
abstract
Adults with Autism Spectrum Conditions (ASC) experience marked difficulties in recognising the emotions of others and responding appropriately. The clinical characteristics of ASC mean that face to face or group interventions may not be appropriate for this clinical group. This article explores the potential of a new interactive technology, converting text to emotionally expressive speech, to improve emotion processing ability and attention to faces in adults with ASC. We demonstrate a method for generating a near-videorealistic avatar (XpressiveTalk), which can produce a video of a face uttering inputted text, in a large variety of emotional tones. We then demonstrate that general population adults can correctly recognize the emotions portrayed by XpressiveTalk. Adults with ASC are significantly less accurate than controls, but still above chance levels for inferring emotions from XpressiveTalk. Both groups are significantly more accurate when inferring sad emotions from XpressiveTalk compared to the original actress, and rate these expressions as significantly more preferred and realistic. The potential applications for XpressiveTalk as an assistive technology for adults with ASC is discussed.
Sarah A. Cassidy, Björn Stenger, L. Van Dongen, Kayoko Yanagisawa, Vincent Wan, Simon Baron-Cohen, Roberto Cipolla
Comput. Vis. Image Underst.2
2016 Visual change detection on tunnel linings
Simon Stent, Riccardo Gherardi, Björn Stenger, Kenichi Soga, Roberto Cipolla
Mach. Vis. Appl.3
2015 Automatic Topic Discovery for Multi-Object Tracking
abstract
This paper proposes a new approach to multi-object tracking by semantic topic discovery. We dynamically cluster frame-by-frame detections and treat objects as topics, allowing the application of the Dirichlet Process Mixture Model (DPMM). The tracking problem is cast as a topic-discovery task where the video sequence is treated analogously to a document. This formulation addresses tracking issues such as object exclusivity constraints as well as cannot-link constraints which are integrated without the need for heuristic thresholds. The video is temporally segmented into epochs to model the dynamics of word (superpixel) co-occurrences and to model the temporal damping effect. In experiments on public data sets we demonstrate the effectiveness of the proposed algorithm.
Wenhan Luo, Björn Stenger, Tae-Kyun Kim 0001
AAAI2
2015 Detecting Change for Multi-View, Long-Term Surface Inspection
abstract
We describe a system for the detection of changes in multiple views of a tunnel surface. From data gathered by a robotic inspection rig, we use a structure-from-motion pipeline to build panoramas of the surface and register images from different time instances. Reliably detecting changes such as hairline cracks, water ingress and other surface damage between the registered images is a challenging problem: achieving the best possible performance for a given set of data requires sub-pixel precision and careful modelling of the noise sources. The task is further complicated by factors such as unavoidable registration error and changes in image sensors, capture settings and lighting. Our contribution is a novel approach to change detection using a two-channel convolutional neural network. The network accepts pairs of approximately registered image patches taken at different times and classifies them to detect anomalous changes. To train the network, we take advantage of synthetically generated training examples and the homogeneity of the tunnel surfaces to eliminate most of the manual labelling effort. We evaluate our method on field data gathered from a live tunnel over several months, demonstrating it to outperform existing approaches from recent literature and industrial practice.
Simon Stent, Riccardo Gherardi, Björn Stenger, Roberto Cipolla
BMVC3
2015 Editorial
Björn Stenger, Norimichi Ukita, Yoichi Sato 0001, Pascal Fua, David J. Fleet
Comput. Vis. Image Underst.1
2015 Distances and Means of Direct Similarities
Minh-Tri Pham, Oliver J. Woodford, Frank Perbet, Atsuto Maki, Riccardo Gherardi, Björn Stenger, Roberto Cipolla
Int. J. Comput. Vis.6
2014 Reconstructing Fukushima: A Case Study
abstract
We present the application of 3D reconstruction technology to the inspection and decommissioning work at the damaged Fukushima Daiichi nuclear power station in Japan. We discuss the challenges of this project, such as the difficult image capture conditions (including under water), required use of limited imaging hardware, and capture by personnel inexperienced in 3D reconstruction. We present an overview of the system developed for this project, a real-time reconstruction pipeline with robust camera pose estimation, low-latency probabilistic dense depth estimation and a novel descriptor for point cloud alignment - the Co-occurrence Histogram of Angle and Distance (CHAD). We discuss the modifications required to standard algorithms in order to perform reliably in such a scenario. As well as quantitative evaluations of these components on existing datasets, we show qualitative 3D reconstruction results of debris from the damaged plant and its spent fuel pool. Such results have enabled planning of the critical process of debris removal, without the harmful requirement of extensive human presence on site.
Akihito Seki, Oliver J. Woodford, Björn Stenger, Makoto Hatakeyama, Junichi Shimamura
3DV4
2014 Full-Angle Quaternions for Robustly Matching Vectors of 3D Rotations
abstract
In this paper we introduce a new distance for robustly matching vectors of 3D rotations. A special representation of 3D rotations, which we coin full-angle quaternion (FAQ), allows us to express this distance as Euclidean. We apply the distance to the problems of 3D shape recognition from point clouds and 2D object tracking in color video. For the former, we introduce a hashing scheme for scale and translation which outperforms the previous state-of-the-art approach on a public dataset. For the latter, we incorporate online subspace learning with the proposed FAQ representation to highlight the benefits of the new representation.
Stephan Liwicki, Minh-Tri Pham, Stefanos Zafeiriou, Maja Pantic, Björn Stenger
CVPR5
2014 Bi-label Propagation for Generic Multiple Object Tracking
abstract
In this paper, we propose a label propagation framework to handle the multiple object tracking (MOT) problem for a generic object type (cf. pedestrian tracking). Given a target object by an initial bounding box, all objects of the same type are localized together with their identities. We treat this as a problem of propagating bi-labels, i.e. a binary class label for detection and individual object labels for tracking. To propagate the class label, we adopt clustered Multiple Task Learning (cMTL) while enforcing spatio-temporal consistency and show that this improves the performance when given limited training data. To track objects, we propagate labels from trajectories to detections based on affinity using appearance, motion, and context. Experiments on public and challenging new sequences show that the proposed method improves over the current state of the art on this task.
Wenhan Luo, Tae-Kyun Kim 0001, Björn Stenger, Roberto Cipolla
CVPR3
2014 Human Body Shape Estimation Using a Multi-resolution Manifold Forest
abstract
This paper proposes a method for estimating the 3D body shape of a person with robustness to clothing. We formulate the problem as optimization over the manifold of valid depth maps of body shapes learned from synthetic training data. The manifold itself is represented using a novel data structure, a Multi-Resolution Manifold Forest (MRMF), which contains vertical edges between tree nodes as well as horizontal edges between nodes across trees that correspond to overlapping partitions. We show that this data structure allows both efficient localization and navigation on the manifold for on-the-fly building of local linear models (manifold charting). We demonstrate shape estimation of clothed users, showing significant improvement in accuracy over global shape models and models using pre-computed clusters. We further compare the MRMF with alternative manifold charting methods on a public dataset for estimating 3D motion from noisy 2D marker observations, obtaining state-of-the-art results.
Frank Perbet, Sam Johnson, Minh-Tri Pham, Björn Stenger
CVPR4
2014 Using Bounded Diameter Minimum Spanning Trees to Build Dense Active Appearance Models
Björn Stenger, Roberto Cipolla
Int. J. Comput. Vis.2
2014 Demisting the Hough Transform for 3D Shape Recognition and Registration
Oliver J. Woodford, Minh-Tri Pham, Atsuto Maki, Frank Perbet, Björn Stenger
Int. J. Comput. Vis.5
2013 Expressive Visual Text-to-Speech Using Active Appearance Models
abstract
This paper presents a complete system for expressive visual text-to-speech (VTTS), which is capable of producing expressive output, in the form of a 'talking head', given an input text and a set of continuous expression weights. The face is modeled using an active appearance model (AAM), and several extensions are proposed which make it more applicable to the task of VTTS. The model allows for normalization with respect to both pose and blink state which significantly reduces artifacts in the resulting synthesized sequences. We demonstrate quantitative improvements in terms of reconstruction error over a million frames, as well as in large-scale user studies, comparing the output of different systems.
Björn Stenger, Vincent Wan, Roberto Cipolla
CVPR2
2013 Photo-realistic expressive text to talking head synthesis
Vincent Wan, Art Blokland, Norbert Braunschweiler, Langzhou Chen, BalaKrishna Kolluru, Javier Latorre, Ranniery Maia, Björn Stenger, Kayoko Yanagisawa, Yannis Stylianou, Masami Akamine, Mark J. F. Gales, Roberto Cipolla
INTERSPEECH9
2013 Detecting bipedal motion from correlated probabilistic trajectories
Atsuto Maki, Frank Perbet, Björn Stenger, Roberto Cipolla
Pattern Recognit. Lett.3
2012 Dense Active Appearance Models Using a Bounded Diameter Minimum Spanning Tree
abstract
We present a method for producing dense Active Appearance Models (AAMs), suitable for video-realistic synthesis. To this end we estimate a joint alignment of all training images using a set of pairwise registrations and ensure that these pairwise registrations are only calculated between similar images. This is achieved by defining a graph on the image set whose edge weights correspond to registration errors and computing a bounded diameter minimum spanning tree (BDMST). Dense optical flow is used to compute pairwise registration and we introduce a flow refinement method to align small scale texture. Once registration between training images has been established we propose a method to add vertices to the AAM in a way that minimises error between the observed flow fields and a flow field interpolated between the AAM mesh points. We demonstrate a significant improvement in model compactness using the proposed method and show it dealing with cases that are problematic for current state-of-the-art approaches.
Björn Stenger, Roberto Cipolla
BMVC2
2012 Contraction Moves for Geometric Model Fitting
Oliver J. Woodford, Minh-Tri Pham, Atsuto Maki, Riccardo Gherardi, Frank Perbet, Björn Stenger
ECCV (7)6
2011 Demisting the Hough Transform for 3D Shape Recognition and Registration
abstract
In applying the Hough transform to the problem of 3D shape recognition and registration, we develop two new and powerful improvements to this popular inference method. The first, intrinsic Hough, solves the problem of exponential memory requirements of the standard Hough transform by exploiting the sparsity of the Hough space. The second, minimum-entropy Hough, explains away incorrect votes, substantially reducing the number of modes in the posterior distribution of class and pose, and improving precision. Our experiments demonstrate that these contributions make the Hough transform not only tractable but also highly accurate for our example application. Both contributions can be applied to other tasks that already use the standard Hough transform.
Oliver J. Woodford, Minh-Tri Pham, Atsuto Maki, Frank Perbet, Björn Stenger
BMVC5
2011 Color photometric stereo for multicolored surfaces
abstract
We present a multispectral photometric stereo method for capturing geometry of deforming surfaces. A novel photometric calibration technique allows calibration of scenes containing multiple piecewise constant chromaticities. This method estimates per-pixel photometric properties, then uses a RANSAC-based approach to estimate the dominant chromaticities in the scene. A likelihood term is developed linking surface normal, image intensity and photometric properties, which allows estimating the number of chromaticities present in a scene to be framed as a model estimation problem. The Bayesian Information Criterion is applied to automatically estimate the number of chromaticities present during calibration. A two-camera stereo system provides low resolution geometry, allowing the likelihood term to be used in segmenting new images into regions of constant chromaticity. This segmentation is carried out in a Markov Random Field framework and allows the correct photometric properties to be used at each pixel to estimate a dense normal map. Results are shown on several challenging real-world sequences, demonstrating state-of-the-art results using only two cameras and three light sources. Quantitative evaluation is provided against synthetic ground truth data.
Björn Stenger, Roberto Cipolla
ICCV2
2011 A new distance for scale-invariant 3D shape recognition and registration
abstract
This paper presents a method for vote-based 3D shape recognition and registration, in particular using mean shift on 3D pose votes in the space of direct similarity transforms for the first time. We introduce a new distance between poses in this space-the SRT distance. It is left-invariant, unlike Euclidean distance, and has a unique, closed-form mean, in contrast to Riemannian distance, so is fast to compute. We demonstrate improved performance over the state of the art in both recognition and registration on a real and challenging dataset, by comparing our distance with others in a mean shift framework, as well as with the commonly used Hough voting approach.
Minh-Tri Pham, Oliver J. Woodford, Frank Perbet, Atsuto Maki, Björn Stenger, Roberto Cipolla
ICCV5
2011 Incremental Linear Discriminant Analysis Using Sufficient Spanning Sets and Its Applications
Tae-Kyun Kim 0001, Björn Stenger, Josef Kittler, Roberto Cipolla
Int. J. Comput. Vis.2
2011 Video Normals from Colored Lights
abstract
We present an algorithm and the associated single-view capture methodology to acquire the detailed 3D shape, bends, and wrinkles of deforming surfaces. Moving 3D data has been difficult to obtain by methods that rely on known surface features, structured light, or silhouettes. Multispectral photometric stereo is an attractive alternative because it can recover a dense normal field from an untextured surface. We show how to capture such data, which in turn allows us to demonstrate the strengths and limitations of our simple frame-to-frame registration over time. Experiments were performed on monocular video sequences of untextured cloth and faces with and without white makeup. Subjects were filmed under spatially separated red, green, and blue lights. Our first finding is that the color photometric stereo setup is able to produce smoothly varying per-frame reconstructions with high detail. Second, when these 3D reconstructions are augmented with 2D tracking results, one can register both the surfaces and relax the homogenous-color restriction of the single-hue subject. Quantitative and qualitative experiments explore both the practicality and limitations of this simple multispectral capture system.
Gabriel J. Brostow, Carlos Hernández 0002, George Vogiatzis, Björn Stenger, Roberto Cipolla
IEEE Trans. Pattern Anal. Mach. Intell.4
2009 Random Forest Clustering and Application to Video Segmentation
abstract
This paper considers the problem of clustering large data sets in a high-dimensional space. Using a random forest, we first generate multiple partitions of the same input space, one per tree. The partitions from all trees are merged by intersecting them, resulting in a partition of higher resolution. A graph is then constructed by assigning a node to each region and linking adjacent nodes. This Graph of Superimposed Partitions (GSP) represents a remapped space of the input data where regions of high density are mapped to a larger number of nodes. Generating such a graph turns the clustering problem in the feature space into a graph clustering task which we solve with the Markov cluster algorithm (MCL). The proposed algorithm is able to capture non-convex structure while being computationally efficient, capable of dealing with large data sets. We show the clustering performance on synthetic data and apply the method to the task of video segmentation.
Frank Perbet, Björn Stenger, Atsuto Maki
BMVC2
2009 Learning to track with multiple observers
abstract
We propose a novel approach to designing algorithms for object tracking based on fusing multiple observation models. As the space of possible observation models is too large for exhaustive on-line search, this work aims to select models that are suitable for a particular tracking task at hand. During an off-line training stage observation models from various off-the-shelf trackers are evaluated. From this data different methods of fusing the observers on-line are investigated, including parallel and cascaded evaluation. Experiments on test sequences show that this evaluation is useful for automatically designing and assessing algorithms for a particular tracking task. Results are shown for face tracking with a handheld camera and hand tracking for gesture interaction. We show that for these cases combining a small number of observers in a sequential cascade results in efficient algorithms that are both robust and precise.
Björn Stenger, Thomas Woodley, Roberto Cipolla
CVPR1
2009 Correlated probabilistic trajectories for pedestrian motion detection
abstract
This paper introduces an algorithm for detecting walking motion using point trajectories in video sequences. Given a number of point trajectories, we identify those which are spatio-temporally correlated as arising from feet in walking motion. Unlike existing techniques we do not assume clean point tracks but instead propose “probabilistic trajectories” as new features to classify. These are extracted from directed acyclic graphs whose edges represent temporal point correspondences and are weighted with their matching probability in terms of appearance and location. This representation tolerates the inherent trajectory ambiguity, for example due to occlusions. We then learn the correlation between the movement of two feet using a random forest classifier. The effectiveness of the algorithm is demonstrated in experiments on image sequences captured with a static camera.
Frank Perbet, Atsuto Maki, Björn Stenger
ICCV3
2008 AIDIA - Adaptive Interface for Display InterAction
abstract
This paper presents a vision-based system for interaction with a display via hand pointing. An attention mechanism based on face and hand detection allows users in the camera’s field of view to take control of the interface. Face recognition is used for identification and customisation. The system allows the user to control the screen pointer by tracking their fist. On-screen items can be selected using one of four activation mechanisms. Current sample applications include browsing image and video collections as well as viewing a gallery of 3D objects. In experiments we demonstrate the performance of the vision components in challenging conditions and compare it to that of other systems. 1
Björn Stenger, Thomas Woodley, Tae-Kyun Kim 0001, Carlos Hernández 0002, Roberto Cipolla
BMVC1
2008 Discriminative Feature Co-Occurrence Selection for Object Detection
abstract
This paper describes an object detection framework that learns the discriminative co-occurrence of multiple features. Feature co-occurrences are automatically found by Sequential Forward Selection at each stage of the boosting process. The selected feature co-occurrences are capable of extracting structural similarities of target objects leading to better performance. The proposed method is a generalization of the framework proposed by Viola and Jones, where each weak classifier depends only on a single feature. Experimental results obtained using four object detectors, for finding faces and three different hand gestures, respectively, show that detectors trained with the proposed algorithm yield consistently higher detection rates than those based on their framework while using the same number of features.
Takeshi Mita, Toshimitsu Kaneko, Björn Stenger, Osamu Hori
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Pose estimation and tracking using multivariate regression
Arasanathan Thayananthan, Ramanan Navaratnam, Björn Stenger, Philip Torr 0001, Roberto Cipolla
Pattern Recognit. Lett.3
2007 Tracking Using Online Feature Selection and a Local Generative Model
abstract
This paper proposes an algorithm for online feature selection which improves robustness to occlusions by referring to a localized generative appearance model. Discriminative classifiers based on feature extraction have classically either prepared a fixed prior model by training offline, or continually adapted their classification parameters to any apparent appearance changes. By combining the attractive qualities of each approach, our framework can cope with appearance changes of a target object and will maintain proximity to a static appearance model. Our main contribution is the use of a generative model to guide the online feature selection to regions of an image which maintain a valid appearance. The generative model exhibits the properties of non-negativity, localization and orthogonality. We demonstrate the system in a tracking framework to show improved tracking performance through occlusions. 1
Thomas Woodley, Björn Stenger, Roberto Cipolla
BMVC2
2007 Incremental Linear Discriminant Analysis Using Sufficient Spanning Set Approximations
abstract
This paper presents a new incremental learning solution for linear discriminant analysis (LDA). We apply the concept of the sufficient spanning set approximation in each update step, i.e. for the between-class scatter matrix, the projected data matrix as well as the total scatter matrix. The algorithm yields a more general and efficient solution to incremental LDA than previous methods. It also significantly reduces the computational complexity while providing a solution which closely agrees with the batch LDA result. The proposed algorithm has a time complexity of O(Nd2) and requires O(Nd) space, where d is the reduced subspace dimension and N the data dimension. We show two applications of incremental LDA: First, the method is applied to semi-supervised learning by integrating it into an EM framework. Secondly, we apply it to the task of merging large databases which were collected during MPEG standardization for face image retrieval.
Tae-Kyun Kim 0001, Shu-Fai Wong, Björn Stenger, Josef Kittler, Roberto Cipolla
CVPR3
2007 Non-rigid Photometric Stereo with Colored Lights
abstract
We present an algorithm and the associated capture methodology to acquire and track the detailed 3D shape, bends, and wrinkles of deforming surfaces. Moving 3D data has been difficult to obtain by methods that rely on known surface features, structured light, or silhouettes. Multispec- tral photometric stereo is an attractive alternative because it can recover a dense normal field from an un-textured surface. We show how to capture such data and register it over time to generate a single deforming surface. Experiments were performed on video sequences of un- textured cloth, filmed under spatially separated red, green, and blue light sources. Our first finding is that using zero- depth-silhouettes as the initial boundary condition already produces rather smoothly varying per-frame reconstructions with high detail. Second, when these 3D reconstructions are augmented with 2D optical flow, one can register the first frame's reconstruction to every subsequent frame.
Carlos Hernández 0002, George Vogiatzis, Gabriel J. Brostow, Björn Stenger, Roberto Cipolla
ICCV4
2007 Estimating 3D hand pose using hierarchical multi-label classification
Björn Stenger, Arasanathan Thayananthan, Philip Torr 0001, Roberto Cipolla
Image Vis. Comput.1
2006 A Framework for 3D Object Recognition Using the Kernel Constrained Mutual Subspace Method
Kazuhiro Fukui, Björn Stenger, Osamu Yamaguchi
ACCV (2)2
2006 Virtual Fashion Show Using Real-Time Markerless Motion Capture
Ryuzo Okada, Björn Stenger, Tsukasa Ike, Nobuhiro Kondoh
ACCV (2)2
2006 Template-Based Hand Pose Recognition Using Multiple Cues
Björn Stenger
ACCV (2)1
2006 Multivariate Relevance Vector Machines for Tracking
Arasanathan Thayananthan, Ramanan Navaratnam, Björn Stenger, Philip Torr 0001, Roberto Cipolla
ECCV (3)3
2006 Model-Based Hand Tracking Using a Hierarchical Bayesian Filter
abstract
This paper sets out a tracking framework, which is applied to the recovery of three-dimensional hand motion from an image sequence. The method handles the issues of initialization, tracking, and recovery in a unified way. In a single input image with no prior information of the hand pose, the algorithm is equivalent to a hierarchical detection scheme, where unlikely pose candidates are rapidly discarded. In image sequences, a dynamic model is used to guide the search and approximate the optimal filtering equations. A dynamic model is given by transition probabilities between regions in parameter space and is learned from training data obtained by capturing articulated motion. The algorithm is evaluated on a number of image sequences, which include hand motion with self-occlusion in front of a cluttered background.
Björn Stenger, Arasanathan Thayananthan, Philip Torr 0001, Roberto Cipolla
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Learning a Kinematic Prior for Tree-Based Filtering
abstract
The aim in this paper is to track articulated hand motion from monocular video. Bayesian filtering is implemented by using a tree-based representation of the posterior distribution. Each tree node corresponds to a partition of the state space with piecewise constant density. In a hierarchical search regions with low probability mass can be rapidly discarded, while the modes of the posterior can be approximated to high precision. Large sets of training data are captured using a data glove, and two techniques for constructing the tree are described: One method is to cluster the collected data points using a hierarchical clustering algorithm, and use the cluster centres as nodes. Alternatively, a lower dimensional eigenspace can be partitioned using a grid at multiple resolutions, and each partition centre corresponds to a node in the tree. The effectiveness of these techniques is demonstrated by using them for tracking 3D articulated hand motion in front of a cluttered background. 1
Arasanathan Thayananthan, Björn Stenger, Philip Torr 0001, Roberto Cipolla
BMVC2
2003 Shape Context and Chamfer Matching in Cluttered Scenes
abstract
This paper compares two methods for object localization from contours: shape context and chamfer matching of templates. In the light of our experiments, we suggest improvements to the shape context: shape contexts are used to find corresponding features between model and image. In real images it is shown that the shape context is highly influenced by clutters; furthermore, even when the object is correctly localized, the feature correspondence may be poor. We show that the robustness of shape matching can be increased by including a figural continuity constraint. The combined shape and continuity cost is minimized using the Viterbi algorithm on features, resulting in improved localization and correspondence. Our algorithm can be generally applied to any feature based shape matching method. Chamfer matching correlates model templates with the distance transform of the edge image. This can be done efficiently using a coarse-to-fine search over the transformation parameters. The method is robust in clutter, however, multiple templates are needed to handle scale, rotation and shape variation. We compare both methods for locating hand shapes in cluttered images, and applied to word recognition in EZ-Gimpy images.
Arasanathan Thayananthan, Björn Stenger, Philip Torr 0001, Roberto Cipolla
CVPR (1)2
2003 Filtering Using a Tree-Based Estimator
abstract
Within this paper a new framework for Bayesian tracking is presented, which approximates the posterior distribution at multiple resolutions. We propose a tree-based representation of the distribution, where the leaves define a partition of the state space with piecewise constant density. The advantage of this representation is that regions with low probability mass can be rapidly discarded in a hierarchical search, and the distribution can be approximated to arbitrary precision. We demonstrate the effectiveness of the technique by using it for tracking 3D articulated and nonrigid motion in front of cluttered background. More specifically, we are interested in estimating the joint angles, position and orientation of a 3D hand model in order to drive an avatar.
Björn Stenger, Arasanathan Thayananthan, Philip Torr 0001, Roberto Cipolla
ICCV1
2001 Model-Based Hand Tracking Using an Unscented Kalman Filter
abstract
This paper presents a novel method for hand tracking. It uses a 3D model built from quadrics which approximates the anatomy of a human hand. This approach allows for the use of results from projective geometry that yield an elegant technique to generate the projection of the model as a set of conics, as well as providing an efficient ray tracing algorithm to handle self-occlusion. Once the model is projected, an Unscented Kalman Filter is used to update its pose in order to minimise the geometric error between the model projection and a video sequence on the background. Results from experiments with real data show the accuracy of the technique. 1
Björn Stenger, Paulo R. S. Mendonça, Roberto Cipolla
BMVC1
2001 Model-Based 3D Tracking of an Articulated Hand
abstract
This paper presents a practical technique for model-based 3D hand tracking. An anatomically accurate hand model is built from truncated quadrics. This allows for the generation of 2D profiles of the model using elegant tools from projective geometry, and for an efficient method to handle self-occlusion. The pose of the hand model is estimated with an Unscented Kalman filter (UKF), which minimizes the geometric error between the profiles and edges extracted from the images. The use of the UKF permits higher frame rates than more sophisticated estimation methods such as particle filtering, whilst providing higher accuracy than the extended Kalman filter The system is easily scalable from single to multiple views, and from rigid to articulated models. First experiments on real data using one and two cameras demonstrate the quality of the proposed method for tracking a 7 DOF hand model.
Björn Stenger, Paulo R. S. Mendonça, Roberto Cipolla
CVPR (2)1
2001 Topology Free Hidden Markov Models: Application to Background Modeling
Björn Stenger, Visvanathan Ramesh, Nikos Paragios, Frans Coetzee, Joachim M. Buhmann
ICCV1