EDBT 2026 Demo / reviewers in the wild / expert
Peter H. N. de With
dblp:w/PeterHNdeWith
· DBLP profile ↗
183ranked-venue papers
2as first author
32since 2021 · last 2026
0000-0002-7639-7716ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 138 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 29 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 11 since 2021Software engineering, systems software and programming languages · 8Systems, architecture and hardware · 3Human-computer interaction and ubiquitous computing · 2Computer networks · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative ModelsabstractFlow Matching has emerged as a powerful framework for learning continuous transformations between distributions, enabling high-fidelity generative modeling. This work introduces Symmetrical Flow Matching (SymmFlow), a new formulation that unifies semantic segmentation, classification, and image generation within a single model. Using a symmetric learning objective, SymmFlow models forward and reverse transformations jointly, ensuring bi-directional consistency, while preserving sufficient entropy for generative diversity. A new training objective is introduced to explicitly retain semantic information across flows, featuring efficient sampling while preserving semantic structure, allowing for one-step segmentation and classification without iterative refinement. Unlike previous approaches that impose strict one-to-one mapping between masks and images, SymmFlow generalizes to flexible conditioning, supporting both pixel-level and image-level class labels. Experimental results on various benchmarks demonstrate that SymmFlow achieves state-of-the-art performance on semantic image synthesis, obtaining FID scores of 11.9 on CelebAMask-HQ and 7.0 on COCO-Stuff with only 25 inference steps. Additionally, it delivers competitive results on semantic segmentation and shows promising capabilities in classification tasks. Francisco Caetano, Christiaan G. A. Viviers, Peter H. N. de With, Fons van der Sommen |
AAAI | 3 |
| 2026 | Scaling up self-supervised learning for improved surgical foundation modelsabstract• Demonstration of effectiveness of SSL for surgical computer vision using the largest dataset reported to date. • Strong generalization and robust evaluation are shown across six surgical datasets, four procedures, and three tasks, outperforming current SOTA foundation models. • Providing insights into large-scale SSL for surgical computer vision in terms of scaling, pretraining time, dataset composition, and model architecture. • Release of the models and a curated dataset of 2.1 million surgical video frames, establishing a critical resource for advancing surgical foundation model training Foundation models have revolutionized computer vision by achieving vastly superior performance across diverse tasks through large-scale pretraining on extensive datasets. However, their application in surgical computer vision has been limited. This study addresses this gap by introducing SurgeNetXL, a novel surgical foundation model that sets a new benchmark in surgical computer vision. Trained on the largest reported surgical dataset to date, comprising over 4.7 million video frames, SurgeNetXL achieves consistent top-tier performance across six datasets spanning four surgical procedures and three tasks, including semantic segmentation, surgical phase recognition, and critical view of safety (CVS) classification. Compared with the best-performing surgical foundation model, SurgeNetXL shows mean improvements of 4.0%, 8.9%, and 11.4% for semantic segmentation, phase recognition, and CVS classification, respectively. Additionally, SurgeNetXL outperforms ImageNet1k by 16.1%, 8.0%, and 4.3% for the respective tasks. In addition to advancing model performance, this study provides key insights into scaling pretraining datasets, extending training durations, and optimizing model architectures specifically for surgical computer vision. These findings pave the way for improved generalization and robustness in data-scarce scenarios, offering a comprehensive framework for future research in this domain. All models and a subset of the SurgeNetXL dataset, including over 2 million video frames, are publicly available at: https://github.com/TimJaspers0801/SurgeNet . Tim J. M. Jaspers, Ronald L. P. D. de Jong, Yiping Li 0002, Carolus H. J. Kusters, Franciscus H. A. Bakker, Romy C. van Jaarsveld, Gino M. Kuiper, Richard van Hillegersberg, Jelle P. Ruurda, Willem M. Brinkman, Josien P. W. Pluim, Peter H. N. de With, Marcel Breeuwer, Yasmina Alkhalil, Fons van der Sommen |
Medical Image Anal. | 12 |
| 2026 | S2DENet: Shallow suppression and deep enhancement network for general ultrasound image segmentation
Xintao Pang, Jinlin Yang, Zhifan Gao, Chuan Lin 0003, Yue Sun 0001, Shuo Li 0001, Peter H. N. de With, Tao Tan 0002 |
Medical Image Anal. | 7 |
| 2026 | Multi-Granularity Query Network With Adaptive Category Feature Embedding for Behavior RecognitionabstractBehavior recognition is a highly challenging task, particularly in scenarios requiring unified recognition across both human and animal subjects. Most existing approaches primarily focus on single-species datasets or rely heavily on prior information such as species labels, positional annotations, or skeletal keypoints, which limits their applicability in real-world scenarios where species labels may be ambiguous or annotations are insufficient. To address these limitations, we propose a query-based Multi-Granularity Behavior Recognition Network that directly mines cross-species shared spatiotemporal behavior patterns from raw video inputs. Specifically, we design a Multi-Granularity Query module to effectively fuse fine-grained and coarse-grained features, thereby enhancing the model's capability in capturing spatiotemporal dynamics at different granularities. Additionally, we introduce a Category Query Decoder that leverages learnable category query vectors to achieve explicit behavior category modeling and mapping. Without relying on any extra annotations, the proposed method achieves unified recognition of multi-species and multi-category behaviors, setting a new state-of-the-art on the Animal Kingdom dataset and demonstrating strong generalization ability on the Charades dataset. Nuoer Long, Yonghao Dang, Chengpeng Xiong, Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Chan-Tong Lam, Jianqin Yin, Peter H. N. de With, Yue Sun 0001 |
IEEE Trans. Multim. | 10 |
| 2025 | DisCoPatch: Taming Adversarially-Driven Batch Statistics for Improved Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection holds significant importance across many applications. While semantic and domain-shift OOD problems are well-studied, this work focuses on covariate shifts - subtle variations in the data distribution that can degrade machine learning performance. We hypothesize that detecting these subtle shifts can improve our understanding of in-distribution boundaries, ultimately improving OOD detection. In adversarial discriminators trained with Batch Normalization (BN), real and adversarial samples form distinct domains with unique batch statistics - a property we exploit for OOD detection. We introduce DisCoPatch, an unsupervised Adversarial Variational Autoencoder (VAE) framework that harnesses this mechanism. During inference, batches consist of patches from the same image, ensuring a consistent data distribution that allows the model to rely on batch statistics. DisCoPatch uses the VAE's suboptimal outputs (generated and reconstructed) as negative samples to train the discriminator, thereby improving its ability to delineate the boundary between in-distribution samples and covariate shifts. By tightening this boundary, DisCoPatch achieves state-of-the-art results in public OOD detection benchmarks. The proposed model not only excels in detecting covariate shifts, achieving 95.5% AUROC on ImageNet-1K(-C), but also outperforms all prior methods on public Near-OOD (95.0%) benchmarks. With a compact model size of 25MB, it achieves high OOD detection performance at notably lower latency than existing methods, making it an efficient and practical solution for real-world OOD detection applications. The code is available at github.com/caetas/DisCoPatch. Francisco Caetano, Christiaan G. A. Viviers, Luis Albert Zavala-Mondragón, Peter H. N. de With, Fons van der Sommen |
ICCV | 4 |
| 2025 | Learning to recognize correctly completed procedure steps in egocentric assembly videos through spatio-temporal modelingabstractProcedure step recognition (PSR) aims to identify all correctly completed steps and their sequential order in videos of procedural tasks. The existing state-of-the-art models rely solely on detecting assembly object states in individual video frames. By neglecting temporal features, model robustness and accuracy are limited, especially when objects are partially occluded. To overcome these limitations, we propose Spatio-Temporal Occlusion-Resilient Modeling for Procedure Step Recognition (STORM-PSR), a dual-stream framework for PSR that leverages both spatial and temporal features. The assembly state detection stream operates effectively with unobstructed views of the object, while the spatio-temporal stream captures both spatial and temporal features to recognize step completions even under partial occlusion. This stream includes a spatial encoder, pre-trained using a novel weakly supervised approach to capture meaningful spatial representations, and a transformer-based temporal encoder that learns how these spatial features relate over time. STORM-PSR is evaluated on the MECCANO and IndustReal datasets, reducing the average delay between actual and predicted assembly step completions by 11.2% and 26.1%, respectively, compared to prior methods. We demonstrate that this reduction in delay is driven by the spatio-temporal stream, which does not rely on unobstructed views of the object to infer completed steps. The code for STORM-PSR, along with the newly annotated MECCANO labels, is made publicly available at https://timschoonbeek.github.io/stormpsr . • Spatio-temporal features reduce prediction delay for procedure step recognition (PSR). • Weakly-supervised key-frame selection enables meaningful spatial features. • Key-clip aware sampling significantly improves training of the temporal encoder. • PSR annotations and benchmark on the MECCANO dataset to stimulate further research. Tim J. Schoonbeek, Shao-Hsuan Hung, Dan Lehman, Hans Onvlee, Jacek Kustra, Peter H. N. de With, Fons van der Sommen |
Comput. Vis. Image Underst. | 6 |
| 2025 | Will Transformers change gastrointestinal endoscopic image analysis? A comparative analysis between CNNs and Transformers, in terms of performance, robustness and generalizationabstractGastrointestinal endoscopic image analysis presents significant challenges, such as considerable variations in quality due to the challenging in-body imaging environment, the often-subtle nature of abnormalities with low interobserver agreement, and the need for real-time processing. These challenges pose strong requirements on the performance, generalization, robustness and complexity of deep learning-based techniques in such safety-critical applications. While Convolutional Neural Networks (CNNs) have been the go-to architecture for endoscopic image analysis, recent successes of the Transformer architecture in computer vision raise the possibility to update this conclusion. To this end, we evaluate and compare clinically relevant performance, generalization and robustness of state-of-the-art CNNs and Transformers for neoplasia detection in Barrett's esophagus. We have trained and validated several top-performing CNNs and Transformers on a total of 10,208 images (2,079 patients), and tested on a total of 7,118 images (998 patients) across multiple test sets, including a high-quality test set, two internal and two external generalization test sets, and a robustness test set. Furthermore, to expand the scope of the study, we have conducted the performance and robustness comparisons for colonic polyp segmentation (Kvasir-SEG) and angiodysplasia detection (Giana). The results obtained for featured models across a wide range of training set sizes demonstrate that Transformers achieve comparable performance as CNNs on various applications, show comparable or slightly improved generalization capabilities and offer equally strong resilience and robustness against common image corruptions and perturbations. These findings confirm the viability of the Transformer architecture, particularly suited to the dynamic nature of endoscopic video analysis, characterized by fluctuating image quality, appearance and equipment configurations in transition from hospital to hospital. The code is made publicly available at: https://github.com/BONS-AI-VCA-AMC/Endoscopy-CNNs-vs-Transformers. Carolus H. J. Kusters, Tim J. M. Jaspers, T. G. W. Boers, Martijn R. Jong, Jelmer Jukema, Kiki Fockens, Albert Jeroen de Groof, Jacques J. Bergman, Fons van der Sommen, Peter H. N. de With |
Medical Image Anal. | 10 |
| 2025 | Investigating and Improving Latent Density Segmentation Models for Aleatoric Uncertainty Quantification in Medical ImagingabstractData uncertainties, such as sensor noise, occlusions or limitations in the acquisition method can introduce irreducible ambiguities in images, which result in varying, yet plausible, semantic hypotheses. In Machine Learning, this ambiguity is commonly referred to as aleatoric uncertainty. In image segmentation, latent density models can be utilized to address this problem. The most popular approach is the Probabilistic U-Net (PU-Net), which uses latent Normal densities to optimize the conditional data log-likelihood Evidence Lower Bound. In this work, we demonstrate that the PU-Net latent space is severely sparse and heavily under-utilized. To address this, we introduce mutual information maximization and entropy-regularized Sinkhorn Divergence in the latent space to promote homogeneity across all latent dimensions, effectively improving gradient-descent updates and latent space informativeness. Our results show that by applying this on public datasets of various clinical segmentation problems, our proposed methodology receives up to 11% performance gains compared against preceding latent variable models for probabilistic segmentation on the Hungarian-Matched Intersection over Union. The results indicate that encouraging a homogeneous latent space significantly improves latent density modeling for medical image segmentation. M. M. Amaan Valiuddin, Christiaan G. A. Viviers, Ruud van Sloun, Peter H. N. de With, Fons van der Sommen |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Retaining Informative Latent Variables in Probabilistic SegmentationabstractConditional latent-variable models can successfully quantify annotation variability in segmentation. Training such models involves tuning the dimensionality of the latent space to optimally capture the inherent data ambiguity. Nevertheless, we discover after careful tuning, that the latent space does not always reflect this. In fact, some latent dimensions are completely neglected. For such segmentation models the latent dimensionality is often poorly motivated or based on computational constraints. In this paper, we offer an information-theoretic approach to optimally leverage all latent dimensions. We adapt and improve the Probabilistic U-Net to maximize the mutual information between the latent and output variables, leading to improved latent space properties and higher segmentation performance. M. M. Amaan Valiuddin, Christiaan G. A. Viviers, Ruud van Sloun, Peter H. N. de With, Fons van der Sommen |
ICASSP | 4 |
| 2024 | TeG: Temporal-Granularity Method for Anomaly Detection with Attention in Smart City SurveillanceabstractAnomaly detection in video surveillance has recently gained interest from the research community. Temporal duration of anomalies vary within video streams, leading to complications in learning the temporal dynamics of specific events. This paper presents a temporal-granularity method for an anomaly detection model (TeG) in real-world surveillance, combining spatio-temporal features at different time-scales. The TeG model employs multi-head cross-attention (MCA) blocks and multi-head self-attention (MSA) blocks for this purpose. Additionally, we extend the UCF-Crime dataset with new anomaly types relevant to Smart City research project. The TeG model is deployed and validated in a city surveillance system, achieving successful real-time results in industrial settings. Erkut Akdag, Egor Bondarev, Peter H. N. de With |
VCIP | 3 |
| 2024 | Automated Camera Calibration via Homography Estimation with GNNsabstractOver the past few decades, a significant rise of camera-based applications for traffic monitoring has occurred. Governments and local administrations are increasingly relying on the data collected from these cameras to enhance road safety and optimize traffic conditions. However, for effective data utilization, it is imperative to ensure accurate and automated calibration of the involved cameras. This paper proposes a novel approach to address this challenge by leveraging the topological structure of intersections.We propose a framework involving the generation of a set of synthetic intersection viewpoint images from a bird’seye-view image, framed as a graph of virtual cameras to model these images. Using the capabilities of Graph Neural Networks, we effectively learn the relationships within this graph, thereby facilitating the estimation of a homography matrix. This estimation leverages the neighbourhood representation for any real-world camera and is enhanced by exploiting multiple images instead of a single match. In turn, the homography matrix allows the retrieval of extrinsic calibration parameters. As a result, the proposed framework demonstrates superior performance on both synthetic datasets and real-world cameras, setting a new state-of-the-art benchmark. Giacomo D'Amicantonio, Egor Bondarev, Peter H. N. de With |
WACV | 3 |
| 2024 | IndustReal: A Dataset for Procedure Step Recognition Handling Execution Errors in Egocentric Videos in an Industrial-Like SettingabstractAlthough action recognition for procedural tasks has received notable attention, it has a fundamental flaw in that no measure of success for actions is provided. This limits the applicability of such systems especially within the industrial domain, since the outcome of procedural actions is often significantly more important than the mere execution. To address this limitation, we define the novel task of procedure step recognition (PSR), focusing on recognizing the correct completion and order of procedural steps. Alongside the new task, we also present the multi-modal IndustReal dataset. Unlike currently available datasets, IndustReal contains procedural errors (such as omissions) as well as execution errors. A significant part of these errors are exclusively present in the validation and test sets, making IndustReal suitable to evaluate robustness of algorithms to new, unseen mistakes. Additionally, to encourage reproducibility and allow for scalable approaches trained on synthetic data, the 3D models of all parts are publicly available. Annotations and benchmark performance are provided for action recognition and assembly state detection, as well as the new PSR task. IndustReal, along with the code and model weights, is available at: https://github.com/TimSchoonbeek/IndustReal. Tim J. Schoonbeek, Tim Houben, Hans Onvlee, Peter H. N. de With, Fons van der Sommen |
WACV | 4 |
| 2024 | Foundation models in gastrointestinal endoscopic AI: Impact of architecture, pre-training approach and data efficiencyabstractPre-training deep learning models with large data sets of natural images, such as ImageNet, has become the standard for endoscopic image analysis. This approach is generally superior to training from scratch, due to the scarcity of high-quality medical imagery and labels. However, it is still unknown whether the learned features on natural imagery provide an optimal starting point for the downstream medical endoscopic imaging tasks. Intuitively, pre-training with imagery closer to the target domain could lead to better-suited feature representations. This study evaluates whether leveraging in-domain pre-training in gastrointestinal endoscopic image analysis has potential benefits compared to pre-training on natural images. To this end, we present a dataset comprising of 5,014,174 gastrointestinal endoscopic images from eight different medical centers (GastroNet-5M), and exploit self-supervised learning with SimCLRv2, MoCov2 and DINO to learn relevant features for in-domain downstream tasks. The learned features are compared to features learned on natural images derived with multiple methods, and variable amounts of data and/or labels (e.g. Billion-scale semi-weakly supervised learning and supervised learning on ImageNet-21k). The effects of the evaluation is performed on five downstream data sets, particularly designed for a variety of gastrointestinal tasks, for example, GIANA for angiodyplsia detection and Kvasir-SEG for polyp segmentation. The findings indicate that self-supervised domain-specific pre-training, specifically using the DINO framework, results into better performing models compared to any supervised pre-training on natural images. On the ResNet50 and Vision-Transformer-small architectures, utilizing self-supervised in-domain pre-training with DINO leads to an average performance boost of 1.63% and 4.62%, respectively, on the downstream datasets. This improvement is measured against the best performance achieved through pre-training on natural images within any of the evaluated frameworks. Moreover, the in-domain pre-trained models also exhibit increased robustness against distortion perturbations (noise, contrast, blur, etc.), where the in-domain pre-trained ResNet50 and Vision-Transformer-small with DINO achieved on average 1.28% and 3.55% higher on the performance metrics, compared to the best performance found for pre-trained models on natural images. Overall, this study highlights the importance of in-domain pre-training for improving the generic nature, scalability and performance of deep learning for medical image analysis. The GastroNet-5M pre-trained weights are made publicly available in our repository: huggingface.co/tgwboers/GastroNet-5M_Pretrained_Weights. T. G. W. Boers, Kiki Fockens, Joost van der Putten, Tim J. M. Jaspers, Carolus H. J. Kusters, Jelmer Jukema, Martijn R. Jong, Maarten R. Struyvenberg, Jeroen de Groof, Jacques J. Bergman, Peter H. N. de With, Fons van der Sommen |
Medical Image Anal. | 11 |
| 2024 | Robustness evaluation of deep neural networks for endoscopic image analysis: Insights and strategiesabstractComputer-aided detection and diagnosis systems (CADe/CADx) in endoscopy are commonly trained using high-quality imagery, which is not representative for the heterogeneous input typically encountered in clinical practice. In endoscopy, the image quality heavily relies on both the skills and experience of the endoscopist and the specifications of the system used for screening. Factors such as poor illumination, motion blur, and specific post-processing settings can significantly alter the quality and general appearance of these images. This so-called domain gap between the data used for developing the system and the data it encounters after deployment, and the impact it has on the performance of deep neural networks (DNNs) supportive endoscopic CAD systems remains largely unexplored. As many of such systems, for e.g. polyp detection, are already being rolled out in clinical practice, this poses severe patient risks in particularly community hospitals, where both the imaging equipment and experience are subject to considerable variation. Therefore, this study aims to evaluate the impact of this domain gap on the clinical performance of CADe/CADx for various endoscopic applications. For this, we leverage two publicly available data sets (KVASIR-SEG and GIANA) and two in-house data sets. We investigate the performance of commonly-used DNN architectures under synthetic, clinically calibrated image degradations and on a prospectively collected dataset including 342 endoscopic images of lower subjective quality. Additionally, we assess the influence of DNN architecture and complexity, data augmentation, and pretraining techniques for improved robustness. The results reveal a considerable decline in performance of 11.6% (±1.5) as compared to the reference, within the clinically calibrated boundaries of image degradations. Nevertheless, employing more advanced DNN architectures and self-supervised in-domain pre-training effectively mitigate this drop to 7.7% (±2.03). Additionally, these enhancements yield the highest performance on the manually collected test set including images with lower subjective quality. By comprehensively assessing the robustness of popular DNN architectures and training strategies across multiple datasets, this study provides valuable insights into their performance and limitations for endoscopic applications. The findings highlight the importance of including robustness evaluation when developing DNNs for endoscopy applications and propose strategies to mitigate performance loss. Tim J. M. Jaspers, T. G. W. Boers, Carolus H. J. Kusters, Martijn R. Jong, Jelmer Jukema, Albert Jeroen de Groof, Jacques J. Bergman, Peter H. N. de With, Fons van der Sommen |
Medical Image Anal. | 8 |
| 2024 | Advancing 6-DoF Instrument Pose Estimation in Variable X-Ray Imaging GeometriesabstractAccurate 6-DoF pose estimation of surgical instruments during minimally invasive surgeries can substantially improve treatment strategies and eventual surgical outcome. Existing deep learning methods have achieved accurate results, but they require custom approaches for each object and laborious setup and training environments often stretching to extensive simulations, whilst lacking real-time computation. We propose a general-purpose approach of data acquisition for 6-DoF pose estimation tasks in X-ray systems, a novel and general purpose YOLOv5-6D pose architecture for accurate and fast object pose estimation and a complete method for surgical screw pose estimation under acquisition geometry consideration from a monocular cone-beam X-ray image. The proposed YOLOv5-6D pose model achieves competitive results on public benchmarks whilst being considerably faster at 42 FPS on GPU. In addition, the method generalizes across varying X-ray acquisition geometry and semantic image complexity to enable accurate pose estimation over different domains. Finally, the proposed approach is tested for bone-screw pose estimation for computer-aided guidance during spine surgeries. The model achieves a 92.41% by the 0.1·d ADD-S metric, demonstrating a promising approach for enhancing surgical precision and patient outcomes. The code for YOLOv5-6D is publicly available at https://github.com/cviviers/YOLOv5-6D-Pose. Christiaan G. A. Viviers, Lena Filatova, Maurice Termeer, Peter H. N. de With, Fons van der Sommen |
IEEE Trans. Image Process. | 4 |
| 2023 | Homography Estimation for Camera Calibration in Complex Topological ScenesabstractSurveillance videos and images are used for a broad set of applications, ranging from traffic analysis to crime detection. Extrinsic camera calibration data is important for most analysis applications. However, security cameras are susceptible to environmental conditions and small camera movements, resulting in a need for an automated re-calibration method that can account for these varying conditions. In this paper, we present an automated camera-calibration process leveraging a dictionary-based approach that does not require prior knowledge on any camera settings. The method consists of a custom implementation of a Spatial Transformer Network (STN) and a novel topological loss function. Experiments reveal that the proposed method improves the IoU metric by up to 12% w.r.t. a state-of-the-art model across five synthetic datasets and the World Cup 2014 dataset. Giacomo D'Amicantonio, Egor Bondarau, Peter H. N. de With |
IV | 3 |
| 2023 | Performance-Efficiency Comparisons of Channel Attention Modules for ResNetsabstractAbstract Attention modules can be added to neural network architectures to improve performance. This work presents an extensive comparison between several efficient attention modules for image classification and object detection, in addition to proposing a novel Attention Bias module with lower computational overhead. All measured attention modules have been efficiently re-implemented, which allows an objective comparison and evaluation of the relationship between accuracy and inference time. Our measurements show that single-image inference time increases far more (5–50%) than the increase in FLOPs suggests (0.2–3%) for a limited gain in accuracy, making computation cost an important selection criterion. Despite this increase in inference time, adding an attention module can outperform a deeper baseline ResNet in both speed and accuracy. Finally, we investigate the potential of adding attention modules to pretrained networks and show that fine-tuning is possible and superior to training from scratch. The choice of the best attention module strongly depends on the specific ResNet architecture, input resolution, batch size and inference framework. Sander Klomp, Rob G. J. Wijnhoven, Peter H. N. de With |
Neural Process. Lett. | 3 |
| 2022 | Floor-plan generation from noisy point cloudsabstractThis paper proposes a growing-based floor-plan generation method that creates the global layout of buildings from noisy point clouds obtained by a stereo camera. We introduce a PCA-based line-growing concept with a subsequent filtering step, which is able to robustly handle the high noise levels in input point clouds. Experimental results show that this method outperforms the state-of-the-art techniques in floor-plan generation. The average F1 score for building layouts has increased from 0.38 to 0.66 on our test dataset, compared to the previous best floor-plan generation method. Furthermore, the resulting floor plans are multiple thousands of times smaller in memory size than the input point clouds, while still preserving the main building structures. Egor Bondarev, Peter H. N. de With |
ICMV | 3 |
| 2022 | Critical Vehicle Detection for Intelligent Transportation SystemsabstractAn intelligent transportation system (ITS) is one of the core elements of smart cities, enhancing public safety and relieving traffic congestion. Detection and classification of critical vehicles, such as police cars and ambulances, passing through roadways form crucial use cases for ITS. This paper proposes a solution for detecting and classifying safety-critical vehicles on urban roadways using deep learning models. At present, a large-scale dataset for critical vehicles is not publicly available. The appearance scarcity of emergency vehicles and different coloring standards in various countries are significant challenges. To cope with the mentioned drawbacks and to address the unique requirements of our smart city project, we first generate a large-scale critical vehicle dataset, combining images retrieved from various sources with the support of the YOLO vehicle detection model. The classes of the generated dataset are: fire truck, police car, ambulance, military police car, dang erous truck, and standard vehicle. Second, we compare the performance of the Vision in Transformer (ViT) network against the traditional convolutional neural networks (CNNs) for the task of critical vehicle classification. Experimental results on our dataset reveal that the ViT-based solution reaches an average accuracy and recall of 99.39% and 99.34%, respectively. Erkut Akdag, Egor Bondarev, Peter H. N. de With |
VEHITS | 3 |
| 2022 | Depth estimation from a single SEM image using pixel-wise fine-tuning with multimodal dataabstractAbstract To support the ongoing size reduction in integrated circuits, the need for accurate depth measurements of on-chip structures becomes increasingly important. Unfortunately, present metrology tools do not offer a practical solution. In the semiconductor industry, critical dimension scanning electron microscopes (CD-SEMs) are predominantly used for 2D imaging at a local scale. The main objective of this work is to investigate whether sufficient 3D information is present in a single SEM image for accurate surface reconstruction of the device topology. In this work, we present a method that is able to produce depth maps from synthetic and experimental SEM images. We demonstrate that the proposed neural network architecture, together with a tailored training procedure, leads to accurate depth predictions. The training procedure includes a weakly supervised domain adaptation step, which is further referred to as pixel-wise fine-tuning. This step employs scatterometry data to address the ground-truth scarcity problem. We have tested this method first on a synthetic contact hole dataset, where a mean relative error smaller than 6.2% is achieved at realistic noise levels. Additionally, it is shown that this method is well suited for other important semiconductor metrics, such as top critical dimension (CD), bottom CD and sidewall angle. To the extent of our knowledge, we are the first to achieve accurate depth estimation results on real experimental data, by combining data from SEM and scatterometry measurements. An experiment on a dense line space dataset yields a mean relative error smaller than 1%. Tim Houben, Thomas Huisman, Maxim Pisarenco, Fons van der Sommen, Peter H. N. de With |
Mach. Vis. Appl. | 5 |
| 2022 | Toward Multilabel Image Retrieval for Remote SensingabstractThe availability of large-scale remote sensing (RS) data facilitates a wide range of applications, such as disaster management and urban planning. An approach for such problems is image retrieval, where, given a query image, the goal is to find the most relevant match from a database. Most RS literature has been focused on single-label retrieval, where we assume an image has a single label. The primary challenge in single-label RS retrieval is that performance in most datasets is saturated, and it has become difficult to compare the performance of different methods. In this work, we extend the major multilabel classification datasets to the multilabel retrieval problem. We also define protocols, provide evaluation metrics, and study the impact of commonly used loss functions and reranking methods for multilabel retrieval. To this end, a novel multilabel loss function and a reranking technique are proposed, which circumvent the challenges present in conventional single-label image retrieval. The developed loss function considers both class and feature similarity. The proposed reranking technique achieves high performance with computation cost that is well-suited for fast online retrieval. Raffaele Imbriaco, Clint Sebastian, Egor Bondarev, Peter H. N. de With |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Medical Instrument Segmentation in 3D US by Hybrid Constrained Semi-Supervised LearningabstractMedical instrument segmentation in 3D ultrasound is essential for image-guided intervention. However, to train a successful deep neural network for instrument segmentation, a large number of labeled images are required, which is expensive and time-consuming to obtain. In this article, we propose a semi-supervised learning (SSL) framework for instrument segmentation in 3D US, which requires much less annotation effort than the existing methods. To achieve the SSL learning, a Dual-UNet is proposed to segment the instrument. The Dual-UNet leverages unlabeled data using a novel hybrid loss function, consisting of uncertainty and contextual constraints. Specifically, the uncertainty constraints leverage the uncertainty estimation of the predictions of the UNet, and therefore improve the unlabeled information for SSL training. In addition, contextual constraints exploit the contextual information of the training images, which are used as the complementary information for voxel-wise uncertainty estimation. Extensive experiments on multiple ex-vivo and in-vivo datasets show that our proposed method achieves Dice score of about 68.6%-69.1% and the inference time of about 1 sec. per volume. These results are better than the state-of-the-art SSL methods and the inference time is comparable to the supervised approaches. Hongxu Yang, Caifeng Shan, R. Arthur Bouwman, Lukas R. C. Dekker, Alexander F. Kolen, Peter H. N. de With |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Noise Reduction in CT Using Learned Wavelet-Frame Shrinkage NetworksabstractEncoding-decoding (ED) CNNs have demonstrated state-of-the-art performance for noise reduction over the past years. This has triggered the pursuit of better understanding the inner workings of such architectures, which has led to the theory of deep convolutional framelets (TDCF), revealing important links between signal processing and CNNs. Specifically, the TDCF demonstrates that ReLU CNNs induce low-rankness, since these models often do not satisfy the necessary redundancy to achieve perfect reconstruction (PR). In contrast, this paper explores CNNs that do meet the PR conditions. We demonstrate that in these type of CNNs soft shrinkage and PR can be assumed. Furthermore, based on our explorations we propose the learned wavelet-frame shrinkage network, or LWFSN and its residual counterpart, the rLWFSN. The ED path of the (r)LWFSN complies with the PR conditions, while the shrinkage stage is based on the linear expansion of thresholds proposed Blu and Luisier. In addition, the LWFSN has only a fraction of the training parameters (<1%) of conventional CNNs, very small inference times, low memory footprint, while still achieving performance close to state-of-the-art alternatives, such as the tight frame (TF) U-Net and FBPConvNet, in low-dose CT denoising. Luis Albert Zavala-Mondragón, Peter M. J. Rongen, Javier Oliván Bescós, Peter H. N. de With, Fons van der Sommen |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Safe Fakes: Evaluating Face Anonymizers for Face DetectorsabstractSince the introduction of the GDPR and CCPA privacy legislation, both public and private facial image datasets are increasingly scrutinized. Several datasets have been taken offline completely and some have been anonymized. However, it is unclear how anonymization impacts face detection performance. To our knowledge, this paper presents the first empirical study on the effect of image anonymization on supervised training of face detectors. We compare conventional face anonymiz-ers with three state-of-the-art Generative Adversarial Network-based (GAN) methods, by training an off-the-shelf face detector on anonymized data. Our experiments investigate the suitability of anonymization methods for maintaining face detector performance, the effect of detectors overtraining on anonymization artefacts, dataset size for training an anonymizer, and the effect of training time of anonymization GANs. A final experiment investigates the correlation between common GAN evaluation metrics and the performance of a trained face detector. Although all tested anonymization methods lower the performance of trained face detectors, faces anonymized using GANs cause far smaller performance degradation than conventional methods. As the most important finding, the best-performing GAN, DeepPrivacy, removes identifiable faces for a face detector trained on anonymized data, resulting in a modest decrease from 91.0 to 88.3 mAP. In the last few years, there have been rapid improvements in realism of GAN-generated faces. We expect that further progression in GAN research will allow the use of Deep Fake technology for privacy-preserving Safe Fakes, without any performance degradation for training face detectors. Sander Klomp, Matthew van Rijn, Rob G. J. Wijnhoven, Cees Snoek, Peter H. N. de With |
FG | 5 |
| 2021 | Evaluating Self-Supervised Learning Methods for Downstream Classification of Neoplasia in Barrett's EsophagusabstractA major problem in applying machine learning for the medical domain is the scarcity of labeled data, which results in the demand for methods that enable high-quality models trained with little to no labels. Self-supervised learning methods present a plausible solution to this problem, enabling the use of large sets of unlabeled data for model pretraining. In this study, multiple of these methods and training strategies are employed on a large dataset of endoscopic images from the gastrointestinal tract (GastroNet). The suitability of these methods is assessed for an intra-domain downstream classification task on a small endoscopic dataset, involving neoplasia in Barrett’s esophagus. The classification performances are compared against pretraining on ImageNet and training from scratch. This yields promising results for domain-specific self-supervised methods, where super-resolution outperforms pretraining on ImageNet with a mean classification accuracy of 83.8% (cf. 79.2%). This implies that the large amounts of unlabeled data in hospitals could be employed in combination with self-supervised learning methods to improve models for downstream tasks. Stefan Cornelissen, Joost van der Putten, T. G. W. Boers, Jelmer Jukema, Kiki Fockens, Jacques J. Bergman, Fons van der Sommen, Peter H. N. de With |
ICIP | 8 |
| 2021 | Depth estimation from a single CD-SEM image using domain adaptation with multimodal dataabstractThere is a growing need for accurate depth measurements of on-chip structures, fueled by the ongoing size reduction of integrated circuits. However, current metrology methods do not offer a satisfactory solution. As Critical Dimension Scanning Electron Microscopes (CD-SEMs) are already being used for fast and local 2D imaging, it would be beneficial to leverage the 3D information hidden in these images. In this paper, we present a method that can predict depth maps from top-down CD-SEM images. We demonstrate that the proposed neural network architecture, together with a tailored training procedure, leads to accurate depth predictions on synthetic and real experimental data. Our training procedure includes a domain adaptation step, which utilizes data from a different modality (scatterometry), in the absence of ground truth data in the experimental CD-SEM domain. The mean relative error of the proposed method is smaller than 6.2% on a contact-hole dataset of synthetic CD-SEM images with realistic noise levels. Furthermore, we show that the method performs well in terms of important semiconductor metrics. To the extent of our knowledge, we are the first to achieve accurate depth estimation results on experimental data, by combining data from the aforementioned modalities. We achieve a mean relative error smaller than 1%. Thomas Huisman, Maxim Pisarenco, Peter H. N. de With |
ICMV | 3 |
| 2021 | Automatic image and text-based description for colorectal polyps using BASIC classificationabstractColorectal polyps (CRP) are precursor lesions of colorectal cancer (CRC). Correct identification of CRPs during in-vivo colonoscopy is supported by the endoscopist's expertise and medical classification models. A recent developed classification model is the Blue light imaging Adenoma Serrated International Classification (BASIC) which describes the differences between non-neoplastic and neoplastic lesions acquired with blue light imaging (BLI). Computer-aided detection (CADe) and diagnosis (CADx) systems are efficient at visually assisting with medical decisions but fall short at translating decisions into relevant clinical information. The communication between machine and medical expert is of crucial importance to improve diagnosis of CRP during in-vivo procedures. In this work, the combination of a polyp image classification model and a language model is proposed to develop a CADx system that automatically generates text comparable to the human language employed by endoscopists. The developed system generates equivalent sentences as the human-reference and describes CRP images acquired with white light (WL), blue light imaging (BLI) and linked color imaging (LCI). An image feature encoder and a BERT module are employed to build the AI model and an external test set is used to evaluate the results and compute the linguistic metrics. The experimental results show the construction of complete sentences with an established metric scores of BLEU-1 = 0.67, ROUGE-L = 0.83 and METEOR = 0.50. The developed CADx system for automatic CRP image captioning facilitates future advances towards automatic reporting and may help reduce time-consuming histology assessment. Roger Fonolla, Quirine E. W. van der Zander, Ramon-Michel Schreuder, Sharmila Subramaniam, Pradeep Bhandari, Ad A. M. Masclee, Erik J. Schoon, Fons van der Sommen, Peter H. N. de With |
Artif. Intell. Medicine | 9 |
| 2021 | Mask-MCNet: Tooth instance segmentation in 3D point clouds of intra-oral scansabstractComputational dentistry uses computerized methods and mathematical models for dental image analysis. One of the fundamental problems in computational dentistry is accurate tooth instance segmentation in high-resolution mesh data of intra-oral scans (IOS). This paper presents a new computational model based on deep neural networks, called Mask-MCNet, for end-to-end learning of tooth instance segmentation in 3D point cloud data of IOS. The proposed Mask-MCNet localizes each tooth instance by predicting its 3D bounding box and simultaneously segments the points that belong to each individual tooth instance. The proposed model processes the input raw 3D point cloud in its original spatial resolution without employing a voxelization or down-sampling technique. Such a characteristic preserves the finely detailed context in data like fine curvatures in the border between adjacent teeth and leads to a highly accurate segmentation as required for clinical practice (e.g. orthodontic planning). The experiments show that the Mask-MCNet outperforms state-of-the-art models by achieving 98% Intersection over Union (IoU) score on tooth instance segmentation which is very close to human expert performance. Farhad G. Zanjani, Arash Pourtaherian, Svitlana Zinger, David Anssari Moin, Frank Claessen, Teo Cherici, Sarah Parinussa, Peter H. N. de With |
Neurocomputing | 8 |
| 2021 | Efficient and Robust Instrument Segmentation in 3D Ultrasound Using Patch-of-Interest-FuseNet with Hybrid LossabstractInstrument segmentation plays a vital role in 3D ultrasound (US) guided cardiac intervention. Efficient and accurate segmentation during the operation is highly desired since it can facilitate the operation, reduce the operational complexity, and therefore improve the outcome. Nevertheless, current image-based instrument segmentation methods are not efficient nor accurate enough for clinical usage. Lately, fully convolutional neural networks (FCNs), including 2D and 3D FCNs, have been used in different volumetric segmentation tasks. However, 2D FCN cannot exploit the 3D contextual information in the volumetric data, while 3D FCN requires high computation cost and a large amount of training data. Moreover, with limited computation resources, 3D FCN is commonly applied with a patch-based strategy, which is therefore not efficient for clinical applications. To address these, we propose a POI-FuseNet, which consists of a patch-of-interest (POI) selector and a FuseNet. The POI selector can efficiently select the interested regions containing the instrument, while FuseNet can make use of 2D and 3D FCN features to hierarchically exploit contextual information. Furthermore, we propose a hybrid loss function, which consists of a contextual loss and a class-balanced focal loss, to improve the segmentation performance of the network. With the collected challenging ex-vivo dataset on RF-ablation catheter, our method achieved a Dice score of 70.5%, superior to the state-of-the-art methods. In addition, based on the pre-trained model from ex-vivo dataset, our method can be adapted to the in-vivo dataset on guidewire and achieves a Dice score of 66.5% for a different cardiac operation. More crucially, with POI-based strategy, segmentation efficiency is reduced to around 1.3 seconds per volume, which shows the proposed method is promising for clinical use. Hongxu Yang, Caifeng Shan, R. Arthur Bouwman, Alexander F. Kolen, Peter H. N. de With |
Medical Image Anal. | 5 |
| 2021 | Maritime vessel re-identification: novel VR-VCA dataset and a multi-branch architecture MVR-netabstractAbstract Maritime vessel re-identification (re-ID) is a computer vision task of vessel identity matching across disjoint camera views. Prominent applications of vessel re-ID exist in the fields of surveillance and maritime traffic flow analysis. However, the field suffers from the absence of a large-scale dataset that enables training of deep learning models. In this study, we present a new dataset that includes 4614 images of 729 vessels along with 5-bin orientation and 8-class vessel-type annotations to promote further research. A second contribution of this study is the baseline re-ID analysis of our new dataset. Performances of 10 recent deep learning architectures are quantitatively compared to reveal the best practices. Lastly, we propose a novel multi-branch deep learning architecture, Maritime Vessel Re-ID network (MVR-net), to address the challenging problem of vessel re-ID. Evaluation of our approach on the new dataset yields 74.5% mAP and 77.9% Rank-1 score, providing a performance increase of 5.7% mAP and 5.0% Rank-1 over the best-performing baseline. MVR-net also outperforms the PRN (a pioneering vehicle re-ID network), by 2.9% and 4.3% higher mAP and Rank-1, respectively. Amir Ghahremani, Tunç Alkanat, Egor Bondarev, Peter H. N. de With |
Mach. Vis. Appl. | 4 |
| 2021 | Image Noise Reduction Based on a Fixed Wavelet Frame and CNNs Applied to CTabstractRadiation exposure in CT imaging leads to increased patient risk. This motivates the pursuit of reduced-dose scanning protocols, in which noise reduction processing is indispensable to warrant clinically acceptable image quality. Convolutional Neural Networks (CNNs) have received significant attention as an alternative for conventional noise reduction and are able to achieve state-of-the art results. However, the internal signal processing in such networks is often unknown, leading to sub-optimal network architectures. The need for better signal preservation and more transparency motivates the use of Wavelet Shrinkage Networks (WSNs), in which the Encoding-Decoding (ED) path is the fixed wavelet frame known as Overcomplete Haar Wavelet Transform (OHWT) and the noise reduction stage is data-driven. In this work, we considerably extend the WSN framework by focusing on three main improvements. First, we simplify the computation of the OHWT that can be easily reproduced. Second, we update the architecture of the shrinkage stage by further incorporating knowledge of conventional wavelet shrinkage methods. Finally, we extensively test its performance and generalization, by comparing it with the RED and FBPConvNet CNNs. Our results show that the proposed architecture achieves similar performance to the reference in terms of MSSIM (0.667, 0.662 and 0.657 for DHSN2, FBPConvNet and RED, respectively) and achieves excellent quality when visualizing patches of clinically important structures. Furthermore, we demonstrate the enhanced generalization and further advantages of the signal flow, by showing two additional potential applications, in which the new DHSN2 is used as regularizer: (1) iterative reconstruction and (2) ground-truth free training of the proposed noise reduction architecture. The presented results prove that the tight integration of signal processing and deep learning leads to simpler models with improved generalization. Luis Albert Zavala-Mondragón, Peter H. N. de With, Fons van der Sommen |
IEEE Trans. Image Process. | 2 |
| 2021 | Infant Facial Expression Analysis: Towards a Real-Time Video Monitoring System Using R-CNN and HMMabstractThe manual monitoring of young infants suffering from diseases like reflux is significant, since infants can hardly articulate their feelings. In this work, we propose a video-based infant monitoring system for the analysis of infant expressions and states, approaching real-time performance. The expressions of interest consist of discomfort, unhappy, joy and neutral, whereas states include sleep, pacifier and open mouth. Benefiting from the expression analysis, the discomfort moments can also be used and correlated with a symptom-related disease, such as a reflux measurement for the diagnosis of gastroesophageal reflux. The system consists of three components: infant expressions and states detection, object tracking and detection compensation. The proposed system is based on combining expression detection using Fast R-CNN with a compensated detection using analyzing information from the previous frame and utilizing a Hidden Markov Model. The experimental results show a mean average precision of 81.9% and 84.8% for 4 infant expressions and 3 states evaluated with both clinical and daily datasets. Meanwhile, the average precision for discomfort detection achieves up to 90%. Cheng Li 0042, Arash Pourtaherian, Lonneke van Onzenoort, Walther E. Tjon a Ten, Peter H. N. de With |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Inter-Center Cross-Validation and Finetuning without Patient Data Sharing for Predicting Transcatheter Aortic Valve Implantation OutcomeabstractTranscatheter aortic valve implantation (TAVI) is the routine treatment worldwide for aortic valve stenosis in low-to high-risk patients. Assessing patient risk is essential to identify the most suitable candidates that could benefit from the procedure. Despite the broad use of statistical predictors in patient selection, current machine learning predictors have only been validated on retrospective data collected in single centers. Further, external validation is needed to assess the improvement in accuracy, which is offered by machine learning and deep learning techniques. In this study, we propose a finetuning approach for deep learning models by performing an inter-center cross-validation and finetuning technique, in order to improve the cross-validation accuracy results. We aimed to overcome data exchange and policy-related issues of two medical centers with a dedicated protocol, exploiting the exchange of deep learning models, data processing and validation steps which does not require any patient data sharing. The finetuning is based on the other center's data for further training of the initial model. After finetuning the model, we obtain an average AUC improvement of 13% and 7% with respect to the initial models. This research demonstrates that the predicting capabilities of deep learning models can be extended to and cross-validated with other centers, independent of limitations in data-sharing policies. Moreover, the study shows that finetuning can be exploited to considerably improve the accuracy of the prediction models. Ricardo R. Lopes, Marco Mamprin, Jo M. Zelis, Pim A. L. Tonino, Martijn S. van Mourik, Marije M. Vis, Svitlana Zinger, Bas A. J. M. de Mol, Peter H. N. de With, Henk A. Marquering |
CBMS | 9 |
| 2020 | Deep Q-Network-Driven Catheter Segmentation in 3D US by Hybrid Constrained Semi-supervised Learning and Dual-UNet
Hongxu Yang, Caifeng Shan, Alexander F. Kolen, Peter H. N. de With |
MICCAI (1) | 4 |
| 2020 | Multi-stage domain-specific pretraining for improved detection and localization of Barrett's neoplasia: A comprehensive clinically validated studyabstractPatients suffering from Barrett's Esophagus (BE) are at an increased risk of developing esophageal adenocarcinoma and early detection is crucial for a good prognosis. To aid the endoscopists with the early detection for this preliminary stage of esophageal cancer, this work concentrates on the development and extensive evaluation of a state-of-the-art computer-aided classification and localization algorithm for dysplastic lesions in BE. To this end, we have employed a large-scale endoscopic data set, consisting of 494,355 images, in combination with a novel semi-supervised learning algorithm to pretrain several instances of the proposed neural network architecture. Next, several Barrett-specific data sets that are increasingly closer to the target domain with significantly more data compared to other related work, were used in a multi-stage transfer learning strategy. Additionally, the algorithm was evaluated on two prospectively gathered external test sets and compared against 53 medical professionals. Finally, the model was also evaluated in a live setting without interfering with the current biopsy protocol. Results from the performed experiments show that the proposed model improves on the state-of-the-art on all measured metrics. More specifically, compared to the best performing state-of-the-art model, the specificity is improved by more than 20% points while simultaneously preserving high sensitivity and reducing the false positive rate substantially. Our algorithm yields similar scores on the localization metrics, where the intersection of all experts is correctly indicated in approximately 92% of the cases. Furthermore, the live pilot study shows great performance in a clinical setting with a patient level accuracy, sensitivity, and specificity of 90%. Finally, the proposed algorithm outperforms each individual medical expert by at least 5% and the average assessor by more than 10% over all assessor groups with respect to accuracy. Joost van der Putten, Jeroen de Groof, Maarten R. Struyvenberg, T. G. W. Boers, Kiki Fockens, Wouter L. Curvers, Erik J. Schoon, Jacques J. Bergman, Fons van der Sommen, Peter H. N. de With |
Artif. Intell. Medicine | 10 |
| 2020 | Modeling clinical assessor intervariability using deep hypersphere encoder-decoder networksabstractAbstract In medical imaging, a proper gold-standard ground truth as, e.g., annotated segmentations by assessors or experts is lacking or only scarcely available and suffers from large intervariability in those segmentations. Most state-of-the-art segmentation models do not take inter-observer variability into account and are fully deterministic in nature. In this work, we propose hypersphere encoder–decoder networks in combination with dynamic leaky ReLUs, as a new method to explicitly incorporate inter-observer variability into a segmentation model. With this model, we can then generate multiple proposals based on the inter-observer agreement. As a result, the output segmentations of the proposed model can be tuned to typical margins inherent to the ambiguity in the data. For experimental validation, we provide a proof of concept on a toy data set as well as show improved segmentation results on two medical data sets. The proposed method has several advantages over current state-of-the-art segmentation models such as interpretability in the uncertainty of segmentation borders. Experiments with a medical localization problem show that it offers improved biopsy localizations, which are on average 12% closer to the optimal biopsy location. Joost van der Putten, Fons van der Sommen, Jeroen de Groof, Maarten R. Struyvenberg, Svitlana Zinger, Wouter L. Curvers, Erik J. Schoon, Jacques J. Bergman, Peter H. N. de With |
Neural Comput. Appl. | 9 |
| 2019 | Privacy Protection in Street-View Panoramas Using Depth and Multi-View ImageryabstractThe current paradigm in privacy protection in street-view images is to detect and blur sensitive information. In this paper, we propose a framework that is an alternative to blurring, which automatically removes and inpaints moving objects (e.g. pedestrians, vehicles) in street-view imagery. We propose a novel moving object segmentation algorithm exploiting consistencies in depth across multiple street-view images that are later combined with the results of a segmentation network. The detected moving objects are removed and inpainted with information from other views, to obtain a realistic output image such that the moving object is not visible anymore. We evaluate our results on a dataset of 1000 images to obtain a peak noise-to-signal ratio (PSNR) and L 1 loss of 27.2 dB and 2.5%, respectively. To assess overall quality, we also report the results of a survey conducted on 35 professionals, asked to visually inspect the images whether object removal and inpainting had taken place. The inpainting dataset will be made publicly available for scientific benchmarking purposes at https://research.cyclomedia.com/. Ries Uittenbogaard, Clint Sebastian, Julien A. Vijverberg, Bas Boom, Dariu Gavrila, Peter H. N. de With |
CVPR | 6 |
| 2019 | Image Features for Automated Colorectal Polyp Classification Based on Clinical Prediction ModelsabstractAccurate endoscopic differentiation on resection of colorectal polyps (CRPs) (resect-discard or diagnose-leave strategies) increases cost-efficiency and reduces patient risk. We aim to develop a classification algorithm for automated differentiation of CRPs, by following the validated clinical Work-group serrAted polypS and Polyposis (WASP) classification scheme. Quantitative image features are investigated for each individual WASP criterion and classification is performed by conventional SVM. The technical WASP model results in areas under the curve of 0.87-0.95 and accuracies of 78-89%. Predicting polyp histology using model-based learning out-performs medical experts (accuracy, 87-93% vs 86 87%). Direct classification predicts more premalignant polyps-as being benign, compared to the automated WASP scheme. These errors do not occur when including ROC characteristics to the WASP model. The proposed WASP model is the first automated system, competing with medical expert classification. Michelle C. A. van Grinsven, Thom Scheeve, Ramon-Michel Schreuder, Fons van der Sommen, Erik J. Schoon, Peter H. N. de With |
ICIP | 6 |
| 2019 | Informative Frame Classification of Endoscopic Videos Using Convolutional Neural Networks and Hidden Markov ModelsabstractThe goal of endoscopic analysis is to find abnormal lesions and determine further therapy from the obtained information. For example, in case of Barrett's esophagus, the objective of endoscopy is to timely detect dysplastic lesions, before endoscopic resection is no longer possible. However, the procedure produces a variety of non-informative frames and lesions can be missed due to poor video quality. Especially when analyzing entire endoscopic videos made by non-expert endoscopists, informative frame classification is crucial to e.g. video quality grading. This analysis involves classification problems such as polyp detection or dysplasia detection in Barrett's Esophagus. This work concentrates on the design of an automated indication of informativeness of video frames. We propose an algorithm consisting of state-of-the-art deep learning techniques, to initialize frame-based classification, followed by a hidden Markov model to incorporate temporal information and control consistent decision making. Results from the performed experiments show that the proposed model improves on the state-of-the-art with an F1-score of 91%, and a substantial increase in sensitivity of 10%, thereby indicating improved labeling consistency. Additionally, the algorithm is capable of processing 261 frames per second, which is multiple times faster compared to other informative frame classification algorithms, thus enabling real-time computation. Joost van der Putten, Jeroen de Groof, Fons van der Sommen, Maarten R. Struyvenberg, Svitlana Zinger, Wouter L. Curvers, Erik J. Schoon, Jacques J. Bergman, Peter H. N. de With |
ICIP | 9 |
| 2019 | Efficient Catheter Segmentation in 3D Cardiac Ultrasound using Slice-Based FCN With Deep Supervision and F-Score LossabstractFast and accurate catheter segmentation in 3D ultrasound (US) can improve the outcome and efficiency of cardiac interventions. In this paper, we propose an efficient catheter segmentation method based on a fully convolutional neural network (FCN). The FCN is based on a pre-trained VGG-16 model, which processes the 3D US volumes slice by slice. To enhance its performance, we modify its structure by skipping connections under a deep supervision structure, which is learned with an F-score loss function. Our method can exploit more contextual information and increase the detection of catheter-like voxels. We collected a challenging ex-vivo dataset (92 3D US images) from porcine hearts with an RF-ablation catheter inside. Our experiments on this dataset show that the proposed method achieves a segmentation performance with an F2score of 65.2% with a highly efficient inference around 1.1 sec. per volume. Hongxu Yang, Caifeng Shan, Alexander F. Kolen, Peter H. N. de With |
ICIP | 4 |
| 2019 | Automated Catheter Localization in Volumetric Ultrasound Using 3D Patch-Wise U-Net with Focal Lossabstract3D ultrasound (US) imaging has become an attractive option for image-guided interventions. Fast and accurate catheter localization in 3D cardiac US can improve the outcome and efficiency of the cardiac interventions. In this paper, we propose a catheter localization method for 3D cardiac US using the patch-wise semantic segmentation with model fitting. Our 3D U-Net is trained with the focal loss of cross-entropy, which makes the network to focus more on samples that are difficult to classify. Moreover, we adopt a dense sampling strategy to overcome the extremely imbalanced catheter occupation in the 3D US data. Extensive experiments on our challenging ex-vivo dataset show that the proposed method achieves an F-1 score of 65.1% for catheter segmentation, outperforming the state-of-the-art methods. With this, our method can localize RF-ablation catheters with an average error of 1.28 mm. Hongxu Yang, Caifeng Shan, Alexander F. Kolen, Peter H. N. de With |
ICIP | 4 |
| 2019 | Improving open-set person re-identification by statistics-driven gallery refinementabstractPerson re-identification (re-ID) is a valuable tool for multi-camera tracking of persons. Up till now, research on person re-ID has mainly focused on the closed-set case, where a given query is assumed to always have a correct match in the gallery set, which does not hold for practical scenarios. In this study, we explore the open-set person re-ID problem with queries not always included in the gallery set. First, we convert the popular closed-set person re-ID datasets into the open-set scenario. Second, we compare the performances of six state-of-the-art closed-set person re-ID methods under open-set conditions. Third, we investigate the impact of a simple and fast statistics-driven gallery refinement approach on the open-set person re-ID performance. Extensive experimental evaluations show that, gallery refinement increases the performance of existing methods in the low false-accept rate (FAR) region, while simultaneously reducing the computational demands of retrieval. Results show an average detection and identification rate (DIR) increase of 7.91% and 3.31% on the DukeMTMC-reID and Market1501 datasets, respectively, for an FAR of 1%. Tunç Alkanat, Egor Bondarev, Peter H. N. de With |
ICMV | 3 |
| 2019 | Transferring from ex-vivo to in-vivo: Instrument Localization in 3D Cardiac Ultrasound Using Pyramid-UNet with Hybrid Loss
Hongxu Yang, Caifeng Shan, Tao Tan 0002, Alexander F. Kolen, Peter H. N. de With |
MICCAI (5) | 5 |
| 2019 | Mask-MCNet: Instance Segmentation in 3D Point Cloud of Intra-oral Scans
Farhad G. Zanjani, David Anssari Moin, Frank Claessen, Teo Cherici, Sarah Parinussa, Arash Pourtaherian, Svitlana Zinger, Peter H. N. de With |
MICCAI (5) | 8 |
| 2019 | Video-based discomfort detection for infants
Yue Sun 0001, Caifeng Shan, Tao Tan 0002, Xi Long 0001, Arash Pourtaherian, Svitlana Zinger, Peter H. N. de With |
Mach. Vis. Appl. | 7 |
| 2018 | Multispectral Image Analysis for Patient Tissue Tracking During Complex InterventionsabstractDuring complex interventions, patient tracking is needed for optimal motion compensation, in order to guide the physician in a minimally invasive way. Nowadays, optical tracking systems are used for tracking of markers. Despite the unobtrusiveness of this technology, the approach is cumbersome because it requires manual placement of markers, which can alter during the operations due to the presence of liquids. To improve the clinical workflow, a new feature tracking algorithm is designed, involving feature detection and tracking without optical markers. The new markers are created with the multispectral imaging. Maximally stable extremal regions (MSER) and Speeded Up Robust Feature (SURF) methods are applied to design features and track natural landmarks, e.g. moles and veins. Both methods are tested and compared in accuracy with the mean shift tracking method. SURF reaches the highest accuracy of 0.257 pixels for images at 430 nm and 0.562 pixels for images at 970 nm. This study shows that incorporating multispectral imaging in the surgical scenario leads to an attractive benefit minimizing the risk of marker obstruction and displacement. Francesca Manni, Marco Mamprin, Svitlana Zinger, Caifeng Shan, Ronald Holthuizen, Peter H. N. de With |
ICIP | 6 |
| 2018 | Automatic Detection of Early Esophageal Cancer with CNNS Using Transfer LearningabstractThe incidence of Esophageal Adenocarcinoma (EAC), a form of esophageal cancer, has rapidly increased in recent years. Dysplastic tissue can be removed endoscopically at an early stage, and since survival chances of patients are limited at later stages of the disease, early detection is of key impor- tance. Recently, several CAD systems for HD endoscopic images have been proposed, but these are computationally expensive, making them unfit for clinical use requiring real- time analysis. In this paper, we present a novel approach for early esophageal cancer detection using Transfer Learning with CNNs. Given the small amount of annotated data, CNN Codes are applied, where intermediate layers of the net- work are used as features for conventional classifiers. Various classifiers are combined with four of the most widely-used networks. Additionally, sliding windows are used to obtain a coarse-grained annotation indicating any possible cancerous regions. This approach outperforms the current state-of-the-art with a frame-based AUC of 0.92, while allowing both near real-time prediction and annotation at 2 fps, in a MATLAB-based framework. Sjors van Riel, Fons van der Sommen, Svitlana Zinger, Erik J. Schoon, Peter H. N. de With |
ICIP | 5 |
| 2018 | Catheter Detection in 3D Ultrasound Using Triplanar-Based Convolutional Neural Networksabstract3D Ultrasound (US) image-based catheter detection can potentially decrease the cost on extra equipment and training. Meanwhile, accurate catheter detection enables to decrease the operation duration and improves its outcome. In this paper, we propose a catheter detection method based on convolutional neural networks (CNNs) in 3D US. Voxels in US images are classified as catheter (or not) using triplanar-based CNNs. Our proposed CNN employs two-stage training with weighted loss function, which can cope with highly imbalanced training data and improves classification accuracy. When compared to state-of-the-art handcrafted features on ex-vivo datasets, our proposed method improves the F2-score with at least 31%. Based on classified volumes, the catheters are localized with an average position error of smaller than 3 voxels in the examined datasets, indicating that catheters are always detected in noisy and low-resolution images. Hongxu Yang, Caifeng Shan, Alexander F. Kolen, Peter H. N. de With |
ICIP | 4 |
| 2018 | Towards multi-class detection: a self-learning approach to reduce inter-class noise from training datasetabstractThis paper proposes a novel self-learning framework, which converts a noisy, pre-labeled multi-class object dataset into a purified multi-class object dataset with object bounding-box annotations, by iteratively removing noise samples from the low-quality dataset, which may contain a high level of inter-class noise samples. The framework iteratively purifies the noisy training datasets for each class and updates the classification model for multiple classes. The procedure starts with a generic single-class object model which changes to a multi-class model in an iterative procedure of which the F-1 score is evaluated to reach a sufficiently high score. The proposed framework is based on learning the used models with CNNs. As a result, we obtain a purified multi-class dataset and as a spin-off, the updated multi-class object model. The proposed framework is evaluated on maritime surveillance, where vessels need to be classified into eight different types. The experimental results on the evaluation dataset show that the proposed framework improves the F-1 score approximately by 5% and 25% at the end of the third iteration, while the initial training datasets contain 40% and 60% inter-class noise samples (erroneously classified labels of vessels and without annotations), respectively. Additionally, the recall rate increases nearly by 38% (for the more challenging 60% inter-class noise case), while the mean Average Precision (mAP) rate remains stable. Amir Ghahremani, Egor Bondarev, Peter H. N. de With |
ICMV | 3 |
| 2018 | Towards accurate camera geopositioning by image matchingabstractIn this work, we present a camera geopositioning system based on matching a query image against a database with panoramic images. For matching, our system uses memory vectors aggregated from global image descriptors based on convolutional features to facilitate fast searching in the database. To speed up searching, a clustering algorithm is used to balance geographical positioning and computation time. We refine the obtained position from the query image using a new outlier removal algorithm. The matching of the query image is obtained with a recall@5 larger than 90% for panorama-to-panorama matching. We cluster available panoramas from geographically adjacent locations into a single compact representation and observe computational gains of approximately 50% at the cost of only a small (approximately 3%) recall loss. Finally, we present a coordinate estimation algorithm that reduces the median geopositioning error by up to 20%. Raffaele Imbriaco, Clint Sebastian, Egor Bondarev, Peter H. N. de With |
ICMV | 4 |
| 2018 | Enhanced face alignment using an unsupervised roll estimation initializationabstractWe propose a novel and efficient initialization method for generalized facial landmark localization with an unsupervised roll-angle estimation based on B-spline models. We first show that the roll angle is crucial for an accurate landmark localization. Therefore, we develop an unsupervised roll-angle estimation by adopting a joint 1st -order B-spline model, which is robust to intensity variations and generic for application to various face detectors. The method consists of three steps. First, the scaled-normalized Laplacian of Gaussian operator is applied to a bounding box generated by a face detector for extracting facial feature segments. Second, a joint 1 st -order B-spline model is fitted to the extracted facial feature segments, using an iterative optimization method. Finally, the roll angle is estimated through the aligned segments. We evaluate four state-of-the-art landmark localization schemes with the proposed roll-angle estimation initialization in the benchmark dataset. The proposed method boosts the performance of landmark localization in general, especially for cases with large head pose. Moreover, the proposed unsupervised roll-angle estimation method outperforms the standard supervised methods, such as random forest and support vector regression by 41.6% and 47.2%, respectively. Cheng Li 0042, Arash Pourtaherian, Walther E. Tjon a Ten, Peter H. N. de With |
ICMV | 4 |
| 2018 | Conditional Transfer with Dense Residual Attention: Synthesizing traffic signs from street-view imageryabstractObject detection and classification of traffic signs in street-view imagery is an essential element for asset management, map making and autonomous driving. However, some traffic signs occur rarely and consequently, they are difficult to recognize automatically. To improve the detection and classification rates, we propose to generate images of traffic signs, which are then used to train a detector/classifier. In this research, we present an end-to-end framework that generates a realistic image of a traffic sign from a given image of a traffic sign and a pictogram of the target class. We propose a residual attention mechanism with dense concatenation called Dense Residual Attention, that preserves the background information while transferring the object information. We also propose to utilize multi-scale discriminators, so that the smaller scales of the output guide the higher resolution output. We have performed detection and classification tests across a large number of traffic sign classes, by training the detector using the combination of real and generated data. The newly trained model reduces the number of false positives by 1.2 - 1.5% at 99% recall in the detection tests and an absolute improvement of 4.65% (top-l accuracy) in the classification tests. Ries Uittenbogaard, Clint Sebastian, Julien A. Vijverberg, Bas Boom, Peter H. N. de With |
ICPR | 5 |
| 2018 | Deep Convolutional Gaussian Mixture Model for Stain-Color Normalization of Histopathological Images
Farhad G. Zanjani, Svitlana Zinger, Peter H. N. de With |
MICCAI (2) | 3 |
| 2017 | Improved Barrett's Cancer Detection in Volumetric Laser Endomicroscopy Scans Using Multiple-Frame VotingabstractThis paper explores the feasibility of using multiframe analysis to increase the classification performance of machine learning methods for cancer detection in Volumetric Laser Endomicroscopy (VLE). VLE is a novel and promising modality for the detection of neoplasia in patients with Baretts Esophagus (BE). It produces hundreds of high-resolution, cross-sectional images of the esophagus and offers considerable advantages compared to current methods. While some recent studies have proposed cancer detection algorithms for single VLE frames, the study described in this paper is the first to make use of VLE volumes for the differentiation between dysplastic and non-dysplastic tissue. We explore the use of various voting schemes for a broad range of features and classification methods. Our results demonstrate that multi-frame analysis leads to superior performance, irrespective of the chosen feature-classifier combination. By using multi-frame analysis with straightforward voting methods, the Area Under the receiver operating Curve (AUC) is increased by an average of over 12% compared to using single VLE frames. When only considering methods that achieve expert performance or higher (AUC≥0.81), an even larger performance improvement of up to 16.9% is observed. Furthermore, with many feature/classifier combinations showing AUC values ranging from 0.90 to 0.98, our experiments indicate that computeraided methods can considerably outperform medical experts, who demonstrate an AUC of 0.81 using a recently proposed clinical prediction model. Alexandros Rikos, Fons van der Sommen, Anne-Fré Swager, Svitlana Zinger, Erik J. Schoon, Wouter L. Curvers, Jacques J. Bergman, Peter H. N. de With |
CBMS | 8 |
| 2017 | Improving Needle Detection in 3D Ultrasound Using Orthogonal-Plane Convolutional Networks
Arash Pourtaherian, Farhad G. Zanjani, Svitlana Zinger, Nenad Mihajlovic, Gary C. Ng, Hendrikus H. M. Korsten, Peter H. N. de With |
MICCAI (2) | 7 |
| 2017 | Medical Instrument Detection in 3-Dimensional Ultrasound Data VolumesabstractUltrasound-guided medical interventions are broadly applied in diagnostics and therapy, e.g., regional anesthesia or ablation. A guided intervention using 2-D ultrasound is challenging due to the poor instrument visibility, limited field of view, and the multi-fold coordination of the medical instrument and ultrasound plane. Recent 3-D ultrasound transducers can improve the quality of the image-guided intervention if an automated detection of the needle is used. In this paper, we present a novel method for detecting medical instruments in 3-D ultrasound data that is solely based on image processing techniques and validated on various ex vivo and in vivo data sets. In the proposed procedure, the physician is placing the 3-D transducer at the desired position, and the image processing will automatically detect the best instrument view, so that the physician can entirely focus on the intervention. Our method is based on the classification of instrument voxels using volumetric structure directions and robust approximation of the primary tool axis. A novel normalization method is proposed for the shape and intensity consistency of instruments to improve the detection. Moreover, a novel 3-D Gabor wavelet transformation is introduced and optimally designed for revealing the instrument voxels in the volume, while remaining generic to several medical instruments and transducer types. Experiments on diverse data sets, including in vivo data from patients, show that for a given transducer and an instrument type, high detection accuracies are achieved with position errors smaller than the instrument diameter in the 0.5-1.5-mm range on average. Arash Pourtaherian, Harm J. Scholten, Lieneke Kusters, Svitlana Zinger, Nenad Mihajlovic, Alexander F. Kolen, Fei Zuo, Gary C. Ng, Hendrikus H. M. Korsten, Peter H. N. de With |
IEEE Trans. Medical Imaging | 10 |
| 2016 | R3P: Real-time RGB-D Registration Pipeline
Hani Javan Hemmat, Egor Bondarev, Peter H. N. de With |
ACIVS | 3 |
| 2016 | Dense-Hog-based 3D face tracking for infant pain monitoringabstractThis paper presents a new algorithm for 3D face tracking intended for clinical infant pain monitoring under challenging conditions. The algorithm uses a cylinder head model and head pose recovery by alignment of dynamically extracted templates based on dense-HOG features. The algorithm is motivated from the application context and compared with a variant based on intensities. The paper reports experimental results on videos of moving infants in hospital who are relaxed or in pain. Results show good short-term tracking behavior for poses up to 50 degrees from upright-frontal, with significantly higher accuracy resulting from the use of dense-HOG features. Walther E. Tjon a Ten, Peter H. N. de With |
ICIP | 3 |
| 2016 | Dense-HOG-based drift-reduced 3D face tracking for infant pain monitoringabstractThis paper presents a new algorithm for 3D face tracking intended for clinical infant pain monitoring. The algorithm uses a cylinder head model and 3D head pose recovery by alignment of dynamically extracted templates based on dense-HOG features. The algorithm includes extensions for drift reduction, using re-registration in combination with multi-pose state estimation by means of a square-root unscented Kalman filter. The paper reports experimental results on videos of moving infants in hospital who are relaxed or in pain. Results show good tracking behavior for poses up to 50 degrees from upright-frontal. In terms of eye location error relative to inter-ocular distance, the mean tracking error is below 9%. Walther E. Tjon a Ten, Peter H. N. de With |
ICMV | 3 |
| 2016 | ProMARTES: Accurate network and computation delay prediction for component-based distributed systems
Konstantinos Triantafyllidis, Waqar Aslam, Egor Bondarev, Johan J. Lukkien, Peter H. N. de With |
J. Syst. Softw. | 5 |
| 2015 | Solidarity Filter for Noise Reduction of 3D Edges in Depth Images
Hani Javan Hemmat, Egor Bondarev, Peter H. N. de With |
ACIVS | 3 |
| 2015 | Dynamic focus control for preventing motion blurabstractThis paper introduces motion focus as a complement to conventional camera focusing. We show how a relative shift between the image sensor and lens of a camera offers the ability to focus on specific motion in a scene. Focusing on motion allows the use of a longer exposure time while preventing motion blur of the object of interest. As a result, the object is captured with a significantly improved image quality. We derive a theoretical performance gain and compare it to existing imaging methods. Furthermore, we demonstrate an experimental motion focus implementation using modified off-the-shelf optical image stabilization hardware, which can obtain an effective exposure time extension of a factor of 94. Bart Kofoed, Eric Janssen, Peter H. N. de With |
AVSS | 3 |
| 2015 | Real-time semantic context labeling for image understandingabstractThe use of context information in a scene is an important aid for full semantic scene understanding in security and surveillance applications. To this end, this paper presents an innovative semantic context-labeling algorithm for three context classes, trading-off quality and real-time execution. Our system consists of three consecutive stages: image segmentation, region-based feature extraction and classification. We propose the joint use of the features color in HSV space, texture from Gabor filters and spatial context, in combination with the Directional Nearest Neighbor (DNN) method for constructing the undirected graph for segmentation. Compared to recent literature, this combination is over 35 times faster and achieves a coverability rate that is 65% higher. Martin A. R. Pieck, Fons van der Sommen, Svitlana Zinger, Peter H. N. de With |
ICIP | 4 |
| 2015 | Multi-resolution Gabor wavelet feature extraction for needle detection in 3D ultrasoundabstractUltrasound imaging is employed for needle guidance in various minimally invasive procedures such as biopsy guidance, regional anesthesia and brachytherapy. Unfortunately, a needle guidance using 2D ultrasound is very challenging, due to a poor needle visibility and a limited field of view. Nowadays, 3D ultrasound systems are available and more widely used. Consequently, with an appropriate 3D image-based needle detection technique, needle guidance and interventions may significantly be improved and simplified. In this paper, we present a multi-resolution Gabor transformation for an automated and reliable extraction of the needle-like structures in a 3D ultrasound volume. We study and identify the best combination of the Gabor wavelet frequencies. High precision in detecting the needle voxels leads to a robust and accurate localization of the needle for the intervention support. Evaluation in several ex-vivo cases shows that the multi-resolution analysis significantly improves the precision of the needle voxel detection from 0.23 to 0.32 at a high recall rate of 0.75 (gain 40%), where a better robustness and confidence were confirmed in the practical experiments. Arash Pourtaherian, Svitlana Zinger, Nenad Mihajlovic, Peter H. N. de With, Gary C. Ng, Hendrikus H. M. Korsten |
ICMV | 4 |
| 2014 | Context-based object-of-interest detection for a generic traffic surveillance analysis systemabstractWe present a new traffic surveillance video analysis system, focusing on building a framework with robust and generic techniques, based on both scene understanding and moving object-of-interest detection. Since traffic surveillance is widely applied, we want to design a single system that can be reused for various traffic surveillance applications. Scene understanding provides contextual information, which improves object detection and can be further used for other applications in a traffic surveillance system. Our framework consists of two main stages: Semantic Hypothesis Generation (SHG) and Context-Based Hypothesis Verification (CBHV). In the SHG stage, a semantic region labeling engine and an appearance-based detector jointly generate the visual regions with specific features or of specific interests. The regions may also contain objects of interest, either moving or static. In the CBHV stage, a cascaded verification is performed to refine the results and smooth the detection by temporal filtering. We model the context by jointly considering spatial and scale constraints and motion saliency. Our proposed framework is validated on real-life road surveillance videos, in which objects-of-interest are moving vehicles. The results of the obtained vehicle detection outperform a recent object detection algorithm, in both precision (92.7%) and recall (92.0%). The framework is both conceptually and in the applied techniques of a generic nature and can be reused in various traffic surveillance applications, that operate, e.g. on a road crossing or in a harbor. Xinfeng Bao, Solmaz Javanbakhti, Svitlana Zinger, Rob G. J. Wijnhoven, Peter H. N. de With |
AVSS | 5 |
| 2014 | Gabor-based needle detection and tracking in three-dimensional ultrasound data volumesabstractDuring needle interventions for e.g. regional anaesthesia or biopsy, it is very important to visualize the needle and its tip with respect to important structures in the body. In this work, we propose a novel image-based needle detection technique in a 3D ultrasound volume dataset, which can improve the intervention. We present a novel application of the 3D Gabor transformation, which exploits needle-like structures with appropriate designs. Furthermore, we introduce a needle tracking algorithm based on Gradient Descent and show that it limits the computational complexity and detection error. Finally, we visualize the needle on 2D cross-sections of the volume in order to be presented to the physician. Evaluation of our system in challenging cases shows a high detection score (up to 100% but needs larger sets) and accurate visualization. Arash Pourtaherian, Svitlana Zinger, Peter H. N. de With, Hendrikus H. M. Korsten, Nenad Mihajlovic |
ICIP | 3 |
| 2014 | Perimeter-intrusion event classification for on-line detection using multiple instance learning solving temporal ambiguitiesabstractThis paper describes a novel model for training an event detection system based on object tracking. We propose to model the training as a multiple instance learning problem, which allows us to train the classifier from annotated events despite temporal ambiguities. We apply this technique to realize a Perimeter Intrusion Detection (PID) algorithm and employ image-based features to distinguish real objects from moving vegetation and other distractions. An earlier developed tracking system is extended with the proposed technique to create an on-line PID-event detection system. Experiments with challenging videos show a reduction of the number of false positives by a factor 2-3 and improve the F1 detection performance from 0.15 to 0.28, when compared to a commercially available PID algorithm. Julien A. Vijverberg, Roel T. M. Janssen, Remco de Zwart, Peter H. N. de With |
ICIP | 4 |
| 2014 | Optimal Performance-Efficiency Trade-off for Bag of Words Classification of Road SignsabstractThis paper focuses on road sign classification for creating accurate and up-to-date inventories of traffic signs, which is important for road safety and maintenance. This is a challenging multi-class classification task, as a large number of different sign types exist which only differ in minor details. Moreover, changes in viewpoint, capturing conditions and partial occlusions result in large intra-class variations. Ideally, road sign classification systems should be robust against these variations, while having an acceptable computational load. This paper presents a classification approach based on the popular Bag Of Words (BOW) framework, which we optimize towards the best trade-off between performance and execution time. We analyze the performance aspects of PCA-based dimensionality reduction, soft and hard assignment for BOW codebook matching and the codebook size. Furthermore, we provide an efficient implementation scheme. We compare these techniques to design a fast and accurate BOW-based classification scheme. This approach allows for the selection of a fast but accurate classification methodology. This BOW approach is compared against structural classification, and we show that their combination outperforms both individual methods. This combination, exploiting both BOW and structural information, attains high classification scores (96.25% to 98%) on our challenging real-world datasets. Lykele B. Hazelhoff, Ivo M. Creusen, Peter H. N. de With |
ICPR | 3 |
| 2014 | Exploring Distance-Aware Weighting Strategies for Accurate Reconstruction of Voxel-Based 3D Synthetic Models
Hani Javan Hemmat, Egor Bondarev, Peter H. N. de With |
MMM (1) | 3 |
| 2014 | System for semi-automated surveying of street-lighting poles from street-level panoramic imagesabstractAccurate and up-to-date inventories of lighting poles are of interest to energy companies, beneficial for the transition to energy-efficient lighting and may contribute to a more adequate lighting of streets. This potentially improves social security and reduces crime and vandalism during nighttime. This paper describes a system for automated surveying of lighting poles from street-level panoramic images. The system consists of two independent detectors, focusing at the detection of the pole itself and at the detection of a specific lighting fixture type. Both follow the same approach, and start with detection of the feature of interest (pole or fixture) within the individual images, followed by a multi-view analysis to retrieve the real-world coordinates of the poles. Afterwards, the detection output of both algorithms is merged. Large-scale validations, covering about 135 km of road, show that over 91% of the lighting poles is found, while the precision remains above 50%. When applying this system in a semi-automated fashion, high-quality inventories can be created up to 5 times more efficiently compared to manually surveying all poles from the images. Lykele B. Hazelhoff, Ivo M. Creusen, Peter H. N. de With |
WACV | 3 |
| 2014 | Latency optimization for autostereoscopic volumetric visualization in image-guided interventions
Daniel Ruijters, Svitlana Zinger, Luat Do, Peter H. N. de With |
Neurocomputing | 4 |
| 2014 | Supportive automatic annotation of early esophageal cancer using local gabor and color features
Fons van der Sommen, Svitlana Zinger, Erik J. Schoon, Peter H. N. de With |
Neurocomputing | 4 |
| 2014 | Exploiting street-level panoramic images for large-scale automated surveying of traffic signs
Lykele B. Hazelhoff, Ivo M. Creusen, Peter H. N. de With |
Mach. Vis. Appl. | 3 |
| 2013 | Flexible Multi-modal Graph-Based Segmentation
Willem P. Sanberg, Luat Do, Peter H. N. de With |
ACIVS | 3 |
| 2013 | Towards real-time and low-latency video object tracking by linking tracklets of incomplete detectionsabstractThis paper considers tracking of objects for video-based intrusion detection systems. Current tracking algorithms can be used for surveillance, but in that use-case, these algorithms execute with too high latency and are not suitable for real-time applications. In this paper, we propose novel techniques for tracking algorithms based on tracklets in order to improve the execution time by limiting the number of tracklets and connection updates between tracklets. An additional improvement is that tracklet clustering has previously been applied to tracking with complete detections, i.e. a detection has a one-to-one correspondence to an object, while our proposed algorithm can handle incomplete detections as well. We show that the algorithm yields only two avoidable false positives on the i-LIDS SZTE dataset. To show that the algorithm can be executed in real-time, we have measured the worst-case execution time on a popular DSP which is only 31 ms per frame. Furthermore, the tracking algorithm requires only 35 seconds to process the complete i-LIDS dataset on a PC. Julien A. Vijverberg, Cornelis J. Koeleman, Peter H. N. de With |
AVSS | 3 |
| 2013 | On photo-realistic 3D reconstruction of large-scale and arbitrary-shaped environmentsabstractThis paper presents a system architecture for reconstructing photorealistic and accurate 3D models of indoor environments. The system specifically targets large-scale and arbitrary-shaped environments and enables processing of data obtained with an arbitrary-chosen capturing path. The system extends the baseline Kinect Fusion algorithm with a buffering algorithm to remove scene-size capturing limitations. Beside this, the paper presents the complete chain of advanced algorithms for point cloud segmentation/decimation, camera pose correction and texture mapping with post-processing filters. The presented architecture features memory- and processor-efficient processing, such that it can be executed on a conventional PC with a mainstream GPU card at the consumer premises. Egor Bondarev, Francisco Heredia, Rafael Favier, Lingni Ma, Peter H. N. de With |
CCNC | 5 |
| 2013 | Plane segmentation and decimation of point clouds for 3D environment reconstructionabstractThree-dimensional (3D) models of environments are a promising technique for serious gaming and professional engineering applications. In this paper, we introduce a fast and memory-efficient system for the reconstruction of large-scale environments based on point clouds. Our main contribution is the emphasis on the data processing of large planes, for which two algorithms have been designed to improve the overall performance of the 3D reconstruction. First, a flatness-based segmentation algorithm is presented for plane detection in point clouds. Second, a quadtree-based algorithm is proposed for decimating the point cloud involved with the segmented plane and consequently improving the efficiency of triangulation. Our experimental results have shown that the proposed system and algorithms have a high efficiency in speed and memory for environment reconstruction. Depending on the amount of planes in the scene, the obtained efficiency gain varies between 20% and 50%. Lingni Ma, Raphael Favier, Luat Do, Egor Bondarev, Peter H. N. de With |
CCNC | 5 |
| 2013 | Performance analysis method for RT systems: Promartes for autonomous robot
Konstantinos Triantafyllidis, Egor Bondarev, Peter H. N. de With |
FDL | 3 |
| 2013 | Complexity reduction of wavelet codecs through modified quality controlabstractInteger-to-integer wavelets are employed for low-complexity image encoders in embedded applications and used in combination with wavelet coefficient coders, such as EZW, SPIHT, SPECK and TSSP. In scalable coders bitstream creation is computationally expensive when individual coefficients are manipulated. In this paper, we study two options that move the quality control to earlier stages in the encoding process to alleviate this complexity problem: (1) to stage one of the TSSP and (2) to the integer wavelet transform. By inserting special functions based on premature bit-plane dropping, we effectively implement a quality-control step prior to the coefficient coding. For a typical usage scenario with full HD images, with option (1) we achieve an improvement of the processing speed of TSSP by 28%, while retaining the original bitstream and thus the ratedistortion performance. Furthermore, we have found that option (2) is generically applicable to other bit-plane based codecs, while offering nearly the same processing speed. Marijn J. H. Loomans, Peter H. N. de With |
ICIP | 2 |
| 2013 | Robust automatic ship tracking in harbours using active camerasabstractRadar is commonly used to detect and track ships in maritime surveillance. Unfortunately the systems are costly and do not provide any visual information about the object's type. To complement the ship identity information given by a radar system, we propose a supplementary system using active visual cameras that can robustly detect and track ships in harbours. By combining a high-quality, non real-time robust object detector with a feature point tracker with low computational complexity, it is possible to track ships in real time over long intervals and large distances. In addition to controlling pan and tilt, we dynamically control camera zoom to provide a high resolution image of the tracked object over a large range of distances. The tracking system is improved by a special motion estimation model for the feature points, which also incorporates zooming of the camera. The system is robust and sustains tracking even under challenging conditions, such as multiple viewpoints, a large variety of ships and various weather conditions. During experiments, various types of ships were successfully tracked for up to 18 minutes, and over a distance of almost 1.5km in the port of Rotterdam. The proposed system is generic and can be utilized in various tracking applications, by training the detector for a different object class. Marijn J. H. Loomans, Peter H. N. de With, Rob G. J. Wijnhoven |
ICIP | 2 |
| 2013 | TROD: Tracking with occlusion handling and drift correctionabstractWe present a tracking framework in which we learn a HOG-based object detector in the first video frame and use this detector to localize the object in subsequent frames. We contribute and improve the tracking on the three following points. First, an occlusion-handling algorithm exploits discriminative information from the detector by dividing the object bounding box into patches and comparing each patch to the object model. Second, a drift-correction technique uses descriptive information of the object by calculating the similarity between the object in the previous frame and its shifted versions in the current frame. Third, a stochastic learning algorithm updates the object detector using single object and single background samples for selected frames only. Experiments with benchmark sequences show that the proposed tracker outperforms state-of-the-art methods on several sequences and has the smallest average location error. Arash Pourtaherian, Rob G. J. Wijnhoven, Peter H. N. de With |
ICIP | 3 |
| 2013 | Robust moving ship detection using context-based motion analysis and occlusion handlingabstractThis paper proposes an original moving ship detection approach in video surveillance systems, especially con- centrating on occlusion problems among ships and vegetation using context information. Firstly, an over- segmentation is performed to divide and classify by SVM (Support Vector Machine) segments into water or non-water, while exploiting the context that ships move only in water. We assume that the ship motion to be characterized by motion saliency and consistency, such that each ship distinguish itself. Therefore, based on the water context model, non-water segments are merged into regions with motion similarity. Then, moving ships are detected by measuring the motion saliency of those regions. Experiments on real-life surveillance videos prove the accuracy and robustness of the proposed approach. We especially pay attention to testing in the cases of severe occlusions between ships and between ship and vegetation. The proposed algorithm outperforms, in terms of precision and recall, our earlier work and a proposal using SVM-based ship detection. Xinfeng Bao, Svitlana Zinger, Rob G. J. Wijnhoven, Peter H. N. de With |
ICMV | 4 |
| 2013 | Robust classification system with reliability prediction for semi-automatic traffic-sign inventory systemsabstractInventories of traffic signs are acquired from street-level images in a semi-automated fashion, employing object detection and classification techniques. This is a challenging task, as signs are captured from different viewpoints and under various weather conditions. Furthermore, many similar signs exist, only differing in minor details, and moreover, sign-like objects occur frequently. Consequently, current state-of-the-art systems are unable to reach the required quality level, implying the need for manual corrections. This involves checking all classification results to correct the small minority of misclassifications. This paper presents a classification approach aiming at both high recognition scores and predicting the reliability of the classification output, enabling selective manual intervention. Two reliability prediction methods are compared, analyzing either the classifier scores, or matching the input samples with predefined templates. Large-scale experiments performed for three sign classes, each containing numerous sign types, show that over 80% of the correctly classified results can be marked as reliable, while not marking any misclassifications as reliable. Hence, our research shows that a reliable prediction is possible and that manual invention can be concentrated to the about 25% remaining samples only. Overall, 92.7% of the 8, 159 signs are classified correctly. Lykele B. Hazelhoff, Ivo M. Creusen, Peter H. N. de With |
WACV | 3 |
| 2013 | Sparse-plus-dense-RANSAC for estimation of multiple complex curvilinear models in 2D and 3D
Chrysi Papalazarou, Peter H. N. de With, Peter M. J. Rongen |
Pattern Recognit. | 2 |
| 2013 | Extracting semantics from multi-spectrum video
Jungong Han, Eric J. Pauwels, Feng Wu 0001, Peter H. N. de With |
Pattern Recognit. Lett. | 4 |
| 2012 | Water Region Detection Supporting Ship Identification in Port Surveillance
Xinfeng Bao, Svitlana Zinger, Rob G. J. Wijnhoven, Peter H. N. de With |
ACIVS | 4 |
| 2012 | Color transformation for improved traffic sign detectionabstractThis paper considers large scale traffic sign detection on a dataset consisting of high-resolution street-level panoramic photographs. Traffic signs are automatically detected and classified with a set of state-of-the-art algorithms. We introduce a color transformation to extend a Histogram of Oriented Gradients (HOG) based detection algorithm to further improve the performance. This transformation uses a specific set of reference colors that aligns with traffic sign characteristics, and measures the distance of each pixel to these reference colors. This results in an improved consistency on the gradients at the outer edge of the traffic sign. In an experiment with 33, 400 panoramic images, the number of misdetections decreased by 53.6% and 51.4% for red/blue circular signs, and by 19.6% and 28.4% for yellow speed bump signs, measured at a realistic detector operating point. Ivo M. Creusen, Lykele B. Hazelhoff, Peter H. N. de With |
ICIP | 3 |
| 2012 | Robust classification of traffic signs using multi-view cuesabstractTraffic sign inventories are created for road safety and maintenance based on street-level panoramic images. Due to the large capturing interval, large viewpoint deviations between the different capturings occur. These viewpoint variations complicate the classification procedure, which aims at the selection of the correct sign type, out of a high number of nearly similar sign types, typically resulting in misclassifications. This paper describes a novel approach for incorporating viewpoint information to the classification procedure, where the sign orientation is estimated based on dense matching. Afterwards, each sample is corrected to a frontal viewpoint, which is then classified. Finally, the sign type is obtained by weighted voting. Large-scale experiments including 2, 224 traffic signs show that this approach reduces the misclassification rate by about 33% compared to the single-view case. Lykele B. Hazelhoff, Ivo M. Creusen, Peter H. N. de With |
ICIP | 3 |
| 2012 | Depth-guided inpainting algorithm for Free-Viewpoint VideoabstractFree-Viewpoint Video (FVV) is a novel technique which creates virtual images of multiple direction by view synthesis. In this paper, an exemplar-based depth-guided inpainting algorithm is proposed to fill disocclusions due to uncovered areas after projection. We develop an improved priority function which uses the depth information to impose a desirable inpainting order. We also propose an efficient background-foreground separation technique to enhance the accuracy of hole filling. Furthermore, a gradient-based searching approach is developed to reduce the computational cost and the location distance is incorporated into patch matching criteria to improve the accuracy. The experimental results have shown that the gradient-based search in our algorithm requires a much lower computational cost (factor of 6 compared to global search), while producing significantly improved visual results. Lingni Ma, Luat Do, Peter H. N. de With |
ICIP | 3 |
| 2012 | Robust detection, classification and positioning of traffic signs from street-level panoramic images for inventory purposesabstractAccurate inventories of traffic signs are required for road maintenance and increase of the road safety. These inventories can be performed efficiently based on street-level panoramic images. However, this is a challenging problem, as these images are captured under a wide range of weather conditions. Besides this, occlusions and sign deformations occur and many sign look-a-like objects exist. Our approach is based on detecting present signs in panoramic images, both to derive a classification code and to combine multiple detections into an accurate position of the signs. It starts with detecting the present signs in each panoramic image. Then, all detections are classified to obtain the specific sign type, where also false detections are identified. Afterwards, detections from multiple images are combined to calculate the sign positions. The performance of this approach is extensively evaluated in a large, geographical region, where over 85% of the 3; 341 signs are automatically localized, with only 3:2% false detections. As nearly all missed signs are detected in at least a single image, only very limited manual interactions have to be supplied to safeguard the performance for highly accurate inventories. Lykele B. Hazelhoff, Ivo M. Creusen, Peter H. N. de With |
WACV | 3 |
| 2012 | Intelligent trainee behavior assessment system for medical training employing video analysis
Jungong Han, Peter H. N. de With, Ashley Merien, Guid Oei |
Pattern Recognit. Lett. | 2 |
| 2012 | View Interpolation for Medical Images on Autostereoscopic DisplaysabstractWe present an approach for efficient rendering and transmitting views to a high-resolution autostereoscopic display for medical purposes. Displaying biomedical images on an autostereoscopic display poses different requirements than in a consumer case. For medical usage, it is essential that the perceived image represents the actual clinical data and offers sufficiently high quality for diagnosis or understanding. Autostereoscopic display of multiple views introduces two hurdles: transmission of multi-view data through a bandwidth-limited channel and the computation time of the volume rendering algorithm. We address both issues by generating and transmitting limited set of views enhanced with a depth signal per view. We propose an efficient view interpolation and rendering algorithm at the receiver side based on texture+depth data representation, which can operate with a limited amount of views. We study the main artifacts that occur during rendering-occlusions, and we quantify them first for a synthetic model and then for real-world biomedical data. The experimental results allow us to quantify the peak signal-to-noise ratio for rendered texture and depth as well as the amount of disoccluded pixels as a function of the angle between surrounding cameras. Svitlana Zinger, Daniel Ruijters, Luat Do, Peter H. N. de With |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2011 | Detection of Human Groups in Videos
Selçuk Sandikci, Svitlana Zinger, Peter H. N. de With |
ACIVS | 3 |
| 2011 | Robust Model-Based Detection of Gable Roofs in Very-High-Resolution Aerial Images
Lykele B. Hazelhoff, Peter H. N. de With |
CAIP (1) | 2 |
| 2011 | Clustering of tracklets for on-line multi-target tracking in networked camera systemsabstractThis paper considers the problem of tracking a variable number of objects through a surveillance site monitored by multiple cameras with slightly overlapping field-of-views. To this end, we propose to cluster tracklets generated by a commercially available single-camera video-analysis algorithm which is solely based on the position of objects. A first contribution of this paper is the proposal of a novel, extended energy function representing the confidence that two tracklets correspond to the same object. In contrast to previous work, the proposed motion-consistency error enables the clustering of tracklets from arbitrary views and temporal overlap. A second contribution is to evaluate the performance of several clustering algorithms. The results show that the clustering techniques employing only the merging of tracklets yield 10-15% higher F1 score than clustering techniques using various types of clustering moves including split and swap moves. Julien A. Vijverberg, Cornelis J. Koeleman, Peter H. N. de With |
CISDA | 3 |
| 2011 | Real-time free-viewpoint DIBR on GPUs for large base-line multi-view 3DTV videosabstractGPUs are ample utilized for their parallel processing capabilities and are also interesting in the area of 3D image processing where computational demanding tasks are essential. We concentrate particularly on free-view 3DTV applications, based on image warping to the new view and image artifact reduction techniques. In this paper, we report on the implementation of an efficient free-viewpoint DIBR algorithm with an off-the-shelf GPU, which can be readily integrated into future 3DTV systems. For exploiting maximal parallelism, we have mapped the processing of each pixel to a separate thread and grouped multiple threads to larger blocks as much as possible. Using a combination of the highly parallel programming architecture CUDA and a graphics API, we have achieved a real-time performance operating on 1080p HD multi-view video with a rendering quality that is comparable to the software implementation. Luat Do, Germán Bravo, Svitlana Zinger, Peter H. N. de With |
VCIP | 4 |
| 2011 | Real-time Two-Stage SPECK (TSSP) design and implementation for scalable video coding on embedded systemsabstractIn this paper, we discuss the design and real-time implementation of a novel wavelet coefficient encoder for Scalable Video Coding (SVC) in a surveillance environment. The novel wavelet coefficient encoder, called Two-Stage SPECK (TSSP), is designed for efficient hardware/software implementations on parallel architectures. As the encoder needs to be integrated into an embedded system, we will elaborate on its implementation on a common VLIW DSP (DM642). Optimization techniques such as SIMD (Single Instruction Multiple Data) and DMA (Direct Memory Access) are applied to maximize parallelism with concurrent calculations and memory transfers. We have realized the execution of the TSSP at 4CIF (CCIR-601) broad- cast resolution and YCbCr 4:2:0 color in 7.884 Mcycles per frame, including memory stalls. This translates to more than 75 full-color encodings per second at a clock rate of 600 MHz. Marijn J. H. Loomans, Cornelis J. Koeleman, Peter H. N. de With |
VCIP | 3 |
| 2011 | Real-time multiple people tracking for automatic group-behavior evaluation in delivery simulation trainingabstractThis paper aims at generating an automated way to evaluate the team-behavior of trainees in a delivery simulation course using video-processing techniques with emphasis on multiple people tracking. The paper is composed of two interacting, but clearly separated stages: moving people segmentation and multiple people tracking. At people segmentation stage, the combination of the Gaussian Mixture Model (GMM) and the Dynamic Markov Random Fields (DMRF) technique helps to extract the foreground pixels. For a better extraction of the human silhouettes, the energy function of DMRF is extended with texture information. At multiple people tracking stage, we concentrate on solving human-occlusion problem caused by interacting persons based on silhouette data and a non-linear regression model. Our model effectively transfers the person location problem during the occlusion into the finding of the local maximum points on a smooth curve, so that visual persons in the partial or complete occlusion can still be precisely captured. We have compared our algorithm with two other popular tracking algorithms: mean-shift and particle-filter. Experimental results reveal that the correctness of our method is much higher than the mean-shift algorithm and slightly lower than a particle-filter, however, with the major benefit of being a factor of 10–15 faster in computing. Jungong Han, Peter H. N. de With |
Multim. Tools Appl. | 2 |
| 2010 | Video Quality Analysis for Concert Video Mashup Generation
Prarthana Shrestha, Hans Weda, Mauro Barbieri, Peter H. N. de With |
ACIVS (2) | 4 |
| 2010 | Color exploitation in hog-based traffic sign detectionabstractWe study traffic sign detection on a challenging large-scale real-world dataset of panoramic images. The core processing is based on the Histogram of Oriented Gradients (HOG) algorithm which is extended by incorporating color information in the feature vector. The choice of the color space has a large influence on the performance, where we have found that the CIELab and YCbCr color spaces give the best results. The use of color significantly improves the detection performance. We compare the performance of a specific and HOG algorithm, and show that HOG outperforms the specific algorithm by up to tens of percents in most cases. In addition, we propose a new iterative SVM training paradigm to deal with the large variation in background appearance. This reduces memory consumption and increases utilization of background information. Ivo M. Creusen, Rob G. J. Wijnhoven, Ernst Herbschleb, Peter H. N. de With |
ICIP | 4 |
| 2010 | Objective quality analysis for free-viewpoint DIBRabstractInteractive free-viewpoint selection applied to a 3D multi-view video signal is an attractive feature of the rapidly developing 3DTV media. In recent years, significant research has been done on free-viewpoint rendering algorithms which mostly have similar building blocks. In this paper, we analyze the four principal building blocks of most recent rendering algorithms and their contribution to the overall rendering quality. We have discovered that the rendering quality is dominated by the first step, Warping which determines the basic quality level of the complete rendering chain. The third step, Blending, is a valuable step which further increases the rendering quality by as much as 1.4 dB while reducing the disocclusions to less than 1% of the total image. Varying the angle between the two reference cameras, we notice that the quality of each principal building block degrades with a similar rate, 0.1-0.3dB/degree for real-life sequences. While experimenting with synthetic data of higher accuracy, we conclude that for developing better free-viewpoint algorithms, it is necessary to generate depth maps with more quantization levels so that the Warping and Blending steps can further contribute to the quality enhancement. Luat Do, Svitlana Zinger, Peter H. N. de With |
ICIP | 3 |
| 2010 | Enhanced prediction for motion estimation in Scalable Video CodingabstractIn this paper, we present a temporal candidate generation scheme that can be applied to motion estimators in Scalable Video Codecs (SVCs). For bidirectional motion estimation, usually a test is made for each block to determine which motion compensation direction is preferred: forward, bidirectional or backward. Instead of simply using the last computed motion vector field (backward or forward), giving an asymmetry in the estimation, we involve both vector fields to generate a single candidate field for a more stable and improved prediction. This field is generated with the aid of mode decision information of the codec. This single field of motion vector candidates serves two purposes: (1) it initializes the next recursion and (2) it is the foundation for the succeeding scale in the scalable coding. We have implemented this improved candidate system for both HPPS as EPZS motion estimators in a scalable video codec. We have found that it reduces the errors caused by occlusion of moving objects or image boundaries. For EPZS, only a small improvement is observed compared to the simple candidate scheme. However, for HPPS improvements are more significant: when looking at individual levels, motion compensation performance improves by up to 0.84 dB and when implemented in SVC, HPPS slightly outperforms EPZS. Marijn J. H. Loomans, Cornelis J. Koeleman, Peter H. N. de With |
ICIP | 3 |
| 2010 | Surgical needle reconstruction using small-angle multi-view X-rayabstractIn biopsies, drainages, vertebroplasty, and other needle-based procedures, insight on the 3D position of a needle is crucial for correct navigation by the clinician. In this paper, we present a method for the reconstruction of surgical needles using multi-view X-ray imaging with a small motion of the C-arm. It is required that the extent of the motion is limited (<; 30 degrees) to allow use of this method during an intervention. This small motion provides sufficient multi-view information, which is used in combination with a needle model for the 3D reconstruction of the needle. To this end, we describe a system comprising the steps of (a) needle detection in a novel, RANSAC-based framework, (b) tracking of needles in subsequent views using geometric constraints and (c) needle reconstruction. Results are presented in comparison to a volume reconstruction using a full rotation of the C-arm (≈207 degrees), showing good accuracy of the proposed method. Chrysi Papalazarou, Peter M. J. Rongen, Peter H. N. de With |
ICIP | 3 |
| 2010 | Geometric averaging of X-ray signals in automatic exposure controlabstractImproper dose control in X-ray cardio-vascular systems leads to a reduced Signal-to-Noise Ratio (SNR) in regions of interest of the X-ray image. We aim at reducing the influence of direct radiation, entering a measuring field for X-ray dose control in a Flat Detector which gives too bright areas (highlights) in the image. It is our desire to use a norm-like signal size that represents a minimal dose value while maximizing information transfer and thus image quality. In a dose control system, it is common practice to employ a special averaging technique for computing a representative signal level controlling the X-ray. We have found that the geometric averaging outperforms the existing techniques and significantly improves the image quality. Our approach reduces the highlight influence and guarantees an adequate Contrast-to-Noise ratio for decentered objects. We provide convincing experimental results showing a strongly improved image quality with respect to contrast and detail. Rudolph M. Snoeren, Peter H. N. de With |
ICIP | 2 |
| 2010 | Conversion of free-viewpoint 3D multi-view video for stereoscopic displaysabstractThis paper presents our ongoing research on view synthesis of free-viewpoint 3D multi-view video for 3DTV. With the emerging breakthrough of stereoscopic 3DTV, we have extended a reference free-viewpoint rendering algorithm to generate stereoscopic views. Two similar solutions for converting free-viewpoint 3D multi-view video into a stereoscopic vision have been developed. These solutions take into account the complexity of the algorithms by exploiting the redundancy in stereo images, since we aim at a real-time hardware implementation. Both solutions are based on applying a horizontal shift instead of double execution of the reference free-viewpoint rendering algorithm for stereo generation (FVP stereo generation), so that the rendering time can be reduced by as much as 30-40 %. The trade-off however, is that the rendering quality is 0.5-0.9 dB lower than when applying FVP stereo generation. Our results show that stereoscopic views can be efficiently generated from 3D multi-view video by using unique properties in stereoscopic views, such as identical orientation, similarities in textures and small baseline. Luat Do, Svitlana Zinger, Peter H. N. de With |
ICME | 3 |
| 2010 | Multiple Model Estimation for the Detection of Curvilinear Segments in Medical X-ray Images Using Sparse-plus-dense-RANSACabstractIn this paper, we build on the RANSAC method to detect multiple instances of objects in an image, where the objects are modeled as curvilinear segments with distinct endpoints. Our approach differs from previously presented work in that it incorporates soft constraints, based on a dense image representation, that guide the estimation process in every step. This enables (1) better correspondence with image content, (2) explicit endpoint detection and (3) a reduction in the number of iterations required for accurate estimation. In the case of curvilinear objects examined in this paper, these constraints are formulated as binary image labels, where the estimation proved to be robust to mislabeling, e.g. in case of intersections. Results for both synthetic and real data from medical X-ray images show the improvement from incorporating soft image-based constraints. Chrysi Papalazarou, Peter M. J. Rongen, Peter H. N. de With |
ICPR | 3 |
| 2010 | Fast Training of Object Detection Using Stochastic Gradient DescentabstractTraining datasets for object detection problems are typically very large and Support Vector Machine (SVM) implementations are computationally complex. As opposed to these complex techniques, we use Stochastic Gradient Descent (SGD) algorithms that use only a single new training sample in each iteration and process samples in a stream-like fashion. We have incorporated SGD optimization in an object detection framework. The object detection problem is typically highly asymmetric, because of the limited variation in object appearance, compared to the background. Incorporating SGD speeds up the optimization process significantly, requiring only a single iteration over the training set to obtain results comparable to state-of-the-art SVM techniques. SGD optimization is linearly scalable in time and the obtained speedup in computation time is two to three orders of magnitude. We show that by considering only part of the total training set, SGD converges quickly to the overall optimum. Rob G. J. Wijnhoven, Peter H. N. de With |
ICPR | 2 |
| 2010 | Automatic mashup generation from multiple-camera concert recordingsabstractA large number of videos are captured and shared by the audience from musical concerts. However, such recordings are typically perceived as boring mainly because of their limited view, poor visual quality and incomplete coverage. It is our objective to enrich the viewing experience of these recordings by exploiting the abundance of content from multiple sources. In this paper, we propose a novel \Virtual Director system that automatically combines the most desirable segments from different recordings resulting in a single video stream, called mashup. We start by eliciting requirements from focus groups, interviewing professional video editors and consulting film grammar literature. We design a formal model for automatic mashup generation based on maximizing the degree of fulfillment of the requirements. Various audio-visual content analysis techniques are used to determine how well the requirements are satisfied by a recording. To validate the system, we compare our mashups with two other mashups: manually created by a professional video editor and machine generated by random segment selection. The mashups are evaluated in terms of visual quality, content diversity and pleasantness by 40 subjects. The results show that our mashups and the manual mashups are perceived as comparable, while both of them are significantly higher than the random mashups in all three terms. Prarthana Shrestha, Peter H. N. de With, Hans Weda, Mauro Barbieri, Emile H. L. Aarts |
ACM Multimedia | 2 |
| 2010 | Temporal signal energy correction and low-complexity encoder feedback for lossy scalable video codingabstractIn this paper, we address two problems found in embedded implementations of Scalable Video Codecs (SVCs): the temporal signal energy distribution and frame-to-frame quality fluctuations. The unequal energy distribution between the low- and high-pass band with integer-based wavelets leads to sub-optimal rate-distortion choices coupled with quantization-error accumulations. The second problem is the quality fluctuation between frames within a Group Of Pictures (GOP). To solve these two problems, we present two modifications to the SVC. The first modification aims at a temporal energy correction of the lifting scheme in the temporal wavelet decomposition. By moving this energy correction to the leaves of the temporal tree, we can save on required memory size, bandwidth and computations, while reducing floating/fixed-point conversion errors. The second modification feeds back the decoded first frame of the GOP (the temporal low-pass) into the temporal coding chain. The decoding of the first frame is achieved without entropy decoding while avoiding any required modifications at the decoder. Experiments show that quality fluctuations within the GOP are significantly reduced, thereby significantly increasing the subjective visual quality. On top of this, a small quality improvement is achieved on average. Marijn J. H. Loomans, Cornelis J. Koeleman, Peter H. N. de With |
PCS | 3 |
| 2010 | Detecting critical configurations for Euclidean 3D reconstruction by analyzing the scaled measurement matrixabstract3D reconstruction is ambiguous under so-called critical motions or critical surfaces. This paper proposes an
algorithm to detect a few critical configurations where Euclidean reconstruction degenerates. Assuming that the
focal lengths are the only unknown intrinsic parameters, the following critical configurations are detected: (1)
coplanar 3D points, (2) pure rotation; (3) rotation around two camera centers; (4) presence of excessive noise
and outliers in the measurements. The configurations in Cases (1), (2) and (4) will affect the rank of the scaled
measurement matrix (SMM). The number of camera centers in Case (3) will affect the number of independent
rows of the SMM. By examining the rankness and the number of independent rows of the SMM, we are able
to detect the above-mentioned critical configurations. Experimental results on both synthetic and real data
demonstrate the effectiveness of the proposed algorithm on detecting the critical situations for factorizationbased
3D reconstruction. Ping Li 0002, Rene Klein Gunnewiek, Peter H. N. de With |
VCIP | 3 |
| 2010 | Free-viewpoint depth image based rendering
Svitlana Zinger, Luat Do, Peter H. N. de With |
J. Vis. Commun. Image Represent. | 3 |
| 2009 | Detecting Critical Configurations for Dividing Long Image Sequences for Factorization-Based 3-D Scene Reconstruction
Ping Li 0002, Rene Klein Gunnewiek, Peter H. N. de With |
ACCV (2) | 3 |
| 2009 | Behavioral State Detection of Newborns Based on Facial Expression Analysis
Lykele B. Hazelhoff, Jungong Han, Sidarto Bambang-Oetomo, Peter H. N. de With |
ACIVS | 4 |
| 2009 | Evaluation of Interest Point Detectors for Non-planar, Transparent Scenes
Chrysi Papalazarou, Peter M. J. Rongen, Peter H. N. de With |
ACIVS | 3 |
| 2009 | Comparing Feature Matching for Object Categorization in Video Surveillance
Rob G. J. Wijnhoven, Peter H. N. de With |
ACIVS | 2 |
| 2009 | Global Illumination Compensation for Background Subtraction Using Gaussian-Based Background Difference ModelingabstractThis paper presents a background segmentation technique, which is able to process acceptable segmentation masks under fast global illumination changes. The histogram of the frame-based background difference is modeled with multiple kernels. The model that represents the histogram at best, is used to determine the shift in luminance due to global illumination or diaphragm changes, such that the background difference can be compensated. Experimental results have revealed that the number of incorrectly classified pixels using global illumination compensation instead of only the approximated median method reduces from 77% to 19% shortly after a fast change. The performance of the proposed technique is similar to state-of-the-art related work for global illumination changes, despite the fact that only luminance information is used. The algorithm is computationally simple and can operate at 30 frames-per-second for VGA resolution on a P-IV 3-GHz PC. Julien A. Vijverberg, Marijn J. H. Loomans, Cornelis J. Koeleman, Peter H. N. de With |
AVSS | 4 |
| 2009 | Resource usage prediction for groups of dynamic image-processing tasks using Markov modelingabstractWith the introduction of dynamic image processing, such as in image analysis, the computational complexity has become data dependent and memory usage irregular. Therefore, the possibility of runtime estimation of resource usage would be highly attractive and would enable quality-of-service (QoS) control for dynamic image-processing applications with shared resources. A possible solution to this problem is to characterize the application execution using model descriptions of the resource usage. In this paper, we attempt to predict resource usage for groups of dynamic image-processing tasks based on Markov-chain modeling. As a typical application, we explore a medical imaging application to enhance a wire mesh tube (stent) under X-ray fluoroscopy imaging during angioplasty. Simulations show that Markov modeling can be successfully applied to describe the resource usage function even if the flow graph dynamically switches between groups of tasks. For the evaluated sequences, an average prediction accuracy of 97% is reached with sporadic excursions of the prediction error up to 20-30%. Rob Albers, Eric Suijs, Peter H. N. de With |
ICASSP | 3 |
| 2009 | Resource prediction and quality control for parallel execution of heterogeneous medical imaging tasksabstractWe have established a novel control system for combining the parallel execution of deterministic and non-deterministic medical imaging applications on a single platform, sharing the same constrained resources. The control system aims at avoiding resource overload and ensuring throughput and latency of critical applications, by means of accurate resource-usage prediction. Our approach is based on modeling the required computation tasks, by employing a combination of weighted moving-average filtering and scenario-based Markov chains to predict the execution. Experimental validation on medical image processing shows an accuracy of 97%. As a result, the latency variation within non-deterministic analysis applications is reduced by 70% by adaptively splitting/merging of tasks. Furthermore, the parallel execution of a deterministic live-viewing application features constant throughput and latency by dynamically switching between quality modes. Interestingly, our solution can successfully be reused for alternative applications with several parallel streams, like in surveillance. Rob Albers, Eric Suijs, Peter H. N. de With |
ICIP | 3 |
| 2009 | Highly-parallelized motion estimation for Scalable Video CodingabstractIn this paper, we discuss the design of a highly-parallel motion estimator for real-time Scalable Video Coding (SVC). In an SVC, motion is commonly estimated bidirectionally and over various temporal distances. Current motion estimators are optimized for frame-by-frame estimation, and such estimators are designed without serious implementation constraints. To support efficient embedded applications, we propose a Highly Parallel Predictive Search (HPPS) motion estimator while preserving an accurate estimation performance. The motion estimation algorithm is optimized for processing on parallel cores and utilizes a novel recursive search strategy. This strategy is based on hierarchically increasing the temporal distance in the estimation algorithm while using the state of the previous hierarchical layer as an input. Due to the absence of local recursions in the algorithm, the proposed motion estimator has a constant computational load, regardless of video activity or temporal distance. We compared our proposed motion estimator to the well-known full search, ARPS3, 3DRS and EPZS motion estimators for the SVC case, and obtain a performance close to full search (0.2dB), while outperforming other algorithms in prediction. Marijn J. H. Loomans, Cornelis J. Koeleman, Peter H. N. de With |
ICIP | 3 |
| 2009 | Region-based all-in-focus light field renderingabstractLight field rendering is an approach to synthesize virtual views of a scene from a set of original images. When minimizing the number of images for rendering, the light field may become under-sampled, leading to aliasing artifacts. To render an under-sampled light field in high quality and without aliasing artifacts is a challenge. We present a light field rendering algorithm with region-based focus to create all-in-focus virtual views. Our algorithm was compared to: (a) ground truth images and (b) a state-of-the-art technique for rendering under-sampled light fields. Extensive rendering experiments confirm that our algorithm provides visible quality improvement, quantified as about 10% RMSE reduction. However, the subjective improvement is larger and it produces images comparable to the ground truth. This algorithm contributes to practical applications of light field rendering, such as image generation in stereoscopic displays and free-viewpoint video (FVV). Goran Petrovic, Aneez K. Shahulhameed, Svitlana Zinger, Peter H. N. de With |
ICIP | 4 |
| 2009 | Real-time scalable video codec implementation for surveillanceabstractIn this paper, we discuss the design and real-time implementation of a Scalable Video Codec (SVC) for surveillance applications. We present a complexity-scalable temporal wavelet transform and the implementation of a multi-level 2D 5/3 wavelet transform, using the lifting framework. We have employed SIMD (Single Instruction Multiple Data) and DMA (Direct Memory Access) techniques, where the proposed process of background DMA transfers is so effective, that the ALUs are always supplied with input data. We have realized the execution of a 4-level transform at 4CIF (CCIR-601) broadcast resolution in 3.65 Mcycles, including memory stalls, on a TMS320DM642 DSP. At a clock rate of 600 MHz, this translates to more than 160 transforms per second. For our complete SVC, we achieve a frame rate of 12.5-15 fps depending on scene activity. Marijn J. H. Loomans, Cornelis J. Koeleman, Peter H. N. de With |
ICME | 3 |
| 2009 | Triple-C: Resource-usage prediction for semi-automatic parallelization of groups of dynamic image-processing tasksabstractWith the emergence of dynamic video processing, such as in image analysis, runtime estimation of resource usage would be highly attractive for automatic parallelization and QoS control with shared resources. A possible solution is to characterize the application execution using model descriptions of the resource usage. In this paper, we introduce Triple-C, a prediction model for Computation, Cache-memory and Communication-bandwidth usage with scenario-based Markov chains. As a typical application, we explore a medical imaging function to enhance objects of interest in X-ray angiography sequences. Experimental results show that our method can be successfully applied to describe the resource usage for dynamic image-processing tasks, even if the flow graph dynamically switches between groups of tasks. An average prediction accuracy of 97% is reached with sporadic excursions of the prediction error up to 20–30%. As a case study, we exploit the prediction results for semi-automatic parallelization. Results show that with Triple-C prediction, dynamic processing tasks can be executed in real-time with a constant low latency. Rob Albers, Eric Suijs, Peter H. N. de With |
IPDPS | 3 |
| 2009 | The effects of multiview depth video compression on multiview rendering
Philipp Merkle, Yannick Morvan, Aljoscha Smolic, Dirk Farin, Karsten Müller 0001, Peter H. N. de With, Thomas Wiegand 0001 |
Signal Process. Image Commun. | 6 |
| 2008 | Video-Based Fall Detection in the Home Using Principal Component Analysis
Lykele B. Hazelhoff, Jungong Han, Peter H. N. de With |
ACIVS | 3 |
| 2008 | Scene Reconstruction Using MRF Optimization with Image Content Adaptive Energy Functions
Ping Li 0002, Rene Klein Gunnewiek, Peter H. N. de With |
ACIVS | 3 |
| 2008 | A real-time video surveillance system with human occlusion handling using nonlinear regressionabstractThis paper presents a real-time single-camera surveillance system, aiming at detecting and partly analyzing a group of people. A set of moving persons is segmented using a combination of the Gaussian Mixture Model (GMM) and the Dynamic Markov Random Fields (DMRF) technique. For a better extraction of the human silhouettes, the energy function of DMRF is extended with texture information. The mean-shift algorithm is utilized to track multiple people over the sequence. To address the human-occlusion problem, we model the horizontal projection histograms of the human silhouettes using a nonlinear regression algorithm. This model enables to automatically locate the people during the occlusions. Experiments show that the proposal has nearly same performance (also with occlusion) as the particle-filter with the benefit of being a factor of 10–20 faster in computing. Jungong Han, Minwei Feng, Peter H. N. de With |
ICME | 3 |
| 2008 | Facial feature extraction by a cascade of model-based algorithms
Fei Zuo, Peter H. N. de With |
Signal Process. Image Commun. | 2 |
| 2008 | Broadcast Court-Net Sports Video Analysis Using Fast 3-D Camera ModelingabstractThis paper addresses the automatic analysis of court-net sports video content. We extract information about the players, the playing-field in a bottom-up way until we reach scene-level semantic concepts. Each part of our framework is general, so that the system is applicable to several kinds of sports. A central point in our framework is a camera calibration module that relates thea-prioriinformation of the geometric layout in the form of a court model to the input image. Exploiting this information, several novel algorithms are proposed, including playing-frame detection, players segmentation and tracking. To address the player-occlusion problem, we model the contour map of the player silhouettes using a nonlinear regression algorithm, which enables to locate the players during the occlusions caused by players in the same team. Additionally, a Bayesian-based classifier helps to recognize predefined key events, where the input is a number ofreal-worldvisual features. We illustrate the performance and efficiency of the proposed system by evaluating it for a variety of sports videos containing badminton, tennis and volleyball, and we show that our algorithm can operate with more than 91% feature detection accuracy and 90% event detection. Jungong Han, Dirk Farin, Peter H. N. de With |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | Texture-Independent Feature-Point Matching (TIFM) from Motion Coherence
Ping Li 0002, Dirk Farin, Rene Klein Gunnewiek, Peter H. N. de With |
ACCV (1) | 4 |
| 2007 | Descriptor-Free Smooth Feature-Point Matching for Images Separated by Small/Mid Baselines
Ping Li 0002, Dirk Farin, Rene Klein Gunnewiek, Peter H. N. de With |
ACIVS | 4 |
| 2007 | Multiview Depth-Image Compression Using an Extended H.264 Encoder
Yannick Morvan, Dirk Farin, Peter H. N. de With |
ACIVS | 3 |
| 2007 | Patch-Based Experiments with Object Classification in Video Surveillance
Rob G. J. Wijnhoven, Peter H. N. de With |
ACIVS | 2 |
| 2007 | Grass Detection for Picture Quality Enhancement of TV Video
Bahman Zafarifar, Peter H. N. de With |
ACIVS | 2 |
| 2007 | Experiments with patch-based object classificationabstractWe present and experiment with a patch-based algorithm for the purpose of object classification in video surveillance. A feature vector is calculated based on template matching of a large set of image patches, within detected regions-of-interest (ROIs, also called blobs), of moving objects. Instead of matching direct image pixels, we use Gabor-filtered versions of the input image at several scales. We present results for a new typical video surveillance dataset containing over 9,000 object images. Additionally, we show results for the PETS 2001 dataset and another dataset from literature. Because our algorithm is not invariant to the object orientation, the set was split into four subsets with different orientation. We show the improvements, resulting from taking the object orientation into account. Using 50 training samples or higher, our resulting detection rate is on the average above 95%, which improves with the orientation consideration to 98%. Because of the inherent scalability of the algorithm, an embedded system implementation is well within reach. Rob G. J. Wijnhoven, Peter H. N. de With |
AVSS | 2 |
| 2007 | CARAT: a toolkit for design and performance analysis of component-based embedded systemsabstractSolid frameworks and toolkits for design and analysis of embedded systems are of high importance, since they enable early reasoning about critical properties of a system. This paper presents a software toolkit that supports the design and performance analysis of real-time component-based software architectures deployed on heterogeneous multiprocessor platforms. The tooling environment contains a set of integrated tools for (a) component storage and retrieval, (b) graphics-based design of software and hardware architectures, (c) performance analysis of the designed architectures and, (d) automated code generation. The cornerstone of the toolkit is a performance analysis framework that automates composition of the individual component models into a system executable model, allows simulation of the system model and gives design-time predictions of key performance properties like response time, data throughput, and usage of hardware resources. The efficiency of this toolkit was illustrated on a car radio navigation benchmark system Egor Bondarev, Michel R. V. Chaudron, Peter H. N. de With |
DATE | 3 |
| 2007 | Contrast-Invariant Feature Point CorrespondenceabstractMost existing feature-matching methods utilize texture correlation for feature matching, which is usually sensitive to contrast changes. This paper proposes a new feature-point matching algorithm that does not rely on the image texture. Instead, only the smoothness assumption, which states that the displacement field in a neighborhood is coherent (smooth), is used. In the proposed method, the collected correspondences of a group of feature points within a neighborhood are efficiently determined such that the coherence measure of the displacement field in the neighborhood is maximized. The experimental results show that the proposed method is invariant to contrast changes and significantly outperforms the conventional block-matching technique. Ping Li 0002, Dirk Farin, Rene Klein Gunnewiek, Peter H. N. de With |
ICASSP (1) | 4 |
| 2007 | High-Level Traffic-Violation Detection for Embedded Traffic AnalysisabstractThis paper presents the design of a robust and real-time traffic-violation detection system for cameras on intersections. We use background segmentation and a novel road-model to obtain the candidate traffic participants. A region-based tracking system, equipped with static occlusion-reasoning, tracks the positions of the objects in the scene. A computationally efficient camera model is defined which only requires three input parameters and enables the extraction of key object parameters like vehicle type and speed. Experiments have shown that an impressive average processing rate of 63-150 Hz is achieved, with high average correct road detection and object-type classification rates of 93-94% and event detection accuracy of 85%. Julien A. Vijverberg, Nick A. H. M. de Koning, Jungong Han, Peter H. N. de With, Dion Cornelissen |
ICASSP (2) | 4 |
| 2007 | Locally-Adaptive Image Contrast Enhancement without Noise and Ringing ArtifactsabstractFor real-time imaging in surveillance applications, visibility of details is of primary importance to ensure customer confidence. Additional constraints are the absence of human interaction and low computational complexity. Usually, image quality is improved by enhancing contrast and sharpness. Many complex scenes require local contrast improvements that should bring details to the best possible visibility. However, local enhancement methods mainly suffer from ringing artifacts and noise over-enhancement. In this paper, we present a new multi-window real-time high-frequency enhancement scheme, in which gain is a nonlinear function of the detail energy. Our algorithm controls perceived sharpness, ringing artifacts (contrast) and noise, resulting in a good balance between visibility of details and non-disturbance of artifacts. Its advantage is that gains can be set now much higher than usual and the algorithm will reduce them only at places where it is really needed. Sascha D. Cvetkovic, Johan Schirris, Peter H. N. de With |
ICIP (3) | 3 |
| 2007 | Incorporating Depth-Image Based View-Prediction into H.264 for Multiview-Image CodingabstractWe investigate the coding of multiview images obtained from a set of multiple cameras. To exploit the inter-view correlation, two view-prediction tools have been implemented and used in parallel: a block-based motion compensation scheme and a depth image based rendering technique (DIBR). Whereas DIBR relies on an accurate depth image, the block-based motion-compensation scheme can be performed without any geometry information. Our encoder adaptively selects the most appropriate prediction scheme using a rate-distortion criterion for an optimal prediction-mode selection. The attractiveness of the algorithm is that the compression algorithm is robust against inaccurately estimated depth images and requires only one single reference camera for fast random-access to different views. We present experimental results for several multiview sequences, that result in a quality improvement of up to 1.4 dB as compared to H.264 compression. Yannick Morvan, Dirk Farin, Peter H. N. de With |
ICIP (1) | 3 |
| 2007 | Depth-Image Compression Based on an R-D Optimized Quadtree Decomposition for the Transmission of Multiview ImagesabstractThis paper presents a novel depth-image coding algorithm that concentrates on the special characteristics of depth images: smooth regions delineated by sharp edges. The algorithm models these smooth regions using piecewise-linear functions and sharp edges by a straight line. To define the area of support for each modeling function, we employ a quadtree decomposition that divides the image into blocks of variable size, each block being approximated by one modeling function containing one or two surfaces. The subdivision of the quadtree and the selection of the type of modeling function is optimized such that a global rate-distortion trade-off is realized. Additionally, we present a predictive coding scheme that improves the coding performance of the quadtree decomposition by exploiting the correlation between each block of the quadtree. Experimental results show that the described technique improves the resulting quality of compressed depth images by 1.5-4 dB when compared to a JPEG-2000 encoder. Yannick Morvan, Dirk Farin, Peter H. N. de With |
ICIP (5) | 3 |
| 2007 | Floor-plan reconstruction from panoramic imagesabstractThe capturing of panoramic 360° images has become a popular photographic technique. While a panoramic image gives an impressive view of the environment, many people have difficulties to understand the spatial scene arrangement from this flat image. In this paper, we present a new visualization technique for panoramic images based on a coarse reconstruction of the indoor environment, in which the panorama was captured. Applications of this are, for example, real-estate or hotel advertising, featuring virtual tours through the apartment. We use a semi-automatic reconstruction process, in which the user marks the room corners in the panoramic images. These can be translated into viewing-angle measurements, from which our algorithms can compute the exact sizes of the walls, based on a pre-defined geometric model. Dirk Farin, Wolfgang Effelsberg, Peter H. N. de With |
ACM Multimedia | 3 |
| 2007 | A real-time augmented-reality system for sports broadcast video enhancementabstractThis paper presents a new augmented-reality system designed to generate visual enhancements for TV broadcasted court-net sports. A probabilistic method based on the Expectation Maximization (EM) procedure is utilized to find the optimal feature points, thereby enabling the automatic acquisition of the camera parameters from the TV image with high accuracy. A virtual camera derived from the original camera, helps to synthesize a variety of virtual scenes, such as the scene from the viewpoint of a player, depending on the intention of the user. To preserve the visual nature of the original human motion, the player's shape and texture are extracted from the real video and texture-mapped onto the virtual video. The system was tested over a set of court-net sports videos containing tennis, badminton and volleyball and demonstrated promising results. Jungong Han, Dirk Farin, Peter H. N. de With |
ACM Multimedia | 3 |
| 2007 | Generic 3-D Modeling for Content Analysis of Court-Net Sports Sequences
Jungong Han, Dirk Farin, Peter H. N. de With |
MMM (2) | 3 |
| 2007 | A Matching-Based Approach for Human Motion Analysis
Weilun Lao, Jungong Han, Peter H. N. de With |
MMM (2) | 3 |
| 2006 | Content-Based Model Template Adaptation and Real-Time System for Behavior Interpretation in Sports Video
Jungong Han, Peter H. N. de With |
ACIVS | 2 |
| 2006 | Blue Sky Detection for Picture Quality Enhancement
Bahman Zafarifar, Peter H. N. de With |
ACIVS | 2 |
| 2006 | Background Estimation and Adaptation Model with Light-Change Removal for Heavily Down-Sampled Video Surveillance SignalsabstractThis paper describes a background-subtraction system with light change-detection which works on a luminance QCIF-size video signal for surveillance applications. The new proposed pixel background model is controlled by a statistical threshold and is robust for cluttered background and small object motions. Moreover, (or light-change detection, we introduce temporal prediction of pixel values to estimate trends while quickly adapting to scene changes to facilitate a very sensitive detection of moving targets. Experiments show that a local contrast enhancement applied prior to down-sampling improves detection sensitivity, arid combined with the shifted sealed difference and me Wronskian determinant operators provides the best background/foreground detection. Sascha D. Cvetkovic, Peter Bakker, Johan Schirris, Peter H. N. de With |
ICIP | 4 |
| 2006 | Reuse of Motion Processing for Camera Stabilization and Video CodingabstractThe low bit rate of existing video encoders relies heavily on the accuracy of estimating actual motion in the input video sequence. In this paper, we propose a Video Stabilization and Encoding (ViSE) system to achieve a higher coding efficiency through a preceding motion processing stage (to the compression), of which the stabilization part should compensate for vibrating camera motion. The improved motion prediction is obtained by differentiating between the temporal coherent motion and a more noisy motion component which is orthogonal to the coherent one. The system compensates the latter undesirable motion, so that it is eliminated prior to video encoding. To reduce the computational complexity of integrating a digital stabilization algorithm with video encoding, we propose a system that reuses the already evaluated motion vector from the stabilization stage in the compression. As compared to H.264, our system shows a 14% reduction in bit rate yet obtaining an increase of about 0.5 dB in SNR. Bao Lei, Rene Klein Gunnewiek, Peter H. N. de With |
ICME | 3 |
| 2006 | Near-Future Streaming Framework for 3D-TV ApplicationsabstractThis paper presents a layered framework for 3D-TV applications, combining multiview and depth-image based approaches in a scalable fashion. To solve the problem of missing data due to disocclusions, we add specific layers for coded occlusion data and the edge-mask information for high-quality 3D rendering of key objects in the scene. We show how the same framework can be extended towards FTV applications by jointly addressing simulcast and multicast transmission. By adopting a distributed delivery architecture, new interesting properties can be realized such as shared processing for the creation and streaming of virtual viewpoints. Goran Petrovic, Peter H. N. de With |
ICME | 2 |
| 2006 | A Toolkit for Design and Performance Analysis of Real-Time Component-Based Software SystemsabstractSoftware tools supporting the design and analysis of complex software-intensive systems are highly desirable, since they enable earlier decision making about system realization. This paper presents a tooling environment that supports the design and performance analysis of time-critical component-based software architectures deployed on complex multiprocessor platforms. The tooling environment contains a set of integrated tools for (a) component storage and retrieval, (b) graphics-based design of software and hardware architectures, (c) performance analysis of the defined architectures and, (d) automated code generation. The cornerstone of the toolkit is a performance analyzer that provides efficient simulation of the designed architectures and enables design-time prediction of key performance properties like response time, data throughout, and usage of hardware resources (processor, memory and bus). For every architecture alternative, the performance predictions can be quickly obtained, thereby enabling a fast and yet broad design space exploration. We demonstrate the efficiency and robustness of this toolkit on a Car Radio Navigation benchmark case. Egor Bondarev, Michel R. V. Chaudron, Heorhiy Byelas, Peter H. N. de With |
ICSEA | 4 |
| 2006 | Realization of QoS management using negotiation algorithms for multiprocessor NoCabstractWith the growing complexity of new multimedia applications and the parallel execution enabled by multiprocessor architectures, quality-of-service (QoS) management for modern network-on-chip (NoC) systems becomes relevant. This paper presents an approach for efficient usage of platform resources and simultaneously controlling overall system performance, when multiple applications are active. Therefore, on top of the conventional QoS, we introduce an algorithm for setting quality levels of parallel executed applications on a multiprocessor network-on-chip. The new concept is a hierarchical system, where an application manager negotiates with the resource management to jointly optimize the quality of individual and overall application settings. The proposed QoS manager concept was successfully evaluated by an experimental set-up, employing two independent video-object decoders, a background sprite decoder from the MPEG-4 standard and an abstract audio object. This simultaneous execution of four independent applications was mapped on a simulator based on AEthereal network and eight ARM7 processors executed with a clock-cycle true simulator. The proposed distributed QoS management decreases the complexity of the overall management and improves the transparency and responsibilities in the decision taking. Milan Pastrnak, Peter H. N. de With, Jef L. van Meerbergen |
ISCAS | 2 |
| 2006 | Enabling arbitrary rotational camera motion using multisprites with minimum coding costabstractObject-oriented coding in the MPEG-4 standard enables the separate processing of foreground objects and the scene background (sprite). Since the background sprite only has to be sent once, transmission bandwidth can be saved. We have found that the counter-intuitive approach of splitting the background into several independent parts can reduce the overall amount of data. Furthermore, we show that in the general case, the synthesis of a single background sprite is even impossible and that the scene background must be sent as multiple sprites instead. For this reason, we propose an algorithm that provides an optimal partitioning of a video sequence into independent background sprites (a multisprite), resulting in a significant reduction of the involved coding cost. Additionally, our sprite-generation algorithm ensures that the sprite resolution is kept high enough to preserve all details of the input sequence, which is a problem especially during camera zoom-in operations. Even though our sprite generation algorithm creates multiple sprites instead of only a single background sprite, it is fully compatible with the existing MPEG-4 standard. The algorithm has been evaluated with several test sequences, including the well-known Table-tennis and Stefan sequences. The total coding cost for the sprite VOP is reduced by a factor of about 2.6 or even higher, depending on the sequence. Dirk Farin, Peter H. N. de With |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2005 | Fast Face Detection Using a Cascade of Neural Network Ensembles
Fei Zuo, Peter H. N. de With |
ACIVS | 2 |
| 2005 | Multistage Face Recognition Using Adaptive Feature Selection and Classification
Fei Zuo, Peter H. N. de With, Michiel van der Veen |
ACIVS | 2 |
| 2005 | Facial feature extraction using a cascade of model-based algorithmsabstractWe present a cascaded framework for robust and accurate facial feature extraction. In this framework, we propose the following three model-based algorithms: (1) constrained global deformation using a sparse feature representation; (2) component texture fitting using direct parameter estimation by SVR, and (3) component feature refinement by direct optimization. The algorithms capture different characteristics of facial features, giving various extraction performances in terms of robustness (convergence) and accuracy. To achieve both high accuracy and robustness, we cascade these algorithms into a chain, where each algorithm progressively 'pulls' the model closer to the correct position. Experiments show that the combined algorithm achieves a large convergence area and high accuracy. Fei Zuo, Peter H. N. de With |
AVSS | 2 |
| 2005 | Misregistration errors in change detection algorithms and how to avoid themabstractBackground subtraction is a popular algorithm for video object segmentation. It identifies foreground objects by comparing the input images with a pure background image. In camera-motion compensated sequences, small errors in the motion estimation can lead to large image differences along sharp edges. Consequently, the errors in the image registration finally lead to segmentation errors. This paper proposes a computationally efficient approach to detect image areas having a high risk of showing misregistration errors. Furthermore, we describe how existing change detection algorithms can be modified to avoid segmentation errors in these areas. Experiments show that our algorithm can improve the segmentation quality. The algorithm is memory efficient and suitable for real-time processing. Dirk Farin, Peter H. N. de With |
ICIP (2) | 2 |
| 2005 | Fast camera calibration for the analysis of sport sequencesabstractSemantic analysis of sport sequences requires camera calibration to obtain player and ball positions in real-world coordinates. For court sports like tennis, the marker lines on the field can be used to determine the calibration parameters. We propose a real-time calibration algorithm that can be applied to all court sports simply by exchanging the court model. The algorithm is based on (1) a specialized court-line detector, (2) a RANSAC-based line parameter estimation, (3) a combinatorial optimization step to localize the court within the set of detected line segments, and (4) an iterative court-model tracking step. Our results show real-time calibration of, e.g., tennis and soccer sequences with a computation time of only about 6 ms per frame. Dirk Farin, Jungong Han, Peter H. N. de With |
ICME | 3 |
| 2005 | Real-Time and Distributed AV Content Analysis System for Consumer Electronics NetworksabstractThe ever-increasing complexity of generic multimedia-content-analysis-based (MCA) solutions, their processing power demanding nature and the need to prototype and assess solutions in a fast and cost-saving manner motivated the development of the Cassandra framework. The combination of state-of-the-art network and grid-computing solutions and recently standardized interfaces facilitated the set-up of this framework, forming the basis for multiple cross-domain and cross-organizational collaborations. It enables distributed computing scenario simulations for e.g. distributed content analysis (DCA) across consumer electronics (CE) in-home networks, but also the rapid development and assessment of complex multi-MCA-algorithm-based applications and system solutions. Furthermore, the framework's modular nature-logical MCA units are wrapped into so-called service units (SU)-ease the split between system-architecture- and algorithmic-related work and additionally facilitate reusability, extensibility and upgrade ability of those SUs Jan Nesvadba, Pedro Fonseca 0002, Alexander Sinitsyn, Fons de Lange, Martijn Thijssen, Patrick van Kaam, Hong Liu 0008, Rien van Leeuwen, Johan J. Lukkien, Andrei Korostelev, Jan Ypma, Bart Kroon, Hasan Celik, Alan Hanjalic, Suphi Umut Naci, Jenny Benois-Pineau, Peter H. N. de With, Jungong Han |
ICME | 17 |
| 2005 | Extended abstract: estimation times of on-chip multiprocessor stream-oriented applicationsabstractThis paper focuses on stream-oriented applications with real-time constraints for on-chip multiprocessors, e.g. video/audio coding. In such an application the same function, e.g. frame decoding, is executed over and over again. In general, the execution time is data-dependent. Thus, at run-time a situation may arise when the execution time exceeds the deadline specified in the real-time constraints and a preventive action should be taken. For stream-oriented applications pipelined execution of the function is very important for achieving the required throughput. An accurate execution time estimation method supporting pipelined execution on a multiprocessor architecture is proposed in this paper Peter Poplavko, Twan Basten, Milan Pastrnak, Jef L. van Meerbergen, Marco Bekooij, Peter H. N. de With |
MEMOCODE | 6 |
| 2005 | Successful Architecture for Short Message Service CenterabstractThis paper presents and analyzes the key architectural decisions in the design of a successful Short Message Service Center as part of a GSM network. Eltjo R. Poort, Hans Adriaanse, Arie Kuijt, Peter H. N. de With |
WICSA | 4 |
| 2004 | Corridor scissors: a semi-automatic segmentation tool employing minimum-cost circular pathsabstractWe present a new semiautomatic segmentation tool, which is motivated by the intelligent scissors algorithm, but which uses a modified concept of user-interaction. This new interface provides better capabilities for modifying previous segmentation results. The advantage of the new approach is that it enables to gradually increase the quality of the segmentation. The segmentation tool is based on a shortest circular path search within a corridor that is drawn by the user along the object boundary. For this purpose, we present a new algorithm for computing the shortest circular paths. Our algorithm is so fast that it almost achieves the speed of a regular noncircular shortest path search, while still ensuring an optimal solution. Dirk Farin, Magnus Pfeffer, Peter H. N. de With, Wolfgang Effelsberg |
ICIP | 3 |
| 2004 | Fast facial feature extraction using a deformable shape model with haar-wavelet based local texture attributes
Fei Zuo, Peter H. N. de With |
ICIP | 2 |
| 2004 | Video-object segmentation using multi-sprite background subtractionabstractThe background subtraction algorithm is a frequently-used object segmentation technique because of its algorithmic simplicity. However, we show that, for general rotational camera-motion, it is impractical, or even impossible, to use a single background image. As a solution, we propose to use multi-sprite backgrounds which enables processing of arbitrary rotational camera-motion. The paper describes a complete video-object segmentation system employing multi-sprites. The system generates object masks and background sprites that are compatible with MPEG-4 object-oriented video-coding tools. Good segmentation results are also obtained for sequences which cannot be processed with ordinary background images. Dirk Farin, Peter H. N. de With, Wolfgang Effelsberg |
ICME | 2 |
| 2004 | Real-time facial feature extraction using statistical shape model and Haar-wavelet based feature searchabstractWe propose a fast facial feature extraction technique for an embedded face recognition system. The novel key element is a combination of a statistical shape model and the application of a Haar-wavelet based feature matching. Our statistical face model is based on the active shape model (ASM). However ASM lacks robustness to illumination changes and it has a limited convergence area. Instead of a 1D profile analysis, we propose a 2D texture pattern search-and-fitting scheme, which provides more robustness and faster convergence than conventional ASM. Furthermore, we employ Haar-wavelets to model local-facial textures, which yields two improvements: faster processing and more robustness with respect to low-quality images. Our proposed approach shows good results dealing with test face images, which are quite dissimilar with the faces used for statistical training. The convergence area of our proposed method almost quadruples compared to ASM, and the extraction accuracy is also improved. The total processing requires 30 - 70 ms, which is comparable to ASM, but faster than the active appearance model (AAM). Fei Zuo, Peter H. N. de With |
ICME | 2 |
| 2004 | Image enhancement circuit using nonlinear processing curve and constrained histogram range equalizationabstractFor real-time imaging in surveillance applications, image fidelity is of primary importance to ensure customer confidence. The obtained image fidelity is a result from amongst others dynamic range expansion and video signal enhancement. The dynamic range of the signal needs adaptation, because the sensor signal has a much larger range than the standard CRT display. The signal enhancement should accommodate for the widely varying light and scene conditions and user scenarios of the equipment. This paper proposes a new system to combine dynamic range and enhancement processing, offering a strongly improved picture quality for surveillance applications. The key to our solution is that we use Non-Linear Processing (NLP) with a so-called Constrained Histogram Range Equalization (CHRE). The NLP transforms the digitized high-dynamic luminance sensor signal such that details of the low-luminance parts are enhanced, while avoiding detail losses in the high-luminance areas. The CHRE technique enhances visibility of the global contrast for the camera signal without significant information loss in the statistically less relevant areas. Evaluations of this proposal have shown clear improvements of the perceptual image quality. An additional advantage is that the new scheme is adaptable and allows the concatenation of further enhancement techniques without sacrificing the obtained picture quality improvement. Sascha D. Cvetkovic, Peter H. N. de With |
VCIP | 2 |
| 2004 | Minimizing MPEG-4 sprite coding cost using multi-spritesabstractObject-oriented coding in the MPEG-4 standard enables the separate processing of foreground objects and the scene background (sprite). Since the background sprite only has to be sent once, transmission bandwidth can be saved. This paper shows that the concept of merging several views of a non-changing scene background into a single background sprite is usually not the most efficient way to transmit the background image. We have found that the counter-intuitive approach of splitting the background into several independent parts can reduce the overall amount of data. For this reason, we propose an algorithm that provides an optimal partitioning of a video sequence into independent background sprites (a multi-sprite), resulting in a significant reduction of the involved coding cost. Additionally, our algorithm results in background sprites with better quality by ensuring that the sprite resolution has at least the final display resolution throughout the sequence. Even though our sprite generation algorithm creates multiple sprites instead of a single background sprite, it is fully compatible with the existing MPEG-4 standard. The algorithm has been evaluated with several test-sequences, including the well-known Table-tennis and Stefan sequences. The total coding cost could be reduced by factors of about 2.7 or even higher. Dirk Farin, Peter H. N. de With, Wolfgang Effelsberg |
VCIP | 2 |
| 2004 | Resource-aware complexity scalability for mobile MPEG encodingabstractComplexity scalability attempts to scale the required resources of an algorithm with the chose quality settings, in order to broaden the application range. In this paper, we present complexity-scalable MPEG encoding of which the core processing modules are modified for scalability. Scalability is basically achieved for the computational complexity by varying the number of computed DCT coefficients and the number of evaluated motion vectors, while other modules are designed such that they scale with the previous parameters. Resource usage such as power-consuming memory accesses scale accordingly. The interdependencies of the scalable modules and the system performance are evaluated. Experimental results show scalability giving a smooth change in complexity and coresponding video quality. The elapsed execution time of the scalable encoder, reflecting the computational complexity, can be gradually reduced to roughly 50% of its execution time when operating at high quality. The video quality scaled betwee 21.5 dB and 38.5 dB PSNR for different sequences targeting 1500 kbps. The implemented encoder and the scalability techniques can be successfully applied in mobile systems based on MPEG video compression, and also has advantages in multi-tasking environments like high-end TV sets. The obtained scalability techniques can be applied to other coding standards like MPEG-4 and H.264 as well. Stephan Mietens, Peter H. N. de With, Christian Hentschel |
VCIP | 2 |
| 2004 | Multistage feature extraction for accurate face alignmentabstractWe propose a novel multistage facial feature extraction approach using a combination of 'global' and 'local' techniques. At the first stage, we use template matching, based on an Edge-Orientation-Map for fast feature position estimation. Using this result, a statistical framework applying the Active Shape Model (ASM) is initialized and deformed to fit the real face image. In our proposal, we use a 2-D pattern search-and-fitting scheme guiding the deformation process, which provides more robustness and faster convergence than the traditional ASM. Our proposed approach for feature extraction shows good results dealing with a test set composed of faces images which are quite dissimilar with the faces used for the statistical training of the face model. The convergence area of our proposed technique almost quadruples compared to the ASM, while the amount of faces doubles for which the convergence is reached. The total processing for feature extraction takes less than 1 second for 250x250 face images on a Pentium-IV PC (1.7GHz). Fei Zuo, Peter H. N. de With |
VCIP | 2 |
| 2004 | Resolving Requirement Conflicts through Non-Functional DecompositionabstractA lack of insight into the relationship between (non) functional requirements and architectural solutions often leads to problems in real life projects. This paper presents a model that concentrates on the mapping of nonfunctional requirements onto functional requirements for architecture design. We build a framework that both provides a model and a repeatable method to transform conflicting requirements into a system decomposition. This paper presents the framework, and discusses two cases onto which the method is applied. In one case, the method is successfully used to reconstruct the high-level structure of a system from its requirements. The second case is one in which the method was actually used to create a system design fitting the stakeholders' needs, and that is reproducible from its requirements. Eltjo R. Poort, Peter H. N. de With |
WICSA | 2 |
| 2003 | Robust background estimation for complex video sequencesabstractKnowing the background image of a video scene simplifies the general video-object segmentation problem and therefore it is required by several automatic segmentation algorithms. This paper presents a new background estimation algorithm which is applicable to complex video sequences where many objects are simultaneously visible and the background is visible for a short time period only. The algorithm applies a rough segmentation of the input images into foreground and background regions to exclude the foreground objects from background synthesis. This prevents a bias of the synthesized background image towards the color of foreground objects. Experiments show that the obtained background images differ significantly less from the real background than those obtained with previous algorithms. Dirk Farin, Peter H. N. de With, Wolfgang Effelsberg |
ICIP (1) | 2 |
| 2003 | Evolution of a Software Maintenance Organization from Cost Center to Service CenterabstractThe paper describes experiences with the evolution of a software maintenance organization for digital set-top boxes of a leading electronics company from a cost center towards a service center. Several years ago a dedicated software maintenance group was constituted. As the costs for software maintenance were not recovered from the customers, the software maintenance group was merely considered a cost center. Through starting a metrics program for software maintenance and defining a service strategy with various service levels, the software maintenance group generated sufficient revenues to become self-supporting. An important conclusion is that the use of ITIL (IT infrastructure library) service support has helped to develop a better customer focused approach, which is considered as the most important critical success factor for a professional, self-supporting maintenance organization. Sander Smit, Peter H. N. de With, Gert-Jan van Dijk |
ICSM | 2 |
| 2003 | A segmentation system with model-assisted completion of video objects
Dirk Farin, Peter H. N. de With, Wolfgang Effelsberg |
VCIP | 2 |
| 2003 | Toward fast feature adaptation and localization for real-time face recognition systems
Fei Zuo, Peter H. N. de With |
VCIP | 2 |
| 2002 | New flexible motion estimation technique for scalable MPEG encoding using display frame order and multi-temporal referencesabstractThe applicability of MPEG video coding can be improved by scaling both the algorithmic complexity and resource usage appropriately for the intended device and application. For this purpose, we present a new technique for motion estimation, based on a scalable three-stage process including frame processing in display order, approximation of motion vector fields using multiple references and optional quality refinements. Experiments show that the computational effort is scalable with a factor of 14, resulting in a global variation of 7 dB SNR in picture quality. At full processing, our technique slightly outperforms a 32/spl times/32 full search motion estimation. The technique forms a valuable contribution to mobile MPEG coding applications, following the scalability concepts introduced by Mietens, de With and Hentschel (see IEEE Int. Conf. on Image Proc. (ICIP 2001), vol.3, p.462-465, Oct. 2001). Stephan Mietens, Gerben J. Hekstra, Peter H. N. de With, Christian Hentschel |
ICIP (1) | 3 |
| 2002 | Robust clustering-based video-summarization with integration of domain-knowledgeabstractClustering techniques have been widely used in automatic video-summarization applications to group shots with comparable content. We enhance the popular k-means clustering algorithm to integrate user-supplied domain-knowledge into the cluster generation step. This provides a convenient way to exclude scenes from the summary which are a-priori known to be irrelevant. Furthermore, we added an additional, time-constrained clustering step preceding the scene clustering step to exclude short ranges with transitional content. This makes the algorithm robust to fading and wipe-effects in the input without requiring explicit cut detection. Dirk Farin, Wolfgang Effelsberg, Peter H. N. de With |
ICME (1) | 3 |
| 2002 | New scalable three-stage motion estimation technique for mobile MPEG encodingabstractThe paper presents a new scalable three-stage motion estimation technique, which includes processing of frames in display order and approximating motion vector fields using multiple references. Quality refinement is added as an optional stage. The complete system provides a flexible framework with a large scalability range in computational effort, resulting in different picture-quality levels or bitrates. Experiments show a scalable computational effort with a factor of 14, resulting in a global variation of 7 dB SNR in picture quality (with the "Stefan" sequence). In high-quality operation, the new method is comparable to a full-search motion estimation with a search window of 32/spl times/32 pixels (or even outperforms it). The innovation provides an excellent starting point for scalability in resource-constrained mobile system design (see Mietens, S. et al., IEEE Int. Conf. on Image Proc., ICIP 2001, vol.3, p.462-5, 2001). Stephan Mietens, Peter H. N. de With, Christian Hentschel |
ICME (1) | 2 |
| 2001 | New DCT computation algorithm for video quality scalingabstractThe application of video coding systems, such as MPEG, in portable systems like organizers and mobile phones can be scaled down to a reduced complexity that matches with the desired video quality and/or display. A new DCT computation algorithm is presented, based on an analysis for optimizing the number of computations involved at each computing stage, using existing fast DCT calculation algorithms. The analysis is used to scale down the video quality, thereby lowering the computing power and resource usage. Compared to a diagonally oriented computation of coefficients that matches with the conventional MPEG scanning, a 2 to 4 dB SNR improvement is obtained when scaling down the video quality to half the computing resources. Stephan Mietens, Peter H. N. de With, Christian Hentschel |
ICIP (3) | 2 |
| 2001 | SAMPEG: a scene-adaptive parallel MPEG-2 software encoder
Dirk Farin, Niels Mache, Peter H. N. de With |
VCIP | 3 |
| 2000 | Synchronization of video in distributed computing systems
Egbert G. T. Jaspers, Bert S. Visser, Peter H. N. de With |
VCIP | 3 |
| 1999 | Architecture of Embedded Video Processing in a Multimedia Chip-SetabstractA new chip-set for video display processing in a consumer television or set-top box is presented. Key aspect of the chip-set is a high flexibility and programmability of multi-window features with, for example, full-motion video, Internet and Teletext. To provide a large amount of computational power for such a full-featured application domain and to prevent the system from a communication bottleneck to the external memory, a heterogenous multi-processor architecture is implemented. The architecture offers a minimum of external communication overhead and enables programming on high functional level. The chip-set can simultaneously display e.g. two full-motion video windows, an Internet page and an additional mail indicator in front of a pixel-based wallpaper background. In addition, the video can be noise reduced and enhanced in sharpness together with a 50-100 Hz conversion to reduce field flicker. Egbert G. T. Jaspers, Peter H. N. de With |
ICIP (2) | 2 |
| 1992 | Data Compression Systems for Home-Use Digital Video RecordingabstractThe authors focus on image data compression techniques for digital recording. Image coding for storage equipment covers a large variety of systems because the applications differ considerably in nature. Video coding systems suitable for digital TV and HDTV recording and digital electronic still picture storage are considered. In addition, attention is paid to picture coding for interactive systems, such as the compact-disc interactive system. The relation between the recording system boundary conditions and the applied coding techniques is outlined. The main emphasis is on picture coding techniques for digital consumer recording.> Peter H. N. de With, Marcel Breeuwer, Peter A. M. van Grinsven |
IEEE J. Sel. Areas Commun. | 1 |
| 1992 | Digital consumer HDTV recording based on motion-compensated DCT coding of video signals
Peter H. N. de With, A. M. A. Rijckaert, J. Kaaden, H.-W. Keesen |
Signal Process. Image Commun. | 1 |
| 1990 | Source coding of HDTV with compatibility to TVabstractGradual introduction of HDTV is considered to be important. In this paper, a bit-rate reduction system is introduced, which decreases the bit rate of digital HDTV from 664 Mbit/s to about 80 Mbit/s, while ensuring compatibility with TV. The system is based on first subband splitting the interlaced HDTV signal into an interlaced compatible TV signal and three surplus signals. Then, the compatible TV signal is coded with intraframe DCT coding, whereas the surplus signals are coded with quantization, variable-length coding and runlength coding. Marcel Breeuwer, Peter H. N. de With |
VCIP | 2 |