Cláudio R. Jung

dblp:07/5754 · also Cláudio Rosito Jung · DBLP profile ↗
← Back
82ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-4711-5783ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 62 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 20 · 6 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1
YearPublicationVenuePosition
2026 Neurosymbolic computer vision: a survey and perspective
Márcio Nicolau, Cláudio R. Jung, Luís C. Lamb
Neural Comput. Appl.2
2025 GauCho: Gaussian Distributions with Cholesky Decomposition for Oriented Object Detection
abstract
Oriented Object Detection (OOD) has received increased attention in the past years, being a suitable solution for detecting elongated objects in remote sensing analysis. In particular, using regression loss functions based on Gaussian distributions has become attractive since they yield simple and differentiable terms. However, existing solutions are still based on regression heads that produce Oriented Bounding Boxes (OBBs), and the known problem of angular boundary discontinuity persists. In this work, we propose a regression head for OOD that directly produces Gaussian distributions based on the Cholesky matrix decomposition. The proposed head, named GauCho, theoretically mitigates the boundary discontinuity problem and is fully compatible with recent Gaussian-based regression loss functions. Furthermore, we advocate using Oriented Ellipses (OEs) to represent oriented objects, which relates to GauCho through a bijective function and alleviates the encoding ambiguity problem for circular objects. Our experimental results show that GauCho can be a viable alternative to the traditional OBB head, achieving results comparable to or better than state-of-the-art detectors for the challenging dataset DOTA. Our code will be available at https://github.com/jhlmarques/GauCho.
Jose Henrique Lima Marques, Jeffri Murrugarra-Llerena, Cláudio R. Jung
CVPR3
2025 Anchor-Based Gravity Alignment for Panoramas
abstract
A crucial step for improving the usability of panoramas is the accurate alignment of the upright vector to match the direction of gravity. This paper introduces an anchor-based method for estimating the upright vector in panoramas. Our method, named AnchorsUp, combines the strengths of classification and regression by selecting a rough estimate of the upright vector from a predefined set of anchor vectors, followed by a fine-grained adjustment through regression to achieve precise alignment. We evaluate AnchorsUp on the SUN360 dataset and show superior performance compared to state-of-the-art methods, particularly for angular errors below 5 and 12 degrees, considered the most relevant in practical applications. The source code and dataset information are available at https://github.com/mabergmann/anchorsup.
Matheus A. Bergmann, Rômulo Marconato Stringhini, Thiago L. T. da Silveira, Cláudio R. Jung
ICIP4
2025 PetsRS - a Dataset and Benchmark for Pet Recognition on a Climate Disaster Scenario
abstract
In the first half of 2024, thousands of pets were separated from their guardians due to the floods in Rio Grande do Sul, Brazil. Tools that help guardians search for their lost animals are pivotal in such situations. Such a tool might allow users (in this case, guardians or volunteers from shelters) to send images of the animals, allowing manual or automated searches to match lost and rescued animals. This work presents a novel dataset – the PetsRS dataset1– containing organized tuples of pictures of dogs and cats for pet retrieval. We show pet recognition results using pre-trained off-the-shelf image encoders and the K-NN algorithm with embeddings as data points. In our results, DINOv2 achieves the best recall@K on all tests. This work aims to create a benchmark and baseline for evaluating techniques that tackle pet recognition.
Paulo G. L. Pinto, Thiago L. T. da Silveira, Cláudio R. Jung
ICIP3
2025 Oriented Cell Dataset: A Dataset and Benchmark for Oriented Cell Detection and Applications
abstract
This work presents a new public dataset for cell detection in bright-field microscopy images annotated with Oriented Bounding Boxes (OBBs), named Oriented Cell Dataset (OCD). Our dataset also contains a subset of images with five independent expert annotations, which allows inter-annotation analysis to determine a suitable IoU acceptance threshold for evaluating cell detectors. We show that OBBs and a derived representation, Oriented Ellipses (OEs), provide a more accurate shape representation than standard Horizontal Bounding Boxes (HBBs) with a slight overhead of one extra click in the annotation process. We benchmarked OCD using 14 state-of-the-art oriented object detectors, and explored two main problems in cancer biology: cell confluence and polarity determination. Our code and dataset are available at https://github.com/LucasKirsten/Deep-Cell-Tracking-EBB.
Lucas N. Kirsten, Angelo Angonezi, Jose Marques, Fernanda Oliveira, Juliano Faccioni, Camila Cassel, Débora Santos de Sousa, Samlai Vedovatto, Guido Lenz, Cláudio R. Jung
WACV10
2025 Noise-Aware Evaluation of Object Detectors
abstract
Supervised object detection requires annotated datasets for training and evaluation purposes. However, human annotation of large datasets is error-prone, and frequent mistakes are erroneous labels, missing objects, and impre-cise bounding boxes. The main goals of this work are to quantify the extent of annotation noise in terms of corner-wise discrepancies, assess how it impacts evaluation met-rics for object detection, and propose noise-aware alter-natives that serve as upper and lower bounds for a base-line metric. We focus our analysis on the Microsoft COCO dataset and re-evaluate several state-of-the-art object de-tectors using the proposed metrics. We show that the Average Precision (AP) metric might be considerably over or under-estimated, particularly for small objects and restrictive IoU acceptance thresholds. Our code is available at https://github.com/Artcsl/Error-Aware.
Jeffri Murrugarra-Llerena, Cláudio R. Jung
WACV2
2025 Cell tracking-by-detection using elliptical bounding boxes
Lucas N. Kirsten, Cláudio R. Jung
J. Vis. Commun. Image Represent.2
2024 Single-Panorama Classification of 3D Objects Using Horizontally Stacked Dilated Convolutions
abstract
This paper presents a single-image approach for classifying 3D objects represented as meshes. Our method centers a virtual spherical camera at the object’s centroid and casts omnidirectional rays. Then, it computes local geometry information of each ray’s first and last intersection points, generating a single multi-channel equirectangular (ERP) image per object. We propose a convolutional block named Horizontally Stacked Dilated Convolution (HSDC) to handle ERP distortions and introduce a classifier built upon these blocks. Our experiments in popular datasets show that the results produced by our method are competitive or better than state-of-the-art voxel-and point-based methods, being the best among single-view approaches. Code is available at https://github.com/rmstringhini/HSDCNet.
Rômulo Marconato Stringhini, Thiago S. Lermen, Thiago L. T. da Silveira, Cláudio R. Jung
ICIP4
2024 Probabilistic Intersection-Over-Union for Training and Evaluation of Oriented Object Detectors
abstract
Oriented object detection is a challenging and relatively new problem. Most existing approaches are based on deep learning and explore Oriented Bounding Boxes (OBBs) to represent the objects. They are typically based on adaptations of traditional detectors that work with Horizontal Bounding Boxes (HBBs), which have been exploring IoU-like loss functions to regress the HBBs. However, extending this idea for OBBs is challenging due to complex formulations or requirement for customized backpropagation implementations. Furthermore, using OBBs presents limitations for irregular or roughly circular objects, since the definition of the ideal OBB is an ambiguous and ill-posed problem. In this work, we jointly tackle the problem of training, representing, and evaluating oriented detectors. We explore Gaussian distributions-called Gaussian Bounding Boxes (GBBs)-as fuzzy representations for oriented objects and propose using a similarity metric between two GBBs based on the Hellinger distance. We show that this metric leads to a differentiable closed-form expression that can be directly used as a localization loss term to train OBB object detectors. We also show that GBBs present a natural representation as elliptical regions (called EBBs), which inherently mitigate ambiguity representation for circular objects. Finally, we empirically show that the proposed similarity metric computed between two GBBs strongly correlates with the IoU between the corresponding EBBs, motivating the name Probabilistic Intersection-over-Union (ProbIoU). Our experiments show that results using ProbIoU as a regression loss are competitive with state-of-the-art alternatives without requiring additional hyperparameters or customized implementations, and that ProbIoU is a promising alternative to evaluate oriented object detectors. Our code is available at https://github.com/ProbIOU/.
Jeffri Murrugarra-Llerena, Lucas N. Kirsten, Luis Felipe de Araujo Zeni, Cláudio R. Jung
IEEE Trans. Image Process.4
2023 Omnidirectional visual computing: Foundations, challenges, and applications
Thiago L. T. da Silveira, Cláudio R. Jung
Comput. Graph.2
2023 Omnidirectional 2.5D representation for COVID-19 diagnosis using chest CTs
Thiago L. T. da Silveira, Paulo G. L. Pinto, Thiago S. Lermen, Cláudio R. Jung
J. Vis. Commun. Image Represent.4
2022 SVBR-Net: A Non-Blind Spatially Varying Defocus Blur Removal Network
abstract
Defocus blur is a physical consequence of the optical sensors used in most cameras. Although it can be used as a photographic style, it is commonly viewed as an image degradation modeled as the convolution of a sharp image with a spatially-varying blur kernel. Motivated by the advance of blur estimation methods in the past years, we propose a non-blind approach for image deblurring that can deal with spatially-varying kernels. We introduce two encoder-decoder sub-networks that are fed with the blurry image and the estimated blur map, respectively, and produce as output the deblurred (deconvolved) image. Each sub-network presents several skip connections that allow data propagation from layers spread apart, and also inter-subnetwork skip connections that ease the communication between the modules. The network is trained with synthetically blur kernels that are augmented to emulate blur maps produced by existing blur estimation methods, and our experimental results show that our method works well when combined with a variety of blur estimation methods. Our code will be available at https: //github.com/alikaraali/SVBR_Net_icip.
Ali Karaali, Cláudio R. Jung
ICIP2
2022 A Temporally Coherent Background Model for DIBR View Synthesis
abstract
Aligned color and depth images allow exploring a scene by projecting the virtual camera to arbitrary positions. Depth-image-based rendering (DIBR) methods are often adopted in practical solutions, requiring fewer data than other approaches that explore stereo or multi-view setups. As the virtual camera moves, artifacts are expected to appear, which must be tackled adequately. For video sequences, not only spatial but also temporal inconsistencies affect the user experience. This paper proposes a background model that can be coupled to reduce the flickering of state-of-the-art still-image DIBR methods applied to video sequences captured by static cameras. Qualitative and quantitative results show the potential of the use of our method in practical applications.
Adriano Q. de Oliveira, Thiago L. T. da Silveira, Marcelo Walter, Cláudio R. Jung
ICIP4
2022 Deep Multi-Scale Feature Learning for Defocus Blur Estimation
abstract
This paper presents an edge-based defocus blur estimation method from a single defocused image. We first distinguish edges that lie at depth discontinuities (called depth edges, for which the blur estimate is ambiguous) from edges that lie at approximately constant depth regions (called pattern edges, for which the blur estimate is well-defined). Then, we estimate the defocus blur amount at pattern edges only, and explore an interpolation scheme based on guided filters that prevents data propagation across the detected depth edges to obtain a dense blur map with well-defined object boundaries. Both tasks (edge classification and blur estimation) are performed by deep convolutional neural networks (CNNs) that share weights to learn meaningful local features from multi-scale patches centered at edge locations. Experiments on naturally defocused images show that the proposed method presents qualitative and quantitative results that outperform state-of-the-art (SOTA) methods, with a good compromise between running time and accuracy.
Ali Karaali, Naomi Harte, Cláudio R. Jung
IEEE Trans. Image Process.3
2022 A Flexible Approach for Automatic License Plate Recognition in Unconstrained Scenarios
abstract
Automatic License Plate Recognition is a crucial task for several applications related to Intelligent Transportation Systems, from access control to traffic monitoring. Most existing approaches are focused on a specific setup (e.g., toll control) or a single license plate (LP) region (e.g., European, US, Brazilian, Taiwanese, etc.), which limits their application. This work proposes a complete ALPR system focusing on unconstrained capture scenarios, where the LP might be considerably distorted due to oblique views. We present an Improved Warped Planar Object Detection Network (IWPOD-NET) that is able to detect the four corners of an LP in a variety of conditions, so that it can be warped to a fronto-parallel view and alleviate perspective-related distortions. Given the rectified LP, we test two different Optical Character Recognition (OCR) methods based on object detection. Our experimental results show that the proposed detector is competitive with state-of-the-art (SOTA) methods using a very limited training set. Regarding the full ALPR results, our method achieves top-scoring results for several datasets that include a variety of capture conditions and vehicle types (in particular, motorcycles).
Sergio Montazzolli Silva, Cláudio R. Jung
IEEE Trans. Intell. Transp. Syst.2
2021 Improving the Classification of Rare Chords With Unlabeled Data
abstract
In this work, we explore techniques to improve performance for rare classes in the task of Automatic Chord Recognition (ACR). We first explored the use of the focal loss in the context of ACR, which was originally proposed to improve the classification of hard samples. In parallel, we adapted a self-learning technique originally designed for image recognition to the musical domain. Our experiments show that both approaches individually (and their combination) improve the recognition of rare chords, but using only self-learning with noise addition yields the best results.
Marcelo Bortolozzo, Rodrigo Schramm, Cláudio R. Jung
ICASSP3
2021 Crowd flow estimation from calibrated cameras
Igor Rodrigues de Almeida, Cláudio R. Jung
Mach. Vis. Appl.2
2021 Fast and accurate superpixel algorithms for 360∘ images
Thiago L. T. da Silveira, Adriano Q. de Oliveira, Marcelo Walter, Cláudio R. Jung
Signal Process.4
2021 A Hierarchical Superpixel-Based Approach for DIBR View Synthesis
abstract
View synthesis allows observers to explore static scenes using aligned color images and depth maps captured in a preset camera path. Among the options, depth-image-based rendering (DIBR) approaches have been effective and efficient since only one pair of color and depth map is required, saving storage and bandwidth. The present work proposes a novel DIBR pipeline for view synthesis that properly tackles the different artifacts that arise from 3D warping, such as cracks, disocclusions, ghosts, and out-of-field areas. A key aspect of our contributions relies on the adaptation and usage of a hierarchical image superpixel algorithm that helps to maintain structural characteristics of the scene during image reconstruction. We compare our approach with state-of-the-art methods and show that it attains the best average results in two common assessment metrics under public still-image and video-sequence datasets. Visual results are also provided, illustrating the potential of our technique in real-world applications.
Adriano Q. de Oliveira, Thiago L. T. da Silveira, Marcelo Walter, Cláudio R. Jung
IEEE Trans. Image Process.4
2020 Real-time license plate detection and recognition using deep convolutional neural networks
Sergio Montazzolli Silva, Cláudio R. Jung
J. Vis. Commun. Image Represent.2
2019 Optimal Group Distribution based on Thermal and Psycho-Social Aspects
abstract
In crowds, one important aspect that has been studied in literature is the sociability of groups dealing with aspects based on personality and emotions. In this paper we contribute to the space design area while considering the cultural, personality and thermal aspects to provide spatial group distribution. Our method applies a thermal comfort method together with cultural and personality model to optimally distribute the groups in a virtual environment. Results indicate that obtained groups distribution are coherent with expected based on literature.
Paulo Knob, Gabriel Rockenbach, Cláudio R. Jung, Soraia Raupp Musse
CASA3
2019 Perturbation Analysis of the 8-Point Algorithm: A Case Study for Wide FoV Cameras
abstract
This paper presents a perturbation analysis for the estimate of epipolar matrices using the 8-Point Algorithm (8-PA). Our approach explores existing bounds for singular subspaces and relates them to the 8-PA, without assuming any kind of error distribution for the matched features. In particular, if we use unit vectors as homogeneous image coordinates, we show that having a wide spatial distribution of matched features in both views tends to generate lower error bounds for the epipolar matrix error. Our experimental validation indicates that the bounds and the effective errors tend to decrease as the camera Field of View (FoV) increases, and that using the 8-PA for spherical images (that present 360°x180° FoV) leads to accurate essential matrices. As an additional contribution, we present bounds for the direction of the translation vector extracted from the essential matrix based on singular subspace analysis.
Thiago L. T. da Silveira, Cláudio R. Jung
CVPR2
2019 On the Performance of DIBR Methods When Using Depth Maps from State-of-the-art Stereo Matching Algorithms
abstract
In this paper we compare the quality of synthesized views produced by four DIBR methods when fed by depth maps estimated by five state-of-the-art stereo matching algorithms. Also, we compute the correlation between four popular metrics for ranking stereo matching algorithms and two metrics commonly used to evaluate synthesized views (PSNR and SSIM) plus one specific for DIBR. Among our findings, we highlight that (i) PSNR and SSIM have a weak correlation with common stereo matching metrics, (ii) using ground-truth depth does not lead necessarily to the best DIBR result; and (iii) estimated depth maps present artifacts that affect differently DIBR methods.
Adriano Q. de Oliveira, Thiago L. T. da Silveira, Marcelo Walter, Cláudio R. Jung
ICASSP4
2019 Dense 3D Scene Reconstruction from Multiple Spherical Images for 3-DoF+ VR Applications
abstract
We propose a novel method for estimating the 3D geometry of indoor scenes based on multiple spherical images. Our technique produces a dense depth map registered to a reference view so that depth-image-based-rendering (DIBR) techniques can be explored for providing three-degrees-of-freedom plus immersive experiences to virtual reality users. The core of our method is to explore large displacement optical flow algorithms to obtain point correspondences, and use cross-checking and geometric constraints to detect and remove bad matches. We show that selecting a subset of the best dense matches leads to better pose estimates than traditional approaches based on sparse feature matching, and explore a weighting scheme to obtain the depth maps. Finally, we adapt a fast image-guided filter to the spherical domain for enforcing local spatial consistency, improving the 3D estimates. Experimental results indicate that our method quantitatively outperforms competitive approaches on computer-generated images and synthetic data under noisy correspondences and camera poses. Also, we show that the estimated depth maps obtained from only a few real spherical captures of the scene are capable of producing coherent synthesized binocular stereoscopic views by using traditional DIBR methods.
Thiago L. T. da Silveira, Cláudio R. Jung
VR2
2018 License Plate Detection and Recognition in Unconstrained Scenarios
Sergio Montazzolli Silva, Cláudio R. Jung
ECCV (12)2
2018 Temporal Consistency for Still Image Based Defocus Blur Estimation Methods
abstract
Many Defocus blur estimation methods have been proposed in recent years but, when applied to video sequences in a frame-by-frame manner, they typically exhibit temporal inconsistencies or flickering. This paper presents a temporal coherence scheme that can be coupled to any existing defocus blur estimation for still images, aiming to produce spatiotemporally coherent defocus blur map videos. The proposed method is based on the design of a Kalman Filter which is applied on a patch level. Experimental results show that the proposed method can smooth out undesirable temporal fluctuations whilst still being able to preserve the abrupt local appearance changes due to motion, occlusions or dis-occlusions.
Ali Karaali, Cláudio R. Jung, François Pitié
ICIP2
2018 Indoor Depth Estimation from Single Spherical Images
abstract
In this paper we propose a framework for inferring depth from a single spherical image, which can be coupled to any generic planar image monocular depth estimation algorithm. It consists of first inferring depth from overlapping planar patches extracted from the spherical image, and then using a regularized minimization scheme to stitch the patches back to the sphere. We test three state-of-the-art convolutional neural network (CNN)-based methodologies as baseline methods, and for all of them the proposed approach presented better results than applying the CNN directly to the equirectangular projection and to disjoint sections of the sphere according to the scale-invariant mean squared error (SIMSE) metric.
Thiago L. T. da Silveira, Lorenzo P. Dal'Aqua, Cláudio R. Jung
ICIP3
2018 Facial expression recognition using temporal POEM features
Edwin Alberto Silva Cruz, Cláudio R. Jung, Carlos Humberto Esparza Franco
Pattern Recognit. Lett.2
2018 An Artifact-Type Aware DIBR Method for View Synthesis
abstract
Depth-image-based rendering is a popular way to produce content for three-dimensional television and free viewpoint video, allowing the synthesis of numerous viewpoints using a single reference view and its depth map. Due to the synthesis process and nature of the input, artifacts and holes appear, and solving these problems becomes a challenge. In this letter, we propose solutions to remove those artifacts and apply different filling strategies depending on the nature of each hole. Cracks are identified and filled using very local neighborhood information. Regions classified as ghosts are projected to their correct place. The remaining holes are classified as disocclusions or out-of-field areas, and filled with an appropriate adaptation of a popular inpainting method. In both adaptations, patch matching explores the spatial locality concept, using dynamically adaptive patch sizes from the reference image. For disocclusions we propose a filling order using depth and background terms, and a searching process that considers only background patches. We show that our method outperforms several view synthesis methods in the quantitative evaluation, besides presenting consistent visual results for both large baselines and severely occluded scenes.
Adriano Q. de Oliveira, Marcelo Walter, Cláudio R. Jung
IEEE Signal Process. Lett.3
2018 Edge-Based Defocus Blur Estimation With Adaptive Scale Selection
abstract
Objects that do not lie at the focal distance of a digital camera generate defocused regions in the captured image. This paper presents a new edge-based method for spatially varying defocus blur estimation using a single image based on reblurred gradient magnitudes. The proposed approach initially computes a scale-consistent edge map of the input image and selects a local reblurring scale aiming to cope with noise, edge mis-localization, and interfering edges. An initial blur estimate is computed at the detected scale-consistent edge points and a novel connected edge filter is proposed to smooth the sparse blur map based on pixel connectivity within detected edge contours. Finally, a fast guided filter is used to propagate the sparse blur map through the whole image. Experimental results show that the proposed approach presents a very good compromise between estimation error and running time when compared with the state-of-the-art methods. We also explore our blur estimation method in the context of image deblurring, and show that metrics typically used to evaluate blur estimation may not correlate as expected with the visual quality of the deblurred image.
Ali Karaali, Cláudio R. Jung
IEEE Trans. Image Process.2
2017 Predicting Future Crowd Motion Including Event Treatment
Cliceres Mack Dal Bianco, Soraia Raupp Musse, Adriana Braun, Rodrigo Poli Caetani, Cláudio R. Jung, Norman I. Badler
IVA5
2017 Disparity map estimation and view synthesis using temporally adaptive triangular meshes
Guilherme P. Fickel, Cláudio R. Jung
Comput. Graph.2
2017 Detection of Global and Local Motion Changes in Human Crowds
abstract
Crowds arise in a variety of situations, such as public concerts and sporting matches. In typical conditions, the crowd moves in an orderly manner, but panic situations may lead to catastrophic results. We propose a computer vision method to identify motion pattern changes in human crowds that can be related to an unusual event. The proposed approach can identify global changes, by evaluating 2D motion histograms in time, and also local effects, by identifying clusters that present similar spatial locations and velocity vectors. The method is tested both on publicly available data sets involving crowded scenarios and on synthetic data produced by a crowd simulation algorithm, which allows the creation of controlled environments with known motion patterns that are particularly suitable for multicamera scenarios.
Igor Rodrigues de Almeida, Vinícius Jurinic Cassol, Norman I. Badler, Soraia Raupp Musse, Cláudio R. Jung
IEEE Trans. Circuits Syst. Video Technol.5
2017 Camera Self-Calibration Based on Nonlinear Optimization and Applications in Surveillance Systems
abstract
This paper presents a new approach for self-calibration of static cameras in the context of surveillance applications. Initially, a pedestrian detector is applied and the responses are validated using background removal. Then, foreground-related pixels within the detection results are used to estimate the feet-head line segments of each person (called poles), which are used to find a linear estimate for the camera matrix. Finally, a nonlinear cost function is used to refine the initial estimate, aiming to mostly improve the orientation of the reprojected poles. We also present different applications of self-calibration in tasks related to video surveillance itself, such as improvements to pedestrian detection and tracking algorithms, and augmented reality applications, such as the insertion of virtual cameras to aid the placement of real cameras in the scene.
Gustavo Führ, Cláudio R. Jung
IEEE Trans. Circuits Syst. Video Technol.2
2016 Image retargeting based on spatially varying defocus blur map
abstract
This paper presents a new image retargeting method that explores blur information. Given the input image, we compute the blur map and estimate in-focus regions. For retargeting, we first try to crop image boundaries as much as possible (preserving in-focus regions). If cropping is not enough, we use seam carving exploring a novel blur-aware energy function that concentrates the seams in blurred regions of the image. Experimental results show that the proposed blur-aware retargeting scheme works better at preserving in-focus objects than other competitive retargeting algorithms.
Ali Karaali, Cláudio R. Jung
ICIP2
2016 Fast-Forwarding Crowd Simulations
Cliceres Mack Dal Bianco, Adriana Braun, Soraia Raupp Musse, Cláudio R. Jung, Norman I. Badler
IVA4
2016 Evaluation of Histogram of Oriented Gradients Soft Errors Criticality for Automotive Applications
abstract
Pedestrian detection reliability is a key problem for autonomous or aided driving, and methods that use Histogram of Oriented Gradients (HOG) are very popular. Embedded Graphics Processing Units (GPUs) are exploited to run HOG in a very efficient manner. Unfortunately, GPUs architecture has been shown to be particularly vulnerable to radiation-induced failures. This article presents an experimental evaluation and analytical study of HOG reliability. We aim at quantifying and qualifying the radiation-induced errors on pedestrian detection applications executed in embedded GPUs. We analyze experimental results obtained executing HOG on embedded GPUs from two different vendors, exposed for about 100 hours to a controlled neutron beam at Los Alamos National Laboratory. We consider the number and position of detected objects as well as precision and recall to discriminate critical erroneous computations. The reported analysis shows that, while being intrinsically resilient (65% to 85% of output errors only slightly impact detection), HOG experienced some particularly critical errors that could result in undetected pedestrians or unnecessary vehicle stops. Additionally, we perform a fault-injection campaign to identify HOG critical procedures. We observe that Resize and Normalize are the most sensitive and critical phases, as about 20% of injections generate an output error that significantly impacts HOG detection. With our insights, we are able to find those limited portions of HOG that, if hardened, are more likely to increase reliability without introducing unnecessary overhead.
Fernando Santos 0001, Lucas Weigel, Cláudio R. Jung, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech
ACM Trans. Archit. Code Optim.3
2016 Audiovisual Tool for Solfège Assessment
abstract
Solfège is a general technique used in the music learning process that involves the vocal performance of melodies, regarding the time and duration of musical sounds as specified in the music score, properly associated with the meter-mimicking performed by hand movement. This article presents an audiovisual approach for automatic assessment of this relevant musical study practice. The proposed system combines the gesture of meter-mimicking (video information) with the melodic transcription (audio information), where hand movement works as a metronome, controlling the time flow (tempo) of the musical piece. Thus, meter-mimicking is used to align the music score (ground truth) with the sung melody, allowing assessment even in time-dynamic scenarios. Audio analysis is applied to achieve the melodic transcription of the sung notes and the solfège performances are evaluated by a set of Bayesian classifiers that were generated from real evaluations done by experts listeners.
Rodrigo Schramm, Helena de Souza Nunes, Cláudio R. Jung
ACM Trans. Multim. Comput. Commun. Appl.3
2015 Selective hole-filling for depth-image based rendering
abstract
One of the biggest challenges in view interpolation is to fill the regions without projective information in the synthesized view. In this paper, we present a new approach that identifies and corrects different types of missing information. In the first stage, we propose a fast solution to tackle the problems of cracks and ghost, common artifacts in the view interpolation process. Then, we complete larger holes by exploring the disparity map as an additional cue to select the best patch in a patch-based inpainting procedure. Our experimental results indicate that we were able to outperform current state of the art hole filling techniques for view interpolation.
Adriano Q. de Oliveira, Guilherme P. Fickel, Marcelo Walter, Cláudio R. Jung
ICASSP4
2015 Audiovisual voice activity detection using off-the-shelf cameras
abstract
This paper presents a new audiovisual voice activity detection (VAD) method for off-the-shelf cameras presenting a color sensor and two microphones. The motion of particles in the mouth region of each face detected by the camera is used as video cue, while the Generalized Cross Correlation with the PHase Transform (GCC-PHAT) is used as audio cue. We then estimate the distribution of the audiovisual cues and perform the final VAD result for each detected face using a Hidden Markov Model (HMM). Experimental results indicated that our method achieves an average 87% accuracy for a set of test videos.
Sergio Montazzolli Silva, Cláudio R. Jung, Dan Gelb
ICIP2
2015 Low-cost license plate detection using a calibrated camera
abstract
This paper presents a new approach for automatic license plate detection using an embedded camera inside a moving vehicle, implemented in a low-cost prototype. The key idea of the proposed method is to initially identify shadowed regions under the rear of visible vehicles, and then explore information from a calibrated camera to reduce the search space for license plates. Within the search space, any “baseline” detector can be used, and the results are validated exploring the expected size of a license plate in world coordinates and the known camera parameters. Our experimental results show that we can increase the speed of a “baseline” detector by 93%, and at the same time improving the f-score (increase in precision without significant loss in recall).
Henrique Weber, Cláudio R. Jung
ICIP2
2015 Hand and object segmentation from RGB-D images for interaction with planar surfaces
abstract
We introduce a new approach for hand and object segmentation using RGB-D cameras suitable for gesture-based Human-Computer Interfaces (HCIs) that involve an interaction plane. The technique consists of detecting the interaction plane using a temporally coherent version of RANSAC, followed by segmenting off-plane objects using a markers-based watershed transform with an energy function that combines depth and chromaticity gradients. Experimental results show that our approach can segment the hand even when it is very close to the interaction plane, unlike traditional approaches based on distance thresholds.
Henrique Weber, Cláudio R. Jung, Dan Gelb
ICIP2
2015 Automatic Detection and Classification of Road Lane Markings Using Onboard Vehicular Cameras
abstract
This paper presents a new approach for road lane classification using an onboard camera. Initially, lane boundaries are detected using a linear-parabolic lane model, and an automatic on-the-fly camera calibration procedure is applied. Then, an adaptive smoothing scheme is applied to reduce noise while keeping close edges separated, and pairs of local maxima-minima of the gradient are used as cues to identify lane markings. Finally, a Bayesian classifier based on mixtures of Gaussians is applied to classify the lane markings present at each frame of a video sequence as dashed, solid, dashed solid, solid dashed, or double solid. Experimental results indicate an overall accuracy of over 96% using a variety of video sequences acquired with different devices and resolutions.
Mauricio Braga de Paula, Cláudio R. Jung
IEEE Trans. Intell. Transp. Syst.2
2015 Multimodal Multi-Channel On-Line Speaker Diarization Using Sensor Fusion Through SVM
abstract
Speaker diarization (SD) is the process of assigning speech segments of an audio stream to its corresponding speakers, thus comprising the problem of voice activity detection (VAD), speaker labeling/identification, and often sound source localization (SSL). Most research activities in the past aimed towards applications as broadcast news, meetings, conversational telephony, and automatic multimodal data annotation, where SD may be performed off-line. However, a recent research focus is human-computer interaction (HCI) systems where SD must be performed on-line, and in real-time, as in modern gaming devices and interaction with large displays. Often, such applications further suffer from noise, reverberations, and overlapping speech, making them increasingly challenging. In such situations, multimodal/multisensory approaches can provide more accurate results than unimodal ones, given a data stream may compensate for occasional instabilities of other modalities. Accordingly, this paper presents an on-line multimodal SD algorithm designed to work in a realistic environment with multiple, overlapping speakers. Our work employs a microphone array, a color camera, and a depth sensor as input streams, from which speech-related features are extracted to be later merged through a support vector machine approach consisting of VAD and SSL modules. Speaker identification is incorporated through a hybrid technique of face positioning history and face recognition. Our final SD approach experimentally achieves an average diarization error rate of 11.48% in scenarios with up to three simultaneous speakers, and is able to run 3.2 × real-time.
Vicente P. Minotto, Cláudio R. Jung, Bowon Lee
IEEE Trans. Multim.2
2015 Dynamic Time Warping for Music Conducting Gestures Evaluation
abstract
Musical performance by an ensemble of performers often requires a conductor. This paper presents a tool to aid the study of basic conducting gestures, also known as meter- mimicking gestures, performed by beginners. It is based on the automatic detection of musical metrics and their subdivisions by analysis of hand gestures. Musical metrics are represented by visual conducting patterns performed by hands, which are tracked using an RGB-D camera. These patterns are recognized and evaluated using a probabilistic framework based on dynamic time warping (DTW). There are two main contributions in this work. Firstly, a new metric is proposed for the DTW, allowing better alignment between two gesture movements without the use of explicit maxima local points. Secondly, the time precision of the conducting gesture is extracted directly from the warping path and its accuracy is evaluated by a confidence measure. Experimental results indicate that the classification scheme represents an improvement over other existing related approaches.
Rodrigo Schramm, Cláudio R. Jung, Eduardo Reck Miranda
IEEE Trans. Multim.2
2014 Temporally coherent stereo matching using kinematic constraints
abstract
This paper explores a simple yet effective way to generate temporally coherent disparity maps from binocular video sequences based on kinematic constraints. Given the disparity map at a certain frame, the proposed approach computes the set of possible disparity values for each pixel in the subsequent frame, assuming a maximum displacement constraint (in world coordinates) allowed for each object. These disparity sets are then used to guide the stereo matching procedure in the subsequent frame, generating a temporally coherent disparity map. Experimental results indicate that the proposed approach produces temporally coherent disparity maps comparable to or better than competitive methods.
Rodrigo Schramm, Cláudio R. Jung
ICASSP2
2014 Multiview image and video interpolation using weighted vector median filters
abstract
In Depth Image-Based Rendering (DIBR), interpolated views generated using one or two cameras usually present artifacts and holes due to occlusions and/or inconsistencies in the input disparity maps. In this paper we propose a multiple (3 or more) camera view interpolation technique that is able to combine redundant projections in a single interpolated view by using Weighted Vector Median Filters (WVMFs). By expressing the weights of the WVMF using both the distance from each reference view to the synthetic view and a measure of consonance of each projection to the others, we achieve a high quality view interpolation without holes and common visual artifacts, such as cracks and ghost effects. Additionally, we present an extension to multiview video sequences by imposing temporal coherence in the estimated disparity maps.
Guilherme P. Fickel, Cláudio R. Jung, Bowon Lee
ICIP2
2014 Head-shoulder human contour estimation in still images
abstract
In this paper we propose a head-shoulder contour estimation model for human figures in still images, captured in a frontal pose. The contour estimation is guided by a learned head-shoulder shape model, initialized automatically by a face detector. A graph is generated around the detected face with an omega-like shape, and the estimated head-shoulder contour is a path in the graph with maximal cost. A dataset with labeled data is used to create the head-shoulder shape model and to quantitatively analyze the results. The proposed model is scaled according to the detected face size to be scale invariant. Experimental results indicate that the proposed technique works well in non trivial images, effectively estimating the contour of the head-shoulder even under partial occlusions.
Júlio C. S. Jacques Júnior, Cláudio R. Jung, Soraia Raupp Musse
ICIP2
2014 Adaptive scale selection for multiresolution defocus blur estimation
abstract
This paper presents a new method for defocus blur estimation using a single image. The proposed method exploits the ratio of gradient magnitude images computed at multiple scales, using the scale-space theory to estimate the number of reliable scales. Experimental results on synthetic and real images show that the proposed method is robust to noise, edge mis-localization and neighboring edge interference. We also show a new application of blur estimation algorithms to perform image re-targeting algorithms, leading to in-focus object preservation.
Ali Karaali, Cláudio R. Jung
ICIP2
2014 Automatic on-the-fly extrinsic camera calibration of onboard vehicular cameras
Mauricio Braga de Paula, Cláudio R. Jung, Luiz Gonzaga 0001
Expert Syst. Appl.2
2014 Combining SRP-PHAT and two Kinects for 3D Sound Source Localization
Lucas Adams Seewald, Luiz Gonzaga 0001, Maurício Roberto Veronez, Vicente P. Minotto, Cláudio R. Jung
Expert Syst. Appl.5
2014 Combining patch matching and detection for robust pedestrian tracking in monocular calibrated cameras
Gustavo Führ, Cláudio R. Jung
Pattern Recognit. Lett.2
2014 Simultaneous-Speaker Voice Activity Detection and Localization Using Mid-Fusion of SVM and HMMs
abstract
Humans can extract speech signals that they need to understand from a mixture of background noise, interfering sound sources, and reverberation for effective communication. Voice Activity Detection (VAD) and Sound Source Localization (SSL) are the key signal processing components that humans perform by processing sound signals received at both ears, sometimes with the help of visual cues by locating and observing the lip movements of the speaker. Both VAD and SSL serve as the crucial design elements for building applications involving human speech. For example, systems with microphone arrays can benefit from these for robust speech capture in video conferencing applications, or for speaker identification and speech recognition in Human Computer Interfaces (HCIs). The design and implementation of robust VAD and SSL algorithms in practical acoustic environments are still challenging problems, particularly when multiple simultaneous speakers exist in the same audiovisual scene. In this work we propose a multimodal approach that uses Support Vector Machines (SVMs) and Hidden Markov Models (HMMs) for assessing the video and audio modalities through an RGB camera and a microphone array. By analyzing the individual speakers' spatio-temporal activities and mouth movements, we propose a mid-fusion approach to perform both VAD and SSL for multiple active and inactive speakers. We tested the proposed algorithm in scenarios with up to three simultaneous speakers, showing an average VAD accuracy of 95.06% with an average error of 10.9 cm when estimating the three-dimensional locations of the speakers.
Vicente P. Minotto, Cláudio R. Jung, Bowon Lee
IEEE Trans. Multim.2
2013 An evaluation of stereo matching methods for view interpolation
abstract
Stereo matching has a long history in image processing and computer vision. In fact, there are inumerous approaches reported in the literature, and quantitative evaluation is usually performed by comparing the obtained disparity maps with ground truth data (using the MSE, for instance). One important application of stereo matching is view interpolation, where it is desired to produce a new synthetic view from (at least) a pair of images and the corresponding disparity maps. In view interpolation, evaluation is mostly qualitative (visual quality of the synthesized image), and quantitative approaches compute objective similarity metrics between the synthesized image and the actual image at the same position (e.g. PSNR). The main goal of this paper is to evaluate the impact of several different stereo matching algorithms in a view interpolation context, relating the quality of the disparity maps with the quality of the corresponding synthesized views using standardized datasets. In this paper, experiments using the MPEG reference software for view interpolation and more than twenty datasets are presented and discussed. Our results indicate that the use of the common percentage of bad pixels as a metric for stereo matching methods does not translate well to the quality of view interpolation.
Gustavo Führ, Guilherme P. Fickel, Lorenzo P. Dal'Aqua, Cláudio R. Jung, Thomas Malzbender, Ramin Samadani
ICIP4
2013 Self-occlusion and 3D pose estimation in still images
abstract
In this paper we propose a self-occlusion and 3D pose estimation model for human figures in still images based on a user-provided 2D skeleton. An initial segmentation model is used to capture labeled human body parts in a 2D image. Then, occluded body parts are detected when different body parts overlap, and are disambiguated by analyzing the energy of the corresponding contours around the intersection points. The estimated occlusion results feed the 3D pose estimation algorithm, which reconstructs a set of plausible 3D postures. Experimental results indicate that the proposed technique works well in non trivial images, effectively estimating the occluded body parts and reducing the number of possible 3D postures.
Júlio C. S. Jacques Júnior, Leandro Dihl, Cláudio R. Jung, Soraia Raupp Musse
ICIP3
2013 Stereo Matching and View Interpolation Based on Image Domain Triangulation
abstract
This paper presents a new approach for stereo matching and view interpolation problems based on triangular tessellations suitable for a linear array of rectified cameras. The domain of the reference image is initially partitioned into triangular regions using edge and scale information, aiming to place vertices along image edges and increase the number of triangles in textured regions. A region-based matching algorithm is then used to find an initial disparity for each triangle, and a refinement stage is applied to change the disparity at the vertices of the triangles, generating a piecewise linear disparity map. A simple post-processing procedure is applied to connect triangles with similar disparities generating a full 3D mesh related to each camera (view), which are used to generate new synthesized views along the linear camera array. With the proposed framework, view interpolation reduces to the trivial task of rendering polygonal meshes, which can be done very fast, particularly when GPUs are employed. Furthermore, the generated views are hole-free, unlike most point-based view interpolation schemes that require some kind of post-processing procedures to fill holes.
Guilherme P. Fickel, Cláudio R. Jung, Thomas Malzbender, Ramin Samadani, W. Bruce Culbertson
IEEE Trans. Image Process.2
2012 Stereo matching based on image triangulation for view synthesis
abstract
In this paper we propose a new disparity map estimation algorithm from multiple rectified images. The reference image is initially segmented into triangular regions, and each triangle is assigned to an initial disparity, leading to a piece-wise constant disparity map. Then, a refinement step is applied, where the disparity of the triangle vertices is adjusted aiming to impose spatial consistency to the disparity map, smoothing the map within the objects and keeping the discontinuities between them. The final result is a piece-wise linear depth map on a triangular domain that can be explored for view synthesis.
Guilherme P. Fickel, Cláudio R. Jung, Ramin Samadani, Thomas Malzbender
ICIP2
2012 Skeleton-based human segmentation in still images
abstract
In this paper we propose a skeleton-based model for human segmentation in static images. Our approach explores edge information, orientation coherence and anthropometric-estimated parameters to generate a graph, and the desired contour is a path with maximal cost. Experimental results show that the proposed technique works well in non trivial images.
Júlio C. S. Jacques Júnior, Cláudio R. Jung, Soraia Raupp Musse
ICIP2
2012 Simulating crowds based on a space colonization algorithm
Alessandro de Lima Bicho, Rafael Araújo Rodrigues, Soraia Raupp Musse, Cláudio R. Jung, Marcelo Paravisi, Léo Pini Magalhães
Comput. Graph.4
2012 Towards a quantitative approach for comparing crowds
abstract
ABSTRACT In this paper, we propose a new model to quantitatively compare global flow characteristics of two crowds. The proposed approach explores a 4‐D histogram that contains information on the local velocity (speed and orientation) of each spatial position, and the comparison is made using histogram distances. The 4‐D histogram also allows the comparison of specific characteristics, such as distribution of orientations only, speed only, relative spatial occupancy only, and combinations of such features. Experimental results indicate that the proposed quantitative metric correlates with visual inspection. Copyright © 2012 John Wiley & Sons, Ltd.
Soraia Raupp Musse, Vinícius Jurinic Cassol, Cláudio R. Jung
Comput. Animat. Virtual Worlds3
2012 Voice activity detection and speaker localization using audiovisual cues
Dante A. Blauth, Vicente P. Minotto, Cláudio R. Jung, Bowon Lee, Ton Kalker
Pattern Recognit. Lett.3
2011 Color-based lips extraction applied to voice activity detection
abstract
The lip motion detection stands out as relevant visual feature for detecting the active speaker and speech recognition. In this paper, a new approach for lips and visual voice activity detection is proposed. First, the algorithm performs skin segmentation to reduce the search area for lip extraction, and the most likely lip and non-lip regions are detected using a Bayesian approach within the delimited area. Then, the final lip segmentation is obtained by thresholding the calculated probability regions and applying simple morphological operators. Finally, the temporal motion of the lips is explored using Hidden Markov Models (HMMs) to detect the likely occurrence of active speech within a temporal window.
Carlos B. O. Lopes, Andre L. Gonçalves, Jacob Scharcanski, Cláudio R. Jung
ICIP4
2011 Bilayer video segmentation for videoconferencing applications
abstract
This paper presents a new bilayer video segmentation algorithm focusing on videoconferencing application. A face tracking algorithm is used to guide a generic Ω-shaped template of the head and shoulders. A region of interest (ROI) is created around the generic template, and an energy function based on edge, color and motion cues is used to define the boundary between the person and the background. Our experimental results indicate that the silhouettes can be effectively extracted in common videoconferencing scenarios.
Alessandro Parolin, Guilherme P. Fickel, Cláudio R. Jung, Thomas Malzbender, Ramin Samadani
ICME3
2011 Image processing, computer vision and pattern recognition in Latin America
Cláudio R. Jung
Pattern Recognit. Lett.1
2010 Human upper body identification from images
abstract
Estimating human pose in static images is challenging due to the high dimensional state space, presence of image clutter and ambiguities of image observations. In this paper we propose a method to automatically segment human subjects in images, based on dominant colors, and given the face captured by a face detector. The posture is estimated using a 2D model combined with anthropometric data. Experimental results showed that the proposed technique performs well in non trivial images.
Júlio C. S. Jacques Júnior, Leandro Dihl, Cláudio R. Jung, Marcelo Thielo, Renato Keshet, Soraia Raupp Musse
ICIP3
2009 Object tracking using multiple fragments
abstract
This paper presents a low-cost tracking algorithm based on multiple multiple fragments, increasing robustness with respect to partial occlusions. Given the initial template representing the desired target, each pixel is classified into a different cluster based on a Mixture of Gaussians (MOG) model, and a set of disjoint fragments is created. The mean vector and covariance matrix of each fragment are computed, and the Mahalanobis distance is used to decide which pixels of the adjacent frame within a neighborhood are associated with each fragment. The template is then placed at the position that maximizes a similarity measure based on the number of matched points.
Cláudio R. Jung, Amir Said
ICIP1
2009 Feature-Based Face Tracking for Videoconferencing Applications
abstract
This paper proposes a new approach for face tracking based on the individual tracking of KLT features. The face is initially detected using a face detection scheme, and KLT features are distributed along the face. Each feature is tracked individually, and the displacement of the center of the face is obtained using a Weighted Vector Median Filter (WVMF) of the individual displacements. The scale change is then computed based on the position of each feature w.r.t. the center of the face. The experimental results indicate that the proposed approach is fast and robust in the presence of partial occlusions.
José Bins, Cláudio R. Jung, Leandro Dihl, Amir Said
ISM2
2009 Video Based VAD Using Adaptive Color Information
abstract
This paper presents a new approach for voice activity detection (VAD) in videoconferencing applications based only on video information. In the proposed approach, a face detection scheme is applied to locate the face of the participant, and anthropometric measures are used to identify one patch below each eye, as well as a neighborhood around the mouth. The patches below the eyes are used to train a skin color model suited for that specific person, which is then used to identify non-skin pixels in the mouth neighborhood. The number of non-skin pixels is taken as an estimate of the mouth openness, and its evolution across time is explored for VAD.
Dario Scott, Cláudio R. Jung, José Bins, Amir Said, Ton Kalker
ISM2
2009 Tree Paths: A New Model for Steering Behaviors
Rafael Araújo Rodrigues, Alessandro de Lima Bicho, Marcelo Paravisi, Cláudio R. Jung, Léo Pini Magalhães, Soraia Raupp Musse
IVA4
2009 Efficient Background Subtraction and Shadow Removal for Monochromatic Video Sequences
abstract
This letter presents a new method for background subtraction and shadow removal for grayscale video sequences. The background image is modeled using robust statistical descriptors, and a noise estimate is obtained. Foreground pixels are extracted, and a statistical approach combined with geometrical constraints are adopted to detect and remove shadows.
Cláudio R. Jung
IEEE Trans. Multim.1
2008 Event Detection Using Trajectory Clustering and 4-D Histograms
abstract
In this paper, we propose a framework for event detection based on trajectory clustering and 4-D histograms. In the training period, captured trajectories are grouped into coherent clusters according to global motion flows. Within each cluster, the position and instantaneous velocity of each tracked object are used to build a 4-D motion histogram for the cluster. In the test period, each new trajectory is compared against the 4-D histograms of all clusters, so that its coherence with previously tracked objects can be evaluated. Experimental results showed that these criteria can be effectively used to measure the coherence of test trajectories with those in the training stage, allowing a range of events to be detected in surveillance and traffic applications.
Cláudio R. Jung, Luciano Hennemann, Soraia Raupp Musse
IEEE Trans. Circuits Syst. Video Technol.1
2007 Combining wavelets and watersheds for robust multiscale image segmentation
Cláudio R. Jung
Image Vis. Comput.1
2007 Using computer vision to simulate the motion of virtual agents
abstract
Abstract In this paper, we propose a new model to simulate the movement of virtual humans based on trajectories captured automatically from filmed video sequences. These trajectories are grouped into similar classes using an unsupervised clustering algorithm, and an extrapolated velocity field is generated for each class. A physically‐based simulator is then used to animate virtual humans, aiming to reproduce the trajectories fed to the algorithm and at the same time avoiding collisions with other agents. The proposed approach provides an automatic way to reproduce the motion of real people in a virtual environment, allowing the user to change the number of simulated agents while keeping the same goals observed in the filmed video. Copyright © 2007 John Wiley & Sons, Ltd.
Soraia Raupp Musse, Cláudio R. Jung, Júlio C. S. Jacques Júnior, Adriana Braun
Comput. Animat. Virtual Worlds2
2007 Understanding people motion in video sequences using Voronoi diagrams
Júlio C. S. Jacques Júnior, Adriana Braun, John Soldera, Soraia Raupp Musse, Cláudio R. Jung
Pattern Anal. Appl.5
2007 Unsupervised multiscale segmentation of color images
Cláudio R. Jung
Pattern Recognit. Lett.1
2007 Block-based image inpainting in the wavelet domain
Ubiratã A. Ignácio, Cláudio R. Jung
Vis. Comput.2
2006 A Background Subtraction Model Adapted to Illumination Changes
abstract
This paper presents a new adaptive background model for grayscale video sequences, that includes shadows and highlight detection. In the training period, statistics are computed for each image pixel to obtain the initial background model and an estimate of the image global noise, even in the presence of several moving objects. Each new frame is then compared to this background model, and spatio-temporal features are used to obtain foreground pixels. Local statistics are then used to detect shadows and highlights, and pixels that are detected as either shadow or highlight for a certain number of frames are adapted to become part of the background. Experimental results indicate that the proposed algorithm can effectively detect shadows and highlights, adapting the background with respect to illumination changes.
Júlio C. S. Jacques Júnior, Cláudio R. Jung, Soraia Raupp Musse
ICIP2
2006 A Randomized Approach for Patch-based Texture Synthesis using Wavelets
abstract
Abstract We present a wavelet‐based approach for selecting patches in patch‐based texture synthesis. We randomly select the first block that satisfies a minimum error criterion, computed from the wavelet coefficients (using 1D or 2D wavelets) for the overlapping region. We show that our wavelet‐based approach improves texture synthesis for samples where previous work fails, mainly textures with prominent aligned features. Also, it generates similar quality textures when compared against texture synthesis using feature maps with the advantage that our proposed method uses implicit edge information (since it is embedded in the wavelet coefficients) whereas feature maps rely explicitly on edge features. In previous work, the best patches are selected among all possible using a L2 norm on the RGB or grayscale pixel values of boundary zones. The L2 metric provides the raw pixel‐to‐pixel difference, disregarding relevant image structures — such as edges — that are relevant in the human visual system and therefore on synthesis of new textures.
Leandro Tonietto, Marcelo Walter, Cláudio R. Jung
Comput. Graph. Forum3
2005 Lane following and lane departure using a linear-parabolic model
Cláudio R. Jung, Christian Roberto Kelber
Image Vis. Comput.1
2005 Robust watershed segmentation using wavelets
Cláudio R. Jung, Jacob Scharcanski
Image Vis. Comput.1
2003 Adaptive image denoising and edge enhancement in scale-space using the wavelet transform
Cláudio R. Jung, Jacob Scharcanski
Pattern Recognit. Lett.1
2002 Adaptive image denoising using scale and space consistency
abstract
This paper proposes a new method for image denoising with edge preservation, based on image multiresolution decomposition by a redundant wavelet transform. In our approach, edges are implicitly located and preserved in the wavelet domain, whilst image noise is filtered out. At each resolution level, the image edges are estimated by gradient magnitudes (obtained from the wavelet coefficients), which are modeled probabilistically, and a shrinkage function is assembled based on the model obtained. Joint use of space and scale consistency is applied for better preservation of edges. The shrinkage functions are combined to preserve edges that appear simultaneously at several resolutions, and geometric constraints are applied to preserve edges that are not isolated. The proposed technique produces a filtered version of the original image, where homogeneous regions appear separated by well-defined edges. Possible applications include image presegmentation, and image denoising.
Jacob Scharcanski, Cláudio R. Jung, Robin T. Clarke
IEEE Trans. Image Process.2