EDBT 2026 Demo / reviewers in the wild / expert
Dinei A. F. Florêncio
dblp:31/926 · also Dinei Florêncio
· DBLP profile ↗
94ranked-venue papers
21as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 71 · 11 first-author · 3 since 2021Artificial intelligence and machine learning · 17 · 7 since 2021Security and privacy · 9 · 7 first-authorSystems, architecture and hardware · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Vision and language · 46% Language models and text generation · 17% Image recognition and object detection · 13% | |
| Computer graphics and multimedia
13 papers |
Audio and music processing · 36% Image and video coding · 31% Image and video processing · 27% | |
| Network and information security
2 papers |
Authentication and access control · 84% Usable security · 16% |
Topics — the 30 heaviest of 54, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 |
Computer vision › Vision and language › visual reasoning
visual chain-of-thought |
0.9 | 1 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 |
Computer vision › Vision and language
visual reasoning |
0.9 | 1 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 |
Computer vision › Vision and language
open-vocabulary models |
0.7 | 1 | 2023 | From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language Understanding · ACL (1) 2023 |
Computer vision › Image recognition and object detection › text recognition
optical character recognition |
0.7 | 1 | 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained Models · AAAI 2023 |
Computer vision › Image recognition and object detection
text recognition |
0.7 | 1 | 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained Models · AAAI 2023 |
Natural language and speech › Language models and text generation › text generation › neural text generation
transformer-based text generation |
0.7 | 1 | 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained Models · AAAI 2023 |
Image and video coding › shape coding
contour coding |
0.6 | 2 | 2018 | Joint Denoising/Compression of Image Contours via Shape Prior and Context Tree · IEEE Trans. Image Process. 2018 Context Tree-Based Image Contour Coding Using a Geometric Prior · IEEE Trans. Image Process. 2017 |
Natural language and speech › Information extraction and text analysis
document understanding |
0.5 | 1 | 2021 | LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding · ACL/IJCNLP (1) 2021 |
Computer vision › Vision and language › visual question answering
text-based visual question answering |
0.5 | 1 | 2021 | TAP: Text-Aware Pre-Training for Text-VQA and Text-Caption · CVPR 2021 |
Computer vision › Vision and language
vision-language pretraining |
0.5 | 1 | 2021 | TAP: Text-Aware Pre-Training for Text-VQA and Text-Caption · CVPR 2021 |
Machine learning › Deep learning architectures and training › convolutional neural network
convolutional neural network training |
0.4 | 1 | 2019 | RePr: Improved Training of Convolutional Filters · CVPR 2019 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2019 | RePr: Improved Training of Convolutional Filters · CVPR 2019 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.4 | 1 | 2019 | RePr: Improved Training of Convolutional Filters · CVPR 2019 |
Audio and music processing
spatial audio |
0.4 | 2 | 2016 | A Manifold Learning Approach for Personalizing HRTFs from Anthropometric Features · IEEE ACM Trans. Audio Speech Lang. Process. 2016 An Interactive 3-D Audio System With Loudspeakers · IEEE Trans. Multim. 2011 |
Image and video processing › image restoration
image denoising |
0.3 | 1 | 2018 | Joint Denoising/Compression of Image Contours via Shape Prior and Context Tree · IEEE Trans. Image Process. 2018 |
Image and video coding › entropy coding
arithmetic coding |
0.3 | 1 | 2017 | Context Tree-Based Image Contour Coding Using a Geometric Prior · IEEE Trans. Image Process. 2017 |
Audio and music processing
room acoustics |
0.3 | 2 | 2012 | Geometrically Constrained Room Modeling With Compact Microphone Arrays · IEEE Trans. Speech Audio Process. 2012 Using Reverberation to Improve Range and Elevation Discrimination for Small Array Sound Source Localization · IEEE Trans. Speech Audio Process. 2010 |
Image and video processing › image enhancement
bit-depth enhancement |
0.2 | 1 | 2016 | Image Bit-Depth Enhancement via Maximum A Posteriori Estimation of AC Signal · IEEE Trans. Image Process. 2016 |
Audio and music processing › spatial audio › head-related transfer function
head-related transfer function personalization |
0.2 | 1 | 2016 | A Manifold Learning Approach for Personalizing HRTFs from Anthropometric Features · IEEE ACM Trans. Audio Speech Lang. Process. 2016 |
Image and video processing
image enhancement |
0.2 | 1 | 2016 | Image Bit-Depth Enhancement via Maximum A Posteriori Estimation of AC Signal · IEEE Trans. Image Process. 2016 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.2 | 1 | 2023 | From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language Understanding · ACL (1) 2023 |
Audio and music processing
sound source localization |
0.2 | 2 | 2010 | Using Reverberation to Improve Range and Elevation Discrimination for Small Array Sound Source Localization · IEEE Trans. Speech Audio Process. 2010 Maximum Likelihood Sound Source Localization and Beamforming for Directional Microphone Arrays in Distributed Meetings · IEEE Trans. Multim. 2008 |
Computer vision › 3D vision
3d shape reconstruction |
0.2 | 1 | 2014 | Rate-Constrained 3D Surface Estimation From Noise-Corrupted Multiview Depth Videos · IEEE Trans. Image Process. 2014 |
Image and video coding › video compression › 3d video coding
depth map coding |
0.2 | 1 | 2014 | Rate-Constrained 3D Surface Estimation From Noise-Corrupted Multiview Depth Videos · IEEE Trans. Image Process. 2014 |
Image and video coding › video compression › 3d video coding
depth video coding |
0.2 | 1 | 2014 | Arbitrarily Shaped Motion Prediction for Depth Video Compression Using Arithmetic Edge Coding · IEEE Trans. Image Process. 2014 |
Image and video coding
edge encoding |
0.2 | 1 | 2014 | Arbitrarily Shaped Motion Prediction for Depth Video Compression Using Arithmetic Edge Coding · IEEE Trans. Image Process. 2014 |
Computer animation and physical simulation › motion modeling
motion prediction |
0.2 | 1 | 2014 | Arbitrarily Shaped Motion Prediction for Depth Video Compression Using Arithmetic Edge Coding · IEEE Trans. Image Process. 2014 |
Authentication and access control › password security
password management |
0.2 | 1 | 2014 | Password Portfolios and the Finite-Effort User: Sustainably Managing Large Numbers of Accounts · USENIX Security Symposium 2014 |
Methods — techniques the papers use, named apart from their topics
visual editing · 0.9code generation · 0.9transformer · 0.7pre-trained text transformer · 0.7pre-trained image transformer · 0.7hierarchical modeling · 0.7maximum a posteriori · 0.6dynamic programming · 0.6multimodal pretraining · 0.5masked language modeling · 0.5layout-aware transformers · 0.5contrastive matching · 0.5arithmetic coding · 0.5wavelet transform · 0.3total suffix tree · 0.3foreground and background modeling · 0.3burst error model · 0.3geometric priors · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image UnderstandingabstractStructured image understanding, such as interpreting tables and charts, requires strategically refocusing across various structures and texts within an image, forming a reasoning sequence to arrive at the final answer. However, current multimodal large language models (LLMs) lack this multihop selective attention capability. In this work, we introduce ReFocus, a simple yet effective framework that equips multimodal LLMs with the ability to generate ``visual thoughts'' by performing visual editing on the input image through code, shifting and refining their visual focuses. Specifically, ReFocus enables multimodal LLMs to generate Python codes to call tools and modify the input image, sequentially drawing boxes, highlighting sections, and masking out areas, thereby enhancing the visual reasoning process. We experiment upon a wide range of structured image understanding tasks involving tables and charts. ReFocus largely improves performance on all tasks over GPT-4o without visual editing, yielding an average gain of 11.0% on table tasks and 6.8% on chart tasks. We present an in-depth analysis of the effects of different visual edits, and reasons why ReFocus can improve the performance without introducing additional information. Further, we collect a 14k training set using ReFocus, and prove that such visual chain-of-thought with intermediate information offers a better supervision than standard VQA data, reaching a 8.0% average gain over the same model trained with QA pairs and 2.6% over CoT. Minqian Liu, Zhengyuan Yang, John Corring, Yijuan Lu, Dan Roth 0001, Dinei A. F. Florêncio, Cha Zhang |
ICML | 8 |
| 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsabstractText recognition is a long-standing research problem for document digitalization. Existing approaches are usually built based on CNN for image understanding and RNN for char-level text generation. In addition, another language model is usually needed to improve the overall accuracy as a post-processing step. In this paper, we propose an end-to-end text recognition approach with pre-trained image Transformer and text Transformer models, namely TrOCR, which leverages the Transformer architecture for both image understanding and wordpiece-level text generation. The TrOCR model is simple but effective, and can be pre-trained with large-scale synthetic data and fine-tuned with human-labeled datasets. Experiments show that the TrOCR model outperforms the current state-of-the-art models on the printed, handwritten and scene text recognition tasks. The TrOCR models and code are publicly available at https://aka.ms/trocr. Minghao Li 0004, Tengchao Lv, Jingye Chen, Lei Cui 0001, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Zhoujun Li 0001, Furu Wei |
AAAI | 6 |
| 2023 | From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language UnderstandingabstractCurrent state-of-the-art models for natural language understanding require a preprocessing step to convert raw text into discrete tokens.This process known as tokenization relies on a pre-built vocabulary of words or sub-word morphemes.This fixed vocabulary limits the model's robustness to spelling errors and its capacity to adapt to new domains.In this work, we introduce a novel open-vocabulary language model that adopts a hierarchical two-level approach: one at the word level and another at the sequence level.Concretely, we design an intraword module that uses a shallow Transformer architecture to learn word representations from their characters, and a deep inter-word Transformer module that contextualizes each word representation by attending to the entire word sequence.Our model thus directly operates on character sequences with explicit awareness of word boundaries, but without biased sub-word or word-level vocabulary.Experiments on various downstream tasks show that our method outperforms strong baselines.We also demonstrate that our hierarchical model is robust to textual corruption and domain shift. Li Sun 0010, Florian Luisier, Kayhan Batmanghelich, Dinei A. F. Florêncio, Cha Zhang |
ACL (1) | 4 |
| 2023 | Diffusion-Based Document Layout Generation
Yijuan Lu, John Corring, Dinei A. F. Florêncio, Cha Zhang |
ICDAR (1) | 4 |
| 2022 | Layer Separation via a Spatial-Attention GANabstractModern OCR systems are achieving excellent results on difficult, even physically degraded, documents. One persistent limitation is the inability to disentangle overlapping ink, from various sources like handwriting, stamps, watermarks, and printed ink. In this work, we develop an approach to layer separation of documents into distinct layers which contain dedicated ink for these respective components. Our work builds on recent progress in image decomposition and employs a spatial-attention bypass enabling us to overwrite and in-paint challenging overlap regions while still maintaining near perfect reconstruction for undisturbed regions of the documents. Compared to baseline approaches, our approach improves on both perceptual metrics on the layer reconstruction task and end-to-end character error rate (CER) measured on an OCR system. John Corring, Jinsol Lee, Florian Luisier, Dinei A. F. Florêncio |
ICPR | 4 |
| 2021 | LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingabstractYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, Lidong Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yang Xu 0049, Yiheng Xu, Tengchao Lv, Lei Cui 0001, Furu Wei, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Wanxiang Che, Min Zhang 0005, Lidong Zhou |
ACL/IJCNLP (1) | 8 |
| 2021 | TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionabstractIn this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answering and image caption generation, respectively. In contrast to conventional vision-language pretraining that fails to capture scene text and its relationship with the visual and text modalities, TAP explicitly incorporates scene text (generated from OCR engines) during pretraining. With three pre-training tasks, including masked language modeling (MLM), image-text (contrastive) matching (ITM), and relative (spatial) position prediction (RPP), pre-training with scene text effectively helps the model learn a better aligned representation among the three modalities: text word, visual object, and scene text. Due to this aligned representation learning, even pre-trained on the same downstream task dataset, TAP already boosts the absolute accuracy on the TextVQA dataset by +5:4%, compared with a non-TAP baseline. To further improve the performance, we build a large-scale scene text-related imagetext dataset based on the Conceptual Caption dataset, named OCR-CC, which contains 1:4 million images with scene text. Pre-trained on this OCR-CC dataset, our approach outperforms the state of the art by large margins on multiple tasks, i.e., +8:3% accuracy on TextVQA, +8:6% accuracy on ST-VQA, and +10:2 CIDEr score on TextCaps. Zhengyuan Yang, Yijuan Lu, Xi Yin 0006, Dinei A. F. Florêncio, Cha Zhang, Lei Zhang 0001, Jiebo Luo 0001 |
CVPR | 5 |
| 2019 | RePr: Improved Training of Convolutional FiltersabstractA well-trained Convolutional Neural Network can easily be pruned without significant loss of performance. This is because of unnecessary overlap in the features captured by the network's filters. Innovations in network architecture such as skip/dense connections and inception units have mitigated this problem to some extent, but these improvements come with increased computation and memory requirements at run-time. We attempt to address this problem from another angle - not by changing the network structure but by altering the training method. We show that by temporarily pruning and then restoring a subset of the model's filters, and repeating this process cyclically, overlap in the learned features is reduced, producing improved generalization. We show that the existing model-pruning criteria are not optimal for selecting filters to prune in this context, and introduce inter-filter orthogonality as the ranking criteria to determine under-expressive filters. Our method is applicable both to vanilla convolutional networks and more complex modern architectures, and improves the performance across a variety of tasks, especially when applied to smaller networks. Aaditya Prakash, James A. Storer, Dinei A. F. Florêncio, Cha Zhang |
CVPR | 3 |
| 2018 | Deep Learning Based Speech BeamformingabstractMulti-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the speech models tend to be oversimplified or the inference would otherwise be too complicated. On the other hand, deep learning based enhancement approaches are able to learn complicated speech distributions and perform efficient inference, but they are unable to deal with variable number of input channels. Also, deep learning approaches introduce a lot of errors, particularly in the presence of unseen noise types and settings. We have therefore proposed an enhancement framework called DEEPBEAM, which combines the two complementary classes of algorithms. DEEPBEAM introduces a beamforming filter to produce natural sounding speech, but the filter coefficients are determined with the help of a monaural speech enhancement neural network. Experiments on synthetic and real-world data show that DEEPBEAM is able to produce clean, dry and natural sounding speech, and is robust against unseen noise. Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Xuesong Yang, Dinei A. F. Florêncio, Mark Hasegawa-Johnson |
ICASSP | 5 |
| 2018 | Distance-Based Probability Model for Octree CodingabstractWe present a context-driven method to encode nodes of an octree, which is typically used to encode the point-cloud geometry. Instead of using one bit per node of the tree, the context allows for deriving probabilities for that node based on distances of the actual voxel to the voxels in a reference point cloud. Accurate probabilities of the node state allow for the use of an arithmetic coder to reduce the bit rate. Results point to potentially large reductions in rate if there is a good model from which to derive the context, i.e., one can get large reductions if the reference-cloud geometry is close enough to the one being encoded. Ricardo L. de Queiroz, Diogo C. Garcia, Philip A. Chou, Dinei A. F. Florêncio |
IEEE Signal Process. Lett. | 4 |
| 2018 | A Fusion Framework for Camouflaged Moving Foreground Detection in the Wavelet DomainabstractDetecting camouflaged moving foreground objects has been known to be difficult due to the similarity between the foreground objects and the background. Conventional methods cannot distinguish the foreground from background due to the small differences between them and thus suffer from underdetection of the camouflaged foreground objects. In this paper, we present a fusion framework to address this problem in the wavelet domain. We first show that the small differences in the image domain can be highlighted in certain wavelet bands. Then the likelihood of each wavelet coefficient being foreground is estimated by formulating foreground and background models for each wavelet band. The proposed framework effectively aggregates the likelihoods from different wavelet bands based on the characteristics of the wavelet transform. Experimental results demonstrated that the proposed method significantly outperformed existing methods in detecting camouflaged foreground objects. Specifically, the average F-measure for the proposed algorithm was 0.87, compared to 0.71 to 0.8 for the other stateof- the-art methods. Shuai Li 0005, Dinei A. F. Florêncio, Wanqing Li 0001, Yaqin Zhao, Chris Cook |
IEEE Trans. Image Process. | 2 |
| 2018 | Joint Denoising/Compression of Image Contours via Shape Prior and Context TreeabstractThe advent of depth sensing technologies means that the extraction of object contours in images-a common and important pre-processing step for later higher level computer vision tasks like object detection and human action recognition-has become easier. However, captured depth images contain acquisition noise and the detected contours suffer from errors as a result. In this paper, we propose to jointly denoise and compress detected contours in an image for bandwidth-constrained transmission to a client, who can then carry out aforementioned application-specific tasks using the decoded contours as input. First, we prove theoretically that in general a joint denoising/compression approach can outperform a separate two-stage approach that first denoises then encodes contours lossily. Adopting a joint approach, we propose a burst error model that models typical errors encountered in an observed string of directional edges. We then formulate a rate-constrained maximum a posteriori problem that trades off the posterior probability of an estimated string given with its code rate. We design a dynamic programming algorithm that solves the posed problem optimally, and propose a compact context representation called total suffix tree that can reduce complexity of the algorithm dramatically. To the best of our knowledge, we are the first in the literature to study the problem of joint denoising/compression of image contours and offer a computation-efficient optimization algorithm. Experimental results show that our joint denoising/compression scheme can reduce bitrate by up to 18% compared with a competing separate scheme at comparable visual quality. Amin Zheng, Gene Cheung, Dinei A. F. Florêncio |
IEEE Trans. Image Process. | 3 |
| 2017 | Foreground detection in camouflaged scenesabstractForeground detection has been widely studied for decades due to its importance in many practical applications. Most of the existing methods assume foreground and background show visually distinct characteristics and thus the foreground can be detected once a good background model is obtained. However, there are many situations where this is not the case. Of particular interest in video surveillance is the camouflage case. For example, an active attacker camouflages by intentionally wearing clothes that are visually similar to the background. In such cases, even given a decent background model, it is not trivial to detect foreground objects. This paper proposes a texture guided weighted voting (TGWV) method which can efficiently detect foreground objects in camouflaged scenes. The proposed method employs the stationary wavelet transform to decompose the image into frequency bands. We show that the small and hardly noticeable differences between foreground and background in the image domain can be effectively captured in certain wavelet frequency bands. To make the final foreground decision, a weighted voting scheme is developed based on intensity and texture of all the wavelet bands with weights carefully designed. Experimental results demonstrate that the proposed method achieves superior performance compared to the current state-of-the-art results. Shuai Li 0005, Dinei A. F. Florêncio, Yaqin Zhao, Chris Cook, Wanqing Li 0001 |
ICIP | 2 |
| 2017 | Progressive graph-signal sampling and encoding for static 3D geometry representationabstractCompression of arbitrary 3D geometry like a human figure in 3D space is challenging. Existing 3D representations like point cloud require encoding of input-specified 3D coordinates, resulting in a large overhead. In this paper, assuming that there exists an underlying smooth 2D manifold in 3D space that describes the geometric shape of a target object, we develop a new progressive 3D geometry representation that signal-adaptively identifies new samples on the manifold surface and encodes them efficiently as graph-signals. Specifically, at each iteration, using previous encoded samples in 3D space, the encoder and decoder first synchronously interpolate a continuous sampling kernel (a 3D mesh) - an approximation of the target surface. We next distribute new sample locations on the continuous kernel based on locally computed kernel curvatures, and compute the signed distances between sample locations and the target surface as sample values. Finally, we connect new discrete samples into a graph for graph-based transform coding of the sample values, which are transmitted to the decoder to refine 3D reconstruction. Experimental results show that our coding scheme outperforms an existing mesh-bsed approach significantly at the low-bitrate region for two different datasets. Gene Cheung, Dinei A. F. Florêncio, Xiangyang Ji |
ICIP | 3 |
| 2017 | Speech Enhancement Using Bayesian Wavenet
Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Xuesong Yang, Dinei A. F. Florêncio, Mark Hasegawa-Johnson |
INTERSPEECH | 5 |
| 2017 | Glottal Model Based Speech Beamforming for ad-hoc Microphone Arrays
Yang Zhang 0001, Dinei A. F. Florêncio, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2017 | Interpolation of Head-Related Transfer Functions Using Manifold LearningabstractWe propose a new head-related transfer function (HRTF) interpolation method using Isomap, a nonlinear dimensionality reduction technique. First, we construct a single manifold for all subjects across both azimuth and elevation angles through the construction of an intersubject graph (ISG) that includes important prior knowledge of the HRTFs such as correlations across individuals, directions, and ears. Then, for a new direction, we predict its corresponding low-dimensional HRTF by interpolating over same subject low-dimensional measured HRTFs. Finally, we use a local neighborhood mapping in the manifold to reconstruct the high-dimensional HRTF from measured HRTFs of all subjects. We show that a single manifold representation obtained through the ISG is a powerful way to allow measured HRTFs from different subjects to contribute for reconstructing the HRTFs for new directions. Moreover, our results suggest that a small number of spatial measurements capture most of acoustical properties of HRTFs. Finally, our approach outperforms other linear and nonlinear dimensionality reduction techniques such as principal component analysis, locally linear embedding, and Laplacian eigenmaps. Felipe Grijalva, Luiz Martini, Dinei A. F. Florêncio, Siome Goldenstein |
IEEE Signal Process. Lett. | 3 |
| 2017 | A Kinect-Based Wearable Face Recognition System to Aid Visually Impaired UsersabstractIn this paper, we introduce a real-time face recognition (and announcement) system targeted at aiding the blind and low-vision people. The system uses a Microsoft Kinect sensor as a wearable device, performs face detection, and uses temporal coherence along with a simple biometric procedure to generate a sound associated with the identified person, virtualized at his/her estimated 3-D location. Our approach uses a variation of the K-nearest neighbors algorithm over histogram of oriented gradient descriptors dimensionally reduced by principal component analysis. The results show that our approach, on average, outperforms traditional face recognition methods while requiring much less computational resources (memory, processing power, and battery life) when compared with existing techniques in the literature, deeming it suitable for the wearable hardware constraints. We also show the performance of the system in the dark, using depth-only information acquired with Kinect's infrared camera. The validation uses a new dataset available for download, with 600 videos of 30 people, containing variation of illumination, background, and movement patterns. Experiments with existing datasets in the literature are also considered. Finally, we conducted user experience evaluations on both blindfolded and visually impaired users, showing encouraging results. Laurindo de Sousa Britto Neto, Felipe Grijalva, Vanessa Regina Margareth Lima Maike, Luiz Martini, Dinei A. F. Florêncio, Maria Cecília Calani Baranauskas, Anderson Rocha 0001, Siome Goldenstein |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2017 | Context Tree-Based Image Contour Coding Using a Geometric PriorabstractEfficient encoding of object contours in images can facilitate advanced image/video compression techniques, such as shape-adaptive transform coding or motion prediction of arbitrarily shaped pixel blocks. We study the problem of lossless and lossy compression of detected contours in images. Specifically, we first convert a detected object contour into a sequence of directional symbols drawn from a small alphabet. To encode the symbol sequence using arithmetic coding, we compute an optimal variable-length context tree (VCT) T via a maximum a posterior (MAP) formulation to estimate symbols' conditional probabilities. MAP can avoid overfitting given a small training set X of past symbol sequences by identifying a VCT T with high likelihood P(X|T) of observing X given T , using a geometric prior P(T) stating that image contours are more often straight than curvy. For the lossy case, we design fast dynamic programming (DP) algorithms that optimally trade off coding rate of an approximate contour [Formula: see text] given a VCT T with two notions of distortion of [Formula: see text] with respect to the original contour x. To reduce the size of the DP tables, a total suffix tree is derived from a given VCT T for compact table entry indexing, reducing complexity. Experimental results show that for lossless contour coding, our proposed algorithm outperforms state-of-the-art context-based schemes consistently for both small and large training datasets. For lossy contour coding, our algorithms outperform comparable schemes in the literature in rate-distortion performance. Amin Zheng, Gene Cheung, Dinei A. F. Florêncio |
IEEE Trans. Image Process. | 3 |
| 2016 | Joint denoising / compression of image contours via geometric prior and variable-length context treeabstractThe advent of depth sensing technologies has eased the detection of object contours in images. For efficient image compression, coded contours can enable edge-adaptive coding techniques such as graph Fourier transform (GFT) and arbitrarily shaped sub-block motion prediction. However, acquisition noise in captured depth images means that detected contours also suffer from errors. In this paper, we propose to jointly denoise and compress detected contours in an image. Specifically, we first propose a burst error model that models typical errors encountered in an observed string y of directional edges. We then formulate a rate-constrained maximum a posteriori (MAP) problem that trades off the posterior probability P(x|y) of an estimated string x given y with its code rate R(x). Given our burst error model, we show that the negative log of the likelihood P(y|x) can be written as a simple sum of burst error events, error symbols and burst lengths, while the geometric prior P(x) states intuitively that contours are more likely straight than curvy. We design a dynamic programming (DP) algorithm that solves the posed problem optimally. Experimental results show that our joint denoising / compression scheme outperformed a competing separate scheme in rate-distortion performance noticeably. Amin Zheng, Gene Cheung, Dinei A. F. Florêncio |
ICIP | 3 |
| 2016 | Speech Enhancement in Multiple-Noise Conditions Using Deep Neural NetworksabstractIn this paper we consider the problem of speech enhancement in real-world like conditions where multiple noises can simultaneously corrupt speech. Most of the current literature on speech enhancement focus primarily on presence of single noise in corrupted speech which is far from real-world environments. Specifically, we deal with improving speech quality in office environment where multiple stationary as well as non-stationary noises can be simultaneously present in speech. We propose several strategies based on Deep Neural Networks (DNN) for speech enhancement in these scenarios. We also investigate a DNN training strategy based on psychoacoustic models from speech coding for enhancement of noisy speech Anurag Kumar 0003, Dinei A. F. Florêncio |
INTERSPEECH | 2 |
| 2016 | A Manifold Learning Approach for Personalizing HRTFs from Anthropometric FeaturesabstractWe present a new anthropometry-based method to personalize head-related transfer functions (HRTFs) using manifold learning in both azimuth and elevation angles with a single nonlinear regression model. The core element of our approach is a domain-specific nonlinear dimensionality reduction technique, denominated Isomap, over the intraconic component of HRTFs resulting from a spectral decomposition. HRTF intraconic components encode the most important cues for HRTF individualization, leaving out subject-independent cues. First, we modify the graph construction procedure of Isomap to integrate relevant prior knowledge of spatial audio into a single manifold for all subjects by exploiting the existing correlations among HRTFs across individuals, directions, and ears. Then, with the aim of preserving the multifactor nature of HRTFs (i.e. subject, direction and frequency), we train a single artificial neural network to predict low-dimensional HRTFs from anthropometric features. Finally, we reconstruct the HRTF from its estimated low-dimensional version using a neighborhood-based reconstruction approach. Our findings show that introducing prior knowledge in Isomap's manifold is a powerful way to capture the underlying factors of spatial hearing. Our experiments show, with p-values less than 0.05, that our approach outperforms using, either a PCA linear reduction, or the full HTRF, in its intermediate stages. Felipe Grijalva, Luiz Martini, Dinei A. F. Florêncio, Siome Goldenstein |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Image Bit-Depth Enhancement via Maximum A Posteriori Estimation of AC SignalabstractWhen images at low bit-depth are rendered at high bit-depth displays, missing least significant bits needs to be estimated. We study the image bit-depth enhancement problem: estimating an original image from its quantized version from a minimum mean squared error (MMSE) perspective. We first argue that a graph-signal smoothness prior-one defined on a graph embedding the image structure-is an appropriate prior for the bit-depth enhancement problem. We next show that directly solving for the MMSE solution is, in general, too computationally expensive to be practical. We then propose an efficient approximation strategy. In particular, we first estimate the ac component of the desired signal in a maximum a posteriori formulation, efficiently computed via convex programming. We then compute the dc component with an MMSE criterion in a closed form given the computed ac component. Experiments show that our proposed two-step approach has improved performance over the conventional bit-depth enhancement schemes in both objective and subjective comparisons. Pengfei Wan 0001, Gene Cheung, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Trans. Image Process. | 3 |
| 2015 | Maximum a posteriori estimation of room impulse responsesabstractEstimating room impulse responses (RIRs) has a number of applications, including personalized audio, analyzing and improving acoustic behavior of concert halls, listening room compensation, sound source localization, and many others. RIRs have been estimated in essentially the same fashion for the last 50 years: Compute the cross correlation between a signal played at point A, and the signal received at point B. Best results are obtained when the signal played is white noise, or a maximum length sequence. No prior knowledge is exploited in computing the RIR, which is simply assumed to be the cross correlation between played and received signals. In contrast, research in adaptive RIR estimation (a.k.a. adaptive Acoustic Echo Cancellation) has made huge progress by (among other things) incorporating models for the RIR. In this paper we propose a new RIR estimation technique, based on a maximum a posteriori formulation. More specifically, we estimate the room reverberation time, as well as the room noise level, and use those as priors for the RIR estimation. Comparison with ground truth shows an average improvement of 12 dB compared to traditional methods. Dinei A. F. Florêncio, Zhengyou Zhang |
ICASSP | 1 |
| 2015 | 3D numerical modeling of parametric speaker using finite-difference time-domainabstractParametric speakers produce sound by emitting ultrasound, and using the small nonlinearity in air to demodulate it back to audible sound. The use of ultrasound allows for producing very narrow audio beams, which finds application in a number of military and consumer scenarios. However, designing better parametric speakers has been hard: closed-form solution of the nonlinear wave equation for generic geometries is nearly impossible, and the only existing solution was derived for the simple case of a cylindrical beam. FDTD methods were considered not practical since the desired (audible) signal is orders of magnitude weaker than the ultrasound signal, and thus the noise floor (from the numerical approximation) will dwarf the audible signal. In this paper, we introduce a novel FDTD scheme that models nonlinear sound propagation for parametric speakers. By taking the difference between linear and nonlinear FDTD simulations, we successfully suppress the numerical noise floor and extract the audible signal. Both spectrum and radiation pattern of simulation match the measurements well. This offers a simulation tool for further research in creating advanced parametric speakers. Dinei A. F. Florêncio |
ICASSP | 2 |
| 2015 | Precision Enhancement of 3-D Surfaces from Compressed Multiview Depth MapsabstractTransmitting depth maps captured from multiple viewpoints of a 3-D scene enables a wide range of receiver-side 3-D applications, including virtual view synthesis via depth-image-based rendering (DIBR). Observing that compressed depth maps from different viewpoints constitute multiple descriptions (MD) of the same signal, we propose to reconstruct 3-D surfaces of the scene by considering multiple compressed depth maps jointly. Specifically, we propose an alternating projection algorithm, inspired by the theory of projection onto convex sets (POCS), which at convergence returns a 3-D surface that satisfies three sets of conditions: spatial smoothness prior, quantization bin constraints in the block transform domain, and inter-view consistency. We present a theoretical proof that shows convergence of our algorithm under benign conditions. Compared to existing multiview depth map denoising schemes and single image de-quantization schemes, our proposed solution achieves higher objective quality for both reconstructed depth maps and synthesized virtual views. Pengfei Wan 0001, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Signal Process. Lett. | 4 |
| 2014 | Anthropometric-based customization of head-related transfer functions using Isomap in the horizontal planeabstractIn this paper, we introduce a new anthropometric-based method for customizing of Head-Related Transfer Functions (HRTF) in the horizontal plane. The method uses Isomap, artificial neural networks (ANN), and a neighborhood-based reconstruction procedure. We first modify Isomap's graph construction step to emphasize the individuality of HRTFs and perform a customized nonlinear dimensionality reduction of the HTRFs. We then use an ANN to model the nonlinear relationship between anthropometric features and our low-dimensional HRTFs. Finally, we use a neighborhood-based reconstruction approach to reconstruct the HRTF from the estimated low-dimensional version. Simulations show that our approach performs better than PCA and confirm that Isomap is capable of discovering the underlying nonlinear relationships of sound perception. Felipe Grijalva, Luiz Martini, Siome Goldenstein, Dinei A. F. Florêncio |
ICASSP | 4 |
| 2014 | Image bit-depth enhancement via maximum-a-posteriori estimation of graph AC componentabstractWhile modern displays offer high dynamic range (HDR) with large bit-depth for each rendered pixel, the bulk of legacy image and video contents were captured using cameras with shallower bit-depth. In this paper, we study the bit-depth enhancement problem for images, so that a high bit-depth (HBD) image can be reconstructed from an input low bit-depth (LBD) image. The key idea is to apply appropriate smoothing given the constraints that reconstructed signal must lie within the per-pixel quantization bins. Specifically, we first define smoothness via a signal-dependent graph Laplacian, so that natural image gradients can nonetheless be interpreted as low frequencies. Given defined smoothness prior and observed LBD image, we then demonstrate that computing the most probable signal via maximum a posteriori (MAP) estimation can lead to large expected distortion. However, we argue that MAP can still be used to efficiently estimate the AC component of the desired HBD signal, which along with a distortion-minimizing DC component, can result in a good approximate solution that minimizes the expected distortion. Experimental results show that our proposed method outperforms existing bit-depth enhancement methods in terms of reconstruction error. Pengfei Wan 0001, Gene Cheung, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
ICIP | 3 |
| 2014 | Point cloud attribute compression with graph transformabstractCompressing attributes on 3D point clouds such as colors or normal directions has been a challenging problem, since these attribute signals are unstructured. In this paper, we propose to compress such attributes with graph transform. We construct graphs on small neighborhoods of the point cloud by connecting nearby points, and treat the attributes as signals over the graph. The graph transform, which is equivalent to Karhunen-Loève Transform on such graphs, is then adopted to decorrelate the signal. Experimental results on a number of point clouds representing human upper bodies demonstrate that our method is much more efficient than traditional schemes such as octree-based methods. Cha Zhang, Dinei A. F. Florêncio, Charles T. Loop |
ICIP | 2 |
| 2014 | An Administrator's Guide to Internet Password Research
Dinei A. F. Florêncio, Cormac Herley, Paul C. van Oorschot |
LISA | 1 |
| 2014 | Password Portfolios and the Finite-Effort User: Sustainably Managing Large Numbers of Accounts
Dinei A. F. Florêncio, Cormac Herley, Paul C. van Oorschot |
USENIX Security Symposium | 1 |
| 2014 | Guest editorial: Advances in 3D video processing
Shang-Hong Lai, Gene Cheung, Dinei A. F. Florêncio, Peter Eisert, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Sparse Array-Based Room Transfer Function Estimation for Echo CancellationabstractA number of applications in acoustics, such as echo cancellation, require learning the acoustic impulse response from each deployed loudspeaker to each microphone- the room transfer function. This has conventionally been done separately at each microphone for each loudspeaker. However, the signals arriving at the array share a common structure, which can be exploited to improve the impulse response estimates. In this work, we propose an algorithm that takes advantage of the array structure, as well as the sparsity of the reflections arriving at the array in order to form reliable estimates of the impulse response between each loudspeaker and microphone. The algorithm is shown to improve performance over the matched filter algorithm in echo cancellation applications, using both synthetic and real data. Atulya Yellepeddi, Dinei A. F. Florêncio |
IEEE Signal Process. Lett. | 2 |
| 2014 | Arbitrarily Shaped Motion Prediction for Depth Video Compression Using Arithmetic Edge CodingabstractDepth image compression is important for compact representation of 3D visual data in texture-plus-depth format, where texture and depth maps from one or more viewpoints are encoded and transmitted. A decoder can then synthesize a freely chosen virtual view via depth-image-based rendering using nearby coded texture and depth maps as reference. Further, depth information can be used in other image processing applications beyond view synthesis, such as object identification, segmentation, and so on. In this paper, we leverage on the observation that neighboring pixels of similar depth have similar motion to efficiently encode depth video. Specifically, we divide a depth block containing two zones of distinct values (e.g., foreground and background) into two arbitrarily shaped regions (sub-blocks) along the dividing boundary before performing separate motion prediction (MP). While such arbitrarily shaped sub-block MP can lead to very small prediction residuals (resulting in few bits required for residual coding), it incurs an overhead to transmit the dividing boundaries for sub-block identification at decoder. To minimize this overhead, we first devise a scheme called arithmetic edge coding (AEC) to efficiently code boundaries that divide blocks into sub-blocks. Specifically, we propose to incorporate the boundary geometrical correlation in an adaptive arithmetic coder in the form of a statistical model. Then, we propose two optimization procedures to further improve the edge coding performance of AEC for a given depth image. The first procedure operates within a code block, and allows lossy compression of the detected block boundary to lower the cost of AEC, with an option to augment boundary depth pixel values matching the new boundary, given the augmented pixels do not adversely affect synthesized view distortion. The second procedure operates across code blocks, and systematically identifies blocks along an object contour that should be coded using sub-block MP via a rate-distortion optimized trellis. Experimental results show an average overall bitrate reduction of up to 33% over classical H.264/AVC. Ismaël Daribo, Dinei A. F. Florêncio, Gene Cheung |
IEEE Trans. Image Process. | 2 |
| 2014 | Rate-Constrained 3D Surface Estimation From Noise-Corrupted Multiview Depth VideosabstractTransmitting compactly represented geometry of a dynamic 3D scene from a sender can enable a multitude of imaging functionalities at a receiver, such as synthesis of virtual images at freely chosen viewpoints via depth-image-based rendering. While depth maps—projections of 3D geometry onto 2D image planes at chosen camera viewpoints-can nowadays be readily captured by inexpensive depth sensors, they are often corrupted by non-negligible acquisition noise. Given depth maps need to be denoised and compressed at the encoder for efficient network transmission to the decoder, in this paper, we consider the denoising and compression problems jointly, arguing that doing so will result in a better overall performance than the alternative of solving the two problems separately in two stages. Specifically, we formulate a rate-constrained estimation problem, where given a set of observed noise-corrupted depth maps, the most probable (maximum a posteriori (MAP)) 3D surface is sought within a search space of surfaces with representation size no larger than a prespecified rate constraint. Our rate-constrained MAP solution reduces to the conventional unconstrained MAP 3D surface reconstruction solution if the rate constraint is loose. To solve our posed rate-constrained estimation problem, we propose an iterative algorithm, where in each iteration the structure (object boundaries) and the texture (surfaces within the object boundaries) of the depth maps are optimized alternately. Using the MVC codec for compression of multiview depth video and MPEG free viewpoint video sequences as input, experimental results show that rate-constrained estimated 3D surfaces computed by our algorithm can reduce coding rate of depth maps by up to 32% compared with unconstrained estimated surfaces for the same quality of synthesized virtual views at the decoder. Wenxiu Sun, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Trans. Image Process. | 4 |
| 2013 | Rate-distortion optimized 3D reconstruction from noise-corrupted multiview depth videosabstractTransmitting compactly represented geometry of a dynamic scene from a sender can enable a multitude of 3D imaging functionalities at a receiver, such as synthesis of virtual images from freely chosen viewpoints via depth-image-based rendering (DIBR). While depth maps can now be readily captured using inexpensive depth sensors, they are often corrupted by non-negligible acquisition noise. In this paper, we derive 3D surfaces of a dynamic scene from noise-corrupted depth maps in a rate-distortion (RD) optimal manner. Specifically, unlike previous work that finds the most likely (e.g., maximum likelihood) 3D surface from noisy observations regardless of representation size, we judiciously search for the best fitting (i.e., minimum distortion) 3D surface subject to a bitrate constraint. Our RD-optimal solution reduces to the maximum likelihood solution as the rate constraint is loosened. Using the MVC codec for compression of multiview depth video and MPEG free viewpoint test sequences as input, experimental results show that RD-optimized 3D reconstructions computed by our algorithm outperform unprocessed depth maps by up to 2:42dB in PSNR of synthesized virtual views at the decoder for the same bitrate. Wenxiu Sun, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
ICME | 4 |
| 2013 | Autonomous person following for telepresence robotsabstractWe present a method for a mobile robot to follow a person autonomously where there is an interaction between the robot and human during following. The planner takes into account the predicted trajectory of the human and searches future trajectories of the robot for the path with the highest utility. Contrary to traditional motion planning, instead of determining goal points close to the person, we introduce a task dependent goal function which provides a map of desirable areas for the robot to be at, with respect to the person. The planning framework is flexible and allows encoding of different social situations with the help of the goal function. We implemented our approach on a telepresence robot and conducted a controlled user study to evaluate the experiences of the users on the remote end of the telepresence robot. The user study compares manual teleoperation to our autonomous method for following a person while having a conversation. By designing a behavior specific to a flat screen telepresence robot, we show that the person following behavior is perceived as safe and socially acceptable by remote users. All 10 participants preferred our autonomous following method over manual teleoperation. Akansel Cosgun, Dinei A. F. Florêncio, Henrik I. Christensen |
ICRA | 2 |
| 2013 | Learning how to increase the chance of human-robot engagementabstractThe increasing use of mobile robots in social contexts makes it important to provide them with the ability to behave in the most socially acceptable way possible. In this paper we investigate the problem of making a robot learn how to approach a person in order to increase the chance of a successful engagement. We propose the use of Gaussian Process Regression (GPR), combined with ideas from reinforcement learning to make sure the space is properly and continuously explored. In the proposed example scenario, this is used by the robot to predict the best decisions in relation to its position in the environment and approach distance, each one accordingly to a certain time of the day. Numerical simulations show a significant performance improvement when compared with a random technique. The robot is able to improve performance after just one day of interaction (a few dozens of trials), and achieves the maximum expected value for the proposed approach within sixty days. Douglas G. Macharet, Dinei A. F. Florêncio |
IROS | 2 |
| 2013 | Attention-Weighted Rate Allocation in Free-Viewpoint TelevisionabstractAn architecture for free-viewpoint broadcast television transmission is proposed where all the views are transmitted at potentially different qualities and watched by a large number of viewers. The quality (or bit-rate) of each view is controlled by the distribution of viewpoints chosen by the viewers. For example, if most viewers are watching synthetic views in between viewsnandn+1, those views are allocated more transmission bits than views that are scarcely watched. We developed an attention-weighted bit-rate-allocation method that is optimal in the total observer distortion sense. The optimality of the method relies on knowing the viewpoint probability distribution at every moment. Simulation results show that overall transmission rate can be reduced for the same total observed distortion. Thacio Scandarolli, Ricardo L. de Queiroz, Dinei A. F. Florêncio |
IEEE Signal Process. Lett. | 3 |
| 2013 | Analyzing the Optimality of Predictive Transform Coding Using Graph-Based ModelsabstractIn this letter, we provide a theoretical analysis of optimal predictive transform coding based on the Gaussian Markov random field (GMRF) model. It is shown that the eigen-analysis of the precision matrix of the GMRF model is optimal in decorrelating the signal. The resulting graph transform degenerates to the well-known 2-D discrete cosine transform (DCT) for a particular 2-D first order GMRF, although it is not a unique optimal solution. Furthermore, we present an optimal scheme to perform predictive transform coding based on conditional probabilities of a GMRF model. Such an analysis can be applied to both motion prediction and intra-frame predictive coding, and may lead to improvements in coding efficiency in the future. Cha Zhang, Dinei A. F. Florêncio |
IEEE Signal Process. Lett. | 2 |
| 2012 | Arithmetic edge coding for arbitrarily shaped sub-block motion prediction in depth video compressionabstractDepth map compression is important for compact representation of 3D visual data in “texture-plus-depth” format, where texture and depth maps of multiple closely spaced viewpoints are encoded and transmitted. A decoder can then freely synthesize any chosen inter-mediate view via depth-image-based rendering (DIBR) using neighboring coded texture and depth maps as anchors. In this work, we leverage on the observation that “pixels of similar depth have similar motion” to efficiently encode depth video. Specifically, we divide a depth block containing two zones of distinct values (e.g., foreground and background) into two sub-blocks along the dividing edge before performing separate motion prediction. While doing such arbitrarily shaped sub-block motion prediction can lead to very small prediction residuals (resulting in few bits required to code them), it incurs an overhead to losslessly encode dividing edges for sub-block identification. To minimize this overhead, we first devise an edge prediction scheme based on linear regression to predict the next edge direction in a contiguous contour. From the predicted edge direction, we assign probabilities to each possible edge direction using the von Mises distribution, which are subsequently inputted to a conditional arithmetic codec for entropy coding. Experimental results show an average overall bitrate reduction of up to 30% over classical H.264 implementation. Ismaël Daribo, Gene Cheung, Dinei A. F. Florêncio |
ICIP | 3 |
| 2012 | A collaborative control system for telepresence robotsabstractInterest in telepresence robots is at an all time high, and several companies are already commercializing early or basic versions. There seems to be a huge potential for their use in professional applications, where they can help address some of the challenges companies have found in integrating a geographically distributed work force. However, teleoperation of these robots is typically a difficult task. This difficulty can be attributed to limitations on the information provided to the operator and to communication delay and failures. This may compromise the safety of the people and of the robot during its navigation through the environment. Most commercial systems currently control this risk by reducing size and weight of their robots. Research effort in addressing this problem is generally based on “assisted driving”, which typically adds a “collision avoidance” layer, limiting or avoiding movements that would lead to a collision. In this article, we bring assisted driving to a new level, by introducing concepts from collaborative driving to telepresence robots. More specifically, we use the input from the operator as a general guidance to the target direction, then couple that with a variable degree of autonomy to the robot, depending on the task and the environment. Previous work has shown collision avoidance makes operation easier and reduce the number of collisions. In addition (and in contrast to traditional collision avoidance systems), our approach also reduces the time required to complete a circuit, making navigation easier, safer, and faster. The methodology was evaluated through a controlled user study (N=18). Results show that the use of the proposed collaborative control helped reduce the number of collisions (none in most cases) and also decreased the time to complete the designated task. Douglas G. Macharet, Dinei A. F. Florêncio |
IROS | 2 |
| 2012 | Auditory augmented reality: Object sonification for the visually impairedabstractAugmented reality applications have focused on visually integrating virtual objects into real environments. In this paper, we propose an auditory augmented reality, where we integrate acoustic virtual objects into the real world. We sonify objects that do not intrinsically produce sound, with the purpose of revealing additional information about them. Using spatialized (3D) audio synthesis, acoustic virtual objects are placed at specific real-world coordinates, obviating the need to explicitly tell the user where they are. Thus, by leveraging the innate human capacity for 3D sound source localization and source separation, we create an audio natural user interface. In contrast with previous work, we do not create acoustic scenes by transducing low-level (for instance, pixel-based) visual information. Instead, we use computer vision methods to identify high-level features of interest in an RGB-D stream, which are then sonified as virtual objects at their respective real-world coordinates. Since our visual and auditory senses are inherently spatial, this technique naturally maps between these two modalities, creating intuitive representations. We evaluate this concept with a head-mounted device, featuring modes that sonify flat surfaces, navigable paths and human faces. Flavio P. Ribeiro, Dinei A. F. Florêncio, Philip A. Chou, Zhengyou Zhang |
MMSP | 2 |
| 2012 | Arbitrarily shaped sub-block motion prediction in texture map compression using depth informationabstractWhen transmitting the so-called “texture-plus-depth” video format, texture and depth maps from the same viewpoint exhibit high correlation. Coded bits from one map can then be used as side information to encode the other. In this paper, we propose to use the depth information to divide the corresponding block in texture map into arbitrarily shaped regions (sub-blocks) for separate motion estimation (ME) and motion compensation (MC). We implemented our proposed sub-block motion prediction (MP) method for texture map coding using depth information as a new coding mode (z-mode) in H.264. Nonetheless, in practical experiments one can observe either a misalignment between texture and depth edges, or an aliasing effect at the texture boundaries. To overcome this issue, z-mode offers two MC types: i) non-overlapping MC, and ii) overlapping MC. In the latter case, overlapped sub-blocks after ME are alpha-blended using a properly designed filter. Moreover, the MV of each sub-block in z-mode is predicted using a Laplacian-weighted average of MVs of neighboring blocks of similar depth. Experimental results show that using z-mode, coding performance of the texture map can be improved by up to 0.7dB compared to native H.264 implementation at high bitrate. Ismaël Daribo, Dinei A. F. Florêncio, Gene Cheung |
PCS | 2 |
| 2012 | Geometrically Constrained Room Modeling With Compact Microphone ArraysabstractThe geometry of an acoustic environment can be an important information in many audio signal processing applications. To estimate such a geometry, previous work has relied on large microphone arrays, multiple test sources, moving sources or the assumption of a 2-D room. In this paper, we lift these requirements and present a novel method that uses a compact microphone array to estimate a 3-D room geometry, delivering effective estimates with low-cost hardware. Our approach first probes the environment with a known test signal emitted by a loudspeaker co-located with the array, from which the room impulse responses (RIRs) are estimated. It then uses an ℓ1-regularized least-squares minimization to fit synthetically generated reflections to the RIRs, producing a sparse set of reflections. By enforcing structural constraints derived from the image model, these are classified into first-, second-, and third-order reflections, thereby deriving the room geometry. Using this method, we detect walls using off-the-shelf teleconferencing hardware with a typical range resolution of about 1 cm. We present results using simulations and data from real environments. Flavio P. Ribeiro, Dinei A. F. Florêncio, Demba Ba 0001, Cha Zhang |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Introduction to the ICME 2011 Special IssueabstractThe 14 papers in this special issue are extended versions of papers presented at ICME 2011, held in Barcelona, Spain, on 11-15 July 2011. Dinei A. F. Florêncio, Sethuraman Panchanathan, Philippe Salembier, Mohamed Hefeeda, Alexander C. Loui, Mrinal Mandal 0001 |
IEEE Trans. Multim. | 2 |
| 2011 | CROWDMOS: An approach for crowdsourcing mean opinion score studiesabstractMOS (mean opinion score) subjective quality studies are used to evaluate many signal processing methods. Since laboratory quality studies are time consuming and expensive, researchers often run small studies with less statistical significance or use objective measures which only approximate human perception. We propose a cost-effective and convenient measure called crowdMOS, obtained by having internet users participate in a MOS-like listening study. Workers listen and rate sentences at their leisure, using their own hardware, in an environment of their choice. Since these individuals cannot be supervised, we propose methods for detecting and discarding inaccurate scores. To automate crowdMOS testing, we offer a set of freely distributable, open-source tools for Amazon Mechanical Turk, a platform designed to facilitate crowdsourcing. These tools implement the MOS testing methodology described in this pa per, providing researchers with a user-friendly means of performing subjective quality evaluations without the overhead associated with laboratory studies. Finally, we demonstrate the use of crowdMOS using data from the Blizzard text-to-speech competition, showing that it delivers accurate and repeatable results. Flavio P. Ribeiro, Dinei A. F. Florêncio, Cha Zhang, Michael L. Seltzer |
ICASSP | 2 |
| 2011 | Crowdsourcing subjective image quality evaluationabstractSubjective tests are generally regarded as the most reliable and definitive methods for assessing image quality. Nevertheless, laboratory studies are time consuming and expensive. Thus, researchers often choose to run informal studies or use objective quality measures, producing results which may not correlate well with human perception. In this paper we propose a cost-effective and convenient subjective quality measure called crowdMOS, obtained by having Internet workers participate in MOS (mean opinion score) subjective quality studies. Since these workers cannot be supervised, we propose methods for detecting and discarding inaccurate or malicious scores. To facilitate this process, we offer an open source set of tools for Amazon Mechanical Turk, which is an Internet marketplace for crowdsourcing. These tools completely automate the test design, score retrieval and statistical analysis, abstracting away the technical details of Mechanical Turk and ensuring a user-friendly, affordable and consistent test methodology. We demonstrate crowdMOS using data from the LIVE subjective quality image dataset, showing that it delivers accurate and repeatable results. Flavio P. Ribeiro, Dinei A. F. Florêncio, Vítor H. Nascimento |
ICIP | 2 |
| 2011 | Enhanced adaptive playout scheduling and loss concealment techniques for Voice over IP networksabstractVoice over IP (VoIP) is already of commercial quality when traversing corporate, or other high-quality networks. Nevertheless, to make the final bridge to toll quality over standard internet connections, a few problems remains to be solved. These include packet loss, jitter control, and clock drift compensation. Previous research in adaptive playout mechanisms has significantly contributed towards that end. In this paper, we present a novel technique that enhances previous adaptive playout mechanisms, and virtually eliminate late losses. Additionally, enhanced stretching and loss concealment algorithms are also presented. These work in harmony with the proposed technique to alleviate network jitter and any clock drift. Dinei A. F. Florêncio, Li-wei He |
ISCAS | 1 |
| 2011 | Region of interest determination using human computationabstractThe ability to identify and track visually interesting regions has many practical applications - for example, in image and video compression, visual marketing and foveal machine vision. Due to challenges in modeling the peculiarities of human physiological and psychological responses, automatic detection of fixation points is an open problem. Indeed, no objective methods are currently capable of fully modeling the human perception of regions of interest (ROIs). Thus, research often relies on user studies with eye tracking systems. In this paper we propose a cost-effective and convenient alternative, obtained by having internet workers annotate videos with ROI coordinates. The workers use an interactive video player with a simulated mouse-driven fovea, which models the fall-off in resolution of the human visual system. Since this approach is not supervised, we implement methods for identifying inaccurate or malicious results. Using this proposal, one can collect ROI data in an automated fashion, and at a much lower cost than laboratory studies. Flavio P. Ribeiro, Dinei A. F. Florêncio |
MMSP | 2 |
| 2011 | An Interactive 3-D Audio System With LoudspeakersabstractTraditional 3-D audio systems using two loudspeakers often have a limited sweet spot and may suffer from poor performance in reverberant environments. This paper presents a novel binaural 3-D audio system that actively combines head tracking and room modeling into 3-D audio synthesis. The user's head position and orientation are first tracked by a webcam-based 3-D head tracker. The system then improves its robustness to head movement and strong early reflections by incorporating the tracking information and an explicit room model into the binaural synthesis and crosstalk cancellation process. Sensitivity analysis on the room model shows that the method is reasonably robust to modeling errors. Subjective listening tests confirm that the proposed 3-D audio system significantly improves the users' perception and ability for localization. Myung-Suk Song, Cha Zhang, Dinei A. F. Florêncio, Hong-Goo Kang |
IEEE Trans. Multim. | 3 |
| 2010 | L1 regularized room modeling with compact microphone arraysabstractAcoustic room modeling has several applications. Recent results using large microphone arrays show good performance, and are helpful in many applications. For example, when designing a better acoustic treatment for a concert hall, these large arrays can be used to help map the acoustic environment and aid in the design. However, in real-time applications - including de-reverberation, sound source localization, speech enhancement and 3D audio - it is desirable to model the room with existing small arrays and existing loudspeakers. In this paper we propose a novel room modeling algorithm, which uses a constrained room model and ℓ1-regularized least-squares to achieve good estimation of room geometry. We present experimental results on both real and synthetic data. Demba Ba 0001, Flavio P. Ribeiro, Cha Zhang, Dinei A. F. Florêncio |
ICASSP | 4 |
| 2010 | Turning enemies into friends: Using reflections to improve sound source localizationabstractSound Source Localization (SSL) based on microphone arrays has numerous applications, and has received significant research attention. Common to all published research is the observation that the accuracy of SSL degrades with reverberation. Indeed, early (strong) reflections can have amplitudes similar to the direct signal, and will often interfere with the estimation. In this paper, we show that reverberation is not the enemy, and can be used to improve estimation. More specifically, we are able to use early reflections to significantly improve range and elevation estimation. The process requires two steps: during setup, a loudspeaker integrated with the array emits a probing sound, which is used to obtain estimates of the ceiling height, as well as the locations of the walls. In a second step (e.g., during a meeting), the device incorporates this knowledge into a maximum likelihood SSL algorithm. Experimental results on both real and synthetic data show huge improvements in range estimation accuracy. Flavio P. Ribeiro, Demba Ba 0001, Cha Zhang, Dinei A. F. Florêncio |
ICME | 4 |
| 2010 | Personal 3D audio system with loudspeakersabstractTraditional 3D audio systems often have a limited sweet spot for the user to perceive 3D effects successfully. In this paper, we present a personal 3D audio system with loudspeakers that has unlimited sweet spots. The idea is to have a camera track the user's head movement, and recompute the crosstalk canceller filters accordingly. As far as the authors are aware of, our system is the first non-intrusive 3D audio system that adapts to both the head position and orientation with six degrees of freedom. The effectiveness of the proposed system is demonstrated with subjective listening tests comparing our system against traditional non-adaptive systems. Myung-Suk Song, Cha Zhang, Dinei A. F. Florêncio, Hong-Goo Kang |
ICME | 3 |
| 2010 | Enhancing loudspeaker-based 3D audio with room modelingabstractFor many years, spatial (3D) sound using headphones has been widely used in a number of applications. A rich spatial sensation is obtained by using head related transfer functions (HRTF) and playing the appropriate sound through headphones. In theory, loudspeaker audio systems would be capable of rendering 3D sound fields almost as rich as headphones, as long as the room impulse responses (RIRs) between the loudspeakers and the ears are known. In practice, however, obtaining these RIRs is hard, and the performance of loudspeaker based systems is far from perfect. New hope has been recently raised by a system that tracks the user's head position and orientation, and incorporates them into the RIRs estimates in real time. That system made two simplifying assumptions: it used generic HRTFs, and it ignored room reverberation. In this paper we tackle the second problem: we incorporate a room reverberation estimate into the RIRs. Note that this is a nontrivial task: RIRs vary significantly with the listener's positions, and even if one could measure them at a few points, they are notoriously hard to interpolate. Instead, we take an indirect approach: we model the room, and from that model we obtain an estimate of the main reflections. Position and characteristics of walls do not vary with the users' movement, yet they allow to quickly compute an estimate of the RIR for each new user position. Of course the key question is whether the estimates are good enough. We show an improvement in localization perception of up to 32% (i.e., reducing average error from 23.5° to 15.9°). Myung-Suk Song, Cha Zhang, Dinei A. F. Florêncio, Hong-Goo Kang |
MMSP | 3 |
| 2010 | Where do security policies come from?abstractWe examine the password policies of 75 different websites. Our goal is understand the enormous diversity of requirements: some will accept simple six-character passwords, while others impose rules of great complexity on their users. We compare different features of the sites to find which characteristics are correlated with stronger policies. Our results are surprising: greater security demands do not appear to be a factor. The size of the site, the number of users, the value of the assets protected and the frequency of attacks show no correlation with strength. In fact we find the reverse: some of the largest, most attacked sites with greatest assets allow relatively weak passwords. Instead, we find that those sites that accept advertising, purchase sponsored links and where the user has a choice show strong inverse correlation with strength. Dinei A. F. Florêncio, Cormac Herley |
SOUPS | 1 |
| 2010 | Joint tracking and multiview video compressionabstractIn immersive communication applications, knowing the user's viewing position can help improve the efficiency of multiview compression and streaming significantly, since often only a subset of the views are needed to synthesize the desired view(s). However, uncertainty regarding the viewer location can have negative impacts on the rendering quality. In this paper, we propose an algorithm to improve the robustness of view-dependent compression schemes by jointly performing user tracking and compression. A face tracker tracks the user's head location and sends the probability distribution of the face locations as one or many particles. The server then applies motion model to the particles and compresses the multiview video accordingly in order to improve the expected rendering quality of the viewer. Experimental results show significantly improved robustness against tracking errors. Cha Zhang, Dinei A. F. Florêncio |
VCIP | 2 |
| 2010 | Using Reverberation to Improve Range and Elevation Discrimination for Small Array Sound Source LocalizationabstractSound source localization (SSL) is an essential task in many applications involving speech capture and enhancement. As such, speaker localization with microphone arrays has received significant research attention. Nevertheless, existing SSL algorithms for small arrays still have two significant limitations: lack of range resolution, and accuracy degradation with increasing reverberation. The latter is natural and expected, given that strong reflections can have amplitudes similar to that of the direct signal, but different directions of arrival. Therefore, correctly modeling the room and compensating for the reflections should reduce the degradation due to reverberation. In this paper, we show a stronger result. If modeled correctly, early reflections can be used to provide more information about the source location than would have been available in an anechoic scenario. The modeling not only compensates for the reverberation, but also significantly increases resolution for range and elevation. Thus, we show that under certain conditions and limitations, reverberation can be used to improve SSL performance. Prior attempts to compensate for reverberation tried to model the room impulse response (RIR). However, RIRs change quickly with speaker position, and are nearly impossible to track accurately. Instead, we build a 3-D model of the room, which we use to predict early reflections, which are then incorporated into the SSL estimation. Simulation results with real and synthetic data show that even a simplistic room model is sufficient to produce significant improvements in range and elevation estimation, tasks which would be very difficult when relying only on direct path signal components. Flavio P. Ribeiro, Cha Zhang, Dinei A. F. Florêncio, Demba Ba 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Multiview video compression and streaming based on predicted viewer positionabstractTechnological advances have made possible a number of new applications in the area of 3D video. One of the enabling technologies for many of these 3D applications is multiview video coding, which has received significant attention in the last several years. However, the fundamental need of multiview coding for applications like immersive tele-conferencing has not been addressed. In this paper we define the boundaries of the problem, and show how a simple algorithm can yield gains of up to 2times reduction in bitrate with similar PSNR in the synthesized view. Our algorithm is based on using an estimate of the viewer position to compute the expected contribution of each pixel to the synthesized view, and encoding each macroblock of each camera views with quality proportional to the likelihood that the pixel will be used in the synthetic image. Dinei A. F. Florêncio, Cha Zhang |
ICASSP | 1 |
| 2009 | Background recovery from video sequences using motion parametersabstractThis paper presents a novel scheme for extracting a still background occluded by a number of foreground objects, moving in different directions and velocities in a video sequence, such that every background pixel is exposed in at least one of the frames. Each identified foreground object is decomposed into blocks. The proposed scheme is able to efficiently estimate, for each foreground block, a source frame from which the occluded background pixels can be extracted. The pixels of the identified source frames are used to populate the co-located occluded pixels in the initial frame. The efficacy and the simplicity of the algorithm lie in its capacity to recover the background directly from the estimated source frames instead of performing a foreground-background classification for every frame. The proposed algorithm is robust to variations in lighting and is effective in removing both rigid and deformable foreground objects. Simulation results are presented to illustrate the performance of the proposed scheme. Srenivas Varadarajan, Lina J. Karam, Dinei A. F. Florêncio |
ICASSP | 3 |
| 2009 | Improving depth perception with motion parallax and its application in teleconferencingabstractDepth perception, or 3D perception, can add a lot to the feeling of immersiveness in many applications such as 3D TV, 3D teleconferencing, etc. Stereopsis and motion parallax are two of the most important cues for depth perception. Most of the 3D displays today rely on stereopsis to create 3D perception. In this paper, we propose to improve user's depth perception by tracking their motions and creating motion parallax for the rendered image, which can be done even with legacy displays. Two enabling technologies, face tracking and foreground/background segmentation, are discussed in detail. In particular, we propose an efficient and robust feature based face tracking algorithm that is capable of estimating the face's location and scale accurately. We also propose a novel foreground/background segmentation and matting algorithm with time-of-flight camera, which is robust to moving background, lighting variations, moving camera, etc. We demonstrate the application of the above technologies in teleconferencing on legacy displays to create pseudo-3D effects. Cha Zhang, Zhaozheng Yin, Dinei A. F. Florêncio |
MMSP | 3 |
| 2008 | Why does PHAT work well in lownoise, reverberative environments?abstractAmong many existing time difference of arrival (TDOA) based sound source localization (SSL) algorithms, the Phase Transform (PHAT) is extremely popular for its excellent performance in low noise environments, even under relatively heavy reverberation. However, PHAT was developed as a heuristic approach and its working principle has not been completely understood. In this paper, we present the relationship between PHAT and a maximum likelihood (ML) framework for multi-microphone sound source localization. We show that when the environment noise approaches zero, PHAT is indeed a special case of the ML algorithm, which explains its good performance under low noise environments. In addition, we show that as long as the noise stays low, PHAT remains optimal in ML sense even when the room reverberation is heavy, which explains its robustness over reverberation. Cha Zhang, Dinei A. F. Florêncio, Zhengyou Zhang |
ICASSP | 2 |
| 2008 | One-Time Password Access to Any Server without Changing the Server
Dinei A. F. Florêncio, Cormac Herley |
ISC | 1 |
| 2008 | A profitless endeavor: phishing as tragedy of the commonsabstractConventional wisdom is that phishing represents easy money. In this paper we examine the economics that underly the phenomenon, and find a very different picture. Phishing is a classic example of tragedy of the commons, where there is open access to a resource that has limited ability to regenerate. Since each phisher independently seeks to maximize his return, the resource is over-grazed and yields far less than it is capable of. The situation stabilizes only when the average phisher is making only as much as he gives up in opportunity cost. Cormac Herley, Dinei A. F. Florêncio |
NSPW | 2 |
| 2008 | Protecting Financial Institutions from Brute-Force Attacks
Cormac Herley, Dinei A. F. Florêncio |
SEC | 2 |
| 2008 | Maximum Likelihood Sound Source Localization and Beamforming for Directional Microphone Arrays in Distributed MeetingsabstractIn distributed meeting applications, microphone arrays have been widely used to capture superior speech sound and perform speaker localization through sound source localization (SSL) and beamforming. This paper presents a unified maximum likelihood framework of these two techniques, and demonstrates how such a framework can be adapted to create efficient SSL and beamforming algorithms for reverberant rooms and unknown directional patterns of microphones. The proposed method is closely related to steered response power-based algorithms, which are known to work extremely well in real-world environments. We demonstrate the effectiveness of the proposed method on challenging synthetic and real-world datasets, including over six hours of recorded meetings. Cha Zhang, Dinei A. F. Florêncio, Demba Ba 0001, Zhengyou Zhang |
IEEE Trans. Multim. | 2 |
| 2007 | Maximum Likelihood Sound Source Localization for Multiple Directional MicrophonesabstractThis paper presents a maximum likelihood (ML) framework for multi-microphone sound source localization (SSL). Besides deriving the framework, we focus on making the connection and contrast between the ML-based algorithm and popular steered response power (SRP) SSL algorithms such as phase transform (SRP-PHAT). We also show under our ML framework how challenging conditions such as directional microphone arrays and reverberations can be handled. The computational cost of our method is low-similar to SRP-PHAT. The effectiveness of the proposed method is shown on a large dataset with 99 real-world audio sequences recorded by directional circular microphone arrays in over 50 different meeting rooms. Cha Zhang, Zhengyou Zhang, Dinei A. F. Florêncio |
ICASSP (1) | 3 |
| 2007 | Enhanced MVDR Beamforming for Arrays of Directional MicrophonesabstractMicrophone arrays based on the minimum variance distortionless response (MVDR) beamformer are among the most popular for speech enhancement applications. The original MVDR is excessively sensitive to source location and microphone gains. Previous research has made MVDR practical by successfully increasing the robustness of MVDR to source location, and MVDR-based microphone arrays are already commercially available. Nevertheless, MVDR performance is still weak in cases where microphone gain variations are too large, e.g., for circular arrays of directional microphones. In this paper we propose an improved MVDR beamformer which takes into account the effect of sensors (e.g. microphones) with arbitrary, potentially directional responses. Specifically, we form estimates of the relative magnitude responses of the sensors based on the data received at the array and include those in the original formulation of the MVDR beamforming problem. Experimental results on real-world audio data show an average 2.4 dB improvement over conventional MVDR beamforming, which does not account for the magnitude responses of the sensors. Demba Ba 0001, Dinei A. F. Florêncio, Cha Zhang |
ICME | 2 |
| 2007 | Multi-Party Audio Conferencing Based on a Simpler MCU and Client-Side ECHO CancellationabstractTraditional multiparty audio conferencing uses a star-shaped topology where all the clients connect to a central MCU (multipoint control unit). The MCU mixes the signals from the speakers, encodes it, and sends back the encoded signal to each client. To prevent the speakers from hearing their own voices, the MCU has to produce and encode a different mixed signal for each speaker. As a result, the CPU load on the MCU increases proportionally to the number of speakers in the conference. In this paper, we introduce a new conferencing architecture, where the MCU produces a single encoded signal sum of all received signals and each client is responsible for removing its own signal if necessary. This architecture can substantially reduce CPU load on the MCU. The major challenge, however, is that the client's original speech is non-linearly distorted by the MCU encoding process. Simply subtracting the original speech from the mixed signal would produce an echo-like distortion. We solve that problem using a novel algorithm which completely removes the echo with minimal artifacts. Mean opinion score (MOS) results imply that the proposed algorithm works well, making the proposed multiparty audio conferencing architecture promising. Li-wei He, Dinei A. F. Florêncio |
ICME | 3 |
| 2007 | Do Strong Web Passwords Accomplish Anything?
Dinei A. F. Florêncio, Cormac Herley, Baris Coskun |
HotSec | 1 |
| 2007 | A large-scale study of web password habitsabstractWe report the results of a large scale study of password use andpassword re-use habits. The study involved half a million users over athree month period. A client component on users' machines recorded a variety of password strength, usage and frequency metrics. This allows us to measure or estimate such quantities as the average number of passwords and average number of accounts each user has, how many passwords she types per day, how often passwords are shared among sites, and how often they are forgotten. We get extremely detailed data on password strength, the types and lengths of passwords chosen, and how they vary by site. The data is the first large scale study of its kind, and yields numerous other insights into the role the passwords play in users' online experience. Dinei A. F. Florêncio, Cormac Herley |
WWW | 1 |
| 2006 | KLASSP: Entering Passwords on a Spyware Infected Machine Using a Shared-Secret ProxyabstractIn this paper we examine the problem of entering sensitive data, such as passwords, from an untrusted machine. By untrusted we mean that it is suspected to be infected with spyware which snoops on the user's activity. Using such a machine is obviously undesirable, and yet roaming users often have no choice. They are in no position to judge the security status of Internet cafe, airport lounge or business center machines. Either malice or negligence on the part of an administrator means that any such machine can easily be running a keylogger. The roaming user has no reliable way of determining whether it is safe, and has no alternative to typing the password. We consider whether it is possible to enter data to confound spyware assumed to be running on the machine in question. The difficulty of mounting a collusion attack on a single user's password makes the problem more tractable than it might appear. We explore several approaches. In the first, we show how the user can embed a password in random keystrokes to confuse spyware, while leaving the actual login unaffected. In the second we employ a proxy server to strip random keys. In the third we again employ a proxy that inverts a key mapping performed by the user. We examine also several potential attacks Dinei A. F. Florêncio, Cormac Herley |
ACSAC | 1 |
| 2006 | Acoustic Echo Cancelation for High Noise EnvironmentsabstractAcoustic echo cancellation (AEC) is highly imperative for enhanced communication in noisy environments such as a car or a conference room. In this work, we present a dual-structured AEC architecture that improves both the convergence time and misadjustment of a conventional adaptive sub-band AEC algorithm in high noise environments. In this architecture, one part performs smooth adaptation while the other part performs fast adaptation; a convergence detector is implemented to facilitate switching between the fast and smooth adaptations. We propose the momentum normalized least mean square (MNLMS) algorithm for smooth adaptation and we implement the NLMS algorithm for fast adaptation. The current architecture provides up to 3-4 dB echo reduction improvement over a conventional adaptive subband AEC algorithm and it helps minimize near-end distortion and artifacts in the post-processed AEC output Amit Chhetri, Jack W. Stokes, Dinei A. F. Florêncio |
ICME | 3 |
| 2006 | PASS: Peer-Aware Silence Suppression for Internet Voice ConferencesabstractA novel tandem-free solution for multiparty VoIP conferences called PASS (peer-aware silence suppression) is presented. Similar to traditional tandem-free solutions, PASS introduces a limit on the number of concurrent speakers in a conference. But in contrast to traditional solutions, PASS silence suppression and speaker selection are completely distributed, running on each client. No speaker selection is performed at the bridge at all. This configuration leads to better scalability, lower bandwidth occupation and jitter buffer delay, and higher compatibility with a wide variety of network topologies. The key component of PASS, distributed silence suppression and speaker selection, is realized through a robust approach proposed in this paper. Based on a voice activity measure derived using machine learning techniques, this approach is able to reliably suppress silence in complex environments, and perform accurate and transparent speaker selection as well Xun Xu 0003, Li-wei He, Dinei A. F. Florêncio, Yong Rui |
ICME | 3 |
| 2006 | Is IEEE 802.11 ready for VoIP?abstractIn this paper, we empirically explore voice communication over IEEE 802.11 networks (VoWiFi). The objective is to understand the limitations of the current WiFi network for VoWiFi deployment. Our experiment finds two major problems of VoWiFi: unstable and excessively long handoffs and unpredictable occurrence of bursts. We also discuss several other minor factors that could hinder VoWiFi deployment, such as network capacity, fairness, and interference susceptibility. Finally, we describe the scenarios where VoWiFi could be used. We conclude that VoWiFi is feasible if used moderately, with low mobility and good signal strength. Arlindo Flávio da Conceição, Dinei A. F. Florêncio, Fabio Kon |
MMSP | 3 |
| 2006 | Analysis and Improvement of Anti-Phishing Schemes
Dinei A. F. Florêncio, Cormac Herley |
SEC | 1 |
| 2006 | Password Rescue: A New Approach to Phishing Prevention
Dinei A. F. Florêncio, Cormac Herley |
HotSec | 1 |
| 2005 | Sound source localization for circular arrays of directional microphonesabstractPrevious research in sound source localization has helped increase the robustness of estimates to noise and reverberation. Circular arrays are of particular interest for a number of scenarios, particularly because they can be placed in the center of the sources. First, that improves the sound capture due to the reduced distance. Second, it helps on the direction estimation, not only because of the reduced distance, but also because it increases the angle differences. Nevertheless, most research on circular arrays focused on the case of omni-directional microphones. In this paper, we present a new algorithm for sound source localization developed specifically for directional microphones. Results obtained from real meeting room setups show a typical error of less than 3 degrees. Yong Rui, Dinei A. F. Florêncio, Warren Lam, Jinyan Su |
ICASSP (3) | 2 |
| 2005 | Image de-noising by selective filtering based on double-shot picturesabstractThe quality of photographs is often reduced by sensor noise. This was a problem with film cameras, and is still a problem with current CCD and CMOS sensors, particularly in low-lighting conditions. De-noising techniques do not always perform satisfactorily. Typical de-noise techniques reduce the sharpness of the image. In this paper we propose a new de-noising technique, which is based on a dual-shot technique. The proposed algorithm is based on selectively removing high frequencies that do not correlate well between the frames. The algorithm borrows from video processing noise-removal techniques, but the final picture is derived from filtering a single shot, avoiding double-contouring and other artifacts that may happen with video techniques. While the decision is made based on both frames, the filtering itself is done using exclusively one of the frames. For this reason, the second (auxiliary) shot may be of much lower quality. Dinei A. F. Florêncio |
ICIP (3) | 1 |
| 2004 | Time delay estimation in the presence of correlated noise and reverberationabstractWe propose a new two-stage framework for time delay estimation in the presence of correlated noise and reverberation. The new framework allows us to develop a set of new approaches as well as to unify existing ones. We further develop the maximum likelihood estimation when reverberation is present. The corresponding weighting function is a more accurate form of the weighting function proposed by H. Wang and P. Chu (Proc. ICASSP, 1997), one of the best existing techniques. We compare our new algorithms with the existing ones and report superior performance. Yong Rui, Dinei A. F. Florêncio |
ICASSP (2) | 2 |
| 2003 | Can the sample being transmitted be used to refine its own PDF estimate?abstractMany image coders map the input image to a transform domain, and encode coefficients in that domain. Usually, encoders use previously transmitted samples to help estimate the probabilities for the next sample to be encoded. This backward probability estimation provides a better PDF estimation, without need to send any side information. A new method of encoding was proposed that goes one step further: besides past samples, it also uses information about the current sample in computing the PDF for the current sample. Yet, no side information was transmitted. The initial PDF estimate is based on a Tarp filter, but probabilities are then progressively refined for non-zero samples. Results are superior to JPEG2000, and to bit-plane Tarp. Dinei A. F. Florêncio, Patrice Y. Simard |
DCC | 1 |
| 2003 | New direct approaches to robust sound source localizationabstractWhen more than two microphones are used, the traditional time-delay-of-arrival (TDOA) based sound source localization (SSL) approach involves two steps. The first step computes TDOA for each microphone pair, and the second step combines these estimates. This two-step process discards relevant information in the first step, thus degrading the SSL accuracy and robustness. Although less used, one-step processes do exist. In this paper, we review these processes, create a unified framework, and introduce two new one-step algorithms. We compare our proposed approaches against existing 1and 2-step approaches and demonstrate significantly better SSL performance. Yong Rui, Dinei A. F. Florêncio |
ICME | 2 |
| 2002 | An improved spread spectrum technique for robust watermarkingabstractThis paper introduces a new watermarking modulation technique, which we call improved spread spectrum (ISS). Unlike in traditional spread spectrum (SS), in ISS the carrier signal does not act as a noise source; leading significant performance gains. In typical examples, the gain of ISS over SS is 20 dB or more in signal-to-noise ratio or 20 orders of magnitude or more in error probability. The proposed method achieves roughly the same noise robustness gain as quantization index modulation (QIM). Nevertheless, while QIM is quite sensitive to amplitude variations, ISS is as robust in practice as traditional SS. Henrique S. Malvar, Dinei A. F. Florêncio |
ICASSP | 2 |
| 2001 | Multichannel filtering for optimum noise reduction in microphone arraysabstractIntroduces an optimization criterion for the design of microphone arrays, and derives an optimum filter based on this criterion. The algorithm computes two separate correlation matrices for the signal: one for when only background noise is present, and one for when both noise and signal are present. A filter is then computed based on these matrices, optimizing the proposed weighted mean-square error criterion. A block-recursive version of the algorithm is presented, using LMS-like adaptation of the multichannel filters, with a computational complexity under 40 MIPS for a typical application with four microphones. Simulation results with typical office noise show improvements of up to 20 dB in signal-to-noise ratio, even in low-noise environments. Dinei A. F. Florêncio, Henrique S. Malvar |
ICASSP | 1 |
| 2001 | Speech dereverberation via maximum-kurtosis subband adaptive filteringabstractThis paper presents an efficient algorithm for high-quality speech capture in applications such as hands-free teleconferencing or voice recording by personal computers. We process the microphone signals by a subband adaptive filtering structure using a modulated complex lapped transform (MCLT), in which the subband filters are adapted to maximize the kurtosis of the linear prediction (LP) residual of the reconstructed speech. In this way, we attain good solutions to the problem of blind speech dereverberation. Experimental results with actual data, as well as with artificially difficult reverberant situations, show very good performance, both in terms of a significant reduction of the perceived reverberation, as well as improvement in spectral fidelity. Bradford W. Gillespie, Henrique S. Malvar, Dinei A. F. Florêncio |
ICASSP | 3 |
| 2001 | Motion sensitive pre-processing for videoabstractDue to the capture process, video signals are generally contaminated by noise. Furthermore, overall video quality can be improved by reducing sharpness in fast moving areas, except in cases where the eye is able to track the motion. We propose a relatively simple pre-processing method that is able to address both these situations. The proposed. pre-processing algorithm is based on selectively removing high frequencies that are not well predicted by the motion compensation. While the decision is made based on temporal tracking, the filtering is done exclusively in the spatial domain, thus avoiding the artifacts produced by other pre-processing methods. Dinei A. F. Florêncio |
ICIP (2) | 1 |
| 1998 | Nonexpansive pyramid for image coding using a nonlinear filterbankabstractA nonexpansive pyramidal decomposition is proposed for low-complexity image coding. The image is decomposed through a nonlinear filterbank into low- and highpass signals and the recursion of the filterbank over the lowpass signal generates a pyramid resembling that of the octave wavelet transform. The structure itself guarantees perfect reconstruction and we have chosen nonlinear filters for performance reasons. The transformed samples are grouped into square blocks and used to replace the discrete cosine transform (DCT) in the Joint Photographic Expert Group (JPEG) coder. The proposed coder has some advantages over the DCT-based JPEG: computation is greatly reduced, image edges are better encoded, blocking is eliminated, and it allows lossless coding. Ricardo L. de Queiroz, Dinei A. F. Florêncio, Ronald W. Schafer |
IEEE Trans. Image Process. | 2 |
| 1996 | The motion transform: a new motion compensation techniqueabstractMotion estimation plays an important role in video coding schemes by removing temporal redundancies that exist in image sequences. Motion information that is required in motion compensated prediction is transmitted as side information, as in forward motion estimation, or can be computed at the receiver in backward motion estimation. The former method has the advantage of more accurate prediction but, suffers in that the motion information is transmitted as side information. We introduce the motion transform (MT), a new motion estimation and compensation technique that does not require the transmission of motion vectors yet, performs well under a variety of conditions. In the proposed method, motion information is obtained in a bottom-up hierarchical fashion using a two-dimensional wavelet-like decomposition. By using a nonlinear non-expansive filter bank a hierarchical decomposition of the image allows coarse-to-fine motion compensation at the decoder without motion vector transmission. Robert M. Armitano, Dinei A. F. Florêncio, Ronald W. Schafer |
ICASSP | 2 |
| 1996 | Perfect reconstructing nonlinear filter banksabstractPerfect reconstruction (PR) filter banks have found numerous applications, and have received much attention in the literature. For linear filter banks, necessary and sufficient conditions for PR have been established for most practical situations. Recently, nonlinear filter banks have been proposed for image coding applications. These filters are generally simple, and produce better results than linear filters of same complexity. Nevertheless, the lack of general PR conditions limits these filters to cases where one of the filters is the identity. In this paper, we present a framework that allows, for the first time, the design of PR nonlinear filter banks including (non-trivial) filters on all channels. Although the framework does not include all nonlinear PR filter banks, it does include all previously published nonlinear filter banks, as well as all linear ones. This framework suggest new possibilities for the design of nonlinear PR filter banks. Dinei A. F. Florêncio, Ronald W. Schafer |
ICASSP | 1 |
| 1996 | A pyramidal coder using a nonlinear filter bankabstractWe propose a novel pyramidal coder, which resembles the general structure of a JPEG coder, but which uses a nonlinear transform to replace the DCT. The nonlinear transform is obtained by the hierarchical application of a median filter predictor at subsampled versions of the original signal. The transformed samples are grouped into square blocks and used to replace the DCT in the JPEG baseline coder. The proposed coder shows several advantages: computation is greatly reduced compared to the DCT, image edges are better encoded, blocking is eliminated, and it allows lossless coding. Objective comparisons show the superiority of the proposed coder against both baseline and lossless JPEG. Ricardo L. de Queiroz, Dinei A. F. Florêncio |
ICASSP | 2 |
| 1996 | Motion transforms for video codingabstractVideo coding standards use motion compensated prediction to reduce temporal redundancies. In general this is performed by computing motion vectors for predefined regions in the image (rectangular blocks), and transmitting these vectors as side-information. The motion transform (MT) provides a new motion compensation technique that does not require the transmission of motion vectors and yet, performs as well as forward motion estimation for many image sequences. In the MT, motion information is obtained in a bottom-up hierarchical fashion using a two-dimensional nonlinear filter bank. In this paper we elaborate on MT implementation details and discuss a new variation of the MT. Results obtained using MTs on typical image sequences are compared to results for the full search block matching algorithm (BMA). Dinei A. F. Florêncio, Robert M. Armitano, Ronald W. Schafer |
ICIP (2) | 1 |
| 1995 | Post-sampling aliasing control for natural imagesabstractSampling and reconstruction are usually analyzed under the framework of linear signal processing. Powerful tools like the Fourier transform and optimum linear filter design techniques, allow for a very precise analysis of the process. In particular, an optimum linear filter of any length can be derived under most situations. Many of these tools are not available for non-linear systems, and it is usually difficult to find an optimum non-linear system under any criteria. The authors analyze the possibility of using non-linear filtering in the interpolation of subsampled images. They show that a very simple (5/spl times/5) non-linear reconstruction filter outperforms (for the images analyzed) linear filters of up to 256/spl times/256, including optimum (separable) Wiener filters of any size. Dinei A. F. Florêncio, Ronald W. Schafer |
ICASSP | 1 |
| 1994 | A Non-Expansive Pyramidal Morphological Image CoderabstractPyramid image coding is a natural coding scheme for applications where progressive transmission is desired. In this kind of coder, versions of the original image at several resolution levels are formed by successive filtering and subsampling. Then, beginning from the coarsest image, the image is used to produce an estimate for the next (higher resolution) level and the error is coded and transmitted. While expansive pyramids (e.g., Burt's Laplacian pyramid) are usually less efficient, non-expansive pyramids tend to produce ringing (e.g., subband/wavelet) and/or blocking (e.g., DCT). We introduce a non-expansive pyramid that does not, produce ringing or blocking effects. Instead, the main artifact is texture removal. The simulations have produced images with entropies in the range of 0.2 to 1.5 bpp, with SNR figures similar to or better than JPEG at equivalent rates. The proposed coder has several attractive features, including 8 bit integer operations only, a perfect reconstruction mode, progressive transmission and an interesting progressive computation property.> Dinei A. F. Florêncio, Ronald W. Schafer |
ICIP (2) | 1 |
| 1991 | On the use of asymmetric windows for reducing the time delay in real-time spectral analysisabstractThe idea of asymmetric windowing as a way of reducing time delay in spectral analysis is explored. Traditional windows are analyzed in relation to time delay. Some asymmetric windows are introduced and a comparative analysis with traditional windows is made. It is shown that it is possible to reduce the time delay up to about 25% of the analysis window size. Application to linear predictive speech coding is analyzed, and the experimental results of using an asymmetric window on a standard LPC vocoder are presented.> Dinei A. F. Florêncio |
ICASSP | 1 |