EDBT 2026 Demo / reviewers in the wild / expert
Cha Zhang
dblp:74/6770
· DBLP profile ↗
97ranked-venue papers
39as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 82 · 38 first-author · 4 since 2021Artificial intelligence and machine learning · 22 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 3Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
19 papers |
Vision and language · 34% Information extraction and text analysis · 14% Efficient and distributed learning · 11% | |
| Computer graphics and multimedia
18 papers |
Audio and music processing · 32% Image and video processing · 28% Image and video coding · 10% |
Topics — the 30 heaviest of 87, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.5 | 2 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 Unifying Vision, Text, and Layout for Universal Document Processing · CVPR 2023 |
Natural language and speech › Information extraction and text analysis
document understanding |
1.2 | 2 | 2023 | Unifying Vision, Text, and Layout for Universal Document Processing · CVPR 2023 LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 |
Computer vision › Vision and language › visual reasoning
visual chain-of-thought |
0.9 | 1 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 |
Computer vision › Vision and language
visual reasoning |
0.9 | 1 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 2 | 2020 | Towards Efficient Model Compression via Learned Global Ranking · CVPR 2020 RePr: Improved Training of Convolutional Filters · CVPR 2019 |
Natural language and speech › Information extraction and text analysis
document AI |
0.7 | 1 | 2023 | Unifying Vision, Text, and Layout for Universal Document Processing · CVPR 2023 |
Computer vision › Vision and language
open-vocabulary models |
0.7 | 1 | 2023 | From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language Understanding · ACL (1) 2023 |
Computer vision › Image recognition and object detection › text recognition
optical character recognition |
0.7 | 1 | 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained Models · AAAI 2023 |
Computer vision › Image recognition and object detection
text recognition |
0.7 | 1 | 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained Models · AAAI 2023 |
Natural language and speech › Language models and text generation › text generation › neural text generation
transformer-based text generation |
0.7 | 1 | 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained Models · AAAI 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling |
0.6 | 1 | 2022 | DiT: Self-supervised Pre-training for Document Image Transformer · ACM Multimedia 2022 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.6 | 1 | 2022 | DiT: Self-supervised Pre-training for Document Image Transformer · ACM Multimedia 2022 |
Computer vision › Vision and language › visual question answering
text-based visual question answering |
0.5 | 1 | 2021 | TAP: Text-Aware Pre-Training for Text-VQA and Text-Caption · CVPR 2021 |
Computer vision › Vision and language
vision-language pretraining |
0.5 | 1 | 2021 | TAP: Text-Aware Pre-Training for Text-VQA and Text-Caption · CVPR 2021 |
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
channel pruning |
0.4 | 1 | 2020 | Towards Efficient Model Compression via Learned Global Ranking · CVPR 2020 |
Image and video processing
image enhancement |
0.4 | 2 | 2016 | Image Bit-Depth Enhancement via Maximum A Posteriori Estimation of AC Signal · IEEE Trans. Image Process. 2016 Video Enhancement of People Wearing Polarized Glasses: Darkening Reversal and Reflection Reduction · CVPR 2013 |
Machine learning › Deep learning architectures and training › convolutional neural network
convolutional neural network training |
0.4 | 1 | 2019 | RePr: Improved Training of Convolutional Filters · CVPR 2019 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.4 | 1 | 2019 | RePr: Improved Training of Convolutional Filters · CVPR 2019 |
Computer vision › Video understanding and tracking
video analytics |
0.3 | 1 | 2017 | Deep Learning for Intelligent Video Analysis · ACM Multimedia 2017 |
Computer vision › Video understanding and tracking
video representation learning |
0.3 | 1 | 2017 | Deep Learning for Intelligent Video Analysis · ACM Multimedia 2017 |
Audio and music processing
sound source localization |
0.3 | 3 | 2010 | Using Reverberation to Improve Range and Elevation Discrimination for Small Array Sound Source Localization · IEEE Trans. Speech Audio Process. 2010 Boosting-Based Multimodal Speaker Detection for Distributed Meeting Videos · IEEE Trans. Multim. 2008 Maximum Likelihood Sound Source Localization and Beamforming for Directional Microphone Arrays in Distributed Meetings · IEEE Trans. Multim. 2008 |
Audio and music processing
room acoustics |
0.3 | 2 | 2012 | Geometrically Constrained Room Modeling With Compact Microphone Arrays · IEEE Trans. Speech Audio Process. 2012 Using Reverberation to Improve Range and Elevation Discrimination for Small Array Sound Source Localization · IEEE Trans. Speech Audio Process. 2010 |
Image and video processing › image enhancement
bit-depth enhancement |
0.2 | 1 | 2016 | Image Bit-Depth Enhancement via Maximum A Posteriori Estimation of AC Signal · IEEE Trans. Image Process. 2016 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.2 | 1 | 2023 | From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language Understanding · ACL (1) 2023 |
Computer vision › 3D vision
3d shape reconstruction |
0.2 | 1 | 2014 | Rate-Constrained 3D Surface Estimation From Noise-Corrupted Multiview Depth Videos · IEEE Trans. Image Process. 2014 |
Image and video coding › video compression › 3d video coding
depth map coding |
0.2 | 1 | 2014 | Rate-Constrained 3D Surface Estimation From Noise-Corrupted Multiview Depth Videos · IEEE Trans. Image Process. 2014 |
Virtual and augmented reality › telepresence
immersive communication |
0.2 | 1 | 2014 | Immersive 3D Communication · ACM Multimedia 2014 |
Natural language and speech › Information extraction and text analysis › document understanding
document layout analysis |
0.2 | 1 | 2022 | DiT: Self-supervised Pre-training for Document Image Transformer · ACM Multimedia 2022 |
Natural language and speech › Information extraction and text analysis › document understanding › table recognition
table detection |
0.2 | 1 | 2022 | DiT: Self-supervised Pre-training for Document Image Transformer · ACM Multimedia 2022 |
Methods — techniques the papers use, named apart from their topics
self-supervised pretraining · 1.2visual editing · 0.9code generation · 0.9transformer · 0.7prompt-based sequence generation · 0.7pre-trained text transformer · 0.7pre-trained image transformer · 0.7masked image reconstruction · 0.7hierarchical modeling · 0.7vision transformer · 0.6graph-signal smoothness · 0.2convex programming · 0.2MMSE estimation · 0.2rate-constrained maximum a posteriori estimation · 0.2depth-image-based rendering · 0.23d display · 0.23d capture · 0.2boosting · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image UnderstandingabstractStructured image understanding, such as interpreting tables and charts, requires strategically refocusing across various structures and texts within an image, forming a reasoning sequence to arrive at the final answer. However, current multimodal large language models (LLMs) lack this multihop selective attention capability. In this work, we introduce ReFocus, a simple yet effective framework that equips multimodal LLMs with the ability to generate ``visual thoughts'' by performing visual editing on the input image through code, shifting and refining their visual focuses. Specifically, ReFocus enables multimodal LLMs to generate Python codes to call tools and modify the input image, sequentially drawing boxes, highlighting sections, and masking out areas, thereby enhancing the visual reasoning process. We experiment upon a wide range of structured image understanding tasks involving tables and charts. ReFocus largely improves performance on all tasks over GPT-4o without visual editing, yielding an average gain of 11.0% on table tasks and 6.8% on chart tasks. We present an in-depth analysis of the effects of different visual edits, and reasons why ReFocus can improve the performance without introducing additional information. Further, we collect a 14k training set using ReFocus, and prove that such visual chain-of-thought with intermediate information offers a better supervision than standard VQA data, reaching a 8.0% average gain over the same model trained with QA pairs and 2.6% over CoT. Minqian Liu, Zhengyuan Yang, John Corring, Yijuan Lu, Dan Roth 0001, Dinei A. F. Florêncio, Cha Zhang |
ICML | 9 |
| 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsabstractText recognition is a long-standing research problem for document digitalization. Existing approaches are usually built based on CNN for image understanding and RNN for char-level text generation. In addition, another language model is usually needed to improve the overall accuracy as a post-processing step. In this paper, we propose an end-to-end text recognition approach with pre-trained image Transformer and text Transformer models, namely TrOCR, which leverages the Transformer architecture for both image understanding and wordpiece-level text generation. The TrOCR model is simple but effective, and can be pre-trained with large-scale synthetic data and fine-tuned with human-labeled datasets. Experiments show that the TrOCR model outperforms the current state-of-the-art models on the printed, handwritten and scene text recognition tasks. The TrOCR models and code are publicly available at https://aka.ms/trocr. Minghao Li 0004, Tengchao Lv, Jingye Chen, Lei Cui 0001, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Zhoujun Li 0001, Furu Wei |
AAAI | 7 |
| 2023 | From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language UnderstandingabstractCurrent state-of-the-art models for natural language understanding require a preprocessing step to convert raw text into discrete tokens.This process known as tokenization relies on a pre-built vocabulary of words or sub-word morphemes.This fixed vocabulary limits the model's robustness to spelling errors and its capacity to adapt to new domains.In this work, we introduce a novel open-vocabulary language model that adopts a hierarchical two-level approach: one at the word level and another at the sequence level.Concretely, we design an intraword module that uses a shallow Transformer architecture to learn word representations from their characters, and a deep inter-word Transformer module that contextualizes each word representation by attending to the entire word sequence.Our model thus directly operates on character sequences with explicit awareness of word boundaries, but without biased sub-word or word-level vocabulary.Experiments on various downstream tasks show that our method outperforms strong baselines.We also demonstrate that our hierarchical model is robust to textual corruption and domain shift. Li Sun 0010, Florian Luisier, Kayhan Batmanghelich, Dinei A. F. Florêncio, Cha Zhang |
ACL (1) | 5 |
| 2023 | Unifying Vision, Text, and Layout for Universal Document ProcessingabstractWe propose Universal Document Processing (UDOP), a foundation Document AI model which unifies text, image, and layout modalities together with varied task formats, including document understanding and generation. UDOP leverages the spatial correlation between textual content and document image to model image, text, and layout modalities with one uniform representation. With a novel Vision-Text-Layout Transformer, UDOP unifies pretraining and multi-domain downstream tasks into a prompt-based sequence generation scheme. UDOP is pretrained on both large-scale unlabeled document corpora using innovative self-supervised objectives and diverse labeled data. UDOP also learns to generate document images from text and layout modalities via masked image reconstruction. To the best of our knowledge, this is the first time in the field of document AI that one model simultaneously achieves high-quality neural document editing and content customization. Our method sets the state-of-the-art on 8 Document AI tasks, e.g., document understanding and QA, across diverse data domains like finance reports, academic papers, and web-sites. UDOP ranks first on the leaderboard of the Document Understanding Benchmark.11Code and models: https://github.com/microsoft/i-Code/tree/main/i-Code-Doc Zineng Tang, Ziyi Yang 0011, Yuwei Fang, Yang Liu 0124, Chenguang Zhu 0001, Michael Zeng 0001, Cha Zhang, Mohit Bansal |
CVPR | 8 |
| 2023 | Diffusion-Based Document Layout Generation
Yijuan Lu, John Corring, Dinei A. F. Florêncio, Cha Zhang |
ICDAR (1) | 5 |
| 2022 | DiT: Self-supervised Pre-training for Document Image TransformerabstractImage Transformer has recently achieved significant progress for natural image understanding, either using supervised (ViT, DeiT, etc.) or self-supervised (BEiT, MAE, etc.) pre-training techniques. In this paper, we propose DiT, a self-supervised pre-trained Document Image Transformer model using large-scale unlabeled text images for Document AI tasks, which is essential since no supervised counterparts ever exist due to the lack of human-labeled document images. We leverage DiT as the backbone network in a variety of vision-based Document AI tasks, including document image classification, document layout analysis, table detection as well as text detection for OCR. Experiment results have illustrated that the self-supervised pre-trained DiT model achieves new state-of-the-art results on these downstream tasks, e.g. document image classification (91.11 - 92.69), document layout analysis (91.0 - 94.9), table detection (94.23 - 96.55) and text detection for OCR (93.07 - 94.29). The code and pre-trained models are publicly available at https://aka.ms/msdit. Yiheng Xu, Tengchao Lv, Lei Cui 0001, Cha Zhang, Furu Wei |
ACM Multimedia | 5 |
| 2022 | Editorial for Special Issue on Computer Vision in the Wild
Cha Zhang, Katsushi Ikeuchi |
Int. J. Comput. Vis. | 2 |
| 2021 | LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingabstractYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, Lidong Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yang Xu 0049, Yiheng Xu, Tengchao Lv, Lei Cui 0001, Furu Wei, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Wanxiang Che, Min Zhang 0005, Lidong Zhou |
ACL/IJCNLP (1) | 9 |
| 2021 | TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionabstractIn this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answering and image caption generation, respectively. In contrast to conventional vision-language pretraining that fails to capture scene text and its relationship with the visual and text modalities, TAP explicitly incorporates scene text (generated from OCR engines) during pretraining. With three pre-training tasks, including masked language modeling (MLM), image-text (contrastive) matching (ITM), and relative (spatial) position prediction (RPP), pre-training with scene text effectively helps the model learn a better aligned representation among the three modalities: text word, visual object, and scene text. Due to this aligned representation learning, even pre-trained on the same downstream task dataset, TAP already boosts the absolute accuracy on the TextVQA dataset by +5:4%, compared with a non-TAP baseline. To further improve the performance, we build a large-scale scene text-related imagetext dataset based on the Conceptual Caption dataset, named OCR-CC, which contains 1:4 million images with scene text. Pre-trained on this OCR-CC dataset, our approach outperforms the state of the art by large margins on multiple tasks, i.e., +8:3% accuracy on TextVQA, +8:6% accuracy on ST-VQA, and +10:2 CIDEr score on TextCaps. Zhengyuan Yang, Yijuan Lu, Xi Yin 0006, Dinei A. F. Florêncio, Cha Zhang, Lei Zhang 0001, Jiebo Luo 0001 |
CVPR | 7 |
| 2020 | Towards Efficient Model Compression via Learned Global RankingabstractPruning convolutional filters has demonstrated its effectiveness in compressing ConvNets. Prior art in filter pruning requires users to specify a target model complexity (e.g., model size or FLOP count) for the resulting architecture. However, determining a target model complexity can be difficult for optimizing various embodied AI applications such as autonomous robots, drones, and user-facing applications. First, both the accuracy and the speed of ConvNets can affect the performance of the application. Second, the performance of the application can be hard to assess without evaluating ConvNets during inference. As a consequence, finding a sweet-spot between the accuracy and speed via filter pruning, which needs to be done in a trial-and-error fashion, can be time-consuming. This work takes a first step toward making this process more efficient by altering the goal of model compression to producing a set of ConvNets with various accuracy and latency trade-offs instead of producing one ConvNet targeting some pre-defined latency constraint. To this end, we propose to learn a global ranking of the filters across different layers of the ConvNet, which is used to obtain a set of ConvNet architectures that have different accuracy/latency trade-offs by pruning the bottom-ranked filters. Our proposed algorithm, LeGR, is shown to be 2× to 3× faster than prior work while having comparable or better performance when targeting seven pruned ResNet-56 with different accuracy/FLOPs profiles on the CIFAR-100 dataset. Additionally, we have evaluated LeGR on ImageNet and Bird-200 with ResNet-50 and Mo- bileNetV2 to demonstrate its effectiveness. Code available at https://github.com/cmu-enyac/LeGR. Ting-Wu Chin, Ruizhou Ding, Cha Zhang, Diana Marculescu |
CVPR | 3 |
| 2020 | Multimodal Active Speaker Detection and Virtual Cinematography for Video ConferencingabstractActive speaker detection (ASD) and virtual cinematography (VC) can significantly improve the experience of a video conference by automatically panning, tilting and zooming of a camera: subjectively users rate an expert video cinematographer significantly higher than the unedited video. We describe a new automated ASD and VC that performs within 0.3 MOS of an expert cinematographer based on subjective ratings with a 1-5 scale. This system uses a 4K wide-FOV camera, a depth camera, and a microphone array, extracts features from each modality and trains an ASD using an AdaBoost machine learning system that is very efficient and runs in real-time. A VC is similarly trained using machine learning. To avoid distracting the room participants the system has no moving parts - the VC works by cropping and zooming the 4K wide-FOV video stream. The system was tuned and evaluated using extensive crowdsourcing techniques and evaluated on a system with N=100 meetings, each 25 minutes in length. Ross Cutler, Ramin Mehran, Sam Johnson, Cha Zhang, Adam Kirk, Oliver Whyte, Adarsh Kowdle |
ICASSP | 4 |
| 2019 | RePr: Improved Training of Convolutional FiltersabstractA well-trained Convolutional Neural Network can easily be pruned without significant loss of performance. This is because of unnecessary overlap in the features captured by the network's filters. Innovations in network architecture such as skip/dense connections and inception units have mitigated this problem to some extent, but these improvements come with increased computation and memory requirements at run-time. We attempt to address this problem from another angle - not by changing the network structure but by altering the training method. We show that by temporarily pruning and then restoring a subset of the model's filters, and repeating this process cyclically, overlap in the learned features is reduced, producing improved generalization. We show that the existing model-pruning criteria are not optimal for selecting filters to prune in this context, and introduce inter-filter orthogonality as the ranking criteria to determine under-expressive filters. Our method is applicable both to vanilla convolutional networks and more complex modern architectures, and improves the performance across a variety of tasks, especially when applied to smaller networks. Aaditya Prakash, James A. Storer, Dinei A. F. Florêncio, Cha Zhang |
CVPR | 4 |
| 2017 | Automatic speech emotion recognition using recurrent neural networks with local attentionabstractAutomatic emotion recognition from speech is a challenging task which relies heavily on the effectiveness of the speech features used for classification. In this work, we study the use of deep learning to automatically discover emotionally relevant features from speech. It is shown that using a deep recurrent neural network, we can learn both the short-time frame-level acoustic features that are emotionally relevant, as well as an appropriate temporal aggregation of those features into a compact utterance-level representation. Moreover, we propose a novel strategy for feature pooling over time which uses local attention in order to focus on specific regions of a speech signal that are more emotionally salient. The proposed solution is evaluated on the IEMOCAP corpus, and is shown to provide more accurate predictions compared to existing emotion recognition algorithms. Seyedmahdad Mirsamadi, Emad Barsoum, Cha Zhang |
ICASSP | 3 |
| 2017 | Deep Learning for Intelligent Video AnalysisabstractAnalyzing videos is one of the fundamental problems of computer vision and multimedia content analysis for decades. The task is very challenging as video is an information-intensive media with large variations and complexities. Thanks to the recent development of deep learning techniques, researchers in both computer vision and multimedia communities are now able to boost the performance of video analysis significantly and initiate new research directions to analyze video content. This tutorial will present recent advances under the umbrella of video understanding, which start from a unified deep learning toolkit--Microsoft Cognitive Toolkit (CNTK) that supports popular model types such as convolutional nets and recurrent networks, to fundamental challenges of video representation learning and video classification, recognition, and finally to an emerging area of video and language. Tao Mei 0001, Cha Zhang |
ACM Multimedia | 2 |
| 2016 | Emotion recognition in the wild from videos using imagesabstractThis paper presents the implementation details of the proposed solution to the Emotion Recognition in the Wild 2016 Challenge, in the category of video-based emotion recognition. The proposed approach takes the video stream from the audio-video trimmed clips provided by the challenge as input and produces the emotion label corresponding to this video sequence. This output is encoded as one out of seven classes: the six basic emotions (Anger, Disgust, Fear, Happiness, Sad, Surprise) and Neutral. Overall, the system consists of several pipelined modules: face detection, image pre-processing, deep feature extraction, feature encoding and, finally, an SVM classification. Sarah Adel Bargal, Emad Barsoum, Cristian Canton, Cha Zhang |
ICMI | 4 |
| 2016 | Training deep networks for facial expression recognition with crowd-sourced label distributionabstractCrowd sourcing has become a widely adopted scheme to collect ground truth labels. However, it is a well-known problem that these labels can be very noisy. In this paper, we demonstrate how to learn a deep convolutional neural network (DCNN) from noisy labels, using facial expression recognition as an example. More specifically, we have 10 taggers to label each input image, and compare four different approaches to utilizing the multiple labels: majority voting, multi-label learning, probabilistic label drawing, and cross-entropy loss. We show that the traditional majority voting scheme does not perform as well as the last two approaches that fully leverage the label distribution. An enhanced FER+ data set with multiple labels for each face image will also be shared with the research community. Emad Barsoum, Cha Zhang, Cristian Canton, Zhengyou Zhang |
ICMI | 2 |
| 2016 | Image Bit-Depth Enhancement via Maximum A Posteriori Estimation of AC SignalabstractWhen images at low bit-depth are rendered at high bit-depth displays, missing least significant bits needs to be estimated. We study the image bit-depth enhancement problem: estimating an original image from its quantized version from a minimum mean squared error (MMSE) perspective. We first argue that a graph-signal smoothness prior-one defined on a graph embedding the image structure-is an appropriate prior for the bit-depth enhancement problem. We next show that directly solving for the MMSE solution is, in general, too computationally expensive to be practical. We then propose an efficient approximation strategy. In particular, we first estimate the ac component of the desired signal in a maximum a posteriori formulation, efficiently computed via convex programming. We then compute the dc component with an MMSE criterion in a closed form given the computed ac component. Experiments show that our proposed two-step approach has improved performance over the conventional bit-depth enhancement schemes in both objective and subjective comparisons. Pengfei Wan 0001, Gene Cheung, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Trans. Image Process. | 4 |
| 2015 | Image based Static Facial Expression Recognition with Multiple Deep Network LearningabstractWe report our image based static facial expression recognition method for the Emotion Recognition in the Wild Challenge (EmotiW) 2015. We focus on the sub-challenge of the SFEW 2.0 dataset, where one seeks to automatically classify a set of static images into 7 basic emotions. The proposed method contains a face detection module based on the ensemble of three state-of-the-art face detectors, followed by a classification module with the ensemble of multiple deep convolutional neural networks (CNN). Each CNN model is initialized randomly and pre-trained on a larger dataset provided by the Facial Expression Recognition (FER) Challenge 2013. The pre-trained models are then fine-tuned on the training set of SFEW 2.0. To combine multiple CNN models, we present two schemes for learning the ensemble weights of the network responses: by minimizing the log likelihood loss, and by minimizing the hinge loss. Our proposed method generates state-of-the-art result on the FER dataset. It also achieves 55.96% and 61.29% respectively on the validation and test set of SFEW 2.0, surpassing the challenge baseline of 35.96% and 39.13% with significant gains. Zhiding Yu, Cha Zhang |
ICMI | 2 |
| 2015 | A survey on face detection in the wild: Past, present and future
Stefanos Zafeiriou, Cha Zhang, Zhengyou Zhang |
Comput. Vis. Image Underst. | 2 |
| 2015 | Precision Enhancement of 3-D Surfaces from Compressed Multiview Depth MapsabstractTransmitting depth maps captured from multiple viewpoints of a 3-D scene enables a wide range of receiver-side 3-D applications, including virtual view synthesis via depth-image-based rendering (DIBR). Observing that compressed depth maps from different viewpoints constitute multiple descriptions (MD) of the same signal, we propose to reconstruct 3-D surfaces of the scene by considering multiple compressed depth maps jointly. Specifically, we propose an alternating projection algorithm, inspired by the theory of projection onto convex sets (POCS), which at convergence returns a 3-D surface that satisfies three sets of conditions: spatial smoothness prior, quantization bin constraints in the block transform domain, and inter-view consistency. We present a theoretical proof that shows convergence of our algorithm under benign conditions. Compared to existing multiview depth map denoising schemes and single image de-quantization schemes, our proposed solution achieves higher objective quality for both reconstructed depth maps and synthesized virtual views. Pengfei Wan 0001, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Signal Process. Lett. | 5 |
| 2014 | Image bit-depth enhancement via maximum-a-posteriori estimation of graph AC componentabstractWhile modern displays offer high dynamic range (HDR) with large bit-depth for each rendered pixel, the bulk of legacy image and video contents were captured using cameras with shallower bit-depth. In this paper, we study the bit-depth enhancement problem for images, so that a high bit-depth (HBD) image can be reconstructed from an input low bit-depth (LBD) image. The key idea is to apply appropriate smoothing given the constraints that reconstructed signal must lie within the per-pixel quantization bins. Specifically, we first define smoothness via a signal-dependent graph Laplacian, so that natural image gradients can nonetheless be interpreted as low frequencies. Given defined smoothness prior and observed LBD image, we then demonstrate that computing the most probable signal via maximum a posteriori (MAP) estimation can lead to large expected distortion. However, we argue that MAP can still be used to efficiently estimate the AC component of the desired HBD signal, which along with a distortion-minimizing DC component, can result in a good approximate solution that minimizes the expected distortion. Experimental results show that our proposed method outperforms existing bit-depth enhancement methods in terms of reconstruction error. Pengfei Wan 0001, Gene Cheung, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
ICIP | 4 |
| 2014 | Point cloud attribute compression with graph transformabstractCompressing attributes on 3D point clouds such as colors or normal directions has been a challenging problem, since these attribute signals are unstructured. In this paper, we propose to compress such attributes with graph transform. We construct graphs on small neighborhoods of the point cloud by connecting nearby points, and treat the attributes as signals over the graph. The graph transform, which is equivalent to Karhunen-Loève Transform on such graphs, is then adopted to decorrelate the signal. Experimental results on a number of point clouds representing human upper bodies demonstrate that our method is much more efficient than traditional schemes such as octree-based methods. Cha Zhang, Dinei A. F. Florêncio, Charles T. Loop |
ICIP | 1 |
| 2014 | Facial expression tracking from head-mounted, partially observing camerasabstractHead-mounted displays (HMDs) have gained more and more interest recently. They can enable people to communicate with each other from anywhere, at anytime. However, since most HMDs today are only equipped with cameras pointing outwards, the remote party would not be able to see the user wearing the HMD. In this paper, we present a system for facial expression tracking based on head-mounted, inward looking cameras, such that the user can be represented with animated avatars at the remote party. The main challenge is that the cameras can only observe partial faces since they are very close to the face. We experiment with multiple machine learning algorithms to estimate facial expression parameters based on training data collected with the assistance of a Kinect depth sensor. Our results show that we can reliably track people's facial expression even from very limited view angles of the cameras. Bernardino Romera-Paredes, Cha Zhang, Zhengyou Zhang |
ICME | 2 |
| 2014 | Video face beautificationabstractThis paper presents a novel system framework of face beautification. Unlike prior works that deal with single images, the proposed beautification framework is designed for an input video and it is able to improve both the appearance and the shape of a face. Our system adopts a state-of-the-art algorithm to synthesize and track 3D face models using blendshapes. The personalized 3D model can be edited to satisfy personal preference. This interactive process is needed only once per subject. Based on the tracking result and the modified face model, we present an algorithm to beautify the face video efficiently and consistently. Furthermore we develop a variant of content preserving warping to reduce warping distortions along the face boundary. Finally we adopt real time bilateral filtering to remove wrinkles, freckles, and unwanted blemishes. This framework is evaluated on a set of videos. The experiments demonstrate that our framework can generate consistent and pleasant results over video frames while the original expressions and features are persevered naturally. Xinyu Huang 0001, Jizhou Gao, Alade O. Tokuta, Cha Zhang, Ruigang Yang |
ICME | 5 |
| 2014 | Immersive 3D CommunicationabstractThe last few decades have witnessed tremendous advances in telecommunication, with the invention of technologies such as radio, telephone, voice-over-IP, and video conferencing. While all these communication tools are useful and valuable, the ultimate goal of telecommunication is to enable fully immersive remote interaction in a way that simulates or even surpasses the face-to-face experience. Immersive 3D communication technologies are developed aiming at that goal. The objective of this tutorial is to present an overview of the recent advances in immersive 3D communication. Topics include the basics of human 3D perception, new systems and algorithms in real-time 3D scene capture and reconstruction, 3D data compression and dissemination, 3D displays, etc. We intend to provide insights into the latest immersive 3D communication technologies, and highlight some open research challenges for the future. Wanmin Wu, Cha Zhang |
ACM Multimedia | 2 |
| 2014 | Improving multiview face detection with multi-task deep convolutional neural networksabstractMultiview face detection is a challenging problem due to dramatic appearance changes under various pose, illumination and expression conditions. In this paper, we present a multi-task deep learning scheme to enhance the detection performance. More specifically, we build a deep convolutional neural network that can simultaneously learn the face/nonface decision, the face pose estimation problem, and the facial landmark localization problem. We show that such a multi-task learning scheme can further improve the classifier's accuracy. On the challenging FDDB data set, our detector achieves over 3% improvement in detection rate at the same false positive rate compared with other state-of-the-art methods. Cha Zhang, Zhengyou Zhang |
WACV | 1 |
| 2014 | Iterative transductive learning for automatic image segmentation and matting with RGB-D data
Bei He, Guijin Wang, Cha Zhang |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | A robust optical/inertial data fusion system for motion tracking of the robot manipulatorabstractWe present an optical/inertial data fusion system for motion tracking of the robot manipulator, which is proved to be more robust and accurate than a normal optical tracking system (OTS). By data fusion with an inertial measurement unit (IMU), both robustness and accuracy of OTS are improved. The Kalman filter is used in data fusion. The error distribution of OTS provides an important reference on the estimation of measurement noise using the Kalman filter. With a proper setup of the system and an effective method of coordinate frame synchronization, the results of experiments show a significant improvement in terms of robustness and position accuracy. Jens Hofschulte, Wan-li Jiang, Cha Zhang |
J. Zhejiang Univ. Sci. C | 5 |
| 2014 | Rate-Constrained 3D Surface Estimation From Noise-Corrupted Multiview Depth VideosabstractTransmitting compactly represented geometry of a dynamic 3D scene from a sender can enable a multitude of imaging functionalities at a receiver, such as synthesis of virtual images at freely chosen viewpoints via depth-image-based rendering. While depth maps—projections of 3D geometry onto 2D image planes at chosen camera viewpoints-can nowadays be readily captured by inexpensive depth sensors, they are often corrupted by non-negligible acquisition noise. Given depth maps need to be denoised and compressed at the encoder for efficient network transmission to the decoder, in this paper, we consider the denoising and compression problems jointly, arguing that doing so will result in a better overall performance than the alternative of solving the two problems separately in two stages. Specifically, we formulate a rate-constrained estimation problem, where given a set of observed noise-corrupted depth maps, the most probable (maximum a posteriori (MAP)) 3D surface is sought within a search space of surfaces with representation size no larger than a prespecified rate constraint. Our rate-constrained MAP solution reduces to the conventional unconstrained MAP 3D surface reconstruction solution if the rate constraint is loose. To solve our posed rate-constrained estimation problem, we propose an iterative algorithm, where in each iteration the structure (object boundaries) and the texture (surfaces within the object boundaries) of the depth maps are optimized alternately. Using the MVC codec for compression of multiview depth video and MPEG free viewpoint video sequences as input, experimental results show that rate-constrained estimated 3D surfaces computed by our algorithm can reduce coding rate of depth maps by up to 32% compared with unconstrained estimated surfaces for the same quality of synthesized virtual views at the decoder. Wenxiu Sun, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Trans. Image Process. | 5 |
| 2013 | Wide-Baseline Hair Capture Using Strand-Based RefinementabstractWe propose a novel algorithm to reconstruct the 3D geometry of human hairs in wide-baseline setups using strand-based refinement. The hair strands are first extracted in each 2D view, and projected onto the 3D visual hull for initialization. The 3D positions of these strands are then refined by optimizing an objective function that takes into account cross-view hair orientation consistency, the visual hull constraint and smoothness constraints defined at the strand, wisp and global levels. Based on the refined strands, the algorithm can reconstruct an approximate hair surface: experiments with synthetic hair models achieve an accuracy of ~3mm. We also show real-world examples to demonstrate the capability to capture full-head hair styles as well as hair in motion with as few as 8 cameras. Linjie Luo, Cha Zhang, Zhengyou Zhang, Szymon Rusinkiewicz |
CVPR | 2 |
| 2013 | Video Enhancement of People Wearing Polarized Glasses: Darkening Reversal and Reflection ReductionabstractWith the wide-spread of consumer 3D-TV technology, stereoscopic videoconferencing systems are emerging. However, the special glasses participants wear to see 3D can create distracting images. This paper presents a computational framework to reduce undesirable artifacts in the eye regions caused by these 3D glasses. More specifically, we add polarized filters to the stereo camera so that partial images of reflection can be captured. A novel Bayesian model is then developed to describe the imaging process of the eye regions including darkening and reflection, and infer the eye regions based on Classification Expectation-Maximization (EM). The recovered eye regions under the glasses are brighter and with little reflections, leading to a more nature videoconferencing experience. Qualitative evaluations and user studies are conducted to demonstrate the substantial improvement our approach can achieve. Mao Ye 0005, Cha Zhang, Ruigang Yang |
CVPR | 2 |
| 2013 | Rate-distortion optimized 3D reconstruction from noise-corrupted multiview depth videosabstractTransmitting compactly represented geometry of a dynamic scene from a sender can enable a multitude of 3D imaging functionalities at a receiver, such as synthesis of virtual images from freely chosen viewpoints via depth-image-based rendering (DIBR). While depth maps can now be readily captured using inexpensive depth sensors, they are often corrupted by non-negligible acquisition noise. In this paper, we derive 3D surfaces of a dynamic scene from noise-corrupted depth maps in a rate-distortion (RD) optimal manner. Specifically, unlike previous work that finds the most likely (e.g., maximum likelihood) 3D surface from noisy observations regardless of representation size, we judiciously search for the best fitting (i.e., minimum distortion) 3D surface subject to a bitrate constraint. Our RD-optimal solution reduces to the maximum likelihood solution as the rate constraint is loosened. Using the MVC codec for compression of multiview depth video and MPEG free viewpoint test sequences as input, experimental results show that RD-optimized 3D reconstructions computed by our algorithm outperform unprocessed depth maps by up to 2:42dB in PSNR of synthesized virtual views at the decoder for the same bitrate. Wenxiu Sun, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
ICME | 5 |
| 2013 | Analyzing the Optimality of Predictive Transform Coding Using Graph-Based ModelsabstractIn this letter, we provide a theoretical analysis of optimal predictive transform coding based on the Gaussian Markov random field (GMRF) model. It is shown that the eigen-analysis of the precision matrix of the GMRF model is optimal in decorrelating the signal. The resulting graph transform degenerates to the well-known 2-D discrete cosine transform (DCT) for a particular 2-D first order GMRF, although it is not a unique optimal solution. Furthermore, we present an optimal scheme to perform predictive transform coding based on conditional probabilities of a GMRF model. Such an analysis can be applied to both motion prediction and intra-frame predictive coding, and may lead to improvements in coding efficiency in the future. Cha Zhang, Dinei A. F. Florêncio |
IEEE Signal Process. Lett. | 1 |
| 2012 | 3D scene reconstruction by multiple structured-light based commodity depth camerasabstractCommodity depth cameras have attracted a lot of research interest recently, in particular the structured-light based Kinect cameras available on the mass market. One important application of such cameras is 3D scene reconstruction and view synthesis. However, a single depth camera often has limited field of view and there is missing depth information when synthesizing a virtual view from a new viewpoint. In this paper, we study the problem of 3D scene reconstruction from multiple structured-light based depth cameras. Since multiple cameras may cause severe interference in the regions where the projected light overlaps, we present a novel planesweeping based algorithm to handle such interference. The proposed algorithm takes into account the correlation between multiple projectors and the infrared images as well as the correlation between the infrared images, thereby recovering the depth information for both overlapped and non-overlapped regions. Simulation results demonstrate that the proposed solution is very effective on various scenes. Cha Zhang, Wenwu Zhu 0001, Zhengyou Zhang, Zixiang Xiong, Philip A. Chou |
ICASSP | 2 |
| 2012 | See-through Image Enhancement through Sensor FusionabstractMany hardware designs have been developed to allow a camera to be placed optically directly behind the screen. The purpose of such setups is to enable two-way video teleconferencing that maintains eye-contact. However, the image from the see-through camera usually exhibits a number of imaging artifacts such as low signal to noise ratio, incorrect color balance, and lost of details. We develop a novel image enhancement framework that utilizes an auxiliary color+depth camera that is mounted on the side of the screen. By fusing the information from both cameras, we are able to significantly improve the quality of the see-through image. Experimental results have demonstrated that our fusion method compares favorably against traditional image enhancement/warping methods that uses only a single image. Mao Ye 0005, Ruigang Yang, Cha Zhang |
ICME | 4 |
| 2012 | Virtual View Reconstruction Using Temporal InformationabstractThe most significant problem in generating virtual views from a limited number of video camera views is handling areas that have become dis-occluded by shifting the virtual view away from the camera view. We propose using temporal information to address this problem, based on the notion that dis-occluded areas may have been seen by some camera in some previous frames. We formulate the problem as one of estimating the underlying state of the object in a stochastic dynamical system, given a sequence of observations. We apply the formulation to improving the visual quality of virtual views generated from a single “color plus depth” camera, and show that our algorithm achieves better results than depth image based rendering using standard inpainting. Shujie Liu 0001, Philip A. Chou, Cha Zhang, Zhengyou Zhang, Chang Wen Chen |
ICME | 3 |
| 2012 | Automatic Real-Time Video Matting Using Time-of-Flight Camera and Multichannel Poisson Equations
Liang Wang 0002, Minglun Gong, Ruigang Yang, Cha Zhang, Yee-Hong Yang |
Int. J. Comput. Vis. | 5 |
| 2012 | Geometrically Constrained Room Modeling With Compact Microphone ArraysabstractThe geometry of an acoustic environment can be an important information in many audio signal processing applications. To estimate such a geometry, previous work has relied on large microphone arrays, multiple test sources, moving sources or the assumption of a 2-D room. In this paper, we lift these requirements and present a novel method that uses a compact microphone array to estimate a 3-D room geometry, delivering effective estimates with low-cost hardware. Our approach first probes the environment with a known test signal emitted by a loudspeaker co-located with the array, from which the room impulse responses (RIRs) are estimated. It then uses an ℓ1-regularized least-squares minimization to fit synthetically generated reflections to the RIRs, producing a sparse set of reflections. By enforcing structural constraints derived from the image model, these are classified into first-, second-, and third-order reflections, thereby deriving the room geometry. Using this method, we detect walls using off-the-shelf teleconferencing hardware with a typical range resolution of about 1 cm. We present results using simulations and data from real environments. Flavio P. Ribeiro, Dinei A. F. Florêncio, Demba Ba 0001, Cha Zhang |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | CROWDMOS: An approach for crowdsourcing mean opinion score studiesabstractMOS (mean opinion score) subjective quality studies are used to evaluate many signal processing methods. Since laboratory quality studies are time consuming and expensive, researchers often run small studies with less statistical significance or use objective measures which only approximate human perception. We propose a cost-effective and convenient measure called crowdMOS, obtained by having internet users participate in a MOS-like listening study. Workers listen and rate sentences at their leisure, using their own hardware, in an environment of their choice. Since these individuals cannot be supervised, we propose methods for detecting and discarding inaccurate scores. To automate crowdMOS testing, we offer a set of freely distributable, open-source tools for Amazon Mechanical Turk, a platform designed to facilitate crowdsourcing. These tools implement the MOS testing methodology described in this pa per, providing researchers with a user-friendly means of performing subjective quality evaluations without the overhead associated with laboratory studies. Finally, we demonstrate the use of crowdMOS using data from the Blizzard text-to-speech competition, showing that it delivers accurate and repeatable results. Flavio P. Ribeiro, Dinei A. F. Florêncio, Cha Zhang, Michael L. Seltzer |
ICASSP | 3 |
| 2011 | A novel see-through screen based on weave fabricsabstractSee-through screens (STS) have found important applications in remote collaboration systems to enhance non-verbal communication and gaze awareness. Existing STS designs often sacrifice the display quality significantly, rendering low-contrast images that discount the overall user experience. In this paper, we present a novel see-through screen solution based on weave fabrics. Such fabrics are known to be acoustically transparent and used to build professional projection screens for Hollywood studios. We place a cam-era immediately behind the screen and synchronize it with a 120Hz projector to perform time-multiplexing display and video capture. By focusing the camera at the user 4–5 feet away from the screen, the image of the weave fabric will be severely blurred. We present the imaging principle of the setup, and derive image processing techniques to enhance the quality of the captured video. The overall system is low cost, has much better display quality than existing systems, and can be used to build wall-size see-through screens for various applications. Cha Zhang, Ruigang Yang, Tim Large, Zhengyou Zhang |
ICME | 1 |
| 2011 | Calibration between depth and color sensors for commodity depth camerasabstractCommodity depth cameras have created many interesting new applications in the research community recently. These applications often require the calibration information between the color and the depth cameras. Traditional checkerboard based calibration schemes fail to work well for the depth camera, since its corner features cannot be reliably detected in the depth image. In this paper, we present a maximum likelihood solution for the joint depth and color calibration based on two principles. First, in the depth image, points on the checker board shall be co-planar, and the plane is known from color camera calibration. Second, additional point correspondences between the depth and color images may be manually specified or automatically established to help improve calibration accuracy. Uncertainty in depth values has been taken into account systematically. The proposed algorithm is reliable and accurate, as demonstrated by extensive experimental results on simulated and real-world examples. Cha Zhang, Zhengyou Zhang |
ICME | 1 |
| 2011 | Low-complexity, near-lossless coding of depth maps from kinect-like depth camerasabstractDepth cameras are gaining interest rapidly in the market as depth plus RGB is being used for a variety of applications ranging from foreground/background segmentation, face tracking, activity detection, and free viewpoint video rendering. In this paper, we present a low-complexity, near-lossless codec for coding depth maps. This coding requires no buffering of video frames, is table-less, can encode or decode a frame in close to 5ms with little code optimization, and provides between 7:1 to 16:1 compression ratio for near-lossless coding of 16-bit depth maps generated by the Kinect camera. Sanjeev Mehrotra, Zhengyou Zhang, Qin Cai, Cha Zhang, Philip A. Chou |
MMSP | 4 |
| 2011 | An Interactive 3-D Audio System With LoudspeakersabstractTraditional 3-D audio systems using two loudspeakers often have a limited sweet spot and may suffer from poor performance in reverberant environments. This paper presents a novel binaural 3-D audio system that actively combines head tracking and room modeling into 3-D audio synthesis. The user's head position and orientation are first tracked by a webcam-based 3-D head tracker. The system then improves its robustness to head movement and strong early reflections by incorporating the tracking information and an explicit room model into the binaural synthesis and crosstalk cancellation process. Sensitivity analysis on the room model shows that the method is reasonably robust to modeling errors. Subjective listening tests confirm that the proposed 3-D audio system significantly improves the users' perception and ability for localization. Myung-Suk Song, Cha Zhang, Dinei A. F. Florêncio, Hong-Goo Kang |
IEEE Trans. Multim. | 2 |
| 2010 | 3D Deformable Face Tracking with a Commodity Depth Camera
Qin Cai, David Gallup, Cha Zhang, Zhengyou Zhang |
ECCV (3) | 3 |
| 2010 | L1 regularized room modeling with compact microphone arraysabstractAcoustic room modeling has several applications. Recent results using large microphone arrays show good performance, and are helpful in many applications. For example, when designing a better acoustic treatment for a concert hall, these large arrays can be used to help map the acoustic environment and aid in the design. However, in real-time applications - including de-reverberation, sound source localization, speech enhancement and 3D audio - it is desirable to model the room with existing small arrays and existing loudspeakers. In this paper we propose a novel room modeling algorithm, which uses a constrained room model and ℓ1-regularized least-squares to achieve good estimation of room geometry. We present experimental results on both real and synthetic data. Demba Ba 0001, Flavio P. Ribeiro, Cha Zhang, Dinei A. F. Florêncio |
ICASSP | 3 |
| 2010 | Turning enemies into friends: Using reflections to improve sound source localizationabstractSound Source Localization (SSL) based on microphone arrays has numerous applications, and has received significant research attention. Common to all published research is the observation that the accuracy of SSL degrades with reverberation. Indeed, early (strong) reflections can have amplitudes similar to the direct signal, and will often interfere with the estimation. In this paper, we show that reverberation is not the enemy, and can be used to improve estimation. More specifically, we are able to use early reflections to significantly improve range and elevation estimation. The process requires two steps: during setup, a loudspeaker integrated with the array emits a probing sound, which is used to obtain estimates of the ceiling height, as well as the locations of the walls. In a second step (e.g., during a meeting), the device incorporates this knowledge into a maximum likelihood SSL algorithm. Experimental results on both real and synthetic data show huge improvements in range estimation accuracy. Flavio P. Ribeiro, Demba Ba 0001, Cha Zhang, Dinei A. F. Florêncio |
ICME | 3 |
| 2010 | Personal 3D audio system with loudspeakersabstractTraditional 3D audio systems often have a limited sweet spot for the user to perceive 3D effects successfully. In this paper, we present a personal 3D audio system with loudspeakers that has unlimited sweet spots. The idea is to have a camera track the user's head movement, and recompute the crosstalk canceller filters accordingly. As far as the authors are aware of, our system is the first non-intrusive 3D audio system that adapts to both the head position and orientation with six degrees of freedom. The effectiveness of the proposed system is demonstrated with subjective listening tests comparing our system against traditional non-adaptive systems. Myung-Suk Song, Cha Zhang, Dinei A. F. Florêncio, Hong-Goo Kang |
ICME | 2 |
| 2010 | Enhancing loudspeaker-based 3D audio with room modelingabstractFor many years, spatial (3D) sound using headphones has been widely used in a number of applications. A rich spatial sensation is obtained by using head related transfer functions (HRTF) and playing the appropriate sound through headphones. In theory, loudspeaker audio systems would be capable of rendering 3D sound fields almost as rich as headphones, as long as the room impulse responses (RIRs) between the loudspeakers and the ears are known. In practice, however, obtaining these RIRs is hard, and the performance of loudspeaker based systems is far from perfect. New hope has been recently raised by a system that tracks the user's head position and orientation, and incorporates them into the RIRs estimates in real time. That system made two simplifying assumptions: it used generic HRTFs, and it ignored room reverberation. In this paper we tackle the second problem: we incorporate a room reverberation estimate into the RIRs. Note that this is a nontrivial task: RIRs vary significantly with the listener's positions, and even if one could measure them at a few points, they are notoriously hard to interpolate. Instead, we take an indirect approach: we model the room, and from that model we obtain an estimate of the main reflections. Position and characteristics of walls do not vary with the users' movement, yet they allow to quickly compute an estimate of the RIR for each new user position. Of course the key question is whether the estimates are good enough. We show an improvement in localization perception of up to 32% (i.e., reducing average error from 23.5° to 15.9°). Myung-Suk Song, Cha Zhang, Dinei A. F. Florêncio, Hong-Goo Kang |
MMSP | 2 |
| 2010 | Joint tracking and multiview video compressionabstractIn immersive communication applications, knowing the user's viewing position can help improve the efficiency of multiview compression and streaming significantly, since often only a subset of the views are needed to synthesize the desired view(s). However, uncertainty regarding the viewer location can have negative impacts on the rendering quality. In this paper, we propose an algorithm to improve the robustness of view-dependent compression schemes by jointly performing user tracking and compression. A face tracker tracks the user's head location and sends the probability distribution of the face locations as one or many particles. The server then applies motion model to the particles and compresses the multiview video accordingly in order to improve the expected rendering quality of the viewer. Experimental results show significantly improved robustness against tracking errors. Cha Zhang, Dinei A. F. Florêncio |
VCIP | 1 |
| 2010 | Using Reverberation to Improve Range and Elevation Discrimination for Small Array Sound Source LocalizationabstractSound source localization (SSL) is an essential task in many applications involving speech capture and enhancement. As such, speaker localization with microphone arrays has received significant research attention. Nevertheless, existing SSL algorithms for small arrays still have two significant limitations: lack of range resolution, and accuracy degradation with increasing reverberation. The latter is natural and expected, given that strong reflections can have amplitudes similar to that of the direct signal, but different directions of arrival. Therefore, correctly modeling the room and compensating for the reflections should reduce the degradation due to reverberation. In this paper, we show a stronger result. If modeled correctly, early reflections can be used to provide more information about the source location than would have been available in an anechoic scenario. The modeling not only compensates for the reverberation, but also significantly increases resolution for range and elevation. Thus, we show that under certain conditions and limitations, reverberation can be used to improve SSL performance. Prior attempts to compensate for reverberation tried to model the room impulse response (RIR). However, RIRs change quickly with speaker position, and are nearly impossible to track accurately. Instead, we build a 3-D model of the room, which we use to predict early reflections, which are then incorporated into the SSL estimation. Simulation results with real and synthetic data show that even a simplistic room model is sufficient to produce significant improvements in range and elevation estimation, tasks which would be very difficult when relying only on direct path signal components. Flavio P. Ribeiro, Cha Zhang, Dinei A. F. Florêncio, Demba Ba 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Efficient Scale-Space Spatiotemporal Saliency Tracking for Distortion-Free Video Retargeting
Gang Hua 0001, Cha Zhang, Zicheng Liu 0001, Zhengyou Zhang, Ying Shan |
ACCV (2) | 2 |
| 2009 | Boosted multi-task learning for face verification with applications to web image and video searchabstractFace verification has many potential applications including filtering and ranking image/video search results on celebrities. Since these images/videos are taken under uncontrolled environments, the problem is very challenging due to dramatic lighting and pose variations, low resolutions, compression artifacts, etc. In addition, the available number of training images for each celebrity may be limited, hence learning individual classifiers for each person may cause overfitting. In this paper, we propose two ideas to meet the above challenges. First, we propose to use individual bins, instead of whole histograms, of Local Binary Patterns (LBP) as features for learning, which yields significant performance improvements and computation reduction in our experiments. Second, we present a novel Multi-Task Learning (MTL) framework, called Boosted MTL, for face verification with limited training data. It jointly learns classifiers for multiple people by sharing a few boosting classifiers in order to avoid overfitting. The effectiveness of Boosted MTL and LBP bin features is verified with a large number of celebrity images/videos from the web. Cha Zhang, Zhengyou Zhang |
CVPR | 2 |
| 2009 | Multiview video compression and streaming based on predicted viewer positionabstractTechnological advances have made possible a number of new applications in the area of 3D video. One of the enabling technologies for many of these 3D applications is multiview video coding, which has received significant attention in the last several years. However, the fundamental need of multiview coding for applications like immersive tele-conferencing has not been addressed. In this paper we define the boundaries of the problem, and show how a simple algorithm can yield gains of up to 2times reduction in bitrate with similar PSNR in the synthesized view. Our algorithm is based on using an estimate of the viewer position to compute the expected contribution of each pixel to the synthesized view, and encoding each macroblock of each camera views with quality proportional to the likelihood that the pixel will be used in the synthetic image. Dinei A. F. Florêncio, Cha Zhang |
ICASSP | 2 |
| 2009 | ACM 2009 workshop on ambient media computing (AMC'09) overviewabstractNo abstract available. Howard Leung, Cha Zhang, Qing Li 0001, Rynson W. H. Lau, Benjamin W. Wah, Abdulmotaleb El Saddik, K. Selçuk Candan, Irene Cheng 0001 |
ACM Multimedia | 2 |
| 2009 | Improving depth perception with motion parallax and its application in teleconferencingabstractDepth perception, or 3D perception, can add a lot to the feeling of immersiveness in many applications such as 3D TV, 3D teleconferencing, etc. Stereopsis and motion parallax are two of the most important cues for depth perception. Most of the 3D displays today rely on stereopsis to create 3D perception. In this paper, we propose to improve user's depth perception by tracking their motions and creating motion parallax for the rendered image, which can be done even with legacy displays. Two enabling technologies, face tracking and foreground/background segmentation, are discussed in detail. In particular, we propose an efficient and robust feature based face tracking algorithm that is capable of estimating the face's location and scale accurately. We also propose a novel foreground/background segmentation and matting algorithm with time-of-flight camera, which is robust to moving background, lighting variations, moving camera, etc. We demonstrate the application of the above technologies in teleconferencing on legacy displays to create pseudo-3D effects. Cha Zhang, Zhaozheng Yin, Dinei A. F. Florêncio |
MMSP | 1 |
| 2008 | Taylor expansion based classifier adaptation: Application to person detectionabstractBecause of the large variation across different environments, a generic classifier trained on extensive data-sets may perform sub-optimally in a particular test environment. In this paper, we present a general framework for classifier adaptation, which improves an existing generic classifier in the new test environment. Viewing classifier learning as a cost minimization problem, we perform classifier adaptation by combining the cost function on the old data-sets with the cost function on the data-set collected from the new environment. The former term is further approximated with its second order Taylor expansion to reduce the amount of information that needs to be saved for adaptation. Unlike traditional approaches that are often designed for a specific application or classifier, our scheme is applicable to various types of classifiers and user labels. We demonstrate this property on two popular classifiers (logistic regression and boosting), while using two types of user labels (direct labels and similarity labels). Extensive experiments conducted for the task of person detection in conference-room environments show that significant performance improvement can be achieved with our proposed method. Cha Zhang, Raffay Hamid, Zhengyou Zhang |
CVPR | 1 |
| 2008 | Why does PHAT work well in lownoise, reverberative environments?abstractAmong many existing time difference of arrival (TDOA) based sound source localization (SSL) algorithms, the Phase Transform (PHAT) is extremely popular for its excellent performance in low noise environments, even under relatively heavy reverberation. However, PHAT was developed as a heuristic approach and its working principle has not been completely understood. In this paper, we present the relationship between PHAT and a maximum likelihood (ML) framework for multi-microphone sound source localization. We show that when the environment noise approaches zero, PHAT is indeed a special case of the ML algorithm, which explains its good performance under low noise environments. In addition, we show that as long as the noise stays low, PHAT remains optimal in ML sense even when the room reverberation is heavy, which explains its robustness over reverberation. Cha Zhang, Dinei A. F. Florêncio, Zhengyou Zhang |
ICASSP | 1 |
| 2008 | Requirements and recommendations for an enhanced meeting viewing experienceabstractWe have found that viewing recorded meetings using traditional meeting viewers whose interfaces consist of an automatic speaker and a fixed context view does not provide sufficient information and control to the users. In particular, a survey of users who watch meeting recordings on a regular basis revealed that it is also useful to provide (1) speaker-related information, including who the speaker is talking to, looking at, and being interrupted by, and (2) more control of the interface, including changing the relative sizes of the speaker and context views and navigating within the context view. We present a 3D interface prototype designed specifically to meet these requirements when viewing recorded meetings. We describe in detail the results of a user study comparing the effectiveness of the new and traditional style interfaces with respect to these requirements. Based on this study, we present a set of guidelines for future interfaces. Sasa Junuzovic, Rajesh Hegde, Zhengyou Zhang, Philip A. Chou, Zicheng Liu 0001, Cha Zhang |
ACM Multimedia | 6 |
| 2008 | Semantic saliency driven camera control for personal remote collaborationabstractThis paper presents a camera combo system for personal remote collaboration applications. The system consists of two different cameras. One camera has a wide field of view, and the other can pan/tilt/zoom (PTZ) based on analysis of the images captured by the wide angle camera. Unlike traditional approaches which usually drive the PTZ camera to follow the person or his/her head, our system is capable of capturing general objects of interest in remote collaboration. For instance, when the user raises something trying to show it to the remote person, our system will automatically position the PTZ camera to zoom in at the object. At the core of our system is a semantic saliency map that overcomes many limitations of low-level saliency maps computed from preliminary image features. We demonstrate how such a semantic saliency map can be computed through contextual analysis, sign analysis and transitional analysis, and how it can be used for PTZ camera control with a novel information loss optimization based virtual director. The effectiveness of the proposed method is demonstrated with real-world sequences. Cha Zhang, Zicheng Liu 0001, Zhengyou Zhang |
MMSP | 1 |
| 2008 | Maximum Likelihood Sound Source Localization and Beamforming for Directional Microphone Arrays in Distributed MeetingsabstractIn distributed meeting applications, microphone arrays have been widely used to capture superior speech sound and perform speaker localization through sound source localization (SSL) and beamforming. This paper presents a unified maximum likelihood framework of these two techniques, and demonstrates how such a framework can be adapted to create efficient SSL and beamforming algorithms for reverberant rooms and unknown directional patterns of microphones. The proposed method is closely related to steered response power-based algorithms, which are known to work extremely well in real-world environments. We demonstrate the effectiveness of the proposed method on challenging synthetic and real-world datasets, including over six hours of recorded meetings. Cha Zhang, Dinei A. F. Florêncio, Demba Ba 0001, Zhengyou Zhang |
IEEE Trans. Multim. | 1 |
| 2008 | Boosting-Based Multimodal Speaker Detection for Distributed Meeting VideosabstractIdentifying the active speaker in a video of a distributed meeting can be very helpful for remote participants to understand the dynamics of the meeting. A straightforward application of such analysis is to stream a high resolution video of the speaker to the remote participants. In this paper, we present the challenges we met while designing a speaker detector for the Microsoft RoundTable distributed meeting device, and propose a novel boosting-based multimodal speaker detection (BMSD) algorithm. Instead of separately performing sound source localization (SSL) and multiperson detection (MPD) and subsequently fusing their individual results, the proposed algorithm fuses audio and visual information at feature level by using boosting to select features from a combined pool of both audio and visual features simultaneously. The result is a very accurate speaker detector with extremely high efficiency. In experiments that includes hundreds of real-world meetings, the proposed BMSD algorithm reduces the error rate of SSL-only approach by 24.6%, and the SSL and MPD fusion approach by 20.9%. To the best of our knowledge, this is the first real-time multimodal speaker detection algorithm that is deployed in commercial products. Cha Zhang, Pei Yin, Yong Rui, Ross Cutler, Paul A. Viola, Xinding Sun, Nelson Pinto, Zhengyou Zhang |
IEEE Trans. Multim. | 1 |
| 2008 | An automated end-to-end lecture capture and broadcasting systemabstractRemote viewing of lectures presented to a live audience is becoming increasingly popular. At the same time, the lectures can be recorded for subsequent on-demand viewing over the Internet. Providing such services, however, is often prohibitive due to the labor-intensive cost of capturing and pre/post-processing. This article presents a complete automated end-to-end system that supports capturing, broadcasting, viewing, archiving and searching of presentations. Specifically, we describe a system architecture that minimizes the pre- and post-production time, and a fully automated lecture capture system called iCam2 that synchronously captures all contents of the lecture, including audio, video, and presentation material. No staff is needed during lecture capture and broadcasting, so the operational cost of the system is negligible. The system has been used on a daily basis for more than 4 years, during which 522 lectures have been captured. These lectures have been viewed over 20,000 times. Cha Zhang, Yong Rui, Jim Crawford, Li-wei He |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2007 | Maximum Likelihood Sound Source Localization for Multiple Directional MicrophonesabstractThis paper presents a maximum likelihood (ML) framework for multi-microphone sound source localization (SSL). Besides deriving the framework, we focus on making the connection and contrast between the ML-based algorithm and popular steered response power (SRP) SSL algorithms such as phase transform (SRP-PHAT). We also show under our ML framework how challenging conditions such as directional microphone arrays and reverberations can be handled. The computational cost of our method is low-similar to SRP-PHAT. The effectiveness of the proposed method is shown on a large dataset with 99 real-world audio sequences recorded by directional circular microphone arrays in over 50 different meeting rooms. Cha Zhang, Zhengyou Zhang, Dinei A. F. Florêncio |
ICASSP (1) | 1 |
| 2007 | Enhanced MVDR Beamforming for Arrays of Directional MicrophonesabstractMicrophone arrays based on the minimum variance distortionless response (MVDR) beamformer are among the most popular for speech enhancement applications. The original MVDR is excessively sensitive to source location and microphone gains. Previous research has made MVDR practical by successfully increasing the robustness of MVDR to source location, and MVDR-based microphone arrays are already commercially available. Nevertheless, MVDR performance is still weak in cases where microphone gain variations are too large, e.g., for circular arrays of directional microphones. In this paper we propose an improved MVDR beamformer which takes into account the effect of sensors (e.g. microphones) with arbitrary, potentially directional responses. Specifically, we form estimates of the relative magnitude responses of the sensors based on the data received at the array and include those in the original formulation of the MVDR beamforming problem. Experimental results on real-world audio data show an average 2.4 dB improvement over conventional MVDR beamforming, which does not account for the magnitude responses of the sensors. Demba Ba 0001, Dinei A. F. Florêncio, Cha Zhang |
ICME | 3 |
| 2007 | Learning-Based Perceptual Image Quality Improvement for Video ConferencingabstractIt is well known that in professional TV show filming, stage lighting has to be carefully designed in order to make the host and the scene look visually appealing. The lighting affects not only the brightness but also the color tone which plays a critical role in the perceived look of the host and the mood of the stage. In contrast, during video conferencing, the lighting is usually far from ideal thus the perceived image quality is low. There has been a lot of research on improving the brightness of the captured images. In this paper, we propose a learning-based technique to improve the perceptual image quality by enhancing both brightness and color tone. The basic idea is to learn the color statistics from a training set of images which look visually appealing, and adjust the color of an input image so that its color statistics matches those in the training set. To validate our approach, we have conducted user study and the results show that our technique significantly improves the perceived image quality. Zicheng Liu 0001, Cha Zhang, Zhengyou Zhang |
ICME | 2 |
| 2007 | Multiple-Instance Pruning For Learning Efficient Cascade DetectorsabstractCascade detectors have been shown to operate extremely rapidly, with high accuracy, and have important applications such as face detection. Driven by this success, cascade earning has been an area of active research in recent years. Nevertheless, there are still challenging technical problems during the training process of cascade detectors. In particular, determining the optimal target detection rate for each stage of the cascade remains an unsolved issue. In this paper, we propose the multiple instance pruning (MIP) algorithm for soft cascades. This algorithm computes a set of thresholds which aggressively terminate computation with no reduction in detection rate or increase in false positive rate on the training dataset. The algorithm is based on two key insights: i) examples that are destined to be rejected by the complete classifier can be safely pruned early; ii) face detection is a multiple instance learning problem. The MIP process is fully automatic and requires no assumptions of probability distributions, statistical independence, or ad hoc intermediate rejection targets. Experimental results on the MIT+CMU dataset demonstrate significant performance advantages. Cha Zhang, Paul A. Viola |
NIPS | 1 |
| 2007 | Active Rearranged Capturing of Image-Based Rendering Scenes-Theory and PracticeabstractIn this paper, we propose to capture image-based rendering scenes using a novel approach called active rearranged capturing (ARC). Given the total number of available cameras, ARC moves them strategically on the camera plane in order to minimize the sum of squared rendering errors for a given set of light rays to be rendered. Assuming the scene changes slowly so that the optimized camera locations are valid in the next time instance, we formulate the problem as a recursive weighted vector quantization problem, which can be solved efficiently. The ARC approach is verified on both synthetic and real-world scenes. In particular, a large self-reconfigurable camera array is built to demonstrate ARC's performance on real-world scenes. The system renders virtual views at 5-10 frames/s depending on the scene complexity on a moderately equipped computer. Given the virtual view point, the cameras move on a set of rails to perform ARC and improve the rendering quality on the fly Cha Zhang, Tsuhan Chen |
IEEE Trans. Multim. | 1 |
| 2006 | Light Weight Background Blurring for Video Conferencing ApplicationsabstractBackground blurring is an effective way to both preserving privacy and keeping communication effective during video conferencing. This paper proposes a light weight real-time algorithm to perform background blurring using a fast background modeling algorithm combined with a face detector/tracker. A soft decision is made at each pixel whether it belongs to the foreground or the background based on multiple vision features. The classification results are mapped to a per-pixel blurring radius image to blur the background. The algorithm produces satisfactory results under a wide range of conditions, and occupies less than 30% of the CPU cycles on a 3 GHz Pentium 4 machine without further optimization. Cha Zhang, Yong Rui, Li-wei He |
ICIP | 1 |
| 2006 | A Three-Layer Virtual Director Model for Supporting Automated Multi-Site Distributed EducationabstractIn multi-site distributed education (MSDE), video streams from multiple sites are available. To best utilize the limited screen space at each site, we develop a customizable, automated display management system in this paper, i.e., only user-preferred streams will be shown as triggered by events and timers. The configuration of such user preference, however, is challenging because it has to be both human-friendly and machine-friendly. To address this challenge, we propose a three-layer virtual director model. In the user layer, we identify three categories of parameters that can represent a wide range of user preferences yet are easy to use. These preferences are then automatically translated into a machine-friendly timed automaton in the execution layer. The automaton is simulated dynamically, which selects a subset of streams to show on the screen through a display layer. Evaluation results demonstrate the correctness and efficiency of the proposed framework Bin Yu 0010, Cha Zhang, Yong Rui, Klara Nahrstedt |
ICME | 2 |
| 2006 | Boosting-Based Multimodal Speaker Detection for Distributed MeetingsabstractSpeaker detection is a very important task in distributed meeting applications. This paper discusses a number of challenges we met while designing a speaker detector for the Microsoft RoundTable distributed meeting device, and proposes a boosting-based multimodal speaker detection (BMSD) algorithm. Instead of performing sound source localization (SSL) and multi-person detection (MPD) separately and subsequently fusing their individual results, the proposed algorithm uses boosting to select features from a combined pool of both audio and visual features simultaneously. The result is a very accurate speaker detector with extremely high efficiency. The algorithm reduces the error rate of SSL-only approach by 47%, and the SSL and MPD fusion approach by 27% Cha Zhang, Pei Yin, Yong Rui, Ross Cutler, Paul A. Viola |
MMSP | 1 |
| 2005 | Light field capturing with lensless camerasabstractWe introduce a novel approach to capturing light field with lensless cameras. By moving the cameras back and forth, we capture a set of images. We show that it is possible to reconstruct the light field from these blurry images. The problem is formulized in a way similar to computer tomography, so that the light field can be reconstructed using existing algorithms. The light field can then be used to render 3D scenes. Synthetic examples are presented to show the effectiveness of the proposed method. Cha Zhang, Tsuhan Chen |
ICIP (3) | 1 |
| 2005 | Hybrid speaker tracking in an automated lecture roomabstractWe present a hybrid speaker tracking scheme based on a single pan/tilt/zoom (PTZ) camera in an automated lecture capturing system. Given that the camera's video resolution is higher than the required output resolution, we frame the output video as a sub-region of the camera's input video. This allows us to track the speaker both digitally and mechanically. Digital tracking has the advantage of being smooth, and mechanical tracking can cover a wide area. The hybrid tracking achieves the benefits of both worlds. In addition to hybrid tracking, we present an intelligent pan/zoom selection scheme to improve the aestheticity of the lecture scene. Cha Zhang, Yong Rui, Li-wei He, Michael N. Wallick |
ICME | 1 |
| 2005 | An automated end-to-end lecture capturing and broadcasting systemabstractWe present a complete end-to-end system that is fully automated and supports capturing, broadcasting, viewing, archiving and search. Specifically, we describe a system architecture that minimizes the pre- and post-production time, and a fully automated lecture capturing system called iCam2, which synchronously captures all the contents of the lecture, including audio, video and visual aids. As no staff is needed during the capturing and broadcasting process, the operation cost of our system is negligible. The system has been used on a daily basis for more than 4 years, during which 467 lectures were captured with 17,000+ online viewers. Cha Zhang, Jim Crawford, Yong Rui, Li-wei He |
ACM Multimedia | 1 |
| 2005 | Multiple Instance Boosting for Object DetectionabstractA good image object detection algorithm is accurate, fast, and does not require exact locations of objects in a training set. We can create such an object detector by taking the architecture of the Viola-Jones detector cascade and training it with a new variant of boosting that we call MILBoost. MILBoost uses cost functions from the Multiple Instance Learning literature combined with the AnyBoost framework. We adapt the feature selection criterion of MILBoost to optimize the performance of the Viola-Jones cascade. Experiments show that the detection rate is up to 1.6 times better using MILBoost. This increased detection rate shows the advantage of simultaneously learning the locations and scales of the objects in the training set along with the parameters of the classifier. Paul A. Viola, John C. Platt, Cha Zhang |
NIPS | 3 |
| 2005 | On the compression and streaming of concentric mosaic data for free wandering in a realistic environment over the InternetabstractIn this paper, we describe a system for wandering in a realistic environment over the Internet. The environment is captured by the concentric mosaic, compressed via the reference block coder (RBC), and accessed and delivered over the Internet through the virtual media (Vmedia) access protocol. One of the key contributions of the paper is the proposal of the RBC concentric mosaic coder. The RBC coder not only compresses the huge dataset of the concentric mosaic very efficiently, but also produces a well-organized bitstream that can be accessed just-in-time (JIT). To reconstruct a virtual view, only a portion of the RBC bitstream needs to be accessed and decoded. This greatly reduces the memory and computation requirement of the viewer compared with first decoding the entire concentric mosaic data set and then rendering from the decoded data. Our second contribution is the employment of the Vmedia protocol to deliver the compressed concentric mosaic bitstream just-in-time over the Internet. Only the bitstream segments corresponding to the current view are streamed over the Internet. The delivered bitstream segments are managed by a local Vmedia cache, so that frequently used bitstream segments do not need to be streamed over the Internet repeatedly, and a RBC bitstream larger than the memory capacity can be easily handled. Combining RBC and Vmedia, a concentric mosaic interactive browser is developed through which the user can freely wander in a realistic environment, e.g., rotate around, walk forward/backward, and sidestep, even under a tight bandwidth. Cha Zhang |
IEEE Trans. Multim. | 1 |
| 2004 | View-dependent non-uniform sampling for image-based renderingabstractIn this paper, we propose an algorithm for view-dependent nonuniform sampling for image-based rendering (IBR). Given a set of virtual views, the positions of the capturing cameras are rearranged in order to obtain the optimal rendering quality. The resulting arrangement of the cameras is effectively nonuniform sampling of the plenoptic function. We formulate the above sampling problem as a recursive weighted vector quantization problem, which can be solved efficiently. Experimental results show that the nonuniform sampling scheme renders much better images than traditional uniform sampling methods. Cha Zhang, Tsuhan Chen |
ICIP | 1 |
| 2004 | Security analysis for key generation systems using face imagesabstractIn this paper, we analyze the security problem of user information associated key generation (UIAKG) systems. We consider three kinds of attacks from the hacker: the exhaustive search attack, the authentic key statistics attack and the device key statistics attack. Under each attack, we give the estimate of the number of guesses the hacker has to make in order to access the system. Such analysis provides useful guidelines in designing a UIAKG system. The analysis also suggests that a user-dependent UIAKG is more secure than a user-independent one. Two UIAKG systems using face images as inputs are designed and compared to support the above theoretical analysis. Wende Zhang, Cha Zhang, Tsuhan Chen |
ICIP | 2 |
| 2004 | Semantic propagation from relevance feedbacksabstractRelevance feedback has been a very useful tool to enhance the performance of content-based information retrieval (CBIR) systems. To fully make use of the precious user feedback provided to a system, we propose an approach named semantic propagation, which reveals the deep semantic relationships among objects in the database, given a set of relevance feedbacks between object pairs. In particular, we present two semantic propagation algorithms that are applicable to CIBR systems with a feature vector space model and a general metric space model, respectively. Experiments on a 3D model retrieval system and a logo image retrieval system are performed to show the effectiveness of the proposed methods. Hoon Yul Bang, Cha Zhang, Tsuhan Chen |
ICME | 2 |
| 2004 | Distributed hosting of Web content with erasure coding and unequal weight assignmentabstractBy sharing and hosting personal Web content in a peer-to-peer (P2P) fashion, we may increase the bandwidth of retrieval and improve the reliability of retrieval. The cost is that the Web content has to be replicated to and stored in the peers, consuming valuable network bandwidth and peer storage space. In this work, we develop two technologies, namely hierarchical content organization with unequal weight assignment and erasure coding, to reduce the amount of content to be distributed, yet still maintain the retrieval speed up and reliability. Significant improvement over the ordinary Web server is demonstrated. Cha Zhang |
ICME | 2 |
| 2004 | Non-Uniform Sampling for Image-Based Rendering: Convergence of Image, Vision, and GraphicabstractRecent convergence of image processing, computer vision, and computer graphics has resulted in an exciting research topic referred to as image-based rendering (IBR). Widely used in applications ranging from movie special effects (e.g., "the Matrix") to building virtual environments, IBR has become a critical tool for creating visually exciting content. With IBR, real-world scenes can be captured and rendered directly from images captured by cameras, eliminating the need for computationally expensive modeling of 3D geometry or surface reflectance, as is often done in traditional computer graphics. Various approaches to IBR have been proposed to render the scenes correctly and effectively given the captured images. We propose an active scene-capturing algorithm to efficiently capture the images. Based on the images that have been taken and the geometry information known so far, the algorithm intelligently determines where to pose the camera to capture the scene for best rendering performance. This results in a nonuniform but optimal set of captured images. Cha Zhang, Tsuhan Chen |
MMM | 1 |
| 2004 | A survey on image-based rendering - representation, sampling and compression
Cha Zhang, Tsuhan Chen |
Signal Process. Image Commun. | 1 |
| 2003 | On generalized sampling for image-based rendering dataabstractWe apply generalized sampling to image-based rendering (IBR) data, more specifically, the lightfield. We show that, in theory, the lowest sampling rate of a lightfield when we use generalized sampling can be as low as half of that when we use rectangular sampling. However, in practice, rectangular sampling has several advantages over generalized sampling. We analyze the pros and cons for each sampling approach, and explain why, in practice, rectangular sampling is still preferable. Cha Zhang, Tsuhan Chen |
ICASSP (3) | 1 |
| 2003 | Surface plenoptic function: a tool for the sampling analysis of image-based renderingabstractIn this paper, we introduce the surface plenoptic function (SPF) as a tool for the sampling analysis of image-based rendering (IBR). SPF is a plenoptic function defined on the object surface. It can be mapped onto the normal IBR representation via a coordinate transform. By assuming some properties of the SPF, we can provide insightful analysis on the sampling of the IBR representation, for both uniform and non-uniform methods. Cha Zhang, Tsuhan Chen |
ICASSP (4) | 1 |
| 2003 | Annotating retrieval database with active learningabstractIn this paper, we describe a retrieval system that uses hidden annotation to improve the performance. The contribution of this paper is a novel active learning framework that can improve the annotation efficiency. For each object in the database, we maintain a list of probabilities, each indicating the probability of this object having one of the attributes. This list of probabilities serves as the basis of our active learning algorithm, as well as semantic features to determine the similarity between objects in the database. We show active learning has better performance than random sampling in all our experiments. Cha Zhang, Tsuhan Chen |
ICIP (2) | 1 |
| 2003 | A system for active image-based renderingabstractIn this paper, we develop a system for active image-based rendering (IBR). Active IBR is a framework that is capable of estimating the final rendering quality and capturing the next view at the position where the rendering quality can be improved the most. The result is a non-uniform capturing scheme for IBR. Experimental results on synthetic scenes have shown that active IBR outperforms uniformly captured IBR. In this paper, we set up an active IBR system that can capture real objects with off-the-shelf components. We show that our system can capture objects more intelligently than uniform capturing. Cha Zhang, Tsuhan Chen |
ICME | 1 |
| 2003 | Nonuniform sampling of image-based rendering data with the position-interval-error (PIE) function
Cha Zhang, Tsuhan Chen |
VCIP | 1 |
| 2003 | Color image sharpening based on collective time-evolution of simultaneous nonlinear reaction-diffusion
Cha Zhang, Tsuhan Chen |
VCIP | 1 |
| 2003 | Spectral analysis for sampling image-based rendering dataabstractImage-based rendering (IBR) has become a very active research area in recent years. The spectral analysis problem for IBR has not been completely solved. In this paper, we present a new method to parameterize the problem, which is applicable for general-purpose IBR spectral analysis. We notice that any plenoptic function is generated by light ray emitted/reflected/refracted from the object surface. We introduce the surface plenoptic function (SPF), which represents the light rays starting from the object surface. Given that radiance along a light ray does not change unless the light ray is blocked, SPF reduces the dimension of the original plenoptic function to 6D. We are then able to map or transform the SPF to IBR representations captured along any camera trajectory. Assuming some properties on the SPF, we can analyze the properties of IBR for generic scenes such as scenes with Lambertian or non-Lambertian surfaces and scenes with or without occlusions, and for different sampling strategies such as lightfield/concentric mosaic. We find that in most cases, even though the SPF may be band-limited, the frequency spectrum of IBR is not band-limited. We show that non-Lambertian reflections, depth variations and occlusions can all broaden the spectrum, with the latter two being more significant. SPF is defined for scenes with known geometry. When the geometry is unknown, spectral analysis is still possible. We show that with the "truncating windows" analysis and some conclusions obtained with SPF, the spectrum expansion caused by non-Lambertian reflections and occlusions can be quantatively estimated, even when the scene geometry is not explicitly known. Given the spectrum of IBR, we also study how to sample IBR data more efficiently. Our analysis is based on the generalized periodic sampling theory with arbitrary geometry. We show that the sampling efficiency can be up to twice of that when we use rectangular sampling. The advantages and disadvantages of generalized periodic sampling for IBR are also discussed. Cha Zhang, Tsuhan Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Towards optimal least square filters using the eigenfilter approachabstractIn this paper, we propose a new eigenfilter approach to designing least square error filters. The filters are obtained by finding an eigenvector of a real, symmetric and positive definite matrix, which is numerically stable. The proposed algorithm has two advantages. First, we show that the least-square solution, which can only be obtained through matrix inversion in the literature, can be asymptotically reached with our algorithm. Second, when numerical errors break. the matrix inversion method, our algorithm can still find some ìoptimalî filter through tuning an internal parameter. Cha Zhang, Tsuhan Chen |
ICASSP | 1 |
| 2002 | Smart rebinning for the compression of concentric mosaicabstractThe concentric mosaic offers a quick solution to construct a virtual copy of a real environment, and to navigate in that virtual environment. However, the huge amount of data associated with the concentric mosaic is a heavy burden for its application. A three-dimensional (3D) wavelet-based compressor has been proposed in a previous work to compress the concentric mosaic. In this paper, we greatly improve the compression efficiency of the 3D wavelet coder by a data rearrangement mechanism called "smart rebinning". The proposed scheme first aligns the concentric mosaic image shots along the horizontal direction and then rebins the shots into multiperspective panoramas. Smart rebinning effectively enhances the correlation in the 3D data volume, translating the data into a representation that is more suitable for 3D wavelet compression. Experimental results show that the performance of the 3D wavelet coder is improved by an average of 4.3 dB with the use of the smart rebinning. The proposed coder outperforms MPEG-2 coding of the concentric mosaic by an average of 3.7 dB. Yunnan Wu, Cha Zhang, Jin Li 0001 |
IEEE Trans. Multim. | 2 |
| 2002 | An active learning framework for content-based information retrievalabstractWe propose a general active learning framework for content-based information retrieval. We use this framework to guide hidden annotations in order to improve the retrieval performance. For each object in the database, we maintain a list of probabilities, each indicating the probability of this object having one of the attributes. During training, the learning algorithm samples objects in the database and presents them to the annotator to assign attributes. For each sampled object, each probability is set to be one or zero depending on whether or not the corresponding attribute is assigned by the annotator. For objects that have not been annotated, the learning algorithm estimates their probabilities with biased kernel regression. Knowledge gain is then defined to determine, among the objects that have not been annotated, which one the system is the most uncertain. The system then presents it as the next sample to the annotator to which it is assigned attributes. During retrieval, the list of probabilities works as a feature vector for us to calculate the semantic distance between two objects, or between the user query and an object in the database. The overall distance between two objects is determined by a weighted sum of the semantic distance and the low-level feature distance. The algorithm is tested on both synthetic databases and real databases of 3D models. In both cases, the retrieval performance of the system improves rapidly with the number of annotated samples. Furthermore, we show that active learning outperforms learning based on random sampling. Cha Zhang, Tsuhan Chen |
IEEE Trans. Multim. | 1 |
| 2001 | Efficient feature extraction for 2D/3D objects in mesh representationabstractMeshes are dominantly used to represent 3D models as they fit well with graphics rendering hardware. Features such as volume, moments, and Fourier transform coefficients need to be calculated from the mesh representation efficiently. We propose an algorithm to calculate these features without transforming the mesh into other representations such as the volumetric representation. To calculate a feature for a mesh, we show that we can first compute it for each elementary shape such as a triangle or a tetrahedron, and then add up all the values for the mesh. The algorithm is simple and efficient, with many potential applications. Cha Zhang, Tsuhan Chen |
ICIP (3) | 1 |
| 2001 | Indexing and retrieval of 3D models aided by active learningabstractWe demonstrate a system for indexing and retrieval of 3D models aided by active learning. We propose a new set of region-based features for 3D models. Each model is treated as a solid volume with a uniform density. Features such as the volume-surface ratio, the moment invariants and the Fourier transform coefficients are efficiently calculated from the mesh model directly. Comparable retrieval performance is achieved with other features such as the cord histogram, the 3D shape spectrum, etc. To further improve the performance, we incorporate hidden annotation into our system. We propose to use active learning to improve the annotation efficiency. We show that with active learning, the system can perform better than random annotation, and the retrieval result improves rapidly with the number of annotated samples. Moreover, relevance feedback is included in the system and combined with active learning, which provides better user-adoptive retrieval results. Cha Zhang, Tsuhan Chen |
ACM Multimedia | 1 |
| 2001 | Interactive browsing of 3D environment over the Internet
Cha Zhang |
VCIP | 1 |
| 2000 | Compression of Lumigraph with Multiple Reference Frame (MRF) Prediction and Just-in-Time RenderingabstractIn the form of a 2D image array, Lumigraph captures the complete appearance of an object or a scene, and is able to quickly render a novel view independent of the scene/object complexity. Since the data amount of Lumigraph is huge, efficient storage and access for Lumigraph are essential. In this paper, we propose a multiple reference frame (MRF) structure to compress the Lumigraph data. By predicting the Lumigraph view from multiple neighbor views, a higher compression ratio is achieved. We also implement the key functionality of just-in-time (JIT) Lumigraph rendering, in which only a small portion of the compressed bitstream necessary for rendering the current view is accessed and decoded. JIT rendering eliminates the need to predecode the entire Lumigraph data set, thus greatly reduces the memory requirement of Lumigraph rendering. A decoder cache has been implemented to speed up rendering by reusing the decoded data. The trade off between the computational speed and cache size of the decoder is discussed in the paper. Cha Zhang |
Data Compression Conference | 1 |
| 2000 | Smart rebinning for compression of concentric mosaicsabstractConcentric mosaics offer a quick solution to construct a virtual copy of a real environment, and navigate in the virtual environment. However, the huge amount of data associated with concentric mosaics is a heavy burden for its application. A 3D wavelet transform-based compressor has been proposed in previous work to compress the concentric mosaics. In this paper, we greatly improve the performance of the 3D wavelet coder with a data rearrangement mechanism called “smart rebinning”. The proposed scheme first aligns the concentric mosaic image shots along the horizontal direction and then rebins the shots into multi-perspective panoramas. Smart rebinning greatly improves the cross shot correlation and enables the coder to better explore the redundancy among shots. Experimental results show that the performance of the 3D wavelet coder improves an average of 4.3dB with the use of smart rebinning. The proposed coder outperforms MPEG-2 coding of concentric mosaics by an average of 3.7dB. Yunnan Wu, Cha Zhang, Jin Li 0001, Jizheng Xu |
ACM Multimedia | 2 |
| 2000 | Compression and rendering of concentric mosaics with reference block codec (RBC)
Cha Zhang |
VCIP | 1 |