EDBT 2026 Demo / reviewers in the wild / expert
Edward J. Delp
dblp:d/EdwardJDelp · also Edward J. Delp III
· DBLP profile ↗
246ranked-venue papers
6as first author
26since 2021 · last 2025
0000-0002-2909-7323ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 186 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 27 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 6 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-authorSystems, architecture and hardware · 7Computer networks · 7 · 2 first-authorSecurity and privacy · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiffSSD: A Diffusion-Based Dataset For Speech ForensicsabstractDiffusion-based speech generators are ubiquitous. These methods can generate very high quality synthetic speech and several recent incidents report their malicious use. To counter such misuse, synthetic speech detectors have been developed. Many of these detectors are trained on datasets which do not include diffusion-based synthesizers. In this paper, we demonstrate that existing detectors trained on one such dataset, ASVspoof2019, do not perform well in detecting synthetic speech from recent diffusion-based synthesizers. We propose the Diffusion-Based Synthetic Speech Dataset (DiffSSD), a dataset consisting of about 200 hours of labeled speech, including synthetic speech generated by 8 diffusion-based open-source and 2 commercial generators. We also examine the performance of existing synthetic speech detectors on DiffSSD in both closed-set and open-set scenarios. The results highlight the importance of this dataset in detecting synthetic speech generated from recent open-source and commercial speech generators. Kratika Bhagtani, Amit Kumar Singh Yadav, Paolo Bestagini, Edward J. Delp |
ICASSP | 4 |
| 2024 | A Natural Approach for Synthetic Short-Form Text AnalysisabstractDetecting synthetically generated text in the wild has become increasingly difficult with advances in Natural Language Generation techniques and the proliferation of freely available Large Language Models (LLMs). Social media and news sites can be flooded with synthetically generated misinformation via tweets and posts while authentic users can inadvertently spread this text via shares and retweets. Most modern natural language processing techniques designed to detect synthetically generated text focus primarily on long-form content, such as news articles, or incorporate stylometric characteristics and metadata during their analysis. Unfortunately, for short form text like tweets, this information is often unavailable, usually detached from its original source, displayed out of context, and is often too short or informal to yield significant information from stylometry. This paper proposes a method of detecting synthetically generated tweets via a Transformer architecture and incorporating unique style-based features. Additionally, we have created a new dataset consisting of human-generated and Large Language Model generated tweets for 4 topics and another dataset consisting of tweets paraphrased by 3 different paraphrase models. Ruiting Shao, Ryan Schwarz, Christopher Clifton, Edward J. Delp |
LREC/COLING | 4 |
| 2024 | Mdrt: Multi-Domain Synthetic Speech LocalizationabstractWith recent advancements in generating synthetic speech, tools to generate high-quality synthetic speech impersonating any human speaker are easily available. Several incidents report misuse of high-quality synthetic speech for spreading misinformation and for large-scale financial frauds. Many methods have been proposed for detecting synthetic speech; however, there is limited work on localizing the synthetic segments within the speech signal. In this work, our goal is to localize the synthetic speech segments in a partially synthetic speech signal. Most existing methods for synthetic speech localization obtain features from either the time domain waveform or the spectrogram representation of the speech signal. In this work, we propose Multi-Domain ResNet Transformer (MDRT) that obtains multi-domain features from both the time domain and the spectrogram representation of a speech signal to localize synthetic speech segments. MDRT uses transformer neural networks to obtain multi-domain features and processes them using a ResNet-style neural network. We use the PartialSpoof dataset to examine the performance of MDRT on localizing synthetic speech segments of varying duration. Our results show that MDRT performs better than several existing synthetic speech localization methods. Amit Kumar Singh Yadav, Kratika Bhagtani, Sriram Baireddy, Paolo Bestagini, Stefano Tubaro, Edward J. Delp |
ICASSP | 6 |
| 2024 | Are Recent Deepfake Speech Generators Detectable?abstractDeep learning methods can generate high-quality synthetic speech which is perceptually indistinguishable from real human speech. Synthetic speech can be maliciously used for fraud. Synthetic speech detection methods have been proposed which perform well on ASVspoof2019 and ASVspoof2021 Datasets. These datasets consist of synthetic speech from conventional neural network speech generators. Recently, many voice cloning methods have been proposed which use diffusion models and generative adversarial networks for high-quality speech synthesis. In this work, we present a new synthetic speech dataset containing 25,000 synthetic speech signals for 11 distinct speakers, with a total duration of 52 hours. We have developed this dataset using 5 recent diffusion model-based synthetic speech generators. These generators can clone a speaker's voice from text using only a few minutes of their real speech. We evaluate 6 of the best synthetic speech detectors that work well on the ASVspoof2019 Dataset on this new dataset, and demonstrate their performance using Equal Error Rate (EER). Kratika Bhagtani, Amit Kumar Singh Yadav, Paolo Bestagini, Edward J. Delp |
IH&MMSec | 4 |
| 2024 | A Quantitative Metric of Confidence for Segmentation of Nuclei in Large Spatially Variable Image VolumesabstractNuclei segmentation is an important step for quantitative analysis of fluorescence microscopy images. A large volume generally has many different regions containing nuclei with varying spatial characteristics. Automatically identifying nuclei that are challenging to segment can speed up the analysis of biological tissues. Here we show a segmentation technique that provides a metric of segmentation “confidence” for each segmented object in an image volume. This confidence metric can be used either to generate a “confidence map” for visual distinction of reliable from unreliable regions, or in the data space to identify questionable measurements that can be analyzed separately or eliminated from analysis. In an analysis of nuclei in a 3-dimensional image volume, we show that the confidence map correlates well with visual evaluations of segmentation quality, and that the confidence metric correlates well with F1 scores within subregions of the image volume. In addition, we also describe three visualization methods that can visualize the segmentation differences between a segmented volume and a reference volume. Liming Wu, Alain Chen, Paul Salama, Kenneth W. Dunn, Seth Winfree, Edward J. Delp |
MMSP | 6 |
| 2023 | ASSD: Synthetic Speech Detection in the AAC Compressed DomainabstractSynthetic human speech signals have become very easy to generate given modern text-to-speech methods. When these signals are shared on social media they are often compressed using the Advanced Audio Coding (AAC) standard. Our goal is to study if a small set of coding metadata contained in the AAC compressed bit stream is sufficient to detect synthetic speech. This would avoid decompressing of the speech signals before analysis. We call our proposed method AAC Synthetic Speech Detection (ASSD). ASSD extracts information from the AAC compressed bit stream without decompressing the speech signal. ASSD analyzes the information using a transformer neural network. In our experiments, we compressed the ASVspoof2019 dataset according to the AAC standard using different data rates. We compared the performance of ASSD to a time domain based and a spectrogram based synthetic speech detection methods. We evaluated ASSD on approximately 71k compressed speech signals. The results show that our proposed method typically only requires 1000 bits per speech block/frame from the AAC compressed bit stream to detect synthetic speech. This is much lower than other reported methods. Our method also had a 9.7 percentage points higher detection accuracy compared to existing methods. Amit Kumar Singh Yadav, Ziyue Xiang, Emily R. Bartusiak, Paolo Bestagini, Stefano Tubaro, Edward J. Delp |
ICASSP | 6 |
| 2023 | DSVAE: Disentangled Representation Learning for Synthetic Speech DetectionabstractTools to generate high quality synthetic speech that is perceptually indistinguishable from speech recorded from hu-man speakers are easily available. Many incidents report misuse of synthetic speech for spreading misinformation and committing financial fraud. Several approaches have been proposed for detecting synthetic speech. Many of these approaches use deep learning methods without providing reasoning for the decisions they make. This limits the explainability of these approaches. In this paper, we use disentangled representation learning for developing a synthetic speech detector. We propose Disentangled Spectrogram Variational Auto Encoder (DSVAE) which is a two stage trained variational autoencoder that processes spectrograms of speech to generate features that disentangle synthetic and bona fide speech. We evaluated DSVAE using the ASVspoof2019 dataset. Our experimental results show high accuracy (> 98%) on detecting synthetic speech from 6 known and 10 unknown speech synthesizers. Further, the visualization of disentangled features obtained from DSVAE provides rea-soning behind the working principle of DSVAE, improving its explainability. DSVAE performs well compared to several existing methods. Additionally, DSVAE works in practical scenarios such as detecting synthetic speech uploaded on social platforms and against simple attacks such as removing silence regions. Amit Kumar Singh Yadav, Kratika Bhagtani, Ziyue Xiang, Paolo Bestagini, Stefano Tubaro, Edward J. Delp |
ICMLA | 6 |
| 2023 | PS3DT: Synthetic Speech Detection Using Patched Spectrogram TransformerabstractMany deep learning synthetic speech generation tools are readily available. The use of synthetic speech has caused financial fraud, impersonation of people, and misinformation to spread. For this reason forensic methods that can detect synthetic speech have been proposed. Existing methods often overfit on one dataset and their performance reduces substantially in practical scenarios such as detecting synthetic speech shared on social platforms. In this paper we propose, Patched Spectrogram Synthetic Speech Detection Transformer (PS3DT), a synthetic speech detector that converts a time domain speech signal to a mel-spectrogram and processes it in patches using a trans-former neural network. We evaluate the detection performance of PS3DT on ASVspoof2019 dataset. Our experiments show that PS3DT performs well on ASVspoof2019 dataset compared to other approaches using spectrogram for synthetic speech detection. We also investigate generalization performance of PS3DT on In-the-Wild dataset. PS3DT generalizes well than several existing methods on detecting synthetic speech from an out-of-distribution dataset. We also evaluate robustness of PS3DT to detect telephone quality synthetic speech and synthetic speech shared on social platforms (compressed speech). PS3DT is robust to compression and can detect telephone quality synthetic speech better than several existing methods. Amit Kumar Singh Yadav, Ziyue Xiang, Kratika Bhagtani, Paolo Bestagini, Stefano Tubaro, Edward J. Delp |
ICMLA | 6 |
| 2023 | Semi-Supervised Object Detection for Sorghum Panicles in UAV ImageryabstractThe sorghum panicle is an important trait related to grain yield and plant development. Detecting and counting sorghum panicles can provide significant information for plant phenotyping. Current deep-learning-based object detection methods for panicles require a large amount of training data. The data labeling is time-consuming and not feasible for real application. In this paper, we present an approach to reduce the amount of training data for sorghum panicle detection via semi-supervised learning. Results show we can achieve similar performance as supervised methods for sorghum panicle detection by only using 10% of original training data. Enyu Cai, Changye Yang, Edward J. Delp |
IGARSS | 4 |
| 2023 | Rotation Adaptive Plot Extraction from UAV RGB ImagesabstractUnmanned Aerial Vehicles (UAVs) have the ability to acquire high resolution RGB images from plant fields. A field used for plant research is usually divided into smaller groups of plants known as "plots" to evaluate varieties or management practices. In this paper, we propose an optimization-based, rotation-adaptive approach for extracting plots in a UAV RGB orthomosaic image. From the experiments, our proposed method is robust against range/plot rotation and achieves higher segmentation accuracy compared with existing plot extraction approaches. Changye Yang, Enyu Cai, Edward J. Delp |
IGARSS | 4 |
| 2023 | Synthesized Speech Attribution Using The Patchout Spectrogram Attribution TransformerabstractThe malicious use of synthetic speech has increased with the recent availability of speech generation tools. It is important to determine whether a speech signal is authentic (spoken by a human) or is synthesized and to determine the generation method used to create it. Identifying the synthesis method is known as synthetic speech attribution. In this paper, we propose the use of a transformer deep learning method that analyzes mel-spectrograms for synthetic speech attribution. Our method known as Patchout Spectrogram Attribution Transformer (PSAT) can distinguish new, unseen speech generation methods from those seen during training. PSAT demonstrates high performance in attributing synthetic speech signals. Evaluation on the DARPA SemaFor Audio Attribution Dataset and the ASVSpoof2019 Dataset shows that our method achieves more than 95% accuracy in synthetic speech attribution and performs better than existing deep learning approaches. Kratika Bhagtani, Emily R. Bartusiak, Amit Kumar Singh Yadav, Paolo Bestagini, Edward J. Delp |
IH&MMSec | 5 |
| 2023 | Extracting Efficient Spectrograms From MP3 Compressed Speech Signals for Synthetic Speech DetectionabstractMany speech signals are compressed with MP3 to reduce the data rate. In many synthetic speech detection methods the spectrogram of the speech signal is used. This usually requires the speech signal to be fully decompressed. We show that the design of MP3 compression allows one to approximate the spectrogram of the MP3 compressed speech efficiently without fully decoding the compressed speech. We denote the spectograms obtained using our proposed approach by Efficient Spectrograms (E-Specs). E-Spec can reduce the complexity of spectrogram computation by ~77.60 percentage points (p.p.) and save ~37.87 p.p. of MP3 decoding time. E-Spec bypasses the reconstruction artifacts introduced by the MP3 synthesis filterbank, which makes it useful in speech forensics tasks. We tested E-Spec in the synthetic speech detection, where a detector is asked to determine whether a speech signal is synthesized or recorded from a human. We examined 4 different neural network architectures to evaluate the performance of E-Spec compared to speech features extracted from the fully decoded speech signal. E-Spec achieved the best synthetic speech detection performance for 3 architectures; it also achieved the best overall detection performance across architectures. The computation of E-Spec is an approximation to Short Time Fourier Transform (STFT). E-Spec can be extended to other audio compression methods. Ziyue Xiang, Amit Kumar Singh Yadav, Stefano Tubaro, Paolo Bestagini, Edward J. Delp |
IH&MMSec | 5 |
| 2023 | Gallery-Query Protocol for Evaluating Face Image Quality MetricsabstractAs more automated face recognition systems are integrated into society, face image quality estimators (QE) become important. These QEs help quantify whether an input face image contains reliable information necessary for face recognition. However, to assess the effectiveness of face QEs, it is essential to have reliable evaluation protocols. Current face QE evaluation protocols require long computation times and do not have explicit real-world implications. In this paper, we propose a novel face QE evaluation protocol named “Gallery-Query (GQ) Protocol”. The GQ protocol is significantly faster in evaluating face QEs when compared to previous approaches. Furthermore, it has a very explicit real-world use case in constructing an optimal gallery set for face recognition tasks. In addition to this, we used these evaluation protocols to investigate the generalizability of face Quality Estimators (QEs) across various face recognition models, which also has implications for real-world use cases. Haoyu Chen 0006, Praneet Singh, Edward J. Delp, Amy R. Reibman |
MMSP | 3 |
| 2022 | Panchromatic Imagery Copy-Paste Localization Through Data-Driven Sensor AttributionabstractOverhead images can be obtained using different acquisition and processing techniques, and they are becoming more and more popular. As with common photographs, they can be forged and manipulated by malicious users. However, not all image forensics methods tailored to normal photos can be successfully applied out of the box to overhead images. In this paper we consider the problem of localizing copy-paste forgeries on panchromatic images acquired with different satellites. We leverage a set of Convolutional Neural Networks (CNNs) that extract traces of the acquisition satellite directly from image patches. We then determine whether an image region appears to have been acquired with a different satellite than the rest of the picture. Results show that the proposed technique outperforms more sophisticated image forensics tools tailoring common photographs. Edoardo Daniele Cannas, János Horváth, Sriram Baireddy, Paolo Bestagini, Edward J. Delp, Stefano Tubaro |
ICASSP | 5 |
| 2022 | Forensic Analysis and Localization of Multiply Compressed MP3 Audio Using TransformersabstractAudio signals are often stored and transmitted in compressed formats. Among the many available audio compression schemes, MPEG-1 Audio Layer III (MP3) is very popular and widely used. Since MP3 is lossy it leaves characteristic traces in the compressed audio which can be used forensically to expose the past history of an audio file. In this paper, we consider the scenario of audio signal manipulation done by temporal splicing of compressed and uncompressed audio signals. We propose a method to find the temporal location of the splices based on transformer networks. Our method identifies which temporal portions of a audio signal have undergone single or multiple compression at the temporal frame level, which is the smallest temporal unit of MP3 compression. We tested our method on a dataset of 486,743 MP3 audio clips. Our method achieved higher performance and demonstrated robustness with respect to different MP3 data when compared with existing methods. Ziyue Xiang, Paolo Bestagini, Stefano Tubaro, Edward J. Delp |
ICASSP | 4 |
| 2022 | Video-Analytics Task-Aware Quad-Tree Partitioning and Quantization for HEVCabstractVideo analytics systems designed for computer vision tasks use deep learning models that rely on high-quality input data to maximize performance. However, in a real-world system, these inputs are often compressed using video codecs such as HEVC. Video compression degrades the quality of the inputs, thereby degrading the performance of these models. Region-of-interest (ROI) coding enables bits to be allocated to improve performance; however, the method to select regions should be computationally simple since it must occur during or before the video is compressed and transmitted for further processing. In this paper, we propose a task-aware quad-tree (TA-QT) partitioning and quantization method to achieve ROI coding for HEVC and other video coding standards. TA-QT uses a lightweight edge-based model to guide task-aware video encoding to improve end-stage video analytics (ESVA) performance while reducing both bit-rate and encoding time. We demonstrate the effectiveness of our approach in terms of (a) the performance of the ESVA on compressed inputs, (b) transmission bit-rates, and (c) encoding time. Praneet Singh, Edward J. Delp, Amy R. Reibman |
ICIP | 2 |
| 2022 | 3D Centroidnet: Nuclei Centroid Detection with Vector Flow VotingabstractAutomated microscope systems are increasingly used to collect large-scale 3D image volumes of biological tissues. Since cell boundaries are seldom delineated in these images, detection of nuclei is a critical step for identifying and analyzing individual cells. Due to the large intra-class variability in nuclei morphology and the difficulty of generating ground truth annotations, accurate nuclei detection remains a challenging task. We propose a 3D nuclei centroid detection method by estimating the "vector flow" volume where each voxel represents a 3D vector pointing to its nearest nuclei centroid in the corresponding microscopy volume. We then use a voting mechanism to estimate the 3D nuclei centroids from the "vector flow" volume. Our system is trained on synthetic microscopy volumes and tested on real microscopy volumes. The evaluation results indicate our method outperforms other methods both visually and quantitatively. Liming Wu, Alain Chen, Paul Salama, Kenneth W. Dunn, Edward J. Delp |
ICIP | 5 |
| 2022 | Leaf Tar Spot Detection Using RGB ImagesabstractTar spot disease is a fungal disease that appears as a series of black circular spots containing spores on corn leaves. Tar spot has proven to be an impactful disease in terms of reducing crop yield. To quantify disease progression, experts usually have to visually phenotype leaves from the plant. This process is very time-consuming and difficult to incorporate into any high-throughput phenotyping system. Deep neural networks could provide quick, automated tar spot detection with sufficient ground truth. However, manually labeling tar spots in images to serve as ground truth is also tedious and time-consuming. In this paper we first describe an approach that uses automated image analysis tools to generate ground truth images that are then used for training a Mask R-CNN. We show that a Mask R-CNN can be used effectively to detect tar spots in close-up images of leaf surfaces. We additionally show that the Mask R-CNN can be used for in-field images of whole leaves to capture the number of tar spots and area of the leaf infected by the disease. Sriram Baireddy, Da-Young Lee, Carlos Gongora-Canul, Christian Cruz, Edward J. Delp |
ICMLA | 5 |
| 2022 | Transformer-Based Speech Synthesizer Attribution in an Open Set ScenarioabstractSpeech synthesis methods can create realistic-sounding speech, which may be used for fraud, spoofing, and mis-information campaigns. Forensic methods that detect synthesized speech are important for protection against such attacks. Forensic attribution methods provide even more information about the nature of synthesized speech signals because they identify the specific speech synthesis method (i.e., speech synthesizer) used to create a speech signal. Due to the increasing number of realistic-sounding speech synthesizers, we propose a speech attribution method that generalizes to new synthesizers not seen during training. To do so, we investigate speech synthesizer attribution in both a closed set scenario and an open set scenario. In other words, we consider some speech synthesizers to be "known" synthesizers (i.e., part of the closed set) and others to be "unknown" synthesizers (i.e., part of the open set). We represent speech signals as spectrograms and train our proposed method, known as compact attribution transformer (CAT), on the closed set for multi-class classification. Then, we extend our analysis to the open set to attribute synthesized speech signals to both known and unknown synthesizers. We utilize a t-distributed stochastic neighbor embedding (tSNE) on the latent space of the trained CAT to differentiate between each unknown synthesizer. Additionally, we explore poly-1 loss formulations to improve attribution results. Our proposed approach successfully attributes synthesized speech signals to their respective speech synthesizers in both closed and open set scenarios. Emily R. Bartusiak, Edward J. Delp |
ICMLA | 2 |
| 2022 | Label-Free Mammalian Cell Tracking Enhanced by Precomputed Velocity FieldsabstractLabel-free cell imaging, where the cell is not "labeled" or modified by fluorescent chemicals, is an important research area in the field of biology. It avoids altering the cell’s properties which typically happens in the process of chemical labeling. However, without the contrast enhancement from the label, the analysis of label-free imaging is more challenging than label-based imaging. In addition, it provides few human interpretable features, and thus needs machine learning approaches to help with the identification and tracking of specific cells. We are interested in label-free phase contrast imaging to track cells flowing in a cell sorting device where images are acquired at 500 frames/s. Existing Multiple Object Tracking (MOT) methods face four major challenges when used for tracking cells in a microfluidic sorting device: (i) most of the cells have large displacements between frames without any overlap; (ii) it is difficult to distinguish between cells as they are visually similar to each other; (iii) the velocities of cells vary with the location in the device; (iv) the appearance of cells may change as they move in and out of the focal plane of the imaging sensor that observes the isolation process. In this paper, we introduce a method for tracking cells in a predefined flow in the sorting device via phase contrast microscopy. Our proposed method is based on DeepSORT and YOLOv4 and exploits prior knowledge of a cell’s velocity to assist tracking. We modify the Kalman filter in DeepSORT to accommodate a non-constant velocity motion model and integrate a representative velocity field obtained from fluid dynamics into the Kalman filter. The experimental results show that our proposed method outperforms several MOT methods for tracking cells in the sorting device. Viktor Shkolnikov, Daisy Xin, Steven Barcelo, Jan P. Allebach, Edward J. Delp |
ICMLA | 7 |
| 2022 | Infrared Small Target Detection Enhancement Using a Lightweight Convolutional Neural NetworkabstractDetection of small, point targets is fundamental in applications such as early warning systems, surveillance, astronomy, and microscopy. The presence of noise and clutter can make it challenging to detect small targets while minimizing false detections. This paper presents a method for infrared small target detection using convolutional neural networks. The proposed method augments a conventional space-based detection processing chain with a lightweight neural network to predict the probability that a detection is a target. The proposed network is trained on 7 × 7 pixel windows using both the image sequence and the respective background-subtracted images. Results show that our method improves probability of detection at low false detection rates. Mitchell Krouss, Greg Furlich, Paul Martens, Moses W. Chan, Mary L. Comer, Edward J. Delp |
IEEE Geosci. Remote. Sens. Lett. | 8 |
| 2021 | Improving Building Segmentation Using Uncertainty Modeling and Metadata InjectionabstractAutomatic building segmentation is an important task for satellite imagery analysis and scene understanding. Most existing segmentation methods focus on the case where the images are taken from directly overhead (i.e., low off-nadir/viewing angle). These methods often fail to provide accurate results on satellite images with larger off-nadir angles due to the higher noise level and lower spatial resolution. In this paper, we propose a method that is able to provide accurate building segmentation for satellite imagery captured from a large range of off-nadir angles. Based on Bayesian deep learning, we explicitly design our method to learn the data noise via aleatoric and epistemic uncertainty modeling. Satellite image metadata (e.g., off-nadir angle and ground sample distance) is also used in our model to further improve the result. We show that with uncertainty modeling and metadata injection, our method achieves better performance than the baseline method, especially for noisy images taken from large off-nadir angles1. Hanxiang Hao, Sriram Baireddy, Kevin J. LaTourette, Latisha Konz, Moses W. Chan, Mary L. Comer, Edward J. Delp |
SIGSPATIAL/GIS | 7 |
| 2021 | Open-Set Source Attribution for Panchromatic Satellite ImageryabstractIn the last few years, several companies started offering the possibility of buying different kinds of overhead images acquired by satellites orbiting around the planet. This market is interesting for several customers, from those who simply fancy a shot of their house from space, to those aiming to acquire strategic information on portions of land. Due to the sensitive nature of this data, which can be maliciously altered by anyone, the forensic community has started investigating methodologies to verify overhead imagery authenticity and integrity. Within this context, in this paper we investigate the possibility of using Convolutional Neural Networks (CNNs) to attribute a panchromatic satellite image to the satellite used to acquire it. In our investigation we tackle both closed-set and, adapting Deep Ensemble (DE) and Monte Carlo Dropout (MCD) techniques, open-set image attribution problems. Edoardo Daniele Cannas, Sriram Baireddy, Emily R. Bartusiak, Sri Yarlagadda, Daniel Mas Montserrat, Paolo Bestagini, Stefano Tubaro, Edward J. Delp |
ICIP | 8 |
| 2021 | Panicle Counting in UAV Images for Estimating Flowering Time in SorghumabstractFlowering time (time to flower after planting) is important for estimating plant development and grain yield for many crops including sorghum. Flowering time of sorghum can be approximated by counting the number of panicles (clusters of grains on a branch) across multiple dates. Traditional manual methods for panicle counting are time-consuming and tedious. In this paper, we propose a method for estimating flowering time and rapidly counting panicles using RGB images acquired by an Unmanned Aerial Vehicle (UAV). We evaluate three different deep neural network structures for panicle counting and location. Experimental results demonstrate that our method is able to accurately detect panicles and estimate sorghum flowering time. Enyu Cai, Sriram Baireddy, Changye Yang, Edward J. Delp, Melba M. Crawford |
IGARSS | 4 |
| 2021 | An Attention-Based System for Damage Assessment Using Satellite ImageryabstractWhen a disaster strikes, accurate situational information and a fast, effective response are critical to save lives. High resolution satellite images enable emergency responders to estimate the location, cause, and severity of damage. In this paper, we present a Siam-U-Net-Attn model with an attention mechanism to assess the damage level of buildings given a pair of satellite images showing a scene before and after a disaster. We evaluate the proposed method on xView2, a building damage assessment dataset, and demonstrate that the proposed approach achieves accurate damage scale classification and building segmentation. Hanxiang Hao, Sriram Baireddy, Emily R. Bartusiak, Latisha Konz, Kevin J. LaTourette, Michael Gribbons, Moses W. Chan, Edward J. Delp, Mary L. Comer |
IGARSS | 8 |
| 2021 | Estimating Image Quality for Person Re-IdentificationabstractTask-based image quality is an important component in designing a real-life analytics system, and has been studied in many fields such as biometric recognition and pedestrian detection. Person re-identification (re-id) is an application in surveillance where one attempts to identify a person after they leave the surveillance system and then re-enter it. As a newer field in recognition applications, re-id lacks image quality-related research. In this paper, we propose an unsupervised and automated quality measure for the query images used in re-id, which we call "identifiability". Images that are less identifiable are more challenging for any re-identification system, creating output results that would likely be less reliable. Our proposed method measures feature consistency in the presence of perturbations as an indicator of identifiability. We then introduce two evaluation protocols for such a quality measure. We demonstrate our proposed quality measure is effective at ranking an image’s usefulness to a recognition system, and at identifying an image’s robustness against further compression. Haoyu Chen 0006, Edward J. Delp, Amy R. Reibman |
MMSP | 2 |
| 2020 | Robustness Analysis of Face ObscurationabstractFace obscuration is needed by law enforcement and mass media outlets to guarantee privacy. Sharing sensitive content where obscuration or redaction techniques have failed to completely remove all identifiable traces can lead to many legal and social issues. Hence, we need to be able to systematically measure the face obscuration performance of a given technique. In this paper we propose to measure the effectiveness of eight obscuration techniques. We do so by attacking the redacted faces in three scenarios: obscured face identification, verification, and reconstruction. Threat modeling is also considered to provide a vulnerability analysis for each studied obscuration technique. Based on our evaluation, we show that the k-same based methods are the most effective. Hanxiang Hao, David Guera, János Horváth, Amy R. Reibman, Edward J. Delp |
FG | 5 |
| 2020 | A Weakly Supervised Deep Learning Approach for Plant Center Detection and CountingabstractDeep learning applications are rapidly advancing in agriculture, and phenotyping in particular. Identifying the locations of plant centers and counting are critical for assessing stand quality. Yet, its precise measurement after the emergence of plants is impractical in large-scale production fields due to the labor required. In this paper, we propose a weakly supervised deep learning framework for detecting the centers and counting maize plants using high resolution georeferenced RGB data acquired from a UAV platform. We evaluate the performance of the proposed method over two dates. The obtained results show that the proposed method can efficiently identify the locations of plant centers, and thereby the plant counts, reducing both the manual field counting and human labeling required for image-based analysis. Azam Karami, Melba M. Crawford, Edward J. Delp |
IGARSS | 3 |
| 2019 | Locating Objects Without Bounding BoxesabstractRecent advances in convolutional neural networks (CNN) have achieved remarkable results in locating objects in images. In these networks, the training procedure usually requires providing bounding boxes or the maximum number of expected objects. In this paper, we address the task of estimating object locations without annotated bounding boxes which are typically hand-drawn and time consuming to label. We propose a loss function that can be used in any fully convolutional network (FCN) to estimate object locations. This loss function is a modification of the average Hausdorff distance between two unordered sets of points. The proposed method has no notion of bounding boxes, region proposals, or sliding windows. We evaluate our method with three datasets designed to locate people's heads, pupil centers and plant centers. We outperform state-of-the-art generic object detectors and methods fine-tuned for pupil tracking. Javier Ribera, David Guera, Yuhao Chen 0001, Edward J. Delp |
CVPR | 4 |
| 2019 | Shadow Removal Detection and Localization for Forensics AnalysisabstractThe recent advancements in image processing and computer vision allow realistic photo manipulations. In order to avoid the distribution of fake imagery, the image forensics community is working towards the development of image authenticity verification tools. Methods based on shadow analysis are particularly reliable since they are part of the physical integrity of the scene, thus detecting forgeries is possible whenever inconsistencies are found (e.g., shadows not coherent with the light direction). An attacker can easily delete inconsistent shadows and replace them with correctly cast shadows in order to fool forensics detectors based on physical analysis. In this paper, we propose a method to detect shadow removal done with state-of-the-art tools. The proposed method is based on a conditional generative adversarial network (cGAN) specifically trained for shadow removal detection. Sri Yarlagadda, David Guera, Daniel Mas Montserrat, Fengqing Zhu 0001, Edward J. Delp, Paolo Bestagini, Stefano Tubaro |
ICASSP | 5 |
| 2019 | Image Anonymization Detection with Deep Handcrafted FeaturesabstractIn recent years, the number of images shared online has continuously grown. The forensics community has kept the pace by developing techniques to both reliably extract information from these images, but also to remove it. In particular, the latest developments in image anonymization methods exposes an attack vector when used by skilled ill-intentioned image producers that may want to elude prosecution. We present an approach to detect whether or not an image has undergone a laundering process, i.e., it has been tampered with so that its unique characterizing features have been changed to avoid detection. We focus on the photo response non uniformity (PRNU) noise unique to every imaging sensor, and we consider that an image has been "laundered" when we detect the absence of PRNU from an image. We propose a per image preprocessing pipeline that generates information-rich features later used as input of fine-tuned convolutional neural networks (CNNs). We study the performance of the proposed approach using various CNN architectures and blind anonymization techniques and show its effectiveness under several training and testing scenarios. Our results also show that CNN models trained with the proposed feature are capable of generalizing over unseen devices and are robust against non-geometric transformations. Nicolò Bonettini, David Guera, Luca Bondi, Paolo Bestagini, Edward J. Delp, Stefano Tubaro |
ICIP | 5 |
| 2018 | Deepfake Video Detection Using Recurrent Neural NetworksabstractIn recent months a machine learning based free software tool has made it easy to create believable face swaps in videos that leaves few traces of manipulation, in what are known as "deepfake" videos. Scenarios where these realistic fake videos are used to create political distress, blackmail someone or fake terrorism events are easily envisioned. This paper proposes a temporal-aware pipeline to automatically detect deepfake videos. Our system uses a convolutional neural network (CNN) to extract frame-level features. These features are then used to train a recurrent neural network (RNN) that learns to classify if a video has been subject to manipulation or not. We evaluate our method against a large set of deepfake videos collected from multiple video websites. We show how our system can achieve competitive results in this task while using a simple architecture. David Guera, Edward J. Delp |
AVSS | 2 |
| 2018 | Single-View Food Portion Estimation: Learning Image-to-Energy Mappings Using Generative Adversarial NetworksabstractDue to the growing concern of chronic diseases and other health problems related to diet, there is a need to develop accurate methods to estimate an individual's food and energy intake. Measuring accurate dietary intake is an open research problem. In particular, accurate food portion estimation is challenging since the process of food preparation and consumption impose large variations on food shapes and appearances. In this paper, we present a food portion estimation method to estimate food energy (kilocalories) from food images using Generative Adversarial Networks (GAN). We introduce the concept of an “energy distribution” for each food image. To train the GAN, we design a food image dataset based on ground truth food labels and segmentation masks for each food image as well as energy information associated with the food image. Our goal is to learn the mapping of the food image to the food energy. We can then estimate food energy based on the energy distribution. We show that an average energy estimation error rate of 10.89% can be obtained by learning the image-to-energy mapping. Shaobo Fang, Zeman Shao, Runyu Mao, Chichen Fu, Edward J. Delp, Fengqing Zhu 0001, Deborah A. Kerr, Carol J. Boushey |
ICIP | 5 |
| 2018 | Reliability Map Estimation for CNN-Based Camera Model AttributionabstractAmong the image forensic issues investigated in the last few years, great attention has been devoted to blind camera model attribution. This refers to the problem of detecting which camera model has been used to acquire an image by only exploiting pixel information. Solving this problem has great impact on image integrity assessment as well as on authenticity verification. Recent advancements that use convolutional neural networks (CNNs) in the media forensic field have enabled camera model attribution methods to work well even on small image patches. These improvements are also important for determining forgery localization. Some patches of an image may not contain enough information related to the camera model (e.g., saturated patches). In this paper, we propose a CNN-based solution to estimate the camera model attribution reliability of a given image patch. We show that we can estimate a reliabilitymap indicating which portions of the image contain reliable camera traces. Testing using a well known dataset confirms that by using this information, it is possible to increase small patch camera model attribution accuracy by more than 8% on a single patch. David Guera, Fengqing Zhu 0001, Sri Yarlagadda, Stefano Tubaro, Paolo Bestagini, Edward J. Delp |
WACV | 6 |
| 2018 | Digital watermarking for camera-captured images based on just noticeable distortion and Wiener filtering
Kharittha Thongkor, Thumrongrat Amornraksa, Edward J. Delp |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Context based image analysis with application in dietary assessment and evaluation
Yu Wang 0034, Ye He 0001, Carol J. Boushey, Fengqing Zhu 0001, Edward J. Delp |
Multim. Tools Appl. | 5 |
| 2017 | A Two Stream Siamese Convolutional Neural Network for Person Re-identificationabstractPerson re-identification is an important task in video surveillance systems. It can be formally defined as establishing the correspondence between images of a person taken from different cameras at different times. In this paper, we present a two stream convolutional neural network where each stream is a Siamese network. This architecture can learn spatial and temporal information separately. We also propose a weighted two stream training objective function which combines the Siamese cost of the spatial and temporal streams with the objective of predicting a person's identity. We evaluate our proposed method on the publicly available PRID2011 and iLIDS-VID datasets and demonstrate the efficacy of our proposed method. On average, the top rank matching accuracy is 4% higher than the accuracy achieved by the cross-view quadratic discriminant analysis used in combination with the hierarchical Gaussian descriptor (GOG+XQDA), and 5% higher than the recurrent neural network method. Dahjung Chung, Khalid Tahboub, Edward J. Delp |
ICCV | 3 |
| 2017 | Plant leaf segmentation for estimating phenotypic traitsabstractIn this paper we propose a method to segment individual leaves of crop plants from Unmanned Aerial Vehicle (UAV) imagery for the purposes of deriving phenotypic properties of the plant. The crop plant used in our study is sorghum [Sorghum bicolor (L.) Moench]. Phenotyping is a set of methodologies for analyzing and obtaining characteristic traits of a plant. In a phenotypic study, leaves are often used to estimate traits such as individual leaf area and Leaf Area Index (LAI). Our approach is to segment the leaves in polar coordinates using the plant center as the origin. The shape of each leaf is estimated by a shape model. Experimental results indicate that this approach can provide good estimates of leaf phenotypic properties. Yuhao Chen 0001, Javier Ribera, Christopher Boomsma, Edward J. Delp |
ICIP | 4 |
| 2017 | Spatial pyramid alignment for sparse coding based object classificationabstractThe bag of visual words (BOW) model is widely used for image representation and classification. Spatial pyramid based feature pooling utilizes the BOW model and is the most popular approach to capture the spatial distribution (layout) of local image features. It makes the assumption that the center of an object is aligned with the center of an image, which can lead to misalignment and degradation in performance. In this paper, we propose a method to utilize max pooled features to estimate objects centers and align the spatial pyramid accordingly. We also propose an image representation descriptor robust to misalignments and objects deformations. The experimental results demonstrate that our spatial pyramid alignment method is simple yet efficient in handling misalignments and achieves high object classification accuracy. Joonsoo Kim, Khalid Tahboub, Edward J. Delp |
ICIP | 3 |
| 2017 | Quality-adaptive deep learning for pedestrian detectionabstractPedestrian detection is a fundamental task for many applications including autonomous vehicles and surveillance systems. In a mobile or networked environment bandwidth is limited and adaptive datarate streaming is used. Video compression can introduce significant quality degradation that impacts the accuracy of video analytics. In this paper, we examine the problem of a changing video data-rate and examine how it affects the performance of video analytics, in particular pedestrian detection, using a two-stage quality-adaptive convolutional neural network system. Our experimental results demonstrate that when adaptive data-rate streaming is used, our proposed quality-adaptive approach reduces the miss rate by 20% compared to the baseline detector. Khalid Tahboub, David Guera, Amy R. Reibman, Edward J. Delp |
ICIP | 4 |
| 2017 | Accuracy prediction for pedestrian detectionabstractIn this paper, we address the problem of predicting accuracy for pedestrian detection. We want to be able to predict the accuracy of a video analytic method without actually executing the method. We propose the use of texture descriptors and random forests to predict the accuracy of various pedestrian detection methods. Our experimental results demonstrate that using the local binary pattern (LBP) or a bank of Schmid and Gabor filters can capture spatial textural information associated with video quality degradation that can be used to predict accuracy. We also demonstrate how predicting the absolute accuracy can save network and computational resources. Khalid Tahboub, Amy R. Reibman, Edward J. Delp |
ICIP | 3 |
| 2017 | Weakly supervised food image segmentation using class activation mapsabstractFood image segmentation plays a crucial role in image-based dietary assessment and management. Successful methods for object segmentation generally rely on a large amount of labeled data on the pixel level. However, such training data are not yet available for food images and expensive to obtain. In this paper, we describe a weakly supervised convolutional neural network (CNN) which only requires image level annotation. We propose a graph based segmentation method which uses the class activation maps trained on food datasets as a top-down saliency model. We evaluate the proposed method for both classification and segmentation tasks. We achieve competitive classification accuracy compared to the previously reported results. Yu Wang 0034, Fengqing Zhu 0001, Carol J. Boushey, Edward J. Delp |
ICIP | 4 |
| 2017 | First Steps Toward Camera Model Identification With Convolutional Neural NetworksabstractDetecting the camera model used to shoot a picture enables to solve a wide series of forensic problems, from copyright infringement to ownership attribution. For this reason, the forensic community has developed a set of camera model identification algorithms that exploit characteristic traces left on acquired images by the processing pipelines specific of each camera model. In this letter, we investigate a novel approach to solve camera model identification problem. Specifically, we propose a data-driven algorithm based on convolutional neural networks, which learns features characterizing each camera model directly from the acquired pictures. Results on a well-known dataset of 18 camera models show that: 1) the proposed method outperforms up-to-date state-of-the-art algorithms on classification of 64 × 64 color image patches; 2) features learned by the proposed network generalize to camera models never used for training. Luca Bondi, Luca Baroffio, David Guera, Paolo Bestagini, Edward J. Delp, Stefano Tubaro |
IEEE Signal Process. Lett. | 5 |
| 2017 | Neighborhood Matching for Image RetrievalabstractIn the last few years, large-scale image retrieval has attracted a lot of attention from the multimedia community. Usual approaches addressing this task first generate an initial ranking of the reference images using fast approximations that do not take into consideration the spatial arrangement of local features in the image (e.g., the bag-of-words paradigm). The top positions of the rankings are then re-estimated with verification methods that deal with more complex information, such as the geometric layout of the image. This verification step allows pruning of many false positives at the expense of an increase in the computational complexity, which may prevent its application to large-scale retrieval problems. This paper describes a geometric method known as neighborhood matching (NM), which revisits the keypoint matching process by considering a neighborhood around each keypoint and improves the efficiency of a geometric verification step in the image search system. Multiple strategies are proposed and compared to incorporate NM into a large-scale image retrieval framework. A detailed analysis and comparison of these strategies and baseline methods have been investigated. The experiments show that the proposed method not only improves the computational efficiency, but also increases the retrieval performance and outperforms state-of-the-art methods in standard datasets, such as the Oxford 5 k and 105 k datasets, for which the spatial verification step has a significant impact on the system performance. Iván González-Díaz 0001, Murat Birinci, Fernando Díaz-de-María, Edward J. Delp |
IEEE Trans. Multim. | 4 |
| 2016 | A comparison of food portion size estimation using geometric models and depth imagesabstractSix of the ten leading causes of death in the United States, including cancer, diabetes, and heart disease, can be directly linked to diet. Dietary intake, the process of determining what someone eats during the course of a day, provides valuable insights for mounting intervention programs for prevention of many of the above chronic diseases. Measuring accurate dietary intake is considered to be an open research problem in the nutrition and health fields. In this paper we compare two techniques to estimating food portion size from images of food. The techniques are based on 3D geometric models and depth images. An expectation-maximization based technique is developed to detect the reference plane in depth images, which is essential for portion size estimation using depth images. Our experimental results indicate that volume estimation based on geometric model is more accurate for objects with well-defined 3D shapes compared to estimation using depth images. Shaobo Fang, Fengqing Zhu 0001, Chufan Jiang, Song Zhang 0002, Carol J. Boushey, Edward J. Delp |
ICIP | 6 |
| 2016 | Shape matching using a self similar affine invariant descriptorabstractIn this paper we introduce a shape descriptor known as Self Similar Affine Invariant (SSAI) descriptor for shape retrieval. The SSAI descriptor is based on the property that two sets of points are transformed by an affine transform, then subsets of each set of points are also related by the same affine transformation. Also, the SSAI descriptor is insensitive to local shape distortions. We use multiple SSAI descriptors based on different sets of neighbor points to improve shape recognition accuracy. We also describe an efficient image matching method for the multiple SSAI descriptors. Experimental results show that our approach achieves very good performance on two publicly available shape datasets. Joonsoo Kim, He Li 0002, Jiaju Yue, Edward J. Delp |
ICIP | 4 |
| 2016 | Person re-identification using a patch-based appearance modelabstractPerson re-identification is the process of recognizing a person across a network of cameras with non-overlapping fields of view. In this paper we present an unsupervised multi-shot approach based on a patch-based dynamic appearance model. We use deformable graph matching for person re-identification using histograms of color and texture as features of nodes. Each graph model spans multiple images and each node is a local patch in the shape of a rectangle. We evaluate our proposed method on publicly available PRID 2011 and iLIDS-VID databases. Khalid Tahboub, Blanca Delgado, Edward J. Delp |
ICIP | 3 |
| 2016 | Efficient superpixel based segmentation for food image analysisabstractIn this paper, we propose a segmentation method based on normalized cut and superpixels. The method relies on color and texture cues for fast computation and efficient use of memory. The method is used for food image segmentation as part of a mobile food record system we have developed for dietary assessment and management. The accurate estimate of nutrients relies on correctly labelled food items and sufficiently well-segmented regions. Our method achieves competitive results using the Berkeley Segmentation Dataset and outperforms some of the most popular techniques in a food image dataset. Yu Wang 0034, Chang Liu 0004, Fengqing Zhu 0001, Carol J. Boushey, Edward J. Delp |
ICIP | 5 |
| 2016 | VPx video coding for lossy transmission channels using error resilience packetsabstractMany popular video coding standards such as VPx, H.26x achieve video compression by using spatial and temporal dependencies in the source video signal. This makes the encoded bitstream vulnerable to errors during transmission. In this paper, we investigate an error resilient video coding for the VPx bitstreams using error resilience packets. An error resilient packet consists of encoded keyframe contents and the prediction signals for each non-keyframe. Experimental results exhibit that our proposed method is effective under typical packet loss conditions. Neeraj Gadgil, Edward J. Delp |
PCS | 3 |
| 2016 | Adaptive error concealment for temporal-spatial multiple description video coding
Meilin Yang, Neeraj Gadgil, Mary L. Comer, Edward J. Delp |
Signal Process. Image Commun. | 4 |
| 2016 | Image-Like 2D Barcodes Using Generalizations of the Kuznetsov-Tsybakov ProblemabstractIn this paper, we propose a novel method for generating visually appealing two-dimensional (2D) barcodes that resemble meaningful images to human observers. The technology of 2D barcodes, currently dominated by quick response codes, is widely adopted in many applications, including product tracking, document management, and general marketing. Such barcodes typically lack user friendly appearance and do not convey any visual significance to human observers. The proposed method addresses this problem by allowing 2D barcodes to resemble an arbitrary image or a logo. Our method is based on a generalization of the Kuznetsov-Tsybakov problem that served as a foundation for wet paper codes, commonly adopted in digital steganography. We introduce weaker statistical constraints to obtain additional flexibility allowing the barcode to assume the appearance of an arbitrary pattern. This paper provides the theoretical analysis of the proposed coding framework and a practical algorithm for rapid approximation of the optimal code. We also discuss the introduction of error correction capabilities, and experimentally evaluate a prototype implementation in a smartphone-based acquisition scenario. Jaroslaw Duda 0001, Pawel Korus, Neeraj Gadgil, Khalid Tahboub, Edward J. Delp |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2015 | Robust local and global shape context for tattoo image matchingabstractTattoos can provide useful information related to criminal gang activity. Law enforcement can use the information embedded in tattoos to identify and track the criminal history of a suspect. For matching processes, tattoo images are difficult to use due to problems such as deformations and weak edge structures. In this paper we describe a tattoo image retrieval and matching system based on a combination of local and global image matching methods to improve matching accuracy. The proposed local shape context combined with SIFT descriptors are used for local features of a tattoo object and global shape is used for overall shape of a tattoo object. The contributions of this paper include the introduction of a multiple different sized-bin polar histograms based local shape context (MHLC) and a global shape descriptor combining the multiple different sized-bin polar histogram and 2D Fourier Transform for robustness of translation, scale, rotation and shape distortions. We also describe robust similarity for local descriptors and a weighted matching method based on local and global descriptors. Our experimental results show that our proposed method performs better than previously published tattoo image retrieval systems. Joonsoo Kim, Albert Parra Pozo, Jiaju Yue, He Li 0002, Edward J. Delp |
ICIP | 5 |
| 2015 | Content based video retrieval on mobile devices: How much content is enough?abstractIn the field of content based video retrieval, recent work has focused on creation and matching of signatures as an effective method for video search, copy or “near-duplicate” detection. In this paper, we address the question of how much content is enough to initiate a content based video retrieval query from a mobile device. We extract and utilize the motion vectors from an HEVC-encoded bitstream and use them as the content-dependent features to create video signatures. We propose an energy function to quantify the mount of motion information contained in video. Our experimental results exhibit that our proposed approach is able to generate a signature robust to common signal processing techniques and that our proposed energy function can be efficiently used to select the length of a video clip used to query the video database from a mobile device. Khalid Tahboub, Neeraj Gadgil, Edward J. Delp |
ICIP | 3 |
| 2015 | Single-View Food Portion Estimation Based on Geometric ModelsabstractIn this paper we present a food portion estimation technique based on a single-view food image used for the estimation of the amount of energy (in kilocalories) consumed at a meal. Unlike previous methods we have developed, the new technique is capable of estimating food portion without manual tuning of parameters. Although single-view 3D scene reconstruction is in general an ill-posed problem, the use of geometric models such as the shape of a container can help to partially recover 3D parameters of food items in the scene. Based on the estimated 3D parameters of each food item and a reference object in the scene, the volume of each food item in the image can be determined. The weight of each food can then be estimated using the density of the food item. We were able to achieve an error of less than 6% for energy estimation of an image of a meal assuming accurate segmentation and food classification. Shaobo Fang, Chang Liu 0004, Fengqing Zhu 0001, Edward J. Delp, Carol J. Boushey |
ISM | 4 |
| 2015 | Report on the evaluation of current and future image compression technologiesabstractThis document reports the conclusions, comments and recommendations resulting from the final panel discussion which happened at the Session on Evaluation of Current and Future Image Compression Technologies held at the Picture Coding Symposium, Cairns, Australia, in June 2015. Fernando Pereira 0001, Ralf Schaefer, Touradj Ebrahimi, Jörn Ostermann, Edward J. Delp |
PCS | 5 |
| 2015 | The use of asymmetric numeral systems as an accurate replacement for Huffman codingabstractEntropy coding is an integral part of most data compression systems. Huffman coding (HC) and arithmetic coding (AC) are two of the most widely used coding methods. HC can process a large symbol alphabet at each step allowing for fast encoding and decoding. However, HC typically provides suboptimal data rates due to its inherent approximation of symbol probabilities to powers of 1 over 2. In contrast, AC uses nearly accurate symbol probabilities, hence generally providing better compression ratios. However, AC relies on relatively slow arithmetic operations making the implementation computationally demanding. In this paper we discuss asymmetric numeral systems (ANS) as a new approach to entropy coding. While maintaining theoretical connections with AC, the proposed ANS-based coding can be implemented with much less computational complexity. While AC operates on a state defined by two numbers specifying a range, an ANS-based coder operates on a state defined by a single natural number such that the x ∈ ℕ state contains ≈ log2(x) bits of information. This property allows to have the entire behavior for a large alphabet summarized in the form of a relatively small table (e.g. a few kilobytes for a 256 size alphabet). The proposed approach can be interpreted as an equivalent to adding fractional bits to a Huffman coder to combine the speed of HC and the accuracy offered by AC. Additionally, ANS can simultaneously encrypt a message encoded this way. Experimental results demonstrate effectiveness of the proposed entropy coder. Jaroslaw Duda 0001, Khalid Tahboub, Neeraj Gadgil, Edward J. Delp |
PCS | 4 |
| 2015 | Spatial subsampling-based multiple description video coding with adaptive temporal-spatial error concealmentabstractThe encoded bitstream of a typical video compression method is vulnerable to errors during transmission over lossy network. Multiple Description Coding (MDC) is an error resilient video coding suitable for real-time applications when retransmission is unacceptable. Subsampling-based MDC with error concealment is often effective for mitigating unpredictable losses. In this paper, a spatial subsampling based MDC with an adaptive concealment method is described. Error concealment is done using the motion information from its spatial-temporal neighbors. Preliminary experimental results exhibit the efficacy of the proposed method. Neeraj Gadgil, He Li 0002, Edward J. Delp |
PCS | 3 |
| 2015 | Visual Saliency Models Based on Spectrum ProcessingabstractSome visual saliency models have been proposed to describe how the human visual system perceives and processes visual information. In this paper we describe four frequency domain visual saliency models based on new spectrum processing methods. The four saliency models are the Gamma Corrected Spectrum (GCS) model, the Gamma Corrected Log Spectrum (GCLS) model, the Gaussian Filtered Spectrum (GFS) model, and the Gaussian Filtered Log Spectrum (GFLS) model. A set of saliency map candidates are generated by inverse transform of a set of modified spectrums. An output saliency map is selected by minimizing the entropy among the set of saliency map candidates. Extension of these models are also described using various color spaces. Experimental results show that four extensions of our GCS, GCLS, GFS, and GFLS models are more accurate and efficient than some state-of-the-art saliency models in predicting eye fixation on standard image datasets. Bin Zhao 0003, Edward J. Delp |
WACV | 2 |
| 2015 | Color difference weighted adaptive residual preprocessing using perceptual modeling for video compression
Mark Q. Shaw, Jan P. Allebach, Edward J. Delp |
Signal Process. Image Commun. | 3 |
| 2015 | Multiple Hypotheses Image Segmentation and Classification With Application to Dietary AssessmentabstractWe propose a method for dietary assessment to automatically identify and locate food in a variety of images captured during controlled and natural eating events. Two concepts are combined to achieve this: a set of segmented objects can be partitioned into perceptually similar object classes based on global and local features; and perceptually similar object classes can be used to assess the accuracy of image segmentation. These ideas are implemented by generating multiple segmentations of an image to select stable segmentations based on the classifier's confidence score assigned to each segmented image region. Automatic segmented regions are classified using a multichannel feature classification system. For each segmented region, multiple feature spaces are formed. Feature vectors in each of the feature spaces are individually classified. The final decision is obtained by combining class decisions from individual feature spaces using decision rules. We show improved accuracy of segmenting food images with classifier feedback. Fengqing Zhu 0001, Marc Bosch, Nitin Khanna, Carol J. Boushey, Edward J. Delp |
IEEE J. Biomed. Health Informatics | 5 |
| 2014 | Generalizations of the Kuznetsov-Tsybakov problem for generating image-like 2D barcodesabstractMany two-dimensional (2D) barcodes, such as quick response (QR) codes, lack user-friendly appearance. Our goal in this paper is to generate 2D barcodes that “look” like recognizable images or logos. Standard steganographic methods hide a message (payload) in an image usually by modifying bits in a specific way using predetermined pixels of the image. This approach cannot be directly used for very low bit rates commonly used in 1 bit per pixel 2D barcodes. It is possible to produce barcodes in which the grayness of a pixel in an image is interpreted as the probability of assigning a value (black or white) to the corresponding pixel of the encoded message (payload). This can be viewed as statistical constraints enforced on the encoded bit-sequence. Using an information theoretic approach, Kuznetsov and Tsybakov have shown that this can be done for a specific case of constraints almost without any loss of capacity. In this paper, we propose generalizations of this approach with weaker constraints as an application to generating 2D barcodes that resemble images. We describe a coding framework, various types of constraints, a practical approximation and some example 2D barcodes generated from our implementation. Jaroslaw Duda 0001, Neeraj Gadgil, Khalid Tahboub, Edward J. Delp |
ICIP | 4 |
| 2014 | Analysis of food images: Features and classificationabstractIn this paper we investigate features and their combinations for food image analysis and a classification approach based on k-nearest neighbors and vocabulary trees. The system is evaluated on a food image dataset consisting of 1453 images of eating occasions in 42 food categories which were acquired by 45 participants in natural eating conditions. The same image dataset is used to test the classification system proposed in the previously reported work [1]. Experimental results indicate that using our combination of features and vocabulary trees for classification improves the food classification performance about 22% for the Top 1 classification accuracy and 10% for the Top 4 classification accuracy. Ye He 0001, Chang Xu 0006, Nitin Khanna, Carol J. Boushey, Edward J. Delp |
ICIP | 5 |
| 2014 | Foundational metadata for image based cognitionabstractThe need to create useful information from full motion video gathered by drones is a significant motivation for devising methods to approximate human cognitive behaviors. Additionally, the regulatory needs associated with drone systems has spawned the requirement to be able to confirm, or audit, the activities of such devices. A conditional approach, as compared with a generalized video processing environment, is presented that associates practical and realistic constraints to simplify the problem of finding useful information from video acquired by a drone into something that is tractable and consistent with real-world requirements. A primary contribution of this paper is to introduce the concept of continuous cognition from a theoretical perspective, followed by a practical application derived from an operational system. Gary L. Viviani, Edward J. Delp |
ICIP | 2 |
| 2013 | Context based food image analysisabstractWe are developing a dietary assessment system that records daily food intake through the use of food images. Recognizing food in an image is difficult due to large visual variance with respect to eating or preparation conditions. This task becomes even more challenging when different foods have similar visual appearance. In this paper we propose to incorporate two types of contextual dietary information, food co-occurrence patterns and personalized learning models, in food image analysis to reduce ambiguity in food visual appearance and improve food recognition accuracy. We evaluate our model on 1453 food images acquired by 45 participants in natural eating conditions. The result shows that incorporating contextual dietary information improves the food categorization accuracy by about 10%. Ye He 0001, Chang Xu 0006, Nitin Khanna, Carol J. Boushey, Edward J. Delp |
ICIP | 5 |
| 2013 | Three dimensional segmentation of fluorescence microscopy images using active surfacesabstractThree dimensional image volumes collected using optical microscopy exhibit many characteristics that cause difficulty in segmentation. These include decreasing image contrast with increasing tissue depth, poor edge detail with regard to cellular structures, and limited spatial resolution. This paper describes a three dimensional segmentation method utilizing active surfaces to segment microscopy image volumes. We demonstrate this method on a three dimensional sequence of images acquired from stationary/stabilized kidney tissue of a rat. Results from this method are compared against a prior pseudo-three dimensional segmentation method that analyzes slices on a single image-by-image basis, as well as against another native three dimensional segmentation method. Kevin S. Lorenz, Paul Salama, Kenneth W. Dunn, Edward J. Delp |
ICIP | 4 |
| 2013 | Hazardous material sign detection and recognitionabstractIn this paper we describe two methods for hazardous material (hazmat) sign recognition. The first method is based on segment detection and grouping using geometric constraints. The second method is based on the use of a saliency map and convex quadrilateral detection. Our experimental results show a detection accuracy of 57.7% on a set of hazmat signs taken in the field under various lightning conditions, distances, and perspectives. Albert Parra Pozo, Bin Zhao 0003, Andrew W. Haddad, Mireille Boutin, Edward J. Delp |
ICIP | 5 |
| 2013 | Model-based food volume estimation using 3D poseabstractWe are developing a dietary assessment system to automatically identify and quantify foods and beverages consumed by analyzing meal images captured with a mobile device. After food items are segmented and identified, accurately estimating the volume of the food in the image is important for determining the nutrient content of the food. In this paper, we proposed a novel food portion size estimation method for rigid food items using a single image. First, we create a 3D graphical model during the training step using 3D reconstruction from multiple views. Then, for each food image, we determine the translation and elevation parameters of each of the food items, which are relative to the camera coordinate through camera calibration. Using these geometric parameters we project the pre-built 3D model of each food item back to the image plane. Subsequently, the remaining degrees-of-freedom (DOF) for the final pose is estimated by image similarity measure. The experimental results of our volume estimation method for four food categories validate the accuracy and reliability of our model-based approach. Chang Xu 0006, Ye He 0001, Nitin Khanna, Carol J. Boushey, Edward J. Delp |
ICIP | 5 |
| 2013 | Food image analysis: Segmentation, identification and weight estimationabstractWe are developing a dietary assessment system that records daily food intake through the use of food images taken at a meal. The food images are then analyzed to extract the nutrient content in the food. In this paper, we describe the image analysis tools to determine the regions where a particular food is located (image segmentation), identify the food type (feature classification) and estimate the weight of the food item (weight estimation). An image segmentation and classification system is proposed to improve the food segmentation and identification accuracy. We then estimate the weight of food to extract the nutrient content from a single image using a shape template for foods with regular shapes and area-based weight estimation for foods with irregular shapes. Ye He 0001, Chang Xu 0006, Nitin Khanna, Carol J. Boushey, Edward J. Delp |
ICME | 5 |
| 2013 | Inter-layer error concealment for scalable video coding based on motion vector averaging and slice interleavingabstractScalable video coding (SVC) is desirable for video applications in heterogeneous environments where users have a variety of terminals. Because motion-compensated prediction is used in SVC codec, when SVC video is transmitted over loss-prone networks, the decompressed video can suffer severe visual degradation across multiple frames. In order to overcome this, one scheme utilized at the decoder is error concealment to enhance the visual quality. This paper assumes that the packet delivery of the base layer is loss-prone the same way as the enhancement layer. Three inter-layer error concealment methods are proposed using two new approaches. (1) Motion vector averaging using adaptively averaging over multiple types of motion vectors in different layers for the recovery of lost motion vectors. (2) Slice interleaving utilizing an optimum ordering technique to make the average distance between two contiguous slices as far as possible. A two-layer spatial-temporal scalable video coding system is developed to evaluate the existing and proposed error concealment methods. The impact of burst packet losses and error propagation on the frames in both layers is investigated. Experimental results verified that the proposed error concealment methods outperform two existing methods. Bin Zhao 0003, Edward J. Delp |
ICME | 2 |
| 2013 | An Extreme Learning Machine-based pedestrian detection methodabstractPedestrian detection is a challenging task due to the high variance of pedestrians and fast changing background, especially for a single in-car camera system. Traditional HOG+SVM methods have two challenges: (1) false positives and (2) processing speed. In this paper, a new pedestrian detection method using multimodal HOG for pedestrian feature extraction and kernel based Extreme Learning Machine (ELM) for classification is presented. The experimental results using our naturalistic driving dataset show that the proposed method outperforms the traditional HOG+SVM method in both recognition accuracy and processing speed. Kai Yang 0009, Yingzi Du, Edward J. Delp, Pingge Jiang, Yaobin Chen, Rini Sherony, Hiroyuki Takahashi |
Intelligent Vehicles Symposium | 3 |
| 2013 | Adaptive error concealment for Multiple Description video coding using error estimationabstractMultiple Description Coding (MDC) is an error resilient video coding method suitable for real-time applications when retransmission is unacceptable. A four description MDC architecture is described with temporal and spatial error concealment capabilities at the decoder. In this paper, we present an adaptive error concealment method using error estimation. Experimental results demonstrate the efficacy of the proposed method. Neeraj Gadgil, Mary L. Comer, Edward J. Delp |
PCS | 3 |
| 2012 | Adaptive error concealment for Multiple Description Video Coding using motion vector analysisabstractMultiple Description Coding (MDC) is an efficient error resilient video coding method especially for real-time applications when retransmission is unacceptable. For applications with scalable, multicast and P2P environments, it is advantageous to use more than two descriptions. In this paper, we present an adaptive temporal-spatial error concealment method using motion vector analysis to improve the performance of MDC. Experimental results demonstrate the efficacy of the proposed method. Neeraj Gadgil, Meilin Yang, Mary L. Comer, Edward J. Delp |
ICIP | 4 |
| 2012 | Snakes assisted food image segmentationabstractIn this paper we describe an image segmentation method for segmenting food items in images used for dietary assessment. Dietary assessment methods used to determine the foods and beverages consumed at a meal are essential for understanding the link between diet and health. Snakes, or active contours, are used extensively to locate object boundaries and segment images. Experimental results using classical snakes on food images show the problems associated with contour initialization and poor detection performance for food images. In this paper, we explore various methods of contour initialization and integrate a background removal method to improve the performance of food image segmentation. We describe the details of the proposed food image segmentation method and also evaluate our segmentation approach on food images. Nitin Khanna, Carol J. Boushey, Edward J. Delp |
MMSP | 4 |
| 2012 | Macroblock-level adaptive error concealment methods for MDCabstractMultiple Description Coding (MDC) is one of the most efficient error resilient video coding methods especially when retransmissions are not acceptable. For applications in scalable, multicast and P2P environments, it is advantageous to use more than two descriptions. In this paper, we present two macroblock-level adaptive error concealment methods to improve the performance of MDC with four descriptions. A Gilbert model is used as the channel model to simulate burst packet loss. Experimental results demonstrate the efficacy of the two proposed methods. Meilin Yang, Mary L. Comer, Edward J. Delp |
PCS | 3 |
| 2012 | Quality fusion based multimodal eye recognitionabstractMultimodal eye recognition can improve the biometric systems recognition accuracy by combining iris and sclera recognition. However, poor quality images can significantly affect the system performance. In this paper, we proposed a quality fusion based multimodal eye recognition. Our quality measure evaluated the entire eye image quality, iris area quality, and sclera area quality. The experimental results show that our overall iris and sclera quality scores are highly correlated to recognition accuracy, and our quality fusion based eye recognition can improve and predict the performance of eye recognition systems. Zhi Zhou 0001, Yingzi Du, Craig Belcher, N. Luke Thomas, Edward J. Delp |
SMC | 5 |
| 2012 | Generalized PCM Coding of ImagesabstractPulse-code modulation (PCM) with embedded quantization allows the rate of the PCM bitstream to be reduced by simply removing a fixed number of least significant bits from each codeword. Although this source coding technique is extremely simple, it has a poor coding efficiency. In this paper, we present a generalized PCM (GPCM) algorithm for images that simply removes bits from each codeword. In contrast to PCM, however, the number and the specific bits that a GPCM encoder removes in each codeword depends on its position in the bitstream and the statistics of the image. Since GPCM allows the encoding to be performed with different degrees of computational complexity, it can adapt to the computational resources that are available in each application. Experimental results show that GPCM outperforms PCM with a gain that depends on the rate, the computational complexity of the encoding, and the degree of inter-pixel correlation of the image. Josep Prades-Nebot, Marleen Morbée, Edward J. Delp |
IEEE Trans. Image Process. | 3 |
| 2012 | A New Human Identification Method: Sclera RecognitionabstractThe blood vessel structure of the sclera is unique to each person, and it can be remotely obtained nonintrusively in the visible wavelengths. Therefore, it is well suited for human identification (ID). In this paper, we propose a new concept for human ID: sclera recognition. This is a challenging research problem because images of sclera vessel patterns are often defocused and/or saturated and, most importantly, the vessel structure in the sclera is multilayered and has complex nonlinear deformations. This paper has several contributions. First, we proposed the new approach for human ID: sclera recognition. Second, we developed a new method for sclera segmentation which works for both color and grayscale images. Third, we designed a Gabor wavelet-based sclera pattern enhancement method to emphasize and binarize the sclera vessel patterns. Finally, we proposed a line-descriptor-based feature extraction, registration, and matching method that is illumination, scale, orientation, and deformation invariant and can mitigate the multilayered deformation effects and tolerate segmentation error. The experimental results show that sclera recognition is a promising new biometrics for positive human ID. Zhi Zhou 0001, Yingzi Du, N. Luke Thomas, Edward J. Delp |
IEEE Trans. Syst. Man Cybern. Part A | 4 |
| 2011 | Crowd flow estimation using multiple visual features for scenes with changing crowd densitiesabstractCrowd estimation and monitoring is an important surveillance task. We address the problem of estimating the “flow,” that is the number of persons passing a designated region in a unit time. We designate an area of the scene as a virtual trip wire and accumulate the total number of foreground pixels (in the trip wire) over a chosen time period. We show that cumulative pixel count is related to the number of persons passing through the trip-wire by a scale factor. This scale factor is highly sensitive to the “crowdedness” (levels of crowd density) of the scene which creates different levels of occlusion of the individuals walking/passing through the trip-wire. We use texture features to determine the crowdedness and choose the most appropriate scaling factor. Our method does not require detection and tracking of individuals and is robust to scene dynamics, background subtraction errors, and different crowd levels. Satyam Srivastava, Ka Ki Ng, Edward J. Delp |
AVSS | 3 |
| 2011 | Secret Sharing in the Encrypted Domain with Secure ComparisonabstractSecret sharing refers to the sharing and reconstruction of a secret among a group of participants. In light of a few disadvantages of the traditional secret sharing, it is desirable to reconstruct the secret directly from the encrypted shares to protect the secret and shares during secret reconstruction. This paper develops several useful techniques for processing encrypted data. A new secure comparison protocol based on the additive homomorphic property and a new scheme of secret sharing in the encrypted domain with secure comparison are proposed. In this scenario, secret reconstruction could be done on an untrusted device while comparison is performed in a trusted environment. Experimental results confirm the effectiveness of the proposed scheme and protocol. Bin Zhao 0003, Edward J. Delp |
GLOBECOM | 2 |
| 2011 | Color correction for object tracking across multiple camerasabstractColor is a powerful attribute that is used to characterize objects for tracking and other surveillance tasks. Since color is dependent on ambient illumination and the imaging equipment, the reliability of color features decreases as the object moves through regions observed with different cameras and with different illumination conditions. Typically, this issue is addressed using alternative color spaces or highly simplified color transformation. In this paper, we propose a more methodical approach to achieving color consistency using colorimetric principles. We model the image analysis system as an observer and develop camera-specific transformations so that images of the same object appear similar to this observer. The transformations thus developed are approximated with 3D look-up tables for fast operation. Experimental results show improved color consistency even under two very different illumination conditions. The results are evaluated both qualitatively and in terms of the effect on a particle filter based object tracker. Satyam Srivastava, Ka Ki Ng, Edward J. Delp |
ICASSP | 3 |
| 2011 | Secret Sharing in the Encrypted DomainabstractSecret sharing refers to dividing a secret into pieces or shares and allocating the shares among a group of participants. The secret can be reconstructed only when a sufficient number of shares are combined. To protect each share during secret reconstruction, it is desirable to reconstruct the secret directly from the encrypted shares. Secure signal processing is typically based on probabilistic and homomorphic encryption in public key cryptosystem. Based on related works, a composite method using binary representation and precomputation is implemented for efficient exponentiation of encrypted signals. A new scheme of secret sharing in the encrypted domain that makes use of the efficient exponentiation method is presented. Experimental results validate the suitability of the proposed method. Bin Zhao 0003, Edward J. Delp |
ICC | 2 |
| 2011 | Combining global and local features for food identification in dietary assessmentabstractMany chronic diseases, such as heart diseases, diabetes, and obesity, can be related to diet. Hence, the need to accurately measure diet becomes imperative. We are developing methods to use image analysis tools for the identification and quantification of food consumed at a meal. In this paper we describe a new approach to food identification using several features based on local and global measures and a "voting" based late decision fusion classifier to identify the food items. Experimental results on a wide variety of food items are presented. Marc Bosch, Fengqing Zhu 0001, Nitin Khanna, Carol J. Boushey, Edward J. Delp |
ICIP | 5 |
| 2011 | An adaptable spatial-temporal error concealment method for Multiple Description Coding based on error trackingabstractMultiple Description Coding (MDC) is one of the most efficient methods to combat error-prone channels especially when retransmission is unacceptable. In applications involving scalable, multicast and P2P environments, it is advantageous to use more than two descriptions. In this paper, we present an adaptable temporal-spatial error concealment method based on error tracking to improve the performance of MDC. A Gilbert model is used as the channel model for burst packet loss simulation. Experimental results demonstrate the efficacy of the proposed method. Meilin Yang, Mary L. Comer, Edward J. Delp |
ICIP | 3 |
| 2011 | Integrated database system for mobile dietary assessment and analysisabstractOf the 10 leading causes of death in the US, 6 are related to diet. Unfortunately, methods for real-time assessment and proactive health management of diet do not currently exist. There are only minimally successful tools for historical analysis of diet and food consumption available. In this paper, we present an integrated database system that provides a unique perspective on how dietary assessment can be accomplished. We have designed three interconnected databases: an image database that contains data generated by food images, an experiments database that contains data related to nutritional studies and results from the image analysis, and finally an enhanced version of a nutritional database by including both nutritional and visual descriptions of each food. We believe that these databases provide tools to the healthcare community and can be used for data mining to extract diet patterns of individuals and/or entire social groups. Marc Bosch, TusaRebecca Schap, Fengqing Zhu 0001, Nitin Khanna, Carol J. Boushey, Edward J. Delp |
ICME | 6 |
| 2011 | A method for translating printed documents using a hand-held deviceabstractWe have developed a system to translate and interpret printed documents (such as periodicals) using a commercially available mobile device (e.g., a mobile telephone) with an embedded camera. The system comprises an automatic layout analysis tool along with an Optical Character Recognition (OCR) engine and a translation engine. The translation engine combines Rule-Based Machine Translation (RBMT) software, a list of context-sensitive words/phrases, and an encyclopedia for a list of words. We implemented the proposed system on a Nokia N900 smartphone for translating newspaper articles from Spanish to English. Our tests show that the accuracy and speed of the system is mostly influenced by the accuracy and speed of the OCR method used. Our application requires 19.43 MB of memory and no more than 28 MB of dynamic memory. We show the ability to complete the process for roughly 410 words in as little as 65 seconds for high accuracy or 17 seconds for high speed, on average. The energy consumption in both cases is minimal. Albert Parra Pozo, Andrew W. Haddad, Mireille Boutin, Edward J. Delp |
ICME | 4 |
| 2011 | A hand-held multimedia translation and interpretation system for diet managementabstractWe propose a system for helping individuals who follow a medical diet maintain this diet while visiting countries where a foreign language is spoken. Our focus is on diets where certain foods must either be restricted (e.g., metabolic diseases), avoided (e.g., food intolerance or allergies), or preferably consumed for medical reasons. However, our framework can be used to manage other diets (e.g., vegan) as well. The system is based on the use of a hand-held multimedia device such as a PDA or mobile telephone to analyze and/or disambiguate the content of foods offered on restaurant menus and interpret them in the context of specific diets. The system also provides the option to communicate diet-related instructions or information to a local person (e.g., a waiter) as well as obtain clarifications through dialogue. All computations are performed within the device and do not require a network connection. Real-time text translation is a challenge. We address this challenge with a light-weight, context-specific machine translation method. This method builds on a modification of existing open source Machine Translation (MT) software to obtain a fast and accurate translation. In particular, we describe a method we call n-gram consolidation that joins words in a language pair and increases the accuracy of the translation. We developed and implemented this system on the iPod Touch for English speakers traveling in Spain. Our tests indicate that our translation method yields the correct translation more often than general purpose translation engines such as Google Translate, and does so almost instantaneously. The memory requirements of the application, including the database of picture, are also well within the limits of the device. Albert Parra Pozo, Andrew W. Haddad, Mireille Boutin, Edward J. Delp |
ICME | 4 |
| 2011 | Co-ordinate mapping and analysis of vehicle trajectory for anomaly detectionabstractBehavioral information can be inferred from the trajectory of a rigid object. We propose a method to detect anomalies in the approach of a vehicle by observing the patterns in its velocity and describe methods for more effective analysis of the velocity trajectory. First we define a hypothetical co-ordinate system in which the axes are specified with respect to the road and distances are true ground measurements. In this co-ordinate system the (axial) velocity is a one dimensional quantity. We estimate the “normal” trends in the velocities by activity path modeling after scaling the velocities by the vehicle's average speed. We detect an anomaly if a vehicle's velocity does not fall in the available path models. Finally, we use the shape of the trajectory to determine turns and other significant maneuvers. We also detect an anomaly if a vehicle's velocity falls in a path model which is inconsistent with the shape of its trajectory. We observe that our proposed co-ordinate system also improves trajectory shape analysis by suppressing false turns. Satyam Srivastava, Ka Ki Ng, Edward J. Delp |
ICME | 3 |
| 2011 | Temporal Dietary Patterns Using Kernel k-Means ClusteringabstractChronic diseases, such as heart disease, diabetes, and obesity, have been linked with diet. Nutrient intake is also associated with diet. However, much of the research completed to elucidate these associations has not incorporated the concept of time. This paper introduces the concept of temporal dietary patterns and demonstrates a novel construct of 24-hour temporal dietary patterns for energy intake, present in a sample of the adult U.S. population 20 years and older (NHANES 1999-2004 dataset). An appropriate distance metric is proposed for comparing 24-hour diet records and is used with kernel k-means clustering to identify the temporal dietary patterns. Nitin Khanna, Heather A. Eicher-Miller, Carol J. Boushey, Saul B. Gelfand, Edward J. Delp |
ISM | 5 |
| 2011 | Low Complexity Image Quality Measures for Dietary Assessment Using Mobile DevicesabstractMany chronic diseases, such as heart disease, diabetes, and obesity, can be related to diet. Hence, the need to accurately measure diet becomes imperative. We are developing image analysis tools for the identification and quantification of foods consumed at a meal. Our system relies on a single meal image from the user for doing food identification and quantity estimation. Therefore, it is very important to assist the user in acquiring a good quality image by providing immediate feedback about the image quality. This paper presents low complexity image quality measures which are deployed on handheld mobile devices. Chang Xu 0006, Nitin Khanna, Carol J. Boushey, Edward J. Delp |
ISM | 4 |
| 2011 | A Low Complexity Sign Detection and Text Localization Method for Mobile ApplicationsabstractWe propose a low complexity method for sign detection and text localization in natural images. This method is designed for mobile applications (e.g., unmanned or handheld devices) in which computational and energy resources are limited. No prior assumption is made regarding the text size, font, language, or character set. However, the text is assumed to be located on a homogeneous background using a contrasting color. We have deployed our method on a Nokia N800 cellular phone as part of a system for automatic detection and translation of outdoor signs. This handheld device is equipped with a 0.3-megapixel camera capable of acquiring images of outdoor signs that typically contain enough details for the sign to be readable by a human viewer. Our experiments show that the text of these images can be accurately localized within the device in a fraction of a second. Katherine L. Bouman, Golnaz Abdollahian, Mireille Boutin, Edward J. Delp |
IEEE Trans. Multim. | 4 |
| 2010 | A low complexity method for detection of text area in natural imagesabstractWe propose a low complexity method for segmentation of text regions in natural images. This algorithm is designed for mobile applications (e.g. unmanned or hand-held devices) in which computational and energy resources are limited. No prior assumption is made regarding the text size, font, language, character set or the camera angle. However, the text is assumed to be located on a piecewise homogeneous background with a contrasting color. We have deployed our method on a Nokia N800 Internet tablet as part of a system for automatic detection and translation of outdoor signs. Our experiments show that the 0.3 megapixel images taken by the phone camera can be accurately segmented within the device in a fraction of a second. Katherine L. Bouman, Golnaz Abdollahian, Mireille Boutin, Edward J. Delp |
ICASSP | 4 |
| 2010 | An image analysis system for dietary assessment and evaluationabstractThere is a growing concern about chronic diseases and other health problems related to diet including obesity and cancer. Dietary intake provides valuable insights for mounting intervention programs for prevention of chronic diseases. Measuring accurate dietary intake is considered to be an open research problem in the nutrition and health fields. In this paper, we describe a novel mobile telephone food record that provides a measure of daily food and nutrient intake. Our approach includes the use of image analysis tools for identification and quantification of food that is consumed at a meal. Images obtained before and after foods are eaten are used to estimate the amount and type of food consumed. The mobile device provides a unique vehicle for collecting dietary information that reduces the burden on respondents that are obtained using more classical approaches for dietary assessment. We describe our approach to image analysis that includes the segmentation of food items, features used to identify foods, a method for automatic portion estimation, and our overall system architecture for collecting the food intake information. Fengqing Zhu 0001, Marc Bosch, Carol J. Boushey, Edward J. Delp |
ICIP | 4 |
| 2010 | Intrinsic signatures for scanned documents forensics : Effect of font shape and sizeabstractRecently there has been a great deal of interest in using features intrinsic to a data-generating sensor for the purpose of source identification. Numerous methods have been proposed for different problems related to sensor forensics, such as source camera identification using sensor noise as intrinsic signatures, printer identification using banding artifacts as intrinsic signatures and so on. The goal of our work is to identify the scanner used for generating a scanned (digital) version of a printed (hard-copy) document. In this paper we do extensive analysis of the effect of font shape and size on our recently proposed, texture analysis based intrinsic signatures for scanned documents. Some improvements in these intrinsic signatures are proposed to make them robust to font shapes and sizes. Nitin Khanna, Edward J. Delp |
ISCAS | 2 |
| 2010 | An Overview of the Technology Assisted Dietary Assessment Project at Purdue UniversityabstractIn this paper, we describe the Technology Assisted Dietary Assessment (TADA) project at Purdue University. Dietary intake, what someone eats during the course of a day, provides valuable insights for mounting intervention programs for prevention of many chronic diseases such as obesity and cancer. Accurate methods and tools to assess food and nutrient intake are essential for research on the association between diet and health. An overview of our methods used in the TADA project is presented. Our approach includes the use of image analysis tools for identification and quantification of food that is consumed at a meal. Images obtained before and after foods are eaten are used to estimate the amount and type of food consumed. Nitin Khanna, Carol J. Boushey, Deborah A. Kerr, Martin Okos, David S. Ebert, Edward J. Delp |
ISM | 6 |
| 2010 | Development of a mobile user interface for image-based dietary assessmentabstractIn this paper, we present a mobile user interface for image-based dietary assessment. The mobile user interface provides a front end to a client-server image recognition and portion estimation software. In the client-server configuration, the user interactively records a series of food images using a built-in camera on the mobile device. Images are sent from the mobile device to the server, and the calorie content of the meal is estimated. In this paper, we describe and discuss the design and development of our mobile user interface features. We discuss the design concepts, through initial ideas and implementations. For each concept, we discuss qualitative user feedback from participants using the mobile client application. We then discuss future designs, including work on design considerations for the mobile application to allow the user to interactively correct errors in the automatic processing while reducing the user burden associated with classical pen-and-paper dietary records. SungYe Kim, TusaRebecca Schap, Marc Bosch, Ross Maciejewski, Edward J. Delp, David S. Ebert, Carol J. Boushey |
MUM | 5 |
| 2010 | A four-description MDC for high loss-rate channelsabstractOne of the most difficult problems in video transmission is communication over error-prone channels, especially when retransmission is unacceptable. To address this problem, Multiple Description Coding (MDC) has been proposed as an effective solution due to its robust error resilience. Considering applications in scalable, multicast and P2P environments, it is advantageous to use more than two descriptions (which is designated multi-description MDC in this paper). In this paper, we present a new four-description MDC for high lossrate channel using a hybrid structure of temporal and spatial correlations. A Gilbert model is used as the channel model for burst packet loss simulation. Experimental results demonstrate the efficacy of the proposed method. Meilin Yang, Mary L. Comer, Edward J. Delp |
PCS | 3 |
| 2010 | Camera Motion-Based Analysis of User Generated VideoabstractIn this paper we propose a system for the analysis of user generated video (UGV). UGV often has a rich camera motion structure that is generated at the time the video is recorded by the person taking the video, i.e., the ¿camera person.¿ We exploit this structure by defining a new concept known ascamera viewfor temporal segmentation of UGV. The segmentation provides a video summary with unique properties that is useful in applications such as video annotation. Camera motion is also a powerful feature for identification of keyframes and regions of interest (ROIs) since it is an indicator of the camera person's interests in the scene and can also attract the viewers' attention. We propose a new location-based saliency map which is generated based on camera motion parameters. This map is combined with other saliency maps generated using features such as color contrast, object motion and face detection to determine the ROIs. In order to evaluate our methods we conducted several user studies. A subjective evaluation indicated that our system produces results that is consistent with viewers' preferences. We also examined the effect of camera motion on human visual attention through an eye tracking experiment. The results showed a high dependency between the distribution of fixation points of the viewers and the direction of camera movement which is consistent with our location-based saliency map. Golnaz Abdollahian, Cüneyt M. Taskiran, Zygmunt Pizlo, Edward J. Delp |
IEEE Trans. Multim. | 4 |
| 2009 | Perceptual quality evaluation for texture and motion based video codingabstractOne approach that can be used to increase compression efficiency beyond the data rates achievable by state-of-the-art video codecs is to use content-based methods whereby not all the pixels are conventionally encoded. An approach to reduce the data rate is to use different coding methods for pixels belonging to areas containing large amount of detail that are costly to encode, for example textures. This can be extended by focusing on the semantic meaning of objects represented in the video sequence and also taking into consideration human visual system properties. The goal is to determine where ¿detail-irrelevant¿ regions are located in the frame and synthesize them with acceptable perceptual quality. In this paper, we discuss the effects and trade-offs of these techniques based on a set of perceptual experiments and analyze how these areas can influence the viewer's attention. Marc Bosch, Fengqing Zhu 0001, Edward J. Delp |
ICIP | 3 |
| 2009 | Model-based methods for developing color transformation between two display devicesabstractWe present a framework for developing a color transformation between two display devices to achieve a desired perceptual match. The framework is realized by first developing an optimal color transformation between the two devices, and then formulating an optimization problem to develop a hardware resource-constrained approximation to the optimal transformation. We employ the framework to investigate how well the system performs when the transformation is constrained to be a single 3×3 linear transformation directly between the non-linear RGB spaces of two devices. The motivation for this constraint is to ease resource requirements in a real-time hardware implementation. Thanh H. Ha, Satyam Srivastava, Edward J. Delp, Jan P. Allebach |
ICIP | 3 |
| 2009 | Segmentation and registration based analysis of microscopy imagesabstractOptical microscopy exhibits many challenges for digital image analysis. In general, microscopy volumes are inherently anisotropic, suffer from decreasing contrast with tissue depth, and characteristically have low signal levels. This paper describes a method to segment multiphoton fluorescent microscopy images via a combination of segmentation and registration methods. In particular, the proposed method utilizes image enhancement and spatial filtering along with registration and temporal filtering. Experimental results demonstrate the efficacy of the proposed method. Kevin S. Lorenz, Francisco Serrano, Paul Salama, Edward J. Delp |
ICIP | 4 |
| 2009 | Generating optimal look-up tables to achieve complex color space transformationsabstractColor space transformations are very common in digital imaging and display systems. Their complex mathematics requires hardware friendly implementations to make them practical. In this paper, look-up tables (LUT) are described for implementation of the color transformations. We show that this is possible because an analytical model of the transform is available for training and cross-validation. We identify parameters that affect the performance of a LUT-based system and then formulate the solution as an optimization problem. Satyam Srivastava, Thanh H. Ha, Edward J. Delp, Jan P. Allebach |
ICIP | 3 |
| 2009 | User generated video annotation using Geo-tagged image databasesabstractIn this paper we propose a system that annotates a user generated video based on the associated location metadata, by exploiting user-tagged image databases. An example of such a database is a photo sharing Web site such as Flickr where users upload their images and annotate them with various tags. The goal is to find the tags that have high probability of being relevant to the video without any complex object or action recognition being done to the video sequence. A video is first segmented into camera views and a set of keyframes are selected to represent the video. We will describe the concept of camera view as the basic element of user generated videos which has special properties suitable for the video annotation application. The keyframes are used to retrieve the most relevant images in the database. A ldquotag processingrdquo step is then used to tag the video. Golnaz Abdollahian, Edward J. Delp |
ICME | 2 |
| 2009 | Forensic Techniques for Image Source Classification: A Comparative Study
Edward J. Delp |
IWDW | 1 |
| 2009 | An overviewof texture and motion based video coding at Purdue UniversityabstractIn recent years there has been a growing interest in developing novel techniques for increasing the coding efficiency of video compression methods. We approach the problem by not encoding all the pixels, in particular, regions belonging to areas that the viewer will not perceive the specific details in the scene could be skipped or encoded at a much lower data rate. This approach can also be expanded by considering a model of the human visual perception system. In this paper we review some of the approaches we have investigated at Purdue University. The goal is to determine where ldquodetail-irrelevantrdquo regions in the frame are located and not encode them. We will also discuss a set of subjective quality evaluation experiments to determine what is the overall perceptual quality of these approaches. Marc Bosch, Fengqing Zhu 0001, Edward J. Delp |
PCS | 3 |
| 2009 | Feedback aided content adaptive unequal error protection based on Wyner-Ziv codingabstractIn prior work, we proposed an unequal error protection algorithm based on Wyner-Ziv coding for error resilient video transmission. In subsequentwork it was demonstrated that using either of content adaptive unequal error protection or feedback aided unequal error protection individually improved error resilience performance. In the current paper we propose to combine the use of a content adaptive function, implemented at the encoder, with the channel loss feedback provided by the decoder. The experimental results demonstrate improved rate distortion performance. Paul Salama, Edward J. Delp |
PCS | 3 |
| 2009 | Comparison of lossy compression performance on natural color imagesabstractIn estimation of the efficiency for lossy image compression methods, standard sets of test images are commonly used. This allows the comparison of new techniques to existing methods without having to actually implement the existing technique. However, this does not allow adequate evaluation of the performance of the methods for compressing natural images. In this paper, we analyze the efficiency of a set of lossy compression techniques (JPEG, JPEG2000, HD Photo and ADCTC) using a set of images obtained by three consumer quality digital cameras. Nikolay N. Ponomarenko, Vladimir Lukin 0001, Karen Egiazarian, Edward J. Delp |
PCS | 4 |
| 2009 | Efficient and Low-Complexity Surveillance Video Compression Using Backward-Channel Aware Wyner-Ziv Video CodingabstractVideo surveillance has been widely used in recent years to enhance public safety and privacy protection. A video surveillance system that deals with content analysis and activity monitoring needs efficient transmission and storage of the surveillance video data. Video compression techniques can be used to achieve this goal by reducing the size of the video with no or small quality loss. State-of-the-art video compression methods such as H.264/AVC often lead to high computational complexity at the encoder, which is generally implemented in a video camera in a surveillance system. This can significantly increase the cost of a surveillance system, especially when a mass deployment of end cameras is needed. In this paper, we discuss the specific considerations for surveillance video compression. We present a surveillance video compression system with low-complexity encoder based on Wyner-Ziv coding principles to address the tradeoff between computational complexity and coding efficiency. In addition, we propose a backward-channel aware Wyner-Ziv (BCAWZ) video coding approach to improve the coding efficiency while maintaining low complexity at the encoder. The experimental results show that for surveillance video contents, BCAWZ can achieve significantly higher coding efficiency than H.264/AVC intra coding as well as existing Wyner-Ziv video coding methods and is close to H.264/AVC inter coding, while maintaining similar coding complexity with intra coding. This shows that the low motion characteristics of many surveillance video contents and the low-complexity encoding requirement make our scheme a particularly suitable candidate for surveillance video compression. We further propose an error resilience scheme for BCAWZ to address the concern of reliable transmission in the backward-channel, which is essential to the quality of video data for real-time and reliable object detection and event analysis. Zhen Li 0008, Edward J. Delp |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Scanner identification using feature-based processing and analysisabstractDigital images can be obtained through a variety of sources including digital cameras and scanners. In many cases, the ability to determine the source of a digital image is important. This paper presents methods for authenticating images that have been acquired using flatbed desktop scanners. These methods use scanner fingerprints based on statistics of imaging sensor pattern noise. To capture different types of sensor noise, a denoising filterbank consisting four different denoising filters is used for obtaining the noise patterns. To identify the source scanner, a support vector machine classifier based on these fingerprints is used. These features are shown to achieve high classification accuracy. Furthermore, the selected fingerprints based on statistical properties of the sensor noise are shown to be robust under postprocessing operations, such as JPEG compression, contrast stretching, and sharpening. Nitin Khanna, Aravind K. Mikkilineni, Edward J. Delp |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2008 | Forensic techniques for classifying scanner, computer generated and digital camera imagesabstractDigital images can be captured or generated by a variety of sources including digital cameras, scanners and computer graphics softwares. In many cases it is important to be able to determine the source of a digital image such as for criminal and forensic investigation. This paper presents methods for distinguishing between an image captured using a digital camera, a computer generated image and an image captured using a scanner. The method proposed here is based on the differences in the image generation processes used in these devices and is independent of the image content. The method is based on using features of the residual pattern noise that exist in images obtained from digital cameras and scanners. The residual noise present in computer generated images does not have structures similar to the pattern noise of cameras and scanners. The experiments show that a feature based approach using an SVM classifier gives high accuracy. Nitin Khanna, George Chiu 0001, Jan P. Allebach, Edward J. Delp |
ICASSP | 4 |
| 2008 | A study on the effect of camera motion on human visual attentionabstractThe aim of this paper is to examine the effect of camera motion in user generated video with respect to human visual attention. Having a more accurate human attention model is particularly useful in applications such as video summarization where the identification of visual importance is crucial to the quality of the results. Most models proposed thus far, have not considered camera motion as an independent factor in identifying visual saliency. In this study eye movement was recorded while subjects watched videos with different types of camera motion. The distribution of the fixation points indicated high correlation between the visual saliency and the type and direction of camera movement. The statistical tests confirmed that the results are statistically significant. Golnaz Abdollahian, Zygmunt Pizlo, Edward J. Delp |
ICIP | 3 |
| 2008 | Video coding using motion classificationabstractIn this paper we present a video coding approach similar to texture- based methods but based on motion models. We consider motion perception properties instead of spatial texture properties of the video sequence. We integrate a motion classification algorithm to separate foreground objects containing noticeable motion from the background. These background areas are labeled as skipped areas that are not encoded. After decoding, frame reconstruction is performed by inserting the skipped background into the decoded frames. We are able to show as much as 15% an improvement over previous texture- based implementations in terms of video compression efficiency. Marc Bosch, Fengqing Zhu 0001, Edward J. Delp |
ICIP | 3 |
| 2008 | Automatic text area segmentation in natural imagesabstractWe present a hierarchical method for segmenting text areas in natural images. The method assumes that the text is written with a contrasting color on a more or less uniform background. No assumption is made regarding the language or character set used to write the text. In particular, the text can contain simple graphics or symbols. The key feature of our approach is that we first concentrate on finding the background of the text, before testing whether there is actually text on the background. Since uniform areas are easy to find in natural images, and since text backgrounds define areas which contain "holes" (where the text is written) we thus look for uniform areas containing "holes" and label them as text backgrounds candidates. Each candidate area is then further tested for the presence of text within its convex hull. We tested our method on a database of 65 images including English and Urdu text. The method correctly segmented all the text areas in 63 of these images, and in only 4 of these were areas that do not contain text also segmented. Syed Ali Raza Jafri, Mireille Boutin, Edward J. Delp |
ICIP | 3 |
| 2008 | A low-complexity iterative mode selection algorithm Forwyner-Ziv video compressionabstractAiming at improving compression performance, we consider mode selection for Wyner-Ziv video compression where a block of pixels in a video frame, after discrete cosine transform(DCT), can be either encoded by using H.264 Intra mode or Wyner-Ziv(WZ) mode with side information processed at the decoder. Under the constraint of encoding complexity, an iterative algorithm is proposed to find the best partition of a video frame into these two modes in the sense of minimizing the overall compression rate. It is shown that the algorithm always converges. Experimental results on standard video test sequences show that by using the proposed algorithm for mode selection, one can achieve about 0.4dB gain for WZ-encoded frames over a WZ video compression system without intra mode at rate 0.2 bits per pixel. Furthermore, in all the experiments our algorithm converges in 3 iterations. Dake He, Ashish Jagmohan, Ligang Lu, Edward J. Delp |
ICIP | 5 |
| 2008 | Modulo-PCM based encoding for high speed video camerasabstractIn this paper, we present a low complexity Modulo-PCM based coding algorithm for high speed video cameras used in applications that demand very high frame rates. By compressing the raw digital video generated at the camera front-end, our algorithm reduces the high data transfer speed required in this type of cameras. At the encoder, our algorithm compress the raw digital video by selectively removing bits. At the decoder, the spatial correlation between pixels is taken into account in order to achieve an efficient decoding. Experimental results show that our algorithm achieves significant improvements with respect to the simple strategy of dropping the same number of bits in all the pixels when high rate reductions are necessary or when the spatial correlation of the frames is high. Josep Prades-Nebot, Antoni Roca 0002, Edward J. Delp |
ICIP | 3 |
| 2008 | Backward channel aware Wyner-Ziv video coding: A study of complexity, rate, and distortion tradeoff
Zhen Li 0008, Edward J. Delp |
Signal Process. Image Commun. | 3 |
| 2007 | Finding Regions of Interest in Home Videos Based on Camera MotionabstractIn this paper, we propose an algorithm for identifying regions of interest (ROIs) in video, particularly for the keyframes extracted from a home video. The camera motion is introduced as a new factor that can influence the visual saliency. The global motion parameters are used to generate location-based importance maps. These maps can be combined with other saliency maps calculated using other visual and high-level features. Here, we employed the contrast-based saliency as an important low level factor along with face detection as a high level feature in our approach. Golnaz Abdollahian, Edward J. Delp |
ICIP (4) | 2 |
| 2007 | Spatial Texture Models for Video CompressionabstractIn this paper we integrate several spatial texture tools into a texture-based video coding scheme. We implemented texture techniques and segmentation strategies in order to detect texture regions in video sequences. These textures are analyzed using temporal motion techniques and are labeled as skipped areas that are not encoded. After the decoding process, frame reconstruction is performed by inserting the skipped texture areas into the decoded frames. We are able to show an improvement over previous texture-based implementations in terms of compression efficiency. Marc Bosch, Fengqing Zhu 0001, Edward J. Delp |
ICIP (1) | 3 |
| 2007 | Complexity-Rate-Distortion Analysis of Backward Channel Aware Wyner-Ziv Video CodingabstractMany Wyner-Ziv video coding (WZVC) schemes encode a video sequence into two types of frames, key frames and Wyner-Ziv frames. We have previously presented a Wyner-Ziv video coding scheme that uses backward channel aware motion estimation to encode the key frames, where motion estimation was performed at the decoder and motion information was sent back to the encoder. We refer to these backward predictively coded frames as BP frames. In this paper, we extend our previous work and propose three types of motion estimators. A model is presented to examine the complexity-rate-distortion performance of BP frames for the three motion estimators. Zhen Li 0008, Edward J. Delp |
ICIP (2) | 3 |
| 2007 | Sensor Forensics: Printers, Cameras and Scanners, They Never LieabstractForensic characterization of devices is important in many situations such as establishing the trust and verifying authenticity of data and the device that created it. Current forensic identification techniques for digital cameras, scanners and printers are highly reliable due to the fact that each of these devices cannot escape inherent electro-mechanical properties which add "signatures" to the data they produce. In this paper we will describe the sensor forensics work going on at Purdue University. Nitin Khanna, Aravind K. Mikkilineni, Pei-Ju Chiang, Maria V. Ortiz Segovia, Sungjoo Suh, George Chiu 0001, Jan P. Allebach, Edward J. Delp |
ICME | 8 |
| 2007 | A low bit-rate video coding approach using modified adaptive warping and long-term spatial memoryabstractIn this paper, an H.264/AVC video coding strategy is introduced that employs a spatial-temporal video sequence representation in which video frames are coded at a low spatial sampling rate and reference I frames are coded at high spatial resolution. High spatial frequency information is re-synthesized at the receiver side using an adaptive motion estimation and warping method. The approach as presented is shown to improve coding quality for sequences with low to moderate motion. Clyde Lettsome, Mark J. T. Smith, Edward J. Delp |
VCIP | 4 |
| 2007 | Unequal error protection using Wyner-Ziv coding for error resilienceabstractCompressed video is very sensitive to channel errors. A few bit losses can derail the entire decoding process. Thus, protecting compressed video is imperative to enable visual communications. Since different elements in a compressed video stream vary in their impact on the quality of the decoded video, unequal error protection can be used to provide efficient protection. This paper describes an unequal error protection method for protecting data elements in a video stream, via a Wyner--Ziv encoder that consists of a coarse quantizer and a Turbo coder based lossless Slepian--wolf encoder. Data elements that significantly impact the visual quality of decoded video, such as modes and motion vectors as used by H.264, are provided more parity bits than coarsely quantized transform coefficients. This results in an improvement in the quality of the decoded video when the transmitted sequence is corrupted by transmission errors, than obtained by the use of equal error protection. Paul Salama, Edward J. Delp |
VCIP | 3 |
| 2007 | Content-adaptive motion estimation for efficient video compressionabstractMotion estimation is the most important step in the video compression. Most of the current video compression systems use forward motion estimation, where motion information is derived at the encoder and sent to the decoder over the channel. Backward motion estimation does not derive an explicit representation of motion at the encoder. Instead, the encoder implicitly embeds the motion information in an alternative subspace. Most recently, an algorithm that adopts least-square prediction (LSP) for backward motion estimation has shown great potential to further improve coding efficiency. Forward motion estimation and backward motion estimation have both their advantages and disadvantages. Each is suitable for handling some specific category of patterns. In this paper, we propose a novel approach that combines both forward motion estimation and backward motion estimation in one framework to adaptively exploit the local motion characteristics in an arbitrary video sequence, thus achieving better coding efficiency. We refer to this as Content-Adaptive Motion Estimation (CoME). The encoder in the proposed system is able to adjust the motion estimation method in a rate-distortion optimized manner. According to the experimental results, CoME reduces the data rate in both lossless and lossy compression. Yuxin Liu 0006, Edward J. Delp |
VCIP | 3 |
| 2007 | Spatial and temporal models for texture-based video codingabstractIn this paper, we investigate spatial and temporal models for texture analysis and synthesis. The goal is to use these models to increase the coding efficiency for video sequences containing textures. The models are used to segment texture regions in a frame at the encoder and synthesize the textures at the decoder. These methods can be incorporated into a conventional video coder (e.g. H.264) where the regions to be modeled by the textures are not coded in a usual manner but texture model parameters are sent to the decoder as side information. We showed that this approach can reduce the data rate by as much as 15%. Fengqing Zhu 0001, Ka Ki Ng, Golnaz Abdollahian, Edward J. Delp |
VCIP | 4 |
| 2007 | Rate Distortion Analysis of Motion Side Estimation in Wyner-Ziv Video CodingabstractWyner-Ziv video coding (WZVC) has gained considerable interests in the research community. In this paper, we present a model to examine the WZVC performance and compare it with conventional motion-compensated prediction (MCP) based video coding. Theoretical results show that although WZVC can achieve as much as 6-dB gain over conventional video coding without motion search, it still falls 6 dB or more behind current best MCP-based INTER-frame video coding. We further study the use of subpixel and multireference motion search methods to improve WZVC efficiency. The analytical results are confirmed by simulations and experiments. Zhen Li 0008, Edward J. Delp |
IEEE Trans. Image Process. | 3 |
| 2006 | Backward Channel Aware Wyner-Ziv Video CodingabstractWyner-Ziv video coding is a low complexity video encoding method that exploits source statistics at the decoder and shifts much of the computational complexity from the encoder to the decoder. Many Wyner-Ziv video coding schemes rely on the use of INTRA frames so that the decoder can derive sufficiently accurate side information. In this paper we extend our previous work on Wyner-Ziv coding that uses backward channel aware motion estimation. The basic idea is to perform motion estimation at the decoder and send the motion information back to the encoder through a backward channel. Experimental results show that this method can significantly improve the coding efficiency with a minimal use of the backward channel bandwidth. We also propose an error resilient technique by using error detection and adaptive coding to handle the situation when the backward channel is subject to erasure errors or delays. Zhen Li 0008, Edward J. Delp |
ICIP | 3 |
| 2006 | A new secure group key management scheme for multicast over wireless cellular networksabstractIn wireless networks, secure multicast protocols are more difficult to implement efficiently due to the dynamic nature of the multicast group and the scarcity of bandwidth at the receiving and transmitting ends. Mobility is one of the most distinct features to be considered in wireless networks. Moving users onto the key tree causes extra key management resources even though they are still in service. To take care of frequent handoff between wireless access networks, it is necessary to reduce the number of rekeying messages and the size of the messages. The multicast protocol used in wired networks does not perform well in wireless networks because multicast structures are fragile as the mobile node moves and connectivity changes. When we choose a key management scheme, the structure of the wireless network should be considered very carefully. In this paper, we design a key management tree such that neighbors on the key tree are also physical neighbors on the cellular network. By tracking the user location, we localize the delivery of rekeying messages to the users who need them. This lessens the amount of traffic in wireless and wired intervals of the network. The group key management scheme uses the prepositioned secret sharing scheme Hwa Young Um, Edward J. Delp |
IPCCC | 2 |
| 2006 | Decomposing parameters of mixture Gaussian model using genetic and maximum likelihood algorithms on dental images
Nariman Majdi-Nasab, Mostafa Analoui, Edward J. Delp |
Pattern Recognit. Lett. | 3 |
| 2006 | Wyner-Ziv Video Coding With Universal PredictionabstractThe coding efficiency of a Wyner-Ziv video codec relies significantly on the quality of side information extracted at the decoder. The construction of efficient side information is difficult thanks in part to the fact that the original video sequence is not available at the decoder. Conventional motion search methods are widely used in the Wyner-Ziv video decoder to extract the side information. This substantially increases the Wyner-Ziv video decoding complexity. In this paper, we propose a new method to construct side estimation based on the idea of universal prediction. This method, referred to as Wyner-Ziv video coding with universal prediction (WZUP), does not perform motion search or assume an underlying model of the original input video sequences at the decoder. Instead, WZUP estimates the side information based on its observations on the past reconstructed video data. We show that WZUP can significantly reduce decoding complexity at the decoder and achieve a fair side estimation performance, thus making it possible to design both the video encoder and the decoder with low computational complexity Zhen Li 0008, Edward J. Delp |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | Rate-distortion analysis of motion-compensated rate scalable videoabstractGenerally speaking, rate scalable video systems today are evaluated operationally, meaning that the algorithm is implemented and the rate-distortion performance is evaluated for an example set of inputs. However, in these cases it is difficult to separate the artifacts caused by the compression algorithm and data set with general trends associated with scalability. In this paper, we derive and evaluate theoretical rate-distortion performance bounds for both layered and continuously rate scalable video compression algorithms which use a single motion-compensated prediction (MCP) loop. These bounds are derived using rate-distortion theory based on an optimum mean-square error (MSE) quantizer, and are thus applicable to all methods of intraframe encoding which use MSE as a distortion measure. By specifying translatory motion and using an approximation of the predicted error frame power spectral density, it is possible to derive parametric versions of the rate-distortion functions which are based solely on the input power spectral density and the accuracy of the motion-compensated prediction. The theory is applicable to systems which allow prediction drift, such as the data-partitioning and SNR-scalability schemes in MPEG-2, as well as those with zero prediction drift such as fine granularity scalability MPEG-4. For systems which allow prediction drift we show that optimum motion compensation is a sufficient condition for stability of the decoding system. Gregory W. Cook, Josep Prades-Nebot, Yuxin Liu 0006, Edward J. Delp |
IEEE Trans. Image Process. | 4 |
| 2006 | An analysis of the efficiency of different SNR-scalable strategies for video codersabstractIn this paper, we analyze the efficiency of three signal-to-noise scalable strategies for video coders using single-loop motion-compensated prediction (MCP). In our analysis, we assume the video sequences have uniform and constant translational motion and we model MCP as a stochastic filter. We also assume an exponential model for the distortion-rate function of the intraframe coding. The analysis is divided into two parts: the steady-state analysis and the transient analysis. In the first part, only the steady-state response of the coders is taken into account, and, thus, this analysis allows us to asses approximately the efficiency of coders with long input sequences. The transitory analysis considers both the transient and the steady-state responses of the coders, which makes it appropriate to analyze coders using periodic intraframes or with short input sequences. To validate our analysis, theoretical results have been compared to results from encodings of real video sequences using the scalable adaptive motion compensated wavelet video coder. We show that our theoretical analysis effectively describes qualitatively the main trends of every video coding strategy. Josep Prades-Nebot, Gregory W. Cook, Edward J. Delp |
IEEE Trans. Image Process. | 3 |
| 2006 | Automated Video Program Summarization Using Speech TranscriptsabstractCompact representations of video data greatly enhances efficient video browsing. Such representations provide the user with information about the content of the particular sequence being examined while preserving the essential message. We propose a method to automatically generate video summaries using transcripts obtained by automatic speech recognition. We divide the full program into segments based on pause detection and derive a score for each segment, based on the frequencies of the words and bigrams it contains. Then, a summary is generated by selecting the segments with the highest score to duration ratios while at the same time maximizing the coverage of the summary over the full program. We developed an experimental design and a user study to judge the quality of the generated video summaries. We compared the informativeness of the proposed algorithm with two other algorithms for three different programs. The results of the user study demonstrate that the proposed algorithm produces more informative summaries than the other two algorithms Cüneyt M. Taskiran, Zygmunt Pizlo, Arnon Amir, Dulce B. Ponceleon, Edward J. Delp |
IEEE Trans. Multim. | 5 |
| 2005 | Wyner-Ziv video side estimator: conventional motion search methods revisitedabstractSide estimation plays an important role in Wyner-Ziv video coding. In this paper, we examine the use of two conventional motion search methods to improve side estimation. Unlike motion estimators used in conventional video encoders, side estimators do not have access to the original frame. This difference leads to some interesting observations concerning conventional motion search methods. Analytical and simulation results show that while multi-reference motion search is still effective, side estimators are not as sensitive to motion search pixel accuracies. Zhen Li 0008, Edward J. Delp |
ICIP (1) | 2 |
| 2005 | Rate allocation for prediction drift reduction in video streamingabstractIn SNR-scalable motion compensated video coders, prediction drift is introduced when decoding below the rate at which the encoder loop operates. Prediction drift reduces the coding efficiency and changes the quality of the video over time which is very annoying visually. In this paper, we propose a rate allocation algorithm to reduce these quality variations. By modeling both the video signal and the coder, we analyze the efficiency of the coder and find the rate allocation which creates frames which have the same distortion when decoded. We have compared numerical simulations of our algorithm with experimental results from real video encodings. The results show that by using a proper rate allocation algorithm quality fluctuations can be reduced. Josep Prades-Nebot, Gregory W. Cook, Edward J. Delp |
ICIP (3) | 3 |
| 2005 | Rate allocation algorithms for motion compensated embedded video codersabstractIn this paper, we present two rate allocation algorithms for embedded motion compensated video coders. The algorithms are based on the modeling of both the video signal and the coder which allow us to express the coding distortion with a recurrence equation. Our algorithms assign rates to the frames of each group of pictures (GOP) of a video sequence in an optimum way. In the first algorithm, the criterion is to minimize the average (MENAVE) distortion and in the second to achieve constant distortion (CD) in all frames. Numerical simulations show the MENAVE criterion can introduce large variations in quality with no significant gains in average distortion with respect to the CD criterion. We also show how the motion estimation accuracy and the GOP length influence in both strategies. Josep Prades-Nebot, Gregory W. Cook, Edward J. Delp |
ICIP (3) | 3 |
| 2005 | Multimedia security: the 22nd century approach
Edward J. Delp |
Multim. Syst. | 1 |
| 2005 | Advances in Digital Video Content ProtectionabstractThe use of digital video offers immense opportunities for creators; however, the ability for anyone to make perfect copies and the ease by which those copies can be distributed also facilitate misuse, illegal copying and distribution ("piracy"), plagiarism, and misappropriation. Popular Internet software based on a peer-to-peer architecture has been used to share copyrighted movies, music, software, and other materials. Concerned about the consequences of illegal copying and distribution on a massive scale, content owners are interested in digital rights management (DRM) systems which can protect their rights and preserve the economic value of digital video. A DRM system protects and enforces the rights associated with the use of digital content. Unfortunately, the technical challenges for securing digital content are formidable and previous approaches have not succeeded. We overview the concepts and approaches for video DRM and describe methods for providing security, including the roles of encryption and video watermarking. Current efforts and issues are described in encryption, watermarking, and key management. Lastly, we identify challenges and directions for further investigation in video DRM. E. I. Lin, Ahmet M. Eskicioglu, Reginald L. Lagendijk, Edward J. Delp |
Proc. IEEE | 4 |
| 2005 | Block artifact reduction using a transform-domain Markov random field modelabstractThe block-based discrete cosine transform (BDCT) is often used in image and video coding. It may introduce block artifacts at low data rates that manifest themselves as an annoying discontinuity between adjacent blocks. In this paper, we address this problem by investigating a transform-domain Markov random field (TD-MRF) model. Based on this model, two block artifact reduction postprocessing methods are presented. The first method, referred to as TD-MRF, provides an efficient progressive transform-domain solution. Our experimental results show that TD-MRF can reduce up to 90% of the computational complexity compared with spatial-domain MRF (SD-MRF) methods while still achieving comparable visual quality improvements. We then discuss a hybrid framework, referred to as TSD-MRF, that exploits the advantages of both TD-MRF and SD-MRF. The experimental results confirm that TSD-MRF can improve visual quality both objectively and subjectively over SD-MRF methods. Zhen Li 0008, Edward J. Delp |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2005 | An enhancement of leaky prediction layered video codingabstractIn this paper, we focus on leaky prediction layered video coding (LPLC). LPLC includes a scaled version of the enhancement layer within the motion compensation (MC) loop to improve the coding efficiency while maintaining graceful recovery in the presence of error drift. However, there exists a deficiency inherent in the LPLC structure, namely that the reconstructed video quality from both the enhancement layer and the base layer cannot be guaranteed to be always superior to that of using the base layer alone, even when no drift occurs. In this paper, we: 1) highlight this deficiency using a formulation that describes LPL; 2) propose a general framework that applies to both LPLC and a multiple description coding scheme using MC and we use this framework to further confirm the existence of the deficiency in LPLC; and 3) furthermore, we propose an enhanced LPLC based on maximum-likelihood estimation to address the previously specified deficiency in LPLC. We then show how our new method performs compared to LPLC. Yuxin Liu 0006, Paul Salama, Zhen Li 0008, Edward J. Delp |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2004 | Rate-distortion bounds for motion compensated rate scalable video codersabstractIn this paper, we derive and evaluate theoretical rate-distortion performance bounds for scalable video compression algorithms which use a single motion-compensated prediction (MCP) loop. These bounds are derived using rate-distortion theory based on an optimum mean-square error (MSE) quantizer. By specifying translatory motion and using an approximation of the predicted error frame power spectral density, it is possible to derive parametric versions of the rate-distortion functions which are based solely on the input power spectral density and the accuracy of the motion-compensated prediction. The theory is applicable to systems which allow prediction drift, such as the SNR-scalability in MPEG-2, as well as those with zero prediction drift such as the MPEG-4 fine grained scalable standard. Gregory W. Cook, Josep Prades-Nebot, Edward J. Delp |
ICIP | 3 |
| 2004 | An image normalization based watermarking scheme robust to general affine transformation
Hwan Il Kang, Edward J. Delp |
ICIP | 2 |
| 2004 | Channel-aware rate-distortion optimized leaky motion prediction
Zhen Li 0008, Edward J. Delp |
ICIP | 2 |
| 2004 | Rate distortion analysis of leaky prediction layered video coding using quantization noise modelingabstractUnlike conventional layered scalable video coding, leaky prediction layered video coding (LPLC) introduces a leaky factor /spl alpha/, which takes on values in the range between 0 and 1, to partially include the enhancement layer in the motion compensation loop, hence obtaining a trade-off between coding efficiency and error resilience performance. In this paper, we use quantization noise modeling to theoretically analyze the rate distortion performance of LPLC. An alternative block diagram of LPLC is first developed, which significantly simplifies the theoretical analysis. Closed form expressions, as a function of the leaky factor, are derived for two scenarios, where drift error occurs in the enhancement layer and no drift occurs within the motion compensation loop. Theoretical results are evaluated with respect to the leaky factor, showing that a leaky factor of 0.4-0.6 is a good choice in terms of the overall rate distortion performance of LPLC. Yuxin Liu 0006, Josep Prades-Nebot, Paul Salama, Edward J. Delp |
ICIP | 4 |
| 2004 | Adaptive lossless video compression using an integer wavelet transformabstractIn this paper, we describe an adaptive lossless compression algorithm for color video sequences utilizing backward adaptive temporal prediction and an integer wavelet transform. We exploit two redundancies in color video sequences, specifically spatial and temporal redundancies. We show that an adaptive scheme exploiting the two redundancies has better compression performance than lossless compression of individual image frames. The result of the proposed scheme is compared to current video compression algorithms. Sahng-Gyu Park, Edward J. Delp, Haoping Yu |
ICIP | 2 |
| 2004 | Analysis of the efficiency of snr-scalable strategies for motion compensated video codersabstractIn this paper, an analysis of the efficiency of three signal-to-noise ratio (SNR) scalable strategies for motion compensated video coders and their non-scalable counterpart is presented. After assuming some models and hypotheses with respect to the signals and systems involved, we have obtained the SNR of each coding strategy as a function of the decoding rate. To validate our analysis, we have compared our theoretical results with data from encodings of real video sequences. Results show that our analysis describes qualitatively the performance of each scalable strategy, and therefore, it can be useful to understand main features of each scalable technique and what factors influence their efficiency. Josep Prades-Nebot, Gregory W. Cook, Edward J. Delp |
ICIP | 3 |
| 2004 | Detection of unique people in news programs using multimodal shot clusteringabstractIn this paper, we describe an approach that uses a combination of visual and audio features to cluster shots belonging to the same person in video programs. We use color histograms extracted from keyframes and faces, as well as cepstral coefficients derived from audio to calculate pairwise shot distances. These distances are then normalized and combined to a single confidence value which reflects our certainty that two shots contain the same person. We then use an agglomerative clustering algorithm to cluster shots based on these confidence values. We report the results of our system on a data set of approximately 8 hours of programming. Cüneyt M. Taskiran, Alberto Albiol, Edward J. Delp |
ICIP | 4 |
| 2004 | Markov random field estimation of lost DCT coefficients in JPEG due to packet errorsabstractInterleaving is used before the encoding of source symbols in JPEG to reduce visual artifacts due to lost packets because interleaving distributes the locations of errors. The recovery of lost DCT coefficients in interleaved image compression is investigated in this paper. To restore the lost coefficients, an Maximum a Posteriori (MAP) estimate for the DCT coefficients is proposed. Under the assumption of a Gauss-Markov Random Field (GMRF) model in the pixel domain, the MAP estimate for the lost DCT coefficients is derived. Jinwha Yang, Edward J. Delp |
ICIP | 2 |
| 2004 | Statistical motion prediction with driftabstractIn this paper we present a statistical analysis of motion prediction with drift in video coding. The drift effect occurs when the video decoder does not have access to the same reference information used in the encoder in a hybrid video codec using motion prediction. Although the drift effect has been known to the video research community for a long time, there has not been a systematic theoretical treatment of this mechanism. Generally the performance of motion prediction with drift is evaluated experimentally. In this paper we derive a closed-form expression for the drift error. Based on this result, we derive an efficient rate distortion optimization given the statistical knowledge of the channel. Zhen Li 0008, Edward J. Delp |
VCIP | 2 |
| 2004 | Universal motion predictionabstractThis paper investigates the motion prediction techniques used in hybrid video coding. We first present a unified interpretation of motion prediction in terms of the prediction of motion threads. It is demonstrated that most current motion prediction techniques can be regarded as linear predictors of motion threads. Based on this new interpretation of motion prediction, we discuss the optimal motion predictor in the framework of Markov universal prediction. We define Markov predictability in a way that it upper bounds the optimal prediction performance in perfect reconstruction scenario. Since most current video applications use lossy coding, this results in imperfect reconstructions of the motion threads used in prediction. However, the optimality with the above perfect reconstruction scenario still holds in this case in an almost sure sense. Zhen Li 0008, Edward J. Delp |
VCIP | 2 |
| 2004 | Rate distortion analysis of layered video coding by leaky predictionabstractLeaky prediction layered video coding (LPLC) partially includes the enhancement layer in the motion compensated prediction loop, by using a leaky factor between 0 and 1, to balance the coding efficiency and error resilience performance. In this paper, rate distortion functions are derived for LPLC from rate distortion theory. Closed form expressions are obtained for two scenarios of LPLC, one where the enhancement layer stays intact and the other where the enhancement layer suffers from data rate truncation. The rate distortion performance of LPLC is then evaluated with respect to different choices of the leaky factor, demonstrating that the theoretical analysis well conforms with the operational results. Yuxin Liu 0006, Paul Salama, Gregory W. Cook, Edward J. Delp |
VCIP | 4 |
| 2004 | An overview of problems in image-based location awareness and navigationabstractIn this paper we describe some of the research issues and challenges in image-based location awareness and navigation. We will describe two systems being developed at Purdue University as testbeds for our ideas. The main system architecture combines image processing, mobility, wireless communication, and location awareness. We will describe two fundamental scenarios for using images to aid in mobile navigation problems. The first provides the ability to use a locally acquired image to determine the identity of an object, for example a building, as one roams in an area. The second problem is the use of images in a database to aid in vehicle navigation. The solution to these problems use location information, such as GPS signals, to compare and search location-annotated images in a database. We believe location information can improve the accuracy in image database search. Yung-Hsiang Lu, Edward J. Delp |
VCIP | 2 |
| 2004 | Adaptive lossless video compressionabstractIn this paper, we describe a new lossless compression algorithm for color video sequences. Our approach is to exploit both spatial and temporal redundancies in the video sequence by developing a new adaptive algorithm. By adaptive selection between spatial and temporal prediction we show that our scheme is better than state-of-the-art lossless compression algorithms. Sahng-Gyu Park, Edward J. Delp |
VCIP | 2 |
| 2004 | Benchmarking of Image Watermarking Algorithms for Digital Rights ManagementabstractWe discuss the issues related to image watermarking benchmarking and scenarios based on digital rights management requirements. We show that improvements are needed in image quality evaluation, especially related to image geometrical deformation assessments, in risk evaluation related to specific delivery scenarios and in multidimensional criteria evaluation. Efficient benchmarking is still an open issue and we suggest the use of open-source Web-based evaluation systems for collective progress in this domain. Benoît Macq, Jana Dittmann, Edward J. Delp |
Proc. IEEE | 3 |
| 2004 | ViBE: a compressed video database structured for active browsing and searchabstractIn this paper, we describe a unique new paradigm for video database management known as ViBE (video indexing and browsing environment). ViBE is a browseable/searchable paradigm for organizing video data containing a large number of sequences. The system first segments video sequences into shots by using a new feature vector known as the Generalized Trace obtained from the DC-sequence of the compressed data. Each video shot is then represented by a hierarchical structure known as the shot tree. The shots are then classified into pseudo-semantic classes that describe the shot content. Finally, the results are presented to the user in an active browsing environment using a similarity pyramid data structure. The similarity pyramid allows the user to view the video database at various levels of detail. The user can also define semantic classes and reorganize the browsing environment based on relevance feedback. We describe how ViBE performs on a database of MPEG sequences. Cüneyt M. Taskiran, Jau-Yuen Chen, Alberto Albiol, Charles A. Bouman, Edward J. Delp |
IEEE Trans. Multim. | 6 |
| 2003 | The indexing of persons in news sequences using audio-visual dataabstractWe describe a video indexing system that automatically searches for a specific person in a news sequence. The proposed approach combines audio and video confidence values extracted from speaker and face recognition analysis. The system also incorporates a shot selection module that seeks for anchors, where the person on the scene is likely speaking. The system has been extensively tested on several news sequences with very good recognition rates. Alberto Albiol, Edward J. Delp |
ICASSP (3) | 3 |
| 2003 | Wavelet video coding via a spatially adaptive lifting structureabstractWe present a spatially adaptive wavelet video coding technique with an update-first lifting structure. A common problem in many adaptive-transform frameworks is the introduction of a large overhead to address side information. We demonstrate that our structure does not need to transmit any side information to synchronize the encoder and decoder. We incorporate this technique in a motion compensated wavelet video codec. The experimental results confirm the performance improvement. Zhen Li 0008, Feng Wu 0001, Shipeng Li 0001, Edward J. Delp |
ICASSP (3) | 4 |
| 2003 | Evaluation of joint source and channel coding over wireless networksabstractTransmission of digital video signals over wireless networks demands efficient compression algorithms as well as reliable coding strategies. In this project, we provide a thorough evaluation of the joint source-channel video coding methodology from two points of view: source coding design for error resilience and channel coding for error detection and recovery. We investigate the current ITU video compression standard, H.263+, for 3G wireless transmission. In particular, we concentrate on error resilient features provided within the standard and forward error correction (FEC) to find the optimal combination of various system parameters under different lossy channel conditions. Yuxin Liu 0006, Christine Podilchuk, Edward J. Delp |
ICASSP (5) | 3 |
| 2003 | L-TFRC: an end-to-end congestion control mechanism for video streaming over the InternetabstractReal-time multimedia applications over the Internet have posed a lot of challenges due to the lack of quality of service (QoS) guarantees, frequent fluctuations in channel bandwidth, and packet losses. To address these issues, a great deal of research has been done in both video coding and video transmission fields. In this paper we present a logarithm-based TCP-friendly rate control (L-TFRC) mechanism, which can estimate the available bandwidth more accurately and improve the smoothness of the multimedia streaming significantly. We also apply it to a progressive fine granularity scalable (PFGS)-based video streaming. Both simulations and experiments over the Internet confirm the performance of L-TFRC. Zhen Li 0008, Guobin Shen, Shipeng Li 0001, Edward J. Delp |
ICME | 4 |
| 2003 | A discussion of leaky prediction based scalable codingabstractIn this paper, we focus on the leaky prediction based scalable coding (LPSC) structure and present a general framework for LPSC. We demonstrate the similarity between LPSC and motion compensation based multiple description coding scheme. We show that since the information contained in the enhancement layer in LPSC is actually a mismatch between two descriptions for each frame, it cannot be guaranteed that the enhancement layer always achieves superior reconstruction quality beyond that achieved by the base layer. We derive three reconstructions for each frame under the LPSC framework, and propose a maximum-likelihood (ML) estimation scheme for LPSC video reconstruction at the decoder. This generally achieves superior decoded video quality than both the enhancement layer and the base layer. Yuxin Liu 0006, Zhen Li 0008, Paul Salama, Edward J. Delp |
ICME | 4 |
| 2003 | Security of digital entertainment content from creation to consumption
Ahmet M. Eskicioglu, John Town, Edward J. Delp |
Signal Process. Image Commun. | 3 |
| 2002 | An automatic face detection and recognition system for video indexing applicationsabstractThe objective of this work is the integration and optimization of an automatic face detection and recognition system for video indexing applications. The system is composed of a face detection stage presented previously which provides good results maintaining a low computational cost. The recognition stage is based on the Principal Components Analysis (PCA) approach which has been modified to cope with the video indexing application. After the integration of the two stages, several improvements are proposed which increase the face detection and recognition rate and the overall performance of the system. Good results have been obtained using the MPEG-7 video content set used in the MPEG-7 evaluation group. Emiliano Acosta, Alberto Albiol, Edward J. Delp |
ICASSP | 4 |
| 2002 | Video preprocessing for audiovisual indexingabstractIn this paper we address the problem of reducing the amount of data to be analyzed when indexing audiovisual secuences. The reduction is based in selecting those parts of the video sequence where it is likely that there is a face talking. This can be useful since usually this kind of scenes contain important and reusable information such as interviews. The proposed technique is based on our a priori knowledge of the editing techniques used in news sequences. The results show that with this algorithm it is possible to discard arround 76% of the news sequence with minimal processing. Alberto Albiol, Edward J. Delp |
ICASSP | 3 |
| 2002 | A comparison of fixed-point 2D 9×7 discrete wavelet transform implementationsabstractWe describe three 2D discrete wavelet transform fixed-point implementations and compare them in terms of quantization error for the Daubechies 9/spl times/7 filter bank. The three implementations are the polyphase form, lifting scheme, and reduced scaling lifting scheme. Experimental results show that the reduced scaling lifting scheme is more robust than the other schemes. Also, the number of cycles the implementations take on a Texas Instruments TMS320C6201 simulator are given as reference. Hyung Cook Kim, Edward J. Delp |
ICIP (1) | 2 |
| 2002 | Nested interleaving transcoder for MPEG-4 Simple Profile bitstreamabstractWe propose a nested interleaving scheme in order to provide three levels of synchronization in a compressed video bitstream for low data rate wireless applications. The first level of interleaving specifies the starting position of each macroblock (MB) in a video frame. The second level of interleaving provides the start position of each MB header codeword and the start position of each DCT block. The start position of each VLC in a DCT block is located in a known position by the third level of interleaving. Our scheme is assumed to operate in the form of a transcoder, placed before and after the channel, of a MPEG-4 Simple Profile bitstream. Since the three level interleaving provides synchronization on a VLC scale, syntax-based bit error detections can be done in a VLC unit. The detected errors can be repaired syntactically so that the transcoder generates a MPEG-4 compliant bitstream which can then be decoded with a standard MPEG-4 decoder such as MoMuSys (FDIS V1.0). Jinwha Yang, Edward J. Delp |
ICIP (1) | 2 |
| 2002 | Combining audio and video for video sequence indexing applicationsabstractWe address the problem of detecting shots of subjects that are interviewed in news sequences. This is useful since usually these kinds of scenes contain important and reusable information that can be used for other news programs. In a previous paper, we presented a technique based on a priori knowledge of the editing techniques used in news sequences which allowed a fast search of news stories. We present a new shot descriptor technique which improves the previous search results by using a simple, yet efficient algorithm, based on the information contained in consecutive frames. Results are provided which prove the validity of the approach. Alberto Albiol, Edward J. Delp |
ICME (2) | 3 |
| 2002 | An integrated approach to encrypting scalable videoabstractScalable video compression is the encoding of a single video stream in multiple layers, each layer with its own bit rate. Because of the computational complexity of full video encryption, partial encryption has emerged as a general trend for both standard and scalable video codecs. Depending on the application, a particular layer of the video stream is chosen for encryption. In some applications, however, more than one video layer may need to be protected. This results in a more complicated key management as multiple keys are needed. In this paper, we present an integrated approach to encrypting multiple layers. Our proposal is a prepositioned shared secret scheme that enables the reconstruction of different keys by communicating different activating shares for the same prepositioned information. It presents certain advantages over three other key management schemes. Ahmet M. Eskicioglu, Edward J. Delp |
ICME (1) | 2 |
| 2002 | MAP-based post processing of video sequences using 3-D Huber-Markov random field modelabstractThe block DCT (BDCT) is by far one of the most popular transforms used in image and video coding. However, it introduces a noticeable blocking artifact at low data rates. A great deal of work has been done to remove the artifact with information extracted from the spatial and frequency domains. In this paper we address the video sequence restoration problem as a 3D Huber-Markov random field model and derive the temporal extension to traditional maximum a posteriori (MAP)-based methods. Two schemes, we call temporal MAP (TMAP) and motion compensated TMAP (MC-TMAP) respectively, are presented. We test our methods on MPEG-2 compressed sequences and evaluate their performances with traditional MAP restoration. Experimental results confirm that our schemes can significantly improve the visual quality of the reconstructed sequences. Zhen Li 0008, Edward J. Delp |
ICME (1) | 2 |
| 2002 | Error resilience of video transmission by rate-distortion optimization and adaptive packetizationabstractWe propose new schemes to introduce error resilience into the compressed video bitstreams for transmission over packet networks. First, we develop an adaptive packetization scheme that prohibits any dependency across packets, for error resilience purposes, while exploiting the dependency within each packet to improve the source coding performance. Secondly, we address a two-layer rate-distortion optimization scheme to serve our packetization method. We also use forward error correction (FEC) coding across packets to provide further error protection. Finally, we present a simplified version of our schemes to make it fully compliant with the current ITU video coding standard - H.263+. Yuxin Liu 0006, Paul Salama, Edward J. Delp |
ICME (2) | 3 |
| 2002 | Rate control for fully fine-grained scalable video coders
Josep Prades-Nebot, Gregory W. Cook, Edward J. Delp |
VCIP | 3 |
| 2002 | MPEG-4 simple profile transcoder for low-data-rate wireless applications
Jinwha Yang, Edward J. Delp |
VCIP | 2 |
| 2001 | Optimum color spaces for skin detectionabstractThe objective of this paper is to show that for every color space there exists an optimum skin detector scheme such that the performance of all these skin detectors schemes is the same. To that end, a theoretical proof is provided and experiments are presented which show that the separability of the skin and no skin classes is independent of the color space chosen. Alberto Albiol, Edward J. Delp |
ICIP (1) | 3 |
| 2001 | An unsupervised color image segmentation algorithm for face detection applicationsabstractThis paper presents an unsupervised color segmentation technique to divide skin detected pixels into a set of homogeneous regions which can be used in face detection applications or any other application which may require color segmentation. The algorithm is carried out in a two stage processing, where the chrominance and luminance information are used consecutively. For each stage a novel algorithm which combines pixel and region based color segmentation techniques is used. The algorithm has proven to be effective under a large number of test images. Alberto Albiol, Edward J. Delp |
ICIP (2) | 3 |
| 2001 | A hybrid embedded video codec using base layer information for enhancement layer codingabstractScalability has become an important feature for video coding algorithm design for heterogeneous networks due to variable bandwidth, variable wired and wireless network conditions and variable terminal capabilities. However, the best video compression algorithms are based on temporal prediction through motion compensation which do not lend themselves naturally to a scalable framework. Hybrid coding schemes are being introduced that combine a base-layer motion-compensation (MC) coder with an enhancement layer which offers fine-grain scalability through an embedded or progressive coder. Such a framework usually results in some compression efficiency loss over the best single layer motion compensated scheme Here, we introduce a hybrid coding scheme which combines a base-layer MC coder and a finely scalable enhancement layer where the information from the base-layer is used to determine the location of the high energy signal in the enhancement layer. This coder provides better compression results than embedded approaches which do not rely on base layer information. Eugene T. Lin, Christine Podilchuk, Arnaud E. Jacquin, Edward J. Delp |
ICIP (2) | 4 |
| 2001 | Error resilience and concealment in embedded zerotree wavelet codecsabstractIn ATM networks cell loss or channel errors can cause data to be dropped in the channel. When digital images/video are transmitted over these networks one must be able to reconstruct the missing data so that the impact of the errors is minimized. We overview the problem of using EZW encoders in channels where data-loss is possible. We also describe an error resilience scheme based on unequal error protection and data interleaving that addresses the problem of using rate scalable encoders over ATM networks. Paul Salama, Ness Shroff, Edward J. Delp |
ICIP (3) | 3 |
| 2001 | Is streaming media becoming mainstream?
Savitha Srinivasan, Dulce B. Ponceleon, Dick C. A. Bulterman, Edward J. Delp, Alexandros Eleftheriadis, Pablo Fernicola, Rob Lanphier, See-Mong Tan |
ACM Multimedia | 4 |
| 2001 | Investigation of robust video streaming using a wavelet-based rate-scalable codec
Gregory W. Cook, Eduardo Asbun, Edward J. Delp |
VCIP | 3 |
| 2001 | An overview of multimedia content protection in consumer electronics devices
Ahmet M. Eskicioglu, Edward J. Delp |
Signal Process. Image Commun. | 2 |
| 2001 | Multiresolution detection of spiculated lesions in digital mammogramsabstractWe present a novel multiresolution scheme for the detection of spiculated lesions in digital mammograms. First, a multiresolution representation of the original mammogram is obtained using a linear phase nonseparable two-dimensional (2-D) wavelet transform. A set of features is then extracted at each resolution in the wavelet pyramid for every pixel. This approach addresses the difficulty of predetermining the neighborhood size for feature extraction to characterize objects that may appear in different sizes. Detection is performed from the coarsest resolution to the finest resolution using a binary tree classifier. This top-down approach requires less computation by starting with the least amount of data and propagating detection results to finer resolutions. Experimental results using the MIAS image database have shown that this algorithm is capable of detecting spiculated lesions of very different sizes at low false positive rates. Charles F. Babbs, Edward J. Delp |
IEEE Trans. Image Process. | 3 |
| 2000 | A Simple and Efficient Face Detection Algorithm for Video Database ApplicationsabstractThe objective of this work is to provide a simple and yet efficient tool to detect human faces in video sequences. This information can be very useful for many applications such as video indexing and video browsing. In particular the paper focuses on the significant improvements made to our face detection algorithm presented by Albiol, Bouman and Delp (see IEEE Int. Conference on Image Processing, Kobe, Japan, 1999). Specifically, a novel approach to retrieve skin-like homogeneous regions is presented, which is later used to retrieve face images. Good results have been obtained for a large variety of video sequences. Alberto Albiol, Charles A. Bouman, Edward J. Delp |
ICIP | 4 |
| 2000 | A Rate-Distortion Approach to Wavelet-Based Encoding of Predictive Error FramesabstractWe develop a framework for efficiently encoding predictive error frames (PEF) as part of a rate scalable, wavelet-based video compression algorithm. We investigate the use of rate-distortion analysis to determine the significance of coefficients in the wavelet decomposition. Based on this analysis, we allocate the bit budget assigned to a PEF to the coefficients that yield the largest reduction in distortion, while maintaining the embedded and rate scalable properties of our video compression algorithm. Eduardo Asbun, Paul Salama, Edward J. Delp |
ICIP | 3 |
| 2000 | Error concealment in MPEG video streams over ATM networksabstractWhen transmitting compressed video over a data network, one has to deal with how channel errors affect the decoding process. This is particularly a problem with data loss or erasures. In this paper we describe techniques to address this problem in the context of asynchronous transfer mode (ATM) networks. Our techniques can be extended to other types of data networks such as wireless networks. In ATM networks channel errors or congestion cause data to be dropped, which results in the loss of entire macroblocks when MPEG video is transmitted. In order to reconstruct the missing data, the location of these macroblocks must be known. We describe a technique for packing ATM cells with compressed data, whereby the location of missing macroblocks in the encoded video stream can be found. This technique also permits the proper decoding of correctly received macroblocks, and thus prevents the loss of ATM cells from affecting the decoding process. The packing strategy can also be used for wireless or other types of data networks. We also describe spatial and temporal techniques for the recovery of lost macroblocks. In particular, we develop several optimal estimation techniques for the reconstruction of missing macroblocks that contain both spatial and temporal information using a Markov random field model. We further describe a sub-optimal estimation technique that can be implemented in real time. Paul Salama, Ness Shroff, Edward J. Delp |
IEEE J. Sel. Areas Commun. | 3 |
| 2000 | The EM/MPM algorithm for segmentation of textured images: analysis and further experimental resultsabstractIn this paper we present new results relative to the "expectation-maximization/maximization of the posterior marginals" (EM/MPM) algorithm for simultaneous parameter estimation and segmentation of textured images. The EM/MPM algorithm uses a Markov random field model for the pixel class labels and alternately approximates the MPM estimate of the pixel class labels and estimates parameters of the observed image model. The goal of the EM/MPM algorithm is to minimize the expected value of the number of misclassified pixels. We present new theoretical results in this paper which show that the algorithm can be expected to achieve this goal, to the extent that the EM estimates of the model parameters are close to the true values of the model parameters. We also present new experimental results demonstrating the performance of the EM/MPM algorithm. Mary L. Comer, Edward J. Delp |
IEEE Trans. Image Process. | 2 |
| 1999 | Face Detection for Pseudo-Semantic Labeling in Video DatabasesabstractPseudo-semantic labeling represents a novel approach for automatic content description of video. This information can be used in the context of a video database to improve browsing and searching. In this paper we describe our work on using face detection techniques for pseudo-semantic labeling. We present our results using a database of MPEG sequences. Alberto Albiol, Charles A. Bouman, Edward J. Delp |
ICIP (3) | 3 |
| 1999 | Encoding of Predictive Error Frames in Rate Scalable Video Codecs Using Wavelet ShrinkageabstractRate scalable video compression is appealing for low bit rate applications, such as video telephony and wireless communication, where bandwidth available to an application cannot be guaranteed. In this paper, we investigate a set of strategies to increase the performance of SAMCoW, a rate scalable encoder. These techniques are based on based on wavelet decomposition, spatial orientation trees, and motion compensation. Eduardo Asbun, Paul Salama, Edward J. Delp |
ICIP (3) | 3 |
| 1999 | Perceptual watermarks for digital images and videoabstractThe growth of new imaging technologies has created a need for techniques that can be used for copyright protection of digital images and video. One approach for copyright protection is to introduce an invisible signal, known as a digital watermark, into an image or video sequence. In this paper, we describe digital watermarking techniques, known as perceptually based watermarks, that are designed to exploit aspects of the the human visual system in order to provide a transparent (invisible), yet robust watermark. In the most general sense, any watermarking technique that attempts to incorporate an invisible mark into an image is perceptually based. However, in order to provide transparency and robustness to attack, two conflicting requirements from a signal processing perspective, more sophisticated use of perceptual information in the watermarking process is required. We describe watermarking techniques ranging from simple schemes which incorporate common-sense rules in using perceptual information in the watermarking process, to more elaborate schemes which adapt to local image characteristics based on more formal perceptual models. This review is not meant to be exhaustive; its aim is to provide the reader with an understanding of how the techniques have been evolving as the requirements and applications become better defined. Raymond B. Wolfgang, Christine Podilchuk, Edward J. Delp |
Proc. IEEE | 3 |
| 1999 | Wavelet based rate scalable video compressionabstractIn this paper, we present a new wavelet based rate scalable video compression algorithm. We will refer to this new technique as the scalable adaptive motion compensated wavelet (SAMCoW) algorithm. SAMCoW uses motion compensation to reduce temporal redundancy. The prediction error frames and the intracoded frames are encoded using an approach similar to the embedded zerotree wavelet (EZW) coder. An adaptive motion compensation (AMC) scheme is described to address error propagation problems. We show that, using our AMC scheme, the quality of the decoded video can be maintained at various data rates. We also describe an EZW approach that exploits the interdependency between color components in the luminance/chrominance color space. We show that, in addition to providing a wide range of rate scalability, our encoder achieves comparable performance to the more traditional hybrid video coders, such as MPEG1 and H.263. Furthermore, our coding scheme allows the data rate to be dynamically changed during decoding, which is very appealing for network-oriented applications. Ke Shen 0002, Edward J. Delp |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | Segmentation of textured images using a multiresolution Gaussian autoregressive modelabstractWe present a new algorithm for segmentation of textured images using a multiresolution Bayesian approach. The new algorithm uses a multiresolution Gaussian autoregressive (MGAR) model for the pyramid representation of the observed image, and assumes a multiscale Markov random field model for the class label pyramid. The models used in this paper incorporate correlations between different levels of both the observed image pyramid and the class label pyramid. The criterion used for segmentation is the minimization of the expected value of the number of misclassified nodes in the multiresolution lattice. The estimate which satisfies this criterion is referred to as the "multiresolution maximization of the posterior marginals" (MMPM) estimate, and is a natural extension of the single-resolution "maximization of the posterior marginals" (MPM) estimate. Previous multiresolution segmentation techniques have been based on the maximum a posterior (MAP) estimation criterion, which has been shown to be less appropriate for segmentation than the MPM criterion. It is assumed that the number of distinct textures in the observed image is known. The parameters of the MGAR model-the means, prediction coefficients, and prediction error variances of the different textures-are unknown. A modified version of the expectation-maximization (EM) algorithm is used to estimate these parameters. The parameters of the Gibbs distribution for the label pyramid are assumed to be known. Experimental results demonstrating the performance of the algorithm are presented. Mary L. Comer, Edward J. Delp |
IEEE Trans. Image Process. | 2 |
| 1998 | Integrating engineering design, signal processing, and community service in the EPICS programabstractOne of the most challenging problems in engineering-and signal processing-education is providing realistic and meaningful design experience. In the Engineering Projects in Community Service (EPICS) program, teams of engineering undergraduates earn academic credit for multi-year projects that solve technology-based problems for community organizations. Key features of EPICS include the long-term nature of the projects; emphasis on "real-world" start-to-finish design; the learning experience embodied in solving ambitious engineering problems; vertical, multidisciplinary teams; development of teamwork and communication skills; and the use of engineering to help the community. We describe the EPICS program and highlight four EPICS signal processing projects: a real-time system to measure speaking rate for Purdue's speech clinic; voice-controlled interactive software to encourage speech in developmentally delayed children; a microphone array hearing aid; and a virtual museum tour and interactive Web-based history games for the Tippecanoe County Historical Association. Leah H. Jamieson, Edward J. Coyle, Mary P. Harper, Edward J. Delp, Patricia N. Davies |
ICASSP | 4 |
| 1998 | Video scene change detection using the generalized sequence traceabstractWe propose an algorithm to detect scene changes in a video sequence in the compressed domain. We define a feature vector extracted from each frame that we call the generalized trace. We examine various ways of processing the generalized trace to determine the temporal location of scene changes in a video stream. Cüneyt M. Taskiran, Edward J. Delp |
ICASSP | 2 |
| 1998 | Very Low Bit Rate Wavelet-based Scalable Video Compression
Eduardo Asbun, Paul Salama, Ke Shen 0002, Edward J. Delp |
ICIP (3) | 4 |
| 1998 | A Gaussian Mixture Model for Edge-Enhanced Images with Application to Sequential Edge Detection and LinkingabstractWe present a new stochastic model for pixels in an edge-enhanced image. The model is robust because it allows for the possibilities of false and multiple edges, and may be efficiently estimated using a expectation-maximization technique with a minimum description length metric. The direct applicability of the model for the sequential edge linking algorithm is investigated. Gregory W. Cook, Edward J. Delp |
ICIP (2) | 2 |
| 1998 | Normal Mammogram Analysis and RecognitionabstractPresents a novel approach to the problem of computer aided analysis of digital mammograms for breast cancer detection: namely, the development of algorithms to recognize unequivocally normal mammograms. First, the authors eliminate amorphous "clouds" or "blobs" in mammograms produced by normal glandular tissue of varying density using local average subtraction. Then the authors identify and remove normal connective tissue markings based on a set of specially designed line detectors. Any abnormality that may exist in the mammogram is therefore enhanced in the residual image, which makes the decision regarding the normality of the mammogram, much easier. Charles F. Babbs, Edward J. Delp |
ICIP (1) | 3 |
| 1998 | A Compressed Video Database Structured for Active Browsing and Search
Cüneyt M. Taskiran, Jau-Yuen Chen, Charles A. Bouman, Edward J. Delp |
ICIP (3) | 4 |
| 1998 | The Effect of Matching Watermark and Compression Transforms in Compressed Color ImagesabstractThe growth of networked multimedia systems has complicated copyright enforcement relative to digital images. One way to protect the copyright of digital images is to add an invisible structure to the image (known as a digital watermark) to identify the owner. In particular, it is important for Internet and image database applications that as much of the watermark as possible remain in the image after compression. Image adaptive watermarks are particularly resistant to removal by signal processing attack such as filtering or compression. Common image adaptive watermarks operate in the transform domain (DCT or wavelet); the same domains are also used for popular image compression techniques (JPEG, EZW). This paper investigates whether matching the watermarking domain to the compression transform domain will make the watermark more robust to compression. Raymond B. Wolfgang, Christine Podilchuk, Edward J. Delp |
ICIP (1) | 3 |
| 1997 | Multiresolution Detection of Stellate Lesions in MammogramsabstractPresents a new multiresolution scheme for the detection of stellate lesions in digital mammograms. First, a multiresolution representation of the original mammogram is obtained using a linear phase nonseparable 2-D wavelet transform. A set of features are then extracted at each resolution for every pixel. This addresses the difficulty of predetermining the neighborhood size for feature extraction to characterize objects that may appear with different sizes. Detection is performed from the coarsest resolution to the finest resolution using binary tree classifiers. This top-down approach requires less computation by starting with the least amount of data and propagating detection results to finer resolutions. Experimental results on the MIAS image database have shown that this algorithm is capable of detecting stellate lesions of very different sizes. Edward J. Delp |
ICIP (2) | 2 |
| 1997 | A Fast Suboptimal Approach to Error Concealment in Encoded Video StreamsabstractIn ATM networks cell loss or channel errors can cause data to be dropped in the channel. When digital video is transmitted over these networks one must be able to reconstruct the missing data so that the impact of these errors is minimized. In this paper we describe a Bayesian approach to concealing these errors by post-processing the received data. In a previous paper (see IEEE Proc. Int. Conf. on Image Processing p.49-52, 1996), each frame in the sequence was modeled as a Markov random field, and maximum a posteriori estimates of the missing macroblocks were obtained. However, the maximum a posteriori estimate is not unique, and the algorithm is also computationally intensive. In this paper we demonstrate, that by using median filtering we arrive at a suboptimal estimate. This will allow real-time nearly optimal reconstruction of the missing data. Paul Salama, Ness Shroff, Edward J. Delp |
ICIP (2) | 3 |
| 1997 | Color Image Compression Using an Embedded Rate Scalable ApproachabstractIn this paper we propose a wavelet based coding algorithm for color images using a luminance/chrominance color space. Data rate scalability is achieved by using an embedded coding scheme, which is similar to Shapiro's (1993) embedded zerotree wavelet (EZW) algorithm. In a luminance/chrominance color space, the three color components have little statistical correlation. However, observations are made that at the spatial locations where chrominance signals have large transitions, it is highly likely for the luminance signal to have large transitions. This interdependence between the color components is exploited in the algorithm. Ke Shen 0002, Edward J. Delp |
ICIP (3) | 2 |
| 1996 | Video and image systems engineering education for the 21st centuryabstractWe are developing a new graduate program at Purdue in Video and Image Systems Engineering (VISE). The project is comprised of three parts: a new curriculum centered around a degree option in VISE to be earned as part of the Masters or Ph.D. degrees; a state-of-the-art lecture/laboratory facility for instruction, laboratory experiments, and project and homework activities in VISE courses; and enhancement of existing courses and development of new courses in the VISE area. Jan P. Allebach, Charles A. Bouman, Edward J. Coyle, Edward J. Delp, David A. Landgrebe, Anthony A. Maciejewski, Zygmunt Pizlo, Ness Shroff, Michael D. Zoltowski |
ICIP (1) | 4 |
| 1996 | The EM/MPM algorithm for segmentation of textured images: analysis and further experimental resultsabstractIn this paper we present new results relative to the "expectation-maximization/maximization of the posterior marginals" (EM/MPM) algorithm for simultaneous parameter estimation and segmentation of textured images. The goal of the EM/MPM algorithm is to minimize the expected value of the number of misclassified pixels. We present new theoretical results in this paper which show that the algorithm can be expected to achieve this goal, to the extent that the EM estimates of the model parameters are close to the true values of the model parameters. We also present new experimental results demonstrating the performance of the algorithm. Mary L. Comer, Edward J. Delp |
ICIP (3) | 2 |
| 1996 | A Bayesian approach to error concealment in encoded video streamsabstractIn ATM networks cell loss causes data to be dropped in the channel. When digital video is transmitted over these networks one must be able to reconstruct the missing data so that the impact of these errors is minimized. In this paper we describe a Bayesian approach to conceal these errors. Assuming that the digital video has been encoded using the MPEG1 or MPEG2 compression scheme, each frame is modeled as a Markov random field. A maximum a posteriori estimate of the missing macroblocks and motion vectors is described based on the model. Paul Salama, Ness Shroff, Edward J. Delp |
ICIP (2) | 3 |
| 1996 | A control scheme for a data rate scalable video codecabstractWe present a control scheme for a rate scalable video codec. We describe a wavelet based video codec with motion compensation used to reduce temporal redundancy. The prediction error frames are encoded using an embedded zerotree wavelet (EZW) approach which allows data rate scalability. Since motion compensation is used in the algorithm, the duality of the decoded video may decay due to the propagation of errors in the temporal domain. An adaptive motion compensation scheme is proposed to address this problem. We show that using our control scheme the quality of the decoded video can be maintained at any data rate. Ke Shen 0002, Edward J. Delp |
ICIP (2) | 2 |
| 1996 | A watermark for digital imagesabstractThe growth of networked multimedia systems has magnified the need for image copyright protection. One approach used to address this problem is to add an invisible structure to an image that can be used to seal or mark it. These structures are known as digital watermarks. We describe two techniques for the invisible marking of images. We analyze the robustness of the watermarks with respect to linear and nonlinear filtering, and JPEG compression. The results show that our watermarks detect all but the most minute changes to the image. Raymond B. Wolfgang, Edward J. Delp |
ICIP (3) | 2 |
| 1996 | Issues in the design of studies to test the effectiveness of stereo imagingabstractRecently, there has been a great increase in interest in using three dimensional stereoscopic displays to provide viewers with realistic 3D views of objects of interest. Some applications where stereoscopic displays are becoming popular include medical visualization, visualization of meteorological data, and various virtual reality applications. To quantify the effectiveness of stereoscopic systems over conventional monoscopic systems, well-designed experiments and data analysis methods are necessary. This task requires the combined effort of application scientists and experts in experimental design. Lack of interdisciplinary collaboration is a primary weakness of many stereoscopic display studies, resulting in the neglect of many important but subtle experimental issues. In this paper, we discuss specific issues that arise in the design of studies to determine the effectiveness of digital stereo imagery. Issues concerning statistical analysis of the experimental data are also discussed. References to related literature from engineering, computer graphics, and psychophysics are given. The issues developed herein provide a guideline for the design of studies to compare observer performance when using different imaging modalities. Jean Hsu, Zygmunt Pizlo, David M. Chelberg, Charles F. Babbs, Edward J. Delp |
IEEE Trans. Syst. Man Cybern. Part A | 5 |
| 1995 | Multiresolution image segmentationabstractIn this paper we present a new algorithm for segmentation of noisy or textured images using a multiresolution Bayesian approach. Our algorithm is different from previously proposed multiresolution segmentation techniques in that we use a multiresolution Gaussian autoregressive (AR) model for the pyramid representation of the observed image. Our algorithm also approximates the "maximization of the posterior marginals" (MPM) estimate of the pixel class labels at each resolution, from coarsest to finest, unlike previously proposed techniques, which have been based on MAP estimation. Experimental results are presented to demonstrate the performance of the new algorithm. Mary L. Comer, Edward J. Delp |
ICASSP | 2 |
| 1995 | Multiresolution sequential edge linkingabstractWe describe a multiresolution approach to edge detection using a sequential search algorithm. The use of a multiresolution image pyramid allows the integration of global edge information contained in lower resolutions to guide the sequential search at higher resolutions. As a consequence, the dependence on a priori knowledge of the image edges is greatly reduced. Estimating the sequential search parameters from lower resolution images provides for a more accurate and less costly search of edge paths in the image. Gregory W. Cook, Edward J. Delp |
ICIP | 2 |
| 1995 | Error concealment techniques for encoded video streamsabstractIn this paper we describe two error-recovery approaches for MPEG encoded video over ATM networks. The first approach aims at reconstructing each lost pixel by spatial interpolation from the nearest undamaged pixels. The second approach recovers lost macroblocks by minimizing intersample variations within each block and across its boundaries. Moreover, a new technique for packing ATM cells with compressed data is also proposed. Paul Salama, Ness Shroff, Edward J. Coyle, Edward J. Delp |
ICIP | 4 |
| 1995 | A fast algorithm for video parsing using MPEG compressed sequencesabstractVideo parsing is a fundamental operation used in many digital video applications such as digital libraries and video servers. The accuracy and execution speed of the parsing algorithm is critical if large amounts of video data are to be processed, particularly in real-time. We present a new algorithm to reconstruct DC coefficient images of a DCT and motion compensation compressed video sequence, e.g. MPEG. The histograms of the DC coefficient images can be used to detect scene changes. Ke Shen 0002, Edward J. Delp |
ICIP | 2 |
| 1995 | An Investigation of Scalable SIMD I/O Techniques with Application to Parallel JPEG Compression
Gregory W. Cook, Edward J. Delp |
J. Parallel Distributed Comput. | 2 |
| 1995 | Edge linking by sequential search
Aly A. Farag, Edward J. Delp |
Pattern Recognit. | 2 |
| 1995 | Color image coding using morphological pyramid decompositionabstractPresents a new algorithm that utilizes mathematical morphology for pyramidal coding of color images. The authors obtain lossy color image compression by using block truncation coding at the pyramid levels to attain reduced data rates. The pyramid approach is attractive due to low computational complexity, simple parallel implementation, and the ability to produce acceptable color images at moderate data rates. In many applications, the progressive transmission capability of the algorithm is very useful. The authors show experimental results for color images at data rates of 1.89 bits/pixel. Lori A. Overturf, Mary L. Comer, Edward J. Delp |
IEEE Trans. Image Process. | 3 |
| 1995 | Preclinical ROC studies of digital stereomammographyabstractReports the diagnostic performance of observers in detecting abnormalities in computer-generated mammogram-like images. A mathematical model of the human breast is defined in which breast tissues are simulated by spheres of different sizes and densities. Images are generated by casting rays from a specified source, through the model, and onto an image plane. Observer performance when using two viewing modalities (stereo versus mono) is compared. In the stereo viewing mode, images are presented to the observer (wearing liquid-crystal display glasses), such that the left eye sees the left image only and the right eye sees the right image only. In this way, the images can be fused by the observer to obtain a sense of depth. In the mono viewing mode, identical images are presented to the left and right eyes so that no binocular disparities will be produced by the images. Observer response data are evaluated using receiver operating characteristic (ROC) analysis to characterize any difference in detectability of abnormalities (in either the density or the arrangement of simulated tissue densities) using the two viewing modes. The authors' experimental results indicate the clear superiority of stereo viewing for detection of arrangement abnormalities. For detection of density abnormalities, the performance of the two viewing modes is similar. These preliminary results suggest that stereomammography may permit easier detection of certain tissue abnormalities, perhaps providing a route to earlier tumor detection in cases of breast cancer. Jean Hsu, David M. Chelberg, Charles F. Babbs, Zygmunt Pizlo, Edward J. Delp |
IEEE Trans. Medical Imaging | 5 |
| 1994 | An investigation of JPEG image and video compression using parallel processingabstractThe problem inherent with any digital image (or digital video) system is the large amount of bandwidth required for transmission or storage. This has driven the research area of image compression to develop more complex algorithms that compress images to lower data rates with better fidelity. One approach that can be used to increase the execution speed of these complex algorithms is through the use of parallel processing. In this paper we address one aspect of the parallel implementation of the JPEG still image compression standard on the MasPar MP-1, a massively parallel SIMD computer. We develop a novel byte alignment algorithm used to efficiently output compressed data from the parallel system.> Gregory W. Cook, Edward J. Delp |
ICASSP (5) | 2 |
| 1994 | Parameter Estimation and Segmentation of Noisy or Textured Images using the EM Algorithm and MPM EstimationabstractPresents a new algorithm for segmentation of noisy or textured images using the expectation-maximization (EM) algorithm for estimating parameters of the probability mass function of the pixel class labels and the maximization of the posterior marginals (MPM) criterion for the segmentation operation. A Markov random field (MRF) model is used for the pixel class labels. The authors present experimental results demonstrating the use of the new algorithm on synthetic images and medical imagery.> Mary L. Comer, Edward J. Delp |
ICIP (2) | 2 |
| 1994 | A Tree Structured Bayesian Scalar Quantizer for Wavelet Based Image CompressionabstractMultiresolution image decompositions (e.g., wavelets), in conjunction with a variety of quantization schemes, have been shown to be very effective for image compression. Recently, several promising tree-structured quantization schemes that exploit the correlation across scales have been proposed. In this paper, we present an image compression algorithm based on a multiresolution Markov random field model used to model the correlations of wavelet coefficients across the scales. We also present experimental results obtained using the algorithm.> Birsen Yazici, Mary L. Comer, Rangasami L. Kashyap, Edward J. Delp |
ICIP (3) | 4 |
| 1994 | Parallel scalable libraries and algorithms for computer visionabstractWe describe a project that integrates applications requirements, parallel algorithm design, models of parallel computing, and software tools in order to improve the ability of applications researchers in the fields of computer vision and image processing (CVIP) to realize the performance potential of high performance parallel computers. This objective is achieved by pursuing four directions of research: the development of efficient and practical scalable algorithms for fundamental CVIP problems; the implementation of realistic CVIP scenarios on high performance computers; the development of a scalable, architecture-independent parallel model and a metric for estimating the performance of an algorithm on an existing parallel machine; and the development of prototype software tools. In this paper we describe work in each of these areas. Leah H. Jamieson, Edward J. Delp, Susanne E. Hambrusch, Ashfaq Khokhar 0001, Gregory W. Cook, Farooq Hameed, Jamshed N. Patel, Ke Shen 0002 |
ICPR (3) | 2 |
| 1994 | Coarse-grained algorithms and implementations of structural indexing-based object recognition on Intel Touchstone DeltaabstractIn this paper, we present efficient parallel solutions for structural indexing-based object recognition on coarse-grained parallel machines. Based on the analysis using the C/sup 6/-model, the parallel algorithm proposed in this paper takes O(Sk/p) computation units and O(p/sup 3/2/) communication units on a p processor coarse-grained machine such that 1/spl les/p/spl les/S, whereas the sequential solution takes O(Sk). The proposed solution is implemented on the Intel Touchstone Delta and performance results are shown. For a scene consisting of 160 feature points, the recognition phase takes 258 ms on a 64 processor Delta when the model database contains 256 models. The sequential algorithm takes approximately 4 seconds on a single node of the Delta. Ashfaq Khokhar 0001, Gregory W. Cook, Leah H. Jamieson, Edward J. Delp |
ICPR (3) | 4 |
| 1994 | Discontinuity preserving regularization of inverse visual problemsabstractThe method of Tikhonov regularization has been widely used to form well-posed inverse problems in low-level vision. The application of this technique usually results in a least squares approximation or a spline fitting of the parameter of interest. This is often adequate for estimating smooth parameter fields. However, when the parameter of interest has discontinuities the estimate formed by this technique will smooth over the discontinuities. Several techniques have been introduced to modify the regularization process to incorporate discontinuities. Many of these approaches however, will themselves be ill-posed or ill-conditioned. This paper presents a technique for incorporating discontinuities into the reconstruction problem while maintaining a well-posed and well-conditioned problem statement. The resulting computational problem is a convex functional minimization problem. This method is compared to previous approaches and examples are presented for the problems of reconstructing curves and surfaces with discontinuities and for estimating image data. Computational issues arising in both analog and digital implementations are also discussed.> Robert L. Stevenson, Barbara E. Schmitz, Edward J. Delp |
IEEE Trans. Syst. Man Cybern. | 3 |
| 1993 | Local estimation of Gaussian-based edge enhancement filters using Fourier analysis
Bin Wang 0002, D. M. Rose, Aly A. Farag, Edward J. Delp |
ICASSP (5) | 4 |
| 1993 | The Nonlinear Prefiltering and Difference of Estimates Approaches to Edge Detection: Applications of Stack Filters
Jisang Yoo, Charles A. Bouman, Edward J. Delp, Edward J. Coyle |
CVGIP Graph. Model. Image Process. | 3 |
| 1993 | Tree-structured piecewise linear adaptive equalizationabstractThe use of a tree-structured piecewise linear filter as an adaptive equalizer is proposed. In the tree equalizer, each node in a tree is associated with a linear filter restricted to a polygonal domain, and each subtree is associated with a piecewise linear filter. A training sequence is used to adaptively update the filter coefficients and domains at each node, and to select the best subtree and corresponding piecewise linear filter. The tree-structured approach offers several advantages. First, it makes use of standard linear adaptive filtering techniques at each node to find the corresponding conditional linear filter. Second, it allows for efficient selection of the subtree and corresponding piecewise linear filter of appropriate complexity. Overall, the approach is computationally efficient and conceptually simple. Numerical experiments are performed to show the advantages of tree-structured piecewise linear and piecewise decision feedback equalizers over linear, polynomial, and decision feedback equalizers for the equalization of channels with severe intersymbol interference.> Saul B. Gelfand, C. S. Ravishankar, Edward J. Delp |
IEEE Trans. Commun. | 3 |
| 1992 | Viewpoint Invariant Recovery of Visual Surfaces from Sparse DataabstractAn algorithm for the reconstruction of visual surfaces from sparse data is proposed. An important aspect of this algorithm is that the surface estimated from the sparse data is approximately invariant with respect to rigid transformation of the surface in 3D space. The algorithm is based on casting the problem as an ill-posed inverse problem that must be stabilized using a priori information related to the image and constraint formation. To form a surface estimate that is approximately invariant with respect to viewpoint, the stabilizing information is based on invariant surface characteristics. With appropriate approximations, this results in a convex functional to minimize, which is then solved using finite element analysis. The relationship of this algorithm to several previously proposed reconstruction algorithms is discussed, and several examples that demonstrate its effectiveness in reconstructing viewpoint-invariant surface estimates are given.> Robert L. Stevenson, Edward J. Delp |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1992 | A Cost Minimization Approach to Edge Detection Using Simulated AnnealingabstractThe authors cast edge detection as a problem in cost minimization. This is achieved by the formulation of a cost function that evaluates the quality of edge configurations. The function is a linear sum of weighted cost factors. The cost factors capture desirable characteristics of edges such as accuracy in localization, thinness, and continuity. Edges are detected by finding the edge configurations that minimize the cost function. The authors give a mathematical description of edges and analyze the cost function in terms of the characteristics of the edges in minimum cost configurations. Through the analysis, guidelines are provided on the choice of weights to achieve certain characteristics of the detected edges. The cost function is minimized by the simulated annealing method. A set of strategies is presented for generating candidate states and to devise a suitable temperature schedule.> Hin Leong Tan, Saul B. Gelfand, Edward J. Delp |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1991 | A tree-structured piecewise linear adaptive filterabstractThe idea of a tree-structured piecewise linear adaptive filter is proposed. The tree-structured filter is constructed as follows. Each node in a binary tree is associated with a linear filter restricted to a polygonal domain, and this is done in such a way that each subtree is associated with a piecewise linear filter. A training sequence is used to adaptively update the filter coefficients and domains at each node, and to select the best subtree and the corresponding piecewise linear filter. The authors summarize the derivation and analysis of the tree-structured piecewise linear adaptive filter. They then investigate the performance of the tree-structured filter as an adaptive echo canceler and compare it with that of linear and truncated Volterra series adaptive echo cancelers through computer simulations.> Saul B. Gelfand, C. S. Ravishankar, Edward J. Delp |
ICASSP | 3 |
| 1991 | Dynamic Intelligent Scheduling and Control of Reconfigurable Parallel Architectures for Computer Vision/Image Processing
Frank J. Weil, Leah H. Jamieson, Edward J. Delp |
J. Parallel Distributed Comput. | 3 |
| 1991 | An Iterative Growing and Pruning Algorithm for Classification Tree DesignabstractA critical issue in classification tree design-obtaining right-sized trees, i.e. trees which neither underfit nor overfit the data-is addressed. Instead of stopping rules to halt partitioning, the approach of growing a large tree with pure terminal nodes and selectively pruning it back is used. A new efficient iterative method is proposed to grow and prune classification trees. This method divides the data sample into two subsets and iteratively grows a tree with one subset and prunes it with the other subset, successively interchanging the roles of the two subsets. The convergence and other properties of the algorithm are established. Theoretical and practical considerations suggest that the iterative free growing and pruning algorithm should perform better and require less computation than other widely used tree growing and pruning algorithms. Numerical results on a waveform recognition problem are presented to support this view.> Saul B. Gelfand, C. S. Ravishankar, Edward J. Delp |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1991 | On detecting dominant points
Nirwan Ansari, Edward J. Delp |
Pattern Recognit. | 2 |
| 1991 | Moment preserving quantization [signal processing]abstractThe general solution to the moment-preserving (MP) quantizer problem is presented. It is shown that the moment preserving quantizer is related to the Gauss-Jacobi mechanical quadrature, the output levels of the N-level MP quantizer are the N zeros of an Nth degree orthogonal polynomial associated with the input probability distribution function, and the N-1 thresholds of the MP quantizer are related to the Christoffel numbers through the Chebyshev-Markov-Stieltjes separation theorem. The statistical convergence of the MP quantizer is investigated. MP quantizer tables are presented for the uniform, Gaussian, and Laplacian density functions. The moment-preserving quantizer is shown to be related to block truncation coding.> Edward J. Delp, Owen Robert Mitchell |
IEEE Trans. Commun. | 1 |
| 1990 | Viewpoint invariant recovery of visual surfaces from sparse dataabstractA new algorithm is described for the reconstruction of visual surfaces from sparse data. An important aspect of this algorithm is that the surface estimated from the sparse data is approximately invariant with respect to rigid transformation of the surface in three-dimensional space. To form a surface estimate which is invariant with respect to viewpoint the stabilizing information is based on invariant surface characteristics. With appropriate approximations this results in a convex functional which must be minimized. The relationship of this algorithm to several previously proposed reconstruction algorithms is discussed, and several examples are given which demonstrate the effectiveness of the proposed algorithm in reconstructing viewpoint invariant surface estimates.> Robert L. Stevenson, Edward J. Delp |
ICCV | 2 |
| 1990 | An Analysis of Fixed-Assignment Hypercube Partitioning
Frank J. Weil, Leah H. Jamieson, Edward J. Delp |
ICPP (1) | 3 |
| 1990 | Dynamic intelligent scheduling and control of reconfigurable parallel architectures for computer vision/image processingabstractA system for dynamic intelligent scheduling and control (DISC) of reconfigurable parallel processors is presented. The purpose of the system is to provide a rapid prototyping capability for computer vision/image processing tasks. The scheduler particularly addresses the problems of algorithms with execution times that depend on the image data and processing scenarios that vary dynamically based on the input image. Since conventional scheduling methods cannot propose schedules for most masks of this type, a dynamic controller is used to schedule the task and reconfigure the machine on the fly. This dynamic scheduling system attempts to balance the overall processing scenario with the needs of the individual routines that make up the task. The implementation of this system is discussed, with emphasis on the scheduling heuristics and the use of the system for prototyping computer vision/image processing tasks. Testing was done on a number of tasks that exercised different aspects of the scheduling strategy.> Frank J. Weil, Leah H. Jamieson, Edward J. Delp |
ICPR (2) | 3 |
| 1990 | The analysis of morphological filters with multiple structuring elements
Jisheng Song, Edward J. Delp |
Comput. Vis. Graph. Image Process. | 2 |
| 1990 | Partial Shape Recognition: A Landmark-Based ApproachabstractA method of recognizing partially occluded objects is presented in which each object is represented by a set of landmarks. Given a scene consisting of partially occluded objects, a model object in the scene is hypothesized by matching the landmarks of the model with those in the scene. A measure of similarity between two landmarks is needed to perform this matching. A local shape measure, sphericity, is introduced. It is shown that any invariant function under a similarity transformation is a function of the sphericity. To match landmarks between the model and the scene, a table of compatibility is constructed. A technique, known as hopping dynamic programming, is described to guide the landmark matching through the compatibility table. The location of the model in the scene is estimated with a least-squares fit among the matched landmarks. A heuristic measure is then computed to decide if the model is in the scene.> Nirwan Ansari, Edward J. Delp |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1990 | On the distribution of a deforming triangle
Nirwan Ansari, Edward J. Delp |
Pattern Recognit. | 2 |
| 1990 | Quantitative analysis of a moment-based edge operatorabstractAn operator that is based on the sample variance of a group of pixels is introduced. It exhibits three unique properties: freedom from a predefined shape, low computational complexity and a rigorous stochastic formulation. The utility of the latter property in applying detection-theoretic principles to the task of edge detection and in theoretical predictions of experimental performance measures is demonstrated. The operator is compared with other common edge operators on natural scenes.> Paul H. Eichel, Edward J. Delp |
IEEE Trans. Syst. Man Cybern. | 2 |
| 1989 | A cost minimization approach to edge detection using simulated annealingabstractEdge detection is analyzed as a problem in cost minimization. A cost function is formulated that evaluates the quality of edge configurations. A mathematical description of edges is given, and the cost function is analyzed in terms of the characteristics of the edges in minimum-cost configurations. The cost function is minimized by the simulated annealing method. A novel set of strategies for generating candidate states and a suitable temperature schedule are presented. Sequential and parallel versions of the annealing algorithm are implemented and compared. Experimental results are presented.> Hin Leong Tan, Saul B. Gelfand, Edward J. Delp |
CVPR | 3 |
| 1989 | Partial shape recognition: a landmark-based approachabstractA technique to guide landmark matching known as hopping dynamic programming is described. The location of the model in the scene is estimated with a least-squares fit. A heuristic measure is then computed to decide if the model is in the scene. The shape features of an object are the landmarks associated with the object. The landmarks of an object are defined as the points of interest of the object that have important shape attributes. Examples of landmarks are corners, holes, protrusions, and high-curvature points.> Nirwan Ansari, Edward J. Delp |
SMC | 2 |
| 1989 | An iterative growing and pruning algorithm for classification tree designabstractAn efficient iterative method is proposed to grow and prune classification trees. This method divides the data sample into two subsets and iteratively grows a tree with one subset and prunes it with the other subset, successively interchanging the roles of the two subsets. The convergence and other properties of the algorithm are established. Theoretical and practical considerations suggest that the iterative tree growing and pruning algorithm should perform better and require less computation than other widely used tree growing and pruning algorithms. Numerical results on a waveform recognition problem are presented to support this view.> Saul B. Gelfand, C. S. Ravishankar, Edward J. Delp |
SMC | 3 |
| 1989 | Morphological based target enhancement algorithms to counter the hostile nuclear environmentabstractIt is noted that a necessary requirement of a strategic defense system is the detection of incoming nuclear warheads in an environment that may include nuclear detonations of undetected or missed target warheads. A computer model is described which simulates incoming warheads as distant endoatmospheric targets. A model of the expected electromagnetic noise present in the nuclear environment is developed; predicted atmospheric effects are included. Various morphological-based image-enhancement algorithms are examined with regard to their ability to suppress the noise and atmospheric effects of the nuclear environment. These algorithms are then tested, using the combined target and noise models, and evaluated in terms of noise removal and the ability to resolve closely spaced targets.> Cameron H. G. Wright, Edward J. Delp, Neal C. Gallagher |
SMC | 2 |
| 1989 | A Model for an Intelligent Operating System for Executing Image Understanding Tasks on a Reconfigurable Parallel Architecture
Chee-Hung Henry Chu, Edward J. Delp, Leah H. Jamieson, Howard Jay Siegel, Frank J. Weil, Andrew B. Whinston |
J. Parallel Distributed Comput. | 2 |
| 1989 | A comparative cost function approach to edge detectionabstractEdge detection is cast as a problem in cost minimization. The concept of an edge that is based on criteria such as accurate localization, thinness, continuity, and length is described. On the basis of this description, a comparative cost function that mathematically captures the intuitive idea of an edge is formulated. The function uses information from both image data and local edge structure in evaluating the relative quality of pairs of edge configurations. The function is a linear combination of weighted cost factors. Computation of the function is performed efficiently by organizing information in the form of a decision tree. Edges are detected using a heuristic iterative search algorithm based on the comparative cost function. The detection process can be implemented largely in parallel. The usefulness of this approach to edge detection is demonstrated by showing experimental results of detected edges for both real and synthetic images.> Hin Leong Tan, Saul B. Gelfand, Edward J. Delp |
IEEE Trans. Syst. Man Cybern. | 3 |
| 1987 | Inspection of machine parts by backprojection reconstructionabstractExtraction of 3D information from 2D views is an important problem in the inspection of manufactured parts. Inspection for defects and correctness of shape and size requires accurate 3D information. In this paper we present a method for 3D reconstruction by backprojection of 2D views. This method uses readily available passive imaging sensors and a computer-controlled positioner to directly produce a 3D reconstruction of the part. Consistency based processing is introduced as a method of increasing reconstruction accuracy by exploiting redundant data in the multiple images. Maximum effciency and accuracy in reconstruction is achieved using optimal view patterns and reliable silhouette extraction techniques. Although this method is applicable to the inspection of general machine parts, our immediate application was in the inspection of laser-drilled rivet holes and laser-drilled cooling holes in jet engine turbine blades. Hin Leong Tan, Eric Viscito, Edward J. Delp, Jan P. Allebach |
ICRA | 3 |
| 1986 | Adaptive gray scale mapping to reduce registration noise in difference images
Thomas F. Knoll, Edward J. Delp |
Comput. Vis. Graph. Image Process. | 2 |
| 1985 | Automatic visual inspection of solder jointsabstractThis paper describes an approach for automatic inspection of solder joints on printed circuit boards using gray-scale images. Common defects in solder joints are recognized using features computed from segmented solder joint subimages. Unacceptable joints are assigned to one of several defective classes. Defect classification, rather than just detection of defective joints, is motivated by the desire to automatically take corrective action on the assembly line. The features used for classification are based on characteristics of intensity surfaces. It is shown that features derived from surface facets are effective in the classification of solder joints using a minimum-distance classification algorithm. Paul J. Besl, Edward J. Delp, Ramesh Jain 0001 |
ICRA | 2 |
| 1985 | Automatic visual solder joint inspectionabstractAn approach is described for the automatic inspection of solder joints on printed circuit boards. Common defects are identified in solder joints and a joint is classified as being good or belonging to one of the defective classes. The motivation for this classification is not just the detection of defective joints, but the desire to automatically take corrective action on the assembly line. The features used for classification are based on characteristics of intensity surfaces. It is shown that features derived from facets and Gaussian curvature are effective in the classification of solder joints using a minimum-distance classification algorithm. Class separation plots are shown to be useful for quickly studying individual effectiveness of a feature or pair of features in classification. Results show the efficacy of the described approach. Paul J. Besl, Edward J. Delp, Ramesh Jain 0001 |
IEEE J. Robotics Autom. | 2 |
| 1985 | Detecting edge segmentsabstractThe performances of three classes of edge operators are evaluated. The three are conventional 3×3 operators with a connectivity test, sequential edge detectors using both edge strength and direction information, and edge detectors using large masks. Implementations of the sequential edge operators using both the edge strength and direction information to detect edge segments with smoothly varying edge direction and physical length are discussed. Results of applying these edge detectors to real-world images are also presented. Edward J. Delp, Chee-Hung Henry Chu |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1979 | Image data compression using autoregressive time series models
Edward J. Delp, Rangasami L. Kashyap, O. Robert Mitcheli |
Pattern Recognit. | 1 |
| 1979 | Image Compression Using Block Truncation CodingabstractA new technique for image compression called Block Truncation Coding (BTC) is presented and compared with transform and other techniques. The BTC algorithm uses a two-level (one-bit) nonparametric quantizer that adapts to local properties of the image. The quantizer that shows great promise is one which preserves the local sample moments. This quantizer produces good quality images that appear to be enhanced at data rates of 1.5 bits/picture element. No large data storage is required, and the computation is small. The quantizer is compared with standard (minimum mean-square error and mean absolute error) one-bit quantizers. Modifications of the basic BTC algorithm are discussed along with the performance of BTC in the presence of channel errors. Edward J. Delp, Owen Robert Mitchell |
IEEE Trans. Commun. | 1 |