VLDB 2026 Research / reviewers in the wild / expert
Thrasyvoulos N. Pappas
dblp:33/6759 · also Thrasos N. Pappas
· DBLP profile ↗
87ranked-venue papers
10as first author
6since 2021 · last 2026
0000-0002-4598-2197ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 80 · 9 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorComputer networks · 3Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Digital Film Grain SimulationabstractWe propose an efficient digital simulation of film grain noise that is based on the random-dot model, applies to a wide range of imaging conditions (magnification factors, optical blur), and accounts for the signal-dependence and spatial correlation of film grain. Our analysis and experimental results demonstrate that the proposed approach produces sharper images with similar film grain compared to Monte Carlo simulations proposed by Newson et al., offering superior rendering efficiency across a wide range of parameter settings including a case where Monte Carlo simulations fail (low magnification with low blur). While recently proposed statistical approximations of film grain by Zhang et al. are significantly faster, they cannot be used for high magnification or low blur. Daizong Tian, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 2 |
| 2024 | Training and Testing Texture Similarity Metrics for Structurally Lossless CompressionabstractWe present a systematic approach for training and testing structural texture similarity metrics (STSIMs) so that they can be used to exploit texture redundancy for structurally lossless image compression. The training and testing is based on a set of image distortions that reflect the characteristics of the perturbations present in natural texture images. We conduct empirical studies to determine the perceived similarity scale across all pairs of original and distorted textures. We then introduce a data-driven approach for training the Mahalanobis formulation of STSIM based on the resulting annotated texture pairs. Experimental results demonstrate that training results in significant improvements in metric performance. We also show that the performance of the trained STSIM metrics is competitive with state of the art metrics based on convolutional neural networks, at substantially lower computational cost. Zhaochen Shi, Jana Zujovic, Huib de Ridder, René van Egmond, David L. Neuhoff, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 7 |
| 2023 | Film Grain Rendering and Parameter EstimationabstractWe propose a realistic film grain rendering algorithm based on statistics derived analytically from a physics-based Boolean model that Newson et al. adopted for Monte Carlo simulations of film grain. We also propose formulas for estimation of the model parameters from scanned film grain images. The proposed rendering is computationally efficient and can be used for real-time film grain simulation for a wide range of film grain parameters when the individual film grains are not visible. Experimental results demonstrate the effectiveness of the proposed approach for both constant and real-world images, for a six orders of magnitude speed-up compared with the Monte Carlo simulations of the Newson et al. approach. Daizong Tian, Thrasyvoulos N. Pappas |
ACM Trans. Graph. | 4 |
| 2022 | Pattern-Based Reconstruction of K-Level Images From CutsetsabstractWe present a pattern-based approach for reconstructing a K-level image from cutsets, dense samples taken along a family of lines or curves in two- or three-dimensional space, which break the image into blocks, each of which is typically reconstructed independently of the others. The pattern-based approach utilizes statistics of human segmentations to generate a codebook of patterns, each of which represents a pair of a block boundary specification and the corresponding pattern in the block interior. We develop the approach for rectangular cutset topologies and show that it can be extended to general periodic sampling topologies. We also show that, for bilevel cutset reconstruction, the pattern-based can be combined with the previously proposed cutset-MRF approach to substantially reduce the size of the codebook with a slight increase in reconstruction error. In addition, we present an algorithm for segmenting the cutset samples of an original grayscale or color image, followed by reconstruction of the full segmentation field via the pattern-based approach. Experimental results show that the proposed approaches outperform the cutset-MRF approaches in terms of both reconstruction error rate and perceptual quality. Moreover, this is accomplished without any side information about the structure of the block interior. Systematic comparisons of the performance of different sampling topologies are also provided. Shengxin Zha, Daizong Tian, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 3 |
| 2021 | Human Skin Gloss Perception Based on Texture StatisticsabstractWe propose objective, image-based techniques for quantitative evaluation of facial skin gloss that is consistent with human judgments. We use polarization photography to obtain separate images of surface and subsurface reflections, and rely on psychophysical studies to uncover and separate the influence of the two components on skin gloss perception. We capture images of facial skin at two levels, macro-scale (whole face) and meso-scale (skin patch), before and after cleansing. To generate a broad range of skin appearances for each subject, we apply photometric image transformations to the surface and subsurface reflection images. We then use linear regression to link statistics of the surface and subsurface reflections to the perceived gloss obtained in our empirical studies. The focus of this paper is on within-subject gloss perception, that is, on visual differences among images of the same subject. Our analysis shows that the contrast of the surface reflection has a strong positive influence on skin gloss perception, while the darkness of the subsurface reflection (skin tone) has a weaker positive effect on perceived gloss. We show that a regression model based on the concatenation of statistics from the two reflection images can successfully predict relative gloss differences. Jing Wang 0069, Carla Kuesten, Jim Mayne, Gopa Majmudar, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 5 |
| 2021 | Hierarchical Lossy Bilevel Image Compression Based on Cutset SamplingabstractWe consider lossy compression of a broad class of bilevel images that satisfy the smoothness criterion, namely, images in which the black and white regions are separated by smooth or piecewise smooth boundaries, and especially lossy compression of complex bilevel images in this class. We propose a new hierarchical compression approach that extends the previously proposed fixed-grid lossy cutset coding (LCC) technique by adapting the grid size to local image detail. LCC was claimed to have the best rate-distortion performance of any lossy compression technique in the given image class, but cannot take advantage of detail variations across an image. The key advantages of the hierarchical LCC (HLCC) is that, by adapting to local detail, it provides constant quality controlled by a single parameter (distortion threshold), independent of image content, and better overall visual quality and rate-distortion performance, over a wider range of bitrates. We also introduce several other enhancements of LCC that improve reconstruction accuracy and perceptual quality. These include the use of multiple connection bits that provide structural information by specifying which black (or white) runs on the boundary of a block must be connected, a boundary presmoothing step, stricter connectivity constraints, and more elaborate probability estimation for arithmetic coding. We also propose a progressive variation that refines the image reconstruction as more bits are transmitted, with very small additional overhead. Experimental results with a wide variety of, and especially complex, bilevel images in the given class confirm that the proposed techniques provide substantially better visual quality and rate-distortion performance than existing lossy bilevel compression techniques, at bitrates lower than lossless compression with the JBIG or JBIG2 standards. Shengxin Zha, Thrasyvoulos N. Pappas, David L. Neuhoff |
IEEE Trans. Image Process. | 2 |
| 2019 | Structural Texture Similarity for Material RecognitionabstractWe propose a new direct approach for material recognition under diverse illumination and viewing conditions based on visual texture. We apply K-means clustering to feature vectors that consist of steerable filter subband statistics and dominant colors of each texture image in order to obtain a small number of exemplars characterizing each material. We then use structural texture similarity metrics and color composition metrics to compare a query texture to the exemplars for material classification. Experimental results using the CUReT database establish the importance of color and demonstrate that five exemplars per texture provide performance comparable to the state of the art. Jue Lin, Thrasyvoulos N. Pappas |
ICIP | 2 |
| 2016 | Generalized k-level cutset sampling and reconstructionabstractWe propose a family of cutset sampling schemes and a generalized k-level image reconstruction approach formulated under a minimum mean squared error (MMSE) framework. The k-level reconstruction approach is a direct generalization of the recently proposed pattern-based approach, and can be applied to periodic samples either on a cutset or on a grid. Our experimental results indicate that the generalization of the k-level reconstruction approach results in only a small performance loss. For rectangular cutsets, we show that the proposed approach outperforms the cutset-MRF approach as well as two inpainting approaches. Moreover, we show that combining the cutset sampling with an additional point sample inside the periodic structure outperforms k-level reconstruction from cutset sampling and point sampling under comparable sampling densities. Shengxin Zha, Thrasyvoulos N. Pappas |
ICASSP | 2 |
| 2016 | A hybrid Markov random field model for bilevel cutset reconstructionabstractWe propose a hybrid Markov random field (MRF) model and a two stage cutset-MRF approach for reconstructing bilevel images from cutsets. We show that the proposed approach leads to substantial improvements over previous cutset-MRF approaches in terms of both reconstruction error and visual quality (continuity of reconstructed segments and preservation of image structure). The proposed approach approaches the performance of pattern-based approaches without the additional memory requirements and training overhead. We also show that it outperforms inpainting approaches adapted to bilevel cutset reconstruction. Shengxin Zha, Thrasyvoulos N. Pappas |
ICIP | 2 |
| 2016 | Perceiving Graphical and Pictorial Information via Hearing and TouchabstractWe propose a dynamic interactive system for conveying visual information via hearing and touch to blind and visually impaired people. The system is implemented with a touch screen that allows the user to actively explore a two-dimensional layout consisting of one or more objects with the finger while listening to auditory feedback. Sound is used as the primary source of information for object localization, identification, and shape, while touch is used for pointing and kinesthetic feedback. A static overlay of raised-dot tactile patterns can also be added. The head-related transfer function is used for rendering sound directionality, and variations of sound intensity or other features are used for rendering proximity. The main focus is on conveying the shape of an object, but the rendering of a simple scene layout, that consists of objects in a linear arrangement, each with a distinct tapping sound, is also considered and compared to a “virtual cane.” We consider a number of acoustic-tactile configurations and use empirical studies with visually blocked sighted participants to compare their effectiveness. Our findings demonstrate the advantages of spatial sound (directionality and proximity cues) for dynamic display of information (localization, identification, shape), while raised-dot patterns provide the best static shape rendition. We also show that the proposed configurations outperform existing techniques. The proposed approach is also expected to impact other applications where vision cannot be used. Pubudu Madhawa Silva, Thrasyvoulos N. Pappas, Joshua Atkins, James E. West |
IEEE Trans. Multim. | 2 |
| 2015 | Pattern-based k-level cutset reconstructionabstractWe propose a pattern-based approach for reconstructing k-level images from cutsets. We construct a database of k-level cutset block patterns and use it for reconstruction based on fuzzy pattern retrieval and Markov random field (MRF) energy minimization. The proposed approach outperforms previously proposed cutset-MRF approaches as well as inpainting approaches. When integrated into a lossy bilevel image coding scheme that utilizes cutsets, the proposed approach outperforms the state of the art with comparable constraints. Shengxin Zha, Thrasyvoulos N. Pappas |
ICIP | 2 |
| 2014 | An adaptive lighting correction method for matched-texture codingabstractMatched-Texture Coding is a novel image coder that utilizes the self-similarity of natural images that include textures, in order to achieve structurally lossless compression. The key to a high compression ratio is replacing large image blocks with previously encoded blocks with similar structure. Adjusting the lighting of the replaced block is critical for eliminating illumination artifacts and increasing the number of matches. We propose a new adaptive lighting correction method that is based on the Poisson equation with incomplete boundary conditions. In order to fully exploit the benefits of the adaptive Poisson lighting correction, we also propose modifications of the side-matching (SM) algorithm and structural texture similarity metric. We show that the resulting matched-texture algorithm achieves better coding performance. Guoxin Jin, Thrasyvoulos N. Pappas, David L. Neuhoff |
ICASSP | 2 |
| 2014 | Subjective similarity evaluation for scenic bilevel imagesabstractIn order to provide ground truth for subjectively comparing compression methods for scenic bilevel images, as well as for judging objective similarity metrics, this paper describes the subjective similarity rating of a collection of distorted scenic bilevel images. Unlike text, line drawings, and silhouettes, scenic bilevel images contain natural scenes, e.g., landscapes and portraits. Seven scenic images were each distorted in forty-four ways, including random bit flipping, dilation, erosion and lossy compression. To produce subjective similarity ratings, the distorted images were each viewed by 77 subjects. These are then used to compare the performance of four compression algorithms and to assess how well percentage error and SmSIM work as bilevel image similarity metrics. These subjective ratings can also provide ground truth for future tests of objective bilevel image similarity metrics. Yuanhao Zhai 0002, David L. Neuhoff, Thrasyvoulos N. Pappas |
ICASSP | 3 |
| 2014 | Structural texture similarity metric based on intra-class variancesabstractTraditional point-by-point image similarity metrics, such as the ℓ2-norm, are not always consistent with human perception, especially in textured regions. We consider the problem of identifying textures that are perceptually identical to a query texture; this is important for image retrieval, compression, and restoration applications. Recently proposed structural texture similarity (STSIM) metrics assign high similarity scores to such perceptually identical textured patches, even though they may have significant pixel-wise deviations. We use an STSIM approach that compares a set of statistical patch descriptors through a weighted distance, and, given a dataset of labeled texture images partitioned into classes of perceptually identical patches, we calculate the weights as the variances of each statistic centered around the mean of its class. Experimental results demonstrate that the proposed approach outperforms existing structural similarity metrics and STSIMs as well as traditional point-by-point metrics when assessing texture similarity in both noisy and noise-free conditions. Matteo Maggioni, Guoxin Jin, Alessandro Foi, Thrasyvoulos N. Pappas |
ICIP | 4 |
| 2014 | Lossy Cutset Coding of Bilevel Images Based on Markov Random FieldsabstractAn effective, low complexity method for lossy compression of scenic bilevel images, called lossy cutset coding, is proposed based on a Markov random field model. It operates by losslessly encoding pixels in a square grid of lines, which is a cutset with respect to a Markov random field model, and preserves key structural information, such as borders between black and white regions. Relying on the Markov random field model, the decoder takes a MAP approach to reconstructing the interior of each grid block from the pixels on its boundary, thereby creating a piecewise smooth image that is consistent with the encoded grid pixels. The MAP rule, which reduces to finding the block interiors with fewest black-white transitions, is directly implementable for the most commonly occurring block boundaries, thereby avoiding the need for brute force or iterative solutions. Experimental results demonstrate that the new method is computationally simple, outperforms the current lossy compression technique most suited to scenic bilevel images, and provides substantially lower rates than lossless techniques, e.g., JBIG, with little loss in perceived image quality. Matthew G. Reyes, David L. Neuhoff, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 3 |
| 2013 | Local radius index - a new texture similarity featureabstractWe develop a new type of statistical texture image feature, called a Local Radius Index (LRI), which can be used to quantify texture similarity based on human perception. Image similarity metrics based on LRI can be applied to image compression, identical texture retrieval and other related applications. LRI extracts texture features by using simple pixel value comparisons in space domain. Better performance can be achieved when LRI is combined with complementary texture features, e.g., Local Binary Patterns (LBP) and the proposed Subband Contrast Distribution. Compared with Structural Texture Similarity Metrics (STSIM), the LRI-based metrics achieve better retrieval performance with much less computation. Applied to the recently developed structurally lossless image coder, Matched Texture Coding, LRI enables similar performance while significantly accelerating the encoding. Yuanhao Zhai 0002, David L. Neuhoff, Thrasyvoulos N. Pappas |
ICASSP | 3 |
| 2013 | Image Analysis: Focus on Texture SimilarityabstractTexture is an important visual attribute both for human perception and image analysis systems. We review recently proposed texture similarity metrics and applications that critically depend on such metrics, with emphasis on image and video compression and content-based retrieval. Our focus is on natural textures and structural texture similarity metrics (STSIMs). We examine the relation of STSIMs to existing models of texture perception, texture analysis/synthesis, and texture segmentation. We emphasize the importance of signal characteristics and models of human perception, both for algorithm development and testing/validation. Thrasyvoulos N. Pappas, David L. Neuhoff, Huib de Ridder, Jana Zujovic |
Proc. IEEE | 1 |
| 2013 | Structural Texture Similarity Metrics for Image Analysis and RetrievalabstractWe develop new metrics for texture similarity that accounts for human visual perception and the stochastic nature of textures. The metrics rely entirely on local image statistics and allow substantial point-by-point deviations between textures that according to human judgment are essentially identical. The proposed metrics extend the ideas of structural similarity and are guided by research in texture analysis-synthesis. They are implemented using a steerable filter decomposition and incorporate a concise set of subband statistics, computed globally or in sliding windows. We conduct systematic tests to investigate metric performance in the context of "known-item search," the retrieval of textures that are "identical" to the query texture. This eliminates the need for cumbersome subjective tests, thus enabling comparisons with human performance on a large database. Our experimental results indicate that the proposed metrics outperform peak signal-to-noise ratio (PSNR), structural similarity metric (SSIM) and its variations, as well as state-of-the-art texture classification metrics, using standard statistical measures. Jana Zujovic, Thrasyvoulos N. Pappas, David L. Neuhoff |
IEEE Trans. Image Process. | 2 |
| 2012 | Subjective and objective texture similarity for image compressionabstractWe focus on the evaluation of texture similarity metrics for structurally lossless or nearly structurally lossless image compression. By structurally lossless we mean that the original and compressed images, while they may have visible differences in a side-by-side comparison, they have similar quality so that one cannot tell which is the original. This is particularly important for textured regions, which can have significant point-by-point differences, even though to the human eye they appear to be the same. As in traditional metrics, texture similarity metrics are expected to provide a monotonic relationship between measured and perceived distortion. To evaluate metric performance according to this criterion, we introduce a systematic approach for generating synthetic texture distortions that model variations that occur in natural textures. Based on such distortions, we conducted subjective experiments with a variety of original texture images and different types and degrees of distortions. Our results indicate that recently proposed structural texture similarity metrics provide the best performance. Jana Zujovic, Thrasyvoulos N. Pappas, David L. Neuhoff, René van Egmond, Huib de Ridder |
ICASSP | 2 |
| 2012 | Matched-texture coding for structurally lossless compressionabstractWe propose a new texture-based compression approach that relies on new texture similarity metrics and is able to exploit texture redundancies for significant compression gains without loss of visual quality, even though there may visible differences with the original image (structurally lossless). Existing techniques rely on point-by-point metrics that cannot account for the stochastic and repetitive nature of textures. The main idea is to encode selected blocks of textures - as well as smooth blocks and blocks containing boundaries between smooth and/or textured regions - by pointing to previously occurring (already encoded) blocks of similar textures, blocks that are not encoded in this way, are encoded by a baseline method, such as JPEG. Experimental results with natural images demonstrate the advantages of the proposed approach. Guoxin Jin, Yuanhao Zhai 0002, Thrasyvoulos N. Pappas, David L. Neuhoff |
ICIP | 3 |
| 2012 | Image reconstruction from a Manhattan grid via piecewise plane fitting and Gaussian Markov random fieldsabstractThis paper builds upon previous work for image reconstruction problems in which samples are taken on evenly spaced rows and columns, i.e., a Manhattan grid. A new reconstruction method is proposed that uses three steps to interpolate the interior of each block under the model that an image can be decomposed into piecewise planar regions plus noise. First, the K-planes algorithm is developed in order to fit several planes to the observed pixel values on the border. Second, one of theK planes is assigned to each pixel of the block interior, by a process of partitioning the block with polygons, thereby creating a piecewise planar approximation. Third, the interior pixels are interpolated by modeling them as a Gauss Markov random field whose mean is the piecewise planar approximation just obtained. The new method is shown to improve significantly upon previous methods, especially in the preservation of “soft” image edges. Matthew A. Prelee, David L. Neuhoff, Thrasyvoulos N. Pappas |
ICIP | 3 |
| 2012 | Hierarchical bilevel image compression based on cutset samplingabstractWe propose a hierarchical lossy bilevel image compression method that relies on adaptive cutset sampling (along lines of a rectangular grid with variable block size) and Markov Random Field based reconstruction. It is an efficient encoding scheme that preserves image structure by using a coarser grid in smooth areas of the image and a finer grid in areas with more detail. Experimental results demonstrate that the proposed method performs as well as or better than the fixed-grid approach, and outperforms other lossy bilevel compression methods in its rate-distortion performance. Shengxin Zha, Thrasyvoulos N. Pappas, David L. Neuhoff |
ICIP | 2 |
| 2011 | Perceiving graphical and pictorial information via touch and hearingabstractWith the ever increasing availability of the Internet and electronic media rich in graphical and pictorial information - for communication, commerce, entertainment, art, education - it has been hard for the visually impaired community to keep up. We propose a non-invasive system that can be used to convey graphical and pictorial information via touch and hearing. The main idea is that the user actively explores a two-dimensional layout consisting of one or more objects on a touch screen with the finger while listening to auditory feedback. We have demonstrated the efficacy of the proposed approach in a range of tasks, from basic shape identification to perceiving a scene with several objects. The proposed approach is also expected to contribute to research in virtual reality, immersive environments, and medicine. Pubudu Madhawa Silva, Thrasyvoulos N. Pappas, Joshua Atkins, James E. West |
ICASSP | 2 |
| 2011 | Cutset sampling and reconstruction of imagesabstractThis paper presents a new approach to sampling images in which samples are taken on a cutset with respect to a graphical image model. The cutsets considered are Manhattan grids, for example every Nth row and column of the image. Cutset sampling is motivated mainly by applications with physical constraints, e.g. a ship taking water samples along its path, but also by the fact that dense sampling along lines might permit better reconstruction of edges than conventional sampling at the same density. The main challenge in cutset sampling lies in the reconstruction of the unsampled blocks. As a first investigation, this paper uses segmentation followed by linear estimation. First, the ACA method [1] is modified to segment the cutset, followed by a binary Markov random field (MRF) inspired segmentation of the unsampled blocks. Finally, block interiors are estimated from the pixels on their boundaries, as well as their segmentation, with methods that include a generalization of bilinear interpolation and linear MMSE methods based on Gaussian MRF models or separable autocorrelation models. The resulting reconstructions are comparable to those obtained with conventional sampling at higher sampling densities, but not generally as good as conventional sampling at lower rates. Ashish Farmer, Awlok Josan, Matthew A. Prelee, David L. Neuhoff, Thrasyvoulos N. Pappas |
ICIP | 5 |
| 2011 | Hybrid light coding for fast and high-accuracy shape acquisitionabstractIn this paper we propose a novel rate control initialization algorithm for real-time H.264/scalable video coding. In particular, a two-step approach is proposed. First, the initial quantization parameter (QP) for each layer is determined by means of a parametric rate-quantization (R-Q) modeling that depends on the layer identifier (base or enhancement) and on the type of scalability (spatial or quality). Second, an intra-frame QP refinement method that allows for adapting the initial QP value when needed is carried out over the three first coded frames in order to take into consideration both the buffer control and the spatio-temporal complexity of the scene. The experimental results show that the proposed R-Q modeling for initial QP estimation, in combination with the intra-frame QP refinement method, provide a good performance in terms of visual quality and buffer control, achieving remarkably similar results to those achieved by using ideal initial QP values. Lulu He, Paul Kane, Thrasyvoulos N. Pappas |
ICIP | 4 |
| 2011 | 3D surface registration using Z-SIFTabstractWe present a Z-SIFT based 3D surface registration algorithm that utilizes the depth information enhanced SIFT features to make initial alignment and the 2D feature weighted Iterative Closest Point (ICP) algorithm to realize accurate registration. The combination of SIFT features and depth information extracts faithful corresponding points between the 2D images and provides good coarse alignment for the 3D surfaces. The 2D feature weighted ICP also outperforms the naive ICP algorithm in terms of speed and accuracy. We use this approach in the context of multiple view alignment for 3D scanners. Experimental results with real objects and human faces indicate the effectiveness of the proposed approach. Lulu He, Thrasyvoulos N. Pappas |
ICIP | 3 |
| 2011 | Design of benchmark imagery for validating facility annotation algorithmsabstractThe design of benchmark imagery for validation of image an notation algorithms is considered. Emphasis is placed on imagery that contains industrial facilities, such as chemical re fineries. An application-level facility ontology is used as a means to define salient objects in the benchmark imagery. Instrinsic and extrinsic scene factors important for comprehensive validation are listed, and variability in the benchmarks discussed. Finally, the pros and cons of three forms of bench mark imagery: real, composite and synthetic, are delineated. Randy S. Roberts, Paul A. Pope, Ranga Raju Vatsavai, Ming Jiang 0005, Lloyd F. Arrowood, Timothy G. Trucano, Shaun S. Gleason, Anil M. Cheriyadat, Alexandre Sorokine, Aggelos K. Katsaggelos, Thrasyvoulos N. Pappas, Lucinda R. Gaines, Lawrence K. Chilton |
IGARSS | 11 |
| 2011 | New Challenges for Image Processing Research
Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 1 |
| 2010 | An adaptive clustering and chrominance-based merging approach for image segmentation and abstractionabstractWe present a novel, computationally efficient approach for natural image segmentation that uses the adaptive clustering algorithm (ACA) to obtain an initial segmentation and chrominance-based region merging to consolidate regions of perceptually uniform texture. The combination of ACA and chrominance-based merging preserves salient edges and smooths out noise and edges within textured regions. It can thus be used for image abstraction. Experimental results with natural images indicate the effectiveness of the proposed approach. Lulu He, Thrasyvoulos N. Pappas |
ICIP | 2 |
| 2010 | Geospatial image mining for nuclear proliferation detection: Challenges and new opportunitiesabstractWith increasing understanding and availability of nuclear technologies, and increasing persuasion of nuclear technologies by several new countries, it is increasingly becoming important to monitor the nuclear proliferation activities. There is a great need for developing technologies to automatically or semi-automatically detect nuclear proliferation activities using remote sensing. Images acquired from earth observation satellites is an important source of information in detecting proliferation activities. High-resolution remote sensing images are highly useful in verifying the correctness, as well as completeness of any nuclear program. DOE national laboratories are interested in detecting nuclear proliferation by developing advanced geospatial image mining algorithms. In this paper we describe the current understanding of geospatial image mining techniques and enumerate key gaps and identify future research needs in the context of nuclear proliferation. Ranga Raju Vatsavai, Budhendra L. Bhaduri, Anil M. Cheriyadat, Lloyd F. Arrowood, Eddie A. Bright, Shaun S. Gleason, Carl Diegert, Aggelos K. Katsaggelos, Thrasyvoulos N. Pappas, Reid B. Porter, Jim Bollinger, Barry Chen, Ryan Hohimer |
IGARSS | 9 |
| 2009 | Structural similarity metrics for texture analysis and retrievalabstractThe development of objective texture similarity metrics for image analysis applications differs from that of traditional image quality metrics because substantial point-by-point deviations are possible for textures that according to human judgment are essentially identical. Thus, structural similarity metrics (SSIM) attempt to incorporate ¿structural¿ information in image comparisons. The recently proposed structural texture similarity metric (STSIM) relies entirely on local image statistics. We extend this idea further by including a broader set of local image statistics, basing the selection on metric performance as compared to subjective evaluations. We utilize both intra- and inter-subband correlations, and also incorporate information about the color composition of the textures into the similarity metrics. The performance of the proposed metrics is compared to PSNR, SSIM, and STSIM on the basis of subjective evaluations using a carefully selected set of 50 texture pairs. Jana Zujovic, Thrasyvoulos N. Pappas, David L. Neuhoff |
ICIP | 2 |
| 2009 | Classifying paintings by artistic genre: An analysis of features & classifiersabstractThis paper describes an approach to automatically classify digital pictures of paintings by artistic genre. While the task of artistic classification is often entrusted to human experts, recent advances in machine learning and multimedia feature extraction has made this task easier to automate. Automatic classification is useful for organizing large digital collections, for automatic artistic recommendation, and even for mobile capture and identification by consumers. Our evaluation uses variable resolution painting data gathered across Internet sources rather than solely using professional high-resolution data. Consequently, we believe this solution better addresses the task of classifying consumer-quality digital captures than other existing approaches. We include a comparison to existing feature extraction and classification methods as well as an analysis of our own approach across classifiers and feature vectors. Jana Zujovic, Lisa Gandy, Bryan Pardo, Thrasyvoulos N. Pappas |
MMSP | 5 |
| 2009 | Perceptual similarity metrics for retrieval of natural texturesabstractWe investigate perceptual similarity metrics for the content-based retrieval of natural textures. The goal is to find perceptually similar textures that may have significant differences on a point-by-point basis. The evaluation of such metrics typically requires extensive and cumbersome subjective tests. The focus of this paper is on the recovery of textures that are "identical" to the query texture, in the sense that they are pieces of the same texture. This is important in content-based image retrieval (CBIR), where one may want to find images that contain a particular texture, as well as in some near-threshold coding applications. The advantage of evaluating metric performance in the context of retrieving identical textures is that the ground truth is known, and therefore no subjective tests are required. We can thus compare the performance of different metrics on large sets of textures, and derive meaningful statistical results.We evaluate the performance of a recently proposed structural texture similarity metric on grayscale textures, and compare it to that of PSNR, as well as space domain and complex wavelet structural similarity metrics. Experimental results with a database of 748 distinct texture images, indicate that the new metric outperforms the other metrics in the retrieval of identical textures, according to a number of standard statistical measures. Jana Zujovic, Thrasyvoulos N. Pappas, David L. Neuhoff |
MMSP | 2 |
| 2009 | Subjective Evaluation of Spatial Resolution and Quantization Noise TradeoffsabstractMost full-reference fidelity/quality metrics compare the original image to a distorted image at the same resolution assuming a fixed viewing condition. However, in many applications, such as video streaming, due to the diversity of channel capacities and display devices, the viewing distance and the spatiotemporal resolution of the displayed signal may be adapted in order to optimize the perceived signal quality. For example, at low bitrate coding applications an observer may prefer to reduce the resolution or increase the viewing distance to reduce the visibility of the compression artifacts. The tradeoff between resolution/viewing conditions and visibility of compression artifacts requires new approaches for the evaluation of image quality that account for both image distortions and image size. In order to better understand such tradeoffs, we conducted subjective tests using two representative still image coders, JPEG and JPEG 2000. Our results indicate that an observer would indeed prefer a lower spatial resolution (at a fixed viewing distance) in order to reduce the visibility of the compression artifacts, but not all the way to the point where the artifacts are completely invisible. Moreover, the observer is willing to accept more artifacts as the image size decreases. The subjective test results we report can be used to select viewing conditions for coding applications. They also set the stage for the development of novel fidelity metrics. The focus of this paper is on still images, but it is expected that similar tradeoffs apply to video. Soo Hyun Bae, Thrasyvoulos N. Pappas, Biing-Hwang Juang |
IEEE Trans. Image Process. | 2 |
| 2008 | Image spam hunterabstractSpammers are constantly creating sophisticated new weapons in their arms race with anti-spam technology, the latest of which is image-based spam. The newest image-based spam uses simple image processing technologies to vary the content of individual messages, e.g. by changing foreground colors, backgrounds, font types, or even rotating and adding artifacts to the images. Thus, they pose great challenges to conventional spam filters. In this paper, we propose a system using a probabilistic boosting tree to determine whether an incoming image is a spam or not based on global image features, i.e. color and gradient orientation histograms. The system identifies spam without the need for OCR and is robust in the face of the kinds of variation found in current spam images. Evaluation results show the system correctly classifies 90% of spam images while mislabeling only 0.86% of non-spam images as spam. Yan Gao 0003, Ming Yang 0007, Xiaonan Zhao, Bryan Pardo, Ying Wu 0001, Thrasyvoulos N. Pappas, Alok N. Choudhary |
ICASSP | 6 |
| 2008 | Structural texture similarity metrics for retrieval applicationsabstractTraditional image similarity metrics compare two images on a point-by-point basis. On the other hand, structural similarity metrics (SSIM) attempt to base image similarity on "structural" information. We evaluate the performance of SSIM metrics in the context of texture similarity, and propose new metrics that incorporate the best features of SSIM and eliminate the most serious drawbacks. We show that the proposed new texture similarity metrics outperform SSIM and its variations, as well as PSNR and other traditional metrics. We demonstrate the advantages of the new metrics on a carefully selected set of 39 texture pairs and comparisons with informal subjective test results. Xiaonan Zhao, Matthew G. Reyes, Thrasyvoulos N. Pappas, David L. Neuhoff |
ICIP | 3 |
| 2008 | Structural Similarity Quality Metrics in a Coding Context: Exploring the Space of Realistic DistortionsabstractPerceptual image quality metrics have explicitly accounted for human visual system (HVS) sensitivity to subband noise by estimating just noticeable distortion (JND) thresholds. A recently proposed class of quality metrics, known as structural similarity metrics (SSIM), models perception implicitly by taking into account the fact that the HVS is adapted for extracting structural information from images. We evaluate SSIM metrics and compare their performance to traditional approaches in the context of realistic distortions that arise from compression and error concealment in video compression/transmission applications. In order to better explore this space of distortions, we propose models for simulating typical distortions encountered in such applications. We compare specific SSIM implementations both in the image space and the wavelet domain; these include the complex wavelet SSIM (CWSSIM), a translation-insensitive SSIM implementation. We also propose a perceptually weighted multiscale variant of CWSSIM, which introduces a viewing distance dependence and provides a natural way to unify the structural similarity approach with the traditional JND-based perceptual approaches. Alan C. Brooks, Xiaonan Zhao, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 3 |
| 2008 | Resource Allocation for Downlink Multiuser Video Transmission Over Wireless Lossy NetworksabstractDemand for multimedia services, such as video streaming over wireless networks, has grown dramatically in recent years. The downlink transmission of multiple video sequences to multiple users over a shared resource-limited wireless channel, however, is a daunting task. Among the many challenges in this area are the time-varying channel conditions, limited available resources, such as bandwidth and power, and the different transmission requirements of different video content. This work takes into account the time-varying nature of the wireless channels, as well as the importance of individual video packets, to develop a cross-layer resource allocation and packet scheduling scheme for multiuser video streaming over lossy wireless packet access networks. Assuming that accurate channel feedback is not available at the scheduler, random channel losses combined with complex error concealment at the receiver make it impossible for the scheduler to determine the actual distortion of the sequence at the receiver. Therefore, the objective of the optimization is to minimize the expected distortion of the received sequence, where the expectation is calculated at the scheduler with respect to the packet loss probability in the channel. The expected distortion is used to order the packets in the transmission queue of each user, and then gradients of the expected distortion are used to efficiently allocate resources across users. Simulations show that the proposed scheme performs significantly better than a conventional content-independent scheme for video transmission. Ehsan Maani, Peshala V. Pahalawatta, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 4 |
| 2007 | Spatiotemporal Algorithm for Background SubtractionabstractBackground modeling and subtraction is a fundamental task in many computer vision and video processing applications. We present a novel probabilistic background modeling and subtraction method that exploits spatial and temporal dependencies between pixels. By using an initial clustering of the background scene, we model each pixel by a mixture of spatiotemporal Gaussian distributions, where each distribution represents locally a region in the neighborhood of the pixel. By extracting the local properties around each pixel, the proposed method obtains accurate models of dynamic backgrounds that are highly effective in detecting foreground objects. Experimental results for indoor and outdoor surveillance videos in comparison with other multimodal methods demonstrate the performance advantages of the proposed method. S. Derin Babacan, Thrasyvoulos N. Pappas |
ICASSP (1) | 2 |
| 2007 | Using Structural Similarity Quality Metrics to Evaluate Image Compression TechniquesabstractPerceptual image quality metrics have explicitly accounted for perceptual characteristics of the human visual system (HVS) by modeling sensitivity to subband noise as the just-noticeable threshold of distortion. While such metrics can successfully account for contrast and luminance masking, they are quite sensitive to spatial shifts, intensity shifts, contrast changes, and scale changes. In contrast, the recently proposed structural similarity (SSIM) metrics account for perception more implicitly with the assumption that the HVS is adapted for extracting structural information (relative spatial covariance) from images. As such, they have the potential to be much more effective in quantifying suprathreshold compression artifacts than traditional perceptual metrics, as such artifacts tend to distort the structure of an image. We use a (perceptually) weighted variation of the complex wavelet SSIM (CWSSIM) to evaluate standard image compression techniques such as JPEG, JPEG 2000, SPIHT, and the Safranek-Johnston perceptual image coder. Our experimental results indicate that the weighted CWSSIM generally agrees with subjective evaluations. Alan C. Brooks, Thrasyvoulos N. Pappas |
ICASSP (1) | 2 |
| 2007 | Content-Aware Resource Allocation for Scalable Video Transmission to Multiple Users Over a Wireless NetworkabstractWireless video transmission is prone to unpredictable degradations due to time-varying channel conditions. Such degradations are difficult to overcome using conventional video coding techniques. Scalable video coding offers a flexible bitstream that can be dynamically adapted to fit the prevailing channel conditions. Within a scalable video coding framework, we develop simple packet prioritization strategies, which, when combined with a reasonable error concealment scheme and a content-aware resource allocation technique, provide for robust video transmission over time-varying channels. The packet prioritization as well as the calculation of the content-aware scheduling metric can be performed offline and signaled to the wireless scheduler. Peshala V. Pahalawatta, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICASSP (1) | 2 |
| 2007 | Resource Allocation for Downlink Multiuser Video Transmission Over Wireless Lossy NetworksabstractThe emergence of 3G and 4G wireless networks brings with it the possibility of streaming high quality video content on-demand to mobile users. Wireless video applications require appropriate scheduling techniques that make use of the specific characteristics of video content, as well as the well known gains from multiuser diversity. While fast and frequent channel feedback is available in the new generation of wireless networks, the channel estimates cannot be perfect, and channel losses should be taken into account in the packet scheduling and resource allocation. The proposed scheme is formulated as a joint optimization over the resource allocation and channel loss protection, in order to minimize the distortion of the received video sequences. The distortion is a function of the packets deliberately dropped at the transmission queue due to congestion, as well as of random channel losses. The scheme makes use of a packet prioritization strategy that orders video packets based on their contribution to reducing the expected distortion of the received video sequence. Simulation results show that the proposed technique significantly outperforms content-independent packet scheduling schemes. Ehsan Maani, Peshala V. Pahalawatta, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
ICIP (5) | 4 |
| 2007 | Lossy Compression of Bilevel Images Based on Markov Random FieldsabstractA new method for lossy compression of bilevel images based on Markov random fields (MRFs) is proposed. It preserves key structural information about the image, and then reconstructs the smoothest image that is consistent with this information. The smoother the original image, the lower the required bit rate, and conversely, the lower the bit rate, the smoother the approximation provided by the decoded image. The main idea is that as long as the key structural information is preserved, then any smooth contours consistent with this information will provide an acceptable reconstructed image. The use of MRFs in the decoding stage is the key to efficient compression. Experimental results demonstrate that the new technique outperforms existing lossy compression techniques, and provides substantially lower rates than lossless techniques (JBIG) with little loss in perceived image quality. Matthew G. Reyes, Xiaonan Zhao, David L. Neuhoff, Thrasyvoulos N. Pappas |
ICIP (2) | 4 |
| 2007 | Content-Aware Resource Allocation and Packet Scheduling for Video Transmission over Wireless NetworksabstractA cross-layer packet scheduling scheme that streams pre-encoded video over wireless downlink packet access networks to multiple users is presented. The scheme can be used with the emerging wireless standards such as HSDPA and IEEE 802.16. A gradient based scheduling scheme is used in which user data rates are dynamically adjusted based on channel quality as well as the gradients of a utility function. The user utilities are designed as a function of the distortion of the received video. This enables distortion-aware packet scheduling both within and across multiple users. The utility takes into account decoder error concealment, an important component in deciding the received quality of the video. We consider both simple and complex error concealment techniques. Simulation results show that the gradient based scheduling framework combined with the content-aware utility functions provides a viable method for downlink packet scheduling as it can significantly outperform current content-independent techniques. Further tests determine the sensitivity of the system to the initial video encoding schemes, as well as to non-real-time packet ordering techniques. Peshala V. Pahalawatta, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
IEEE J. Sel. Areas Commun. | 3 |
| 2006 | Spatial Resolution and Quantization Noise Tradeoffs for Scalable Image CompressionabstractMost full-reference quality metrics compare the original image to a distorted image at the same level of resolution assuming a fixed viewing distance. In video streaming applications, however, the transmitted or received signal may differ from the original in compression as well as spatiotemporal resolution. For example, at low bitrate coding applications the compressed image may be too distorted, and hence the observer may prefer to reduce the resolution or increase the viewing distance in order to reduce the visibility of the compression artifacts. The selection of the best tradeoff between resolution/viewing distance and visibility of compression artifacts requires a quality metric that accounts for both image distortions and image size. Such tradeoffs are not reflected in existing quality metrics, which ignore the signal visibility and only measure the visibility of compression distortions, which decrease with image size. In order to better understand such tradeoffs, with the goal of developing better quality metrics, we conducted subjective tests using a number of existing still image coders (JPEG2000 SPHIT, and JPEG). Our results indicate that the objective quality (perceptually weighted PSNR) of the images that the viewers select decreases with resolution, that is, the viewers are willing to accept more artifacts as image size decreases Soo Hyun Bae, Thrasyvoulos N. Pappas, Biing-Hwang Juang |
ICASSP (2) | 2 |
| 2006 | A Perceptual Approach for Semantic Image RetrievalabstractThe rapid growth of digital imaging technology and the accumulation of large collections of digital images has created the need for efficient and intelligent schemes for image classification and retrieval. Since humans are the ultimate users of most retrieval systems, it is important to organize the contents semantically, according to meaningful categories. We propose a novel approach for assigning semantic labels to image segments, which together segment layout information can lead to content-based image classification and retrieval. The proposed approach relies on a perceptually based, spatially adaptive, color-texture segmentation scheme. We derive segment-wide features (color and spatial texture). These features serve as medium level descriptors that can effectively bridge the "semantic gap" between low and high level descriptors. The segment classification into semantic categories is based on linear discriminant analysis techniques. We demonstrate the effectiveness of the proposed approach on a database that includes 5000 segments from approximately 2000 photographs of natural scenes. Dejan Depalov, Thrasyvoulos N. Pappas, Dongge Li, Bhavan Gandhi |
ICASSP (2) | 2 |
| 2006 | Subjective Image Quality Tradeoffs Between Spatial Resolution and Quantization NoiseabstractThe importance of tradeoffs between spatial resolution and quantization noise has been examined in our previous work. Subjective experiments indicate that as the bitrate decreases, human observers generally prefer to reduce image resolution in order to maintain image quality, but the amount of distortion they are willing to accept increases with decreasing resolution. In this paper, we conducted further experiments with several images, different encoders, and a finer set of bitrates to determine the preferred resolution at each bit rate, and also the resolution at which there are no visible coding artifacts. Analysis of the subjective results using a wavelet-based perceptual quality metric verifies our earlier conclusion that human observers tend to reduce resolution in order to maintain image quality, but are willing to accept more artifacts as image size decreases. Soo Hyun Bae, Thrasyvoulos N. Pappas, Biing-Hwang Juang |
ICIP | 2 |
| 2006 | Perceptual Feature Selection for Semantic Image ClassificationabstractContent-based image retrieval has become an indispensable tool for managing the rapidly growing collections of digital images. The goal is to organize the contents semantically, according to meaningful categories. In recent papers we introduced a new approach for semantic image classification that relies on the adaptive perceptual color-texture segmentation algorithm proposed by Chen et al. This algorithm combines knowledge of human perception and signal characteristics to segment natural scenes into perceptually uniform regions. The resulting segments can be classified into semantic categories using region-wide features as medium level descriptors. Such descriptors are the key to bridging the gap between low-level image primitives and high-level image semantics. The segment classification is based on linear discriminant analysis techniques. In this paper, we examine the classification performance (precision and recall rates) when different sets of region-wide features are used. These include different color composition features, spatial texture, and segment location. We demonstrate the effectiveness of the proposed techniques on a database that includes 9000 segments from approximately 2500 photographs of natural scenes. Dejan Depalov, Thrasyvoulos N. Pappas, Dongge Li, Bhavan Gandhi |
ICIP | 2 |
| 2006 | VAPOR: variance-aware per-pixel optimal resource allocationabstractCharacterizing the video quality seen by an end-user is a critical component of any video transmission system. In packet-based communication systems, such as wireless channels or the Internet, packet delivery is not guaranteed. Therefore, from the point-of-view of the transmitter, the distortion at the receiver is a random variable. Traditional approaches have primarily focused on minimizing the expected value of the end-to-end distortion. This paper explores the benefits of accounting for not only the mean, but also the variance of the end-to-end distortion when allocating limited source and channel resources. By accounting for the variance of the distortion, the proposed approach increases the reliability of the system by making it more likely that what the end-user sees, closely resembles the mean end-to-end distortion calculated at the transmitter. Experimental results demonstrate that variance-aware resource allocation can help limit error propagation and is more robust to channel-mismatch than approaches whose goal is to strictly minimize the expected distortion. Yiftach Eisenberg, Fan Zhai, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2006 | Rate-distortion optimized hybrid error control for real-time packetized video transmissionabstractThe problem of application-layer error control for real-time video transmission over packet lossy networks is commonly addressed via joint source-channel coding (JSCC), where source coding and forward error correction (FEC) are jointly designed to compensate for packet losses. In this paper, we consider hybrid application-layer error correction consisting of FEC and retransmissions. The study is carried out in an integrated joint source-channel coding (IJSCC) framework, where error resilient source coding, channel coding, and error concealment are jointly considered in order to achieve the best video delivery quality. We first show the advantage of the proposed IJSCC framework as compared to a sequential JSCC approach, where error resilient source coding and channel coding are not fully integrated. In the USCC framework, we also study the performance of different error control scenarios, such as pure FEC, pure retransmission, and their combination. Pure FEC and application layer retransmissions are shown to each achieve optimal results depending on the packet loss rates and the round-trip time. A hybrid of FEC and retransmissions is shown to outperform each component individually due to its greater flexibility. Fan Zhai, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 3 |
| 2005 | Advances in Efficient Resource Allocation for Packet-Based Real-Time Video TransmissionabstractMultimedia applications involving the transmission of video over communication networks are rapidly increasing in popularity. Such applications can greatly benefit from adapting video coding parameters to network conditions as well as adapting network parameters to better support the application requirements. These two dimensions can both be viewed as allocating source and network resources to improve video quality. We highlight recent advances in optimal resource allocation for real-time video communications over unreliable and resource constrained communication channels. More specifically, we focus on point-to-point coding and delivery schemes in which the sequences are encoded on the fly. We present a high-level framework for resource-distortion optimization. The framework can be used for jointly considering factors across network layers, including source coding, channel resource allocation, and error concealment. For example, resources can take the form of transmission energy in a wireless channel, and transmission cost in a DiffServ-based Internet channel. This framework can be used to optimally trade off resource consumption with end-to-end video quality in packet-based video transmission. After giving an overview of this framework, we review recent work in two areas-energy efficient wireless video transmission and resource allocation for Internet-based applications. Aggelos K. Katsaggelos, Yiftach Eisenberg, Fan Zhai, Randall Berry, Thrasyvoulos N. Pappas |
Proc. IEEE | 5 |
| 2005 | Joint source-channel coding and power adaptation for energy efficient wireless video communications
Fan Zhai, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
Signal Process. Image Commun. | 3 |
| 2005 | Adaptive Perceptual Color-Texture Image SegmentationabstractWe propose a new approach for image segmentation that is based on low-level features for color and texture. It is aimed at segmentation of natural scenes, in which the color and texture of each segment does not typically exhibit uniform statistical characteristics. The proposed approach combines knowledge of human perception with an understanding of signal characteristics in order to segment natural scenes into perceptually/semantically uniform regions. The proposed approach is based on two types of spatially adaptive low-level features. The first describes the local color composition in terms of spatially adaptive dominant colors, and the second describes the spatial characteristics of the grayscale component of the texture. Together, they provide a simple and effective characterization of texture that the proposed algorithm uses to obtain robust and, at the same time, accurate and precise segmentations. The resulting segmentations convey semantic information that can be used for content-based retrieval. The performance of the proposed algorithms is demonstrated in the domain of photographic images, including low-resolution, degraded, and compressed images. Thrasyvoulos N. Pappas, Aleksandra Mojsilovic, Bernice E. Rogowitz |
IEEE Trans. Image Process. | 2 |
| 2005 | Joint source coding and packet classification for real-time video transmission over differentiated services networksabstractDifferentiated Services (DiffServ) is one of the leading architectures for providing quality of service in the Internet. We propose a scheme for real-time video transmission over a DiffServ network that jointly considers video source coding, packet classification, and error concealment within a framework of cost-distortion optimization. The selections of encoding parameters and packet classification are both used to manage end-to-end delay variations and packet losses within the network. We present two dual formulations of the proposed scheme: the minimum distortion problem, in which the objective is to minimize the end-to-end distortion subject to cost and delay constraints, and the minimum cost problem, which minimizes the total cost subject to end-to-end distortion and delay constraints. A solution to these problems using Lagrangian relaxation and dynamic programming is given. Simulation results demonstrate the advantage of jointly adapting the source coding and packet classification in DiffServ networks. Fan Zhai, Carlos E. Luna, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
IEEE Trans. Multim. | 4 |
| 2004 | DNA-based matching of digital signalsabstractAdleman with his pioneering work set the stage for the new field of bio-computing research (Science, vol.266, p.1021-1024, 1994). His main idea was to use actual chemistry to solve problems that are either unsolvable by conventional computers, or require an enormous amount of computation. The main focus of our research is to consider the application of molecular computing to the domain of digital signal processing (DSP). In this paper, we consider matching problems that arise in signal processing applications and are amenable to a DNA-based solution. Digital data are encoded in DNA sequences using a sophisticated codeword set that satisfies the noise tolerance constraint (NTC) that we introduce. NTC, one of the main contributions of our work, takes into account the presence of noise in digital signals by exploiting the annealing between non-perfect complementary sequences. We propose an algorithm to map binary values into DNA codewords by satisfying a number of constraints, including the NTC. Using that algorithm, we retrieved 128 codewords that enables us to use a DNA based approach to digital signal matching. Sotirios A. Tsaftaris, Aggelos K. Katsaggelos, Thrasyvoulos N. Pappas, Eleftherios T. Papoutsakis |
ICASSP (5) | 3 |
| 2004 | Rate-distortion optimized product code forward error correction for video transmission over IP-based wireless networksabstractThe problem of encoding and transmitting a video sequence over an IP-based wireless network, consisting of both wired and wireless links, is addressed. To combat the different types of packet loss in the heterogeneous network, the use of a product code forward error correction (FEC) scheme capable of providing unequal error protection is considered. At the transport layer, Reed-Solomon (RS) coding is used to provide inter-packet protection. In addition, rate-compatible punctured convolutional (RCPC) coding is used at the link layer to provide unequal intra-packet protection. Optimal bit allocation is performed in a rate-distortion optimized joint source-channel coding and power allocation framework to achieve the best video quality. Simulation results illustrate the advantage of the proposed product code FEC scheme over previously studied approaches. Fan Zhai, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICASSP (5) | 3 |
| 2004 | Rate-distortion optimized hybrid error control for real-time packetized video transmissionabstractIn this paper, hybrid error control for real-time video transmission is studied. The study is carried out using a proposed integrated joint source-channel coding framework, which jointly considers error resilient source coding, channel coding, and error concealment, in order to achieve the best video quality and focuses on the performance comparison of several error correction scenarios, such as forward error correction (FEC), retransmission, and the combination of both. Simulation results show that either FEC or retransmission can be optimal depending on the packet loss rates and network round trip time. The proposed hybrid FEC/retransmission scheme outperforms both. Fan Zhai, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICC | 3 |
| 2004 | Perceptually-tuned multiscale color-texture segmentationabstractWe present a perceptually-tuned multiscale image segmentation algorithm that is based on spatially adaptive color and texture features. The proposed algorithm extends a previously proposed approach to include multiple texture scales. The determination of the multiscale texture features is based on perceptual considerations. We also examine the perceptual tuning of the algorithm and how it is affected by the presence of different texture scales. The multiscale extension is necessary for segmenting higher resolution images and is particularly effective in segmenting objects shown in different perspectives. The performance of the proposed algorithm is demonstrated in the domain of photographic images. Thrasyvoulos N. Pappas, Aleksandra Mojsilovic, Bernice E. Rogowitz |
ICIP | 2 |
| 2004 | Optimal sensor selection for video-based target tracking in a wireless sensor networkabstractThe use of wireless sensor networks for target tracking is an active area of research. Imaging sensors that obtain video-rate images of a scene can have a significant impact in such networks, as they can measure vital information on the identity, position, and velocity of moving targets. Since wireless networks must operate under stringent energy constraints, it is important to identify the optimal set of imagers to be used in a tracking scenario such that the network lifetime is maximized. We formulate this problem as one of maximizing the information utility gained from a set of sensors subject to a constraint on the average energy consumption in the network. We use an unscented Kalman filter framework to solve the tracking and data fusion problem with multiple imaging sensors in a computationally efficient manner, and use a lookahead algorithm to optimize the sensor selection based on the predicted trajectory of the target. Simulation results show the effectiveness of this method of sensor selection. Peshala V. Pahalawatta, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2004 | Channel modeling and its effect on the end-to-end distortion in wireless video communicationsabstractA major limitation faced by a mobile user is their dependence on a limited battery supply. For wireless video communications, joint source coding and transmission power management (JSCPM) has recently been considered as a means of efficiently allocating transmission energy. In order to reduce complexity, the design of many of these adaptive resource allocation algorithms utilizes simplified channel models that do not account for the burstiness of the channel. We analyze the effects of such channel model simplifications on the end-to-end distortion. We present a channel model that is based on information theoretic considerations, which captures the bursty nature of wireless channels and accounts for packet lengths when calculating the probability of loss. Given the source coding and transmission parameters derived using a simplified channel model, our goal is to analyze how the end-to-end distortion is affected when a more realistic complex channel model is used to simulate losses. Experimental results suggest that the performance gain predictions for JSCPM using a simpler channel model are also valid when more sophisticated channel simulations are used, provided that a number of additional steps are taken after the optimization to account for the complex characteristics of wireless channels. Eren Soyak, Yiftach Eisenberg, Fan Zhai, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
ICIP | 5 |
| 2004 | An integrated joint source-channel coding framework for video transmission over packet lossy networksabstractThe problem of application-layer error control for real-time video transmission over packet lossy networks is commonly addressed by joint source-channel coding (JSCC). The traditional JSCC approaches solve this problem in a sequential manner, where source coding and channel coding are not fully integrated. In this paper, we present an integrated joint source-channel coding (IJSCC) framework, where error resilient source coding, channel coding and error concealment are jointly considered in an integrated manner. We show through both analysis and simulations the advantages of the proposed IJSCC approach, in comparison to a sequential JSCC approach. Fan Zhai, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP | 3 |
| 2004 | Motion-compensated wavelet video coding using adaptive mode selectionabstractA motion-compensated wavelet video coder is presented that uses adaptive mode selection (AMS) for each macroblock (MB). The block-based motion estimation is performed in the spatial domain, and an embedded zerotree wavelet coder (EZW) is employed to encode the residue frame. In contrast to other motion-compensated wavelet video coders, where all the MBs are forced to be in INTER mode, we construct the residue frame by combining the prediction residual of the INTER MBs with the coding residual of the INTRA and INTER_ENCODE MBs. Different from INTER MBs that are not coded, the INTRA and INTER_ENCODE MBs are encoded separately by a DCT coder. By adaptively selecting the quantizers of the INTRA and INTER_ENCODE coded MBs, our goal is to equalize the characteristics of the residue frame in order to improve the overall coding efficiency of the wavelet coder. The mode selection is based on the variance of the MB, the variance of the prediction error, and the variance of the neighboring MBs' residual. Simulations show that the proposed motion-compensated wavelet video coder achieves a gain of around 0.7-0.8dB PSNR over MPEG-2 TM5, and a comparable PSNR to other 2D motion-compensated wavelet-based video codecs. It also provides potential visual quality improvement. Fan Zhai, Thrasyvoulos N. Pappas |
VCIP | 2 |
| 2003 | Image segmentation by spatially adaptive color and texture featuresabstractAn image segmentation algorithm that is based on spatially adaptive color and texture features is presented. The proposed algorithm is based on a previously proposed algorithm but introduces a number of new elements. We use a new set of texture features based on a steerable filter decomposition. The steerable filters combined with a new spatial texture segmentation scheme provide a finer and more robust segmentation into texture classes. The proposed algorithm includes an elaborate border estimation procedure, which extends the idea of Pappas (1992) adaptive clustering segmentation algorithm to color texture. The performance of the proposed algorithm is demonstrated in the domain of photographic images, including low resolution compressed images. Thrasyvoulos N. Pappas, Aleksandra Mojsilovic, Bernice E. Rogowitz |
ICIP (1) | 2 |
| 2003 | Variance-aware distortion estimation for wireless video communicationsabstractThe problem of encoding and transmitting a video sequence over a wireless channel is considered. Our objective is to minimize the end-to-end distortion while using a limited amount of transmission energy and delay. In our approach, we jointly adapt the source-coding parameters and transmission power per packet. We introduce the concept of "variance-aware distortion estimation" (VADE), and present a framework for controlling both the expected value and the variance of the end-to-end distortion. This framework is based on knowledge of how the video is compressed, the probability of packet loss, and the concealment strategy. To the best of our knowledge, this paper is the first to address the trade-off between the mean and variance of the end-to-end distortion. Experimental results demonstrate the potential of the proposed approach. Yiftach Eisenberg, Fan Zhai, Carlos E. Luna, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP (1) | 4 |
| 2003 | A novel cost-distortion optimization framework for video streaming over differentiated services networksabstractThis paper presents a novel framework for streaming video over a Differentiated Services (DiffServ) network that jointly considers video source coding, packet classification and error concealment within the scope of cost-distortion optimization. Our formulation incorporates the random network delay for each packet into the calculation of the probability of packet loss and manages the end-to-end packet delay by selecting the encoding parameters and packet priority. We formulate two approaches to evaluate the performance of the proposed framework: a minimum distortion approach and a minimum cost approach, in which the encoding mode and priority class for each packet are optimally selected so as to minimize the total distortion subject to cost constraints, or to minimize the total cost subject to end-to-end distortion constraints. Simulation results demonstrate the advantage of jointly adapting the source coding and packet classification. Fan Zhai, Yiftach Eisenberg, Carlos E. Luna, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP (3) | 4 |
| 2003 | A rate-distortion optimized error control scheme for scalable video streaming over the InternetabstractVideo streaming over the Internet is a challenging task due, in part to the wide range of bandwidth variations caused by network congestion. To deal with this challenge, we propose an optimal error control scheme for scalable video transmission over the Internet. The three major components of error controlerror resilience, forward error correction (FEC), and error concealment- are considered in the proposed framework. Rate-distortion (R-D) optimization is carried out to determine the encoding mode for each packet and the channel coding rates, in order to minimize the overall expected end-to-end distortion. Our simulation study demonstrates that the proposed approach is robust to the wide range channel bandwidth variations and greatly outperforms the classical R-D optimization scheme. Fan Zhai, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
ICME | 3 |
| 2003 | Joint source coding and data rate adaptation for energy efficient wireless video streamingabstractRapid growth in wireless networks is fueling demand for video services from mobile users. While the problem of transmitting video over unreliable channels has received some attention, the wireless network environment poses challenges such as transmission power management that have received little attention previously in connection with video. Transmission power management affects battery life in mobile devices, interference to other users, and network capacity. We consider energy efficient transmission of a video sequence under delay and quality constraints. The selection of source coding parameters is considered jointly with transmitter power and rate adaptation, and packet transmission scheduling. The goal is to transmit a video frame using the minimal required transmission energy under delay and quality constraints. Experimental results are presented that illustrate the advantages of the proposed approach. Carlos E. Luna, Yiftach Eisenberg, Randall Berry, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
IEEE J. Sel. Areas Commun. | 4 |
| 2003 | An efficient rate-distortion optimal shape coding approach utilizing a skeleton-based decompositionabstractIn this paper, we present a new shape-coding approach, which decouples the shape information into two independent signal data sets; the skeleton and the boundary distance from the skeleton. The major benefit of this approach is that it allows for a more flexible tradeoff between approximation error and bit budget. Curves of arbitrary order can be utilized for approximating both the skeleton and distance signals. For a given bit budget for a video frame, we solve the problem of choosing the number and location of the control points for all skeleton and distance signals of all boundaries within a frame, so that the overall distortion is minimized. An operational rate-distortion (ORD) optimal approach using Lagrangian relaxation and a four-dimensional direct acyclic graph (DAG) shortest path algorithm is developed for solving the problem. To reduce the computational complexity from O(N(5)) to O(N(3)), where N is the number of admissible control points for a skeleton, a suboptimal greedy-trellis search algorithm is proposed and compared with the optimal algorithm. In addition, an even more efficient algorithm with computational complexity O(N(2)) that finds an ORD optimal solution using a relaxed distortion criterion is also proposed and compared with the optimal solution. Experimental results demonstrate that our proposed approaches outperform existing ORD optimal approaches, which do not follow the same decomposition of the source data. Haohong Wang, Guido M. Schuster, Aggelos K. Katsaggelos, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 4 |
| 2002 | Adaptive image segmentation based on color and textureabstractWe propose an image segmentation algorithm that is based on spatially adaptive color and texture features. The features are first developed independently, and then combined to obtain an overall segmentation. Texture feature estimation requires a finite neighborhood which limits the spatial resolution of texture segmentation, while color segmentation provides accurate and precise edge localization. We combine a previously proposed adaptive clustering algorithm for color segmentation with a simple but effective texture segmentation approach to obtain an overall image segmentation. Our focus is in the domain of photographic images with an essentially unlimited range of topics. The images are assumed to be of relatively low resolution and may be degraded or compressed. Thrasyvoulos N. Pappas, Aleksandra Mojsilovic, Bernice E. Rogowitz |
ICIP (3) | 2 |
| 2002 | Energy efficient wireless video communications for the digital set-top boxabstractIn the future, digital set-top boxes may serve as the primary access point for wireless home networks, enabling mobile users to use videoconferencing as well as streaming applications on hand-held devices. In this scenario, an important issue that must be addressed is the limited energy supply of a mobile device. This is of course a relevant issue for any wireless device. We focus on methods for efficiently utilizing transmission energy in wireless video communications. We present a general framework for the problem of minimizing the transmission energy required to provide an acceptable level of video quality. We discuss two special cases in which communication resources are adjusted simultaneously with the source coding parameters in order to provide (i) packet loss adaptation and (ii) transmission rate adaptation. Yiftach Eisenberg, Carlos E. Luna, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP (2) | 3 |
| 2002 | Optimal source coding and transmission power management using a min-max expected distortion approachabstractWe consider the problem of compressing a video sequence for transmission over a wireless channel. In our approach we jointly consider error resilience and concealment techniques, at the source coding level, and transmission power management at the physical layer. We formulate a minimum-maximum distortion problem, where our goal is to either (i) minimize the total transmission energy for a given maximum expected distortion, or (ii) minimize the maximum expected distortion at the receiver for a given maximum transmission energy. Experimental results show that simultaneously adjusting the source coding and transmission power is more energy efficient than considering these factors separately. Carlos E. Luna, Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP (1) | 3 |
| 2002 | Joint source coding and transmission power management for energy efficient wireless video communicationsabstractWe consider a situation where a video sequence is to be compressed and transmitted over a wireless channel. Our goal is to limit the amount of distortion in the received video sequence, while minimizing transmission energy. To accomplish this goal, we consider error resilience and concealment techniques at the source coding level, and transmission power management at the physical layer. We jointly consider these approaches in a novel framework. In this setting, we formulate and solve an optimization problem that corresponds to minimizing the energy required to transmit video under distortion and delay constraints. Experimental results show that simultaneously adjusting the source coding and transmission power is more energy efficient than considering these factors separately. Yiftach Eisenberg, Carlos E. Luna, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2001 | Minimizing transmission energy in wireless video communicationsabstractA key constraint in mobile communications is the reliance on a battery with a limited energy supply. Efficiently utilizing the available energy is therefore an important design consideration. We consider a situation where a video sequence is to be compressed and transmitted over a wireless channel. The goal is to limit the amount of distortion in the received video sequence while using the minimum required transmission energy. To accomplish this goal, we consider error resilience and concealment techniques, at the source coding level, as well as the dynamic allocation of physical layer communication resources. We consider these approaches jointly in a novel framework. We formulate an optimization problem that corresponds to minimizing the energy required to transmit a video frame with an acceptable level of distortion. We present methods for solving this problem and other extensions. Yiftach Eisenberg, Thrasyvoulos N. Pappas, Randall Berry, Aggelos K. Katsaggelos |
ICIP (1) | 2 |
| 2001 | A robust and efficient algorithm for bilevel document block classificationabstractWe present a robust and computationally efficient algorithm for the classification of blocks of bilevel machine-printed documents into text and halftone categories. It uses a simple mask that makes use of the different correlation properties between the text and halftone regions, and has comparable or better performance than more sophisticated and computationally intensive spectral analysis techniques. The proposed algorithm is a key component of a document recognition system that segments a document into regions, classifies them into text, halftone, line-art, etc., and then analyzes the regions to obtain a document interpretation. The input data are unusually challenging: multilingual, unoriented (e.g., upside down), and range from ideal (machine-generated) images to very low quality (e.g., copied and faxed) images. We test the proposed algorithm on the University of Washington database and demonstrate its performance on a variety of images from different databases, as well as synthetic images. Thrasyvoulos N. Pappas, Snow H. Tseng, David A. Kosiba |
ICIP (1) | 1 |
| 2001 | Rate-distortion optimal skeleton-based shape codingabstractWe present a new shape-coding approach, which decouples the shape information into two independent data sets, the skeleton and the distance of the boundary from the skeleton. The major benefit of this approach is that it allows a more flexible trade-off between accuracy of the approximation and bit-allocation cost, and thus, provides the possibility of better performance in the operational rate-distortion (ORD) optimal sense than other reported techniques. The characteristics of these data sets are studied and various approximation approaches are applied on each of them to reach an ORD optimal result. We apply, for example, polygonal approximation on both the skeleton and distance data. We demonstrate that the resulting approach outperforms existing ORD optimal approaches. Haohong Wang, Thrasyvoulos N. Pappas, Aggelos K. Katsaggelos |
ICIP (2) | 2 |
| 1999 | Least-squares model-based halftoningabstractA least-squares model-based (LSMB) approach to digital halftoning is proposed. It exploits both a printer model and a model for visual perception. It attempts to produce an optimal halftoned reproduction, by minimizing the squared error between the response of the cascade of the printer and visual models to the binary image and the response of the visual model to the original gray-scale image. It has been shown that the one-dimensional (1-D) least-squares problem, in which each row or column of the image is halftoned independently, can be implemented using the Viterbi algorithm to obtain the globally optimal solution. Unfortunately, the Viterbi algorithm cannot be used in two dimensions. In this paper, the two-dimensional (2-D) least-squares solution is obtained by iterative techniques, which are only guaranteed to produce a total optimum. Experiments show that LSMB halftoning produces better textures and higher spatial and gray-scale resolution than conventional techniques. We also show that the least-squares approach eliminates most of the problems associated with error diffusion. We investigate the performance of the LSMB algorithms over a range of viewing distances, or equivalently, printer resolutions. We also show that the LSMB approach gives us precise control of image sharpness. Thrasyvoulos N. Pappas, David L. Neuhoff |
IEEE Trans. Image Process. | 1 |
| 1997 | Model-based halftoning of color imagesabstractWe present a new class of models for color printers. They form the basis for model-based techniques that exploit the characteristics of the printer and the human visual system to maximize the quality of the printed images. We present two model-based techniques, the modified error diffusion (MED) algorithm and the least-squares model-based (LSMB) algorithm. Both techniques are extensions of the gray-scale model-based techniques and produce images with high spatial resolution and visually pleasant textures. We also examine the use of printer models for designing blue-noise screens. The printer models cam account for a variety of printer characteristics. We propose a specific printer model that accounts for overlap between neighboring dots of ink and the spectral absorption properties of the inks. We show that when we assume a simple "one-minus-RGB" relationship between the red, green, and blue image specification and the corresponding cyan, magenta, and yellow inks, the algorithms are separable. Otherwise, the algorithms are not separable and the modified error diffusion may be unstable, The experimental results consider the separable algorithms that produce high-quality images for applications where the exact colorimetric reproduction of color is not necessary. They are computationally simple and robust to errors in color registration, but the colors are device dependent. Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 1 |
| 1996 | Supra-threshold perceptual image codingabstractWe investigate different algorithms and performance criteria for supra-threshold image compression. The algorithms include JPEG and perceptual JPEG, the Safranek-Johnston perceptual subband image coder (PIC), and the Said-Pearlman algorithm which is based on Shapiro's embedded zerotree wavelet algorithm. We also consider a number of performance criteria. These include mean-squared error, Watson's perceptual metric, a metric based on the PIC coder, as well as an eye-filter weighted mean-squared-error metric. Our experiments indicate that the PIC metric provides the best correlation with subjective evaluations. The metric predicts that at very low bit rates the Said-Pearlman algorithm and the 8/spl times/8 subband PIC coder perform the best, while at high bit rates the 4/spl times/4 subband PIC coder dominates. Thrasyvoulos N. Pappas, Thomas A. Michel, Raynard O. Hinds |
ICIP (1) | 1 |
| 1995 | An adaptive clustering algorithm for segmentation of video sequencesabstractWe present a Bayesian approach for segmenting a sequence of gray-scale images to obtain a binary sketch. We extend a 2-D algorithm to video sequences. The 2-D algorithm is an adaptive thresholding scheme that uses spatial constraints and takes into consideration the local intensity characteristics of the image. We model the segmentation distribution as a 3-D Gibbs random field. We add temporal constraints and temporal local intensity adaptation to ensure a smooth transition of the segmentation from frame to frame. For computational efficiency as well as performance we use a multi-resolution approach. We also consider several suboptimal implementations to reduce the delay as well as the amount of computation. We tested the performance of the algorithm on head and shoulders video sequences. The algorithm achieves accurate rendering of the lip and eye movements and preserves the main characteristics of the face, so that it is easily recognizable. Raynard O. Hinds, Thrasyvoulos N. Pappas |
ICASSP | 2 |
| 1995 | Printer models and error diffusionabstractA new model-based approach to digital halftoning is proposed. It is intended primarily for laser printers, which generate "distortions" such as "dot overlap". Conventional methods, such as clustered-dot ordered dither, resist distortions at the expense of spatial and gray-scale resolution. The proposed approach relies on printer models that predict distortions, and rather than merely resisting them, it exploits them to increase, rather than decrease, both spatial and gray-scale resolution. We propose a general framework for printer models and find a specific model for laser printers. As an example of model-based halftoning, we propose a modification of error diffusion, which is often considered the best halftoning method for CRT displays with no significant distortions. The new version exploits the printer model to extend the benefits of error diffusion to printers. Experiments show that it provides high-quality reproductions with reasonable complexity. The proposed modified error diffusion technique is compared with Stucki's (1981) MECCA, which is a similar but not widely known technique that accounts for dot overlap. Model-based halftoning can be especially useful in transmission of high-quality documents using high-fidelity gray-scale image encoders. Thrasyvoulos N. Pappas, David L. Neuhoff |
IEEE Trans. Image Process. | 1 |
| 1994 | Perceptual coding of images for halftone displayabstractThe authors present a new technique for coding gray-scale images for facsimile transmission and printing on a laser printer. They use a gray-scale image encoder so that it is only at the receiver that the image is converted to a binary pattern and printed. The conventional approach is to transmit the image in halftoned form, using entropy coding (e.g., CCITT Group 3 or JBIG). The main advantages of the new approach are that one can get higher compression rates and that the receiver can tune the halftoning process to the particular printer. They use a perceptually based subband coding approach. It uses a perceptual masking model that was empirically derived for printed images using a specific printer and halftoning technique. In particular, they used a 300 dots/inch write-black laser printer and a standard halftoning scheme ("classical") for that resolution. For nearly transparent coding of gray-scale images, the proposed technique requires lower rates than the standard facsimile techniques. David L. Neuhoff, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 2 |
| 1994 | Perceptual coding of images for halftone displayabstractWe present a new technique for coding gray-scale images for facsimile transmission and printing on a laser printer. We use a gray-scale image encoder so that it is only at the receiver that the image is converted to a binary pattern and printed. The conventional approach is to transmit the image in halftoned form, using entropy coding (e.g. CCITT Group 3 or JBIG). The main advantages of the new approach are that we can get higher compression rates and that the receiver can tune the halftoning process to the particular printer. We use a perceptually based subband coding approach. It uses a perceptual masking model that was empirically derived for printed images using a specific printer and halftoning technique. In particular, we used a 300 dots/inch write-black laser printer and a standard halftoning scheme ("classical") for that resolution. For nearly transparent coding of gray-scale images, the proposed technique requires lower rates than the standard facsimile techniques. David L. Neuhoff, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 2 |
| 1993 | Printer models and colo halftoning
Thrasyvoulos N. Pappas |
ICASSP (5) | 1 |
| 1992 | One-dimensional least-squares model-based halftoningabstractA least-squares model-based approach to digital halftoning is proposed. It exploits both a printer model and a model for visual perception. It obtains an optimal halftoned reproduction by minimizing the squared error between the response of the cascade of the printer and visual models to the binary image and the response of the visual model to the original gray-scale image. Least-squares model-based halftoning uses explicit eye models and relies on printer models that predict distortions and exploit them to increase, rather than decrease, both spatial and gray-scale resolution. The authors examine the one-dimensional case, in which each row or column of the image is halftoned independently. One-dimensional least-squares halftoning is implemented, in closed form, with the Viterbi algorithm. Experiments show that it produces better spatial and gray-scale resolution than conventional one-dimensional techniques and eliminates the problems associated with the modified (to account for printer distortions) error diffusion algorithm.> David L. Neuhoff, Thrasyvoulos N. Pappas, Nambi Seshadri |
ICASSP | 2 |
| 1991 | Perceptual coding of images for halftone displayabstractA new technique is presented for coding gray-scale images for facsimile transmission and printing on a laser printer. The authors use a gray-scale image encoder so that it is only at the receiver that the image is converted to a binary pattern and printed. The conventional approach is to transmit the image in halftoned form, using entropy coding (e.g. CCITT Group 3, or JBIG). The main advantages of the new approach are that it is possible to get higher compression rates and that the receiver can tune the halftoning process to the particular printer. A perceptually based subband coding approach is used.> David L. Neuhoff, Thrasyvoulos N. Pappas |
ICASSP | 2 |
| 1989 | An adaptive clustering algorithm for image segmentationabstractA generalization of the K-means clustering algorithm to include spatial constraints and to account for local intensity variations in the image is proposed. Spatial constraints are included by the use of a Gibbs random field model. Local intensity variations are accounted for in an iterative procedure involving averaging over a sliding window whose size decreases as the algorithm progresses. Results with an eight-neighbor Gibbs random field model applied to pictures of industrial objects and a variety of other images show that the algorithm performs better than the K-means algorithm and its nonadaptive extensions.> Thrasyvoulos N. Pappas, Nikil Jayant |
ICASSP | 1 |
| 1988 | An Adaptive Clustering Algorithm For Image SegmentationabstractA generalization of the K-means clustering algorithm to include spatial constraints and to account for local intensity variations in the image is proposed. Spatial constraints are included by the use of a Gibbs random field model. Local intensity variations are accounted for in an iterative procedure involving averaging over a sliding window whose size decreases as the algorithm progresses. Results with an eight-neighbor Gibbs random field model applied to pictures of industrial objects and a variety of other images show that the algorithm performs better than the K-means algorithm and its nonadaptive extensions. > Thrasyvoulos N. Pappas, Nikil Jayant |
ICCV | 1 |