EDBT 2026 Demo / reviewers in the wild / expert
Amir Said
dblp:s/AmirSaid
· DBLP profile ↗
78ranked-venue papers
30as first author
11since 2021 · last 2025
0000-0002-1802-3513ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 66 · 24 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Systems, architecture and hardware · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 5 · 4 first-authorTheory of computation · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sort-free Gaussian Splatting via Weighted Sum RenderingabstractRecently, 3D Gaussian Splatting (3DGS) has emerged as a significant advancement in 3D scene reconstruction, attracting considerable attention due to its ability to recover high-fidelity details while maintaining low complexity. Despite the promising results achieved by 3DGS, its rendering performance is constrained by its dependence on costly non-commutative alpha-blending operations. These operations mandate complex view dependent sorting operations that introduce computational overhead, especially on the resource-constrained platforms such as mobile phones. In this paper, we propose Weighted Sum Rendering, which approximates alpha blending with weighted sums, thereby removing the need for sorting. This simplifies implementation, delivers superior performance, and eliminates the ``popping'' artifacts caused by sorting. Experimental results show that optimizing a generalized Gaussian splatting formulation to the new differentiable rendering yields competitive image quality. The method was implemented and tested in a mobile device GPU, achieving on average $1.23\times$ faster rendering. Qiqi Hou, Randall Rauwendaal, Hoang Le, Farzad Farhadzadeh, Fatih Porikli, Alex Bourd, Amir Said |
ICLR | 8 |
| 2024 | Low-Latency Neural Stereo StreamingabstractThe rise of new video modalities like virtual reality or autonomous driving has increased the demand for efficient multi-view video compression methods, both in terms of rate-distortion (R-D) performance and in terms of delay and runtime. While most recent stereo video compression approaches have shown promising performance, they compress left and right views sequentially, leading to poor parallelization and runtime performance. This work presents Low-Latency neural codec for Stereo video Streaming (LLSS), a novel parallel stereo video coding method designed for fast and efficient low-latency stereo video streaming. Instead of using a sequential cross-view motion compensation like existing methods, LLSS introduces a bidirectional feature shifting module to directly exploit mutual information among views and encode them effectively with a joint cross-view prior model for entropy coding. Thanks to this design, LLSS processes left and right views in parallel, minimizing latency; all while substantially improving R-D performance compared to both existing neural and conventional codecs. Qiqi Hou, Farzad Farhadzadeh, Amir Said, Guillaume Sautière, Hoang Le |
CVPR | 3 |
| 2024 | Neural Graphics Texture Compression Supporting Random Access
Farzad Farhadzadeh, Qiqi Hou, Hoang Le, Amir Said, Randall Rauwendaal, Alex Bourd, Fatih Porikli |
ECCV (37) | 4 |
| 2024 | MobileNVC: Real-time 1080p Neural Video Compression on a Mobile DeviceabstractNeural video codecs have recently become competitive with standard codecs such as HEVC in the low-delay setting. However, most neural codecs are large floating-point networks that use pixel-dense warping operations for temporal modeling, making them too computationally expensive for deployment on mobile devices. Recent work has demonstrated that running a neural decoder in real time on mobile is feasible, but shows this only for 720p RGB video.This work presents the first neural video codec that decodes 1080p YUV420 video in real time on a mobile device. Our codec relies on two major contributions. First, we design an efficient codec that uses a block-based motion compensation algorithm available on the warping core of the mobile accelerator, and we show how to quantize this model to integer precision. Second, we implement a fast decoder pipeline that concurrently runs neural network components on the neural signal processor, parallel entropy coding on the mobile GPU, and warping on the warping core. Our codec outperforms the previous on-device codec by a large margin with up to 48 % BD-rate savings, while reducing the MAC count on the receiver side by 10×. We perform a careful ablation to demonstrate the effect of the introduced motion compensation scheme, and ablate the effect of model quantization. Ties van Rozendaal, Tushar Singhal, Hoang Le, Guillaume Sautière, Amir Said, Krishna Buska, Anjuman Raha, Dimitrios Kalatzis, Hitarth Mehta, Frank Mayer, Markus Nagel, Auke J. Wiggers |
WACV | 5 |
| 2023 | Bitstream Organization for Parallel Entropy Coding on Neural Network-based Video CodecsabstractVideo compression systems must support increasing bandwidth and data throughput at low cost and power, and can be limited by entropy coding bottlenecks. Efficiency can be greatly improved by parallelizing coding, which can be done at much larger scales with new neural-based codecs, but with some compression loss related to data organization. We analyze the bit rate overhead needed to support multiple bitstreams for concurrent decoding, and for its minimization propose a method for compressing parallel-decoding entry points, using bidirectional bitstream packing, and a new form of jointly optimizing arithmetic coding termination. It is shown that those techniques significantly lower the overhead, making it easier to reduce it to a small fraction of the average bitstream size, like, for example, less than 1% and 0.1% when the average number of bitstream bytes is respectively larger than 95 and 1,200 bytes. Amir Said, Hoang Le, Farzad Farhadzadeh |
ISM | 1 |
| 2023 | Boosting neural video codecs by exploiting hierarchical redundancyabstractIn video compression, coding efficiency is improved by reusing pixels from previously decoded frames via motion and residual compensation. We define two levels of hierarchical redundancy in video frames: 1) first-order: redundancy in pixel space, i.e., similarities in pixel values across neighboring frames, which is effectively captured using motion and residual compensation, 2) second-order: redundancy in motion and residual maps due to smooth motion in natural videos. While most of the existing neural video coding literature addresses first-order redundancy, we tackle the problem of capturing second-order redundancy in neural video codecs via predictors. We introduce generic motion and residual predictors that learn to extrapolate from previously decoded data. These predictors are lightweight, and can be employed with most neural video codecs in order to improve their rate-distortion performance. Moreover, while RGB is the dominant colorspace in neural video coding literature, we introduce general modifications for neural video codecs to embrace the YUV420 colorspace and report YUV420 results. Our experiments show that using our predictors with a well-known neural video codec leads to 38% and 34% bitrate savings in RGB and YUV420 colorspaces measured on the UVG dataset. Reza Pourreza 0002, Hoang Le, Amir Said, Guillaume Sautière, Auke J. Wiggers |
WACV | 3 |
| 2022 | GameCodec: Neural Cloud Gaming Video Codec
Hoang Le, Reza Pourreza 0002, Amir Said, Guillaume Sautière, Auke J. Wiggers |
BMVC | 3 |
| 2022 | Optimized Learned Entropy Coding Parameters for Practical Neural-Based Image and Video CompressionabstractNeural-based image and video codecs are significantly more power-efficient when weights and activations are quantized to low-precision integers. While there are general-purpose techniques for reducing quantization effects, large losses can occur when specific entropy coding properties are not considered. This work analyzes how entropy coding is affected by parameter quantizations, and provides a method to minimize losses. It is shown that, by using a certain type of coding parameters to be learned, uniform quantization becomes practically optimal, also simplifying the minimization of code memory requirements. The mathematical properties of the new representation are presented, and its effectiveness is demonstrated by coding experiments, showing that good results can be obtained with precision as low as 4 bits per network output, and practically no loss with 8 bits. Amir Said, Reza Pourreza 0002, Hoang Le |
ICIP | 1 |
| 2022 | MobileCodec: neural inter-frame video compression on mobile devicesabstractRealizing the potential of neural codecs on real-world mobile devices is a big technological challenge due to the inherent conflict between the computational complexity of deep networks and the power-constrained mobile hardware performance. We demonstrate practical feasibility by leveraging Qualcomm's innovation and technology, bridging the gap from neural network-based model simulations to operation on a mobile device powered by Snapdragon® technology. We show the first-ever inter-frame neural video decoder running on a commercial mobile phone, decompressing high-definition videos in real-time while maintaining a low bitrate and high visual quality, comparable to conventional codecs. Hoang Le, Amir Said, Guillaume Sautière, Yang Yang 0010, Pranav Shrestha, Reza Pourreza 0002, Auke J. Wiggers |
MMSys | 3 |
| 2022 | Differentiable bit-rate estimation for neural-based video codec enhancementabstractNeural networks (NN) can improve standard video compression by pre-and post-processing the encoded video. For optimal NN training, the standard codec needs to be replaced with a codec proxy that can provide derivatives of estimated bit-rate and distortion, which are used for gradient back-propagation. Since entropy coding of standard codecs is designed to take into account non-linear dependencies between transform coefficients, bit-rates cannot be well approximated with simple per-coefficient estimators. This paper presents a new approach for bit-rate estimation that is similar to the type employed in training end-to-end neural codecs, and able to efficiently take into account those statistical dependencies. It is defined from a mathematical model that provides closed-form formulas for the estimates and their gradients, reducing the computational complexity. Experimental results demonstrate the method’s accuracy in estimating HEVC/H.265 codec bit-rates. Amir Said, Manish Kumar Singh 0002, Reza Pourreza 0002 |
PCS | 1 |
| 2021 | Progressive Neural Image Compression With Nested Quantization And Latent OrderingabstractWe present PLONQ, a progressive neural image compression scheme which pushes the boundary of variable bitrate compression by allowing quality scalable coding with a single bitstream. In contrast to existing learned variable bitrate solutions which produce separate bitstreams for each quality, it enables easier rate-control and requires less storage. Leveraging the latent scaling based variable bitrate solution, we introduce nested quantization, a method that defines multiple quantization levels with nested quantization grids, and progressively refines all latents from the coarsest to the finest quantization level. To achieve finer progressiveness in between any two quantization levels, latent elements are incrementally refined with an importance ordering defined in the rate-distortion sense. To the best of our knowledge, PLONQ is the first learning-based progressive image coding scheme and it outperforms SPIHT, a well-known wavelet-based progressive image codec. Yadong Lu, Yinhao Zhu, Yang Yang 0010, Amir Said, Taco Cohen |
ICIP | 4 |
| 2020 | Parametric Graph-Based Separable Transforms For Video CodingabstractIn many video coding systems, separable transforms (such as two-dimensional DCT-2) have been used to code block residual signals obtained after prediction. This paper proposes a parametric approach to build graph-based separable transforms (GBSTs) for video coding. Specifically, a GBST is derived from a pair of line graphs, whose weights are determined based on two non-negative parameters. As certain choices of those parameters correspond to the discrete sine and cosine transform types used in recent video coding standards (including DCT-2, DST-7 and DCT-8), this paper further optimizes these graph parameters to better capture residual block statistics and improve video coding efficiency. The proposed GBSTs are tested on the Versatile Video Coding (VVC) reference software, and the experimental results show that about 0.4% average coding gain is achieved over the existing set of separable transforms constructed based on DCT-2, DST-7 and DCT-8 in VVC. Hilmi E. Egilmez, Oguzhan Teke, Amir Said, Vadim Seregin, Marta Karczewicz |
ICIP | 3 |
| 2020 | Parallelized Rate-Distortion Optimized Quantization Using Deep LearningabstractRate-Distortion Optimized Quantization (RDOQ) has played an important role in the coding performance of recent video compression standards such as H.264/AVC, H.265/HEVC, VP9 and AV1. This scheme yields significant reductions in bit-rate at the expense of relatively small increases in distortion. Typically, RDOQ algorithms are prohibitively expensive to implement on real-time hardware encoders due to their sequential nature and their need to frequently obtain entropy coding costs. This work addresses this limitation using a neural network-based approach, which learns to trade-off rate and distortion during offline supervised training. As these networks are based solely on standard arithmetic operations that can be executed on existing neural network hardware, no additional area-on-chip needs to be reserved for dedicated RDOQ circuitry. We train two classes of neural networks, a fully-convolutional network and an auto-regressive network, and evaluate each as a post-quantization step designed to refine cheap quantization schemes such as scalar quantization (SQ). Both network architectures are designed to have a low computational overhead. After training they are integrated into the HM 16.20 implementation of HEVC, and their video coding performance is evaluated on a subset of the H.266/VVC SDR common test sequences. Comparisons are made to RDOQ and SQ implementations in HM16.20. Our method achieves 1.64% BD-rate savings on luminosity compared to the HM SQ anchor, and on average reaches 45% of the performance of the iterative HM RDOQ algorithm. Dana Kianfar, Auke J. Wiggers, Amir Said, Reza Pourreza 0002, Taco Cohen |
MMSP | 3 |
| 2020 | Hybrid Video Codec Based on Flexible Block Partitioning With Extensions to the Joint Exploration ModelabstractThis article describes the main video coding technologies included in a joint proposal submitted by Qualcomm and Technicolor, in response to a Call for Proposals (CfP) issued by ITU-T SG16 WP3 Q.6 (VCEG) and ISO/IEC JTC1/SC29/WG11 (MPEG) in Oct. 2017. The proposal contains the majority of the tools that have been adopted into the Joint Exploration Model (JEM), developed in the exploratory phase that preceded the CfP. A flexible multi-tree type (MTT) block-partitioning scheme is proposed to extend the quadtree and binary tree (QTBT) based partitioning in JEM by including triple tree (TT) and asymmetric binary tree (ABT) partitions. In addition, several JEM tools in intra and inter prediction, transforms and arithmetic coding are modified, and new tools such as sign prediction and motion compensated padding are proposed. Objective standard dynamic range (SDR) gains of 43.1% and 15.5% in terms of average luma BD-rate improvement have been achieved for the CfP constraint set 1 (random-access configuration) relative to HEVC/H.265 (HM) and JEM anchors, respectively. For the CfP constraint set 2 (low-delay configuration), the average luma BD-rate improvements are 33.7% relative to the HM anchor and 12.7% relative to the JEM anchor. The proposed codec scored highly in both subjective evaluations and objective metrics and was among the best-performing CfP proposals. Wei-Jung Chien, Muhammed Z. Coban, Hilmi E. Egilmez, Marta Karczewicz, Amir Said, Vadim Seregin, Geert Van der Auwera, Philippe Bordes, Franck Galpin, Fabrice Le Léannec, Tangi Poirier, Fabrice Urban |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2019 | Low-Complexity Transform Adjustments for Video CodingabstractRecent video codecs with multiple separable transforms can achieve significant coding gains using asymmetric trigonometric transforms (DCTs and DSTs), because they can exploit diverse statistics of residual block signals. However, they add excessive computational and memory complexity on large transforms (32-point and larger), since their practical software and hardware implementations are not as effi-cient as of the DCT-2. This article introduces a novel technique to design low-complexity approximations of trigonometric transforms. The proposed method uses DCT-2 computations, and applies orthogonal "adjustments" to approximate the most important basis vectors of the desired transform. Experimental results on the Versatile Video Coding (VVC) reference software show that the proposed approach significantly reduces the computational complexity, while providing practically identical coding efficiency. Amir Said, Hilmi E. Egilmez, Yung Hsuan Chao |
ICIP | 1 |
| 2018 | Low-Complexity Intra Prediction Refinements for Video CodingabstractIn existing video coding standards such as H.264/AVC and HEVC, the intra prediction is typically derived using fixed, symmetric prediction filters along the prediction direction, e.g., in planar mode, top-right and bottom-left samples are predicted using symmetric prediction filters. However, in case of asymmetric availability of neighboring reference samples, the performance of intra prediction filters designed in HEVC may not be optimal. To further refine the intra prediction and achieve higher accuracy of prediction samples, this paper proposes low-complexity refinements over HEVC intra prediction, which are applied on frequently used planar, DC, horizontal and vertical modes. The proposed method only requires simple addition and bit-shift operations on top of HEVC's intra prediction implementation. Experimental results show that, an average of 0.7% coding gain is achieved for intra coding with no increase in run-time complexity. Xin Zhao 0003, Vadim Seregin, Amir Said, Kai Zhang 0007, Hilmi E. Egilmez, Marta Karczewicz |
PCS | 3 |
| 2018 | Joint Separable and Non-Separable Transforms for Next-Generation Video CodingabstractThroughout the past few decades, the separable Discrete Cosine Transform (DCT), particularly the DCT type II, has been widely used in image and video compression. It is well known that, under first-order stationary Markov conditions, DCT is an efficient approximation of the optimal Karhunen-Loève transform. However, for natural image and video sources, the adaptivity of a single separable transform with fixed core is rather limited for the highly dynamic image statistics, e.g., textures and arbitrarily directed edges. It is also known that non-separable transforms can achieve better compression efficiency for images with directional texture patterns, yet they are computationally complex, especially when the transform size is large. In order to achieve higher transform coding gains with relatively low-complexity implementations, we propose a joint separable and non-separable transform. The proposed separable primary transform, named Enhanced Multiple Transform (EMT), applies multiple transform cores from a pre-defined subset of sinusoidal transforms, and the transform selection is signaled in a joint block level manner. Moreover, a Non-Separable Secondary Transform (NSST) method is proposed to operate in conjunction with EMT. Unlike the existing non-separable transform schemes which require excessive amounts of memory and computation, the proposed NSST efficiently improves coding gain with much lower complexity. Extensive experimental results show that the proposed methods, in a state-of-the-art video codec, such as HEVC, can provide significant coding gains (average 6.9% and 4.5% bitrate reductions for intra and random-access coding, respectively). Xin Zhao 0003, Jianle Chen, Marta Karczewicz, Amir Said, Vadim Seregin |
IEEE Trans. Image Process. | 4 |
| 2018 | Quality of Experience in a Stereoscopic Multiview EnvironmentabstractIn this paper, we investigate how visualization factors, such as disparity, mobility, angular resolution, and viewpoint interpolation, influence the quality of experience (QoE) in a stereoscopic multiview environment. In order to do so, we set up a dedicated testing room and conducted subjective experiments. We also developed a framework that emulates a supermultiview environment. This framework can be used to investigate and assess the effects of angular resolution and viewpoint interpolation on the QoE produced by multiview systems, and provide relevant cues as to how the baselines of cameras and interpolation strategies in such systems affect user experience. Aspects such as visual comfort, model fluidity, sense of immersion, and the three-dimensional (3D) experience as a whole have been assessed for several test cases. Obtained results suggest that user experience in an motion parallax environment is not as critically influenced by configuration parameters such as disparity as initially thought. In addition, extensive subjective tests have indicated that while users are very sensitive to angular resolution in multiview 3D systems, this sensitivity tends not to be as critical when a user is performing a task that involves a great amount of interaction with the multiview content. These tests have also indicated that interpolating intermediate viewpoints can be effective in reducing the required view density without degrading the perceived QoE. Felipe M. Lopes Ribeiro, José F. L. de Oliveira, Alexandre G. Ciancio, Eduardo A. B. da Silva, Cassius D. Estrada, Luiz G. C. Tavares, Jonathan N. Gois, Amir Said, Marcela C. Martelotte |
IEEE Trans. Multim. | 8 |
| 2016 | Position dependent prediction combination for intra-frame video codingabstractIntra-frame prediction in the High Efficiency Video Coding (HEVC) standard can be empirically improved by applying sets of recursive two-dimensional filters to the predicted values. However, this approach does not allow (or complicates significantly) the parallel computation of pixel predictions. In this work we analyze why the recursive filters are effective, and use the results to derive sets of non-recursive predictors that have superior performance. We present an extension to HEVC intra prediction that combines values predicted using non-filtered and filtered (smoothed) reference samples, depending on the prediction mode, and block size. Simulations using the HEVC common test conditions show that a 2.0% bit rate average reduction can be achieved compared to HEVC, for All Intra (AI) configurations. Amir Said, Xin Zhao 0003, Marta Karczewicz, Jianle Chen |
ICIP | 1 |
| 2016 | Highly efficient non-separable transforms for next generation video codingabstractFor the last few decades, the application of signal-adaptive transform coding to video compression has been stymied by the large computational complexity of matrix-based solutions. In this paper, we propose a novel parametric approach to greatly reduce the complexity without degrading the compression performance. In our approach, instead of following the conventional technique of identifying full transform matrices that yield best compression efficiency, we look for the best transform parameters defining a new class of transforms, called HyGTs, which have low complexity implementations that are easy to parallelize. The proposed HyGTs are implemented as an extension of High Efficiency Video Coding (HEVC), and our comprehensive experimental results demonstrate that proposed HyGTs improve average coding gain by 6% bit rate reduction, while using 6.8 times less memory than KLT matrices. Amir Said, Xin Zhao 0003, Marta Karczewicz, Hilmi E. Egilmez, Vadim Seregin, Jianle Chen |
PCS | 1 |
| 2016 | NSST: Non-separable secondary transforms for next generation video codingabstractIn traditional image and video coding schemes, separable transforms are typically employed due to their low-complexity implementations. However, the compression efficiency of separable transforms is limited for most natural image/video blocks which generally have arbitrarily directed edge and texture patterns. It is well known that non-separable transforms can achieve better compression efficiency for directional texture patterns, yet they are computationally complex, especially for larger block sizes. In order to achieve higher transform coding gains with relatively low-complexity implementations, in this paper, we propose non-separable secondary transforms (NSSTs). The proposed approach applies a secondary non-separable transform on a sub-block of low frequency coefficients generated using a primary separable transform, such as discrete cosine transform (DCT). Since the proposed NSST is a non-separable transform applied on low frequency coefficients in a much smaller block size, which typically captures most of the signal energy, better coding gains can be achieved with at a relatively low-computational cost. Experimental results show that, compared to the latest HEVC reference software (HM16.6), the proposed method achieves up to a significant 12% coding gain for Intra coding. Xin Zhao 0003, Jianle Chen, Amir Said, Vadim Seregin, Hilmi E. Egilmez, Marta Karczewicz |
PCS | 3 |
| 2015 | A Parallelization Framework for High Throughput Entropy CodingabstractWe propose a general framework for parallel entropy coding in media compression, which preserves compression efficiency, and is well matched to future generations of general-purpose or custom processors. Similarly to some previous parallelization methods, it is based on the fact that optimal compression is not affected by the arrangement of coded bits, but it goes further in exploiting the decreasing cost of data processing and memory. We use finite-state-machine models for identifying the best manner of separating data into segments that can be processed independently, while minimizing compression losses. Additional advantages include the ability to use, within this framework, increasingly more complex data modeling techniques, and the freedom to mix different types of coding. We confirm the parallelization effectiveness using coding simulations that run on multi-core processors, and show how throughput scales with the number of cores. Amir Said, Abo Talib Mahfoodh |
DCC | 1 |
| 2015 | Graph-based transforms for inter predicted video codingabstractIn video coding, motion compensation is an essential tool to obtain residual block signals whose transform coefficients are encoded. This paper proposes novel graph-based transforms (GBTs) for coding inter-predicted residual block signals. Our contribution is twofold: (i) We develop edge adaptive GBTs (EA-GBTs) derived from graphs estimated from residual blocks, and (ii) we design template adaptive GBTs (TA-GBTs) by introducing simplified graph templates generating different set of GBTs with low transform signaling overhead. Our experimental results show that proposed methods significantly outperform traditional DCT and KLT in terms of rate-distortion performance. Hilmi E. Egilmez, Amir Said, Yung Hsuan Chao, Antonio Ortega |
ICIP | 2 |
| 2014 | Non-causal encoding of predictively coded samplesabstractWe study the INTRA-prediction operation employed in hybrid video coders where sample values to be compressed are predicted using previously-decoded context values and the prediction errors are transform coded. Given the context values this procedure can be argued to be rate-distortion optimal for Gaussian signals. Yet natural images and video contain many structures that do not readily fit into Gaussian signal assumptions. Targeting such structures we propose a technique that predicts each sample using the context values and samples that are jointly transform coded with the predicted sample. We show that this joint, non-causal encoding can be represented with a nonorthogonal transform whose form and parameters we derive. We augment the HEVC standard with our work and show significant compression improvements on images/video that contain directional structures. Onur G. Guleryuz, Amir Said, Sehoon Yea |
ICIP | 2 |
| 2014 | Improving hybrid coding via control of quantization errors in the spatial and frequency domainsabstractWe propose new techniques to improve hybrid coding of images and video, without increasing decoder complexity, by controlling quantization effects in spatial and frequency domains, and finding optimal quantized coefficients via optimization. Additional compression is achieved by finding sets of optimal parameters, determined using training techniques. Experimental results confirm that the new method is able to provide better compression, both in objective measures, like signal-to-noise ratios, and also subjective, by reducing visually annoying blocking and banding artifacts. Amir Said, Onur G. Guleryuz, Sehoon Yea |
ICIP | 1 |
| 2013 | Quality perception in 3D interactive environmentsabstractIn this paper we investigate how configuration and visualization parameters influence the quality of experience in a 3D interactive environment, more specifically in a motion parallax setup. In order to do so, we designed a dedicated testing room and conducted subjective experiments with a team of evaluators. The tests considered parameters such as different disparities, amount of parallax, monitor sizes and visualization angles. Factors such as visual comfort, sense of immersion and the 3D experience as a whole have been assessed. The users were also asked to execute a task in the 3D motion parallax environment assessing the difficulty to complete the task for different parameter configurations. Obtained results suggest that user experience in an immersive environment is not as critically influenced by configuration parameters such as disparity and amount of parallax as initially thought. They also indicate that a better understanding of how this experience is influenced still requires the design and conduction of more comprehensive testing procedures. Alexandre G. Ciancio, José F. L. de Oliveira, Felipe M. Lopes Ribeiro, Eduardo A. B. da Silva, Amir Said |
ISCAS | 5 |
| 2012 | Integration of eye detection and tracking in videoconference sequences using temporal consistency and geometrical constraintsabstractIn this work, we propose a novel approach to detect and track, in videoconference sequences, six landmarks on eyes: the four corners and the pupils. Detection is based on the Inner Product Detector (IPD), and tracking on the Lucas-Kanade (LK) technique. The novelty of our method consists in the integration between detection and tracking, the evaluation of the temporal consistency to decrease the false positive rates, and the use of geometrical constraints to infer the position of missing points. In our experiments, we use five high definition video sequences with four subjects, different types of background, fast movements, blurring and occlusion. The obtained results have shown that the proposed technique is capable of detecting and tracking landmarks with good reliability. Gabriel M. Araujo, Eduardo A. B. da Silva, Alexandre G. Ciancio, José F. L. de Oliveira, Felipe M. Lopes Ribeiro, Amir Said |
ICIP | 6 |
| 2012 | Impact of encoding configurations on the perceived quality of high definition videoconference sequencesabstractIn this paper we evaluate the impact of different encoding configurations such as compression ratios, frame rates and resolution on the perceived quality of high definition video-conference applications. After generating a high quality video database, degraded sequences had their quality assessed by state-of-the-art automatic metrics. Results have shown that, for low rates, it is preferable to decrease the frame rate and resolution than to increase the compression ratio. Also, a higher frame rate tends to be more important than spatial resolution in terms of perceived quality. Alexandre G. Ciancio, José F. L. de Oliveira, Cassius D. Estrada, Eduardo A. B. da Silva, Amir Said |
ISCAS | 5 |
| 2012 | On the quality-assessment of reverberated speech
Amaro A. de Lima, Thiago de M. Prego, Sergio L. Netto, Bowon Lee, Amir Said, Ronald W. Schafer, Ton Kalker, Majid Fozunbal |
Speech Commun. | 5 |
| 2012 | A Parametric Objective Quality Assessment Tool for Speech Signals Degraded by Acoustic EchoabstractThis paper discusses the automatic quality assessment of echo-degraded speech in the context of teleconference systems. Subjective listening tests conducted over a carefully designed database of signals degraded by acoustic echo have been used to assess how this impairment is perceived and to determine which parameters have a significant impact on speech quality. The results have shown that, similarly to electric transmission line echo, acoustic echo is mainly influenced by echo delay and echo gain. Based on this observation, a mapping between these two parameters and the mean subjective score is devised. Moreover, a signal-based algorithm for the estimation of these parameters is described, and its performance is evaluated. The complete system comprising both the parameter estimators and the mapping function achieves a correlation of 94% between predicted and actual subjective scores, and can be employed as a non-intrusive monitoring tool for in-service quality evaluation of teleconference systems. Further validation indicates the operating range of the proposed quality assessment tool can be extended by proper retraining. Leonardo O. Nunes, Flávio R. Avila, Alan Freihof Tygel, Luiz W. P. Biscainho, Bowon Lee, Amir Said, Ronald W. Schafer |
IEEE Trans. Speech Audio Process. | 6 |
| 2011 | Intrinsic geometric distortions in a type of multi-projector light field displayabstractWe consider a class of projector-based light field 3D displays (horizontal parallax only) that use reflective or refractive anisotropic diffusers, and analyze the implications that at slated projection the light spreading is not in an ideally vertical plane, but instead is shaped like a cone. We demonstrate how this causes geometric distortions, some of which can be corrected by proper choice of projected images, while others, which we call intrinsic distortions, cannot. We show how the amount of distortion can be quantified by angular deviations, which are independent of the reproduced 3D scene, and present some of its properties. Ray-tracing simulations of these displays are shown to demonstrate the important results. Amir Said |
ICIP | 1 |
| 2011 | Virtual object distortions in 3D displays with only horizontal parallaxabstractAn imaging system that can convey a full 3D light field provides viewers with parallax in both the horizontal and vertical directions. However, displays that provide horizontal parallax only (HPO) are commonly preferred because they can greatly reduce cost and complexity, while still providing a good amount of 3D realism. Here we consider one side effect of HPO: the aspect ratio of displayed virtual objects will be observed to change as a viewer moves toward and away from the display. This results in object shape deformations that can only be corrected for viewing positions on a single line. We present an analysis of what happens outside these ideal positions, and how it is related to the range of object depths in the displayed scene. Several numerical examples and simulations of display views are presented to illustrate the effects, and show how they limit the acceptable range of scene depths and viewing positions of HPO 3D displays. Amir Said, W. Bruce Culbertson |
ICME | 1 |
| 2011 | Degradation Type Classifier for Full Band Speech Contaminated With Echo, Broadband Noise, and ReverberationabstractThis paper addresses the problem of identifying impairment types that might be present in a speech signal. In particular, three acoustically induced degradation types that occur in teleconference systems are considered: acoustic echo, reverberation, and broadband noise, as well as combinations among them. The proposed system is double-ended (full reference) and is developed using a database of degraded full-band speech signals created according to a model for teleconference systems. A set of features obtained from both the degraded and non-degraded signals is proposed and shown to adequately capture information associated with each degradation type. A random forest classifier and a support vector machine are successfully employed, achieving a classification error below 2%. Such classifiers can be used to select an appropriate quality assessment tool for a given degraded signal. Leonardo O. Nunes, Luiz W. P. Biscainho, Bowon Lee, Amir Said, Ton Kalker, Ronald W. Schafer |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2011 | No-Reference Blur Assessment of Digital Pictures Based on Multifeature ClassifiersabstractIn this paper, we address the problem of no-reference quality assessment for digital pictures corrupted with blur. We start with the generation of a large real image database containing pictures taken by human users in a variety of situations, and the conduction of subjective tests to generate the ground truth associated to those images. Based upon this ground truth, we select a number of high quality pictures and artificially degrade them with different intensities of simulated blur (gaussian and linear motion), totalling 6000 simulated blur images. We extensively evaluate the performance of state-of-the-art strategies for no-reference blur quantification in different blurring scenarios, and propose a paradigm for blur evaluation in which an effective method is pursued by combining several metrics and low-level image features. We test this paradigm by designing a no-reference quality assessment algorithm for blurred images which combines different metrics in a classifier based upon a neural network structure. Experimental results show that this leads to an improved performance that better reflects the images' ground truth. Finally, based upon the real image database, we show that the proposed method also outperforms other algorithms and metrics in realistic blur scenarios. Alexandre G. Ciancio, André Luiz N. Targino da Costa, Eduardo A. B. da Silva, Amir Said, Ramin Samadani, Pere Obrador |
IEEE Trans. Image Process. | 4 |
| 2010 | Massively parallel processing of signals in dense microphone arraysabstractArrays with large number of microphones can be very effective on audio processing tasks, like denoising, acoustic echo removal, etc. New microphone technologies enable creating large arrays with very low cost per component, but the system can still be very expensive due to costs of transmitting all signals to a single processor, and the computational resources to process the large amount of data. We show how a massively-parallel signal processing approach can solve the cost issues, when applied to the problem of sound source localization. We consider the case where each microphone is coupled to simple processing circuitry, which have full-bandwidth access to data from a few other microphones, while only shared power and low-bandwidth connections are provided between each microphone and a central processor. We discuss implementation issues, and show experimental results obtained in simulations and in microphone array measurements. Amir Said, Ton Kalker, Bowon Lee, Majid Fozunbal |
ISCAS | 1 |
| 2009 | Object tracking using multiple fragmentsabstractThis paper presents a low-cost tracking algorithm based on multiple multiple fragments, increasing robustness with respect to partial occlusions. Given the initial template representing the desired target, each pixel is classified into a different cluster based on a Mixture of Gaussians (MOG) model, and a set of disjoint fragments is created. The mean vector and covariance matrix of each fragment are computed, and the Mahalanobis distance is used to decide which pixels of the adjacent frame within a neighborhood are associated with each fragment. The template is then placed at the position that maximizes a similarity measure based on the number of matched points. Cláudio R. Jung, Amir Said |
ICIP | 2 |
| 2009 | Design of high capacity 3D print codes aiming for robustness to the PS channel and external distortionsabstractThe process of adding high-density information onto printed material enables and improves interesting hardcopy document applications, such as: security, authentication, physical-electronic round tripping, item-level tagging as well as consumer/product interaction. This investigation on robust and high capacity print codes aims to maximize information payload in a given printed page area, subject to robustness to distortions originated by printing and scanning processes and also to degradations introduced by user manipulation of printed documents. The novel approach includes statistical print-and-scan channel characterization, designing of robust segmentation, unsupervised Bayesian color classification with expectation-maximization algorithm for parameters estimation of a mixture of Gaussians model and design of error correction codes. Results illustrate the performance evaluated under real channel and distortions conditions. High payload is achieved with sufficient robustness to distortions resulting of regular office hardcopy document handling: print-and-scan channel and user manipulation. Joceli Mayer, José Carlos M. Bermudez, Andrei Piccinini Legg, Bartolomeu F. Uchôa Filho, Debargha Mukherjee, Amir Said, Ramin Samadani, Steven J. Simske |
ICIP | 6 |
| 2009 | Feature-Based Face Tracking for Videoconferencing ApplicationsabstractThis paper proposes a new approach for face tracking based on the individual tracking of KLT features. The face is initially detected using a face detection scheme, and KLT features are distributed along the face. Each feature is tracked individually, and the displacement of the center of the face is obtained using a Weighted Vector Median Filter (WVMF) of the individual displacements. The scale change is then computed based on the position of each feature w.r.t. the center of the face. The experimental results indicate that the proposed approach is fast and robust in the presence of partial occlusions. José Bins, Cláudio R. Jung, Leandro Dihl, Amir Said |
ISM | 4 |
| 2009 | A Template-Matching Based Method to Perform Iris Detection in Real-Time Using Synthetic TemplatesabstractUnderstanding people attentional focus can be useful for several applications. One important challenge in this area is to determine the iris position in image/video in order to estimate gaze behavior. To this end, this paper presents a robust and non-intrusive method to locate human iris position in low resolution grayscale images, in real-time. The method requires the previous knowledge of the face location and provides iris position in a region of interest (ROI) estimated based on anthropometric parameters. We achieved good results, varying from 90% to 100% of accuracy on images and video sequences. In addition, we tested our algorithm with noisy images. The lowest levels of accuracy are mainly due to light reflection on spectacles and eyes occlusion by eyelids. Júlio C. S. Jacques Júnior, Juliano Lucas Moreira, Adriana Braun, Soraia Raupp Musse, Amir Said |
ISM | 5 |
| 2009 | Video Based VAD Using Adaptive Color InformationabstractThis paper presents a new approach for voice activity detection (VAD) in videoconferencing applications based only on video information. In the proposed approach, a face detection scheme is applied to locate the face of the participant, and anthropometric measures are used to identify one patch below each eye, as well as a neighborhood around the mouth. The patches below the eyes are used to train a skin color model suited for that specific person, which is then used to identify non-skin pixels in the mouth neighborhood. The number of non-skin pixels is taken as an estimate of the mouth openness, and its evolution across time is explored for VAD. Dario Scott, Cláudio R. Jung, José Bins, Amir Said, Ton Kalker |
ISM | 4 |
| 2009 | An objective method for quality assessment of ultra-wideband speech corrupted by echoabstractModern telepresence systems can deliver multimedia signals of unprecedentedly high quality of experience to the user. Setting and maintaining such services call for reliable and automatic tools for multimedia quality probing, in special those targeted at speech data along the transmission path. Most of the objective methods for sound quality assessment (QA) in the literature are intended for either speech signals of 4- to 8-kHz bandwidth or general audio until 24 kHz, but are not specifically designed for speech at high sampling-rates. This work approaches quality evaluation of full-band (24 kHz) high-quality speech corrupted by echo. A simple metric singled out from a standardized double-ended tool for audio QA is proposed as a solution for the problem at hand. Quality measures from a set of speech stimuli corrupted by echo under controlled conditions were obtained via listening tests to allow calibration and evaluation of the proposed method. Experimental results reveal an overall correlation of 0.94 between objective and subjective scores, even in the presence of moderate additive noise. Luiz W. P. Biscainho, Paulo Antonio Andrade Esquef, Fabio P. Freeland, Leonardo O. Nunes, Alan Freihof Tygel, Bowon Lee, Amir Said, Ton Kalker, Ronald W. Schafer |
MMSP | 7 |
| 2009 | Quality assessment of audio: Increasing applicability scope of objective methods via prior identification of impairment typesabstractIn this paper the design of a double-ended (intrusive) diagnostic tool for identifying five types of degradation in audio signals is reported. The impairment types taken into consideration are additive contamination with pink noise, occurrences of signal mutes, distortion by magnitude clipping, and the previous two types mixed with pink noise. As a simple solution to accomplish the established goal, a threshold-based hierarchical classification system is proposed, being completely defined from pre-processing of the input signals, passing through the estimation of a few characteristic features, up to data clustering criteria. Performance evaluation of the classifier is carried out via a validation database containing 60 impaired signals for each type of impairment, with five distinct degradation intensity levels. Considering the types and range of degradation levels considered in this work, excellent results are achieved, scoring above 96% of correctly classified data in the worst case. System performance in identifying mixed impairment types tends to deteriorate as the strength of the noise component increases. Paulo Antonio Andrade Esquef, Luiz W. P. Biscainho, Leonardo O. Nunes, Bowon Lee, Amir Said, Ton Kalker, Ronald W. Schafer |
MMSP | 5 |
| 2009 | Feature analysis for quality assessment of reverberated speechabstractThis paper analyzes the ability of several measurements to quantify the reverberation effect in speech signals. We consider an intrusive scheme, in which the clean and reverberated signals are available, allowing one to estimate the corresponding room impulse response (RIR) signal. An artificial neural network (ANN) is trained for all features and used in a regression approach to estimate the human perceptual evaluation in a mean opinion score (MOS) 1–5 scale. Dimensionality reduction approaches are applied to generate a simpler ANN regression, establishing the most representative features for the problem at hand. A correlation level of 85% with subjective test scores was achieved by reducing the input-vector dimension from 10 to 3, including only the features of reverberation time, room spectral variance, and direct-to-reverberant energy ratio. Amaro A. de Lima, Thiago de M. Prego, Sergio L. Netto, Bowon Lee, Amir Said, Ronald W. Schafer, Ton Kalker, Majid Fozunbal |
MMSP | 5 |
| 2009 | Design of high capacity 3D print codes with visual cues aiming for robustness to the PS channel and external distortionsabstractAdding high-density information to printed materials enables and improves interesting hardcopy document applications involving security, authentication, physical-electronic round tripping, item-level tagging, and consumer/product interaction. This investigation of robust and high capacity print codes aims to maximize information payload in a given printed page area, subject to robustness to channel errors including distortions introduced by the printing and scanning processes and also due to the usual degradations introduced by user manipulation of printed documents. The novel approach includes statistical print-and-scan channel characterization, designing of robust segmentation using visual cues, unsupervised Bayesian color classification with expectation-maximization algorithm for parameters estimation of a mixture of Gaussians model and design of error correction codes. Results illustrate the performance evaluated under real channel and distortions conditions. High payload is achieved with sufficient robustness to distortions resulting of regular office hardcopy document handling: print-and-scan channel and user manipulation. Joceli Mayer, José Carlos M. Bermudez, Andrei Piccinini Legg, Bartolomeu F. Uchôa Filho, Debargha Mukherjee, Amir Said, Ramin Samadani, Steven J. Simske |
MMSP | 6 |
| 2008 | Exploiting patterns of data magnitude for efficient image codingabstractIn order to increase the computational efficiency of compression methods we have to consider that new hardware architectures increasingly rely more on wider data paths and parallel processing (e.g., SIMD and multi-core), than on faster clocks. Higher data throughputs are achieved with entropy coding methods that process larger amounts of information each time, and use context dependencies that are less complicated and that can be quickly updated. We propose a coding method with properties more suited to the new processors, that achieves better compression by exploiting patterns of data magnitude. We present experimental results on image coding implementations that take advantage of the fast decay of transform coefficient variance with frequency. Amir Said, Debargha Mukherjee |
ICIP | 1 |
| 2008 | On the quality assessment of sound signalsabstractThis paper constitutes an introduction to the field of quality evaluation of sound (speech and audio) signals. The need for such an assessment is inherent to modern communications: VoIP, mobile phone, or teleconference systems require meaningful measures of performance, which may ultimately assure good service or profitable business. A brief survey on subjective and objective evaluation methods is provided. Recent developments as well as new topics to be investigated are also addressed. Experiments are conducted to illustrate how to validate quality assessment methods. Amaro A. de Lima, Fabio P. Freeland, Rafael A. de Jesus, Bruno C. Bispo, Luiz W. P. Biscainho, Sergio L. Netto, Amir Said, Ton Kalker, Ronald W. Schafer, Bowon Lee, Mehrban Jam |
ISCAS | 7 |
| 2007 | A New Class of Filters for Image Interpolation and ResizingabstractWe propose a new set of kernels to simplify the design of filters for image interpolation and resizing. Their properties are defined according to two parameters, specifying the width of the transition band and the height of the first sidelobe. By varying these parameters we can get very good approximations of many commonly-used interpolation kernels. Furthermore, because the Fourier transforms of these kernels have very fast decay, they can also be used for downsampling. Amir Said |
ICIP (4) | 1 |
| 2007 | Phase-Domain Statistical Analysis for Audio Source LocalizationabstractWe consider the problem of estimating the time-difference-of-arrival (TDOA) for audio source localization in noisy environments, defining a framework for statistical analysis in the phase domain which enables more reliable estimates. This is motivated by the fact that with complex sources, noise, and interfering signals, different frequency bands have significantly different signal-to-noise ratios, creating non-uniform distributions of errors in the phase measurements. Through a new method for analysis of the variations of the phase in frequency windows, we first estimate the signal-to-noise ratio for frequency, and then use it in a maximum-likelihood estimation of the time difference of arrival. We show that this corresponds to a generalization of the Phase Transform method (PHAT), and provides a theoretical justification of why it works so well. Numerical results show how the proposed technique compares favorably with PHAT. Amir Said, Ton Kalker, Ronald W. Schafer |
MMSP | 1 |
| 2006 | Analysis of Subframe Generation for Superimposed ImagesabstractDisplays and projectors can increase their addressable resolution without changing the expensive spatial light modulators by using a technique called "wobulation," which consists of using mechanical actuators for rapidly shifting projected images (subframes) by fractions of pixel lengths. In this work we provide an analysis of the properties of the resulting super-imposed images, and discuss theoretical issues related to the stability of subframe generation and existence of solutions. Amir Said |
ICIP | 1 |
| 2006 | Worst-case Analysis of the Low-complexity Symbol Grouping Coding TechniqueabstractThe symbol grouping technique is widely used in practice because it allows great reductions on the complexity of entropy coding symbols from large alphabets, at the expense of small losses in compression. While it has been used mostly in an ad hoc manner, it is not known how general this technique is, i.e., in exactly what type of data sources it can be effective. We try to answer this question by searching for worst-case data sources, measuring the performance, and trying to identify trends. We show that finding the worst-case source is a very challenging optimization problem, and propose some solution methods that can be used in alphabets of moderate size. The numerical results provide evidence confirming the hypotheses that all data sources with large number of symbols can be more efficiently coded, with very small loss, using symbol grouping Amir Said |
ISIT | 1 |
| 2005 | Coding the Wavelet Spatial Orientation Tree with Low Computational ComplexityabstractSummary form only given. A very fast, low complexity algorithm for resolution-scalable and random access decoding is presented. The algorithm avoids the multiple passes of bit-plane coding for speed improvement. The decrease in dynamic range of wavelet coefficient magnitudes is efficiently coded. The hierarchical dynamic range coding naturally enables resolution-scalable representation of a wavelet transformed image. The method predicts the dynamic range of energy in each subset based on the dynamic range of energy of a parent set. Speed improvement over SPIHT is up to two times in encoding, and up to four times in decoding. The loss of quality is very small. Our method outperforms the LTW in Oliver et al. (2003), by up to two times in encoding and up to seven times in decoding. Yushin Cho, Amir Said, William A. Pearlman |
DCC | 2 |
| 2005 | Efficient Alphabet Partitioning Algorithms for Low-Complexity Entropy CodingabstractWe analyze the technique for reducing the complexity of entropy coding consisting of the a priori grouping of the source alphabet symbols, and in dividing the coding process in two stages: first coding the number of the symbol's group with a more complex method, followed by coding the symbol's rank inside its group using a less complex method, or simply using its binary representation. Because this method proved to be quite effective it is widely used in practice, and is an important part in standards like MPEG and JPEG. However, a theory to fully exploit its effectiveness had not been sufficiently developed. In this work, we study methods for optimizing the alphabet decomposition, and prove that a necessary optimality condition eliminates most of the possible solutions, and guarantees that dynamic programming solutions are optimal. In addition, we show that the data used for optimization have useful mathematical properties, which greatly reduce the complexity of finding optimal partitions. Finally, we extend the analysis, and propose efficient algorithms, for finding min-max optimal partitions for multiple data sources. Numerical results show the difference in redundancy for single and multiple sources. Amir Said |
DCC | 1 |
| 2005 | Format independent encryption of generalized scalable bit-streams enabling arbitrary secure adaptations [multimedia communication applications]abstractSecure format-independent adaptation of bit-streams during delivery is becoming increasingly important to cope with content piracy while still accommodating diverse networks, terminals and formats. Generalized scalable bit-streams are particularly advantageous in this regard, since they enable a variety of efficient and secure adaptations. Further, by associating such a bit-stream with appropriate metadata, such as those standardized in MPEG-21 Part 7, entitled digital item adaptation (DIA), the adaptation process can be fully format-independent. In this paper, to maximally secure a generalized scalable bit-stream while allowing arbitrary encrypted domain adaptations, strong progressive encryption methods are extended to multiple dimensions. Further it is shown that by appropriate modeling of such bit-streams and re-use of some DIA descriptions, the encryption and decryption engines themselves can be entirely metadata-driven and format-independent. This leads to end-to-end format-independent secure and adaptive delivery architectures for scalable bit-streams. Debargha Mukherjee, Huisheng Wang, Amir Said, Sam Liu |
ICASSP (2) | 3 |
| 2005 | Low complexity resolution progressive image coding algorithm: progres (progressive resolution decompression)abstractA very fast, low complexity algorithm for resolution scalable and random access decoding is presented. The algorithm avoids the multiple passes of bit-plane coding for speed improvement. The decrease in dynamic ranges of wavelet coefficients magnitudes is efficiently coded. The hierarchical dynamic range coding naturally enables a resolution scalable representation of a wavelet transformed image. Yushin Cho, William A. Pearlman, Amir Said |
ICIP (3) | 3 |
| 2005 | Measuring the strength of partial encryption schemesabstractPartial encryption (PE) of compressed multimedia can greatly reduce the computational complexity by encrypting only a fraction of the data bits. It can also easily provide users with low-quality versions, while maintaining the high-quality version inaccessible to unauthorized users. However, it is necessary to realistically evaluate its security strength. Some of the cryptanalysis done for these techniques ignored important characteristics of the multimedia files, and used overly optimistic assumptions. We demonstrate potential weaknesses of such techniques studying attacks that exploit the information provided by non-encrypted bits, and the availability of side information (e.g., from analog signals). We show that a more useful measure of encryption strength is the complexity to reduce distortion, instead of recovering the encryption key. We consider attacks on PE that avoid error propagation (standard-compliant PE), and PE that try to exploit that property for security. In both cases we show that attacks that require complexity much lower than exhaustive enumeration of encrypted/key bits can successfully yield good quality content. Experimental results are shown for images, but the conclusions can be extended to partial encryption of video and other types of media. Amir Said |
ICIP (2) | 1 |
| 2005 | A framework for fully format-independent adaptation of scalable bit streamsabstractThis paper presents a framework for modeling and adaptation of arbitrary scalable multimedia bit streams in a manner that is fully format agnostic ("universal transcoding"). It is entitled structured scalable metaformats (SSM), and is based on establishing a universal model for all scalable bit streams, which in turn allows a compact specification of the set of allowed adaptations, and also how the bit stream adaptation is performed. In order to enable format-agnostic adaptation, two elements are standardized: 1) metadata associated with a scalable bit stream conveying the parameters of its SSM model, and how it is to be manipulated to obtain various adapted versions, as well as information that allows making appropriate adaptation decisions using only the model parameters and 2) a specification of outbound network and terminal constraints, for real-time decisions based on both user preferences and network conditions. By interpreting the descriptions, a universal adaptation engine can adapt the content to suit the specified needs and preferences of recipients, without knowledge of the specifics of the content, its encoding and encryption. With universal adaptation, different adaptation infrastructures are no longer needed for different types of scalable media, eliminating the high costs of traditional transcoding. Many of the SSM concepts have been adopted into the MPEG-21 Part 7 standard entitled Digital Item Adaptation. Debargha Mukherjee, Amir Said, Sam Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Comparative Analysis of Arithmetic Coding Computational ComplexityabstractNew hardware is making obsolete previous assumptions about the most demanding computations for arithmetic coding. The important arithmetic coding tasks include: (a) interval update and arithmetic operations; (b) carry propagation and bit moves; (c) interval renormalization; (d) interval search for decoding; (e) probability estimation (source modeling); (f) support for nonbinary symbol alphabets. In this paper, the performance of binary arithmetic coding implemented with 16, 32 bit integer and 48 bit floating-point arithmetic is benchmarked. Results show that significantly faster adaptation, with small coding loss, can be obtained by periodic updates of the probability estimation. Amir Said |
Data Compression Conference | 1 |
| 2004 | Coding for Fast Access to Image Regions Defined by Pixel RangeabstractMany technical imaging applications, like coding "images" of digital elevation maps, require extracting regions of compressed images in which the pixel values are within a predefined range, and there is a need for coding methods that allow finding these regions efficiently, without having to decompress the whole image. A series of techniques to solve this problem is presented. First, it shows that many of the linear transforms commonly used for image compression can be used for that purpose by proving that the inclusion of nonlinear factors (like minimum or maximum pixel value in a block) does not render the transformation irreversible, and can be made to have very limited impact on the compression efficiency. For example, it shows how the "DC" coefficient of an 8/spl times/8 discrete cosine transform (DCT) can be replaced by the minimum or maximum in the 8/spl times/8 block. This result is valid for a large set of transforms, including the DCT, Walsh-Hadamard, and dyadic Haar transforms, and valid for any type of order-statistic filter output. Next, it shows the results also apply to the quantized transform coefficient cases as well as integer-to-integer transforms. The choices for coding the minimum and maximum values simultaneously, while providing quick access to pixel range and efficient compression were finally studied. Amir Said, Sehoon Yea, William A. Pearlman |
Data Compression Conference | 1 |
| 2004 | Format-independent scalable bit-stream adaptation using mpeg-21 dia
Debargha Mukherjee, Geraldine Kuo, Shih-Ta Hsiang, Sam Liu, Amir Said |
ICIP | 5 |
| 2004 | Efficient and reliable dynamic quality control for compression of compound document imagesabstractCompound images contain a mixture of natural images, text, and graphics. They need special care in the use of compression because text and graphics cannot withstand the significant distortion that is acceptable for natural images. One solution is to identify the areas that need to have higher settings in a dynamic quality control scheme, as supported by JPEG-SPIFF and JPEG 2000. We propose a scheme to identify those regions using a discrimination function that works with a non-linear transform to reliably identify edges, and at the same time avoid false positive detection on regions with complex patterns. It does so by exploiting the properties of histograms of coefficients of this block transform, and their entropy function, which we show can be computed efficiently via table look-up. Experimental results demonstrate the performance of the new scheme, compared to other methods in the literature. Amir Said |
ICIP | 1 |
| 2004 | Deringing and deblocking dct compression artifacts with efficient shifted transforms
Ramin Samadani, Arvind Sundararajan, Amir Said |
ICIP | 3 |
| 2004 | Efficient image coding for access to pixel rangesabstractTechnical imaging applications such as coding "images" of digital elevation maps, require extracting regions of compressed images with the pixel values within a pre-defined range, without having to decompress the whole image. Previously, we introduced a class of nonlinear transforms which are small modifications of linear transforms to facilitate a search for the regions with pixel values below (above) a given 'threshold', without incurring any penalty in coding efficiency. However, coding efficiency had to be somewhat compromised when searching for regions with a given pixel 'range', especially at high coding rates. In this paper, we propose an improved method of pixel 'range' coding to deal with the aforementioned problem. Results show significant improvements in coding efficiency. Sehoon Yea, Amir Said, William A. Pearlman |
ICIP | 2 |
| 2004 | Automatically designed 3D environments for intuitive explorationabstractWhile popular as a visual interface, current interactive 3D environments are often painstakingly created by hand. The paper proposes interactive 3D environments that are automatically designed on-the-fly based on various design rules, thereby separating layout and content. The resulting environments are compelling, customizable, and intuitive for information exploration. The interface also serves as an integrated framework for different types of rich media. The complete system for designing, constructing, and rendering the environments in real lime is discussed and shown to have implications in a variety of applications, including e-commerce. Nelson L. Chang, Amir Said |
ICME | 2 |
| 2004 | Compression of compound images and video for enabling rich media in embedded systemsabstractIt is possible to improve the features supported by devices with embedded systems by increasing the processor computing power, but this always results in higher costs, complexity, and power consumption. An interesting alternative is to use the growing networking infrastructures to do remote processing and visualization, with the embedded system mainly responsible for communications and user interaction. This enables devices to behave as if much more “intelligent” to users, at very low costs and power. In this article we explain how compression can make some of these solutions more bandwidth-efficient, enabling devices to simply decompress very rich graphical information and user interfaces that had been rendered elsewhere. The mixture of natural images and video with text, graphics, and animations simultaneously in the same frame is called compound video. We present a new method for compression of compound images and video, which is able to efficiently identify the different components during compression, and use an appropriate coding method. Our system uses lossless compression for graphics and text, and, on natural images and highly detailed parts, it uses lossy compression with dynamically varying quality. Since it was designed for embedded systems with very limited resources, and it has small executable size, and low complexity for classification, compression and decompression. Other compression methods (e.g., MPEG) can do the same, but are very inefficient for compound content. High-level graphics languages can be bandwidth-efficient, but are much less reliable (e.g., supporting Asian fonts), and are many orders of magnitude more complex. Numerical tests show the very significant gains in compression achieved by these systems. Amir Said |
VCIP | 1 |
| 2004 | Efficient, low-complexity image coding with a set-partitioning embedded block coderabstractWe propose an embedded, block-based, image wavelet transform coding algorithm of low complexity. It uses a recursive set-partitioning procedure to sort subsets of wavelet coefficients by maximum magnitude with respect to thresholds that are integer powers of two. It exploits two fundamental characteristics of an image transform-the well-defined hierarchical structure, and energy clustering in frequency and in space. The two partition strategies allow for versatile and efficient coding of several image transform structures, including dyadic, blocks inside subbands, wavelet packets, and discrete cosine transform (DCT). We describe the use of this coding algorithm in several implementations, including reversible (lossless) coding and its adaptation for color images, and show extensive comparisons with other state-of-the-art coders, such as set partitioning in hierarchical trees (SPIHT) and JPEG2000. We conclude that this algorithm, in addition to being very flexible, retains all the desirable features of these algorithms and is highly competitive to them in compression efficiency. William A. Pearlman, Asad Islam, Nithin Nagaraj, Amir Said |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2003 | Fully scalable video transmission using the SSM adaptation framework
Debargha Mukherjee, Peisong Chen, Shih-Ta Hsiang, John W. Woods, Amir Said |
VCIP | 5 |
| 2002 | Low complexity guaranteed fit compound document compressionabstractWe propose a new, very low complexity, single-pass algorithm for compression of continuous tone compound documents, known as GRAFIT (GuaRAnteed FIT), that can guarantee a minimum compression ratio of as much as 12:1 and even more, for all images in a single pass, while maintaining visually lossless quality when reproduced at resolution 300 dpi or more. The compression ratio is guaranteed in a single pass irrespective of the image being compressed. The complexity of the proposed encoder and decoder is orders of magnitude smaller than all image compression algorithms known today. For electronic compound documents, text is always compressed losslessly, and depending on the type of image, the actual compression ratio achieved may be as high as 200:1 or more. For photographic images, while the compression performance is inferior to DCT or wavelet coders, for documents at resolution 300 dpi and above, the quality is still visually lossless. Overall, performance of GRAFIT is highly competitive with the more expensive algorithms like JPEG2000, JPEG, or JPEG-LS and this performance is achieved in a single pass at much lower cost in both software and hardware. Debargha Mukherjee, Christos Chrysafis, Amir Said |
ICIP (1) | 3 |
| 2002 | JPEG2000-matched MRC compression of compound documentsabstractThe mixed raster content (MRC) ITU document compression standard (T.44) specifies a multilayer decomposition model for compound documents into two contone image layers and a binary mask layer for independent compression. While T.44 does not recommend any procedure for decomposition, it does specify a set of allowable layer codecs to be used after decomposition. While T.44 only allows older standardized codecs such as JPEG/JBIG/G3/G4, higher compression could be achieved if newer contone and bi-level compression standards such as JPEG2000/JBIG2 were used instead. We present an MRC compound document codec using JPEG2000 as the image layer codec and a layer decomposition scheme matched to JPEG2000 for efficient compression. JBIG still codes the mask. Noise removal routines enable efficient coding of scanned documents along with electronic ones. Resolution scalable decoding features are also implemented. The segmentation mask, obtained from layer decomposition, serves to separate text and other features. Debargha Mukherjee, Christos Chrysafis, Amir Said |
ICIP (3) | 3 |
| 2001 | JPEG-matched MRC compression of compound documentsabstractMixed raster content (MRC) is an ITU document compression standard (T.44) specifying both a model for multilayer representation of a compound document, and a set of allowable standardized coders for the individual layers. The model requires decomposition of a document into two image layers and a binary mask layer, but the standard does not recommend any procedure for this task. For best compression results, the decomposition method should be optimized for the layer encoders. In this paper, a high performance MRC compound document codec is presented, where the layer decomposition scheme is matched to the JPEG encoder with arithmetic coding for the foreground and background image layers. JBIG is used to code the mask layer. Integrated noise removal routines enable handling of scanned documents along with electronic ones. Resolution scalable decoding features are also implemented. The page segmenter yields a segmentation mask, which serves to separate text and other features. Debargha Mukherjee, Nasir Memon, Amir Said |
ICIP (3) | 3 |
| 2000 | SBHP-a low complexity wavelet coderabstractWe present a low-complexity entropy coder originally designed to work in the JPEG2000 image compression standard framework. The algorithm is meant for embedded and non-embedded coding of wavelet coefficients inside a subband, and is called subband-block hierarchical partitioning (SBHP). It was extensively tested following the standard experiment procedures, and it was shown to yield a significant reduction in the complexity of entropy coding, with small loss in compression performance. Furthermore, it is able to seamlessly support all JPEG2000 features. We present a description of the algorithm, an analysis of its complexity, and a summary of the results obtained after its integration into the verification model (VM). Christos Chrysafis, Amir Said, Alexander Drukarev, Asad Islam, William A. Pearlman |
ICASSP | 2 |
| 1999 | Simplified Segmentation for Compound Image CompressionabstractThere are three basic segmentation schemes for compound image compression: object-based, layer-based, and block-based. This paper discusses the relative advantages of each scheme and architecture, and studies the use of fast classification techniques for a segmentation that can be used together with a chosen compression architecture. Particularly, we consider classification techniques working on approximate object boundaries, which reaches the localization and precision of the segmentation, but in exchange allows faster, one-pass segmentation, low memory requirements, and segmentation map that is better matched to existing compression methods. We show numerical results obtained on a printer application environment, where rigorous standards of visual quality have to be satisfied. Amir Said, Alexander Drukarev |
ICIP (1) | 1 |
| 1998 | Bandwidth-Efficient Coded Modulation with Optimized Linear Partial-Response SignalsabstractWe study the design of optimal signals for bandwidth-efficient linear coded modulation. Previous results show that for linear channels with intersymbol interference (ISI), reduced-search decoding algorithms have near-maximum-likelihood error performance, but with much smaller complexity than the Viterbi decoder. Consequently, the controlled ISI introduced by a lowpass filter can be practically used for bandwidth reduction. Such spectrum shaping filters comprise an explicit coded modulation, for which we seek the optimal design. We simultaneously constrain the bandwidth and maximize the minimum Euclidean distance between signals. We show that under quite general assumptions the problem can be formulated as a linear program, and solved with well-known efficient optimization techniques. Numerical results are presented, and the performance of the optimal signals, measured by their combined bandwidth and noise immunity, is analyzed. The new codes are comparable to set-partition (TCM) trellis codes. Tests of an M-algorithm decoder confirm this and show that the performance occurs at small detection complexity. Amir Said, John B. Anderson |
IEEE Trans. Inf. Theory | 1 |
| 1996 | A new, fast, and efficient image codec based on set partitioning in hierarchical treesabstractEmbedded zerotree wavelet (EZW) coding, introduced by Shapiro (see IEEE Trans. Signal Processing, vol.41, no.12, p.3445, 1993), is a very effective and computationally simple technique for image compression. We offer an alternative explanation of the principles of its operation, so that the reasons for its excellent performance can be better understood. These principles are partial ordering by magnitude with a set partitioning sorting algorithm, ordered bit plane transmission, and exploitation of self-similarity across different scales of an image wavelet transform. Moreover, we present a new and different implementation based on set partitioning in hierarchical trees (SPIHT), which provides even better performance than our previously reported extension of EZW that surpassed the performance of the original EZW. The image coding results, calculated from actual file sizes and images reconstructed by the decoding algorithm, are either comparable to or surpass previous results obtained through much more sophisticated and computationally complex methods. In addition, the new coding and decoding procedures are extremely fast, and they can be made even faster, with only small loss in performance, by omitting entropy coding of the bit stream by the arithmetic code. Amir Said, William A. Pearlman |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1996 | An image multiresolution representation for lossless and lossy compressionabstractWe propose a new image multiresolution transform that is suited for both lossless (reversible) and lossy compression. The new transformation is similar to the subband decomposition, but can be computed with only integer addition and bit-shift operations. During its calculation, the number of bits required to represent the transformed image is kept small through careful scaling and truncations. Numerical results show that the entropy obtained with the new transform is smaller than that obtained with predictive coding of similar complexity. In addition, we propose entropy-coding methods that exploit the multiresolution structure, and can efficiently compress the transformed image for progressive transmission (up to exact recovery). The lossless compression ratios are among the best in the literature, and simultaneously the rate versus distortion performance is comparable to those of the most efficient lossy compression methods. Amir Said, William A. Pearlman |
IEEE Trans. Image Process. | 1 |
| 1996 | New ternary and quaternary linear block codesabstractWe present several new ternary and quaternary codes found with a heuristic search procedure, originally proposed by the authors to find good binary codes, and which is based on combinatorial optimization. Amir Said, Reginaldo Palazzo Júnior |
IEEE Trans. Inf. Theory | 1 |
| 1993 | Image compression using the spatial-orientation tree
Amir Said, William A. Pearlman |
ISCAS | 1 |
| 1993 | Reversible image compression via multiresolution representation and predictive codingabstractIn this paper a new image transformation suited for reversible (lossless) image compression is presented. It uses a simple pyramid multiresolution scheme which is enhanced via predictive coding. The new transformation is similar to the subband decomposition, but it uses only integer operations. The number of bits required to represent the transformed image is kept small through careful scaling and truncations. The lossless coding compression rates are smaller than those obtained with predictive coding of equivalent complexity. It is also shown that the new transform can be effectively used, with the same coding algorithm, for both lossless and lossy compression. When used for lossy compression, its rate-distortion function is comparable to other efficient lossy compression methods. Amir Said, William A. Pearlman |
VCIP | 1 |
| 1993 | Using combinatorial optimization to design good unit-memory convolutional codesabstractA method for designing good unit-memory convolutional codes is presented. The method is based on the decomposition of the original problem into two easier subproblems that can be formulated as optimization problems and solved by efficient heuristic search algorithms. The efficacy of this method is demonstrated by a table containing 33 new unit-memory convolutional codes (n,k) with 5> Amir Said, Reginaldo Palazzo Júnior |
IEEE Trans. Inf. Theory | 1 |