Yiannis Andreopoulos

dblp:85/3071 · DBLP profile ↗
← Back
91ranked-venue papers
19as first author
10since 2021 · last 2025
0000-0002-2714-4800ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 63 · 18 first-author · 6 since 2021Computer networks · 14 · 1 first-author · 2 since 2021Systems, architecture and hardware · 7Artificial intelligence and machine learning · 6 · 3 since 2021Software engineering, systems software and programming languages · 4
YearPublicationVenuePosition
2025 Perceptual Video Compression with Neural Wrapping
abstract
Standard video codecs are rate-distortion optimization machines, where distortion is typically quantified using PSNR versus the source. However, it is now widely accepted that increasing PSNR does not necessarily translate to better visual quality. In this paper, a better balance between perception and fidelity is targeted, in order to provide for significant rate savings over state-of-the-art standards-based video codecs. Specifically, pre- and post-processing neural networks are proposed that enhance the coding efficiency of standard video codecs when bench-marked with an array of well-established perceptual quality scores. These "neural wrapper" elements are end-to-end trained with a neural codec module serving as a differentiable proxy for standard video codecs. The codec proxy is jointly optimized with the pre- and post components via a novel two-phase pretraining strategy and end-to-end iterative refinement with stop-gradient. This allows the neural pre- and postprocessor to learn to embed, remove and recover information in a codec-aware manner, thus improving its rate-quality performance. A single neural-wrapper model is thereby established and used for the entire rate-quality curve without needing any downscaling or upscaling. The trained model is tested with the AV1 and VVC standard codecs via an array of well-established objective quality scores (SSIM, MS-SSIM, VMAF, AVQT), as well as mean opinion scores (MOS) derived from ITU-T P.910 subjective testing. Experimental results show that the proposed approach improves all quality scores, with -18.5% average Bjontegaard Delta-rate (BD-rate) saving over all objective scores and MOS improvement over both standard codecs. This illustrates the significant potential of neural wrapper components over standards-based video coding.
Muhammad Umar Karim Khan, Aaron Chadha, Mohammad Ashraful Anam, Yiannis Andreopoulos
CVPR4
2025 Content Adaptive Encoding For Interactive Game Streaming
Shakarim Soltanayev, Odysseas Zisimopoulos, Mohammad Ashraful Anam, Man Cheung Kung, Yiannis Andreopoulos
PCS6
2022 PAC-Bayesian Bounds on Rate-Efficient Classifiers
abstract
We derive analytic bounds on the noise invariance of majority vote classifiers operating on compressed inputs. Specifically, starting from recent bounds on the true risk of majority vote classifiers, we extend the applicability of PAC-Bayesian theory to quantify the resilience of majority votes to input noise stemming from compression. The derived bounds are intuitive in binary classification settings, where they can be measured as expressions of voter differentials and voter pair agreement. By combining measures of input distortion with analytic guarantees on noise invariance, we prescribe rate-efficient machines to compress inputs without affecting subsequent classification. Our validation shows how bounding noise invariance can inform the compression stage for any majority vote classifier such that worst-case implications of bad input reconstructions are known, and inputs can be compressed to the minimum amount of information needed prior to inference.
Alhabib Abbas, Yiannis Andreopoulos
ICML2
2022 Advances in Quality Assessment Of Video Streaming Systems: Algorithms, Methods, Tools
abstract
Quality assessment of video has matured significantly in the last 10 years due to a flurry of relevant developments in academia and industry, with relevant initiatives in VQEG, AOMedia, MPEG, ITU-T P.910, and other standardization and advisory bodies . Most advanced video streaming systems are now clearly moving away from good old-fashioned' PSNR and structural similarity type of assessment towards metrics that align better to mean opinion scores from viewers. Several of these algorithms, methods and tools have only been developed in the last 3-5 years and, while they are of significant interest to the research community, their advantages and limitations are not widely known in the research community. This tutorial provides this overview, but also focuses on practical aspects and how to design quality assessment tests that can scale to large datasets.
Yiannis Andreopoulos, Cosmin Stejerean
ACM Multimedia1
2022 Domain-Specific Fusion Of Objective Video Quality Metrics
abstract
Video processing algorithms like video upscaling, denoising, and compression are now increasingly optimized for perceptual quality metrics instead of signal distortion. This means that they may score well for metrics like video multi-method assessment fusion (VMAF), but this may be because of metric overfitting. This imposes the need for costly subjective quality assessments that cannot scale to large datasets and large parameter explorations. We propose a methodology that fuses multiple quality metrics based on small scale subjective testing in order to unlock their use at scale for specific application domains of interest. This is achieved by employing pseudo-random sampling of the resolution, quality range and test video content available, which is initially guided by quality metrics in order to cover the quality range useful to each application. The selected samples then undergo a subjective test, such as ITU-T P.910 absolute categorical rating, with the results of the test postprocessed and used as the means to derive the best combination of multiple objective metrics using support vector regression. We showcase the benefits of this approach in two applications: video encoding with and without perceptual preprocessing, and deep video denoising & upscaling of compressed content. For both applications, the derived fusion of metrics allows for a more robust alignment to mean opinion scores than a perceptually-uninformed combination of the original metrics themselves. The dataset and code is available at https://github.com/isize-tech/VideoQualityFusion.
Aaron Chadha, Ioannis Katsavounidis, Ayan Kumar Bhunia, Cosmin Stejerean, Muhammad Umar Karim Khan, Yiannis Andreopoulos
ACM Multimedia6
2022 Learning-Based Symbol Level Precoding: A Memory-Efficient Unsupervised Learning Approach
abstract
Symbol level precoding (SLP) has been proven to be an effective means of managing the interference in a multiuser downlink transmission and also enhancing the received signal power. This paper proposes an unsupervised-learning based SLP that applies to quantized deep neural networks (DNNs). Rather than simply training a DNN in a supervised mode, our proposal unfolds a power minimization SLP formulation in an imperfect channel scenario using the interior point method (IPM) proximal ‘log’ barrier function. We use binary and ternary quantizations to compress the DNN’s weight values. The results show significant memory savings for our proposals compared to the existing full-precision SLP-DNet with significant model compression of ~ 21× and ~ 13× for both binary DNN-based SLP (RSLP-BDNet) and ternary DNN-based SLP (RSLP-TDNets), respectively.
Abdullahi Mohammad, Christos Masouros, Yiannis Andreopoulos
WCNC3
2022 Sequence-Level Reference Frames in Video Coding
abstract
The proliferation of low-cost DRAM chipsets now begins to allow for the consideration of substantially-increased decoded picture buffers in advanced video coding standards such as HEVC, VVC, and Google VP9. At the same time, the increasing demand for rapid scene changes and multiple scene repetitions in entertainment or broadcast content indicates that extending the frame referencing interval to tens of minutes or even the entire video sequence may offer coding gains, as long as one is able to identify frame similarity in a computationally- and memory-efficient manner. Motivated by these observations, we propose a “stitching” method that defines a reference buffer and a reference frame selection algorithm. Our proposal extends the referencing interval of inter-frame video coding to the entire length of video sequences. Our reference frame selection algorithm uses well-established feature descriptor methods that describe frame structural elements in a compact and semantically-rich manner. We propose to combine such compact descriptors with a similarity scoring mechanism in order to select the frames to be “stitched” to reference picture buffers of advanced inter-frame encoders like HEVC, VVC, and VP9 without breaking standard compliance. Our evaluation on synthetic and real-world video sequences with the HEVC and VVC reference encoders shows that our method offers significant rate gains, with complexity and memory requirements that remain manageable for practical encoders and decoders.
Mohammad K. Jubran, Alhabib Abbas, Yiannis Andreopoulos
IEEE Trans. Circuits Syst. Video Technol.3
2021 A Channel Selection Algorithm Using Reinforcement Learning for Mobile Devices in Massive IoT System
abstract
It is necessary to develop an efficient channel selection method with low power consumption to achieve high communication quality for distributed massive IoT system. To this end, Ma et al. [1] proposed an autonomous distributed channel selection method based on the Tug-of-War (ToW) dynamics. The ToW-based method can achieve equivalent performance to UCB1-tuned [2], [3] with low computational complexity and power consumption, which is recognized as a best practice technique for solving multi-armed bandit (MAB) problems. However, Ref. [1] only considered fixed IoT devices with simplex communication.
Honami Furukawa, Aohan Li, Yozo Shoji, Yoshito Watanabe, Song-Ju Kim, Koya Sato, Yiannis Andreopoulos, Mikio Hasegawa
CCNC7
2021 Deep Perceptual Preprocessing for Video Coding
abstract
We introduce the concept of rate-aware deep perceptual preprocessing (DPP) for video encoding. DPP makes a single pass over each input frame in order to enhance its visual quality when the video is to be compressed with any codec at any bitrate. The resulting bitstreams can be decoded and displayed at the client side without any post-processing component. DPP comprises a convolutional neural network that is trained via a composite set of loss functions that incorporates: (i) a perceptual loss based on a trained no-reference image quality assessment model, (ii) a reference-based fidelity loss expressing L1 and structural similarity aspects, (iii) a motion-based rate loss via block-based transform, quantization and entropy estimates that converts the essential components of standard hybrid video encoder designs into a trainable framework. Extensive testing using multiple quality metrics and AVC, AV1 and VVC encoders shows that DPP+encoder reduces, on average, the bitrate of the corresponding encoder by 11%. This marks the first time a server-side neural processing component achieves such savings over the state-of-the-art in video coding.
Aaron Chadha, Yiannis Andreopoulos
CVPR2
2021 An Unsupervised Learning-Based Approach for Symbol-Level-Precoding
abstract
This paper proposes an unsupervised learning-based precoding framework that trains deep neural networks (DNNs) with no target labels by unfolding an interior point method (IPM) proximal ‘$log$’ barrier function. The proximal ‘$log$’ barrier function is derived from the strict power minimization formulation subject to signal-to-interference-plus-noise ratio (SINR) constraint. The proposed scheme exploits the known interference via symbol-level pre-coding (SLP) to minimize the transmit power and is named strict Symbol-Level-Precoding deep network (SLP-SDNet). The results show that SLP-SDNet out-performs the conventional block-level-precoding (Conventional BLP) scheme while achieving near-optimal performance faster than the SLP optimization-based approach.
Abdullahi Mohammad, Christos Masouros, Yiannis Andreopoulos
GLOBECOM3
2020 Challenges and Perspectives in Neuromorphic-based Visual IoT Systems and Networks
abstract
Neuromorphic sensors, a.k.a. dynamic vision sensors (DVS) or silicon retinas, do not capture full images (frames) at a fixed rate, but asynchronously capture spikes indicating changes of brightness in the scene, following the principles of biological vision and perception in mammals. DVS sensing and processing produces a data representation where the scene can be represented with a very high time resolution with a limited number of bits (an inherent data compression is performed at the time of acquisition). Such representation can be used locally to derive actionable responses and selected parts can be transmitted and then processed in another network location. Due to these features, such sensors represent an excellent choice as visual sensing technology for next-generation Internet-of-Things, e.g. in surveillance, drone technology, and robotics. It is in fact becoming evident that in this framework acquiring, processing, and transmitting frame-based video is inefficient in terms of energy consumption and reaction times, in particular in some scenarios. Hence, we explore here the feasibility of advanced Machine to Machine (M2M) communications systems that directly capture, compress and transmit spike-based visual information to cloud computing services in order to produce content classification or retrieval results with extremely low power and low latency.
Maria G. Martini, Nabeel Khan, Yin Bi, Yiannis Andreopoulos, Hadi Saki, Mohammad Shikh-Bahaei
ICASSP4
2020 Accelerated Learning-Based MIMO Detection through Weighted Neural Network Design
abstract
In this paper, we introduce a framework for a systematic acceleration of deep neural network (DNN) design for MIMO detection. A monotonically non-increasing function is used to scale the values of the layer weights such that only a certain fraction of the inputs is used for feedforward computation. This enables a dynamic weight scaling across and within the network layers, and it is termed as weight-scaling neural network-based MIMO detector (WeSNet). To increase the robustness against the changes in the activation patterns and additional enhancement in the detection accuracy for the same inference complexity, we introduce trainable weight-scaling functions. Experimental results show the superiority of our proposed method over the benchmark model (DetNet) and classical approaches based on semi-definite relaxation in terms of detection accuracy and computational efficiency.
Abdullahi Mohammad, Christos Masouros, Yiannis Andreopoulos
ICC3
2020 Verifiable Event Record Management for a Store-Carry-Forward-Based Data Delivery Platform by Blockchain
abstract
We propose a novel database management framework for data delivery services based on Store-Carry-Forward (SCF) techniques. The platform we present consists of heterogeneous wireless opportunistic networks of long-range narrowband and short-range broadband communications. We introduce a blockchain-based method by which to verify the record of delivery events on a decentralized network. A new consensus mechanism named proof-of-forwarding (PoF) is proposed to substitute the function of previously proposed proof-of-work (PoW) methods, while significantly improving computational complexities of block generation. Specifically, in our proposal a block is generated exclusively when data delivery agents perform node-to-node direct communication using a short-range high-speed wireless standard to deliver data. We additionally propose a digital signature overlay to prevent malicious nodes from producing fake transactions without any effort to carry data content to recipients. Simulation results show that our blockchain-based framework robustly manages data delivery records, where 97% of Distributed Denial-of-Service (DDoS) attacks can be prevented even when half of the entire nodes are assumed to be malicious.
Yoshito Watanabe, Wei Liu 0029, Alhabib Abbas, Yiannis Andreopoulos, Mikio Hasegawa, Yozo Shoji
PIMRC4
2020 Complexity-Scalable Neural-Network-Based MIMO Detection With Learnable Weight Scaling
abstract
This paper introduces a framework for systematic complexity scaling of deep neural network (DNN) based MIMO detectors. The model uses a fraction of the DNN inputs by scaling their values through weights that follow monotonically non-increasing functions. This allows for weight scaling across and within the different DNN layers in order to achieve accuracy-vs.-complexity scalability during inference. In order to further improve the performance of our proposal, we introduce a sparsity-inducing regularization constraint in conjunction with trainable weight-scaling functions. In this way, the network learns to balance detection accuracy versus complexity while also increasing robustness to changes in the activation patterns, leading to further improvement in the detection accuracy and BER performance at the same inference complexity. Numerical results show that our approach is 10-fold and 100-fold less complex than classical approaches based on semi-definite relaxation and ML detection, respectively.
Abdullahi Mohammad, Christos Masouros, Yiannis Andreopoulos
IEEE Trans. Commun.3
2020 Deep Video Precoding
abstract
Several groups worldwide are currently investigating how deep learning may advance the state-of-the-art in image and video coding. An open question is how to make deep neural networks work in conjunction with existing (and upcoming) video codecs, such as MPEG H.264/AVC, H.265/HEVC, VVC, Google VP9 and AOMedia AV1, AV2, as well as existing container and transport formats, without imposing any changes at the client side. Such compatibility is a crucial aspect when it comes to practical deployment, especially when considering the fact that the video content industry and hardware manufacturers are expected to remain committed to supporting these standards for the foreseeable future. We propose to use deep neural networks as precoders for current and future video codecs and adaptive video streaming systems. In our current design, the core precoding component comprises a cascaded structure of downscaling neural networks that operates during video encoding, prior to transmission. This is coupled with a precoding mode selection algorithm for each independently-decodable stream segment, which adjusts the downscaling factor according to scene characteristics, the utilized encoder, and the desired bitrate and encoding configuration. Our framework is compatible with all current and future codec and transport standards, as our deep precoding network structure is trained in conjunction with linear upscaling filters (e.g., the bilinear filter), which are supported by all web video players. Extensive evaluation on FHD (1080p) and UHD (2160p) content and with widely-used H.264/AVC, H.265/HEVC and VP9 encoders, as well as a preliminary evaluation with the current test model of VVC (v.6.2rc1), shows that coupling such standards with the proposed deep video precoding allows for 8% to 52% rate reduction under encoding configurations and bitrates suitable for video-on-demand adaptive streaming systems. The use of precoding can also lead to encoding complexity reduction, which is essential for cost-effective cloud deployment of complex encoders like H.265/HEVC, VP9 and VVC, especially when considering the prominence of high-resolution adaptive video streaming.
Eirina Bourtsoulatze, Aaron Chadha, Ilya Fadeev, Vasileios Giotsas, Yiannis Andreopoulos
IEEE Trans. Circuits Syst. Video Technol.5
2020 Rate-Accuracy Trade-Off in Video Classification With Deep Convolutional Neural Networks
abstract
Advanced video classification systems decode video frames to derive texture and motion representations for ingestion and analysis by spatio-temporal deep convolutional neural networks (CNNs). However, when considering visual Internet-of-Things applications, surveillance systems, and semantic crawlers of large video repositories, the video capture and the CNN-based semantic analysis parts do not tend to be co-located. This necessitates the transport of compressed video over networks and incurs significant overhead in bandwidth and energy consumption, thereby significantly undermining the deployment potential of such systems. In this paper, we investigate the trade-off between the encoding bitrate and the achievable accuracy of CNN-based video classification models that directly ingest AVC/H.264 and HEVC encoded videos. Instead of retaining entire compressed video bitstreams and applying complex optical flow calculations prior to CNN processing, we only retain motion vector and select texture information at significantly reduced bitrates and apply no additional processing prior to CNN ingestion. Based on three CNN architectures and two action recognition datasets, we achieve 11%-94% savings in bitrate with marginal effect on classification accuracy. A model-based selection between multiple CNNs increases these savings further to the point where, if up to 7% loss of accuracy can be tolerated, video classification can take place with as little as 3 kb/s for the transport of the required compressed video information to the system implementing the CNN models.
Mohammad K. Jubran, Alhabib Abbas, Aaron Chadha, Yiannis Andreopoulos
IEEE Trans. Circuits Syst. Video Technol.4
2020 Biased Mixtures of Experts: Enabling Computer Vision Inference Under Data Transfer Limitations
abstract
We propose a novel mixture-of-experts class to optimize computer vision models in accordance with data transfer limitations at test time. Our approach postulates that the minimum acceptable amount of data allowing for highly-accurate results can vary for different input space partitions. Therefore, we consider mixtures where experts require different amounts of data, and train a sparse gating function to divide the input space for each expert. By appropriate hyperparameter selection, our approach is able to bias mixtures of experts towards selecting specific experts over others. In this way, we show that the data transfer optimization between visual sensing and processing can be solved as a convex optimization problem. To demonstrate the relation between data availability and performance, we evaluate biased mixtures on a range of mainstream computer vision problems, namely: (i) single shot detection, (ii) image super resolution, and (iii) realtime video action classification. For all cases, and when experts constitute modified baselines to meet different limits on allowed data utility, biased mixtures significantly outperform previous work optimized to meet the same constraints on available data.
Alhabib Abbas, Yiannis Andreopoulos
IEEE Trans. Image Process.2
2020 Graph-Based Spatio-Temporal Feature Learning for Neuromorphic Vision Sensing
abstract
Neuromorphic vision sensing (NVS) devices represent visual information as sequences of asynchronous discrete events (a.k.a., "spikes") in response to changes in scene reflectance. Unlike conventional active pixel sensing (APS), NVS allows for significantly higher event sampling rates at substantially increased energy efficiency and robustness to illumination changes. However, feature representation for NVS is far behind its APS-based counterparts, resulting in lower performance in high-level computer vision tasks. To fully utilize its sparse and asynchronous nature, we propose a compact graph representation for NVS, which allows for end-to-end learning with graph convolution neural networks. We couple this with a novel end-to-end feature learning framework that accommodates both appearancebased and motion-based tasks. The core of our framework comprises a spatial feature learning module, which utilizes residual-graph convolutional neural networks (RG-CNN), for end-to-end learning of appearance-based features directly from graphs. We extend this with our proposed Graph2Grid block and temporal feature learning module for efficiently modelling temporal dependencies over multiple graphs and a long temporal extent. We show how our framework can be configured for object classification, action recognition and action similarity labeling. Importantly, our approach preserves the spatial and temporal coherence of spike events, while requiring less computation and memory. The experimental validation shows that our proposed framework outperforms all recent methods on standard datasets. Finally, to address the absence of large real-world NVS datasets for complex recognition tasks, we introduce, evaluate and make available the American Sign Language letters (ASL-DVS), as well as human action dataset (UCF101-DVS, HMDB51-DVS and ASLAN-DVS).
Yin Bi, Aaron Chadha, Alhabib Abbas, Eirina Bourtsoulatze, Yiannis Andreopoulos
IEEE Trans. Image Process.5
2020 Improved Techniques for Adversarial Discriminative Domain Adaptation
abstract
Adversarial discriminative domain adaptation (ADDA) is an efficient framework for unsupervised domain adaptation in image classification, where the source and target domains are assumed to have the same classes, but no labels are available for the target domain. While ADDA has already achieved better training efficiency and competitive accuracy on image classification in comparison to other adversarial based methods, we investigate whether we can improve its performance with a new framework and new loss formulations. Following the framework of semi-supervised GANs, we first extend the discriminator output over the source classes, in order to model the joint distribution over domain and task. We thus leverage on the distribution over the source encoder posteriors (which is fixed during adversarial training) and propose maximum mean discrepancy (MMD) and reconstruction-based loss functions for aligning the target encoder distribution to the source domain. We compare and provide a comprehensive analysis of how our framework and loss formulations extend over simple multi-class extensions of ADDA and other discriminative variants of semi-supervised GANs. In addition, we introduce various forms of regularization for stabilizing training, including treating the discriminator as a denoising autoencoder and regularizing the target encoder with source examples to reduce overfitting under a contraction mapping (i.e., when the target per-class distributions are contracting during alignment with the source). Finally, we validate our framework on standard datasets like MNIST, USPS, SVHN, MNIST-M and Office-31. We additionally examine how the proposed framework benefits recognition problems based on sensing modalities that lack training data. This is realized by introducing and evaluating on a neuromorphic vision sensing (NVS) sign language recognition dataset, where the source domain constitutes emulated neuromorphic spike events converted from conventional pixel-based video and the target domain is experimental (real) spike events from an NVS camera. Our results on all datasets show that our proposal is both simple and efficient, as it competes or outperforms the state-of-the-art in unsupervised domain adaptation, such as DIFA and MCDDA, whilst offering lower complexity than other recent adversarial methods.
Aaron Chadha, Yiannis Andreopoulos
IEEE Trans. Image Process.2
2019 UAV Detection: A STDP Trained Deep Convolutional Spiking Neural Network Retina-Neuromorphic Approach
Paul Kirkland, Gaetano Di Caterina, John J. Soraghan, Yiannis Andreopoulos, George Matich
ICANN (1)4
2019 Neuromorphic Vision Sensing for CNN-based Action Recognition
abstract
Neuromorphic vision sensing (NVS) hardware is now gaining traction as a low-power/high-speed visual sensing technology that circumvents the limitations of conventional active pixel sensing (APS) cameras. While object detection and tracking models have been investigated in conjunction with NVS, there is currently little work on NVS for higher-level semantic tasks, such as action recognition. Contrary to recent work that considers homogeneous transfer between flow domains (optical flow to motion vectors), we propose to embed an NVS emulator into a multi-modal transfer learning framework that carries out heterogeneous transfer from optical flow to NVS. The potential of our framework is showcased by the fact that, for the first time, our NVS-based results achieve comparable action recognition performance to motion-vector or optical-flow based methods (i.e., accuracy on UCF-101 within 8.8% of I3D with optical flow), with the NVS emulator and NVS camera hardware offering 3 to 6 orders of magnitude faster frame generation (respectively) compared to standard Brox optical flow. Beyond this significant advantage, our CNN processing is found to have the lowest total GFLOP count against all competing methods (up to 7.7 times complexity saving compared to I3D with optical flow).
Aaron Chadha, Yin Bi, Alhabib Abbas, Yiannis Andreopoulos
ICASSP4
2019 Graph-Based Object Classification for Neuromorphic Vision Sensing
abstract
Neuromorphic vision sensing (NVS) devices represent visual information as sequences of asynchronous discrete events (a.k.a., "spikes'") in response to changes in scene reflectance. Unlike conventional active pixel sensing (APS), NVS allows for significantly higher event sampling rates at substantially increased energy efficiency and robustness to illumination changes. However, object classification with NVS streams cannot leverage on state-of-the-art convolutional neural networks (CNNs), since NVS does not produce frame representations. To circumvent this mismatch between sensing and processing with CNNs, we propose a compact graph representation for NVS. We couple this with novel residual graph CNN architectures and show that, when trained on spatio-temporal NVS data for object classification, such residual graph CNNs preserve the spatial and temporal coherence of spike events, while requiring less computation and memory. Finally, to address the absence of large real-world NVS datasets for complex recognition tasks, we present and make available a 100k dataset of NVS recordings of the American sign language letters, acquired with an iniLabs DAVIS240c device under real-world conditions.
Yin Bi, Aaron Chadha, Alhabib Abbas, Eirina Bourtsoulatze, Yiannis Andreopoulos
ICCV5
2019 Dithen: A Computation-as-a-Service Cloud Platform for Large-Scale Multimedia Processing
abstract
We present Dithen, a novel computation-as-a-service (CaaS) cloud platform specifically tailored to the parallel execution of large-scale multimedia tasks. Dithen handles the upload/download of both multimedia data and executable items, the assignment of compute units to multimedia workloads, and the reactive control of the available compute units to minimize the cloud infrastructure cost under deadline-abiding execution. Dithen combines three key properties: (i) the reactive assignment of individual multimedia tasks to available computing units according to availability and predetermined time-to-completion constraints; (ii) optimal resource estimation based on Kalman-filter estimates; (iii) the use of additive increase multiplicative decrease (AIMD) algorithms (famous for being the resource management in the transport control protocol) for the control of the number of units servicing workloads. The deployment of Dithen over Amazon EC2 spot instances is shown to be capable of processing more than 80,000 video transcoding, face detection and image processing tasks (equivalent to the processing of more than 116 GB of compressed data) for less than $1 in billing cost from EC2. Moreover, the proposed AIMD-based control mechanism, in conjunction with the Kalman estimates, is shown to provide for more than 27 percent reduction in EC2 spot instance cost against methods based on reactive resource estimation. Finally, Dithen is shown to offer a 38 to 500 percent reduction of the billing cost against the current state-of-the-art in CaaS platforms on Amazon EC2 (Amazon Lambda and Amazon Autoscale). A baseline version of Dithen is currently available at dithen.com under the “AutoScale” option.
Joseph Doyle, Vasileios Giotsas, Mohammad Ashraful Anam, Yiannis Andreopoulos
IEEE Trans. Cloud Comput.4
2019 Video Classification With CNNs: Using the Codec as a Spatio-Temporal Activity Sensor
abstract
We investigate video classification via a two-stream convolutional neural network (CNN) design that directly ingests information extracted from compressed video bitstreams. Our approach begins with the observation that all modern video codecs divide the input frames into macroblocks (MBs). We demonstrate that selective access to MB motion vector (MV) information within compressed video bitstreams can also provide for selective, motion-adaptive, MB pixel decoding (a.k.a., MB texture decoding). This in turn allows for the derivation of spatio-temporal video activity regions at extremely high speed in comparison to conventional full-frame decoding followed by optical flow estimation. In order to evaluate the accuracy of a video classification framework based on such activity data, we independently train two CNN architectures on MB texture and MV correspondences and then fuse their scores to derive the final classification of each test video. Evaluation on two standard data sets shows that the proposed approach is competitive with the best two-stream video classification approaches found in the literature. At the same time: 1) a CPU-based realization of our MV extraction is over 977 times faster than GPU-based optical flow methods; 2) selective decoding is up to 12 times faster than full-frame decoding; and 3) our proposed spatial and temporal CNNs perform inference at 5 to 49 times lower cloud computing cost than the fastest methods from the literature.
Aaron Chadha, Alhabib Abbas, Yiannis Andreopoulos
IEEE Trans. Circuits Syst. Video Technol.3
2018 Rate-Accuracy Trade-Off in Video Classification with Deep Convolutional Neural Networks
abstract
Advanced video classification systems decode video frames to derive the necessary texture and motion representations for ingestion and analysis by spatio-temporal deep convolutional neural networks (CNNs). However, when considering visual Internet -of- Things applications, surveillance systems and semantic crawlers of large video repositories, the compressed video content and the CNN-based semantic analysis parts do not tend to be co-located. This necessitates the transport of compressed video over networks and incurs significant overhead in bandwidth and energy consumption, thereby significantly undermining the deployment potential of such systems. In this paper, we investigate the trade-off between the encoding bitrate and the achievable accuracy of CNN-based video classification that ingests AVC/H.264 encoded videos. Instead of entire compressed video bitstreams, we only retain motion vector and selected texture information at significantly reduced bitrates. Based on two CNN architectures and two action recognition datasets, we achieve 38%-59% saving in bitrate with marginal impact in classification accuracy. A simple rate-based selection between the two CNNs shows that even further bitrate savings are possible with graceful degradation in accuracy. This may allow for rate/accuracy-optimized CNN-based video classification over networks.
Alhabib Abbas, Aaron Chadha, Yiannis Andreopoulos, Mohammad K. Jubran
ICIP3
2017 PIX2NVS: Parameterized conversion of pixel-domain video frames to neuromorphic vision streams
abstract
We propose and make available a generic pixel-to-neuromorphic vision stream (PIX2NVS) framework in order to allow for the generation of neuromorphic data streams from conventional pixel-domain video frames. In order to quantify the accuracy of our framework against experimentally-derived NVS data from previous work, we also propose and validate two metrics, the Chamfer distance and e-repeatability. The most important application of PIX2NVS will be in the generation of artificial NVS from large annotated video frame collections used in machine learning research, e.g., YouTube-8M, YFCC100m, YouTube-BoundingBoxes, thereby transferring these datasets to the neuromorphic domain.
Yin Bi, Yiannis Andreopoulos
ICIP2
2017 Compressed-domain video classification with deep neural networks: "There's way too much information to decode the matrix"
abstract
We investigate video classification via a 3D deep convolutional neural network (CNN) that directly ingests compressed bitstream information. This idea is based on the observation that video macroblock (MB) motion vectors (that are very compact and directly available from the compressed bitstream) are inherently capturing local spatio-temporal changes in each video scene. Our results on two standard video datasets show that our approach outperforms pixel-based approaches and remains within 7 percentile points from the best classification results reported by highly-complex optical-flow & deep-CNN methods. At the same time, a CPU-based realization of our approach is found to be more than 2500 times faster in the motion extraction in comparison to GPU-based optical flow methods and also offers 2 to 3.4-fold reduction in the utilized deep CNN weights compared to recent architectures. This indicates that deep learning based on compressed video bitstream information may allow for advanced video classification to be deployed in very large datasets using commodity CPU hardware. Source code is available at http://www.github.com/mvcnn.
Aaron Chadha, Alhabib Abbas, Yiannis Andreopoulos
ICIP3
2017 Mitigating Silent Data Corruptions in Integer Matrix Products: Toward Reliable Multimedia Computing on Unreliable Hardware
abstract
The generic matrix multiply (GEMM) routine comprises the compute and memory-intensive parts of many information retrieval, machine learning, and object recognition systems that process integer inputs. Therefore, it is of paramount importance to ensure that integer GEMM computations remain robust to silent data corruptions (SDCs), which stem from accidental voltage or frequency overscaling, or other hardware nonidealities. In this paper, we introduce a new method for SDC mitigation based on the concept of numerical packing. The key difference between our approach and all existing methods is the production of redundant resultswithinthe numerical representation of the outputs, rather than as a separate set of checksums. Importantly, unlike well-known algorithm-based fault tolerance approaches for GEMM, the proposed approach can reliably detect the locations of the vast majority of all possible SDCs in the results of GEMM computations. An experimental investigation of voltage-scaled integer GEMM computations for visual descriptor matching within state-of-the-art image and video retrieval algorithms running on an Intel i7-4578U 3 GHz processor shows that SDC mitigation based on numerical packing leads to comparable or lower execution and energy-consumption overhead in comparison with all other alternatives.
Ijeoma J. F. Ezika, Mohammad Ashraful Anam, Fabio Verdicchio, Yiannis Andreopoulos
IEEE Trans. Circuits Syst. Video Technol.4
2017 Voronoi-Based Compact Image Descriptors: Efficient Region-of-Interest Retrieval With VLAD and Deep-Learning-Based Descriptors
abstract
We investigate the problem of image retrieval based on visual queries when the latter comprise arbitrary regions-of-interest (ROI) rather than entire images. Our proposal is a compact image descriptor that combines the state-of-the-art in content-based descriptor extraction with a multilevel, Voronoi-based spatial partitioning of each dataset image. The proposed multilevel Voronoi-based encoding uses a spatial hierarchical K-means over interest-point locations, and computes a content-based descriptor over each cell. In order to reduce the matching complexity with minimal or no sacrifice in retrieval performance: 1) we utilize the tree structure of the spatial hierarchical K-means to perform top-to-bottom pruning for local similarity maxima; 2) we propose a new image similarity score that combines relevant information from all partition levels into a single measure for similarity; 3) we combine our proposal with a novel and efficient approach for optimal bit allocation within quantized descriptor representations. By deriving both a Voronoi-based VLAD descriptor (called Fast-VVLAD) and a Voronoi-based deep convolutional neural network (CNN) descriptor (called Fast-VDCNN), we demonstrate that our Voronoi-based framework is agnostic to the descriptor basis, and can easily be slotted into existing frameworks. Via a range of ROI queries in two standard datasets, it is shown that the Voronoi-based descriptors achieve comparable or higher mean average precision against conventional grid-based spatial search, while offering more than twofold reduction in complexity. Finally, beyond ROI queries, we show that Voronoi partitioning improves the geometric invariance of compact CNN descriptors, thereby resulting in competitive performance to the current state-of-the-art on whole image retrieval.
Aaron Chadha, Yiannis Andreopoulos
IEEE Trans. Multim.2
2016 Cloud Instance Management and Resource Prediction for Computation-as-a-Service Platforms
abstract
Computation-as-a-Service (CaaS) offerings have gained traction in the last few years due to their effectiveness in balancing between the scalability of Software-as-a-Service and the customisation possibilities of Infrastructure-as-a-Service platforms. To function effectively, a CaaS platform must have three key properties: (i) reactive assignment of individual processing tasks to available cloud instances (compute units) according to availability and predetermined time-to-completion (TTC) constraints, (ii) accurate resource prediction, (iii) efficient control of the number of cloud instances servicing workloads, in order to optimize between completing workloads in a timely fashion and reducing resource utilization costs. In this paper, we propose three approaches that satisfy these properties (respectively): (i) a service rate allocation mechanism based on proportional fairness and TTC constraints, (ii) Kalman-filter estimates for resource prediction, and (iii) the use of additive increase multiplicative decrease (AIMD) algorithms (famous for being the resource management in the transport control protocol) for the control of the number of compute units servicing workloads. The integration of our three proposals into a single CaaS platform is shown to provide for more than 27% reduction in Amazon EC2 spot instance cost against methods based on reactive resource prediction and 38% to 60% reduction of the billing cost against the current state-of-the-art in CaaS platforms (Amazon Lambda and Autoscale).
Joseph Doyle, Vasileios Giotsas, Mohammad Ashraful Anam, Yiannis Andreopoulos
IC2E4
2016 Core Failure Mitigation in Integer Sum-of-Product Computations on Cloud Computing Systems
abstract
The decreasing mean-time-to-failure estimates in cloud computing systems indicate that multimedia applications running on such environments should be able to mitigate an increasing number of core failures at runtime. We propose a new roll-forward failure-mitigation approach for integer sum-of-product computations, with emphasis on generic matrix multiplication (GEMM) and convolution/crosscorrelation (CONV) routines. Our approach is based on the production of redundant results within the numerical representation of the outputs via the use of numerical packing. This differs from all existing roll-forward solutions that require a separate set of checksum (or duplicate) results. Our proposal imposes 37.5% reduction in the maximum output bitwidth supported in comparison to integer sum-of-product realizations performed on 32-bit integer representations which is comparable to the bitwidth requirement of checksum-methods for multiple core failure mitigation. Experiments with state-of-the-art GEMM and CONV routines running on a c4.8xlarge compute-optimized instance of amazon web services elastic compute cloud (AWS EC2) demonstrate that the proposed approach is able to mitigate up to one quadcore failure while achieving processing throughput that is: 1) comparable to that of the conventional, failure-intolerant, integer GEMM and CONV routines, 2) substantially superior to that of the equivalent roll-forward failure-mitigation method based on checksum streams. Furthermore, when used within an image retrieval framework deployed over a cluster of AWS EC2 spot (i.e., low-cost albeit terminatable) instances, our proposal leads to: 1) 16%-23% cost reduction against the equivalent checksum-based method and 2) more than 70% cost reduction against conventional failure-intolerant processing on AWS EC2 on-demand (i.e., higher-cost albeit guaranteed) instances.
Ijeoma J. F. Ezika, Yiannis Andreopoulos
IEEE Trans. Multim.2
2016 Media Query Processing for the Internet-of-Things: Coupling of Device Energy Consumption and Cloud Infrastructure Billing
abstract
Audio/visual recognition and retrieval applications have recently garnered significant attention within Internet-of-Things-oriented services, given that video cameras and audio processing chipsets are now ubiquitous even in low-end embedded systems. In the most typical scenario for such services, each device extracts audio/visual features and compacts them into feature descriptors, which comprise media queries. These queries are uploaded to a remote cloud computing service that performs content matching for classification or retrieval applications. Two of the most crucial aspects for such services are: (1) controlling the device energy consumption when using the service, and (2) reducing the billing cost incurred from the cloud infrastructure provider. In this paper, we derive analytic conditions for the optimal coupling between the device energy consumption and the incurred cloud infrastructure billing. Our framework encapsulates: the energy consumption to produce and transmit audio/visual queries, the billing rates of the cloud infrastructure, the number of devices concurrently connected to the same cloud server, the query volume constraint of each cluster of devices, and the statistics of the query data production volume per device. Our analytic results are validated via a deployment with: (1) the device side comprising compact image descriptors (queries) computed on Beaglebone Linux embedded platforms and transmitted to Amazon Web Services (AWS) Simple Storage Service, and (2) the cloud side carrying out image similarity detection via AWS Elastic Compute Cloud (EC2) instances, with the AWS Auto Scaling being used to control the number of instances according to the demand.
Francesco Renna, Joseph Doyle, Vasileios Giotsas, Yiannis Andreopoulos
IEEE Trans. Multim.4
2015 Vectors of locally aggregated centers for compact video representation
abstract
We propose a novel vector aggregation technique for compact video representation, with application in accurate similarity detection within large video datasets. The current state-of-the-art in visual search is formed by the vector of locally aggregated descriptors (VLAD) of Jegou et al. VLAD generates compact video representations based on scale-invariant feature transform (SIFT) vectors (extracted per frame) and local feature centers computed over a training set. With the aim to increase robustness to visual distortions, we propose a new approach that operates at a coarser level in the feature representation. We create vectors of locally aggregated centers (VLAC) by first clustering SIFT features to obtain local feature centers (LFCs) and then encoding the latter with respect to given centers of local feature centers (CLFCs), extracted from a training set. The sum-of-differences between the LFCs and the CLFCs are aggregated to generate an extremely-compact video description used for accurate video segment similarity detection. Experimentation using a video dataset, comprising more than 1000 minutes of content from the Open Video Project, shows that VLAC obtains substantial gains in terms of mean Average Precision (mAP) against VLAD and the hyper-pooling method of Douze et al., under the same compaction factor and the same set of distortions.
Alhabib Abbas, Nikos Deligiannis, Yiannis Andreopoulos
ICME3
2015 Region-of-Interest Retrieval in Large Image Datasets with Voronoi VLAD
Aaron Chadha, Yiannis Andreopoulos
ICVS2
2015 Failure mitigation in linear, sesquilinear and bijective operations on integer data streams via numerical entanglement
abstract
A new roll-forward technique is proposed that recovers from any single fail-stop failure in M integer data streams (M ≥ 3) when undergoing linear, sesquilinear or bijective (LSB) operations, such as: scaling, additions/subtractions, inner or outer vector products and permutations. In the proposed approach, the M input integer data streams are linearly superimposed to form M numerically entangled integer data streams that are stored inplace of the original inputs. A series of LSB operations can then be performed directly using these entangled data streams. The output results can be extracted from any M-1 entangled output streams by additions and arithmetic shifts, thereby guaranteeing robustness to a fail-stop failure in any single stream computation. Importantly, unlike other methods, the number of operations required for the entanglement, extraction and recovery of the results is linearly related to the number of the inputs and does not depend on the complexity of the performed LSB operations. We have validated our proposal in an Intel processor (Haswell architecture with AVX2 support) via convolution operations. Our analysis and experiments reveal that the proposed approach incurs only 1.8% to 2.8% reduction in processing throughput in comparison to the failure-intolerant approach. This overhead is 9 to 14 times smaller than that of the equivalent checksum-based method. Thus, our proposal can be used in distributed systems and unreliable processor hardware, or safety-critical applications, where robustness against fail-stop failures becomes a necessity.
Mohammad Ashraful Anam, Yiannis Andreopoulos
IOLTS2
2015 Mitigation of fail-stop failures in integer matrix products via numerical packing
abstract
The decreasing mean-time-to-failure estimates of distributed computing systems indicate that high-performance generic matrix multiply (GEMM) routines running on such environments may need to mitigate an increasing number of fail-stop failures. We propose a new roll-forward solution to this problem that is based on the production of redundant results within the numerical representation of the outputs via the use of numerical packing. This differs from all existing roll-forward solutions that require a separate set of checksum (or duplicate) results. In particular, unlike all existing approaches, the proposed approach does not require additional hardware resources for failure mitigation. Instead, in our proposal the required duplication is inserted in the input matrices themselves. The accommodation of the duplicated inputs imposes 30.6% or 37.5% reduction in the maximum output bitwidth supported in comparison to integer matrix products performed on 32-bit floating-point or integer representations, respectively. Nevertheless, this bitwidth reduction is comparable to the one imposed due to the checksum elements of traditional roll-forward methods, especially for cases where multiple core failures must be mitigated. Experiments performed on an Amazon EC2 instance with 6 Intel Haswell cores dedicated to GEMM computations show that, in comparison to the state-of-the-art failure-intolerant integer GEMM realization, the proposed approach incurs only 5-19.4% drop in the achievable peak performance. This overhead is significantly lower than the 33.3 - 37% overhead incurred by the equivalent checksum-based method.
Ijeoma J. F. Ezika, Yiannis Andreopoulos
IOLTS2
2015 Decentralized multichannel medium access control: viewing desynchronization as a convex optimization method
abstract
Desynchronization algorithms are essential in the design of collision-free medium access control (MAC) mechanisms for wireless sensor networks. Desync is a well-known desynchronization algorithm that operates under limited listening. In this paper, we view Desync as a gradient descent method solving a convex optimization problem. This enables the design of a novel decentralized, collision-free, multichannel medium access control (MAC) algorithm. Moreover, by using Nesterov's fast gradient method, we obtain a new algorithm that converges to the steady network state much faster. Simulations and experimental results on an IEEE 802.15.4-based wireless sensor network deployment show that our algorithms achieve significantly faster convergence to steady network state and substantially higher throughput compared to the recently standardized IEEE 802.15.4e-2012 time synchronized channel hopping (TSCH) scheme. In addition, our mechanism has a comparable power dissipation with respect to TSCH and does not need a coordinator node or coordination channel.
Nikos Deligiannis, João F. C. Mota, George Smart, Yiannis Andreopoulos
IPSN4
2015 Decentralized time-synchronized channel swapping
abstract
We are working on a new concept for decentralized medium access control (MAC), termed decentralized time-synchronized channel swapping (DT-SCS). Under the proposed DT-SCS and its associated MAC-layer protocol, wireless nodes converge to synchronous beacon packet transmissions across all IEEE802.15.4 channels, with balanced numbers of nodes in each channel. This is achieved by reactive listening mechanisms, based on pulse coupled oscillator techniques. Once convergence to the multichannel time-synchronized state is achieved, peer-to-peer channel swapping can then take place via swap requests and acknowledgments made by concurrent transmitters in neighboring channels. Our implementation of DT-SCS reveals that our proposal comprises an excellent candidate for completely decentralized MAC-layer coordination in WSNs by providing for quick convergence to steady state, high bandwidth utilization, high connectivity and robustness to interference and hidden nodes. The demo will showcase the properties of DT-SCS and will also present its behaviour under various scenarios for hidden nodes and interference, both experimentally and with the help of visualization of simulation results.
George Smart, Nikos Deligiannis, João F. C. Mota, Yiannis Andreopoulos
IPSN4
2015 Fast Desynchronization for Decentralized Multichannel Medium Access Control
abstract
Distributed desynchronization algorithms are key to wireless sensor networks as they allow for medium access control in a decentralized manner. In this paper, we view desynchronization primitives as iterative methods that solve optimization problems. In particular, by formalizing a well established desynchronization algorithm as a gradient descent method, we establish novel upper bounds on the number of iterations required to reach convergence. Moreover, by using Nesterov's accelerated gradient method, we propose a novel desynchronization primitive that provides for faster convergence to the steady state. Importantly, we propose a novel algorithm that leads to decentralized time-synchronous multichannel TDMA coordination by formulating this task as an optimization problem. Our simulations and experiments on a densely-connected IEEE 802.15.4-based wireless sensor network demonstrate that our scheme provides for faster convergence to the steady state, robustness to hidden nodes, higher network throughput and comparable power dissipation with respect to the recently standardized IEEE 802.15.4e-2012 time-synchronized channel hopping (TSCH) scheme.
Nikos Deligiannis, João F. C. Mota, George Smart, Yiannis Andreopoulos
IEEE Trans. Commun.4
2014 Energy Consumption of Visual Sensor Networks: Impact of Spatio-Temporal Coverage Based on Single-Hop Topologies
Alessandro Redondi, Dujdow Buranapanichkit, Matteo Cesana, Marco Tagliasacchi, Yiannis Andreopoulos
EWSN5
2014 Bandit framework for systematic learning in wireless video-based face recognition
abstract
In most video-based object or face recognition services on mobile devices, each device captures and transmits video frames over wireless to a remote computing service (a.k.a. “cloud”) that performs the heavy-duty video feature extraction and recognition tasks for a large number of mobile devices. The major challenges of such scenarios stem from the highly-varying contention levels in the wireless local area network (WLAN), as well as the variation in the task-scheduling congestion in the cloud. In order for each device to maximize its object or face recognition rate under such contention and congestion variability, we propose a systematic learning framework based on multi-armed bandits. Unlike well-known reinforcement learning techniques that exhibit very slow convergence rates when operating in highly-dynamic environments, the proposed bandit-based systematic learning quickly approaches the optimal transmission and processing-complexity policies based on feedback on the experienced dynamics (contention and congestion levels). Comparisons against state-of-the-art reinforcement learning methods demonstrate that this makes our proposal especially suitable for the highly-dynamic levels of wireless contention and cloud scheduling congestion.
Onur Atan, Cem Tekin, Mihaela van der Schaar, Yiannis Andreopoulos
ICASSP4
2014 On the stochastic modeling of desynchronization convergence in wireless sensor networks
abstract
Desynchronization is a fundamental approach in wireless sensor networks that allows for convergence to time-division multiple access (TDMA) of the medium without the need for clock synchronization and centralized coordination. The method is based on the concept of reactive listening of periodic fre message broadcasts between nodes sharing the given spectrum. We propose a novel framework to estimate the required iterations for convergence to fair TDMA scheduling. Unlike previous conjectures or bounds found in the literature, our estimation framework is based on a stochastic modeling approach. Experiments via imote2 TinyOS nodes and simulations demonstrate that the proposed estimates characterize the experimental desynchronization convergence iterations signifcantly better than existing conjectures or bounds.
Dujdow Buranapanichkit, Nikos Deligiannis, Yiannis Andreopoulos
ICASSP3
2014 Non-Stationary Resource Allocation Policies for Delay-Constrained Video Streaming: Application to Video over Internet-of-Things-Enabled Networks
abstract
Due to the high bandwidth requirements and stringent delay constraints of multi-user wireless video transmission applications, ensuring that all video senders have sufficient transmission opportunities to use before their delay deadlines expire is a longstanding research problem. We propose a novel solution that addresses this problem without assuming detailed packet-level knowledge, which is unavailable at resource allocation time (i.e. prior to the actual compression and transmission). Instead, we translate the transmission delay deadlines of each sender's video packets into a monotonically-decreasing weight distribution within the considered time horizon. Higher weights are assigned to the slots that have higher probability for deadline-abiding delivery. Given the sets of weights of the senders' video streams, we propose the low-complexity Delay-Aware Resource Allocation (DARA) approach to compute the optimal slot allocation policy that maximizes the deadline-abiding delivery of all senders. A unique characteristic of the DARA approach is that it yields a non-stationary slot allocation policy that depends on the allocation of previous slots. This is in contrast with all existing slot allocation policies such as round-robin or rate-adaptive round-robin policies, which are stationary because the allocation of the current slot does not depend on the allocation of previous slots. We prove that the DARA approach is optimal for weight distributions that are exponentially decreasing in time. We further implement our framework for real-time video streaming in wireless personal area networks that are gaining significant traction within the new Internet-of-Things (IoT) paradigm. For multiple surveillance videos encoded with H.264/AVC and streamed via the 6tisch framework that simulates the IoT-oriented IEEE 802.15.4e TSCH medium access control, our solution is shown to be the only one that ensures all video bitstreams are delivered with acceptable quality in a deadline-abiding manner.
Jie Xu 0001, Yiannis Andreopoulos, Yuanzhang Xiao, Mihaela van der Schaar
IEEE J. Sel. Areas Commun.2
2014 Multimedia modeling
Chong-Wah Ngo, Klaus Schöffmann, Yiannis Andreopoulos, Christian Breiteneder
Multim. Tools Appl.3
2014 Precision-Energy-Throughput Scaling of Generic Matrix Multiplication and Convolution Kernels via Linear Projections
abstract
Generic matrix multiplication (GEMM) and convolution (CONV)/cross-correlation kernels often constitute the bulk of the compute- and memory-intensive processing within image/audio recognition and matching systems. We propose a novel method to scale the energy and processing throughput of GEMM and CONV kernels for such error-tolerant multimedia applications by adjusting the precision of computation. Our technique employs linear projections to the input matrix or signal data during the top-level GEMM and CONV blocking and reordering. The GEMM and CONV kernel processing then uses the projected inputs and the results are accumulated to form the final outputs. Throughput and energy scaling takes place by changing the number of projections computed by each kernel, which in turn produces approximate results, i.e., changes the precision of the performed computation. Results derived from a voltage- and frequency-scaled ARM Cortex A15 processor running face recognition and music-matching algorithms demonstrate that the proposed approach allows for a 280%-440% increase of processing throughput and a 75%-80% decrease of energy consumption against the optimized GEMM and CONV kernels without any impact on the obtained recognition or matching accuracy. Even higher gains can be obtained, if one is willing to tolerate some reduction in the accuracy of the recognition and matching applications.
Mohammad Ashraful Anam, Paul N. Whatmough, Yiannis Andreopoulos
IEEE Trans. Circuits Syst. Video Technol.3
2014 Energy Consumption of Visual Sensor Networks: Impact of Spatio-Temporal Coverage
abstract
Wireless visual sensor networks (VSNs) are expected to play a major role in future IEEE 802.15.4 personal area networks (PANs) under recently established collision-free medium access control (MAC) protocols, such as the IEEE 802.15.4e-2012 MAC. In such environments, the VSN energy consumption is affected by a number of camera sensors deployed (spatial coverage), as well as a number of captured video frames of which each node processes and transmits data (temporal coverage). In this paper we explore this aspect for uniformly formed VSNs, that is, networks comprising identical wireless visual sensor nodes connected to a collection node via a balanced cluster-tree topology, with each node producing independent identically distributed bitstream sizes after processing the video frames captured within each network activation interval. We derive analytic results for the energy-optimal spatio-temporal coverage parameters of such VSNs under a priori known bounds for the number of frames to process per sensor and the number of nodes to deploy within each tier of the VSN. Our results are parametric to the probability density function characterizing the bitstream size produced by each node and the energy consumption rates of the system of interest. Experimental results are derived from a deployment of TelosB motes and reveal that our analytic results are always within 7% of the energy consumption measurements for a wide range of settings. In addition, results obtained via motion JPEG encoding and feature extraction on a multimedia subsystem (BeagleBone Linux Computer) show that the optimal spatio-temporal settings derived by our framework allow for substantial reduction of energy consumption in comparison with ad hoc settings.
Alessandro Redondi, Dujdow Buranapanichkit, Matteo Cesana, Marco Tagliasacchi, Yiannis Andreopoulos
IEEE Trans. Circuits Syst. Video Technol.5
2013 Towards energy neutrality in energy harvesting wireless sensor networks: A case for distributed compressive sensing?
abstract
This paper advocates the use of the emerging distributed compressive sensing (DCS) paradigm in order to deploy energy harvesting (EH) wireless sensor networks (WSN) with practical network lifetime and data gathering rates that are substantially higher than the state-of-the-art. In particular, we argue that there are two fundamental mechanisms in an EH WSN: i) the energy diversity associated with the EH process that entails that the harvested energy can vary from sensor node to sensor node, and ii) the sensing diversity associated with the DCS process that entails that the energy consumption can also vary across the sensor nodes without compromising data recovery. We also argue that such mechanisms offer the means to match closely the energy demand to the energy supply in order to unlock the possibility for energy-neutral WSNs that leverage EH capability. A number of analytic and simulation results are presented in order to illustrate the potential of the approach.
Wei Chen 0016, Yiannis Andreopoulos, Ian J. Wassell, Miguel R. D. Rodrigues
GLOBECOM2
2013 Highly-reliable integer matrix multiplication via numerical packing
abstract
The generic matrix multiply (GEMM) routine comprises the compute and memory-intensive part of many information retrieval, relevance ranking and object recognition systems. Because of the prevalence of GEMM in these applications, ensuring its robustness to transient hardware faults is of paramount importance for highly-efficientlhighly-reliable systems. This is currently accomplished via error control coding (ECC) or via dual modular redundancy (DMR) approaches that produce a separate set of “parity” results to allow for fault detection in GEMM. We introduce a third family of methods for fault detection in integer matrix products based on the concept of numerical packing. The key difference of the new approach against ECC and DMR approaches is the production of redundant results within the numerical representation of the inputs rather than as a separate set of parity results. In this way, high reliability is ensured within integer matrix products while allowing for: (i) in-place storage; (ii) usage of any off-the-shelf 64-bit floating-point GEMM routine; (iii) computational overhead that is independent of the GEMM inner dimension. The only detriment against a conventional (i.e. fault-intolerant) integer matrix multiplication based on 32-bit floating-point GEMM is the sacrifice of approximately 30.6% of the bitwidth of the numerical representation. However, unlike ECC methods that can reliably detect only up to a few faults per GEMM computation (typically two), the proposed method attains more than “12 nines” reliability, i.e. it will only fail to detect 1 fault out of more than 1 trillion arbitrary faults in the GEMM operations. As such, it achieves reliability that approaches that of DMR, at a very small fraction of its cost. Specifically, a single-threaded software realization of our proposal on an Intel i7-3632QM 2.2GHz processor (Ivy Bridge architecture with AVX support) incurs, on average, only 19% increase of execution time against an optimized, fault-intolerant, 32-bit GEMM routine over a range of matrix sizes and it remains more than 80% more efficient than a DMR-based GEMM.
Ijeoma J. F. Ezika, Mohammad Ashraful Anam, Davide Anastasia, Fabio Verdicchio, Yiannis Andreopoulos
IOLTS5
2013 Error Tolerant Multimedia Stream Processing: There's Plenty of Room at the Top (of the System Stack)
abstract
There is a growing realization that the expected fault rates and energy dissipation stemming from increases in CMOS integration will lead to the abandonment of traditional system reliability in favor of approaches that offer reliability to hardware-induced errors across the application, runtime support, architecture, device and integrated-circuit (IC) layers. Commercial stakeholders of multimedia stream processing (MSP) applications, such as information retrieval, stream mining systems, and high-throughput image and video processing systems already feel the strain of inadequate system-level scaling and robustness under the always-increasing user demand. While such applications can tolerate certain imprecision in their results, today's MSP systems do not support a systematic way to exploit this aspect for cross-layer system resilience. However, research is currently emerging that attempts to utilize the error-tolerant nature of MSP applications for this purpose. This is achieved by modifications to all layers of the system stack, from algorithms and software to the architecture and device layer, and even the IC digital logic synthesis itself. Unlike conventional processing that aims for worst-case performance and accuracy guarantees, error-tolerant MSP attempts to provide guarantees for the expected performance and accuracy. In this paper we review recent advances in this field from an MSP and a system (layer-by-layer) perspective, and attempt to foresee some of the components of future cross-layer error-tolerant system design that may influence the multimedia and the general computing landscape within the next ten years.
Yiannis Andreopoulos
IEEE Trans. Multim.1
2013 Guest Editorial: Special Section on New Software/Hardware Paradigms for Error-Tolerant Multimedia Systems
abstract
The five papers in this special section are devoted to the topic of new software and hardware programs and services used for error-tolerant multimedia applications.
Yiannis Andreopoulos, Liang-Gee Chen, Brian L. Evans
IEEE Trans. Multim.1
2013 Analytic Conditions for Energy Neutrality in Uniformly-Formed Wireless Sensor Networks
abstract
Future deployments of wireless sensor network (WSN) infrastructures for environmental or event monitoring are expected to be equipped with energy harvesters (e.g. piezoelectric, thermal, photovoltaic) in order to substantially increase their autonomy. In this paper we derive conditions for energy neutrality, i.e. perpetual energy autonomy per sensor node, by balancing the node's expected energy consumption with its expected energy harvesting capability. Our analysis assumes a uniformly-formed WSN, i.e. a network comprising identical transmitter sensor nodes and identical receiver/relay sensor nodes with a balanced cluster-tree topology. The proposed framework is parametric to: (i) the duty cycle for the network activation; (ii) the number of nodes in the same tier of the cluster-tree topology; (iii) the consumption rate of the receiver node(s) that collect (and possibly relay) data along with their own; (iv) the marginal probability density function (PDF) characterizing the data transmission rate per node; (v) the expected amount of energy harvested by each node. Based on our analysis, we obtain the number of nodes leading to the minimum energy harvestingrequirement for each tier of the WSN cluster-tree topology. We also derive closed-form expressions for the difference in the minimum energy harvesting requirements between four transmission rate PDFs in function of the WSN parameters. Our analytic results are validated via experiments using TelosB sensor nodes and an energy measurement testbed. Our framework is useful for feasibility studies on energy harvesting technologies in WSNs and for optimizing the operational settings of hierarchical WSN-based monitoring infrastructures prior to time-consuming testing and deployment within the application environment.
Hana Besbes, George Smart, Dujdow Buranapanichkit, Christos Kloukinas, Yiannis Andreopoulos
IEEE Trans. Wirel. Commun.5
2012 Throughput Scaling Of Convolution For Error-Tolerant Multimedia Applications
abstract
Convolution and cross-correlation are the basis of filtering and pattern or template matching in multimedia signal processing. We propose two throughput scaling options for any one-dimensional convolution kernel in programmable processors by adjusting the imprecision (distortion) of computation. Our approach is based on scalar quantization, followed by two forms of tight packing in floating-point (one of which is proposed in this paper) that allow for concurrent calculation of multiple results. We illustrate how our approach can operate as an optional pre- and post-processing layer for off-the-shelf optimized convolution routines. This is useful for multimedia applications that are tolerant to processing imprecision and for cases where the input signals are inherently noisy (error tolerant multimedia applications). Indicative experimental results with a digital music matching system and an MPEG-7 audio descriptor system demonstrate that the proposed approach offers up to 175% increase in processing throughput against optimized (full-precision) convolution with virtually no effect in the accuracy of the results. Based on marginal statistics of the input data, it is also shown how the throughput and distortion can be adjusted per input block of samples under constraints on the signal-to-noise ratio against the full-precision convolution.
Mohammad Ashraful Anam, Yiannis Andreopoulos
IEEE Trans. Multim.2
2011 Distortion estimates for adaptive temporal decompositions of video under displacement errors and quantization noise
abstract
In video communication systems, due to quantization and transmission errors, mismatches between the transmitter- and receiver-side information may occur, severely impacting the reconstructed video. Theoretical understanding of the quality degradation ensuing from such mismatches is essential when targeting quality-of-service for video communications. In this paper, by viewing the mismatches in the transform coefficients and the adaptive parameters of the temporal analysis of a video coding system as perturbations in the synthesis system, we derive analytic approximations for the expected reconstruction distortion. Our theoretical results are experimentally assessed using adaptive temporal decompositions within a video coding system based on motion-adaptive temporal lifting decomposition. Since we focus on the generic case of adaptive lifting transforms, our results can provide useful insights for estimation-theoretic resiliency mechanisms to be considered within standardized transform-based codecs.
Fabio Verdicchio, Yiannis Andreopoulos
ICIP2
2011 Distortion estimates for adaptive lifting transforms with noise
Fabio Verdicchio, Yiannis Andreopoulos
Image Vis. Comput.2
2010 Scheduling and energy-distortion tradeoffs with operational refinement of image processing
abstract
Ubiquitous image processing tasks (such as transform decompositions, filtering and motion estimation) do not currently provide graceful degradation when their clock-cycles budgets are reduced, e.g. when delay deadlines are imposed in a multi-tasking environment to meet throughput requirements. This is an important obstacle in the quest for full utilization of modern programmable platforms' capabilities, since: (i) worst-case considerations must be in place for reasonable quality of results; (ii) throughput-distortion tradeoffs are not possible for distortion-tolerant image processing applications without cumbersome (and potentially costly) system customization. In this paper, we extend the functionality of the recently-proposed software framework for operational refinement of image processing (ORIP) and demonstrate its inherent throughput-distortion and energy-distortion scalability. Importantly, our extensions allow for such scalabilities at the software level, without needing hardware-specific customization. Extensive tests on a mainstream notebook computer and on OLPC's subnotebook (¿xo-laptop¿) verify that the proposed designs provide for: (i) seamless quality-complexity scalability per video frame; (ii) up to 60% increase in processing throughput with graceful degradation in output quality; (iii) up to 20% more images captured and filtered for the same power-level reduction on the xo-laptop.
Davide Anastasia, Yiannis Andreopoulos
DATE2
2010 Linear Image Processing Operations With Operational Tight Packing
abstract
Computer hardware with native support for large-bitwidth operations can be used for the concurrent calculation of multiple independent linear image processing operations when these operations map integers to integers. This is achieved by packing multiple input samples in one large-bitwidth number, performing a single operation with that number and unpacking the results. We propose an operational framework for tight packing, i.e., achieve the maximum packing possible by a certain implementation. We validate our framework on floating-point units natively supported in mainstream programmable processors. For image processing tasks where operational tight packing leads to increased packing in comparison to previously-known operational packing, the processing throughput is increased by up to 25%.
Davide Anastasia, Yiannis Andreopoulos
IEEE Signal Process. Lett.2
2010 Software Designs of Image Processing Tasks With Incremental Refinement of Computation
abstract
Software realizations of computationally-demanding image processing tasks (e.g., image transforms and convolution) do not currently provide graceful degradation when their clock-cycles budgets are reduced, e.g., when delay deadlines are imposed in a multitasking environment to meet throughput requirements. This is an important obstacle in the quest for full utilization of modern programmable platforms' capabilities since worst-case considerations must be in place for reasonable quality of results. In this paper, we propose (and make available online) platform-independent software designs performing bitplane-based computation combined with an incremental packing framework in order to realize block transforms, 2-D convolution and frame-by-frame block matching. The proposed framework realizes incremental computation: progressive processing of input-source increments improves the output quality monotonically. Comparisons with the equivalent nonincremental software realization of each algorithm reveal that, for the same precision of the result, the proposed approach can lead to comparable or faster execution, while it can be arbitrarily terminated and provide the result up to the computed precision. Application examples with region-of-interest based incremental computation, task scheduling per frame, and energy-distortion scalability verify that our proposal provides significant performance scalability with graceful degradation.
Davide Anastasia, Yiannis Andreopoulos
IEEE Trans. Image Process.2
2009 Statistical Framework for Video Decoding Complexity Modeling and Prediction
abstract
Video decoding complexity modeling and prediction is an increasingly important issue for efficient resource utilization in a variety of applications, including task scheduling, receiver-driven complexity shaping, and adaptive dynamic voltage scaling. In this paper we present a novel view of this problem based on a statistical framework perspective. We explore the statistical structure (clustering) of the execution time required by each video decoder module (entropy decoding, motion compensation, etc.) in conjunction with complexity features that are easily extractable at encoding time (representing the properties of each module's input source data). For this purpose, we employ Gaussian mixture models (GMMs) and an expectation-maximization algorithm to estimate the joint execution-time-feature probability density function (PDF). A training set of typical video sequences is used for this purpose in an offline estimation process. The obtained GMM representation is used in conjunction with the complexity features of new video sequences to predict the execution time required for the decoding of these sequences. Several prediction approaches are discussed and compared. The potential mismatch between the training set and new video content is addressed by adaptive online joint-PDF re-estimation. An experimental comparison is performed to evaluate the different approaches and compare the proposed prediction scheme with related resource prediction schemes from the literature. The usefulness of the proposed complexity-prediction approaches is demonstrated in an application of rate-distortion-complexity optimized decoding.
N. Kontorinis, Yiannis Andreopoulos, Mihaela van der Schaar
IEEE Trans. Circuits Syst. Video Technol.2
2009 Comments on "Phase-Shifting for Nonseparable 2-D Haar Wavelets"
abstract
In their recent paper, Alnasser and Foroosh derive a wavelet-domain (in-band) method for phase-shifting of 2-D "nonseparable" Haar transform coefficients. Their approach is parametrical to the (a priori known) image translation. In this correspondence, we show that the utilized transform is in fact the separable Haar discrete wavelet transform (DWT). As such, wavelet-domain phase shifting can be performed using previously-proposed phase-shifting approaches that utilize the overcomplete DWT (ODWT), if the given image translation is mapped to the phase component and in-band position within the ODWT.
Yiannis Andreopoulos
IEEE Trans. Image Process.1
2009 Erratum to "Comments on "Phase-Shifting for Nonseparable 2-D Haar Wavelets""
abstract
Typographical errors occurred in equations (2) and (3) of the above-named work. The correct forms of these equations is presented.
Yiannis Andreopoulos
IEEE Trans. Image Process.1
2008 Incremental salient point detection
abstract
In this paper, we investigate an approach that computes salient points, i.e. areas of natural images that contain corners or edges, incrementally. We focus on the popular Harris corner detector and demonstrate how such an approach can operate when the image samples are refined in a bitwise manner, i.e. the image bitplanes are received one-by-one from the image sensor. This has the advantage that the image sensing and the salient point detection can be terminated at any input image precision (e.g. at a bound set by the sensory equipment or by computation, or by the salient point accuracy required by the application) and the obtained salient points under this precision are readily available. We estimate the required energy for image sensing as well as the computation required for the salient point detection and compare them against the conventional salient point detector realization that operates directly on each source precision and cannot refine the result. Our experiments demonstrate the feasibility of incremental approaches for salient point detection in various classes of natural images. In addition, a first comparison between the results obtained by the intermediate detectors is presented.
Ioannis Patras, Yiannis Andreopoulos
ICASSP2
2008 Incremental Refinement of Image Salient-Point Detection
abstract
Low-level image analysis systems typically detect "points of interest", i.e., areas of natural images that contain corners or edges. Most of the robust and computationally efficient detectors proposed for this task use the autocorrelation matrix of the localized image derivatives. Although the performance of such detectors and their suitability for particular applications has been studied in relevant literature, their behavior under limited input source (image) precision or limited computational or energy resources is largely unknown. All existing frameworks assume that the input image is readily available for processing and that sufficient computational and energy resources exist for the completion of the result. Nevertheless, recent advances in incremental image sensors or compressed sensing, as well as the demand for low-complexity scene analysis in sensor networks now challenge these assumptions. In this paper, we investigate an approach to compute salient points of images incrementally, i.e., the salient point detector can operate with a coarsely quantized input image representation and successively refine the result (the derived salient points) as the image precision is successively refined by the sensor. This has the advantage that the image sensing and the salient point detection can be terminated at any input image precision (e.g., bound set by the sensory equipment or by computation, or by the salient point accuracy required by the application) and the obtained salient points under this precision are readily available. We focus on the popular detector proposed by Harris and Stephens and demonstrate how such an approach can operate when the image samples are refined in a bitwise manner, i.e., the image bitplanes are received one-by-one from the image sensor. We estimate the required energy for image sensing as well as the computation required for the salient point detection based on stochastic source modeling. The computation and energy required by the proposed incremental refinement approach is compared against the conventional salient-point detector realization that operates directly on each source precision and cannot refine the result. Our experiments demonstrate the feasibility of incremental approaches for salient point detection in various classes of natural images. In addition, a first comparison between the results obtained by the intermediate detectors is presented and a novel application for adaptive low-energy image sensing based on points of saliency is presented.
Yiannis Andreopoulos, Ioannis Patras
IEEE Trans. Image Process.1
2007 Analytical Complexity Modeling of Wavelet-based Video Coders
abstract
Analytical modeling for video coders can be used in a variety of scenarios where information concerning rate, distortion or complexity is essential for driving system or network interactions with media algorithms. While rate and distortion modeling have been covered extensively in previous works, complexity is not well addressed because it is highly algorithm dependent and hence difficult to model. Based on a stochastic modeling framework for the transform coefficients, we present a novel complexity analysis for state-of-the-art wavelet video coding methods by explicitly modeling several aspects found in operational coders, i.e. embedded quantization and quadtree decompositions of block significance maps. The proposed modeling derives for the first time analytical estimates of the expected number of operations (complexity) for a broad class of wavelet video coders based on stochastic source models, coding algorithm and system parameters.
Brian Foo, Yiannis Andreopoulos, Mihaela van der Schaar
ICASSP (3)2
2007 Incremental Refinement of Computation for the Discrete Wavelet Transform
abstract
Contrary to the conventional paradigm of transform decomposition followed by quantization, we investigate the computation of two-dimensional discrete wavelet transforms (DWT) under quantized representations of the input source. The proposed method builds upon previous research on approximate signal processing and revisits the concept of incremental refinement of computation: Under a refinement of the source description (with the use of an embedded quantizer), the computation of the forward and inverse transform refines the previously-computed result thereby leading to incremental computation of the output. We study for which input sources (and computational-model parameters) can the proposed framework derive identical reconstruction accuracy to the conventional approach without any incurring computational overhead. This is termed successive refinement of computation, since all representation accuracies are produced incrementally under a single (continuous) computation of the refined input source and with no overhead in comparison to the conventional calculation approach that specifically targets each accuracy level and is not refinable.
Yiannis Andreopoulos, Mihaela van der Schaar
ICIP (4)1
2007 Adaptive Linear Prediction for Resource Estimation of Video Decoding
abstract
Current systems often assume "worst case" resource utilization for the design and implementation of compression techniques and standards, thereby neglecting the fact that multimedia coding algorithms require time-varying resources, which differ significantly from the "worst case" requirements. To enable adaptive resource management for multimedia systems, resource-estimation mechanisms are needed. Previous research demonstrated that online adaptive linear prediction techniques typically exhibit superior efficiency to other alternatives for resource prediction of multimedia systems. In this paper, we formulate the problem of adaptive linear prediction of video decoding resources by analytically expressing the possible adaptation parameters for a broad class of video decoders. The resources are measured in terms of the time required for a particular operation of each decoding unit (e.g., motion compensation or entropy decoding of a video frame). Unlike prior research that mainly focuses on estimation of execution time based on previous measurements (i.e., based on autoregressive prediction or platform and decoder-specific off-line training), we propose the use of generic complexity metrics (GCMs) as the input for the adaptive predictor. GCMs represent the number of times the basic building blocks are executed by the decoder and depend on the source characteristics, decoding bit rate, and the specific algorithm implementation. Different GCM granularities (e.g., per video frame or macroblock) are explored. Our previous research indicated that GCMs can be measured or modeled at the encoder or the video server side and they can be streamed to the decoder along with the compressed bitstream. A comparison of GCM-based versus autoregressive adaptive prediction over a large range of adaptation parameters is performed. Our results indicate that GCM-based prediction is significantly superior to the autoregressive approach and also requires less computational resources at the decoder. As a result, a novel resource-prediction tradeoff is explored between: 1) the communication overhead for GCMs and/or the implementation overhead for the realization of the predictor and 2) the improvement of the prediction performance. Since this tradeoff can be significant for the decoder platform (either from the communication or the implementation perspective), we propose complexity (or communication)-bounded adaptive linear prediction in order to derive the best resource estimation under the given implementation (or GCM-communication) bound
Yiannis Andreopoulos, Mihaela van der Schaar
IEEE Trans. Circuits Syst. Video Technol.1
2007 Distortion-Driven Video Streaming over Multihop Wireless Networks with Path Diversity
abstract
Multihop networks provide a flexible infrastructure that is based on a mixture of existing access points and stations interconnected via wireless links. These networks present some unique challenges for video streaming applications due to the inherent infrastructure unreliability. In this paper, we address the problem of robust video streaming in multihop networks by relying on delay- constrained and distortion-aware scheduling, path diversity, and retransmission of important video packets over multiple links to maximize the received video quality at the destination node. To provide an analytical study of this streaming problem, we focus on an elementary multihop network topology that enables path diversity, which we term "elementary cell." Our analysis is considering several cross-layer parameters at the physical and medium access control (MAC) layers, as well as application-layer parameters such as the expected distortion reduction of each video packet and the packet scheduling via an overlay network infrastructure. In addition, we study the optimal deployment of path diversity in order to cope with link failures. The analysis is validated in each case by simulation results with the elementary cell topology, as well as with a larger multihop network topology. Based on the derived results, we are able to establish the benefits of using path diversity in video streaming over multihop networks, as well as to identify the cases where path diversity does not lead to performance improvements.
Xiaolin Tong, Yiannis Andreopoulos, Mihaela van der Schaar
IEEE Trans. Mob. Comput.2
2006 A comparison of 2-D discrete wavelet transform computation schedules on FPGAs
abstract
When it comes to the computation of the 2D discrete wavelet transform (DWT), three major computation schedules have been proposed, namely the row-column, the line-based and the block-based. In this work, the lifting-based designs of these schedules are implemented on FPGA-based platforms to execute the forward 2D DWT, and their comparison is presented. Our implementations are optimized in terms of throughput and memory requirements, in accordance with the specifications of each one of the three computation schedules and the lifting decomposition. All implementations are parameterized with respect to the image size and the number of decomposition levels. Experimental results prove that the suitability of each implementation for a particular application depends on the given specifications, concerning the throughput and the hardware cost
Maria E. Angelopoulou, Kostas Masselos, Peter Y. K. Cheung, Yiannis Andreopoulos
FPT4
2006 Scalable Resource Management for Video Streaming Over IEEE802.11A/E
abstract
Delay-constrained streaming of fully-scalable video over IEEE 802.11a/e wireless (WLANs) is of great interest for many emerging multimedia applications. In this paper, we consider the problem of video transmission over HCF controlled channel access (HCCA), which is part of the new medium access control (MAC) protocol of IEEE 802.11e. A cross-layer optimization across the MAC and application layers is used in order to exploit the features provided by the new HCCA standard, as well as by the versatility of new state-of-the-art scalable video coding algorithms. Under pre-determined delay constraints for streaming, the proposed cross-layer strategy leads to a larger number of stations being simultaneously admitted (without any loss in the video quality) than in systems that utilize application-layer only optimizations. At the same time, the fine-grain layering provided by the scalable bitstream facilitates prioritization and unequal retransmissions of packets at the MAC layer thereby enabling graceful quality degradation under channel-capacity limitations and delay constraints. The expected gains offered by the optimized solutions proposed in this paper are established through simulations
Yiannis Andreopoulos, Mihaela van der Schaar, Zhiping Hu, S. Heo, S. Suh
ICASSP (5)1
2006 Cross-layer Video Streaming Over 802.11e-Enabled Wireless Mesh Networks
abstract
We propose an integrated cross-layer optimization algorithm for maximizing the decoded video quality of delay-constrained streaming in a quality-of-service (QoS) enabled multi-hop wireless mesh network. The key to our algorithm is the synergistic optimization of control parameters at each node of the multi-hop network, across the protocol layers - application, network, medium access control (MAC) and physical (PHY) layers, as well as end-to-end, i.e. across the various network nodes. To drive this optimization, we assume an overlay network infrastructure, which conveys information on the conditions of each link. Quantitative results are presented that demonstrate the merits and the need for cross-layer optimization in an efficient solution for real-time video transmission using existing protocols and infrastructures
Nicholas Mastronarde, Yiannis Andreopoulos, Mihaela van der Schaar, Dilip Krishnaswamy, John B. Vicente
ICASSP (5)2
2006 Execution time comparison of lifting-based 2D wavelet transforms implementations on a VLIW DSP
abstract
Several input-traversal schedules have been proposed for the computation of the 2D discrete wavelet transform (DWT). In this paper, the row-column, the line-based and the block-based schedules for the 2D DWT computation are compared with respect to their execution time on a very long instruction word (VLIW) digital signal processor (DSP). Implementations of the wavelet transform according to the considered schedules have been developed. They are parameterized with respect to filter pair, image size, and number of decomposition levels. All implementations have been mapped on a VLIW DSP. Performance metrics for the implementations for a complete set of parameters have been obtained and compared. The experimental results show that each implementation performs better for different points of the parameter space
Kostas Masselos, Yiannis Andreopoulos, Thanos Stouraitis
ISCAS2
2006 Cross-Layer Optimized Video Streaming Over Wireless Multihop Mesh Networks
abstract
The proliferation of wireless multihop communication infrastructures in office or residential environments depends on their ability to support a variety of emerging applications requiring real-time video transmission between stations located across the network. We propose an integrated cross-layer optimization algorithm aimed at maximizing the decoded video quality of delay-constrained streaming in a multihop wireless mesh network that supports quality-of-service. The key principle of our algorithm lays in the synergistic optimization of different control parameters at each node of the multihop network, across the protocol layers-application, network, medium access control, and physical layers, as well as end-to-end, across the various nodes. To drive this optimization, we assume an overlay network infrastructure, which is able to convey information on the conditions of each link. Various scenarios that perform the integrated optimization using different levels ("horizons") of information about the network status are examined. The differences between several optimization scenarios in terms of decoded video quality and required streaming complexity are quantified. Our results demonstrate the merits and the need for cross-layer optimization in order to provide an efficient solution for real-time video transmission using existing protocols and infrastructures. In addition, they provide important insights for future protocol and system design targeted at enhanced video streaming support across wireless mesh networks
Yiannis Andreopoulos, Nicholas Mastronarde, Mihaela van der Schaar
IEEE J. Sel. Areas Commun.1
2006 Optimized Scalable Video Streaming over IEEE 802.11a/e HCCA Wireless Networks under Delay Constraints
abstract
The quality-of-service (QoS) guarantees enabled by the new IEEE 802.11 a/e Wireless LAN (WLAN) standard are specifically targeting the real-time transmission of multimedia content over the wireless medium. Since video data consume the largest part of the available bitrate compared to other media, optimization of video streaming for this new standard is a significant factor for the successful deployment of practical systems. Delay-constrained streaming of fully-scalable video over IEEE 802.11 a/e WLANs is of great interest for many multimedia applications. The new medium access control (MAC) protocol of IEEE 802.11e is called the Hybrid Coordination Function (HCF) and, in this paper, we will specifically consider the problem of video transmission over HCF Controlled Channel Access (HCCA). A cross-layer optimization across the MAC and application layers of the OSI stack is used in order to exploit the features provided by the combination of the new HCCA standard with new versatile scalable video coding algorithms. Specifically, we propose an optimized and scalable HCCA-based admission control for delay-constrained video streaming applications that leads to a larger number of stations being simultaneously admitted (without quality reduction to any video flow). Subsequently, given the allocated transmission opportunity, each station deploys an optimized Application-MAC-PHY adaptation, scheduling, and protection strategy that is facilitated by the fine-grain layering provided by the scalable bitstream. Given that each video flow needs to always comply with the predetermined (a priori negotiated) traffic specification parameters, this cross-layer strategy enables graceful quality degradation whenever the channel conditions or the video sequence characteristics change. For instance, it is demonstrated that the proposed cross-layer protection and bitstream adaptation strategies facilitate QoS token rate adaptation under link adaptation mechanisms that utilize different physical layer transmission rates. The expected gains offered by the optimized solutions proposed in this paper are established theoretically, as well as through simulations.
Mihaela van der Schaar, Yiannis Andreopoulos, Zhiping Hu
IEEE Trans. Mob. Comput.2
2006 Failure-Aware, Open-Loop, Adaptive Video Streaming With Packet-Level Optimized Redundancy
abstract
A plethora of coding and streaming mechanisms have been proposed for real-time multimedia transmission over the Internet. However, most proposed mechanisms rely only on global (e.g. based on end-to-end measurements), delayed (at least by the round-trip-time), or statistical (often based on simplistic network models) information available about the network state. Based on recently-proposed state-of-the-art open-loop video coding schemes, we propose a new integrated streaming and routing framework for robust and efficient video transmission over networks exhibiting path failures. Our approach explicitly takes into account the network dynamics, path diversity, and the modeled video distortion at the receiver side to optimize the packet redundancy and scheduling. In the derived framework, multimedia streams can be adapted dynamically at the video server based on instantaneous routing-layer information or failure-modeling statistics. The performance of our integrated application and network-layer method is simulated against equivalent approaches that are not optimized based on routing-layer feedback and distortion modeling, and the obtained gains in video quality are quantified
Yiannis Andreopoulos, Ram Keralapura, Mihaela van der Schaar, Chen-Nee Chuah
IEEE Trans. Multim.1
2005 Single-rate calculation of overcomplete discrete wavelet transforms for scalable coding applications
Yiannis Andreopoulos, Adrian Munteanu 0001, Geert Van der Auwera, Jan Cornelis 0001, Peter Schelkens
Signal Process.1
2005 Motion and texture rate-allocation for prediction-based scalable motion-vector coding
Joeri Barbarien, Adrian Munteanu 0001, Fabio Verdicchio, Yiannis Andreopoulos, Jan Cornelis 0001, Peter Schelkens
Signal Process. Image Commun.4
2005 Unconstrained motion compensated temporal filtering (UMCTF) for efficient and flexible interframe wavelet video coding
Deepak S. Turaga, Mihaela van der Schaar, Yiannis Andreopoulos, Adrian Munteanu 0001, Peter Schelkens
Signal Process. Image Commun.3
2005 Rate-distortion-complexity modeling for network and receiver aware adaptation
abstract
Existing research on Universal Multimedia Access has mainly focused on adapting multimedia to the network characteristics while overlooking the receiver capabilities. Alternatively, part 7 of the MPEG-21 standard entitled Digital Item Adaptation (DIA) defines description tools to guide the multimedia adaptation process based on both the network conditions and the available receiver resources. In this paper, we propose a new and generic rate-distortion-complexity model that can generate such DIA descriptions for image and video decoding algorithms running on various hardware architectures. The novelty of our approach is in virtualizing complexity, i.e., we explicitly model the complexity involved in decoding a bitstream by a generic receiver. This generic complexity is translated dynamically into "real" complexity, which is architecture-specific. The receivers can then negotiate with the media server/proxy the transmission of a bitstream having a desired complexity level based on their resource constraints. Hence, unlike in previous streaming systems, multimedia transmission can be optimized in an integrated rate-distortion-complexity setting by minimizing the incurred distortion under joint rate-complexity constraints.
Mihaela van der Schaar, Yiannis Andreopoulos
IEEE Trans. Multim.2
2004 Real-time ubiquitous multimedia streaming using rate-distortion-complexity models
abstract
We discuss the benefits of using realistic rate-distortion-complexity (R-D-C) models to guide ubiquitous multimedia streaming systems. At the server or proxy, on-the-fly bitstream adaptation to varying network conditions and diverse end-devices' processing capabilities can be performed by relying on accurate (pre-computed) R-D-C models. To accommodate various types of receivers with different resources, we introduce the concept of generic complexity metrics (GCMs) which quantify in a generic manner the decoding complexity of a compressed video sequence. A state-of-the-art motion-compensated wavelet-based coding scheme is used to illustrate how R-D-C models can be computed for a particular compression scheme. Finally, we propose a receiver-driven streaming framework and show how R-D-C models can be used to assist real-time wired/wireless multimedia transmission applications.
Mihaela van der Schaar, Yiannis Andreopoulos
GLOBECOM2
2004 Scalable motion vector coding
abstract
Recently proposed scalable wavelet-based video codecs using spatial-domain motion compensated temporal filtering (SDMCTF) offer competitive compression performance when compared to H.264 and generate embedded bit-streams supporting quality, resolution and temporal scalability. To be able to support a large range of bit-rates with optimal compression efficiency, these codecs require a quality-scalable motion vector coding technique. Such an algorithm based on the integer wavelet transform followed by embedded coding of the wavelet coefficients was proposed in the recent past. In this paper, we present a quality-scalable motion vector coding algorithm using median-based motion vector prediction. The compression performance of the proposed algorithm is compared to that of the wavelet-based technique and is found to be superior. Additionally, the proposed motion vector codec is incorporated into an SDMCTF-based video codec and the benefits of using quality-scalable motion vector representations are experimentally demonstrated.
Joeri Barbarien, Adrian Munteanu 0001, Fabio Verdicchio, Yiannis Andreopoulos, Jan Cornelis 0001, Peter Schelkens
ICIP4
2004 Scalable video coding based on motion-compensated temporal filtering: complexity and functionality analysis
abstract
Video coding techniques yielding state-of-the-art compression performance require large amount of computational resources, hence practical implementations, which target a broad market, often tend to trade-off coding efficiency and flexibility for reduced complexity. Scalable video coding instead, not only provides seamless adaptation to bit-rate variation, but also allows the end user to trim down the resources he needs to perform real-time decoding by limiting the process to a subset of the original content. Hence, by choosing the quality, frame-rate and/or resolution of the reconstructed sequence, each decoder can meet its hardware limitations without affecting the encoding process of the media provider. This paper proposes a preliminary analysis of the memory-access behavior of a fully scalable video decoder and investigates the capability of selecting the operational settings in order to adapt to the available hardware resources on the target device.
Fabio Verdicchio, Yiannis Andreopoulos, Tom Clerckx, Joeri Barbarien, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens
ICIP2
2004 In-band motion compensated temporal filtering
Yiannis Andreopoulos, Adrian Munteanu 0001, Joeri Barbarien, Mihaela van der Schaar, Jan Cornelis 0001, Peter Schelkens
Signal Process. Image Commun.1
2003 Fully-scalable wavelet video coding using in-band motion compensated temporal filtering
abstract
This paper presents a novel fully-scalable wavelet video coding scheme that performs efficient open-loop motion compensated temporal filtering (MCTF) in the wavelet domain (in-band). Unlike the conventional spatial-domain MCTF (SDMCTF) schemes, which apply MCTF on the original image data and then encode the residual image using a critically-sampled wavelet transform, the framework presented here applies the in-band MCTF (IBMCTF) after the discrete wavelet transform (DWT) is performed in the spatial dimensions. To overcome the inefficiency of motion estimation (ME) in the wavelet domain, a complete-to-overcomplete DWT (CODWT) is performed. The proposed framework provides improved quality (SNR) and temporal scalability as compared with existing in-band closed-loop temporal prediction schemes with ODWT and improved spatial scalability as compared to SDMCTF. We present a thorough comparison between SDMCTF and the proposed IBMCTF in terms of coding efficiency and scalability. Furthermore, we describe several extensions that enable the filtering of the various bands to be performed independently, based on the resolution, sequence content, complexity requirements and desired scalability.
Yiannis Andreopoulos, Mihaela van der Schaar, Adrian Munteanu 0001, Joeri Barbarien, Peter Schelkens, Jan Cornelis 0001
ICASSP (3)1
2003 Spatio-temporal-SNR scalable wavelet coding with motion-compensated DCT base-layer architectures
abstract
It has been demonstrated recently that 3-D wavelet coding with motion-compensated temporal filtering (MCTF) provides a wide range of spatio-temporal-SNR scalability with state-of-the-art coding performance. However, the coding system is very different from the already standardized motion-compensated DCT (MC-DCT) video coders such as MPEG-2, MPEG-4, or H.26L. Nonetheless, the market acceptance of the new scalable technology will come much easier if backward compatibility to such previous standards would exist. In this paper, we present a new coder architecture where the base layer can be a standard MC-DCT coder, while the enhancement layer is an in-band MCTF codec operating in the overcomplete wavelet domain. First, we propose a simple extension of the scalable 3-D wavelet codec with a standard MC-DCT base layer. As a second step, to improve the performance of the proposed scalable coder over a wide range of bit-rates, we describe several small extensions to the standardized MC-DCT codec structure that can be applied for the base layer coding.
Yiannis Andreopoulos, Mihaela van der Schaar, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001
ICIP (2)1
2003 Motion vector coding for in-band motion compensated temporal filtering
abstract
Recently, a new wavelet-based video codec using in-band motion compensated temporal filtering (IBMCTF) was introduced This codec is fully scalable in resolution, quality aid frame-rate. In comparison to an equivalent video coding scheme based on spatial domain motion compensated temporal filtering (SDMCTF), its compression performance when decoding to lower resolutions is very promising. However, since the IBMCTF scheme is based on in-band motion estimation, considerably more motion vector data is generated than in the SDMCTF scheme. Efficient compression of these motion vectors is therefore of utmost importance. In this paper, several solutions for the compression of motion vectors generated by a video codec based on IBMCTF are presented and compared.
Joeri Barbarien, Yiannis Andreopoulos, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001
ICIP (2)2
2003 Control of the distortion variation in video coding systems based on motion compensated temporal filtering
abstract
The paper proposes a new framework for the control of the distortion variation in video coding schemes based on motion-compensated temporal filtering (MCTF). The distortion in an arbitrary decoded frame at any temporal level in the MCTF pyramid is expressed as a function of the distortions in the reference frames at the same temporal level. The approach is formulated for the bi-directional unconstrained MCTF (UMCTF) scheme of Turaga et al., (2002), which does not include the update-lifting step. The proposed framework can be extended to the generalized form of MCTF by utilizing additional control parameters. Experimental results demonstrate the control of the distortion variation in video coding systems based on spatial-domain and wavelet-domain MCTF. One concludes that the proposed framework provides the means of controlling the tradeoff between the average distortion and the distortion variation in each group-of-pictures (GOPs) within the decoded sequence.
Adrian Munteanu 0001, Yiannis Andreopoulos, Mihaela van der Schaar, Peter Schelkens, Jan Cornelis 0001
ICIP (2)2
2003 Complete-to-overcomplete discrete wavelet transforms for scalable video coding with MCTF
Yiannis Andreopoulos, Mihaela van der Schaar, Adrian Munteanu 0001, Joeri Barbarien, Peter Schelkens, Jan Cornelis 0001
VCIP1
2002 Scalable wavelet video-coding with in-band prediction - implementation and experimental results
abstract
In this paper we elaborate on a recently proposed approach for scalable video-coding based on in-band prediction in the overcomplete wavelet domain. It is shown that, through a new calculation scheme for the level-by-level complete to overcomplete discrete wavelet transform (DWT) that exploits certain symmetries, important reductions in the multiplication budget are obtained in comparison to the fastest-known algorithm of the literature. Based on the derived overcomplete transform-domain coefficients, a pixel-accurate motion estimation and compensation (ME/MC) algorithm is proposed, which provides a hybrid-coding framework that supports full scalability in resolution, quality and frame rate. To give an indication of the coding performance of such a system, some preliminary results are reported.
Yiannis Andreopoulos, Geert Van der Auwera, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001
ICIP (3)1
2001 A wavelet-tree image coding system with efficient memory utilization
abstract
This paper describes an efficient implementation of an image coding system based on the independent wavelet-tree coding concept. The system consists of a transform and a (de)coding engine that operate in a pipelined fashion. The main focus of this paper is on the encoding part since, due to the system architecture, the decoder has identical memory utilization. Experimental results prove that the proposed system achieves comparable coding performance to the state-of-the-art, while it localizes the memory accesses to small memory modules and uses minimal computational resources.
Yiannis Andreopoulos, Peter Schelkens, Nikolaos D. Zervas, Thanos Stouraitis, Constantinos E. Goutis, Jan Cornelis 0001
ICASSP1
2001 Quantization effect on VLSI implementations for the 9/7 DWT filters
abstract
Two basic approaches for implementing the 9/7 filtering unit, used in the discrete wavelet transform, are addressed. The first is the lifting scheme approach and the second is the conventional, convolutional filter approach. Two architectures are examined for each approach, a simple straightforward one and an optimized one, substituting the multipliers used for scaling with shift-add operations. The quantization of the constants used in the calculations is thoroughly explored and the selection of the data-path bit-width is addressed. Experimental results based on hardware implementation, for several quantizations and for the different hardware architectures of the 9/7 filtering units are given.
Vassilis Spiliotopoulos, Nikolaos D. Zervas, Yiannis Andreopoulos, Giorgos P. Anagnostopoulos, Constantinos E. Goutis
ICASSP3
2001 A local wavelet transform implementation versus an optimal row-column algorithm for the 2D multilevel decomposition
abstract
A new method for the implementation of the binary-tree decomposition of the convolution-based wavelet transform, called the local wavelet transform (LWT) has been recently proposed in the literature. While it produces exactly the same results as the classical row-column implementation of the transform, it has many implementation benefits. This fact is shown experimentally for the first time for a general-purpose processor-based architecture, by comparing our C implementation of the LWT with an optimal C implementation of the lifting-scheme row-column algorithm. The comparisons are made for the forward multilevel binary-tree decomposition using the 9/7 filter pair, in the typical Intel Pentium processor family.
Yiannis Andreopoulos, Nikolaos D. Zervas, Gauthier Lafruit, Peter Schelkens, Thanos Stouraitis, Constantinos E. Goutis, Jan Cornelis 0001
ICIP (3)1
2001 Evaluation of design alternatives for the 2-D-discrete wavelet transform
abstract
In this paper, the three main hardware architectures for the 2-D discrete wavelet transform (2-D-DWT) are reviewed. Also, optimization techniques applicable to all three architectures are described. The main contribution of this work is the quantitative comparison among these design alternatives for the 2-D-DWT. The comparison is performed in terms of memory requirements, throughput, and energy dissipation, and is based on a theoretical analysis of the alternative architectures and schedules. Memory requirements, throughput, and energy are expressed by analytical equations with parameters from both the 2-D-DWT algorithm and the implementation platform. The parameterized equations enable the early but efficient exploration of the various tradeoffs related to the selection to the one or the other architecture.
Nikolaos D. Zervas, Giorgos P. Anagnostopoulos, Vassilis Spiliotopoulos, Yiannis Andreopoulos, Constantinos E. Goutis
IEEE Trans. Circuits Syst. Video Technol.4