VLDB 2026 Research / reviewers in the wild / expert
Madhukar Budagavi
dblp:89/3067
· DBLP profile ↗
36ranked-venue papers
11as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 10 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Guided Detail Filter for AVMabstractThis paper proposes an additional filter — that aims at enhancing the fidelity of local details in reconstructed frames and consequently improving coding efficiency — to the in-loop filtering process in the AOM next generation video codec (AVM). This involves correcting each sample through a rectification (i.e., scaling and clipping) process applied to the expected coding error obtained via a multidimensional table look-up with respect to a query vector that is derived by passing the sample’s spatial neighborhood and its intensity gradients through a multichannel filter. Guiding information comprising (multichannel) linear kernels and the table of expected coding errors is chosen according to edge classification and coding settings for optimal adaptation to local sample activity and coding context. Experimental results demonstrate that our approach achieves significant improvements over AVM v7.0.1, with PSNR-Y bitrate saving of -0.57% and -0.52% for All Intra and Random Access configurations respectively. Khanh Quoc Dinh, Yangwoo Kim, Madhukar Budagavi, Rajan Joshi, Kwangpyo Choi |
ICIP | 3 |
| 2023 | 3D Talking Face With Personalized Pose DynamicsabstractRecently, we have witnessed a boom in applications for 3D talking face generation. However, most existing 3D face generation methods can only generate 3D faces with a static head pose, which is inconsistent with how humans perceive faces. Only a few articles focus on head pose generation, but even these ignore the attribute of personality. In this article, we propose a unified audio-driven approach to endow 3D talking faces with personalized pose dynamics. To achieve this goal, we establish an original person-specific dataset, providing corresponding head poses and face shapes for each video. Our framework is composed of two separate modules: PoseGAN and PGFace. Given an input audio, PoseGAN first produces a head pose sequence for the 3D head, and then, PGFace utilizes the audio and pose information to generate natural face models. With the combination of these two parts, a 3D talking head with dynamic head movement can be constructed. Experimental evidence indicates that our method can generate person-specific head pose sequences that are in sync with the input audio and that best match with the human experience of talking heads. Saifeng Ni, Zhipeng Fan 0001, Ming Zeng 0008, Madhukar Budagavi, Xiaohu Guo |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Layered-Garment Net: Generating Multiple Implicit Garment Layers from a Single Image
Alakh Aggarwal, Steven Hogue, Saifeng Ni, Madhukar Budagavi, Xiaohu Guo |
ACCV (1) | 5 |
| 2021 | FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute LearningabstractIn this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and eye blinks that are in-sync with the input audio signal. We note that the synthetic face attributes include not only explicit ones such as lip motions that have high correlations with speech, but also implicit ones such as head poses and eye blinks that have only weak correlation with the input audio. To model such complicated relationships among different face attributes with input audio, we propose a FACe Implicit Attribute Learning Generative Adversarial Network (FACIAL-GAN), which integrates the phonetics-aware, context-aware, and identity-aware information to synthesize the 3D face animation with realistic motions of lips, head poses, and eye blinks. Then, our Rendering-to-Video network takes the rendered face images and the attention map of eye blinks as input to generate the photorealistic output video frames. Experimental results and user studies show our method can generate realistic talking face videos with not only synchronized lip motions, but also natural head movements and eye blinks, with better qualities than the results of state-of-the-art methods. Yifan Zhao 0002, Ming Zeng 0008, Saifeng Ni, Madhukar Budagavi, Xiaohu Guo |
ICCV | 6 |
| 2021 | Predictive Adaptive Streaming to Enable Mobile 360-Degree and VR ExperiencesabstractAs 360-degree videos and virtual reality (VR) applications become popular for consumer and enterprise use cases, the desire to enable truly mobile experiences also increases. Delivering 360-degree videos and cloud/edge-based VR applications require ultra-high bandwidth and ultra-low latency[1], challenging to achieve with mobile networks. A common approach to reduce bandwidth is streaming only the field of view (FOV). However, extracting and transmitting the FOV in response to user head motion can add high latency, adversely affecting user experience. In this paper, we propose a predictive adaptive streaming approach, where the predicted view with high predictive probability is adaptively encoded in relatively high quality according to bandwidth conditions and transmitted in advance, leading to a simultaneous reduction in bandwidth and latency. The predictive adaptive streaming method is based on a deep-learning-based viewpoint prediction model we develop, which uses past head motions to predict where a user will be looking in the 360-degree view. Using a very large dataset consisting of head motion traces from over 36,000 viewers for nineteen 360-degree/VR videos, we validate the ability of our predictive adaptive streaming method to offer high-quality view while simultaneously significantly reducing bandwidth. Xueshi Hou, Sujit Dey, Jianzhong Zhang 0002, Madhukar Budagavi |
IEEE Trans. Multim. | 4 |
| 2020 | Mesh Coding Extensions to MPEG-I V-PCCabstractDynamic point clouds and meshes are used in a wide variety of applications such as gaming, visualization, medicine, and more recently AR/VR/MR. This paper presents two extensions of MPEG-I Video-based Point Cloud Compression (V-PCC) standard to support mesh coding. The extensions are based on Edgebreaker and TFAN mesh connectivity coding algorithms implemented in the Google Draco software and the MPEG SC3DMC software for mesh coding, respectively. Lossless results for the proposed frameworks on top of version 8.0 of the MPEG-I V-PCC test model (TMC2) are presented and compared with Draco for dense meshes. Esmaeil Faramarzi, Rajan Joshi, Madhukar Budagavi |
MMSP | 3 |
| 2019 | Head and Body Motion Prediction to Enable Mobile VR Experiences with Low LatencyabstractAs virtual reality (VR) applications become popular, the desire to enable high-quality, lightweight and mobile VR leads to various edge/cloud-based techniques. This paper introduces a predictive pre-rendering approach to address the ultra-low latency challenge in edge/cloud-based six Degrees of Freedom (6DoF) VR. Compared to 360-degree videos and 3DoF (head motion only) VR, 6DoF VR supports both head and body motions, thus not only viewing direction, but also viewing position changes. In our approach, the predictive view is rendered in advance based on the predicted viewing direction and position, leading to a reduction in latency. The key to achieving this efficient predictive pre-rendering approach is to predict the head and body motion accurately using past head and body motion traces. We develop a deep learning-based model and validate its ability using a dataset of over 840,000 samples for head and body motion. Xueshi Hou, Jianzhong Zhang 0002, Madhukar Budagavi, Sujit Dey |
GLOBECOM | 3 |
| 2017 | Dual-fisheye lens stitching for 360-degree imagingabstractDual-fisheye lens cameras have been increasingly used for 360-degree immersive imaging. However, the limited overlapping field of views and misalignment between the two lenses give rise to visible discontinuities in the stitching boundaries. This paper introduces a novel method for dual-fisheye camera stitching that adaptively minimizes the discontinuities in the overlapping regions to generate full spherical 360-degree images. Results show that this approach can produce good quality stitched images for Samsung Gear 360 - a dual-fisheye camera, even with hard-to-stitch objects in the stitching borders. Tuan Ho, Madhukar Budagavi |
ICASSP | 2 |
| 2017 | 360-degree video stitching for dual-fisheye lens cameras based on rigid moving least squaresabstractDual-fisheye lens cameras are becoming popular for 360-degree video capture, especially for User-generated content (UGC), since they are affordable and portable. Images generated by the dual-fisheye cameras have limited overlap and hence require non-conventional stitching techniques to produce high-quality 360×180-degree panoramas. This paper introduces a novel method to align these images using interpolation grids based on rigid moving least squares. Furthermore, jitter is the critical issue arising when one applies the image-based stitching algorithms to video. It stems from the unconstrained movement of stitching boundary from one frame to another. Therefore, we also propose a new algorithm to maintain the temporal coherence of stitching boundary to provide jitter-free 360-degree videos. Results show that the method proposed in this paper can produce higher quality stitched images and videos than prior work. Tuan Ho, Ioannis D. Schizas, Kamisetty Ramamohan Rao, Madhukar Budagavi |
ICIP | 4 |
| 2017 | Projection based advanced motion model for cubic mapping for 360-degree videoabstractThis paper proposes a novel advanced motion model to handle the irregular motion for the cubic map projection of 360-degree video. Since the irregular motion is mainly caused by the projection from the sphere to the cube map, we first try to project the pixels in both the current picture and reference picture from unfolding cube back to the sphere. Then through utilizing the characteristic that most of the motions in the sphere are uniform, we can derive the relationship between the motion vectors of various pixels in the unfold cube. The proposed advanced motion model is implemented in the High Efficiency Video Coding reference software. Experimental results demonstrate that quite obvious performance improvement can be achieved for the sequences with obvious motions. Li Li 0040, Zhu Li 0001, Madhukar Budagavi, Houqiang Li |
ICIP | 3 |
| 2017 | VR+HDR: A system for view-dependent rendering of HDR video in virtual realityabstractThis paper introduces a view-dependent method for rendering high-dynamic-range (HDR) video in virtual reality (VR) on VR displays such as head mounted displays (HMD), mobile phones, TVs, and computer monitors. The user's view direction is taken into account to design a tone-mapping operator which appropriately displays the HDR content on the display device. The proposed method can be utilized if HDR capturing and playback (i.e. HDR camera and 10-bit video codec) are available. However, it can also be used on the 8-bit pipeline (i.e. 8-bit camera and 8-bit video codec). A VR+HDR prototype was implemented using Samsung Gear360 camera and SamsungVR Android app on Samsung Galaxy Note 5 smartphone. The subjective comparison between the proposed method and rendering of 360-degree VR video through global tone mapping was performed using this prototype and showed that for all test sequences, the proposed technique significantly improves the video quality in VR rendering. Hossein Najaf-Zadeh, Madhukar Budagavi, Esmaeil Faramarzi |
ICIP | 2 |
| 2016 | Introduction to the Special Issue on HEVC Extensions and Efficient HEVC ImplementationsabstractHigh Efficiency Video Coding (HEVC) is the most recent standard in the series of major video coding standards jointly produced by the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG). HEVC was first approved in 2013 in the ITU-T as Recommendation H.265 and in ISO/IEC as International Standard 23008-2, and it offers an unprecedented degree of compression capability for a very wide variety of applications. In the three years since its initial completion, it has been extended in several important ways to further broaden its scope. This special issue on HEVC features two sections: 1) HEVC extensions and 2) efficient HEVC implementations. Jens-Rainer Ohm, Gary J. Sullivan, Vivienne Sze, Thomas Wiegand 0001, Madhukar Budagavi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Improving Intra Prediction in High-Efficiency Video CodingabstractIntra prediction is an important tool in intra-frame video coding to reduce the spatial redundancy. In current coding standard H.265/high-efficiency video coding (HEVC), a copying-based method based on the boundary (or interpolated boundary) reference pixels is used to predict each pixel in the coding block to remove the spatial redundancy. We find that the conventional copying-based method can be further improved in two cases: 1) the boundary has an inhomogeneous region and 2) the predicted pixel is far away from the boundary that the correlation between the predicted pixel and the reference pixels is relatively weak. This paper performs a theoretical analysis of the optimal weights based on a first-order Gaussian Markov model and the effects when the pixel values deviate from the model and the predicted pixel is far away from the reference pixels. It also proposes a novel intra prediction scheme based on the analysis that smoothing the copying-based prediction can derive a better prediction block. Both the theoretical analysis and the experimental results show the effectiveness of the proposed intra prediction method. An average gain of 2.3% on all intra coding can be achieved with the HEVC reference software. Haoming Chen, Tao Zhang 0013, Ming-Ting Sun, Ankur Saxena, Madhukar Budagavi |
IEEE Trans. Image Process. | 5 |
| 2015 | 360 degrees video coding using region adaptive smoothingabstract360 degrees video is emerging as a new way of experiencing immersive video due to the ready availability of powerful handheld devices such as smartphones. 360 degrees video enables immersive “real life”, “being there” experience for consumers by capturing the 360 degree view of the world. Users can change their viewpoint and dynamically view any part of the captured scene they desire. 360 degrees video requires higher bitrate than conventional video due to increased video resolution (4K and beyond) needed to support the wider field of view. Efficient compression of 360 degrees video is thus needed to provide high quality immersive video viewing experience to consumers. This paper provides an overview of 360 degrees video and presents a region adaptive video smoothing technique that exploits the unique characteristics of 360 degrees video to provide up to 20% bitrate savings with minimal perceptual quality degradation while requiring no modifications to the video decoder. Madhukar Budagavi, John Furton, Guoxin Jin, Ankur Saxena, Jeffrey Wilkinson, Andrew Dickerson |
ICIP | 1 |
| 2015 | Motion estimation and compensation for fisheye warped videoabstractVideo captured through fisheye lens is becoming very prevalent due to applications such as surveillance, automotive driver assistance systems, and recreational sports. Virtual reality applications such as 360 degrees video also use fish-eye lens video cameras to capture the 360 degree view. The main advantage of fisheye lenses is that they increases the field of view, however they also introduce warping distortion in the captured video. Due to the warping, the motion in the video is typically non-translational and traditional block motion compensation techniques are not fully effective for such video. This paper presents a warping motion compensation technique that models the fish-eye lens distortion to efficiently code warped video. Given the fisheye lens parameters (e.g., transmitted at a video sequence level), the warping model is implicitly derived in both the encoder and decoder and no block-level warping parameters are transmitted. The proposed approach was integrated into HEVC HM-14.0 and initial results on simulated fish-eye lens distorted video show promise especially when the global motion in the video is fast. Guoxin Jin, Ankur Saxena, Madhukar Budagavi |
ICIP | 3 |
| 2015 | Low-complexity separable multiplier-less loop filter for video codingabstractIn this paper, we present a low-complexity loop filter for video coding. We begin by presenting a set of non-Wiener based loop filters that can complement the Wiener-based adaptive loop filter which was considered as a tool for possible adoption in the HEVC standard. We devise filters which can be adaptively operated on different regions in an image. We devise a quad-tree based signaling for the filters, and present various loop filters: such as bilateral and Gaussian, as well as a separable 3-tap filter which can be implemented by just shifts and adds. The proposed 3-tap filter is thus hardware friendly with minimal complexity. In terms of compression performance, the proposed 3-tap filter can approach other sophisticated filters, albeit at a substantially reduced complexity; and can provide compression gains of 2.3% on average; and upto 7.0% for Low-Delay-P configuration over a data-set of 19 diverse HD, and UHD sequences of upto 8K resolution when implemented on top of the HM14.0 software for HEVC. Finally, we also present results for combining our proposed filters with the Wiener-based adaptive loop filter considered in HEVC, and illustrate that there is a significant amount of compression gain that can be achieved by loop filters for the next generation of video coding standard beyond HEVC. Ankur Saxena, Mohammed A. Aabed, Madhukar Budagavi |
ICIP | 3 |
| 2015 | Improvements on Intra Block Copy in natural content video codingabstractThe Intra Block Copy (IntraBC) is a newly adopted tool in the HEVC extension for the screen content video coding. The IntraBC tool efficiently encodes repeating patterns in a picture. The current IntraBC scheme achieves about 1.0% bit-rate reduction on average and up to 4.3% bitrate reduction on natural content video for a database consisting of 2K, 4K, and 8K sequences. In this paper, we propose to improve the IntraBC with a template matching block vector and a fractional search IntraBC. With these two tools, the gain on natural content video coding can be further improved by 0.5% on average and up to 2.0%. Haoming Chen, Yu-Sheng Chen, Ming-Ting Sun, Ankur Saxena, Madhukar Budagavi |
ISCAS | 5 |
| 2014 | Fast intra block copy (IntraBC) search for HEVC screen content codingabstractCoding of screen content video is becoming important because of applications such as wireless displays, remote desktop, remote gaming, automotive infotainment, cloud computing, distance education etc. Video in these applications often has mixed content consisting of natural video, text and graphics in the same picture. Intra block copy (IntraBC) is a new coding tool being studied for the HEVC Range Extensions (RExt) standard. IntraBC is a block matching technique where in a Coding Unit (CU) is predicted as a displacement from already reconstructed block of samples from neighboring region in the same picture. IntraBC is very effective for screen content video since it removes redundancy from repeating patterns which typically occur in text and graphics regions. IntraBC provides a significant bit rate savings (up to 44%) for screen content sequences but at the cost of encoding time increase because of the search involved in Intra block matching. This paper presents fast encoding techniques for early skipping of IntraBC search which result in about 21% - 24% encoding time reduction for Intra coding. Do-Kyoung Kwon, Madhukar Budagavi |
ISCAS | 2 |
| 2013 | Multi-loop scalable video codec based on high efficiency video coding (HEVC)abstractA multi-loop scalable video coder for high efficiency video coding (HEVC) is proposed in this paper. A coding unit (CU)-level inter-layer sample prediction tool is proposed to exploit redundancy between enhancement-layer and up-sampled base-layer pictures. To reduce decoded picture buffer size and memory bandwidth in a multi-loop decoder, a hierarchical inter-layer prediction tool is proposed as well using two picture-level flags. The proposed solution requires minimum amount of changes relative to the single-layer HEVC codec to support HEVC scalable coding, and provides a good complexity and coding efficiency trade-off as revealed by the experimental results. Do-Kyoung Kwon, Madhukar Budagavi, Minhua Zhou |
ICASSP | 2 |
| 2013 | Intra motion compensation and entropy coding improvements for HEVC screen content codingabstractApplications such as wireless displays, automotive infotainment, remote desktop, remote gaming, distance education, cloud computing etc. are becoming popular. Video in these applications often has mixed content consisting of natural video and screen content (text, graphics etc.) in the same picture. In text and graphics regions, patterns such as text characters, icons, lines etc. can repeat within a picture. Also since the graphics and text regions have sharp edges that are sometimes not predicted well by Intra prediction tools, the probability of prediction error having high amplitude increases. This paper presents two tools for improving the intra coding efficiency of the High Efficiency Video Coding (HEVC) standard when coding screen content. The first tool is a Coding Unit (CU)-level Intra motion compensation tool to eliminate redundancy from repeating patterns in text and graphics regions. The second tool is a modification to HEVC cRiceParam update process of coeff_abs_level_remaining entropy coding to adapt better to larger prediction errors in text and graphics regions. The combination of the two tools achieves an average bit-rate/Bjontegaard Delta-Rate savings in the range of 0.8% to 28.1% over HEVC Range Extensions Test Model on screen content test video sequences used in Range Extensions core experiments by the Joint Collaborative Team for Video Coding (JCT-VC). Madhukar Budagavi, Do-Kyoung Kwon |
PCS | 1 |
| 2013 | Combined scalable and mutiview extension of High Efficiency Video Coding (HEVC)abstractThe scalable and multiview extension of High Efficiency Video Coding (HEVC) is being designed to support scalable layers and multiple views at the same time. Applications of the combined scalable and multiview HEVC coding include scalable stereoscopic video (e.g. 1080p stereo to the emerging 4K stereo), mixed resolution multiview coding, etc. However, there has been no prior work on the coding efficiency of the combined scalable and multiview HEVC coding. In this paper, the scalable multiview HEVC codec is developed based on the scalable HEVC reference software and the performances of inter-layer prediction and inter-view prediction are evaluated in the scalable stereo coding application. Two scalable HEVC coding frameworks, namely TextureRL and RefIdx, are compared as well in the same application. Do-Kyoung Kwon, Madhukar Budagavi |
PCS | 2 |
| 2012 | Unified forward+inverse transform architecture for HEVCabstractThe upcoming HEVC video coding standard supports many transform sizes ranging from 4-point to 32-point in square and rectangular form. Multiple transform sizes improve coding efficiency, but also increase the implementation complexity. Furthermore both forward and inverse transforms need to be supported in various consumer devices. This paper presents a unified forward+inverse transform architecture for HEVC. The unified architecture makes use of symmetry properties that exist in the HEVC forward and inverse transform matrices to achieve hardware sharing across different transform sizes and also between forward and inverse transforms. It uses 43-45% less area than separate forward and inverse core transform implementations. Madhukar Budagavi, Vivienne Sze |
ICIP | 1 |
| 2012 | Parallelization of CABAC transform coefficient coding for HEVCabstractData dependencies in CABAC make it difficult to parallelize and thus limits its throughput. The majority of the bins processed by the CABAC are used to represent the prediction error/residual in terms of quantized transform coefficients. This paper provides an overview of the various improvements to context selection and scans in transform coefficient coding that enable HEVC to potentially achieve higher throughput relative to AVC/H.264. Specifically, it describes changes that remove data dependencies in significance map and coefficient level coding. Proposed and adopted techniques up to HM-4.0 are discussed. This work illustrates that accounting for implementation cost when designing video coding algorithms can result in a design that can enable higher processing speed and reduce hardware cost, while still delivering high coding efficiency. Vivienne Sze, Madhukar Budagavi |
PCS | 2 |
| 2012 | High Throughput CABAC Entropy Coding in HEVCabstractContext-adaptive binary arithmetic coding (CAB-AC) is a method of entropy coding first introduced in H.264/AVC and now used in the newest standard High Efficiency Video Coding (HEVC). While it provides high coding efficiency, the data dependencies in H.264/AVC CABAC make it challenging to parallelize and thus, limit its throughput. Accordingly, during the standardization of entropy coding for HEVC, both coding efficiency and throughput were considered. This paper highlights the key techniques that were used to enable HEVC to potentially achieve higher throughput while delivering coding gains relative to H.264/AVC. These techniques include reducing context coded bins, grouping bypass bins, grouping bins with the same context, reducing context selection dependencies, reducing total bins, and reducing parsing dependencies. It also describes reductions to memory requirements that benefit both throughput and implementation costs. Proposed and adopted techniques up to draft international standard (test model HM-8.0) are discussed. In addition, analysis and simulation results are provided to quantify the throughput improvements and memory reduction compared with H.264/AVC. In HEVC, the maximum number of context-coded bins is reduced by 8×, and the context memory and line buffer are reduced by 3× and 20×, respectively. This paper illustrates that accounting for implementation cost when designing video coding algorithms can result in a design that enables higher processing speed and lowers hardware costs, while still delivering high coding efficiency. Vivienne Sze, Madhukar Budagavi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | HEVC ALF decode complexity analysis and reductionabstractThis paper analyzes the decoder implementation complexity of a new tool called Adaptive Loop Filtering (ALF) being considered for the ITU-T/ISO/IEC High Efficiency Video Coding (HEVC) standard, and proposes new luma filters (Nx7 and Nx5) for ALF that reduce memory bandwidth, memory size requirements, and number of computations. The luma filters in ALF of the initial version HEVC Test Model (HM-1.0) have a maximum vertical size of 9. The vertical size of the ALF filters determines the memory size (line buffers) and memory bandwidth requirements. Accordingly, this paper proposes reducing the vertical size of ALF filters to 7 and 5, which are referred to as Nx7 and Nx5 filter sets respectively. These filters reduce memory bandwidth and size requirements by 25% and 50% respectively with minimal impact on coding efficiency. In addition, the worst case computational complexity is reduced by -10% and -20% respectively. Reduced vertical size luma ALF filters are under consideration for inclusion in HEVC standard with Nx7 being been adopted into HM- 2.0 and Nx5 being under consideration for HM-4.0. Madhukar Budagavi, Vivienne Sze, Minhua Zhou |
ICIP | 1 |
| 2011 | Memory Bandwidth and Power Reduction Using Lossy Reference Frame Compression in Video EncodingabstractLarge external memory bandwidth requirement leads to increased system power dissipation and cost in video coding application. Majority of the external memory traffic in video encoder is due to reference data accesses. We describe a lossy reference frame compression technique that can be used in video coding with minimal impact on quality while significantly reducing power and bandwidth requirement. The low cost transformless compression technique uses lossy reference for motion estimation to reduce memory traffic, and lossless reference for motion compensation (MC) to avoid drift. Thus, it is compatible with all existing video standards. We calculate the quantization error bound and show that by storing quantization error separately, bandwidth overhead due to MC can be reduced significantly. The technique meets key requirements specific to the video encode application. 24-39% reduction in peak bandwidth and 23-31% reduction in total average power consumption are observed for IBBP sequences. Ajit Gupte, Bharadwaj S. Amrutur, Mahesh Mehendale, Ajit V. Rao, Madhukar Budagavi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2010 | Special issue on breakthrough architectures for image and video systems
Rajesh Narasimha, Madhukar Budagavi, Seok-Jun Lee, David D. Wentzloff |
Signal Process. Image Commun. | 2 |
| 2008 | Video coding using compressed reference framesabstractHandheld battery-operated consumer electronics video devices such as camera phones, digital still cameras, digital camcorders, and personal media players have limited system memory bandwidth available because of cost and power consumption constraints. Video coding consumes a significant amount of this limited system memory bandwidth especially at high-definition (HD) resolution. Techniques that reduce memory bandwidth in video coding are crucial for implementing video coding at HD resolutions in portable video devices. Memory bandwidth reduction is also desirable from a power consumption point of view since memory accesses consume a significant amount of power. In this paper we present our technique - in-loop compression of reference frames - for reducing memory bandwidth in video coding. Madhukar Budagavi, Minhua Zhou |
ICASSP | 1 |
| 2008 | Parallel CABAC for low power video codingabstractWith the growing presence of high definition video content on battery-operated handheld devices such as camera phones, digital still cameras, digital camcorders, and personal media players, it is becoming ever more important that video compression be power efficient. A popular form of entropy coding called Context-Based Adaptive Binary Arithmetic Coding (CABAC) provides high coding efficiency but has limited throughput. This can lead to high operating frequencies resulting in high power dissipation. This paper presents a novel parallel CABAC scheme which enables a throughput increase of N-fold (depending on the degree parallelism), reducing the frequency requirement and expected power consumption of the coding engine. Experiments show that this new scheme (with N=2) can deliver ∼2x throughput improvement at a cost of 0.76% average increase in bit-rate or equivalently a decrease in average PSNR of 0.025dB on five 720p resolution video clips when compared with H.264/AVC. Vivienne Sze, Anantha P. Chandrakasan, Madhukar Budagavi, Minhua Zhou |
ICIP | 3 |
| 2007 | Next generation video coding for mobile applications: industry requirements and technologiesabstractHandheld battery-operated consumer electronics devices such as camera phones, digital still cameras, digital camcorders, and personal media players have become very popular in recent years. Video codecs are extensively used in these devices for video capture and/or playback. The annual shipment of such devices already exceeds a hundred million units and is growing, which makes mobile battery-operated video device requirements very important to focus in video coding research and development. This paper highlights the following unique set of requirements for video coding for these applications: low power consumption, high video quality at low complexity, and low cost, and motivates the need for a new video coding standard that enables better trade-offs of power consumption, complexity, and coding efficiency to meet the challenging requirements of portable video devices. This paper also provides a brief overview of some of the video coding technologies being presented in the ITU-T Video Coding Experts Group (VCEG) standardization body for computational complexity reduction and for coding efficiency improvement in a future video coding standard. Madhukar Budagavi, Minhua Zhou |
VCIP | 1 |
| 2005 | Video compression using blur compensationabstractTraditional block-based motion compensation techniques such as those used in the H.264 video coding standards become ineffective when blurring starts to occur in the video sequence. Blurring typically occurs in video sequences when the relative motion between the camera and the scene being captured is faster than the camera exposure time. Blurring is also sometimes used as a special effect for smoothly transitioning between scenes in a movie. Blurring also naturally occurs when objects at different depths in a scene are focused and defocused in a video sequence. In this paper, we propose the blur compensation algorithm that makes use of the blurring information to provide improved compression performance when compared to H.264 in the presence of blur in video sequences. When coding blurred scenes in video sequences, bitrate reductions of up to 64% were achieved. Madhukar Budagavi |
ICIP (2) | 1 |
| 2001 | Multiframe video coding for improved performance over wireless channelsabstractWe propose and evaluate a multi-frame extension to block motion compensation (BMC) coding of videoconferencing-type video signals for wireless channels. The multi-frame BMC (MF-BMC) coder makes use of the redundancy that exists across multiple frames in typical videoconferencing sequences to achieve additional compression over that obtained by using the single frame BMC (SF-BMC) approach, such as in the base-level H.263 codec. The MF-BMC approach also has an inherent ability of overcoming some transmission errors and is thus more robust when compared to the SF-BMC approach. We model the error propagation process in MF-BMC coding as a multiple Markov chain and use Markov chain analysis to infer that the use of multiple frames in motion compensation increases robustness. The Markov chain analysis is also used to devise a simple scheme which randomizes the selection of the frame (amongst the multiple previous frames) used in BMC to achieve additional robustness. The MF-BMC coders proposed are a multi-frame extension of the base level H.263 coder and are found to be more robust than the base level H.263 coder when subjected to simulated errors commonly encountered on wireless channels. Madhukar Budagavi, Jerry D. Gibson |
IEEE Trans. Image Process. | 1 |
| 1999 | Wireless MPEG-4 video on Texas Instruments DSP chipsabstractTechnology has advanced in recent years to the point where multimedia communicators are beginning to emerge. These communicators are low-power, portable devices that can transmit and receive multimedia data through the wireless network. Due to the high computational complexity involved and the low-power constraint in wireless applications, these devices require the use of processors that are powerful and are at the same time very power-efficient. In order to facilitate interoperability, it is important that these devices use standardized compression and communication algorithms. As a first step in implementing multimedia terminals, Texas Instruments (TI) has demonstrated real-time MPEG-4 video decoding (simple profile) on TMS320C54x, TI's low power, high performance DSP chip. In addition, TI has outlined a system-level solution to transmitting video across wireless networks, including channel coding and communication protocols. Madhukar Budagavi, Wendi B. Heinzelman, Jennifer L. H. Webb, Raj Talluri |
ICASSP | 1 |
| 1999 | Unequal Error Protection of MPEG-4 Compressed VideoabstractThe MPEG-4 video compression standard incorporates several techniques to create error-resilient coded video. These techniques are effective against a certain level of errors. However, wireless channels often have very high error rates; thus channel coding is needed to reduce the number of errors in the compressed bitstream that is sent to the MPEG-4 video decoder. The structure of an MPEG-4 compressed bitstream lends itself to using unequal error protection to ensure fewer errors in the important portions of the bitstream. This paper discuses how unequal error protection can be used with the MPEG-4 error resilience tools and describes several experiments using both unequal and equal error protection on video sent through a simulated GSM channel. The results of these experiments show that when the bit error rate is high, unequal error protection can improve reconstructed video quality by as much as 1 dB compared with equal error protection. Wendi B. Heinzelman, Madhukar Budagavi, Raj Talluri |
ICIP (2) | 2 |
| 1998 | Speech coding in mobile radio communicationsabstractSpeech coding, the efficient representation of speech in digital form, is one of the key technologies in current and evolving digital cellular and wireless voice communications offerings. The speech coders in existing standards exhibit a level of sophistication and performance unimaginable just 15 years ago. We outline the characteristics of the mobile communications problem with respect to speech coders and point out the principal issues in speech coder design for these applications. Speech coding methods in existing mobile communications standards are described and contrasted. The limitations imposed by the wireless channel and by background impairments are discussed, and approaches to addressing their resulting effects are presented. Suggestions for future research in speech coding for the mobile communications problem are outlined. Madhukar Budagavi, Jerry D. Gibson |
Proc. IEEE | 1 |
| 1997 | Error Propagation in Motion Compensated Video Over Wireless ChannelsabstractThe use of multiple frames in block motion compensation (BMC) provides us with a new framework for increasing the robustness of video coders. We study the error robustness properties of multiframe BMC (MF-BMC) coders by modeling the error propagation process in the MF-BMC approach as a multiple Markov chain. Using Markov chain analysis we find that the probability of error propagation can decrease (i.e. robustness can increase) by (i) the use of additional frames in BMC, and by (ii) decreasing the probability with which blocks in the immediate previous frame are chosen for motion compensation. We outline a simple scheme that modifies the probabilities of macroblock prediction to achieve additional robustness. The MF-BMC coders proposed in this paper are a multiframe extension of the base level H.263 coder and are found to be more robust than the base level H.263 coder when subjected to simulated errors commonly encountered on wireless channels. Madhukar Budagavi, Jerry D. Gibson |
ICIP (2) | 1 |