Guan-Ming Su

dblp:06/6913 · DBLP profile ↗
← Back
38ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0002-3118-5904ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 5 first-author · 21 since 2021Computer networks · 5 · 4 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Human Pose Aggregation for Multi-View Temporal Video Alignment
abstract
When multiple videos of a scene are taken from differing viewpoints without precise synchronization, it can be difficult to temporally align them after the fact. Often the metadata or audio needed to do so is missing or inaccurate. But human motion in such videos can provide a strong signal for identifying matching time points across videos, through analysis of pose and movement. In this work, we leverage view-invariant human pose features to synchronize videos. Unlike previous human pose-based alignment techniques, our method can align videos containing multiple people without performing tracking or re-identification across views. We achieve this by aggregating pose information from multiple people into a single frame descriptor. This also enables fast ${\mathcal{O}}\left({n\log n}\right)$ search for the optimal alignment. This simple but effective strategy leads to major and consistent improvements over existing human-based and visual feature temporal alignment techniques.
Fabien Delattre, Tsung-Wei Huang, Guan-Ming Su, Erik G. Learned-Miller
WACV3
2026 Gaussian Representations for Video
abstract
We introduce Gaussian representations for videos (GaRV), a novel video encoding and decoding scheme based upon 3D Gaussians. Unlike traditional representations, which encode videos as sequences of frames, or neural representations, which encode videos within the weights of a neural network, we encode videos as a collection of 3D Gaussians within a space-time volume. The key advantage of our approach is that it enables efficient and flexible rasterization-based video decoding. With a slight drop in overall compression rate, GaRV offers an 8-50× improvement in decoding time and 2.5-15× reduction in GPU memory compared with neural counterparts. Existing Gaussian video techniques require 2-30× more disk space, while also using more GPU resources than GaRV. Moreover, GaRV offers unique flexibility in how and when pixels are decoded: One can non-sequentially decode frames/regions without penalty and can selectively decode regions at high-resolution to enable low-cost foveated video decoding.
Sachin Shah, Anustup Choudhury, Guan-Ming Su, Jaclyn Pytlarz, Christopher A. Metzler, Trisha Mittal
WACV3
2026 Standardizing Generative Face Video Compression Using Supplemental Enhancement Information
abstract
This paper proposes a Generative Face Video Compression (GFVC) approach using Supplemental Enhancement Information (SEI), where a series of compact spatial and temporal representations of a face video signal (e.g., 2D/3D key-points, facial semantics and compact features) can be coded using SEI messages and inserted into the coded video bitstream. At the time of writing, the proposed GFVC approach using SEI messages has been included into a draft amendment of the Versatile Supplemental Enhancement Information (VSEI) standard by the Joint Video Experts Team (JVET) of ISO/IEC JTC 1/SC 29 and ITU-T SG21, which will be standardized as a new version of ITU-T H.274$|$ISO/IEC 23002-7. To the best of the authors' knowledge, the JVET work on the proposed SEI-based GFVC approach is the first standardization activity for generative video compression. The proposed SEI approach has not only advanced the reconstruction quality of early-day Model-Based Coding (MBC) via the state-of-the-art generative technique, but also established a new SEI definition for future GFVC applications and deployment. Experimental results illustrate that the proposed SEI-based GFVC approach can achieve remarkable rate-distortion performance compared with the latest Versatile Video Coding (VVC) standard, whilst also potentially enabling a wide variety of functionalities including user-specified animation/filtering and metaverse-related applications.
Yan Ye 0003, Jie Chen 0006, Ru-Ling Liao, Shanzhi Yin, Shiqi Wang 0001, Kaifa Yang, Yue Li 0015, Yiling Xu, Ye-Kui Wang, Shiv Gehlot, Guan-Ming Su, Peng Yin 0002, Sean McCarthy, Gary J. Sullivan
IEEE Trans. Multim.12
2025 Parametric Shadow Control for Portrait Generation in Text-to-Image Diffusion Models
abstract
Text-to-image diffusion models excel at generating diverse portraits, but lack intuitive shadow control. Existing editing approaches, as post-processing, struggle to offer effective manipulation across diverse styles. Additionally, these methods either rely on expensive real-world light-stage data collection or require extensive computational resources for training. To address these limitations, we introduce Shadow Director, a method that extracts and manipulates hidden shadow attributes within well-trained diffusion models. Our approach uses a small estimation network that requires only a few thousand synthetic images and hours of training-no costly real-world light-stage data needed. Shadow Director enables parametric and intuitive control over shadow shape, placement, and intensity during portrait generation while preserving artistic integrity and identity across diverse styles. Despite training only on synthetic data built on real-world identities, it generalizes effectively to generated portraits with diverse styles, making it a more accessible and resource-friendly solution.
Haoming Cai, Tsung-Wei Huang, Shiv Gehlot, Brandon Yushan Feng, Sachin Shah, Guan-Ming Su, Christopher A. Metzler
ICCV6
2025 A Generative Face Video Coding Framework with Disentangled and Consistent Background
abstract
Existing Generative face video coding (GFVC) frameworks enable the ultra-low bandwidth video communication through transmission of compact facial representations. However, non-localization of derived facial representations leads to foreground and background blending (entanglement) in the decoded sequences, resulting in geometry distortions. In this work, we propose a GFVC framework that removes this blending, and suppresses the induced artifacts. To achieve the disentanglement, the proposed framework 1) separately transmits foreground and background of the first frame (base pictures), and 2) performs fusion at the decoder end with background base picture. Further, the proposed methodology supports chroma keying for decoder simplification and background customization. The proposed approach is generic and builds upon existing GFVC frameworks to generate stable and consistent video sequences. Compared to VVC, proposed algorithm reduces the average bit rate by 52.13% and offers a 1.24% average improvement in DISTS over existing GFVC methods at QP 22. Additionally, subjective evaluations reveal a 88.89% preference for the proposed approach, in contrast to 11.11% for current best-performing method.
Shiv Gehlot, Guan-Ming Su, Peng Yin 0002, Sean McCarthy, Gary J. Sullivan
ICIP2
2025 Quanta-Slomo: Single Photon Camera Guided 100x Video Frame Interpolation
abstract
Video Frame Interpolation (VFI) enhances video quality by computationally increasing video frame rates. Conventional VFI methods rely on simplified motion assumptions (e.g. linear motion) between adjacent frames to generate intermediate frames, which often causes error in highly dynamic scenarios. The Single Photon Avalanche Diode (SPAD) array is an emerging type of quanta image sensor that can capture videos at very high frame rates but can be heavily corrupted by Poisson noise. We present Quanta-SloMo, a novel VFI method that leverages the high frame rate but noisy video from a SPAD sensor to guide the VFI of high quality, low frame rate video captured by a CMOS sensor. Our framework utilizes residual learning, feature alignment through deformable convolution, multi-frame merging, and handles the noise-to-blur trade-off between SPAD video frames through a virtual exposure stack. We demonstrate that Quanta-SloMo outperforms state-of-the-art VFI methods by significant margins, especially for scenes with large motion.
Anustup Choudhury, Guan-Ming Su, Andreas Velten
ICIP3
2025 NIVM: Real-time View Morphing via Neural Implicit Function
Tung-I Chen, Dae Yeol Lee, Guan-Ming Su, Mohammad Hajiesmaili, Ramesh K. Sitaraman
ACM Multimedia3
2025 Video Content Restoration in the Wild: Challenges and Opportunities
abstract
With the wide deployment of visual capture systems, video content restoration plays a key component in the video processing pipeline. We will review the main video content restoration technologies developed in the past decades, including both conventional and deep learning-based solutions. We will highlight the limitations of existing methods, discuss the challenges for the real-world captured video use cases, and point out the potential future research directions.
Guan-Ming Su
ACM Multimedia1
2025 INT-DTT+: Low-Complexity Data-Dependent Transforms for Video Coding
Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega, Tsung-Wei Huang, Thuong Nguyen Canh, Guan-Ming Su, Peng Yin 0002
PCS6
2024 The Multiplane Image Information SEI Message and its Use for Distribution of Volumetric Video with Conventional Codecs
abstract
This paper describes the background, design and application of a new SEI message – the Multiplane Image Information SEI message, which has recently been adopted into the Technology under Consideration (TuC) document of the JVET committee for potential inclusion in the VSEI standard (ITU-T H.274 and ISO/IEC 23002-7). The paper also provides preliminary compression experiment results and analysis on the implications of the coding efficiency and functionality of the different packing options supported in the SEI message.
Taoran Lu, Peng Yin 0002, Guan-Ming Su, Dae Yeol Lee, Tsung-Wei Huang, Sejin Oh, Sean McCarthy, Walt Husak, Gary J. Sullivan
DCC3
2024 V-Trans4Style: Visual Transition Recommendation for Video Production Style Adaptation
Pooja Guhan, Tsung-Wei Huang, Guan-Ming Su, Subhadra Gopalakrishnan, Dinesh Manocha
ECCV (80)3
2024 NeRVA: Joint Implicit Neural Representations for Videos and Audios
abstract
Neural fields also known as implicit neural representations (INR) have recently been shown to be quite effective at representing video content. However, video content typically contains audio and existing works on INR do not represent both video and audio. Since INR for audio is not well-explored, we propose a novel neural representation for representing audio (NeRA) that can represent audio using a neural network. We also propose a novel neural representation for jointly representing both videos and audios (NeRVA) using a neural network. We represent multimedia as a neural network that takes timestamp as input and outputs the corresponding RGB image frame and the audio samples. The proposed neural network architecture uses a combination of Multi-Layer Perceptron (MLP) and convolutional blocks. We also demonstrate that our joint representation of multimedia content has better performance than individually representing the components of multimedia (An improvement of +2dB PSNR for video and 10× FAD for audio).
Anustup Choudhury, Praneet Singh, Guan-Ming Su
ICME3
2024 Outdoor Scene Relighting with Diffusion Models
Jinlin Lai, Anustup Choudhury, Guan-Ming Su
ICPR (6)3
2024 Analysis of Coding Gain Due to In-Loop Reshaping
abstract
Reshaping, a point operation that alters the characteristics of signals, has been shown capable of improving the compression ratio in video coding practices. Out-of-loop reshaping that directly modifies the input video signal was first adopted as the supplemental enhancement information (SEI) for the HEVC/H.265 without the need to alter the core design of the video codec. VVC/H.266 further improves the coding efficiency by adopting in-loop reshaping that modifies the residual signal being processed in the hybrid coding loop. In this paper, we theoretically analyze the rate-distortion performance of the in-loop reshaping and use experiments to verify the theoretical result. We prove that the in-loop reshaping can improve coding efficiency when the entropy coder adopted in the coding pipeline is suboptimal, which is in line with the practical scenarios that video codecs operate in. We derive the PSNR gain in a closed form and show that the theoretically predicted gain is consistent with that measured from experiments using standard testing video sequences.
Chau-Wai Wong, Chang-Hong Fu 0002, Mengting Xu, Guan-Ming Su
IEEE Trans. Image Process.4
2024 "May I Speak?": Multi-Modal Attention Guidance in Social VR Group Conversations
abstract
In this paper, we present a novel multi-modal attention guidance method designed to address the challenges of turn-taking dynamics in meetings and enhance group conversations within virtual reality (VR) environments. Recognizing the difficulties posed by a confined field of view and the absence of detailed gesture tracking in VR, our proposed method aims to mitigate the challenges of noticing new speakers attempting to join the conversation. This approach tailors attention guidance, providing a nuanced experience for highly engaged participants while offering subtler cues for those less engaged, thereby enriching the overall meeting dynamics. Through group interview studies, we gathered insights to guide our design, resulting in a prototype that employs light as a diegetic guidance mechanism, complemented by spatial audio. The combination creates an intuitive and immersive meeting environment, effectively directing users' attention to new speakers. An evaluation study, comparing our method to state-of-the-art attention guidance approaches, demonstrated significantly faster response times (p < 0.001), heightened perceived conversation satisfaction (p < 0.001), and preference (p < 0.001) for our method. Our findings contribute to the understanding of design implications for VR social attention guidance, opening avenues for future research and development.
Geonsun Lee, Dae Yeol Lee, Guan-Ming Su, Dinesh Manocha
IEEE Trans. Vis. Comput. Graph.3
2023 Neural Field Real-Time Transmission Using Multiple Description Coding with Random Position Sampling
abstract
Neural fields are a new signal representation and are widely used now to represent various forms of multimedia. However, due to their size, they could introduce latency when delivered over the network. A way to mitigate that is to use Multiple description coding (MDC), that provides multiple descriptions of the same content which can then be transmitted along different paths to improve reliability/efficiency. In this paper, we introduce a novel MDC framework that is based on neural fields. We first apply a randomly sampled neural field (with small model size) to generate multiple descriptions based on random initializations. We leverage the fact that neural fields are continuous functions (constructed using a multi-layer perceptron (MLP)) and use that to reconstruct the entire original content and progressively improve the quality as more descriptions are obtained. We validate the effectiveness of the proposed method by showing results on public data sets.
Anustup Choudhury, Guan-Ming Su
ICIP2
2023 Film Grain Removal Using Metadata
abstract
Film grain has been shown as an effective way to improve the look of video in terms of aesthetic feeling and sharpness. For backward compatibility, a strategy is to inject the film grain before video compression at encoder to ensure that any decoder can acquire the film grain injected images after video decompression. Moreover, for designed decoder, the film grain can be removed, and further video processing can be applied, such as adding other type of film grain or adaptive film grain setting according to the viewing environment. Therefore, in this work, we propose a metadata-aided film grain removal system such that the film grain can be removed efficiently using metadata in video bitstream. We propose three filtering methods (Gaussian, optimal FIR, and guided filter) in the proposed system. Experimental results show that the proposed system can remove film grain effectively and yield high PSNR.
Tsung-Wei Huang, Guan-Ming Su, Peng Yin 0002
ICIP2
2023 Progressive Coding for Neural Field Transmission
abstract
Neural fields are a new signal representation and are widely used now to represent various forms of multimedia. However, due to their size, they could introduce latency when delivered over the network. A way to mitigate that is to use Progressive Coding which encodes the content into a single bitstream that can be decoded at various bitrates. In this paper, we introduce a novel progressive coding framework for neural fields. We propose three progressive scanning mechanisms - first is based on the layers of the model, the second is based on the bit-planes across all layers of the model, and the third is based on the block of coefficients across all layers of the model. We leverage the fact that neural fields are continuous functions (constructed using a multi-layer perceptron (MLP)) and use that to reconstruct the entire original content and progressively improve the quality as more bits are obtained. We validate the effectiveness of the proposed method by showing results on public data sets.
Anustup Choudhury, Guan-Ming Su
ISM2
2023 A Framework for Multi-plane Image Layer Merging
abstract
The multi-plane image (MPI) is a promising framework for enabling novel view synthesis for 3D scene reconstruction. However, inefficient utilization of the MPI layers traditionally causes low compression efficiency, a large memory footprint, and high computational complexity. To address this issue, this paper proposes a framework for decreasing the number of layers in an MPI while producing high quality rendered images. We introduce a principled layer merging algorithm to reduce the number of layers in an MPI by using a constrained Lloyd-Max optimization formulation for determining how the layers should be merged. Since an imperfect MPI generation model may produce undesired artifacts in rendered images, which may be exacerbated after layer merging has been applied, a novel enhancement algorithm is constructed and applied before and after layer merging to remove artifacts and ensure the high quality of the layer-merged MPIs. Experimental results verify the utility of our framework, showing that: (1) our enhancement algorithm can significantly improve the quality of the MPI generated data, with an average increase in peak signal-to-noise ratio (PSNR) of $28 \sim 29$ dB and (2) our layer merging algorithm can successfully reduce the number layers in an MPI while producing high quality rendered images.
Zachary McBride Lazri, Guan-Ming Su, Peng Yin 0002
ISM2
2021 Multi-Scale Feature Guided Low-Light Image Enhancement
abstract
Low-light image enhancement aims at enlarging the intensity of image pixels to better match human perception and to improve the performance of subsequent vision tasks. While it is relatively easy to enlighten a globally low-light image, the lighting condition of realistic scenes is usually non-uniform and complex, e.g., some images may contain both bright and extremely dark regions, with or without rich features and information. Existing methods often generate abnormal light-enhancement results with over-exposure artifacts without proper guidance. To tackle this challenge, we propose a multi-scale feature guided attention mechanism in the deep generator, which can effectively perform a spatially-varying light enhancement. The attention map is fused by both the gray map and extracted feature map of the input image, to focus more on those dark and informative regions. Our baseline is an unsupervised generative adversarial network, which can be trained without any low/normal light image pair. Experimental results demonstrate the superiority in visual quality and performance of subsequent object detection over state-of-the-art alternatives.
Lanqing Guo, Renjie Wan, Guan-Ming Su, Alex Chichung Kot, Bihan Wen
ICIP3
2021 Revertible Guidance Image Based Image Detail Enhancement
abstract
Image detail enhancement is widely used in image processing tasks to give better look and higher contrast to images and photos. However, the reverse process that converts an enhanced image back to its original image usually does not have an explicit representation given only the enhanced image is available. In this work, we propose a generic framework of revertible image detail enhancement so that we can estimate the original image without extra information. We define a family of revertible detail enhancement operators that convert each pixel from original image to enhanced image and vice versa. A guidance image is used to decide which operator to use for each pixel. In the enhancement process, the guidance image is generated from the original image. In the reverse process, the guidance image is estimated from iterative optimization, making the process revertible. Experimental results show that the proposed enhancement framework can convert the enhanced image back to its original image without noticeable difference.
Tsung-Wei Huang, Guan-Ming Su
ICIP2
2020 Efficient Debanding Filtering for Inverse Tone Mapped High Dynamic Range Videos
abstract
When a low or standard dynamic range video is inverse tone mapped to high dynamic range (HDR), there can be banding artifacts in the output HDR video. We design a highly integrated architecture that can detect and alleviate banding artifacts while preserving true edges and details in an extremely efficient way. This is achieved by a 7-tap edge-aware selective sparse filter. Coding artifacts, such as blocky artifacts, can also be reduced by this filter. The filter includes some parameters that depend on the strengths of the banding artifacts. A parameter selection mechanism is presented which considers smoothness of the banding regions and fidelity of the filtering output. The filter yields significant PSNR gain at the regions of artifacts. Subjective tests demonstrate the great quality improvement achieved by the proposed filter, compared to the quality before filtering. The visual quality provided by the filter is better than or similar to that of algorithms which are far more complex.
Qing Song 0005, Guan-Ming Su, Pamela C. Cosman
IEEE Trans. Circuits Syst. Video Technol.2
2018 Transim: Transfer Image Local Statistics Across EOTFS for HDR Image Applications
abstract
Despite the popularity of high dynamic range (HDR) technology in recent years, various algorithms for image and video applications are still designed and optimized for traditional standard dynamic range (SDR) data. Directly applying SDR-optimized algorithms to HDR images and video will result in significant artifacts or coding deficiency. In this work, we present a novel preprocessing method, dubbed TransIm, which transfers local statistics for the images from the desired domain (e.g. SDR) to the current domain (e.g., HDR), while maintaining its current visual presence. It is achieved by controlling the less perceivable “noise” that is orthogonal to the sparsifiable image content, using a unitary sparsifying transform. Numerical results show that the proposed TransIm can effectively transfer local patch variance from Gamma domain to Perceptual Quantizer (PQ) domain for HDR videos. We also demonstrate that the TransIm outputs are more robust to distortions and artifacts in seam carving applications.
Bihan Wen, Guan-Ming Su
ICME2
2018 Single Layer Progressive Coding for High Dynamic Range Videos
abstract
There are different kinds of high dynamic range (HDR) displays in the market today. These displays have different HDR specifications, like, peak/dark brightness levels, electro-optical transfer functions (EOTF), color spaces etc. For the best visual experience on a given HDR screen, colorists have to grade videos for that specific display's luminance range. But simultaneous transmission of multiple video bitstreams graded at different luminance ranges, is inefficient in terms of network utility and server storage. To overcome this problem, we propose transmitting our progressive metadata with a base layer video bitstream. This embedding allows different overlapping portions of metadata to scale the base video to progressively wider luminance ranges. Our progressive metadata format provides a significant design improvement over the existing architectures, preserves colorist intent at all the supported brightness ranges and still keeps the bandwidth or storage overhead minimal.
Harshad Kadu, Qing Song 0005, Guan-Ming Su
PCS3
2016 Content aware quantization: Requantization of high dynamic range baseband signals based on visual masking by noise and texture
abstract
High dynamic range imaging is currently being introduced to television, cinema and computer games. While it has been found that a fixed encoding for high dynamic range imagery needs at least 11 to 12 bits of tonal resolution, current mainstream image transmission interfaces, codecs and file formats are limited to 10 bits. To be able to use current generation imaging pipelines, this paper presents a baseband quantization scheme that exploits content characteristics to reduce the needed tonal resolution per image. The method is of low computational complexity and provides robust performance on a wide range of content types in different viewing environments and applications.
Jan Fröhlich, Guan-Ming Su, Scott Daly, Andreas Schilling 0001, Bernd Eberhardt
ICIP2
2016 Hardware-efficient debanding and visual enhancement filter for inverse tone mapped high dynamic range images and videos
abstract
If a low dynamic range (LDR) image or video is inverse tone mapped to a higher dynamic range, there can be banding artifacts in the output high dynamic range (HDR) image or video. We design a selective sparse filter to remove the banding artifacts and at the same time preserve edges and details. The filter is able to reduce other artifacts, such as blocky artifacts which are due to the compression of the LDR image/video. The filter is computationally efficient.
Qing Song 0005, Guan-Ming Su, Pamela C. Cosman
ICIP2
2016 Impact Analysis of Baseband Quantizer on Coding Efficiency for HDR Video
abstract
Digitally acquired high dynamic range (HDR) video baseband signal can take 10-12 bits per color channel. It is economically important to be able to reuse the legacy 8 or 10-bit video codecs to efficiently compress the HDR video. Linear or nonlinear mapping on the intensity can be applied to the baseband signal to reduce the dynamic range before the signal is sent to the codec, and we refer to this range reduction step as a baseband quantization. We show analytically and verify using test sequences that the use of the baseband quantizer lowers the coding efficiency. Experiments show that as the baseband quantizer is strengthened by 1.6 bits, the drop of PSNR at a high bitrate is up to 1.60 dB. Our result suggests that in order to achieve high coding efficiency, information reduction of videos in terms of quantization error should be introduced in the video codec instead of on the baseband signal.
Chau-Wai Wong, Guan-Ming Su, Min Wu 0001
IEEE Signal Process. Lett.2
2016 QoE in video streaming over wireless networks: perspectives and research challenges
Guan-Ming Su, Xiao Su 0006, Mea Wang, Athanasios V. Vasilakos, Haohong Wang
Wirel. Networks1
2008 Dynamic Resource Allocation for Robust Distributed Multi-Point Video Conferencing
abstract
This paper proposes a distributed multi-point video conferencing system over packet erasure channels, where the aggregation of multiple video streams and resource allocation are performed in a distributed manner. Video stream combiners, which are located in different geographical areas and serve as portals for conferees, aggregate incoming streams supplied by local users with other streams aggregated from nearby video stream combiners. A packet-division multiple-access (PDMA)-based error protection scheme is proposed to be performed at each video stream combiner to minimize the maximal expected video distortion among aggregated streams. The proposed error protection scheme for multi-stream aggregation also supports user preference. In order to deliver video streams to end users with different preferred quality, a consensus algorithm is proposed to adaptively perform resource allocation based on user preference. Simulation results show that the proposed multi-stream aggregation and error protection scheme has significant gains over traditional multi-stream error protection schemes for a multi-point video conferencing system.
Guan-Ming Su, Min Wu 0001
IEEE Trans. Multim.2
2006 Cross-Path PDMA-Based Error Protection for Streaming Multiuser Video over Multiple Paths
abstract
In this paper, we consider aggregating multiple video streams and transmitting the merged stream over multiple error-prone paths. We propose a novel multi-stream error protection scheme based on packet division multiplexing access (PDMA) and cross-path forward error coding. Compared with the traditional time division multiplexing access (TDMA)-based error protection scheme, the proposed scheme outperforms 1.43~1.88 dB for the averaged PSNR of all received streams.
Guan-Ming Su, Min Wu 0001
ICIP1
2006 Robust Distributed Multi-Point Video Conferencing Over Error-Prone Channels
abstract
In this paper, we propose a novel multi-point video conferencing system through error-prone channels, where the aggregation of multiple video streams and resource allocation are performed in a distributed manner. Video stream combiners, who are located in different geographical areas and serve as portals for conferees, aggregate incoming streams supplied by local users with other streams aggregated from nearby video stream combiners. A distributed multi-stream error protection scheme is performed in each video stream combiner to minimize the maximal expected video distortion among all aggregated streams. The simulation results demonstrate that our proposed scheme outperforms the traditional multicasting scheme by 1 dB~1.4 dB in terms of average PSNR
Guan-Ming Su, Min Wu 0001
ICME2
2006 A Scalable Multiuser Framework for Video Over OFDM Networks: Fairness and Efficiency
abstract
In this paper, we propose a framework to transmit multiple scalable video programs over downlink multiuser orthogonal frequency division multiplex (OFDM) networks in real time. The framework explores the scalability of the video codec and multidimensional diversity of multiuser OFDM systems to achieve the optimal service objectives subject to constraints on delay and limited system resources. We consider two essential service objectives, namely, the fairness and efficiency. Fairness concerns the video quality deviation among users who subscribe the same quality of service, and efficiency relates to how to attain the highest overall video quality using the available system resources. We formulate the fairness problem as minimizing the maximal end-to-end distortion received among all users and the efficiency problem as minimizing total end-to-end distortion of all users. Fast suboptimal algorithms are proposed to solve the above two optimization problems. The simulation results demonstrated that the proposed fairness algorithm outperforms a time division multiple (TDM) algorithm by 0.5 ~ 3 dB in terms of the worst received video quality among all users. In addition, the proposed framework can achieve a desired tradeoff between fairness and efficiency. For achieving the same average video quality among all users, the proposed framework can provide fairer video quality with 1 ~ 1.8 dB lower PSNR deviation than a TDM algorithm
Guan-Ming Su, Zhu Han 0001, Min Wu 0001, K. J. Ray Liu
IEEE Trans. Circuits Syst. Video Technol.1
2006 Multiuser Distortion Management of Layered Video over Resource Limited Downlink Multicode-CDMA
abstract
Transmitting multiple real-time encoded videos to multiple users over wireless cellular networks is a key driving force for developing broadband technology. We propose a new framework to transmit multiple users' video programs encoded by MPEG-4 FGS codec over downlink multicode CDMA networks in real time. The proposed framework jointly manages the rate adaptation of source and channel coding, CDMA code allocation, and power control. Subject to the limited system resources, such as the number of pseudo-random codes and the maximal power for CDMA transmission, we develop an adaptive scheme of distortion management to ensure baseline video quality for each user and further reduce the overall distortion received by all users. To efficiently utilize system resources, the proposed scheme maintains a balanced ratio between the power and code usages. We also investigate three special scenarios where demand, power, or code is limited, respectively. Compared with existing methods in the literature, the proposed algorithm can reduce the overall system's distortion by 14% to 26%. In the demand-limited case and the code-limited but power-unlimited case, the proposed scheme achieves the optimal solutions. In the power-limited but code-unlimited case, the proposed scheme has a performance very close to a performance upper bound
Zhu Han 0001, Guan-Ming Su, Andres Kwasinski, Min Wu 0001, K. J. Ray Liu
IEEE Trans. Wirel. Commun.2
2005 Joint uplink and downlink optimization for video conferencing over wireless LAN
abstract
A real-time video conferencing framework is proposed for multiple conferencing pairs by jointly considering the uplink and downlink conditions within IEEE 802.11 networks. We formulate this system so as to minimize the maximal end-to-end expected distortion received by all users by selecting the PHY modes and transmission time. Compared with the strategy of individually optimizing uplink and downlink, the proposed framework outperforms by 3.67-8.65 dB for the average received PSNR among all users.
Guan-Ming Su, Zhu Han 0001, Min Wu 0001, K. J. Ray Liu
ICASSP (2)1
2005 Efficient bandwidth resource allocation for low-delay multiuser video streaming
abstract
This paper studies efficient bandwidth resource allocation for streaming multiple MPEG-4 fine granularity scalability (FGS) video programs to multiple users. We begin with a simple single-user scenario and propose a rate-control algorithm that has low delay and achieves an excellent tradeoff between the average visual distortion and the quality fluctuation. The proposed algorithm employs two weight factors for adjusting the tradeoff, and the optimal choice of these factors is derived. We then extend to the multiuser case and propose a dynamic resource allocation algorithm with low delay and low computational complexity. By exploring the variations in the scene complexity of video programs as well as dynamically and jointly distributing the available system resources among users, our proposed algorithm provides low fluctuation of quality for each user, and can support consistent or differentiated quality among all users to meet applications' needs. Experimental results show that compared to traditional look-ahead sliding-window approaches, our algorithm can achieve comparable visual quality and channel utilization at a much lower cost of delay, computation, and storage.
Guan-Ming Su, Min Wu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2004 Dynamic distortion control for 3-D embedded wavelet video over multiuser OFDM networks
abstract
In this paper, we propose a system to transmit multiple 3D embedded wavelet video programs over downlink multiuser OFDM. We consider the fairness among users and formulate the problem as minimizing the users' maximal distortion subject to power, rate, and subcarrier constraints. By exploring frequency, time, and multiuser diversity in OFDM and flexibility of the 3D embedded wavelet video codec, the proposed algorithm can achieve fair video qualities among all users. Compared to a scheme similar to the current multiuser OFDM standard (IEEE 802.11a), the proposed scheme outperforms it by 1-5 dB on the worst received PSNR among all users and has much smaller PSNR deviation.
Guan-Ming Su, Zhu Han 0001, Min Wu 0001, K. J. Ray Liu
GLOBECOM1
2004 Distortion management of real-time MPEG-4 video over downlink multicode CDMA networks
abstract
In this paper, a protocol is designed to manage source rate/channel coding rate adaptation, code allocation, and power control to transmit real-time MPEG-4 FGS video over downlink multicode CDMA networks. We develop a fast adaptive scheme of distortion management to reduce the overall distortion received by all users subject, to the limited number of codes and maximal transmitted power. Compared with a modified greedy method in literature, our proposed algorithm can reduce the overall system's distortion by at least 45%.
Guan-Ming Su, Zhu Han 0001, Andres Kwasinski, Min Wu 0001, K. J. Ray Liu, Nariman Farvardin
ICC1
2004 Efficient bandwidth resource allocation for low-delay multiuser MPEG-4 video transmission
abstract
An efficient bandwidth resource allocation algorithm with low delay and low fluctuation of quality to transmit multiple MPEG-4 fine granularity scalability (FGS) video programs to multiple users is proposed in this paper. By exploring the variation in the scene complexity of each video program and jointly redistributing available system resources among users, our proposed algorithm provides low fluctuation of quality for each user and consistent quality among all users. Experimental results show that compared to a traditional look-ahead sliding-window approach, our scheme can achieve comparable perceptual quality and channel utilization at a much lower cost of delay, computation, and storage.
Guan-Ming Su, Min Wu 0001
ICC1