Songyu Yu

dblp:86/1513 · DBLP profile ↗
← Back
52ranked-venue papers
0as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 2Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
1 paper
Coding theory · 100%
Computer graphics and multimedia
1 paper
Image and video processing · 67% Rendering · 33%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Coding theory › source coding
rate-distortion theory
0.212014
On Two-Stage Sequential Coding of Correlated Sources · IEEE Trans. Inf. Theory 2014
Coding theory › source coding
sequential coding
0.212014
On Two-Stage Sequential Coding of Correlated Sources · IEEE Trans. Inf. Theory 2014
Rendering
antialiasing
0.112012
Multiscale Semilocal Interpolation With Antialiasing · IEEE Trans. Image Process. 2012
Image and video processing › video frame interpolation › interpolation
image interpolation
0.112012
Multiscale Semilocal Interpolation With Antialiasing · IEEE Trans. Image Process. 2012
Image and video processing
image restoration
0.112012
Multiscale Semilocal Interpolation With Antialiasing · IEEE Trans. Image Process. 2012
Coding theory › source coding › multiterminal source coding
correlated sources
0.112014
On Two-Stage Sequential Coding of Correlated Sources · IEEE Trans. Inf. Theory 2014
Coding theory
source coding
0.112014
On Two-Stage Sequential Coding of Correlated Sources · IEEE Trans. Inf. Theory 2014

Methods — techniques the papers use, named apart from their topics

rate-distortion region characterization · 0.2inner bound · 0.2multiscale semilocal interpolation · 0.1maximum a posteriori estimation · 0.1bilateral total variation · 0.1
YearPublicationVenuePosition
2014 On Two-Stage Sequential Coding of Correlated Sources
abstract
We study the problem of two-stage sequential coding (TSSC), which is an extension of sequential coding of correlated sources. Let X and Y be dependent random variables. The network contains two encoders and two decoders: 1) a Y encoder with input Y; 2) an X encoder with inputs X and Y; 3) a Y decoder that reconstructs Y; and 4) an X decoder that reconstructs X. The first stage is traditional sequential coding, where the Y encoder describes Y to both the X decoder and Y decoder, and the X encoder describes X and Y to the X decoder. At the second stage, the Y encoder refines the description of Y, and the X encoder refines the description of X. The TSSC model is a theoretical abstraction of scalable video coding; here, Y and X represent successive frames of a video sequence, and the two stages together give an embedded description that allows the video to be decoded at two distinct rates. We give an inner bound on the rate distortion region for this TSSC model. The tight bound on the rate distortion region is derived when Y must be reconstructed losslessly (in the usual Shannon sense) in the second stage. We also study the minimum total rate of the TSSC model and show that the minimum total rate of one-stage sequential coding cannot be achieved at both stages for jointly Gaussian sources. This theoretical result can shed light on the rate-distortion performance behavior of scalable video coding widely noted by practitioners.
Jia Wang 0004, Xiaolin Wu 0001, Jun Sun 0005, Songyu Yu
IEEE Trans. Inf. Theory4
2012 Multiscale Semilocal Interpolation With Antialiasing
abstract
Aliasing is a common artifact in low-resolution (LR) images generated by a downsampling process. Recovering the original high-resolution image from its LR counterpart while at the same time removing the aliasing artifacts is a challenging image interpolation problem. Since a natural image normally contains redundant similar patches, the values of missing pixels can be available at texture-relevant LR pixels. Based on this, we propose an iterative multiscale semilocal interpolation method that can effectively address the aliasing problem. The proposed method estimates each missing pixel from a set of texture-relevant semilocal LR pixels with the texture similarity iteratively measured from a sequence of patches of varying sizes. Specifically, in each iteration, top texture-relevant LR pixels are used to construct a data fidelity term in a maximum a posteriori estimation, and a bilateral total variation is used as the regularization term. Experimental results compared with existing interpolation methods demonstrate that our method can not only substantially alleviate the aliasing problem but also produce better results across a wide range of scenes both in terms of quantitative evaluation and subjective visual quality.
Kai Guo 0001, Xiaokang Yang 0001, Hongyuan Zha, Weiyao Lin, Songyu Yu
IEEE Trans. Image Process.5
2011 An Error Resilient Video Coding Scheme Using Embedded Wyner-Ziv Description With Decoder Side Non-Stationary Distortion Modeling
abstract
In this paper, we propose a generic error resilient video coding (ERVC) scheme using embedded Wyner-Ziv (WZ) description. At the encoder side, a joint source-channel R-D optimized mode selection (JSC-RDO-MS) algorithm with WZ-coded anchor frames is statistically studied and developed. Given a stationary first-order Markov Gaussian source, the proposed mode optimization is justified by an analysis of the RD impact on the WZ bit-rate. JSC-RDO-MS involves in the estimation of expected rate and distortion of WZ coding with the unavailable side information, and the WZ bit-rate of each coding mode is determined based on the error correction capability of the specific WZ codec. At the decoder side, an online correlation noise model between the source and the side-information is proposed with a mixture of Laplacians whose parameters are attained to reflect the coherence of the motion field of successive frames and the energy of prediction residual. Each mixture component represents the statistical distribution of prediction residuals, and the mixing coefficients represent the amount of errors in motion compensation. The proposed scheme achieves the so-called classification gain by exploiting the spatially non-stationary characteristics of the motion field and texture. Extensive experimental results show that the proposed WZ-ERVC scheme achieves a better overall RD performance than existing ERVC schemes, and the proposed modeling algorithm also significantly outperforms the conventional Laplacian model by up to 2 dB.
Hongkai Xiong, Zhihai He, Songyu Yu, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.4
2011 Reconstruction for Distributed Video Coding: A Context-Adaptive Markov Random Field Approach
abstract
Within the existing reconstruction process of distributed video coding (DVC), there are two major approaches: the maximum probability reconstruction and the minimum mean square error (MMSE) reconstruction. Both of them assume that each node, a pixel in pixel domain DVC or a coefficient in transform domain DVC, is i.i.d., and reconstruct the value of each node independently by only exploiting statistical correlation between source and side-information. These kinds of models produce considerable amount of artifacts in decoded Wyner-Ziv (WZ) frames and degrade the objective performance. In this paper, we propose a context-adaptive Markov random field (MRF) reconstruction algorithm which exploits both the statistical correlation and the spatio-temporal consistency by modeling the corresponding MRF of a generic DVC architecture, and solve the inference by finding its MRF-based maximum a posteriori (MAP) estimate. The energy function of the MRF model consists of two terms: a data term measuring the statistical correlation, and a geometric regularity term enforcing local spatio-temporal structure consistency which is modeled by optical flow estimation with regard to the critical parameters under a wide variety of DVC scenarios. In case the unreliability of the derived local structure, a confidence parameter is introduced to prevent inappropriate penalizing. To find the reconstructed patch assignment with the largest expected probability in the context-adaptive MRF, the energy minimization for the MRF-based MAP estimate of the WZ frames is solved by global optimization and greedy strategies. Compared to the existing maximum probability and MMSE reconstruction with i.i.d. model, a better subjective and objective performance is validated by extensive experiments.
Hongkai Xiong, Zhihai He, Songyu Yu, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.4
2010 Scene categorization based on heterogeneous features
abstract
In this paper we present a complete framework for scene categorization that builds upon and extends several recent ideas including spatial pyramid representation and a variety of base local descriptors which have different discriminative power and invariance from task to task. Furthermore, we propose two strategies: sum-max and max-max, used to effectively combine diverse source of data in a unified setting way. Our approach shows significantly improved performance on a large, challenging data set of fifteen natural scene categories. Owing to combination of complementary information cues, our approach is expected to equally applicable to a range of tasks.
Fuxiang Lu, Xiaokang Yang 0001, Rui Zhang 0052, Songyu Yu
VCIP4
2010 Reconstruction for distributed video coding: a Markov random field approach with context-adaptive smoothness prior
abstract
An important issue in Wyner-Ziv video coding is the reconstruction of Wyner-Ziv frames with decoded bit-planes. So far, there are two major approaches: the Maximum a Posteriori (MAP) reconstruction and the Minimum Mean Square Error (MMSE) reconstruction algorithms. However, these approaches do not exploit smoothness constraints in natural images. In this paper, we model a Wyner-Ziv frame by Markov random fields (MRFs), and produce reconstruction results by finding an MAP estimation of the MRF model. In the MRF model, the energy function consists of two terms: a data term, MSE distortion metric in this paper, measuring the statistical correlation between side-information and the source, and a smoothness term enforcing spatial coherence. In order to better describe the spatial constraints of images, we propose a context-adaptive smoothness term by analyzing the correspondence between the output of Slepian-Wolf decoding and successive frames available at decoders. The significance of the smoothness term varies in accordance with the spatial variation within different regions. To some extent, the proposed approach is an extension to the MAP and MMSE approaches by exploiting the intrinsic smoothness characteristic of natural images. Experimental results demonstrate a considerable performance gain compared with the MAP and MMSE approaches.
Hongkai Xiong, Zhihai He, Songyu Yu
VCIP4
2009 Event recognition with time varying Hidden Markov Model
abstract
Standard hidden Markov model (HMM) and the more general dynamic Bayesian network (DBN) models assume stationarity of state transition distribution. However, this assumption does not hold for many real life events of interest. In this paper, we propose a new time sequence model that extends HMM to time varying scenario. The time varying property is realized in our model by explicitly allowing the change of state transition density as the time spent in a particular state passes by. Rather than keeping transition densities at different time spots independent of each other, we exploit their temporal correlation by applying a hierarchical Dirichlet prior. This leads to a more robust time varying model, especially when training data are scarce. We also employ Markov chain Monte Carlo (MCMC) sampling in learning the MAP estimate of time varying parameters, with a transition kernel incorporating linear optimization. The proposed model is applied to recognizing real video events, and is shown to outperform existing HMM-based methods.
Ercan E. Kuruoglu, Xiaokang Yang 0001, Yi Xu 0001, Songyu Yu
ICASSP5
2009 Interpolating fine texturess with fields of experts prior
abstract
Traditional image interpolation methods assume that the local spatial structure of the low-resolution (LR) and high-resolution (HR) images are approximately the same, and use edge information of the LR image to estimate the missing pixels. This assumption, however, no longer holds for natural images with fine and dense textures. Consequently, those methods cannot restore dense textures well and tend to generate over-fitting visual effects. In this paper, a learned HR image prior is exploited to overcome the problems. In particular, we use Fields of Experts (FoE) with student's t-distribution experts to model the prior, taking advantage of its representative ability of non-Gaussian natures in images. Then Maximum a Posterior (MAP) estimation incorporating FoE prior is used to estimate the missing pixels. Experimental results compared with traditional interpolation methods demonstrate that our method not only can recover fine details and produce superior PSNR values, but also avoid the visual over-fitting problems.
Kai Guo 0001, Xiaokang Yang 0001, Rui Zhang 0052, Songyu Yu, Hongyuan Zha
ICIP4
2009 Spatial non-stationary correlation noise modeling for Wyner-Ziv error resilience video coding
abstract
Most of the Wyner-Ziv (WZ) video coding schemes in literature model the correlation noise (CN) between original frame and side information (SI) by a given distribution whose parameters are estimated in an offline process. In this paper, an online CN modeling algorithm is proposed towards a more practical WZ-based error resilient video coding (WZ-ERVC). In ERVC scenario, the side-information is typically generated from the error concealed picture instead of bi-directional motion prediction. The proposed online CN modeling algorithm achieves the so-called classification gain by exploiting the spatially non-stationary characteristics of the motion field and texture. The CN between the source and error concealed SI is modeled by a Laplacian mixture model, where each mixture component represents the statistical distribution of prediction residuals and the mixing coefficients portray the motion vectors estimation error. Experimental results demonstrate significant performance gains both in rate and distortion versus the conventional Laplacian model.
Hongkai Xiong, Li Song 0001, Songyu Yu
ICIP4
2009 Learning super resolution with global and local constraints
abstract
In learning based single image super-resolution (SR) approach, the super-resolved image are usually found or combined from training database through patch matching. But because the representation ability of small patch is limited, it is difficult to guarantee that the super-resolved image is best under global view. To tackle this problem, we propose a statistical learning method for SR with both global and local constraints. Firstly, we use maximum a posteriori (MAP) estimation with learned image priors by fields of experts (FoE) model, and regularize SR globally guided by the image priors. Secondly, for each overlapped patch, the higher-order Markov random fields (MRFs) is used to model its local relationship with corresponding high-resolution candidates, then belief propagation is used to find high-resolution image. Compared with traditional patch based learning method without global constraint, our method could not only preserve the global image structure, but also restore the local details well. Experiments verify the idea of our global and local constraint SR method.
Kai Guo 0001, Xiaokang Yang 0001, Rui Zhang 0052, Songyu Yu
ICME4
2009 Image classification based on pyramid histogram of topics
abstract
In this paper we propose PHOTO (pyramid histogram of topics), a new representation for image classification. We partition the image into hierarchical cells and learn the topic histogram using pLSA over each cell with EM algorithm. Then we concatenate the topic histograms over the cells at all levels to form a ldquolongrdquo vector, i.e. pyramid histogram of topics. Finally AdaBoost classifiers are used to select the topics most discriminative for class recognition. Experimental results on two diverse databases show that our method performs significantly better than general topic representation.
Fuxiang Lu, Xiaokang Yang 0001, Rui Zhang 0052, Songyu Yu
ICME4
2009 An improved block size selection method based on macroblock movement characteristic
Jun Sun 0005, Rong Xie 0004, Songyu Yu, Wenjun Zhang 0001
Multim. Tools Appl.4
2009 CamShift guided particle filter for visual tracking
Xiaokang Yang 0001, Yi Xu 0001, Songyu Yu
Pattern Recognit. Lett.4
2008 A Novel Multiple Description Video Codec Based on Slepian-Wolf Coding
abstract
One major task in multiple description video coding is to prevent drift on packet loss channels, where transmission errors occur in each description. We propose a distributed multiple description video coding (DMDVC) scheme excluding any prediction loops. The new codec suffers from no drift problem. In the two-channel mode of symmetry side informations (SI), one side decoder can use the SI of the other without any decoding quality degradation. Thereby, DMDVC achieves high robustness on packet loss channels.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005, Peng Wang 0026, Songyu Yu
DCC5
2008 Confidence based optical flow algorithm for high reliability
abstract
Estimation of optical flow is an important topic to provide motion information for motion analysis. This paper addresses an effective confidence based optical flow algorithm. It considers the bidirectional symmetry of forward and backward flow to compute the confidence measure for each flow estimate. According to the confidence, the reliable flow estimates have greater contribution to local averages while unreliable estimates are suppressed. The errors cannot be propagated. Since the image-driven and flow-driven discontinuity preserving methods have complementary advantages and limitations, we propose a region based method combining these two types of methods to preserve motion boundaries. Experiments on typical sequences have successfully demonstrated the validity of the proposed algorithm.
Songyu Yu
ICASSP2
2008 Variable block size selection for a transcoder based on MB movement information
abstract
To improve coding performance, the new video compression standard H.264 employs seven variable block sizes for one Macro Block (MB) to conduct motion estimation and compensation. MPEG-2 only has one size 16×16. This paper presents a novel fast variable block size selection method for an inter-MB in a video transcoder from MPEG-2 to H.264 with downscaling by a factor two in each dimension based on the MB motion information. Without conducting motion re-estimation, an optimal block size is decided. Experiment results show that this method saves the transcoder complexity dramatically with little compression performance degradation.
Jun Sun 0005, Rong Xie 0004, Shibao Zheng, Songyu Yu
ICME5
2008 Face super-resolution using 8-connected Markov Random Fields with embedded prior
abstract
In patch based face super-resolution method, the patch size is usually very small, and neighbor patchespsila relationship via overlapped regions is only to keep smoothness of reconstructed high-resolution image, so the prior is not always strong enough to regularize super-resolution when observed low-resolution image lose facial structure information. We propose to use Gaussian Mixture Model(GMM) to learn facial prior embedded between un-overlapped regions of neighbor patches. This approach, which has never been used to regularize face super-resolution before, usually works as a potential function in 8-connected Markov Random Fields (MRFs) with belief propagation. In the proposed algorithm, we assign high probability to the neighbor candidate patches that express correct facial structure, and others not. Experiments demonstrate that our method is superior in preserving smoothness and recovers facial structure and local details when low-resolution image lost the details of facial structure.
Kai Guo 0001, Xiaokang Yang 0001, Rui Zhang 0052, Guangtao Zhai, Songyu Yu
ICPR5
2008 Contourlet-based image adaptive watermarking
Songyu Yu, Xiaokang Yang 0001, Li Song 0001, Chen Wang 0053
Signal Process. Image Commun.2
2008 GOP-level transmission distortion modeling for mobile streaming video
Hua Yang 0001, Songyu Yu, Xiaokang Yang 0001
Signal Process. Image Commun.3
2007 On Multi-Stage Sequential Coding of Correlated Sources
abstract
We study the problem of multi-stage sequential coding (MSSC), which is an extension of sequential coding of correlated sources. Consider two correlated random variables X and Y to be coded in two stages. The first stage is sequential coding as referred to in the existing literature. At the second stage, the Y encoder refines the information of Y without any knowledge of X, and X encoder refines the information of X with the knowledge of Y, while all previous outputs are known at the decoder. As the sequential coding problem provides a theoretical abstraction of video coding, the MSSC model is a theoretical abstraction of scalable video coding, which is an important application of network communications. We give an achievable region for the MSSC system. The given achievable region is tight when Y is required to be reconstructed perfectly in the usual Shannon sense at the second stage. We also study the minimum total rate MSSC problem, and derive the minimum total rate for Gaussian sources. This result disproves the possibility that the minimum total rate of one stage sequential coding can be achieved at both stages even for correlated Gaussian sources. Thus we offer a theoretical explanation for the performance loss of scalable video coding widely noted by practitioners
Jia Wang 0004, Xiaolin Wu 0001, Jun Sun 0005, Songyu Yu
DCC4
2007 The Wyner-Ziv Rate-Distortion Function of Multivariate Gaussian Sources and Its Application in Distributed Video Coding
abstract
Wyner-Ziv coding is presented in this paper. It is extended to the scenario of multivariate source and side information, whose rate-distortion function is obtained by a reverse water-filling method for the joint quadratic-Gaussian case.
Peng Wang 0026, Jia Wang 0004, Songyu Yu, Erkang Chen, Xiaokang Yang 0001
DCC3
2007 Multiframe Super-Resolution Reconstruction Based on Cycle-Spinning
abstract
A multiframe super-resolution (SR) reconstruction algorithm based on cycle-spinning (CS) is proposed. We utilize the relative motion information of sequential images to construct a CS-based framework for the resolution enhancement. The unique feature of the proposed algorithm is that it is effective for low-resolution (LR) images with various point spread function (PSF) and noise characteristics, even if the degradation models are unknown for the imaging system. Moreover, the computational complexity is inexpensive. Experiments demonstrate the effectiveness of the proposed method and show the superiority to previous methods in objective and subjective qualities.
Xiangzhong Fang, Songyu Yu
ICASSP (1)4
2007 Cross-Layer Frame Discarding for Cellular Video Coding
abstract
In the case of delivering real-time video over the 3G cellular networks, burst frame losses may be inevitable and unpredictable, which may cause severe quality degradation. Based on cross-layer frame discarding (CLFD), this paper proposes an enhanced error-resilient video coding scheme for cellular video communication. By using unequal retransmission at the radio link (RL) layer, a base station can provide reliable transmission for the relatively important frames in one video sequence. Relying on the unequal protection at the RL layer, the encoder at the application (APP) layer can actively discard a certain number of frames according to the received acknowledgement messages. Thus, unpredictable burst frame losses during transmission can be transformed into selective frame discarding at the encoder. Experiments results show that the proposed scheme can enhance the error resilience of the cellular video communication significantly.
Songyu Yu, Hua Yang 0001, Hongkai Xiong
ICASSP (2)2
2007 A Source-Driven Error Recovery Scheme using Wyner-Ziv Coding
abstract
In this paper, we propose an error recovery scheme to cope with the frame loss problem in large end-to-end delay scenario. Because Wyner-Ziv coding can produce deterministic output with nondeterministic inputs, we adopt the decoder's error-corrupted reference picture as side information to exploit its correlation with source picture, and thus improve the coding efficiency. To prevent retransmission request, we propose a feasible encoder-driven rate estimation scheme by only storing MVs and part of the error pattern at encoder. Furthermore, a new Laplacian parameter computing method is proposed based on discrete Laplacian PDF. The experimental results show that the proposed estimation schemes have quite high precision, and the error recovery scheme outperforms INTRA refresh scheme up to 5 dB.
Hongkai Xiong, Songyu Yu, Hui Lu 0001
ICME3
2007 GOP-Level Transmission Distortion Modeling for Unequal Importance Judgement
abstract
In the case of transmitting stored video streaming over packet-switched networks, unavoidable frame losses may result in error propagation of reconstructed video and thus induce severe quality degradation. In order to minimize the transmission distortion and utilize the limited channel resources efficiently, unequal error protection is usually adopted to exploit the unequal importance of different frames in one group-of-pictures (GOP). In this paper, we develop an estimate model for GOP-level transmission distortion so that the transmitter is able to find the unequal importance of different frames in one GOP. The simulation results demonstrate that the proposed model is accurate and robust.
Hua Yang 0001, Songyu Yu, Xiaokang Yang 0001, Hao Liu 0010
ICME3
2007 Multiple Descriptions with Side Informations Also Known At the Encoder
abstract
We propose a new scheme of multiple descriptions with side information (SI). The two side decoders of the system use two different SI streams. Both SI streams are available to the central decoder and to the encoder. We give an inner bound for this system for general source and SI. The tight bound is obtained for the quadratic Gaussian case. This result is compared with our previous result of the MDWZ (multiple descriptions in the Wyner-Ziv setting) problem in which none of the side information is known at the encoder. It is shown that when side information is absent at the encoder, there is a performance loss. The proposed scheme and its achievable region have practical significance. It offers theoretical insight into the multiple description video coding (MDVC) and suggests an optimal coding strategy.
Jia Wang 0004, Xiaolin Wu 0001, Songyu Yu, Jun Sun 0005
ISIT3
2007 Multiplexed harmonic broadcasting scheme for efficient video-on-demand services
abstract
Periodic broadcasting schemes can improve the efficiency of video-on-demand (VOD) services by reducing the bandwidth requirement to transmit popular videos. The harmonic broadcasting scheme has the best performance in reducing service bandwidth under a given access time, but it uses too many channels. However, a multiplexed harmonic broadcasting scheme that overcomes this drawback is now proposed. This scheme divides each video into equal-sized segments and then broadcasts segments periodically in a small number of sever channels with equal bandwidth. The idea of segment-to-channel mapping in the scheme is inspired by the time division multiplexing system. Each segment is divided equally into several subsegments; subsegments of different segments are multiplexed in a slot with the guarantee of being able to keep playing out continuously for every user. The proposed scheme outperforms the pagoda broadcasting and recursive frequency splitting schemes in reducing the viewers' maximum waiting time, and the scheme requires less client storage.
Yunqiang Liu, Songyu Yu, Xinbing Wang
IET Commun.2
2006 Channel-Aware Frame Dropping for Cellular Video Streaming
abstract
In the case of cellular video streaming over wireless channels, burst frame losses may be unavoidable. Considering the unequal importance of different frames in a group-of-pictures (GOP) and the burst-error characteristics of wireless channels, this paper proposes a channel-aware frame dropping scheme so as to shift burst losses into relatively unimportant frames in the same GOP. By using selective retransmission at the radio link layer, a base station can adaptively assign the unequal transmission attempts to different video frames. Simulation results show that the proposed scheme can be aware of the variation of wireless channel conditions, and thus significantly improve error resilience of cellular video streaming
Hao Liu 0010, Wenjun Zhang 0001, Songyu Yu, Xiaokang Yang 0001
ICASSP (5)3
2006 Motion Vector Smoothing for True Motion Estimation
abstract
This paper proposes a new motion vector (MV) smoothing algorithm to track the real motion in image sequences for MPEG video encoders. First, a pre-checking algorithm is employed to eliminate wrong motion vectors and preserve all possible motion vectors. For each block considered, the motion similarity between the neighboring blocks and the number of candidate motion vectors are jointly exploited to adaptively grow the filtering support, which is supposed to have homogeneous motion and sufficient spatial gradient. Then, all candidate motion vectors are checked within the filtering support using a new motion smoothness-constrained matching criteria. The simulation results show that the proposed algorithm can efficiently track the real motion resulting in smooth motion vector field (MVF).
Hai Bing Yin, Xiangzhong Fang, Hua Yang 0001, Songyu Yu, Xiaokang Yang 0001
ICASSP (2)4
2006 Multi-Rate, Dynamic and Compliant Region of Interest Coding for JPEG2000
abstract
A method is proposed to encode multiple regions of interest (ROI) in JPEG2000 image. It rearranges truncation point for every codeblock in each layer. It assigns higher bitrate to ROI and lower bitrate to non-ROI and combines them to codestream. The proposed strategy produces a fully compliant JPEG2000 codestream. It allows transmission of different ROIs with different priorities and supports dynamic delineation and prioritization of them. Experimental results demonstrating the validity of the proposed approach are presented
Xiangzhong Fang, Haibin Yin, Songyu Yu
ICME5
2006 A Context-Based Error Detection Strategy into H.264/AVC CABAC
abstract
Various error control schemes have been addressed in wireless video stream transmission. By combining an adaptive binary arithmetic coding technique with context modeling, CABAC as a normative part of H.264/AVC has achieved a high degree of adaptation and redundancy reduction. However, error propagation still remains a problem because of the property of arithmetic coding. The presented scheme compares the various error detection methods, and proposes an efficient error detection technique based on CABAC semantics, which is achieved by inserting detective markers denoting by syntax elements. The misdetection probability versus stream size expansion can be easily handled. In addition, placements of markers can vary with regard to specific video content, thus efficiency within this scheme is enhanced. Comparison with other detection scheme is also presented
Hongkai Xiong, Li Song 0001, Songyu Yu
ICME4
2006 A New Deblocking Algorithm Based on Adjusted Contourlet Transform
abstract
A new postprocessing method based on adjusted contourlet transform is introduced in this paper for suppressing blocking artifacts (BA) in block-based discrete cosine transform (BDCT) compressed images. To our best knowledge, this is the first time contourlet is applied to this field. By exploiting scale space edge detector (ss-edge detector), our algorithm can extract and protect blocking map (BM) and edge map (EM) in the compressed image respectively in the same time. By transforming the compressed image into adjusted contourlet domain, the adaptive thresholds are obtained according to BM. According to the adaptive thresholds, the contourlet coefficients in different subbands are filtered. Experimental results show that our deblocking algorithm achieves better performance than the other iterative and noniterative methods reported in the literature
Songyu Yu, Chen Wang 0053, Li Song 0001, Hongkai Xiong
ICME2
2006 Unequal Iterative Decoding for Power Efficient Video Transmission
abstract
We present an unequal iterative decoding (UID) approach for minimization of the receiver power consumption subject to a given quality of service, by exploiting data partitioning and turbo decoding. We assign unequal decoding iterations of forward error correction (FEC) to data partitions with different importance by jointly considering the source coding, channel coding and receiver power allocation. The proposed scheme has been applied to H.264 video over AWGN channel, and achieves excellent tradeoff between video delivery quality and power consumption, and yields significant power saving compared with the classical equal iterative decoding (EID) approach in wireless video transmission
Songyu Yu, Xiaokang Yang 0001
ICME2
2006 Perceptually-adaptive Motion Compensated Temporal Filtering
abstract
We propose a perceptually-adaptive motion compensated temporal filtering (MCTF) method to enhance the visual quality of 3D wavelet video coding schemes with spatial-domain MCTF. In our scheme, a spatio-temporal masking model in image domain is incorporated into the lifting structure of MCTF. The model is used to guide the motion search and the prediction step in MCTF so as to remove the visual redundancy in the video sequence. Experimental results show that the proposed scheme can significantly improve the visual quality of decoded video at different bitrates
Wenjun Zhang 0001, Xiaokang Yang 0001, Songyu Yu
ICME5
2006 Multiple Descriptions in the Wyner-Ziv Setting
abstract
We propose a new scheme of multiple descriptions in the Wyner-Ziv setting (MD-WZ). The two side decoders of MD-WZ use two different side information (SI) streams. Both SI streams are available to the central decoder, but none to the encoder. We derive an achievable region (inner bound) for this MD-WZ system for general source and SI. If the source and SI are correlated Gaussian and for quadratic distortion metric, the tight bound is obtained. Our result is an extension of Ozarow's result on multiple descriptions of Gaussian source without SI. The MD-WZ coding scheme is shown to have a property of practical significance. For symmetric case where the joint distributions of the source and the two SI are the same and the two channels are balanced, interchanging the two channels causes no performance loss for Gaussian source. Considering that the existing multi-description video coding methods suffer from the notorious drifting problem induced by channels interchange, this work lends a theoretical support to distributed multi-description video coding in the Wyner-Ziv setting
Jia Wang 0004, Xiaolin Wu 0001, Songyu Yu, Jun Sun 0005
ISIT3
2006 Error robustness scheme for H.264 based on LDPC code
abstract
With the tremendous increase in the capabilities of wireless multimedia devices and services, the demand to improve quality of such systems within the limited bandwidth resources motivates the interest in robust video transmission. A new unequal error protect (UEP) scheme is proposed based on LDPC code and data partitioning, which can protect much more important bits and correct major error during the transmission over AWGN channel. Our proposed algorithm features assigning an unequal bits number of forward error correction (FEC) to each data partition for H.264, thus, efficiently makes use of available resources. Simulation results demonstrate that this proposed algorithm is more error robust compared to those obtained with the classical EEP technique and common coding mode, such as Turbo code.
Songyu Yu, Xiaokang Yang 0001
MMM2
2006 Adaptive segment-based patching scheme for video streaming delivery system
Yunqiang Liu, Songyu Yu
Comput. Commun.2
2005 A client-driven scalable cross-layer retransmission scheme for 3G video streaming
abstract
The wireless channel is time-varying where burst packet losses often occur during the fading or lossy handovers. In order to avoid unaccepted quality degradation of video streaming over 3G cellular networks, we propose and analyze a client-driven scalable cross-layer (CSC) retransmission scheme. Considering the perceptual importance of different video partitions under the real-time and bandwidth constraints, the proposed scheme uses the radio link-layer retransmission with priority to adapt conventional packet losses in wireless channels; furthermore, it uses the adaptive transport-layer retransmission to provide end-to-end quality-of-service (QoS) guarantees over cellular networks. The simulation experiments show that the proposed scheme can effectively improve the perceptual quality of 3G video streaming as compared to the traditional deadline-based scheme without the prioritized link-layer retransmission.
Hao Liu 0010, Wenjun Zhang 0001, Songyu Yu
ICME3
2005 An Adaptive UEP_BTC_STBC System for Robust H.264 Video Transmission
abstract
A new adaptive UEP_BTC_STBC scheme is proposed to guarantee the robust video transmission according to the channel conditions. This scheme enhanced STBC (space-time block coding) by concatenating BTC code (block turbo code), both the good error correcting capability of BTC and the concurrent large diversity gain characteristic of STBC can be achieved simultaneously. Furthermore, by employing different BTC codes together with different modulation approaches, the scheme is capable of adapting to channel conditions and maximize end-to-end QoS of video transmission. Simulation result shows that the proposed adaptive scheme achieved significant improvement in delivered video quality according to channel condition.
Yinggang Du, Songyu Yu, Kam Tai Chan, Yantao Qiao
ICME3
2005 Streaming Media Delivery with Proxy Cache for Heterogeneous Clients
abstract
Efficient streaming media delivery scheme used multimedia service can largely increase system service ability by reducing the resource requirement. Because of the heterogeneity in the underlying network environments, streaming media delivery scheme should provide different, appropriate video quality to serve clients with different bandwidth. In this paper, we present an efficient streaming media delivery scheme with the aid of proxy caching to delivery layered encoded video for heterogeneous clients. The threshold-based multicast technique is used to delivery video streams form server to proxy via backbone network. In order to reduce the bandwidth requirement of the backbone network, we develop an effective approach to determine which layers of which videos should be cached according to the request rates. A simple and efficient replacement algorithm is presented to deal with the varying request rates. Simulation results demonstrate that our scheme achieve significantly bandwidth reduction
Yunqiang Liu, Songyu Yu
MMSP2
2005 A New Hybrid Delivery Scheme for Efficient Video-on-Demand Services
abstract
In video-on-demand (VOD) system, a periodic broadcast technique like fast broadcasting scheme has been shown to be very effective for serving a popular video in reducing the demand on server bandwidth. A reactive server transmission approach like patching scheme is more suitable for the video that is not popular enough. However, these approaches are not suited for the instance that the level of demand on a video changing greatly with time. In this paper, we propose a novel adaptive hybrid delivery scheme which allocates adaptively transmission resources according to the varying client request rate. The scheme tries to dynamically search the optimal number of channel assigned to the video by the newly updated request rate so as to minimize the bandwidth requirement. Simulation results show that the scheme improves the performance of VOD service significantly in terms of total server bandwidth requirement
Yunqiang Liu, Songyu Yu
MMSP2
2005 An Adaptive UEP_BTC_STBC_OFDM System for Robust Video Transmission
abstract
An adaptive UEP_BTC_STBC_OFDM system is proposed to provide robust video transmission in dispersive fading channel. The system concatenates the Block Turbo Code (BTC) with the Space-Time Block Code (STBC) for an OFDM system, both the good error correcting capability of BTC and the concurrent large diversity gain characteristic of STBC can be achieved simultaneously. Furthermore, by employing the data partition of H.264 and different BTC codes together with different modulation approaches, the scheme is capable of adapting to channel conditions and maximize end-to-end QoS of video transmission with low encoding and decoding complexity. Simulation result shows that the proposed adaptive scheme achieved significant improvement in delivered video quality for specified channel condition and thus has better performance of video transmission
Yinggang Du, Songyu Yu
MMSP3
2005 Flashlight Scene Detection for MPEG Videos
abstract
A new flashlight scene detection approach is presented. It focuses on two techniques: a new flashlight model based on the spatial-temporal characteristics of intensity value for the DC sequence, and a local threshold selection scheme using the sliding window. The advantages of this approach are its good performance for multi-frame and gradual flashlight scene cases, and local threshold selection. Experimental results show that the proposed algorithm is fast, robust and high accuracy
Yanling Xu, Songyu Yu, Yuanhua Zhou
MMSP3
2005 Joint Source-Channel Decoding forH.264 Coded Video Stream
abstract
H.264 is the newest video coding standard and has achieved a significant improvement in coding efficiency. The entropy coding methods used in H.264 are CAVLC and CABAC. Although these two variable length code methods can achieve high compression, they are very sensitive to channel errors. This paper presents a joint source-channel MAP (maximum a posteriori probability) decoding method to dealing with this sensitivity to channel errors and applied it to the decoding of the motion vector in H.264 coded video stream. Although H.264 codec has proposed several error resilience methods, we believe this method could provide additional error resilience to H.264 stream. Experiment indicates that our JSCD achieves significant improvement than a separate scheme
Songyu Yu
MMSP2
2004 A highly efficient, low delay architecture for transporting H.264 video over wireless channel
Hong-Bin Yu, Songyu Yu, Ci Wang
Signal Process. Image Commun.2
2003 1-D and 2-D transforms from integers to integers
abstract
Substituting a real valued linear transform with an integer-to-integer mapping has become very important in lots of applications. This paper introduces a new kind of matrix decomposition method called lifting-like factorization, which leads to a theorem: every 2/sup n/-order real matrix with determinant norm 1 can be expressed as the product of one permutation matrix and at most three unit triangular matrices. Rounding error of this method is analyzed. Realization of 2D integer transform is also studied and it is shown that a 2D integer-to-integer transform cannot be realized by performing two 1D integer transforms separately. Left and right permutation matrices are introduced to reduce rounding error and an application of this method to intDCT is discussed.
Jia Wang 0004, Jun Sun 0005, Songyu Yu
ICASSP (2)3
2003 Video multicast based on estimation of consumed average server bandwidth
abstract
All the previous patching schemes assume the receive-two model (a client receives two streams simultaneously). The proposed scheme in this paper is the first no attempt receive-three model (a client receives three streams simultaneously) for patching. We patch the streams based on estimation of the average consumed server bandwidth. Simulation results show that the proposed patching scheme outperforms the conventional patching technique by a significant margin. It even performs better than the dynamic skyscraper algorithm over a wide range of client request rates. Most importantly, the implementation complexity of our algorithm is much lower than the skyscraper and hierarchical multicast stream merging (HMSM).
Dongliang Guan, Songyu Yu
PIMRC2
2003 A new subspace algorithm of blind channel estimation for OFDM systems
abstract
This paper proposes a statistical subspace algorithm for blind channel estimation in OFDM systems. The algorithm is based on a new input/output relationship which is obtained by a matrix transform on the transmission equation of the OFDM systems. The proposed algorithm is computationally simpler, and extends some results of the existing subspace algorithm. The simulation results are also presented.
Xuejun Huang, Houjie Bi, Songyu Yu
PIMRC3
2002 Modified wavelet coding of arbitrarily shaped objects based on extrapolation and reflection (EAR)
Jia Wang 0004, Jun Sun 0005, Songyu Yu
VCIP3
2002 Wavelet image coding based on directional dilation
Jia Wang 0004, Songyu Yu, Jun Sun 0005
VCIP2
2000 Design and implementation of the second-generation HDTV prototype video encoder of China
Jun Sun 0005, Zhenghua Yu, Songyu Yu
VCIP4
2000 Lost motion vector recovery for digital video communication
Zhenghua Yu, Henry R. Wu, Songyu Yu
VCIP3