EDBT 2026 Demo / reviewers in the wild / expert
Xun Guo 0002
dblp:32/5851-2
· DBLP profile ↗
26ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | I2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video ModelsabstractRecent advances in image-to-video generation have enabled animation of still images and offered pixel-level controllability. While these models hold great potential to transform single images into vivid and dynamic videos, they also carry risks of misuse that could impact privacy, security, and copyright protection. This paper proposes a novel approach that applies imperceptible perturbations on images to degrade the quality of the generated videos, thereby protecting images from misuse in white-box image-to-video diffusion models. Specifically, we function our approach as an adversarial attack, incorporating spatial, temporal, and diffusion attack modules. The spatial attack shifts image features from their original distribution to a lower-quality target distribution, reducing visual fidelity. The temporal attack disrupts coherent motion by interfering with temporal attention maps that guide motion generation. To enhance the robustness of our approach across different models, we further propose a diffusion attack module leveraging contrastive loss. Our approach can be easily integrated with mainstream diffusion-based I2V models. Extensive experiments on SVD, CogVideoX, and ControlNeXt demonstrate that our method significantly impairs generation quality in terms of visual clarity and motion consistency, while introducing only minimal artifacts to the images. To the best of our knowledge, we are the first to explore adversarial attacks on image-to-video generation for security purposes. Dongnan Gui, Xun Guo 0002, Wengang Zhou 0001, Yan Lu 0001 |
CVPR | 2 |
| 2025 | Image as a World: Generating Interactive World from Single Image via Panoramic Video GenerationabstractGenerating an interactive visual world from a single image is both challenging and practically valuable, as single-view inputs are easy to acquire and align well with prompt-driven applications such as gaming and virtual reality. This paper introduces a novel unified framework, Image as a World (**IaaW**), which synthesizes high-quality 360-degree videos from a single image that are both controllable and temporally continuable. Our framework consists of three stages: world initialization, which jointly synthesizes spatially complete and temporally dynamic scenes from a single view; world exploration, which supports user-specified viewpoint rotation; and world continuation, which extends the generated scene forward in time with temporal consistency. To support this pipeline, we design a visual world model based on generative diffusion models modulated with spherical 3D positional encoding and multi-view composition to represent geometry and view semantics. Additionally, a vision-language model (IaaW-VLM) is fine-tuned to produce both global and view-specific prompts, improving semantic alignment and controllability. Extensive experiments demonstrate that our method produces panoramic videos with superior visual quality, minimal distortion and seamless continuation in both qualitative and quantitative evaluations. To the best of our knowledge, this is the first work to generate a controllable, consistent, and temporally expandable 360-degree world from a single image. Dongnan Gui, Xun Guo 0002, Wengang Zhou 0001, Yan Lu 0001 |
NeurIPS | 2 |
| 2024 | MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingabstractRecently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory cost, and long-term temporal connection impose additional challenges. Taking advantage of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination with our specially designed memory mechanism, we propose the MovieChat to overcome these challenges. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1K benchmark with 1K long video and 14K manual annotations for validation of the effectiveness of our method. The code, models and data can be found in https://reself.github.io/MovieChat. Enxin Song, Wenhao Chai, Guanhong Wang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo 0002, Tian Ye 0001, Yanting Zhang 0001, Yan Lu 0001, Jenq-Neng Hwang, Gaoang Wang |
CVPR | 8 |
| 2023 | StableVideo: Text-driven Consistency-aware Diffusion Video EditingabstractDiffusion-based methods can generate realistic images and videos, but they struggle to edit existing objects in a video while preserving their appearance over time. This prevents diffusion models from being applied to natural video editing in practical scenarios. In this paper, we tackle this problem by introducing temporal dependency to existing text-driven diffusion models, which allows them to generate consistent appearance for the edited objects. Specifically, we develop a novel inter-frame propagation mechanism for diffusion video editing, which leverages the concept of layered representations to propagate the appearance information from one frame to the next. We then build up a text-driven video editing framework based on this mechanism, namely StableVideo, which can achieve consistency-aware video editing. Extensive experiments demonstrate the strong editing capability of our approach. Compared with state-of-the-art video editing methods, our approach shows superior qualitative and quantitative results. Our code is available at this https URL. Wenhao Chai, Xun Guo 0002, Gaoang Wang, Yan Lu 0001 |
ICCV | 2 |
| 2022 | Rethinking Minimal Sufficient Representation in Contrastive LearningabstractContrastive learning between different views of the data achieves outstanding success in the field of self-supervised representation learning and the learned representations are useful in broad downstream tasks. Since all supervision information for one view comes from the other view, contrastive learning approximately obtains the minimal sufficient representation which contains the shared information and eliminates the non-shared information between views. Considering the diversity of the downstream tasks, it cannot be guaranteed that all task-relevant information is shared between views. Therefore, we assume the non-shared task-relevant information cannot be ignored and theoretically prove that the minimal sufficient representation in contrastive learning is not sufficient for the downstream tasks, which causes performance degradation. This reveals a new problem that the contrastive learning models have the risk of overfitting to the shared information between views. To alleviate this problem, we propose to increase the mutual information between the representation and input as regularization to approximately introduce more task-relevant information, since we cannot utilize any downstream task information during training. Extensive experiments verify the rationality of our analysis and the effectiveness of our method. It significantly improves the performance of several classic contrastive learning models in downstream tasks. Our code is available at https://github.com/Haoqing-Wang/InfoCL. Haoqing Wang, Xun Guo 0002, Zhi-Hong Deng 0001, Yan Lu 0001 |
CVPR | 2 |
| 2022 | Semantic-aligned Fusion Transformer for One-shot Object DetectionabstractOne-shot object detection aims at detecting novel objects according to merely one given instance. With extreme data scarcity, current approaches explore various feature fusions to obtain directly transferable meta-knowledge. Yet, their performances are often unsatisfactory. In this paper, we attribute this to inappropriate correlation methods that misalign query-support semantics by overlooking spatial structures and scale variances. Upon analysis, we leverage the attention mechanism and propose a simple but effective architecture named Semantic-aligned Fusion Transformer (SaFT) to resolve these issues. Specifically, we equip SaFT with a vertical fusion module (VFM) for cross-scale semantic enhancement and a horizontal fusion module (HFM) for cross-sample feature fusion. Together, they broaden the vision for each feature point from the support to a whole augmented feature pyramid from the query, facilitating semantic-aligned associations. Extensive experiments on multiple benchmarks demonstrate the superiority of our framework. Without fine-tuning on novel classes, it brings significant performance gains to one-stage baselines, lifting state-of-the-art results to a higher level. Xun Guo 0002, Yan Lu 0001 |
CVPR | 2 |
| 2022 | Alignment-guided Temporal Attention for Video Action RecognitionabstractTemporal modeling is crucial for various video learning tasks. Most recent approaches employ either factorized (2D+1D) or joint (3D) spatial-temporal operations to extract temporal contexts from the input frames. While the former is more efficient in computation, the latter often obtains better performance. In this paper, we attribute this to a dilemma between the sufficiency and the efficiency of interactions among various positions in different frames. These interactions affect the extraction of task-relevant information shared among frames. To resolve this issue, we prove that frame-by-frame alignments have the potential to increase the mutual information between frame representations, thereby including more task-relevant information to boost effectiveness. Then we propose Alignment-guided Temporal Attention (ATA) to extend 1-dimensional temporal attention with parameter-free patch-level alignments between neighboring frames. It can act as a general plug-in for image backbones to conduct the action recognition task without any model-specific design. Extensive experiments on multiple benchmarks demonstrate the superiority and generality of our module. Xun Guo 0002, Yan Lu 0001 |
NeurIPS | 3 |
| 2021 | SSAN: Separable Self-Attention Network for Video Representation LearningabstractSelf-attention has been successfully applied to video representation learning due to the effectiveness of modeling long range dependencies. Existing approaches build the dependencies merely by computing the pairwise correlations along spatial and temporal dimensions simultaneously. However, spatial correlations and temporal correlations represent different contextual information of scenes and temporal reasoning. Intuitively, learning spatial contextual information first will benefit temporal modeling. In this paper, we propose a separable self-attention (SSA) module, which models spatial and temporal correlations sequentially, so that spatial contexts can be efficiently used in temporal modeling. By adding SSA module into 2D CNN, we build a SSA network (SSAN) for video representation learning. On the task of video action recognition, our approach outperforms state-of-the-art methods on Something-Something and Kinetics-400 datasets. Our models often outperform counterparts with shallower network and fewer modalities. We further verify the semantic learning ability of our method in visual-language task of video retrieval, which showcases the homogeneity of video representations and text embeddings. On MSR-VTT and Youcook2 datasets, video representations learnt by SSA significantly improve the state-of-the-art performance. Xun Guo 0002, Yan Lu 0001 |
CVPR | 2 |
| 2021 | Self-Supervised Video Representation Learning with Meta-Contrastive NetworkabstractSelf-supervised learning has been successfully applied to pre-train video representations, which aims at efficient adaptation from pre-training domain to downstream tasks. Existing approaches merely leverage contrastive loss to learn instance-level discrimination. However, lack of category information will lead to hard-positive problem that constrains the generalization ability of this kind of methods. We find that the multi-task process of meta learning can provide a solution to this problem. In this paper, we propose a Meta-Contrastive Network (MCN), which combines the contrastive learning and meta learning, to enhance the learning ability of existing self-supervised approaches. Our method contains two training stages based on model-agnostic meta learning (MAML), each of which consists of a contrastive branch and a meta branch. Extensive evaluations demonstrate the effectiveness of our method. For two downstream tasks, i.e., video action recognition and video retrieval, MCN outperforms state-of-the-art approaches on UCF101 and HMDB51 datasets. To be more specific, with R(2+1)D backbone, MCN achieves Top-1 accuracies of 84.8% and 54.5% for video action recognition, as well as 52.5% and 23.7% for video retrieval. Yuanze Lin, Xun Guo 0002, Yan Lu 0001 |
ICCV | 2 |
| 2018 | An End-to-End Compression Framework Based on Convolutional Neural NetworksabstractDeep learning, e.g., convolutional neural networks (CNNs), has achieved great success in image processing and computer vision especially in high-level vision applications, such as recognition and understanding. However, it is rarely used to solve low-level vision problems such as image compression studied in this paper. Here, we move forward a step and propose a novel compression framework based on CNNs. To achieve high-quality image compression at low bit rates, two CNNs are seamlessly integrated into an end-to-end compression framework. The first CNN, named compact convolutional neural network (ComCNN), learns an optimal compact representation from an input image, which preserves the structural information and is then encoded using an image codec (e.g., JPEG, JPEG2000, or BPG). The second CNN, named reconstruction convolutional neural network (RecCNN), is used to reconstruct the decoded image with high quality in the decoding end. To make two CNNs effectively collaborate, we develop a unified end-to-end learning algorithm to simultaneously learn ComCNN and RecCNN, which facilitates the accurate reconstruction of the decoded image using RecCNN. Such a design also makes the proposed compression framework compatible with existing image coding standards. Experimental results validate that the proposed compression framework greatly outperforms several compression frameworks that use existing image coding standards with the state-of-the-art deblocking or denoising post-processing methods. Feng Jiang 0001, Wen Tao, Shaohui Liu, Jie Ren 0016, Xun Guo 0002, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | An End-to-End Compression Framework Based on Convolutional Neural NetworksabstractSummary form only given. Traditional image coding standards (such as JPEG and JPEG2000) make the decoded image suffer from many blocking artifacts or noises since the use of big quantization steps. To overcome this problem, we proposed an end-to-end compression framework based on two CNNs, as shown in Figure 1, which produce a compact representation for encoding using a third party coding standard and reconstruct the decoded image, respectively. To make two CNNs effectively collaborate, we develop a unified end-to-end learning framework to simultaneously learn CrCNN and ReCNN such that the compact representation obtained by CrCNN preserves the structural information of the image, which facilitates to accurately reconstruct the decoded image using ReCNN and also makes the proposed compression framework compatible with existing image coding standards. Wen Tao, Feng Jiang 0001, Shengping Zhang, Jie Ren 0016, Wuzhen Shi, Wangmeng Zuo, Xun Guo 0002, Debin Zhao |
DCC | 7 |
| 2017 | Delay-Rate-Distortion Optimization for Cloud Gaming With Hybrid StreamingabstractCloud gaming as the emerging game service has attracted significant attention. However, traditional video streaming approach suffers from high bandwidth consumption, and traditional graphics streaming approach requires a long initial period to download game models. In this paper, we propose a novel hybrid streaming framework, jointly applying video streaming and graphics streaming to provide a high-quality gaming experience. In the proposed framework, cloud servers not only transmit the encoded video frames but also progressively transmit the graphics data, which are used to render a game frame to provide an additional reference to the video encoder. Based on the proposed framework, we investigate the delay-rate-distortion optimization problem, where the source rate between the video stream and the graphics stream is optimized to minimize the overall distortion under the bandwidth and response delay constraints. The experimental results demonstrate that the proposed hybrid streaming can achieve the lowest distortion under the constraints of bandwidth and response delay, compared with the traditional video streaming and graphics streaming. Xiaoming Nan, Xun Guo 0002, Yan Lu 0001, Ling Guan, Shipeng Li 0001, Baining Guo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | GPU-based optimization for sample adaptive offset in HEVCabstractThe latest high efficiency video coding (HEVC) standard achieves about 50% bit-rate reduction at equivalent visual quality compared to H.264/AVC. Sample adaptive offset (SAO) is one of the newly adopted tools right after deblocking filter, which can improve both coding efficiency and visual quality. However, for real-time encoding scenarios, the complexity of SAO is usually too high to be enabled. In this paper, a GPU-based optimization algorithm is proposed to reduce the complexity of SAO. Experiments are conducted based on the state-of-the-art open source HEVC encoder, i.e. X265. Results show that the proposed algorithm can reduce about 70% processing time of SAO on average without sacrifice of coding efficiency. Yang Wang 0048, Xun Guo 0002, Yan Lu 0001, Xiaopeng Fan 0001, Debin Zhao |
ICIP | 2 |
| 2014 | A novel cloud gaming framework using joint video and graphics streamingabstractAs the popularity of smart phones and tablets, users have an increasing desire to enjoy ubiquitous game playing. The emerging cloud gaming turns this desire into reality, enabling users to play games at anywhere on any devices. However, due to the huge amount of data transmission, it is challenging to provide a high quality game experience under the limited bandwidth capacity. In this paper, we propose a novel cloud gaming framework, in which we introduce two synchronized graphics buffers at both the server and the client sides. The server not only streams the compressed frames captured from game scenes, but also progressively transmits graphics data. The received graphics data is used to generate reference frames. When compressing the next frame, the cloud server will choose the reference frame with a lower residual error, from the previous frame and the current frame rendered from the graphics buffer. With the accumulation of graphics data, the frame rendered from the graphics buffer is close to the captured frame, which greatly reduces the transmission bit rates. Based on the proposed framework, we study the rate allocation problem, in which we optimize the allocated bit rates between the compressed frame and the graphics data to minimize the total distortion under the bandwidth constraint. Experimental results demonstrate that the proposed framework can optimally allocate bit rates to achieve a minimal distortion for cloud gaming compared to the traditional video streaming and graphics streaming approaches. Xiaoming Nan, Xun Guo 0002, Yan Lu 0001, Ling Guan, Shipeng Li 0001, Baining Guo |
ICME | 2 |
| 2014 | A low latency cloud gaming system using edge preserved image homographyabstractThe emerging cloud gaming technology has been growing fast, driving up huge mobile consumer demands. The video streaming based cloud gaming scenario renders the game scenes in the cloud servers, and streams the encoded sequences to the thin clints where the game scenes are decoded and displayed to the players. However, current existing clouding gaming services have some problems, such as the latency and bandwidth limitation. The size of the video stream is usually quite large which requires heavy transmission. Worse still, the frame data rate will burst when the game scenes contain fast translation or rotation, resulting in strong latency problem. In this paper, we propose a novel video streaming based cloud gaming algorithm which reduces the burst of the frame rate significantly. There are mainly two innovations in this paper. Firstly, based on the analysis of the motion estimation strategy in the video codec, we introduce image homography technique for better motion prediction. Meanwhile, according to the rasterization rules of the game engine, we present a special designed interpolation algorithm named Edge Preserved Interpolation (EPI), for more accurate edge interpolation and further reduce the residues in the edge regions. The proposed algorithm is implemented on the x264 platform. Experimental results show that our algorithm has 18.0% BD-rate reduction compared with x264. Lingfeng Xu, Xun Guo 0002, Yan Lu 0001, Shipeng Li 0001, Oscar C. Au, Lu Fang 0001 |
ICME | 2 |
| 2013 | Arbitrary-sized motion detection in screen video codingabstractIn real-time screen remoting system, frame rate is one of essential factors that affect user experience. Therefore, how to compress diversity of screen contents fast and efficiently is a key issue. Existing video codecs such as H.264 are always used in such a system for screen compression. However, arbitrary-sized regions with large motion always exist in typical screen content videos, which lead to a lower encoding speed and higher bit-rate, thereby decrease the frame rate. This paper proposes an efficient motion detection algorithm, which is fast and efficient for large motion regions. In specific, a region-based motion detection is used to find motion vectors instead of traditional block based motion estimation. The motion vectors are then utilized by H.264 encoder for normal motion compensated prediction. Experimental results show that the proposed algorithm can reduce both encoding time and bit-rate significantly. Tao Zhang 0013, Xun Guo 0002, Yan Lu 0001, Shipeng Li 0001, Siwei Ma 0001, Debin Zhao |
ICIP | 2 |
| 2012 | Simplified AMVP for High Efficiency Video CodingabstractIn High Efficiency Video Coding (HEVC), advanced motion vector prediction (AMVP) is adopted to predict current motion vector by utilizing a competition-based scheme from a given candidate set, which include both the spatial and temporal motion vectors. In order to enhance the practicability of the AMVP, a simplified AMVP is proposed. Firstly, by analyzing the importance of the spatial and temporal candidates, we reduce the number of the candidates involved in the competition set and simplify the redundancy checking process, which will decrease the complexity of the decoder as well as improve the robustness of the decoder. Secondly, we simplify the zero motion adding process which will occur only when the number of existing candidates is less than the predefined number. Experimental results show that the proposed scheme provides no loss in random access and low delay conditions. These two simplifications have been proposed and adopted into the HEVC standard. Liang Zhao 0007, Xun Guo 0002, Shawmin Lei, Siwei Ma 0001, Debin Zhao |
VCIP | 2 |
| 2012 | A Single-Pass-Based Localized Adaptive Interpolation Filter for Video CodingabstractRecently, an efficient coding tool named adaptive interpolation filtering (AIF) has been proposed to hybrid video coding scheme. By introducing Wiener filter into the fractional-pixel interpolation procedure, AIF can reduce the inter-prediction error and improve coding efficiency significantly. However, the training-based Wiener filter mechanism brings AIF an inherent multi-pass encoding structure, which imposes big burdens on the encoder in terms of huge computational complexity and memory access. In this paper, we propose a single-pass-based localized adaptive interpolation filtering (SPL-AIF) algorithm for video coding, which can reduce the complexity of AIF dramatically without sacrifice of its outstanding coding performance. The proposed SPL-AIF algorithm is based on the observation that there is a high correlation among optimal interpolation filters of consecutive frames, and different regions in a frame often possess different statistical characteristics. Accordingly, the proposed algorithm can be designed including two major parts. First, a competitive filter set which includes the optimal interpolation filters of several previous frames as well as the fixed H.264/AVC interpolation filters is built up for the coding of the current frame. Then a rate-distortion optimization criterion is used to select the best one at macroblock (MB) level. In order to reduce overhead, a predictive coding method is used to compress the filter signaling flag for each MB. Experimental results show that, by using the proposed algorithm, the encoding complexity can be reduced significantly while the average coding gain in Bjöntegaard distortion bit-rate reduction can be improved about 1% compared with the multi-pass AIF. The proposed method has been adopted into the Video Coding Expert Group Key Technology Area software. Kai Zhang 0007, Xun Guo 0002, Jicheng An, Yu-Wen Huang, Shawmin Lei, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | A single-pass based adaptive interpolation filtering algorithm for video codingabstractAn adaptive interpolation filtering (AIF) algorithm has been proposed to improve the conventional hybrid video coding scheme recently. Although such an algorithm does improve coding efficiency significantly, its encoding complexity, in terms of computational complexity and memory access, increases dramatically due to its inherent two-pass encoding structure. In this paper, we propose a novel single-pass based algorithm which can reduce the complexity of AIF significantly. In our method, a competitive filter set which includes optimal filters trained from several previous frames and the fixed H.264 filters is considered for the current coding frame. A rate-distortion optimization (RDO) criterion is then used to select the best one at macro-block (MB) level. In order to reduce overhead, a predictive coding method is used to compress the filter type for each MB. Experimental results show that, by using the proposed algorithm, the encoding complexity can be significantly reduced without sacrifice of coding gain. Kai Zhang 0007, Xun Guo 0002, Yu-Wen Huang, Shawmin Lei, Wen Gao 0001 |
ICIP | 2 |
| 2010 | Localized multiple adaptive interpolation filters with single-pass encodingabstractAdaptive interpolation filtering (AIF) algorithms have been proposed to enhance the hybrid video coding scheme recently. Although these algorithms can improve coding efficiency significantly, their encoders suffer huge increase in complexity in terms of latency and memory access due to its inherent multi-pass encoding procedure. In this paper, we present a novel single-pass solution for these algorithms, which allows optimal selection among different interpolation filters. In this solution, time-delayed interpolation filters are used to achieve single-pass encoding, and localized ratedistortion (RD) selection is used to compensate the possible coding loss from time-delayed filters. Experimental results show that the proposed method is efficient for All AIF techniques in current ITU-T/SG16 reference software. By using the proposed method, single-pass encoding with multiple AIF filters can be achieved while maintaining similar coding efficiency as multi-pass AIF. Xun Guo 0002, Kai Zhang 0007, Yu-Wen Huang, Jicheng An, Chih-Ming Fu, Shawmin Lei |
VCIP | 1 |
| 2008 | Wyner-Ziv-Based Multiview Video CodingabstractUtilizing video correlations among views would definitely improve multiview video compression in terms of coding efficiency, which usually requests an expensive system to collect videos from different cameras and jointly compress them. Thanks to recent developments on distributed video coding, this paper proposes a new multiview video coding scheme based on Wyner-Ziv (WZ) coding technique, in which the complicated temporal and interview correlation exploration process is shifted from the encoder side to the decoder side so that broadband raw data traffic and high intensive computation for jointly encoding can be avoided. At the encoder side, a wavelet-based WZ scheme is proposed to compress video of every camera. Furthermore, in order to better utilize correlation in wavelet domain, all coefficients are organized as that done in SPIHT bit plane by bit plane. At the decoder side, a more flexible prediction technique that can jointly utilize temporal and view correlations is proposed to generate side information. Finally, experimental results show the proposed scheme significantly outperforms the conventional intra-frame coding for better random access and is even close to the inter-frame coding for better efficiency. Furthermore, compressed data is much robust when it is transmitted over an error-prone channel. Xun Guo 0002, Yan Lu 0001, Feng Wu 0001, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | Wyner-Ziv Video Coding Based on Set Partitioning in Hierarchical TreeabstractIn this paper, we propose a Wyner-Ziv video coding scheme based on set-partitioning in hierarchical trees (SPIHT) which can utilizing not only the spatial and temporal correlations but also the high-order statistical correlations. Wyner-Ziv theory on source coding with side information is employed as the basic coding principle, which makes the independent encoding and joint decoding become possible. In the proposed scheme, wavelet transform is first used to de-correlate the spatial dependency of a Wyner-Ziv frame. Then the quantized transform coefficients are organized by using magnitude with a set partitioning sorting algorithm. The ordered bit planes are coded using the Wyner-Ziv coding based on turbo codes. At the decoder, side information generated by motion compensated interpolation is used to conditionally decode the Wyner-Ziv frame. Since the high order statistical correlation is used, the proposed algorithm owns advantages over the traditional pixel-domain and transform-domain Wyner-ziv video coding schemes. Xun Guo 0002, Yan Lu 0001, Feng Wu 0001, Wen Gao 0001, Shipeng Li 0001 |
ICIP | 1 |
| 2006 | An Optimal Non-Uniform Scalar Quantizer for Distributed Video CodingabstractIn this paper, we propose a novel algorithm to design an optimal non-uniform scalar quantizer for distributed video coding, which aims at achieving a coding rate close to joint conditional entropy of the quantized video frames given the side information. Wyner-Ziv theory on source coding is employed as the basic coding principle and the asymmetric scenario is considered. In this algorithm, a probability distribution model, which considers the influence of the joint distribution of input source and side information to the coding performance, is established and used as the optimality condition firstly. Then, a modified Lloyd Max algorithm is used to design the scalar quantizer to give an optimal quantization for input source before coding. Experimental results show that compared to uniform scalar quantization, proposed algorithm can improve coding performance largely, especially at low bit rate Bo Wu 0016, Xun Guo 0002, Debin Zhao, Wen Gao 0001, Feng Wu 0001 |
ICME | 2 |
| 2006 | Distributed video coding using waveletabstractThis paper proposes a distributed video coding scheme based on the zero tree entropy (ZTE) coding. Wyner-Ziv theory on source coding with side information is taken as the basic coding principle, which makes independent encoding and joint decoding possible. In this scheme, wavelet transform is used to exploit the spatial correlation of a Wyner-Ziv frame. The quantized wavelet coefficients are reorganized in terms of the zero tree structure so as to identify the significant and insignificant coefficients. The significance map is intra-codec and transmitted. In particular, the significant coefficients are independently encoded with turbo coder, and only the parity bits are transmitted. At the decoder, a predictive frame generated through motion-compensated prediction is used as the side information, with which the Wyner-Ziv frame can be conditionally decoded. Experimental results show that, compared to the traditional intra-frame coding and pixel-domain Wnyer-Ziv video coding, the proposed scheme can achieve a better coding performance, especially at low bit rates. Xun Guo 0002, Yan Lu 0001, Feng Wu 0001, Wen Gao 0001 |
ISCAS | 1 |
| 2006 | Inter-View Direct Mode for Multiview Video CodingabstractGlobal disparity between views is usually caused by the displacement between cameras, which can be accurately represented by a global geometric transformation. In this paper, we first propose an inter-view motion model in terms of the global geometric transformation to represent the motion correlation between two adjacent views. Specifically, the motion vector of a pixel in one view may be directly derived from that in another view according to the inter-view motion model. Further, we propose an inter-view direct mode to signal the decoder that the motion of a macroblock (MB) can be achieved from the coded view without any coding bits. The proposed inter-view direct mode is further incorporated in the existing multiview video coding (MVC) schemes (i.e., AVC-based MVC and 4-D wavelet-based MVC), working together with the other classical coding modes. The mode selection at each MB is accomplished with the rate-distortion optimization technique. The proposed inter-view direct mode can significantly reduce bits to code motion vectors especially at low bit rates, thus improving the coding efficiency Xun Guo 0002, Yan Lu 0001, Feng Wu 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | Motion vector prediction in multiview video codingabstractIn video coding, motion vectors always account for a large number of bits and affect coding efficiency largely. In this paper, we propose an efficient motion vector prediction algorithm for multiview video coding (MVC), which can predict motion vectors from adjacent views and achieve good prediction accuracy. We first investigate the correlations among different views and describe the disparity between adjacent views as global motion. The affine model is used to compute the global parameters between frames of adjacent views. At least one view is coded independently without interview prediction. After that, motion vectors of the frame to be coded can be derived from the motion vectors of the co-located coded frame in adjacent view using the global motion information. A rate-distortion optimization scheme is used to choose between the proposed method and traditional motion compensated prediction method. Experimental results show that, compared to simulcast coding, the proposed algorithm can achieve good performance and improve the coding efficiency up to 0.8 dB in PSNR. Xun Guo 0002, Wen Gao 0001, Debin Zhao |
ICIP (2) | 1 |