Kan Chang

dblp:57/7435 · DBLP profile ↗
← Back
45ranked-venue papers
14as first author
25since 2021 · last 2026
0000-0002-6587-0360ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 12 first-author · 13 since 2021Artificial intelligence and machine learning · 15 · 15 since 2021Security and privacy · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Domain adaptive object detection via CLIP-space guidance and LoRA fine-tuning
Enze Qi, Kan Chang, Qingzhi Zhang, Xueyu Zhang, Yehua Ling, Yujian Yuan, Zan Gao 0001
Expert Syst. Appl.2
2026 UECNet: A unified framework for exposure correction utilizing region-level prompts
Shucheng Xia, Kan Chang, Xuxin Tai, Yehua Ling, Yujian Yuan, Zan Gao 0001
Knowl. Based Syst.2
2025 Prompt-Based Two-Stage Enhancement for Low-Light Object Detection
abstract
Detecting objects in low-light conditions is challenging since the primary features of objects may be buried in dark regions. To tackle this issue, we propose a prompt-based two-stage enhancement (PTE) method for low-light object detection. Our approach utilizes a prompt generator to produce prompts from a support set established based on the input low-light image. These prompts encode the specific degradation information of the input, and thus they are further utilized to guide both image-level and feature-level enhancements for object detection. This allows our PTE to adapt to various lighting conditions. With these prompts, the image-level enhancement adjusts the low-light image using an estimated reciprocal curve, while the feature-level enhancement adaptively modulates the multi-scale features extracted by the detection backbone. Experimental results demonstrate that our method outperforms existing methods in terms of detection accuracy, while maintaining a low number of parameters and high running speed. Our source code and pre-trained model will be available at https://github.com/xbh301/PTE-YOLO.
Bohan Xiong, Kan Chang, Shilin Huang, Shucheng Xia, Yujian Yuan
ICME2
2025 Multi-scale wavelet feature fusion network for low-light image enhancement
Xinjie Wei, Shucheng Xia, Kan Chang, Jingxiang Nong
Comput. Graph.4
2025 CSFIN: A lightweight network for camouflaged object detection via cross-stage feature interaction
Minghong Li, Yuqian Zhao 0001, Fan Zhang 0106, Gui Gui, Biao Luo 0001, Chunhua Yang 0001, Weihua Gui 0001, Kan Chang
Expert Syst. Appl.8
2025 Conditional Laplacian pyramid networks for exposure correction
Mengyuan Huang, Kan Chang, Qingpao Qin, Yahui Tang, Guiqing Li
Signal Process. Image Commun.2
2024 Auxiliary Domain-Guided Adaptive Object Detection in Adverse Weather Conditions
Zhuobin Fu, Kan Chang, Qingzhi Zhang, Enze Qi
ACCV (1)2
2024 GPNF:A Point Cloud Registration Framework Using Sharp Global Linear Attention Prior and Neighborhood Filtering Strategy
Congyang Zhu, Mengxiao Yin, Zhijie Liang, Kan Chang
ACCV (10)5
2024 Joint Image and Feature Enhancement for Object Detection under Adverse Weather Conditions
abstract
Object detection under adverse weather conditions remains a challenging problem to date. To address this problem, a joint image and feature enhancement method called JE-YOLO is proposed. Firstly, a lightweight image enhancement network is used to enhance the low-quality image captured under adverse weather conditions. Secondly, to provide rich information for detection, two detection backbones are applied in parallel to extract features from both the low-quality image and its enhanced result. Afterwards, the extracted features are further enhanced by a foreground-guided feature refinement module (FFRM), which introduces a task-driven attention mechanism and explores inter-layer correlation. Finally, the enhanced features from different branches are fused by the adaptive multi-branch weighting (AMW) strategy, and then fed to the neck and head of detector. Experiments are carried out on both the low-light and foggy conditions, and the results demonstrate that compared with state-of-the-art (SOTA) methods, the proposed JE-YOLO is able to achieve the highest accuracy of detection in all cases. Code will be available at https://github.com/Murray-Yin/JE-YOLO.
Mengyu Yin, Kan Chang, Zijian Yuan, Qingpao Qin, Boning Chen
IJCNN3
2024 A Two-Stage Enhancement Method for Object Detection on Low-Resolution Images
abstract
Detecting objects on low-resolution (LR) images is a challenging task, as small structures and details contained in LR images are insufficient. To improve the detection accuracy on LR images, a two-stage enhancement framework is proposed in this paper. In the first stage, a lightweight image super-resolution (SR) sub-network is built by incorporating the external guidance obtained by the detection backbone, so that the potential discriminant features can be well preserved. Moreover, to alleviate the hardship of directly learning a nonlinear mapping from the LR space to the high-resolution (HR) space, an auxiliary module which maps the SR results back to the degraded LR images provides additional constraint. In the second stage, the multi-scale features produced by the detection heads are further enhanced by the lightweight feature adaptation modules, which aim to further bridge the gap between the SR result and the corresponding HR image. During training, the detector is kept frozen in our framework, leading to a flexible plug-and-play mechanism. Experiments demonstrate that our approach achieves better performance on LR images than many state-of-the-art methods. To facilitate further study, our source code will be available at: https://github.com/Zijian-Yuan/TSE-YOLO.
Zijian Yuan, Kan Chang, Mengyu Yin, Minghong Li, Boning Chen
IJCNN3
2024 Object detection on low-resolution images with two-stage enhancement
Minghong Li, Yuqian Zhao 0001, Gui Gui, Fan Zhang 0106, Biao Luo 0001, Chunhua Yang 0001, Weihua Gui 0001, Kan Chang, Hui Wang 0069
Knowl. Based Syst.8
2024 Illumination-aware divide-and-conquer network for improperly-exposed image enhancement
Fenggang Han, Kan Chang, Guiqing Li, Mengyuan Huang
Neural Networks2
2024 Multi-scale feature selection network for lightweight image super-resolution
Minghong Li, Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Chunhua Yang 0001, Weihua Gui 0001, Kan Chang
Neural Networks7
2024 Object and spatial discrimination makes weakly supervised local feature better
Mengxiao Yin, Yunhui Xiong, Pengfei Lai, Kan Chang, Feng Yang 0014
Neural Networks5
2024 FD-Net: Feature Distillation Network for Oral Squamous Cell Carcinoma Lymph Node Segmentation in Hyperspectral Imagery
abstract
Oral squamous cell carcinoma (OSCC) has the characteristics of early regional lymph node metastasis. OSCC patients often have poor prognoses and low survival rates due to cervical lymph metastases. Therefore, it is necessary to rely on a reasonable screening method to quickly judge the cervical lymph metastastic condition of OSCC patients and develop appropriate treatment plans. In this study, the widely used pathological sections with hematoxylin-eosin (H&E) staining are taken as the target, and combined with the advantages of hyperspectral imaging technology, a novel diagnostic method for identifying OSCC lymph node metastases is proposed. The method consists of a learning stage and a decision-making stage, focusing on cancer and non-cancer nuclei, gradually completing the lesions' segmentation from coarse to fine, and achieving high accuracy. In the learning stage, the proposed feature distillation-Net (FD-Net) network is developed to segment the cancerous and non-cancerous nuclei. In the decision-making stage, the segmentation results are post-processed, and the lesions are effectively distinguished based on the prior. Experimental results demonstrate that the proposed FD-Net is very competitive in the OSCC hyperspectral medical image segmentation task. The proposed FD-Net method performs best on the seven segmentation evaluation indicators: MIoU, OA, AA, SE, CSI, GDR, and DICE. Among these seven evaluation indicators, the proposed FD-Net method is 1.75%, 1.27%, 0.35%, 1.9%, 0.88%, 4.45%, and 1.98% higher than the DeepLab V3 method, which ranks second in performance, respectively. In addition, the proposed diagnosis method of OSCC lymph node metastasis can effectively assist pathologists in disease screening and reduce the workload of pathologists.
Xueyu Zhang, Qingxiang Li, Wei Li 0032, Yuxing Guo, Jianyun Zhang, Chuanbin Guo, Kan Chang, Nigel H. Lovell
IEEE J. Biomed. Health Informatics7
2024 PCMG:3D point cloud human motion generation based on self-attention and transformer
Weizhao Ma, Mengxiao Yin, Guiqing Li, Feng Yang 0014, Kan Chang
Vis. Comput.5
2023 DLEN: Deep Laplacian Enhancement Networks for Low-Light Images
abstract
Enhancing low-light images is challenging as it requires simultaneously handling global and local contents. This paper presents a new solution which incorporates the vision transformer (ViT) into Laplacian pyramid and explores cross-layer dependence within the pyramid. It first applies Laplacian pyramid to decompose the low-light image into a low-frequency (LF) component and several high-frequency (HF) components. As the LF component has a low resolution and mainly includes global attributes, ViT is applied on it to explore the interdependence among global contents. Since there exists strong spatial correlation among different frequency components, the refined features from a lower pyramid layer are used to assist the refinement of upper-layer features. Experiments demonstrate that our approach achieves better performance than state-of-the-art methods, while maintaining a relative small model size and low computational complexity. Our source code and trained model will be released at https://github.com/Xinjie-Wei/DLEN.
Xinjie Wei, Kan Chang, Guiqing Li, Mengyuan Huang, Qingpao Qin
ICIP2
2023 Joint Super-Resolution and Classification Based on Bidirectional Mapping and Multiple Constraints
abstract
Since discriminant features are insufficient in the low-resolution (LR) images, it is challenging to accurately classify them. To address this issue, this paper proposes a joint super-resolution (SR) and classification network (JSRCN) based on bidirectional mapping and multiple constraints. In JSRCN, there is a SR sub-network containing a forward mapping path and several backward mapping paths. In the forward mapping path, high resolution (HR) features are progressively reconstructed. On the other hand, the backward mapping paths are used to alleviate the hardship of directly learning a nonlinear mapping from the LR space to the HR space. To effectively restore discriminant features for the image classification sub-network, multiple perceptual loss and multi-scale feature loss are presented to enhance the representation ability for the multi-scale features in images with different resolutions. Experiments show that compared with other competing methods, the proposed method achieves the highest accuracy.
Zijian Yuan, Kan Chang, Zhiquan Liu 0005, Xinjie Wei, Boning Chen
ICME2
2023 CInvISP: Conditional Invertible Image Signal Processing Pipeline
Duanling Guo, Kan Chang, Yahui Tang, Minghong Li
ICONIP (5)2
2023 Hyperspectral pathology image classification using dimension-driven multi-path attention residual network
Xueyu Zhang, Wei Li 0032, Chenzhong Gao, Yue Yang 0041, Kan Chang
Expert Syst. Appl.5
2023 3D mesh pose transfer based on skeletal deformation
abstract
Abstract For 3D mesh pose transfer, the target model is obtained by transferring the pose of the reference mesh to the source mesh, where the shape and pose of the source are usually different from that of the reference. In this paper, pose transfer is considered as a deformation process of the source mesh, and we propose a 3D mesh pose transfer method based on skeletal deformation. First, we design a neural network based on the edge convolution operator to extract the skeleton of the 3D mesh and bind the rigid weights; then, we calculate the bone transformations between the two skeletons with different poses and use the diffusion equation to smooth the rigid weights; finally, the source mesh is deformed according to the bone transformations and the smooth weights to get the target mesh. Experiment results on different datasets show that the pose of the reference mesh can be effectively transferred to the source one while maintaining the shape and high‐quality geometric details of the source mesh by using our method.
Shigeng Yang, Mengxiao Yin, Guiqing Li, Kan Chang, Feng Yang 0014
Comput. Animat. Virtual Worlds5
2023 BMISP: Bidirectional mapping of image signal processing pipeline
Yahui Tang, Kan Chang, Mengyuan Huang, Baoxin Li
Signal Process.2
2022 DENet: Detection-driven Enhancement Network for Object Detection Under Adverse Weather Conditions
Qingpao Qin, Kan Chang, Mengyuan Huang, Guiqing Li
ACCV (3)2
2022 Joint Image Super-Resolution and Low-Light Enhancement in the Dark
Feihu Zhou, Kan Chang, Hengxin Li, Shucheng Xia
ACCV (4)2
2022 A Two-Stage Convolutional Neural Network for Joint Demosaicking and Super-Resolution
abstract
As two practical and important image processing tasks, color demosaicking (CDM) and super-resolution (SR) have been studied for decades. However, most literature studies these two tasks independently, ignoring the potential benefits of a joint solution. In this paper, aiming at efficient and effective joint demosaicking and super-resolution (JDSR), a well-designed two-stage convolutional neural network (CNN) architecture is proposed. For the first stage, by making use of the sampling-pattern information, a pattern-aware feature extraction (PFE) module extracts features directly from the Bayer-sampled low-resolution (LR) image, while keeping the resolution of the extracted features the same as the input. For the second stage, a dual-branch feature refinement (DFR) module effectively decomposes the features into two components with different spatial frequencies, on which different learning strategies are applied. On each branch of the DFR module, the feature refinement unit, namely, densely-connected dual-path enhancement blocks (DDEB), establishes a sophisticated nonlinear mapping from the LR space to the high-resolution (HR) space. To achieve strong representational power, two paths of transformations and the channel attention mechanism are adopted in DDEB. Extensive experiments demonstrate that the proposed method is superior to the sequential combination of state-of-the-art (SOTA) CDM and SR methods. Moreover, with much smaller model size, our approach also surpasses other SOTA JDSR methods.
Kan Chang, Hengxin Li, Yufei Tan, Pak Lun Kevin Ding, Baoxin Li
IEEE Trans. Circuits Syst. Video Technol.1
2020 Lightweight Color Image Demosaicking with Multi-Core Feature Extraction
abstract
Convolutional neural network (CNN)-based color image demosaicking methods have achieved great success recently. However, in many applications where the computation resource is highly limited, it is not practical to deploy large-scale networks. This paper proposes a lightweight CNN for color image demosaicking. Firstly, to effectively extract shallow features, a multi-core feature extraction module, which takes the Bayer sampling positions into consideration, is proposed. Secondly, by taking advantage of inter-channel correlation, an attention-aware fusion module is presented to efficiently reconstruct the full color image. Moreover, a feature enhancement module, which contains several cascading attention-aware enhancement blocks, is designed to further refine the initial reconstructed image. To demonstrate the effectiveness of the proposed network, several state-of-the-art demosaicking methods are compared. Experimental results show that with the smallest number of parameters, the proposed network outperforms the other compared methods in terms of both objective and subjective qualities.
Yufei Tan, Kan Chang, Hengxin Li, Tuanfa Qin
VCIP2
2020 Accurate single image super-resolution using multi-path wide-activated residual network
Kan Chang, Minghong Li, Pak Lun Kevin Ding, Baoxin Li
Signal Process.1
2019 Data-adaptive low-rank modeling and external gradient prior for single image super-resolution
Kan Chang, Xueyu Zhang, Pak Lun Kevin Ding, Baoxin Li
Signal Process.1
2018 Single image super-resolution using collaborative representation and non-local self-similarity
Kan Chang, Pak Lun Kevin Ding, Baoxin Li
Signal Process.1
2018 Single Image Super Resolution Using Joint Regularization
abstract
This letter proposes a reconstruction-based single image super resolution method by using joint regularization, where a group-residual-based regularization (GRR) and a ridge-regression-based regularization (3R) are combined. In GRR, nonlocal similar patches are grouped together, and the group weights are calculated so as to adaptively constrain the residual values in the gradient domain. In 3R, we adopt the ridge-regression-based method to establish the projection matrices from an external high-resolution (HR) training set, so that the external HR information can be utilized. To obtain an estimation of the targeted HR image, an efficient algorithm is designed for solving the joint formulation. Experimental results on different image datasets indicate that the proposed method is able to achieve the state-of-the-art performance.
Kan Chang, Pak Lun Kevin Ding, Baoxin Li
IEEE Signal Process. Lett.1
2017 Convex dictionary learning for single image super-resolution
abstract
In recent years, dictionary learning approaches have been used in image super-resolution, achieving promising results. Such approaches train a dictionary from image patches and reconstruct a new patch by sparse combination of the atoms of the dictionary. Typical training methods do not constrain the dictionary atoms. In this paper, we propose a convex dictionary learning (CDL) algorithm by constraining the dictionary atoms to be formed by non-negative linear combination of the training data, which is a natural, desired property. We evaluate our approach by demonstrating its performance gain over typical approaches.
Pak Lun Kevin Ding, Baoxin Li, Kan Chang
ICIP3
2016 Color image compressive sensing reconstruction by using inter-channel correlation
abstract
This paper proposes a novel algorithm for compressive sensing (CS) reconstruction of color images. First of all, to better describe color image characteristics, we take inter-channel correlation into consideration and present two types of regularization, including inter-channel correlation-based nonlocal low-rank (ICNL) regularization and inter-channel correlation-based total variation (ICTV) regularization. Afterwards, both regularization terms are incorporated into the minimization problem, and an efficient algorithm is proposed to solve the joint formulation, by using a split-Bregman-based technique. To demonstrate the effectiveness of the proposed approach, four benchmark methods are compared, and the experiments are carried out on several color images with different subrates.
Kan Chang, Yun Liang 0005, Tuanfa Qin
VCIP1
2016 Compressive Sensing Reconstruction of Correlated Images Using Joint Regularization
abstract
This letter proposes a novel compressive sensing reconstruction method for correlated images by using joint regularization, where a compensation-based adaptive total variation (CATV) regularization and a multi-image nonlocal low-rank (MNLR) regularization are included. In CATV, local weights are assigned to the residual values in the gradient domain so as to constrain the regularization strength at each pixel. In MNLR, the search of similar patches goes across different images so that both self-similarity and inter-image similarity are explored. Afterward, an efficient algorithm is proposed to solve the joint formulation, using a Split-Bregman-based technique. The effectiveness of the proposed approach is demonstrated with experiments on both multiview images and video sequences.
Kan Chang, Pak Lun Kevin Ding, Baoxin Li
IEEE Signal Process. Lett.1
2015 Joint modeling and reconstruction of a compressively-sensed set of correlated images
Kan Chang, Baoxin Li
J. Vis. Commun. Image Represent.1
2015 Color image demosaicking using inter-channel correlation and nonlocal self-similarity
Kan Chang, Pak Lun Kevin Ding, Baoxin Li
Signal Process. Image Commun.1
2014 Reconstruction of compressed-sensed video using compound regularization
abstract
This paper introduces a novel reconstruction model with compound regularization to recover compressed-sensed video sequences. For a target frame, the compound regularization consists of total variation (TV) norm of the frame, l1norm of the frame in a certain transform domain, and TV norm of the residual between the frame and its prediction. The first two terms in the compound regularization are used to describe image characteristics, while the third term exploits inter-frame correlation within video sequences. To solve the minimization problem, a new splitting objective function is considered, and it is divided into sub-problems that are easy to solve. In addition, bivariate shrinkage method is integrated into the proposed algorithm so that high quality of reconstruction results can be guaranteed. Experimental results show that the proposed algorithms are substantially superior to state-of-the-art reconstruction methods.
Kan Chang, Tuanfa Qin, Miwen Zuo, Jinglan Shi
ICME1
2014 Reconstruction of multi-view compressed imaging using weighted total variation
Kan Chang, Tuanfa Qin
Multim. Syst.1
2013 A joint reconstruction algorithm for multi-view compressed imaging
abstract
As compressed sensing can capture signal at sub-Nyquist rate, it is suitable to apply multi-view compressed imaging framework in vision sensor networks. The image views in such networks are correlated with each other, and therefore the performance of independent view reconstruction can be further improved by joint reconstruction. In this paper, we propose a joint reconstruction algorithm, where disparity estimation and disparity compensation are used to exploit the correlation between views. The target optimization problem is divided into two sub-problems and they are solved alternately by proximal-gradient method. We show by experiments that, for a given sub-rate, the proposed joint reconstruction scheme outperforms the independent reconstruction in terms of image quality.
Kan Chang, Tuanfa Qin, Aidong Men
ISCAS1
2011 Block-level adaptive optimization for inter-layer texture up-sampling in H.264/SVC
abstract
H.264 Scalable Video Coding (SVC) extension has spatial scalability which is able to provide various resolution sequences for a single encoded bit-stream. In order to reduce redundancies between different layers, for spatial scalable intra-coded frames, co-located reconstructed 8×8 sub-macroblock in base layer (BL) is up-sampled to predict the marcoblock (MB) in enhancement layer (EL). Unfortunately, simple 1-D poly-phase up-sampling filter used in current SVC isn't cable of achieving ideal result, which limits the performance of inter-layer intra prediction (ILIP). This paper proposes an adaptive optimization method for inter-layer texture up-sampling by applying wiener filter and controlling it at block level. Working as an additional part of ILIP, the proposed method can greatly reduce the prediction error between the original EL signals and the up-sampled BL signals. Experimental results show that, the proposed method achieves bit rate reduction up to 14.25% and PSNR increment up to 0.97 dB when compared with the traditional method in current SVC.
Kan Chang, Tuanfa Qin, Wenhao Zhang 0001, Aidong Men
MMSP1
2011 An Improved Wyner-Ziv Video Coding for Sensor Network
abstract
Wyner-Ziv video coding is a new compression paradigm based on two key Information Theory results: the Slepian-Wolf and Wyner-Ziv theorems. It shifts the complexity to the decoder, resulting in a low- complexity encoder suitable for mobile video communications and visual sensor networks. This paper presents an improved Wyner-Ziv video coding scheme for sensor network. An improved key frame encoding method based on the correlation noise model (CNM) is proposed, and then a 3DRS-assisted motion estimation algorithm and AOBMC technique are used to improve the rate-distortion performance of the codec. The results show that our coding scheme can achieve 2-4 dB gain compared to state-of-the-art TDWZ codec and be deployed over a real visual sensor platform.
Aidong Men, Kan Chang, Jinhong Di
VTC Spring4
2010 An improved Wyner-Ziv video coding with feedback channel
abstract
This paper presents an improved feedback-assisted low complexity WZVC scheme. The performance of this scheme is improved by two enhancements: an improved mode-based key frame encoding and a 3DRS-assisted (three-dimensional recursive search assisted) motion estimation algorithm for WZ encoding. Experimental results show that our coding scheme can achieve significant gain compared to state-of-the-art TDWZ codec while still low encoding complexity.
Aidong Men, Bo Yang 0007, Manman Fan, Kan Chang
PCS5
2009 Novel Fast Mode Decision Algorithm for P-Slices in H.264/AVC
abstract
H.264/AVC encoder complexity is remarkable mainly due to variable block size motion estimation (ME) and exhaustive rate distortion optimization (RDO). This makes real-time video coding very difficult. In this paper, a novel fast mode decision algorithm for P-slices is presented to reduce the computational load. Based on the temporal and spatial correlation between macroblocks, a four paths prediction structure is proposed with flexible early termination strategy. Experiment results show that with this algorithm, 42%-65% encoding time can be saved with a negligible loss in peak signal-to-noise ratio (PSNR) and very little increment in bit rate. Combined with Choi's method, our algorithm can achieve a time reduction up to 74.31% with little loss in coding efficiency.
Kan Chang, Bo Yang 0007, Wenhao Zhang 0001
IAS1
2009 GOP-Level Transmission Distortion Modeling for Video Streaming over Mobile Networks
abstract
A major challenge in video coding and transmission over mobile networks is that the wireless channel is error-prone and the channel resources are limited. In this work, we analyze the picture distortion caused by channel errors and the distortion propagation behavior in its subsequent frames along the motion prediction path. We propose a linear fitting approach algorithm to achieve a low complexity GOP-level transmission distortion modeling. It is a predictive modeling which allows the encoder to predict the transmission distortion before the whole GOP is compressed and transmitted. The simulation results demonstrate that the proposed modeling has low computational complexity and high accuracy. It can be used in allocating the limited channel resources optimally for mobile video applications.
Aidong Men, Kan Chang, Ziyi Quan
IAS3
2009 Optimal Combination of H.264/AVC Coding Tools for Mobile Multimedia Broadcasting Application
abstract
Through experiments on combinations of coding tools included in H.264/AVC, we obtain an optimal combination of coding tools for the mobile multimedia broadcasting application. Although H.264/AVC has defined several profiles and levels for different applications, the partitions are still imprecise. The mobile communication network is usually with limited bandwidth and different kinds of fading and multipath interference. The mobile terminals with downlink video service in the broadcasting network are battery restricted and low processing power. Considering of these features, the proposal has made an optimal selection of the coding tools in H.264/AVC for the mobile multimedia broadcasting application. This set of tools makes a good balance among compression efficiency, transmission robustness and decoding complexity, therefore provides reliable quality for mobile multimedia broadcasting service.
Bo Yang 0007, Kan Chang, Ziyi Quan
IAS3
2009 Adaptive optimizing filter for inter-layer intra prediction in SVC
abstract
Several inter-layer prediction techniques are adopted in the H.264 Scalable Video Coding Extension to increase the compression efficiency. For spatial scalable intra-coded frames, the prediction of macroblock in the enhancement layer is obtained by upsampling the co-located reconstructed block in the base layer. This inter-layer intra prediction can remove the redundancies between the different layers; however, its performance is limited by the non-ideal upsampling filter and the coding losses of base layer. In order to improve the performance of the inter-layer intra prediction, we propose an optimizing filtering method to enhance the prediction signal in this paper. Adaptive Wiener filters, which are calculated for each slice independently, are used to generate a prediction signal with minimum error energy in a statisitcal way. After this optimizing, the coding efficiency can be increased progressively. Experiments show that up to 6.45% and 15.45% bit rate reduction is achieved for QCIF-CIF scenario and CIF-4CIF scenario, respectively.
Wenhao Zhang 0001, Aidong Men, Kan Chang
ACM Multimedia3