EDBT 2026 Demo / reviewers in the wild / expert
Nan Zhang 0015
dblp:28/6297-15
· DBLP profile ↗
25ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0001-7444-7508ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Point Cloud Semantic Segmentation with Sparse and Inhomogeneous AnnotationsabstractUtilizing uniformly distributed sparse annotations, weakly supervised learning alleviates the heavy reliance on fine-grained annotations in point cloud semantic segmentation tasks. However, few works discuss the inhomogeneity of sparse annotations, albeit it is common in real-world scenarios. Therefore, this work introduces the probability density function into the gradient sampling approximation method to qualitatively analyze the impact of annotation sparsity and inhomogeneity under weakly supervised learning. Based on our analysis, we propose an Adaptive Annotation Distribution Network (AADNet) capable of robust learning on arbitrarily distributed sparse annotations. Specifically, we propose a label-aware point cloud downsampling strategy to increase the proportion of annotations involved in the training stage. Furthermore, we design the multiplicative dynamic entropy as the gradient calibration function to mitigate the gradient bias caused by non-uniformly distributed sparse annotations and explicitly reduce the epistemic uncertainty. Without any prior restrictions and additional information, our proposed method achieves comprehensive performance improvements at multiple label rates and different annotation distributions. Zhiyi Pan 0001, Nan Zhang 0015, Wei Gao 0003, Shan Liu 0001, Ge Li 0002 |
AAAI | 2 |
| 2025 | Performing task automation for surgical robot: A spatial-temporal varying primal-dual neural network with guided obstacle avoidance and null space optimizationabstractPerforming surgical tasks safely and reliably presents significant challenges, including obstacle avoidance, joint limit constraints, and motion smoothness during the tool-target alignment (T-TA) stage, as well as precise tracking of preoperative plans during the execution of the preoperative planning surgery path (EPSP). The traditional inverse kinematics methods fall short in addressing these complex motion planning and control issues within the unstructured and time-varying surgical environment. Therefore, a novel spatial–temporal varying primal–dual neural network (STV-PDNN) that incorporates guided obstacle avoidance and null space optimization to address spatial–temporal constraints during surgery is proposed. Firstly, a velocity control quadratic programming (QP) framework based on target distance and orientation metrics is constructed by considering the relationships among the surgical robot, the environment, and the surgical target. Then, the STV-PDNN enables real-time problem-solving across two specific stages, employing velocity vector projection for obstacle avoidance and joint space obstacle avoidance velocity superposition to enhance the obstacle avoidance guidance. Furthermore, the joint null space optimization and maximum manipulability, along with a preoperative planning path velocity feed-forward and feedback velocity control mechanism, are integrated into the STV-PDNN structure. The improvement facilitates smoother, lower-energy joint movements and effective motion singularity avoidance during the T-TA stage, as well as precise motion control in the EPSP stage. The experiments conducted on the redundant robot Diana7 Med validate the effectiveness of the proposed method in autonomously executing T-TA and EPSP for pedicle screw implantation, offering a promising solution for the task autonomy of surgical robot. Xingqiang Jian, Bo Wu 0016, Yibin Song, Yu Wang 0244, Da He, Nan Zhang 0015 |
Expert Syst. Appl. | 10 |
| 2025 | Motion Planning and Control of Active Robot in Orthopedic Surgery by CDMP-Based Imitation Learning and Constrained OptimizationabstractCurrent orthopedic surgical robots are widely used in pedicle screw implantation tasks due to their precise positioning capabilities. However, the surgical operation processes, including surgical pose alignment and the drilling of the pedicle screw placement path, remain heavily dependent on the surgeon, indicating that the level of automation still needs improvement. This paper aims to enhance automation in pedicle screw implantation tasks through a combination of imitation learning and constraint optimization, thereby ensuring both reliability and safety. Firstly, a high-level motion planning method, leveraging Cartesian space dynamic movement primitives (CDMP) based imitation learning and an image-guided optical navigation system (I-GONS), is proposed to generate the task space path of the rough surgical pose alignment and fine surgical pose alignment, as well as for the drilling of pedicle screw placement path. Secondly, end-effector velocity control based on position and orientation errors (POE-EVC) is employed to follow the high-level planned path. This is achieved by constructing a quadratic programming (QP) problem with the robot kinematic constraints and manipulability optimization. Concurrently, the low-level motion control is addressed online using a modified linear variational inequality based primal dual neural network (mLVI-PDNN). Experimental results demonstrate the distance errors in multiple pedicle screw implantation tasks are$0.738\pm 0.080$mm and$0.154\pm 0.031$mm at the entry and target points, respectively. And the angular errors between the actual drilling path compared to the planned screw placement path are$0.005\pm 0.002$degrees. These results show that the proposed methodology offers a reliable and innovative solution for the higher level of automation in the pedicle screw implantation procedure.Note to Practitioners—The study is driven by the need for higher levels of automation in orthopedic surgery for pedicle screw implantation tasks. Orthopedic surgical robots are usually deployed in complicated surgeries due to their excellent positioning capabilities. However, they primarily rely on basic kinematic constraints and cooperative control modes, thus lacking a reliable mechanism for active surgical pose alignment and screw path drilling, especially in complex surgical scenarios that involve only optical tracking systems (OTS) without additional vision sensors. For example, current popular OTS systems have a limited field of view and cannot fully acquire information about the surgical scene, which impacts collision-free path planning. To address these issues, a novel methodology combining CDMP based imitation learning with QP based motion control for the task of pedicle screw implantation is introduced. Incorporating surgeons’ prior knowledge of the surgical environment and their participation in the robot’s motion planning using the CDMP model enhances the robot’s autonomy and increases reliability in surgeries, without requiring additional vision sensors. This approach allows for accurate reconstruction of multiple surgical targets of automatic surgical pose alignment and pedicle screw drilling, while respecting manipulability optimization and the robot’s kinematic constraints, rather than merely imitating and ignoring the robot’s kinematic structure. This methodology is also applicable to various clinical surgical operations and industrial applications. Xingqiang Jian, Yibin Song, Yu Wang 0244, Xueqian Guo, Bo Wu 0016, Nan Zhang 0015 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2024 | Less Is More: Label Recommendation for Weakly Supervised Point Cloud Semantic SegmentationabstractWeak supervision has proven to be an effective strategy for reducing the burden of annotating semantic segmentation tasks in 3D space. However, unconstrained or heuristic weakly supervised annotation forms may lead to suboptimal label efficiency. To address this issue, we propose a novel label recommendation framework for weakly supervised point cloud semantic segmentation. Distinct from pre-training and active learning, the label recommendation framework consists of three stages: inductive bias learning, recommendations for points to be labeled, and point cloud semantic segmentation learning. In practice, we first introduce the point cloud upsampling task to induct inductive bias from structural information. During the recommendation stage, we present a cross-scene clustering strategy to generate centers of clustering as recommended points. Then we introduce a recommended point positions attention module LabelAttention to model the long-range dependency under sparse annotations. Additionally, we employ position encoding to enhance the spatial awareness of semantic features. Throughout the framework, the useful information obtained from inductive bias learning is propagated to subsequent semantic segmentation networks in the form of label positions. Experimental results demonstrate that our framework outperforms weakly supervised point cloud semantic segmentation methods and other methods for labeling efficiency on S3DIS and ScanNetV2, even at an extremely low label rate. Zhiyi Pan 0001, Nan Zhang 0015, Wei Gao 0003, Shan Liu 0001, Ge Li 0002 |
AAAI | 2 |
| 2024 | MFITrack: Multi-Frame Integration Strategy for Enhanced Motion-Centric Single Object Trackingabstract3D Single Object Tracking (SOT) in LiDAR point clouds is essential for applications like autonomous driving and surveillance, requiring precise tracking for safety and efficiency. Existing methods often face challenges in adapting to changes in target appearance and capturing motion details in complex environments. To overcome these challenges, in this paper, we introduce a novel three-stage tracker employing a multi-frame integration strategy, dubbed MFITrack. It comprises three stages: multi-frame probabilistic data selection, multi-frame motion prediction and bounding box optimization. MFITrack effectively integrates data from multiple frames, enhancing the extraction of motion information and understanding of the target’s trajectory and behavior. Tested on KITTI and Nuscenes datasets, MFI-Track demonstrates superior performance over existing methods, showcasing robustness and precision in varied tracking scenarios, particularly in the more complex Nuscenes dataset. Pochun Chen, Nan Zhang 0015, Ge Li 0002 |
ICME | 2 |
| 2024 | MPVNN: Multi-resolution Point-Voxel Non-parametric Network for 3D Point Cloud ProcessingabstractRecent non-parametric networks have demonstrated the feasibility of point cloud tasks through skillfully combining Farthest point sampling, k-nearest neighbor, and pooling algorithms. However, these point-based algorithms struggle to capture features containing fine-grained spatial structures within a point cloud, thereby limiting the performance of non-parametric networks. To address it, we propose the Multi-resolution Point-Voxel Non-parametric Network (MPVNN) to enhance the structural information of voxelized point clouds at varying resolutions. Specifically, we design a Multi-Resolution Voxel Encoder (MR-VEnc) that processes point clouds with multiple resolutions and adopts trilinear interpolation to synthesize features with respect to 8 voxel grid vertices. Then, we amalgamate features derived from the point-based algorithms with those from MRVEnc to establish a point memory bank for subsequent tasks. Extensive experiments on ModelNet40, 3D-FUTURE, and ShapeNetPart demonstrate the superior performance of MPVNN in both classification and segmentation tasks. Keli Wen, Nan Zhang 0015, Ge Li 0002, Wei Gao 0003 |
ICME | 2 |
| 2023 | Improving Graph Representation for Point Cloud Segmentation via Attentive FilteringabstractRecently, self-attention networks achieve impressive performance in point cloud segmentation due to their superiority in modeling long-range dependencies. However, compared to self-attention mechanism, we find graph convolutions show a stronger ability in capturing local geometry information with less computational cost. In this paper, we employ a hybrid architecture design to construct our Graph Convolution Network with Attentive Filtering (AF-GCN), which takes advantage of both graph convolution and selfattention mechanism. We adopt graph convolutions to aggregate local features in the shallow encoder stages, while in the deeper stages, we propose a self-attention-like module named Graph Attentive Filter (GAF) to better model long-range contexts from distant neighbors. Besides, to further improve graph representation for point cloud segmentation, we employ a Spatial Feature Projection (SFP) module for graph convolutions which helps to handle spatial variations of unstructured point clouds. Finally, a graphshared down-sampling and up-sampling strategy is introduced to make full use of the graph structures in point cloud processing. We conduct extensive experiments on multiple datasets including S3DIS, ScanNetV2, Toronto-3D, and ShapeNetPart. Experimental results show our AF-GCN obtains competitive performance. Nan Zhang 0015, Zhiyi Pan 0001, Thomas H. Li, Wei Gao 0003, Ge Li 0002 |
CVPR | 1 |
| 2023 | Null-Space Diffusion Sampling for Zero-Shot Point Cloud CompletionabstractPoint cloud completion aims at estimating the complete data of objects from degraded observations. Despite existing completion methods achieving impressive performances, they rely heavily on degraded-complete data pairs for supervision. In this work, we propose a novel framework named Null-Space Diffusion Sampling (NSDS) to solve the point cloud completion task in a zero-shot manner. By leveraging a pre-trained point cloud diffusion model as the off-the-shelf generator, our sampling approach can generate desired completion outputs with the guidance of the observed degraded data without any extra training. Furthermore, we propose a tolerant loop mechanism to improve the quality of completion results for hard cases. Experimental results demonstrate our zero-shot framework achieves superior completion performance than unsupervised methods and comparable performance to supervised methods in various degraded situations. Xinhua Cheng, Nan Zhang 0015, Jiwen Yu, Yinhuai Wang, Ge Li 0002 |
IJCAI | 2 |
| 2021 | Virtual view synthesis for the nonuniform illuminated between views in surgical video
Boqi Jia, Nan Zhang 0015, Shiqi Wang 0001, Bo Wu 0016 |
Multim. Tools Appl. | 2 |
| 2017 | Fast intra coding unit size decision for HEVC with GPU based keypoint detectionabstractIn this paper, a fast intra Coding Unit (CU) size decision framework based on keypoint detection on Graphic Processing Unit (GPU) is proposed. In this framework, firstly the original frames are sent to GPU and then keypoint detection is conducted with numerous threads, which is able to avoid bringing in additional computational complexity even in realtime systems. Then, based on the keypoint distribution, whether to split the CU to the next coding depth is efficiently predicted. Experiments show that the proposed algorithm can achieve over 25% time saving under all intra (AI) configuration with ignorable performance loss. Falei Luo, Shanshe Wang, Siwei Ma 0001, Nan Zhang 0015, Wen Gao 0001 |
ISCAS | 4 |
| 2017 | Rate-distortion optimized scan for point cloud color compressionabstractWe propose an adaptive scanning scheme to improve color attribute compression performance of static point cloud. To better exploit the intrinsic correlations of point cloud color data, multiple scanning modes are dedicated to projecting the color attributes into a series of texture blocks before compression. The best mode is subsequently selected in the sense of rate-distortion optimization, such that the compression flexibility and performance can be greatly improved. To achieve rate-distortion optimized scan, the Lagrange multiplier that balances the rate and distortion is off-line derived based on the statistics of color attribute coding. The proposed algorithm is implemented in reference software for point cloud compression (PCC-MP3DG), which is introduced by 3D Graphics (3DG) group of Motion Picture Experts Group (MPEG). Experimental results show that proposed method performs significantly better than the PCC-MP3DG reference software, and on average 6.51% and 9.36% bit-rate savings can be achieved for high and low bit rate codings, respectively. Yiqun Xu, Shanshe Wang, Xinfeng Zhang 0001, Shiqi Wang 0001, Nan Zhang 0015, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 5 |
| 2016 | Adaptive Motion Vector Resolution Scheme for Enhanced Video CodingabstractIn the state-of-the-art H.265/HEVC video coding standard, the motion vector is always fixed to be 1/4-pixel resolution for the entire video sequence regardless of the different video contents, which is not efficient for prediction coding. In this paper, we propose a frame level adaptive motion vector resolution selection scheme based on a rate-distortion model in terms of motion vector resolution. In the proposed rate-distortion model, the relationship between the distortion and the motion vector resolution is approximated with a linear model. And a rate model of motion vector is built, which reflects the relationship between the coding bits of motion vector and its value. With the proposed rate-distortion model, an optimal motion vector resolution minimizing the total rate-distortion cost will be selected for each frame. Experimental results show that the proposed scheme can achieve 1.5%, 1.3% and 2.5% BD-rate gain on average for Random Access, Lowdelay-B and Lowdelay-P configurations without complexity increment. Zhao Wang 0004, Jian Zhang 0018, Nan Zhang 0015, Siwei Ma 0001 |
DCC | 3 |
| 2016 | Structure-driven Adaptive Non-local Filter for High Efficiency Video Coding (HEVC)abstractDeblocking filter (DF) Is High Efficiency Video Coding (HEVC) is Only Applied to all Samples Adjacent to prediction units (PU), or transform units (TU), which actually exists two issues. The first one is that DF in HEVC does not fully exploit nonlocal similarity structure information in video. The second one is that DF is HEVC does not consider the inside pixels, which often suffer from quantization distrotion. To alleviate these issues, in this paper, a structure-driven adaptive non-local filter (SANF) Is Proposed By Simultaneously Enforcing The Intrinsic Local Sparsity And The Non-Local Self-Similarity Of Each Frame. Not only SANF deals with the boundary pixels, but also the inside area, which is able to effectively reduce block artifacts while enhancing the quality of the deblocked frames. Applying SANF to luma and chroma components after DF, simulation results demonstrate that the proposed SANF can save BD-rate reduction up to 10.3% with ALF off. For luma component, SANF achieves 4.1%. 3.3%, 4.4% BD-rate saving for all intra, low delay B and random access configurations, respectively with ALF off. furthermore, the performance with ALF on is also discussed. Jian Zhang 0018, Chuanmin Jia, Nan Zhang 0015, Siwei Ma 0001, Wen Gao 0001 |
DCC | 3 |
| 2016 | Compact and robust video fingerprinting using sparse represented featuresabstractIn this paper, we propose a compact and robust video fingerprinting scheme by using sparse represented features (SRF). The SRF are extracted by a two dimensional matching pursuit decomposition (2D-MPD) method. The motivation of using sparse features is that the sparse coding method can significantly reduce the data dimensionality and effectively retain the structure of the images. To further reduce the length of the fingerprint, a two-stage cascade SVD method is applied. The SVD-based feature extraction method can improve the robustness to certain attacks, such as geometric attack. Then, a locally adaptive quantization (AQ) method which considers the local probability distribution of the sample is applied. This method can quantize the real-valued fingerprints into binary bits without degrading the detection performance too much. At last, a secret key based interleaving method is applied to the binary fingerprints. The interleaving can enlarge the Hamming distance of the fingerprints, so the detection performance is improved. According to the experimental results, the proposed method offers a favourable robustness versus discriminability tradeoff over the state-of-the-art video fingerprint methods. Bo Wu 0016, Sridhar Krishnan 0001, Nan Zhang 0015, Li Su 0003 |
ICME | 3 |
| 2016 | GPU based sample adaptive offset parameter decision and perceptual optimization for HEVCabstractIn this paper, a graphics processing unit (GPU) based sample adaptive offset (SAO) parameters decision scheme is proposed for High Efficiency Video Coding (HEVC). Then, in order to further improve the performance of SAO, a perceptual based optimization scheme is provided according to the adjustment of Lagrange multiplier aiming to improve the subjective performance of SAO. Experimental results demonstrate that the proposed GPU based SAO parameter decision scheme can achieve average 0.76% and 0.78% BD-rate gain in terms of PSNR (Peak Signal to Noise Ratio) and SSIM (Structure Similarity) respectively. Combined with the perceptual optimization scheme, the maximum BD-rate gain in terms of PSNR and SSIM can be up to 1.77% and 3.3% with the average as 1.23% and 1.37%. Moreover, much computation complexity of SAO can be distributed to GPU. Falei Luo, Shanshe Wang, Nan Zhang 0015, Siwei Ma 0001, Wen Gao 0001 |
ISCAS | 3 |
| 2015 | Image guided label map propagation in video sequencesabstractIn this paper, we propose a novel method to transmit the label maps by propagating from a key frame to non-key frames. The label map of a non-key frame is initialized by warping the label map of its corresponding key frame according to the motion estimation between them. Subsequently, the initialized label map is optimized with the guidance of its texture image. The optimization process minimizes an energy function which takes two constraints into consideration: (i) the data term measuring the similarity between an estimated label map and its initialized one, (ii) the regularization term enforcing the local smoothness in the label map and the consistency of region boundaries between the estimated label map and its corresponding texture image. Graph cuts based computation process is finally performed to generate the optimized label map. Experimental results show that our method achieves higher accuracy and better visual quality comparing with the state-of-the-art method. Shuolin Di, Zhebin Zhang, Shiqi Wang 0001, Nan Zhang 0015, Siwei Ma 0001 |
ISCAS | 4 |
| 2015 | Adaptive boundary dependent transform optimization for HEVCabstractHigh Efficiency Video Coding (HEVC) adopts hybrid transform coding scheme to improve the transform performance. However, it does not consider the influence of the prediction unit boundary. This paper proposes an adaptive boundary dependent transform optimization scheme for HEVC. Based on the transform unit boundary type identification by checking whether it lies at a prediction unit boundary or not, some additional transform kernels, including DCT-IV, are incorporated for transform. Moreover, the additional adopted transform kernels are generated by re-using the existing kernels in HEVC. Therefore, the butterfly structure for transform implementation can be preserved to facilitate parallel computing, and computational complexity increasing can be avoided. Experimental results show that the coding performance improvements can be up to 0.96% and 0.64% for low delay P and random access testing configurations respectively. Juanting Fan, Jicheng An, Shanshe Wang, Nan Zhang 0015, Ruiqin Xiong, Siwei Ma 0001, Shawmin Lei |
PCS | 4 |
| 2015 | An optimized probability estimation model for binary arithmetic codingabstractIn this paper, we analyze the binary arithmetic coding of High Efficiency Video Coding (HEVC) and the second generation of audio and video coding standard (AVS2). Then an optimized probability estimation scheme is proposed for arithmetic coder. The proposed scheme is incorporated into the HEVC reference software (HM 16.0) and AVS2 reference software (RD 10.1). Experimental results demonstrate that the proposed scheme can efficiently improve the coding efficiency of entropy coding. The rate-distortion (R-D) performance gain can be up to 0.21% for AVS2 and 0.30% for HEVC respectively. Shanshe Wang, Nan Zhang 0015, Siwei Ma 0001 |
VCIP | 3 |
| 2015 | Parallel intra coding for HEVC on CPU plus GPU platformabstractIn High Efficiency Video Coding (HEVC), the intra coding performance is significantly improved due to the recursive splitting structure and up to 35 intra prediction modes. However, the computational complexity of intra coding increases largely as well. In this paper, a fast intra coding scheme is proposed based on CPU and GPU cooperation. Firstly, the intra prediction of variable blocks is performed in parallel on multi-cores GPU. Secondly, the intra prediction mode with minimum Sum of Absolute Difference (SAD) cost is selected and transmitted to the host CPU. Instead of exhaustively searching all the intra modes in Rough Mode Decision (RMD) process, the mode returned by the GPU is directly selected. Lastly, the texture gradient of each coding unit (CU) is assessed during parallel intra prediction, then used by the CPU for fast CU size decision. Experiment results show that the proposed parallel intra coding method achieves up to 62% complexity reduction with acceptable coding performance loss. Juncheng Ma, Falei Luo, Shanshe Wang, Nan Zhang 0015, Siwei Ma 0001 |
VCIP | 4 |
| 2014 | Optimal entropy-constrained non-uniform scalar quantizer design for low bit-rate pixel domain DVC
Bo Wu 0016, Nan Zhang 0015, Siwei Ma 0001, Debin Zhao, Wen Gao 0001 |
Multim. Tools Appl. | 2 |
| 2013 | Content adaptive in-loop depth map filter for HEVC based 3DV codingabstractIn this paper, a content adaptive in-loop depth map filter is proposed for HEVC based 3DV video coding to improve the quality of the synthesized views. The proposed depth map filtering scheme is block based by reusing the TU (transform unit) split structure in HEVC. Firstly, TU are classified into 3 categories: flat, directional, and textureless, according to the characteristics of the depth map. Then four kinds of filters are designed by considering the characteristics of the current TU and neighbor TUs jointly. The proposed scheme is incorporated into HTM3.0- the reference software of HEVC based 3D video coding. The experimental results show both subjective and objective quality improvement. The average BD-rate reduction is 0.65% for synthesis views in HD test sequences, and the maximum BD-rate reduction is up to 1.6%. Jianqiang He, Siwei Ma 0001, Nan Zhang 0015, Wen Gao 0001 |
ICASSP | 3 |
| 2011 | Joint just noticeable difference model based on depth perception for stereoscopic imagesabstractJust noticeable difference (JND) model can reflect the least perceptible distortion from images, including 2D images and stereoscopic images. As we know, for the perception of human visual system (HVS), stereoscopic images have quite different characteristics from 2D images, since stereoscopic images contain not only planar information, but also depth information. This paper proposes a joint JND (JJND) model based on depth perception for stereoscopic images. Firstly, disparity estimation is performed in order to decompose the image into the occlusion region and the non-overlapped region. Then, different JND thresholds are applied on different regions, according to the depth information of the region, which can be derived from the disparity of the region. Experimental results verified our model's validity for stereoscopic images. Xiaoming Li 0002, Yue Wang 0032, Debin Zhao, Tingting Jiang 0001, Nan Zhang 0015 |
VCIP | 5 |
| 2011 | Improved spatial aided low delay Wyner-Ziv video coding by wavelet shrinkageabstractTo tackle the problem of lacking auxiliary information while generating the side information, we have proposed a novel spatial aided low delay Wyner-Ziv video coding scheme in which the wavelet transformation is applied to generate the auxiliary information. However, there exists an inefficiency that lowers the coding performance. The inefficiency is that many wavelet coefficients having very small absolute value in the high-pass subbands and these small coefficients decrease the correlation between the source and the side information. The diminished correlation makes the bit error rate increased sharply and declines the performance of the channel decoding. This inefficiency is a widespread phenomenon among the wavelet based Wyner-Ziv video coding schemes. Therefore, we propose a scheme to remove these noises like coefficients in high-pass subbands by using a wavelet shrinkage method. In our method, a soft-thresholding operator which is based on the statistical properties of the high-pass coefficients is proposed. The result reveals that the correlation between the smoothed WZ frame and the side information is improved by using the wavelet shrinkage method. Accordingly, the R-D performance is also improved. Bo Wu 0016, Nan Zhang 0015, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 2 |
| 2009 | Compression-Induced Rendering Distortion Analysis for Texture/Depth Rate Allocation in 3D Video CompressionabstractIn 3D video applications, the virtual view is generally rendered by the compressed texture and depth. The texture and depth compression with different bit-rate overheads can lead to different virtual view rendering qualities. In this paper, we analyze the compression-induced rendering distortion for the virtual view. Based on the 3D warping principle, we first address how the texture and depth compression affects the virtual view quality, and then derive an upper bound for the compression-induced rendering distortion. The derived distortion bound depends on the compression-induced depth error and texture intensity error. Simulation results demonstrate that the theoretical upper bound is an approximate indication of the rendering quality and can be used to guide sequence-level texture/depth rate allocation for 3D video compression. Yanwei Liu 0001, Siwei Ma 0001, Qingming Huang, Debin Zhao, Wen Gao 0001, Nan Zhang 0015 |
DCC | 6 |
| 2008 | 2D to 3D convertion based on edge defocus and segmentationabstractThis paper presents a depth estimation method which converts two-dimensional images into three-dimensional data. Based on two-dimensional wavelet analysis of Lipschitz regularity for defocus estimation on edges, this method can effectively eliminate the horizontal stripes in the depth map resulted from traditional one-dimensional wavelet based approaches. Besides, we also propose several techniques such as edge enhancement, color-based segmentation, and depth optimization to obtain a more reliable and smoother depth map. The experimental results demonstrate the effectiveness of our proposed techniques. Ge Guo 0002, Nan Zhang 0015, Longshe Huo, Wen Gao 0001 |
ICASSP | 2 |