EDBT 2026 Demo / reviewers in the wild / expert
Meng Yang 0002
dblp:44/2761-2
· DBLP profile ↗
32ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0002-0525-5059ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RayD3D: Distilling Depth Knowledge Along the Ray for Robust Multi-View 3D Object DetectionabstractMulti-view 3D detection with bird’s eye view (BEV) is crucial for autonomous driving and robotics, but its robustness in real-world is limited as it struggles to predict accurate depth values. A mainstream solution, cross-modal distillation, transfers depth information from LiDAR to camera models but also unintentionally transfers depth-irrelevant information (e.g. LiDAR density). To mitigate this issue, we propose RayD3D, which transfers crucial depth knowledge along the ray: a line projecting from the camera to true location of an object. It is based on the fundamental imaging principle that predicted location of this object can only vary along this ray, which is finally determined by predicted depth value. Therefore, distilling along the ray enables more effective depth information transfer. More specifically, we design two ray-based distillation modules. Ray-based Contrastive Distillation (RCD) incorporates contrastive learning into distillation by sampling along the ray to learn how LiDAR accurately locates objects. Ray-based Weighted Distillation (RWD) adaptively adjusts distillation weight based on the ray to minimize the interference of depth-irrelevant information in LiDAR. For validation, we widely apply RayD3D into three representative types of BEV-based models, including BEVDet, BEVDepth4D, and BEVFormer. Our method is trained on clean NuScenes, and tested on both clean NuScenes and RoboBEV with a variety types of data corruptions. Our method significantly improves the robustness of all the three base models in all scenarios without increasing inference costs, and achieves the best when compared to recently released multi-view and distillation models. Zhaonian Kuang, Zongwei Zhou, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
AAAI | 4 |
| 2026 | Object-Scene-Camera Decomposition and Recomposition for Data Efficient Monocular 3D Object Detection
Zhaonian Kuang, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | Relative depth knowledge distillation for generalizable monocular depth estimation
Mankun Li, Meng Yang 0002, Xuguang Lan, Ce Zhu |
Neurocomputing | 3 |
| 2026 | Multi-Modal Decouple and Recouple Network for Robust 3D Object DetectionabstractMulti-modal 3D object detection with bird’s eye view (BEV) has achieved desired advances on benchmarks. Nonetheless, the accuracy may drop significantly in the real world due to data corruption such as sensor configurations for LiDAR and scene conditions for camera. One design bottleneck of previous models resides in the tightly coupling of multi-modal BEV features during fusion, which may degrade the overall system performance if one modality or both is corrupted. To mitigate, we propose a Multi-Modal Decouple and Recouple Network for robust 3D object detection under data corruption. Different modalities commonly share some high-level invariant features. We observe that these invariant features across modalities do not always fail simultaneously, because different types of data corruption affect each modality in distinct ways. These invariant features can be recovered across modalities for robust fusion under data corruption. To this end, we explicitly decouple Camera/LiDAR BEV features into modality-invariant and modality-specific parts. It allows invariant features to compensate each other while mitigates the negative impact of a corrupted modality on the other. We then recouple these features into three experts to handle different types of data corruption, respectively, i.e., LiDAR, camera, and both. For each expert, we use modality-invariant features as robust information, while modality-specific features serve as a complement. Finally, we adaptively fuse the three experts to exact robust features for 3D object detection. For validation, we collect a benchmark with a large quantity of data corruption for LiDAR, camera, and both based on nuScenes. Our model is trained on clean nuScenes and tested on all types of data corruption. Our model consistently achieves the best accuracy on both corrupted and clean data compared to recent models. Zhaonian Kuang, Yuzhe Ji, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Consistent Feature Alignment for Cross-Modal Knowledge Distillation in Monocular 3D Object DetectionabstractCross-modal knowledge distillation (CMKD) in monocular 3D object detection transfers LiDAR’s accurate depth information to compensate for the limitations of camera model. However, current methods directly align the intermediate features of the teacher and student networks, in which the modality gap between LiDAR and camera hinders their effectiveness. To mitigate this issue, we design two modules, namely, Consistent Alignment Module (CAM) and Deformable Adapter Module (DAM) to reduce the modality gap of CMKD. The CAM transforms intermediate features of LiDAR and camera into some consistent features through a lightweight Target Head. It is based on the observation that some high-level features such as heatmaps and depths are highly correlated in CMKD, though modality gap appears between LiDAR and camera. Therefore, these features can be effectively transferred from teacher to student in CMKD. The DAM introduces a deformable adapter for the intermediate features of the student network to reduce background noise in CMKD. This helps to dynamically align its intermediate features with the teacher network. We then propose a Consistent Feature Alignment network (MonoCFA) for CMKD to boost monocular 3D object detection. Our network integrates the two designed modules at different levels of the teacher and student networks, in order to align the intermediate features of LiDAR and camera more accurately and reliably. Our model can be widely applied to existing monocular 3D object detection models. For validation, we choose the representative MonoDLE, GUPNet, and DID-M3D as base models. Experiments on the KITTI benchmark show that our method significantly outperforms the three base models by 39%, 15.5%, and 15%, respectively, and achieves state-of-the-art when compared to other CMKD models. Meng Yang 0002, Xuguang Lan |
IROS | 3 |
| 2025 | Scale Propagation Network for Generalizable Depth CompletionabstractDepth completion, inferring dense depth maps from sparse measurements, is crucial for robust 3D perception. Although deep learning based methods have made tremendous progress in this problem, these models cannot generalize well across different scenes that are unobserved in training, posing a fundamental limitation that yet to be overcome. A careful analysis of existing deep neural network architectures for depth completion, which are largely borrowing from successful backbones for image analysis tasks, reveals that a key design bottleneck actually resides in the conventional normalization layers. These normalization layers are designed, on one hand, to make training more stable, on the other hand, to build more visual invariance across scene scales. However, in depth completion, the scale is actually what we want to robustly estimate in order to better generalize to unseen scenes. To mitigate, we propose a novel scale propagation normalization (SP-Norm) method to propagate scales from input to output, and simultaneously preserve the normalization operator for easy convergence. More specifically, we rescale the input using learned features of a single-layer perceptron from the normalized input, rather than directly normalizing the input as conventional normalization layers. We then develop a new network architecture based on SP-Norm and the ConvNeXt V2 backbone. We explore the composition of various basic blocks and architectures to achieve superior performance and efficient inference for generalizable depth completion. Extensive experiments are conducted on six unseen datasets with various types of sparse depth maps, i.e., randomly sampled 0.1%/1%/10% valid pixels, 4/8/16/32/64-line LiDAR points, and holes from Structured-Light. Our model consistently achieves the best accuracy with faster speed and lower memory when compared to state-of-the-art methods. Haotian Wang 0009, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Depth Map Super-Resolution via Deep Cross-Modality and Cross-Scale GuidanceabstractGuided depth super-resolution is essential in many applications, which enhances low-resolution (LR) depth maps using high-resolution (HR) RGB images from the same scene. However, the challenge lies in avoiding the texture-copy artifacts issue caused by structural inconsistencies between two modalities. To mitigate, we propose a cross-modality and cross-scale guided depth super-resolution network (D2CNet). We first design a novel two-stage feature integration module to effectively fuse multi-modal RGB and depth while minimizing texture-copy artifacts. That is, a cross-modality fusion stage transfers consistent structures from RGB to depth in a multi-scale manner, and a cross-scale refinement stage mitigates inconsistent structures across modalities. In addition, we design a convolution group as the basic module to well extract high-frequency features and an LR and HR domain projection strategy to enrich features between the fusion and refinement stages. We then develop a new network architecture by progressively repeating the feature integration module and the convolution group, which is flexibly controllable to strike a balance between accuracy and cost for easy implementation in real world. Extensive experiments on multiple benchmarks demonstrate that our D2CNet consistently achieves superior accuracy and generalization ability across sampling scales in both qualitative and quantitative evaluations, when compared to state-of-the-art baselines. Shuzhe Liu, Delong Suzhang, Meng Yang 0002, Xinhu Zheng, Ce Zhu |
IEEE Trans. Multim. | 3 |
| 2024 | G2-MonoDepth: A General Framework of Generalized Depth Inference From Monocular RGB+X DataabstractMonocular depth inference is a fundamental problem for scene perception of robots. Specific robots may be equipped with a camera plus an optional depth sensor of any type and located in various scenes of different scales, whereas recent advances derived multiple individual sub-tasks. It leads to additional burdens to fine-tune models for specific robots and thereby high-cost customization in large-scale industrialization. This article investigates a unified task of monocular depth inference, which infers high-quality depth maps from all kinds of input raw data from various robots in unseen scenes. A basic benchmark G2-MonoDepth is developed for this task, which comprises four components: (a) a unified data representation RGB+X to accommodate RGB plus raw depth with diverse scene scale/semantics, depth sparsity ([0%, 100%]) and errors (holes/noises/blurs), (b) a novel unified loss to adapt to diverse depth sparsity/errors of input raw data and diverse scales of output scenes, (c) an improved network to well propagate diverse scene scales from input to output, and (d) a data augmentation pipeline to simulate all types of real artifacts in raw depth maps for training. G2-MonoDepth is applied in three sub-tasks including depth estimation, depth completion with different sparsity, and depth enhancement in unseen scenes, and it always outperforms SOTA baselines on both real-world data and synthetic data. Haotian Wang 0009, Meng Yang 0002, Nanning Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Selective Transfer Learning of Cross-Modality Distillation for Monocular 3D Object DetectionabstractMonocular 3D object detection is a promising yet ill-posed task for autonomous vehicles due to the lack of accurate depth information. Cross-modality knowledge distillation could effectively transfer depth information from LiDAR to image-based network. However, modality gap between image and LiDAR seriously limits its accuracy. In this paper, we systematically investigate the negative transfer problem induced by modality gap in cross-modality distillation for the first time, including not only the architecture inconsistency issue but more importantly the feature overfitting issue. We propose a selective learning approach named MonoSTL to overcome these issues, which encourages positive transfer of depth information from LiDAR while alleviates the negative transfer on image-based network. On the one hand, we utilize similar architectures to ensure spatial alignment of features between image-based and LiDAR-based networks. On the other hand, we develop two novel distillation modules, namely Depth-Aware Selective Feature Distillation (DASFD) and Depth-Aware Selective Relation Distillation (DASRD), which selectively learn positive features and relationships of objects by integrating depth uncertainty into feature and relation distillations, respectively. Our approach can be seamlessly integrated into various CNN-based and DETR-based models, where we take three recent models on KITTI and a recent model on NuScenes for validation. Extensive experiments show that our approach considerably improves the accuracy of the base models and thereby achieves the best accuracy compared with all recently released SOTA models. The code is released on https://github.com/DingCodeLab/MonoSTL. Meng Yang 0002, Nanning Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | FS-Depth: Focal-and-Scale Depth Estimation From a Single Image in Unseen Indoor SceneabstractIt has long been an ill-posed problem to predict absolute depth maps from single images in unseen scenes. We observe that it is essentially due to not only the scale-ambiguous problem, but more importantly, the focal-ambiguous problem that decreases the generalization ability of monocular depth estimation. That is, images may be captured by cameras of different focal lengths in scenes of different scales. In this paper, we develop a focal-and-scale depth estimation model to well learn absolute depth maps from single images in unseen indoor scenes. First, a relative depth estimation network is adopted to learn relative depths from single images with diverse scales. Second, multi-scale features are generated by mapping a single focal length value to focal length features and concatenating them with intermediate features of different scales in relative depth estimation. Finally, relative depths and multi-scale features are jointly fed into an absolute depth estimation network. Our model is enabled to be well trained on either a single dataset or a mixed dataset with diverse focal lengths and scene scales by a dual-directional alignment strategy. In addition, a new pipeline is developed to augment the diversity of focal lengths of public datasets, which are often captured with cameras of the same or similar focal lengths. The experiments verify that our model trained on NYUDv2 significantly improves the generalization ability of monocular depth estimation by 32%/14% (RMSE) on three unseen datasets with/without data augmentation compared with state-of-the-art (SOTA) baselines, and well alleviates the deformation problem of depth maps in 3D view. The generalization ability is further improved by 16% when the model is trained on a mixture of NYUDv2 and SUNRGBD. In addition, our model maintains a SOTA accuracy, when it is trained and tested on NYUDv2 similar to existing models. The code is released onhttps://github.com/wcrwcrwcr/FS-Depth-v1. Chengrui Wei, Meng Yang 0002, Nanning Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Surface normal and Gaussian weight constraints for indoor depth structure completion
Dongran Ren, Meng Yang 0002, Jiangfan Wu, Nanning Zheng 0001 |
Pattern Recognit. | 2 |
| 2023 | RGB-Guided Depth Map Recovery by Two-Stage Coarse-to-Fine Dense CRF ModelsabstractDepth maps generally suffer from large erroneous areas even in public RGB-Depth datasets. Existing learning-based depth recovery methods are limited by insufficient high-quality datasets and optimization-based methods generally depend on local contexts not to effectively correct large erroneous areas. This paper develops an RGB-guided depth map recovery method based on the fully connected conditional random field (dense CRF) model to jointly utilize local and global contexts of depth maps and RGB images. A high-quality depth map is inferred by maximizing its probability conditioned upon a low-quality depth map and a reference RGB image based on the dense CRF model. The optimization function is composed of redesigned unary and pairwise components, which constraint local structure and global structure of depth map, respectively, with the guidance of RGB image. In addition, the texture-copy artifacts problem is handled by two-stage dense CRF models in a coarse-to-fine way. A coarse depth map is first recovered by embedding RGB image in a dense CRF model in unit of $3\times 3$ blocks. It is refined afterward by embedding RGB image in another model in unit of individual pixels and restricting the model mainly work in discontinued regions. Extensive experiments on six datasets verify that the proposed method considerably outperforms a dozen of baseline methods in correcting erroneous areas and diminishing texture-copy artifacts of depth maps. Haotian Wang 0009, Meng Yang 0002, Ce Zhu, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Subjective low-light image enhancement based on a foreground saliency map model
Pengcheng Hao, Meng Yang 0002, Nanning Zheng 0001 |
Multim. Tools Appl. | 2 |
| 2022 | Depth Map Recovery Based on a Unified Depth Boundary Distortion ModelabstractDepth maps acquired by either physical sensors or learning methods are often seriously distorted due to boundary distortion problems, including missing, fake, and misaligned boundaries (compared with RGB images). An RGB-guided depth map recovery method is proposed in this paper to recover true boundaries in seriously distorted depth maps. Therefore, a unified model is first developed to observe all these kinds of distorted boundaries in depth maps. Observing distorted boundaries is equivalent to identifying erroneous regions in distorted depth maps, because depth boundaries are essentially formed by contiguous regions with different intensities. Then, erroneous regions are identified by separately extracting local structures of RGB image and depth map with Gaussian kernels and comparing their similarity on the basis of the SSIM index. A depth map recovery method is then proposed on the basis of the unified model. This method recovers true depth boundaries by iteratively identifying and correcting erroneous regions in recovered depth map based on the unified model and a weighted median filter. Because RGB image generally includes additional textural contents compared with depth maps, texture-copy artifacts problem is further addressed in the proposed method by restricting the model works around depth boundaries in each iteration. Extensive experiments are conducted on five RGB-depth datasets including depth map recovery, depth super-resolution, depth estimation enhancement, and depth completion enhancement. The results demonstrate that the proposed method considerably improves both the quantitative and visual qualities of recovered depth maps in comparison with fifteen competitive methods. Most object boundaries in recovered depth maps are corrected accurately, and kept sharply and well aligned with the ones in RGB images. Haotian Wang 0009, Meng Yang 0002, Xuguang Lan, Ce Zhu, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | SynBF: A New Bilateral Filter for Postremoval of Noise From Synthesis Views in 3-D VideoabstractIn 3-D video systems, noise in the texture and depth videos of reference views may not be removed (Scenario 1) or not fully removed (Scenario 2) by prefiltering methods before the view synthesis procedure. In these scenarios, the noise is transferred to the generated synthesis view. After investigating the noise model of the synthesis view, we conclude that the noise in the synthesis view not only causes fluctuation in the photometric values of pixels in the range domain but also additionally shifts the positions of neighbor pixels in the spatial domain compared to that in natural images. It consequently damages the textural content near edges in the synthesis view, for which the popular local filters of natural images, that is, the bilateral filter (BF) and the guided filter, do not work well. In this paper, we develop a new local filter for the synthesis view (named SynBF) after it has been generated, which has a similar expression as that of the BF but not the exact same weight terms. On one hand, the spatial term in the classical BF is directly reused due to its robustness to noise, which gives high weights to spatially closed pixels of to-be-filtered pixels. On the other hand, a reliability term is designed that gives high weights to pixels that are unlikely to be affected by noise. It is inspired by the finding that not all pixels are significantly affected by noise in the synthesis view. In this way, true edge profiles are protected in the filtering process. Experiments are conducted on a set of synthesis views for both scenarios above and compared to the two local filters, which verifies its effectiveness in removing noise and protecting edge profiles. The proposed method can be considered as a supplement to prefiltering methods of texture/depth videos in 3-D video systems. Meng Yang 0002, Nanning Zheng 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Efficient Estimation of View Synthesis Distortion for Depth Coding OptimizationabstractDepth coding in depth-based three-dimensional (3-D) video is unique in that its quality is measured by view synthesis distortion (VSD) rather than the depth distortion itself, which further complicates the coding optimization as the VSD is related to quality of both the associated depth and texture videos. In this paper, an efficient VSD estimation scheme is developed to measure the effect of depth errors on the VSD for a block given its depth distortion in mean-squared error. Unlike other relevant VSD models which involve computationally intensive parameter training or Fourier transform, the proposed scheme is free of parameter training, while taking the advantage of integer 4 × 4 discrete Cosine transform to replace Fourier transform, thus well-saving computational cost and diminishing sensitivity to training dataset of video. The proposed scheme is then incorporated on the coding unit basis into the rate-distortion optimization for depth coding optimization, coupled with adapting quantization parameter accordingly to accommodate local effect of the depth errors on the VSD. Experimental results show that our solution obtains better results in depth coding than three testing solutions, on the platform of H.264/AVC reference software JM16.0. Benefiting from the efficiency of the VSD estimation, low coding complexity is obtained as well. The proposed solution is further evaluated on the reference software HTM13.0 of the latest 3-D high-efficiency video coding standard, exhibiting better and comparable results compared against the HTM codec with the view synthesis optimization disabled and enabled, respectively. Meng Yang 0002, Ce Zhu, Xuguang Lan, Nanning Zheng 0001 |
IEEE Trans. Multim. | 1 |
| 2017 | Fast additive quantization for vector compression in nearest neighbor search
Jin Li 0011, Xuguang Lan, Jiang Wang 0001, Meng Yang 0002, Nanning Zheng 0001 |
Multim. Tools Appl. | 4 |
| 2017 | A new compressive sensing video coding framework based on Gaussian mixture model
Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001 |
Signal Process. Image Commun. | 3 |
| 2017 | A Novel Method of Minimizing View Synthesis Distortion Based on Its Non-Monotonicity in 3D VideoabstractIn depth-based 3D video, the view synthesis distortion (VSD), is generally measured by modeling the effect of texture and depth errors separately. With such a development, it has been referred that the VSD changes monotonically with respect to to both the texture and depth distortions. In this paper, we find that the VSD does not always change monotonically with them by both theoretical analysis and experimental test, when the effect of the texture and depth errors is considered together. Specifically, first, we prove that the VSD is non-monotonic with the texture distortion. That is, the VSD increases with the increasing texture distortion at higher distortion range but conversely decreases with it at lower range. It is different from the general scenario that only considering the effect of the texture errors. We also analytically depict their relationship with low computational cost and identify the turning point at which the change of the VSD is converted. Second, we confirm that the VSD is always monotonic with the depth distortion, which is consistent with the general scenario that only considering the effect of the depth errors. The non-monotonicity property of the VSD can be utilized to improve the viewing performance of 3D video in relevant applications, since a minimal value of the VSD exists at the turning point. We conduct two applications for this purpose. First, it is used to generate the synthesis view of minimal distortion, which achieves 0.51-dB gain of PSNR on average for the tested scenarios. Second, it is used for lossy compression of texture videos in 3D video, which reduces the coding rate by 24% on average for the tested scenarios, meanwhile, keeps the VSD not increased simultaneously. Meng Yang 0002, Nanning Zheng 0001, Ce Zhu, Fei Wang 0008 |
IEEE Trans. Image Process. | 1 |
| 2016 | Depth Map Coding by Modeling the Locality and Local Correlation of View Synthesis Distortion in 3-D Video
Qiong Xue, Xuguang Lan, Meng Yang 0002 |
MMM (1) | 3 |
| 2016 | Efficient detail-enhanced exposure correction based on auto-fusion for LDR imageabstractWe consider the problem of how to simultaneously and well correct the over- and under-exposure regions in a single low dynamic range (LDR) image. Recent methods typically focus on global visual quality but cannot well-correct much potential details in extremely wrong exposure areas, and some are also time consuming. In this paper, we propose a fast and detail-enhanced correction method based on automatic fusion which combines a pair of complementarily corrected images, i.e. backlight & highlight correction images (BCI &HCI). A BCI with higher visual quality in details is quickly produced based on a proposed faster multi-scale retinex algorithm; meanwhile, a HCI is generated through contrast enhancement method. Then, an automatic fusion algorithm is proposed to create a color-protected exposure mask for fusing BCI and HCI when avoiding potential artifacts on the boundary. The experiment results show that the proposed method can fast correct over/under-exposed regions with higher detail quality than existing methods. Xuguang Lan, Meng Yang 0002 |
MMSP | 3 |
| 2016 | Efficient compressive sensing video compression method based on Gaussian mixture modelsabstractIn this paper, we propose an efficient lossy compression method for the compressive sensing video that utilizes Gaussian mixture models (GMM). The GMM is used to model the compressive sensing video (CSV) frames. Then we design an efficient lossy compression method based on the GMM. Each CSV frame can be efficiently compressed by the proposed method. The proposed method is for better compromise of compression efficiency and computational complexity. And it achieves a significant Bjontegaard-Delta (BD)-PSNR improvement about 8.84~11.81dB in average compared with existing low complexity compression solutions for compressing the CSV sequence. Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001 |
VCIP | 3 |
| 2015 | Parameter-free view synthesis distortion model with application to depth video codingabstractDepth coding in 3D video is unique in that its quality is measured by the view synthesis distortion (VSD) rather than its own depth distortion, which further complicates the coding optimization as VSD is related to both the depth and texture quality. We propose a parameter-free VSD model to directly estimate the impact of depth errors on VSD on the small block basis, given its depth distortion. The proposed model is incorporated into rate-distortion optimization of depth coding, by adapting the Lagrange multiplier and optimizing the selection of quantization parameters. Simulation results show that our proposed scheme improves the BDPSNR and BDBR by 0.6dB PSNR and 20% bits saving on average compared with H.264 coding standard in depth coding, while keeping the implementation easy. Meng Yang 0002, Ce Zhu, Xuguang Lan, Nanning Zheng 0001 |
ISCAS | 1 |
| 2015 | Optimized truncation model for adaptive compressive sensing acquisition of imagesabstractThe sparsity of the input signal is important for compressive sensing (CS) reconstruction in CS system. In this paper, we establish an optimized truncation model to determine the number of the sparsified coefficients to be truncated in CS acquisition according to the sampling rate. The proposed truncation model suits for signals of any dimension. With the truncation model, the sparsity of the signal can be optimized by properly truncating the small elements of the sparsified coefficients. Furthermore we propose an adaptive CS acquisition solution based on the truncation model to reduce the noise folding effect. The proposed solution is verified for CS acquisition of natural images. Simulation results show that the proposed solution achieves significant improvement of the reconstructed image quality by 0.7~1.4 dB on average compared with existing solutions. Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001 |
VCIP | 3 |
| 2014 | Iterative Pricing-Based Rate Allocation for Video Streams With Fluctuating Bandwidth AvailabilityabstractWe consider rate allocation for video users in the case where the available bandwidth fluctuates. Simply minimizing the objective distortion or optimizing the stability of video qualities does not optimize subjective quality. We formulate a utility-based solution, considering that a user's preference of video quality often varies over a range with upper and lower thresholds of quality. Our iterative pricing-based resource allocation procedure reallocates the bandwidth not only between different users within a time slot but also between different time slots, such that no user suffers quality degradation on average by participating in the multiplexing process. Experimental results show that, compared with equal resource allocation and existing rate allocation solutions, the subjective result becomes increasingly better with the increase of bandwidth fluctuation rate or bandwidth fluctuation range. Moreover, as the number of users increases, the results improve. Meng Yang 0002, Theodore Groves, Nanning Zheng 0001, Pamela C. Cosman |
IEEE Trans. Multim. | 1 |
| 2013 | Universal and low-complexity quantizer design for compressive sensing image codingabstractCompressive sensing imaging (CSI) is a new framework for image coding, which enables acquiring and compressing a scene simultaneously. The CS encoder shifts the bulk of the system complexity to the decoder efficiently. Ideally, implementation of CSI provides lossless compression in image coding. In this paper, we consider the lossy compression of the CS measurements in CSI system. We design a universal quantizer for the CS measurements of any input image. The proposed method firstly establishes a universal probability model for the CS measurements in advance, without knowing any information of the input image. Then a fast quantizer is designed based on this established model. Simulation result demonstrates that the proposed method has nearly optimal rate-distortion (R~D) performance, meanwhile, maintains a very low computational complexity at the CS encoder. Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001 |
VCIP | 3 |
| 2013 | Adaptively post-encoding multiple description video coding
Xuguang Lan, Meng Yang 0002, Yuan Yuan 0001, Songlin Zhao, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2013 | Adaptive multiple description coding for hybrid networks with dynamic PLR and BER
Meng Yang 0002, Xuguang Lan, Nanning Zheng 0001 |
Signal Process. Image Commun. | 1 |
| 2012 | Depth-assisted error concealment for intra frame slices in 3D videoabstractWe propose a depth-assisted error concealment method for slice loss in intra frames of 2D+depth video sequence. Intra frames in the 2D view sequence are offset from intra frames in the depth sequence to guarantee the corresponding frame in the other sequence is not also intra mode. Then for a slice loss in an intra frame in the 2D view sequence, the motion information is extracted from the depth sequence to conceal the slice loss using boundary matching. Experimental results show that the proposed method provides improved performance over existing methods both for PSNR results and computational complexity at the decoder. Meng Yang 0002, Yuhong Yang 0004, Pamela C. Cosman |
ICIP | 1 |
| 2011 | Design of error-resilient M-description codec over wireless broadcasting networksabstractA practical error-resilient M-description codec scheme is designed to combat the bit errors of the wireless broadcasting networks and raise the quality of the reconstructed signal. The signal is coded into large number of mutually refinable descriptions by robust staggered M-description scalar quantizer (RSMDSQ). Then an index assignment method is used to enhance the error-resilient capacity of any subset of all generated descriptions by enlarging the hamming distance of the codewords. Accordingly, at the terminal, an enhanced decoding scheme is used to recover the signal utilizing the robust correlation among the descriptions. Simple simulations have been done on MATLAB. Meng Yang 0002, Xuguang Lan, Nanning Zheng 0001 |
CCNC | 1 |
| 2011 | Explicit Network-Adaptive Robust Multiple Description CodingabstractThe data delivery performance of multiple description coding (MDC) over unreliable network with capacity constraints is related with three factors: redundancy rate, packet loss rate (PLR), and bit error rate (BER). We simplified this network-adaptive delivery problem to only relate with redundancy rate. The proposed scheme is an extension of scalar quantization (SQ) based MDC. Meng Yang 0002, Xuguang Lan, Nanning Zheng 0001 |
DCC | 1 |
| 2011 | Consecutive redundancy control for robust multiple description coding over unreliable networksabstractDifferent system may have different request on coding efficiency and security, even the characteristics of the unreliable networks are fixed, such as bit error rate, packet loss rate and bandwidth. In this paper, a novel multiple description coding (MDC) scheme is presented for the unreliable networks, considering both the network characteristics and system request. An error-resilient problem based on MDSQ method is proposed and analyzed, and then an iterative redundancy-control method is designed based on ERMDC by index pair ordering. Accordingly, a fast IA method is proposed only considering the high redundancy case. Simulation results show that the redundancy can be easily and precisely adjusted to meet all the requests. Meanwhile, the error resilience and R-D bound of MDC is self-adaptively guaranteed well enough for all cases. Meng Yang 0002, Xuguang Lan, Nanning Zheng 0001 |
WCNC | 1 |