Lili Meng

dblp:45/8688 · DBLP profile ↗
← Back
42ranked-venue papers
9as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 7 since 2021Systems, architecture and hardware · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Customizable ROI-Based Deep Image Compression
abstract
Region of Interest (ROI)-based image compression optimizes bit allocation by prioritizing ROI for higher-quality reconstruction. However, as the users (including human clients and downstream machine tasks) become more diverse, ROI-based image compression needs to be customizable to support various preferences. For example, different users may define distinct ROI or require different quality trade-offs between ROI and non-ROI. Existing ROI-based image compression schemes predefine the ROI, making it unchangeable, and lack effective mechanisms to balance reconstruction quality between ROI and non-ROI. This work proposes a paradigm for customizable ROI-based deep image compression. First, we develop a Text-controlled Mask Acquisition (TMA) module, which allows users to easily customize their ROI for compression by just inputting the corresponding semantictext. It makes the encoder controlled bytext. Second, we design a Customizable Value Assign (CVA) mechanism, which masks the non-ROI with a changeable extent decided by users instead of a constant one to manage the reconstruction quality trade-off between ROI and non-ROI. Finally, we present a Latent Mask Attention (LMA) module, where the latent spatial prior of the mask and the latent Rate-Distortion Optimization (RDO) prior of the image are extracted and fused in the latent space, and further used to optimize the latent representation of the source image. Experimental results demonstrate that our proposed customizable ROI-based deep image compression paradigm effectively addresses the needs of customization for ROI definition and mask acquisition as well as the reconstruction quality trade-off management between the ROI and non-ROI. Additionally, even by using the uniform mask as input, our method still outperforms the anchor methods in image reconstruction and machine vision tasks (such as object detection and instance segmentation). Our source code will be available at: https://github.com/hccavgcyv/Customizable-ROI-Based-Deep-Image-Compression.
Fanxin Xia, Feng Ding 0007, Xinfeng Zhang 0001, Meiqin Liu 0002, Yao Zhao 0001, Weisi Lin, Lili Meng
IEEE Trans. Circuits Syst. Video Technol.8
2026 Boosting the No-Reference Image Quality Assessment via Low-Quality Pseudo References
abstract
No-reference image quality assessment (NR-IQA) aims to predict perceptual image quality without access to pristine references, which remains challenging due to diverse and complex distortions. Recent pseudo-reference-based methods attempt to mitigate this challenge but often rely on highfidelity pseudo-reference reconstruction. In contrast, this work shows that improving NR-IQA performance does not depend on reconstruction quality, but on effective representation learning, feature alignment, and deviation modeling between distorted images and pseudo references. To this end, we propose a novel NR-IQA framework that leverages low-quality pseudo references generated by a masked autoencoder with a lightweight decoder. Rather than pursuing detailed reconstruction, the pseudo reference is used to facilitate representation-level deviation modeling in a shared latent space via a cross-attention-based mechanism. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method consistently outperforms state-of-the-art NR-IQA approaches while maintaining modest computational complexity. Our source code will be available at: https://github.com/jianjin008/L-IQA.
Lili Meng, Yingnan Wang, Miaohui Wang, Guosheng Lin, Cheng Liang 0001, Jiande Sun 0001, Huaxiang Zhang 0001, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.1
2025 A New Weighted Nuclear Norm Regularization Model for Removing Salt and Pepper Noise With Applications
abstract
ABSTRACT The application of weighted kernel norm to image denoising has gained significant research interest in recent years by using the non‐local self‐similarity of images. In this paper, we propose a novel model for removing salt and pepper noise that integrates weighted kernel norm with higher‐order total variation regularization. Subsequently, we use the classical method of alternating direction of multipliers and introduce some auxiliary variables to transform the original problem into saddle point problem. To illustrate the analytical results, a series of numerical simulations are conducted. Finally, experimental comparisons demonstrate the superior performance of the proposed model, which outperforms other competitive methods in terms of both signal‐to‐noise ratio and structural similarity index.
Lili Meng, Zhiyi Lu, Huizhong Xue
IET Image Process.1
2025 MFCQA: Multi-Range Feature Cross-Attention Mechanism for no-reference image quality assessment
Nu Sun, Lili Meng, Weisi Lin, Li Liu 0031, Huaxiang Zhang 0001
Knowl. Based Syst.3
2024 ConR: Contrastive Regularizer for Deep Imbalanced Regression
abstract
Imbalanced distributions are ubiquitous in real-world data. They create constraints on Deep Neural Networks to represent the minority labels and avoid bias towards majority labels. The extensive body of imbalanced approaches address categorical label spaces but fail to effectively extend to regression problems where the label space is continuous. Local and global correlations among continuous labels provide valuable insights towards effectively modelling relationships in feature space. In this work, we propose ConR, a contrastive regularizer that models global and local label similarities in feature space and prevents the features of minority samples from being collapsed into their majority neighbours. ConR discerns the disagreements between the label space and feature space, and imposes a penalty on these disagreements. ConR minds the continuous nature of label space with two main strategies in a contrastive manner: incorrect proximities are penalized proportionate to the label similarities and the correct ones are encouraged to model local similarities. ConR consolidates essential considerations into a generic, easy-to-integrate, and efficient method that effectively addresses deep imbalanced regression. Moreover, ConR is orthogonal to existing approaches and smoothly extends to uni- and multi-dimensional label spaces. Our comprehensive experiments show that ConR significantly boosts the performance of all the state-of-the-art methods on four large-scale deep imbalanced regression benchmarks.
Mahsa Keramati, Lili Meng, R. David Evans
ICLR2
2024 AutoCast++: Enhancing World Event Prediction with Zero-shot Ranking-based Context Retrieval
abstract
Machine-based prediction of real-world events is garnering attention due to its potential for informed decision-making. Whereas traditional forecasting predominantly hinges on structured data like time-series, recent breakthroughs in language models enable predictions using unstructured text. In particular, (Zou et al., 2022) unveils AutoCast, a new benchmark that employs news articles for answering forecasting queries. Nevertheless, existing methods still trail behind human performance. The cornerstone of accurate forecasting, we argue, lies in identifying a concise, yet rich subset of news snippets from a vast corpus. With this motivation, we introduce AutoCast++, a zero-shot ranking-based context retrieval system, tailored to sift through expansive news document collections for event forecasting. Our approach first re-ranks articles based on zero-shot question-passage relevance, honing in on semantically pertinent news. Following this, the chosen articles are subjected to zero-shot summarization to attain succinct context. Leveraging a pre-trained language model, we conduct both the relevance evaluation and article summarization without needing domain-specific training. Notably, recent articles can sometimes be at odds with preceding ones due to new facts or unanticipated incidents, leading to fluctuating temporal dynamics. To tackle this, our re-ranking mechanism gives preference to more recent articles, and we further regularize the multi-passage representation learning to align with human forecaster responses made on different dates. Empirical results underscore marked improvements across multiple metrics, improving the performance for multiple-choice questions (MCQ) by 48% and true/false (TF) questions by up to 8%. Code is available at https://github.com/BorealisAI/Autocast-plus-plus.
Raihan Seraj, Lili Meng, Tristan Sylvain
ICLR4
2023 GAN-Based Image Compression with Improved RDO Process
Fanxin Xia, Lili Meng, Huaxiang Zhang 0001
ICIG (3)3
2023 Scaleformer: Iterative Multi-scale Refining Transformers for Time Series Forecasting
Mohammad Amin Shabani, Amir H. Abdi, Lili Meng, Tristan Sylvain
ICLR3
2023 Weight grouping operators selection strategy for a multiobjective evolutionary algorithm based on decomposition
Yanyan Tan, Zeyuan Yan, Lili Meng, Li Liu 0031
Appl. Intell.4
2023 Multi-dimensional constraints-based PPVO for high fidelity reversible data hiding
Wenxiu Liu, Lili Meng, Jiande Sun 0001, Wenbo Wan
Expert Syst. Appl.2
2023 Auto-Weighted Layer Representation Based View Synthesis Distortion Estimation for 3-D Video Coding
abstract
Recently, various view synthesis distortion estimation models have been studied to better serve 3-D video coding. However, they can hardly model the relationship quantitatively among different levels of depth changes, texture degeneration, and view synthesis distortion (VSD), which is crucial for rate-distortion optimization and rate allocation. In this paper, an auto-weighted layer representation based view synthesis distortion estimation model is developed. Firstly, sub-VSD (S-VSD) is defined according to the level of depth changes and their associated texture degeneration. After that, a set of theoretical derivations demonstrate that the VSD can be approximately decomposed into the S-VSDs multiplied by their associated weights. To obtain the S-VSDs efficiently, a layer-based representation method is developed, where all the pixels with the same level of depth changes are represented with a layer. It enables the S-VSD calculation at the layer level. Meanwhile, a nonlinear mapping function is learnt to accurately represent the relationship between the VSD and S-VSDs, automatically providing weights for the S-VSDs during VSD estimation. To learn such a function, a dataset of the VSD and its associated S-VSDs are built, termed as VSDSet. Experimental results show that the VSD can be accurately estimated with the weights learnt by the nonlinear mapping function once its associated S-VSDs are available. The proposed method outperforms the relevant state-of-the-art methods in both accuracy and efficiency. The VSDSet and source code of the proposed method will be available athttps://github.com/jianjin008/.
Xingxing Zhang 0001, Lili Meng, Weisi Lin, Jie Liang 0001, Huaxiang Zhang 0001, Yao Zhao 0001
IEEE Trans. Multim.3
2022 Energy Efficiency Optimization for RIS Assisted RSMA System over Estimated Channel
Caina Gao, Jia Zhang 0028, Linlin Guo, Lili Meng, Jiande Sun 0001
WASA (1)4
2022 Group-pair deep feature learning for multi-view 3d model retrieval
Xiuxiu Chen, Li Liu 0031, Huaxiang Zhang 0001, Lili Meng, Dongmei Liu 0007
Appl. Intell.5
2022 Coordinated rate splitting and power allocation in energy-spectral efficiency tradeoff-based multicell networks
Caina Gao, Jia Zhang 0028, Linlin Guo, Lili Meng, Jiande Sun 0001
Comput. Networks5
2022 An operator pre-selection strategy for multiobjective evolutionary algorithm based on decomposition
Zeyuan Yan, Yanyan Tan, Hongling Chen, Lili Meng, Huaxiang Zhang 0001
Inf. Sci.4
2022 Multiple description coding network based on semantic segmentation
Xue Li 0001, Lili Meng, Yanyan Tan, Jia Zhang 0028, Wenbo Wan, Huaxiang Zhang 0001
Multim. Tools Appl.2
2021 JND-aware robust image watermarking with tri-directional inter-block correlation
abstract
A novel block-level perceptual image watermarking framework is proposed in this study, including tri-directional correlation and a block-level just noticeable difference (JND) model. Specifically, the difference in the discrete cosine transform (DCT) coefficients of two blocks is calculated based on three directions in the neighborhood, called the tri-directional correlation (TriDC). Additionally, the representative alternating current (AC) coefficients along horizontal, vertical, and diagonal directions, which can describe structural patterns, are projected and merged for TriDC differences. Then, the difference of the DCT coefficient is modulated to a predefined zone depending on the JND-based offset. Finally, the extent of the watermarked AC coefficients is determined with perceptual JND adjustment. The experimental results demonstrate that the proposed scheme can protect most common image processing attacks; and has better robustness compared with recent zone modulation watermarking schemes and traditional watermarking methods.
Yunming Zhang, Zhenhua Wang 0004, Yantong Zhan, Lili Meng, Jiande Sun 0001, Wenbo Wan
Int. J. Intell. Syst.4
2021 Leader recommend operators selection strategy for a multiobjective evolutionary algorithm based on decomposition
Zeyuan Yan, Yanyan Tan, Wei Zheng 0004, Lili Meng, Huaxiang Zhang 0001
Inf. Sci.4
2021 Image compression based on octave convolution and semantic segmentation
Lili Meng, Yanyan Tan, Jia Zhang 0028, Huaxiang Zhang 0001
Knowl. Based Syst.2
2021 Deep semantic segmentation-based multiple description coding
Xue Li 0001, Lili Meng, Yanyan Tan, Jia Zhang 0028, Wenbo Wan, Huaxiang Zhang 0001
Multim. Tools Appl.2
2020 G-NOMA for Energy Efficient C-RAN
abstract
As a green wireless network access framework, cloud radio access network (C-RAN) has the advantages of reducing energy consumption and improving spectral efficiency. In this paper, we propose to maximize the energy efficiency (EE) in the downlink of the C-RAN. We propose a design of beamforming and rate allocation based on generalized nonorthogonal multiple access (G-NOMA) by using the central cooperation and local coordination across the base transceiver stations (BTSs) in the C-RAN downlink. Our aim is to maximize the EE with the constraints of each BTS power consumption and the minimum quality of service (QoS) requirements, which we propose to solve by an efficient and low complexity successive convex approximation (SCA) algorithm. Simulation results demonstrate that the G-NOMA scheme can effectively achieve the best energy efficiency.
Jia Zhang 0028, Lili Meng, Jiande Sun 0001
INDIN3
2020 Layout2image: Image Generation from Layout
Bo Zhao 0032, Weidong Yin, Lili Meng, Leonid Sigal
Int. J. Comput. Vis.3
2020 Pattern complexity-based JND estimation for quantization watermarking
Wenbo Wan, Jun Wang 0061, Jing Li 0046, Lili Meng, Jiande Sun 0001, Huaxiang Zhang 0001
Pattern Recognit. Lett.4
2020 Pixel-Level View Synthesis Distortion Estimation for 3D Video Coding
abstract
Recently, region-based 3D video coding has been proposed. However, existing view synthesis distortion estimation (VSDE) methods are performed at the frame level. To guide the rate-distortion optimization process of region-based 3D video coding schemes, this paper proposes the first pixel-level VSDE (PL-VSDE) method. We first give the definition of the pixel-level view synthesis distortion. To estimate it, a backward prediction method is then developed, which starts from the pixels of interest (POIs) in the virtual view and finds their corresponding pixels in the reference view via a coarse-to-fine approach, denoted as coarse-to-fine backward prediction (CFBP) method. Additionally, the CFBP fully considers the details of 3D warping, the rounding operation and the warping competition in view synthesis, leading to improve accuracy of the prediction. Besides, a table-lookup method and a warping property are introduced to speed up the CFBP. After integrating the CFBP into the PL-VSDE, we can estimate the view synthesis distortion at the pixel level. Our method is carried out pixel-by-pixel independently, which is friendly for parallel processing. The experimental results demonstrate that our proposed method has significant advantages in both accuracy and efficiency compared with the state-of-the-art frame-level VSDE methods.
Jie Liang 0001, Yao Zhao 0001, Chunyu Lin, Lili Meng
IEEE Trans. Circuits Syst. Video Technol.6
2019 Image Generation From Layout
abstract
Despite significant recent progress on generative models, controlled generation of images depicting multiple and complex object layouts is still a difficult problem. Among the core challenges are the diversity of appearance a given object may possess and, as a result, exponential set of images consistent with a specified layout. To address these challenges, we propose a novel approach for layout-based image generation; we call it Layout2Im. Given the coarse spatial layout (bounding boxes + object categories), our model can generate a set of realistic images which have the correct objects in the desired locations. The representation of each object is disentangled into a specified/certain part (category) and an unspecified/uncertain part (appearance). The category is encoded using a word embedding and the appearance is distilled into a low-dimensional vector sampled from a normal distribution. Individual object representations are composed together using convolutional LSTM, to obtain an encoding of the complete layout, and then decoded to an image. Several loss terms are introduced to encourage accurate and diverse generation. The proposed Layout2Im model significantly outperforms the previous state of the art, boosting the best reported inception score by 24.66% and 28.57% on the very challenging COCO-Stuff and Visual Genome datasets, respectively. Extensive experiments also demonstrate our method's ability to generate complex and diverse images with multiple objects.
Lili Meng, Weidong Yin, Leonid Sigal
CVPR2
2018 Reversible Architectures for Arbitrarily Deep Residual Neural Networks
abstract
Recently, deep residual networks have been successfully applied in many computer vision and natural language processing tasks, pushing the state-of-the-art performance with deeper and wider architectures. In this work, we interpret deep residual networks as ordinary differential equations (ODEs), which have long been studied in mathematics and physics with rich theoretical and empirical success. From this interpretation, we develop a theoretical framework on stability and reversibility of deep neural networks, and derive three reversible neural network architectures that can go arbitrarily deep in theory. The reversibility property allows a memory-efficient implementation, which does not need to store the activations for most hidden layers. Together with the stability of our architectures, this enables training deeper networks using only modest computational resources. We provide both theoretical analyses and empirical results. Experimental results demonstrate the efficacy of our architectures against several strong baselines on CIFAR-10, CIFAR-100 and STL-10 with superior or on-par state-of-the-art performance. Furthermore, we show our architectures yield superior results when trained using fewer training data.
Bo Chang 0002, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert, Elliot Holtham
AAAI2
2018 Multi-level Residual Networks from Dynamical Systems View
Bo Chang 0002, Lili Meng, Eldad Haber, Frederick Tung, David Begert
ICLR (Poster)2
2018 Learning Motion Predictors for Smart Wheelchair Using Autoregressive Sparse Gaussian Process
abstract
Constructing a smart wheelchair on a commercially available powered wheelchair (PWC) platform avoids a host of seating, mechanical design and reliability issues but requires methods of predicting and controlling the motion of a device never intended for robotics. Analog joystick inputs are subject to black-box transformations which may produce intuitive and adaptable motion control for human operators, but complicate robotic control approaches; furthermore, installation of standard axle mounted odometers on a commercial PWC is difficult. In this work, we present an integrated hardware and software system for predicting the motion of a commercial PWC platform that does not require any physical or electronic modification of the chair beyond plugging into an industry standard auxiliary input port. This system uses an RGB-D camera and an Arduino interface board to capture motion data, including visual odometry and joystick signals, via ROS communication. Future motion is predicted using an autoregressive sparse Gaussian process model. We evaluate the proposed system on real-world short-term path prediction experiments. Experimental results demonstrate the system's efficacy when compared to a baseline neural network model.
Zicong Fan, Lili Meng, Tian Qi Chen, Jingchun Li, Ian M. Mitchell
ICRA2
2018 Exploiting Points and Lines in Regression Forests for RGB-D Camera Relocalization
abstract
Camera relocalization plays a vital role in many robotics and computer vision applications, such as self-driving cars and virtual reality. Recent random forests based methods exploit randomly sampled pixel comparison features to predict 3D world locations for 2D image locations to guide the camera pose optimization. However, these point features are only sampled randomly in images, without considering geometric information such as lines, leading to large errors with the existence of poorly textured areas or in motion blur. Line segments are more robust in these environments. In this work, we propose to jointly exploit points and lines within the framework of uncertainty driven regression forests. The proposed approach is thoroughly evaluated on three publicly available datasets against several strong state-of-the-art baselines in terms of several different error metrics. Experimental results prove the efficacy of our method, showing superior or on-par state-of-the-art performance.
Lili Meng, Frederick Tung, James J. Little, Julien Valentin, Clarence W. de Silva
IROS1
2018 Generating Handwritten Chinese Characters Using CycleGAN
abstract
Handwriting of Chinese has long been an important skill in East Asia. However, automatic generation of handwritten Chinese characters poses a great challenge due to the large number of characters. Various machine learning techniques have been used to recognize Chinese characters, but few works have studied the handwritten Chinese character generation problem, especially with unpaired training data. In this work, we formulate the Chinese handwritten character generation as a problem that learns a mapping from an existing printed font to a personalized handwritten style. We further propose DenseNet CycleGAN to generate Chinese handwritten characters. Our method is applied not only to commonly used Chinese characters but also to calligraphy work with aesthetic values. Furthermore, we propose content accuracy and style discrepancy as the evaluation metrics to assess the quality of the handwritten characters generated. We then use our proposed metrics to evaluate the generated characters from CASIA dataset as well as our newly introduced Lanting calligraphy dataset.
Bo Chang 0002, Shenyi Pan, Lili Meng
WACV4
2018 Camera Selection for Broadcasting Soccer Games
abstract
When broadcasting events such as soccer games, human operators constantly select the camera with the best viewpoint to cover the whole event. Modeling the prediction of which camera should be on air will assist automatic sports broadcasts and influence millions of viewers. In this paper, we propose a proof-of-concept method to automatically select cameras for broadcasting soccer games. First, a random forest based regressor smoothly predicts the visual importance of short video clips using deep convolutional features. Then, the predictions from multiple candidate cameras are regularized by a novel camera duration cumulative distribution function (CDF), naturally guiding the camera selection. We apply our approach to real soccer broadcasts with a professional human operator's result as a reference. The quantitative experiments demonstrate that our method outperforms two alternatives in terms of prediction accuracy. Moreover, the video generated by our method is preferred in the user study experiment, exhibiting its practicality.
Lili Meng, James J. Little
WACV2
2018 An improved MOEA/D design for many-objective optimization problems
Wei Zheng 0004, Yanyan Tan, Lili Meng, Huaxiang Zhang 0001
Appl. Intell.3
2018 Semi-supervised modality-dependent cross-media retrieval
Jiande Sun 0001, Peiyong Duan, Lili Meng, Yanyan Tan, Wenbo Wan, Hongchen Wu, Bin Zhang 0050, Huaxiang Zhang 0001
Multim. Tools Appl.4
2018 Joint graph regularization based modality-dependent cross-media retrieval
Jihong Yan, Huaxiang Zhang 0001, Jiande Sun 0001, Qiang Wang 0015, Peilian Guo, Lili Meng, Wenbo Wan
Multim. Tools Appl.6
2018 Adaptive reconstruction based multiple description coding with randomly offset quantizations
Jingxiu Zong, Lili Meng, Yanyan Tan, Jia Zhang 0028, Huaxiang Zhang 0001
Multim. Tools Appl.2
2017 Backtracking regression forests for accurate camera relocalization
abstract
Camera relocalization plays a vital role in many robotics and computer vision tasks, such as global localization, recovery from tracking failure, and loop closure detection. Recent random forests based methods directly predict 3D world locations for 2D image locations to guide the camera pose optimization. During training, each tree greedily splits the samples to minimize the spatial variance. However, these greedy splits often produce uneven sub-trees in training or incorrect 2D-3D correspondences in testing. To address these problems, we propose a sample-balanced objective to encourage equal numbers of samples in the left and right sub-trees, and a novel backtracking scheme to remedy the incorrect 2D-3D correspondence predictions. Furthermore, we extend the regression forests based methods to use local features in both training and testing stages for outdoor RGB-only applications. Experimental results on publicly available indoor and outdoor datasets demonstrate the efficacy of our approach, which shows superior or on-par accuracy with several state-of-the-art methods.
Lili Meng, Frederick Tung, James J. Little, Julien Valentin, Clarence W. de Silva
IROS1
2017 Autonomous mobile robot navigation in uneven and unstructured indoor environments
abstract
Robots are increasingly operating in indoor environments designed for and shared with people. However, robots working safely and autonomously in uneven and unstructured environments still face great challenges. Many modern indoor environments are designed with wheelchair accessibility in mind. This presents an opportunity for wheeled robots to navigate through sloped areas while avoiding staircases. In this paper, we present an integrated software and hardware system for autonomous mobile robot navigation in uneven and unstructured indoor environments. This modular and reusable software framework incorporates capabilities of perception and navigation. Our robot first builds a 3D OctoMap representation for the uneven environment with the 3D mapping using wheel odometry, 2D laser and RGB-D data. Then we project multilayer 2D occupancy maps from OctoMap to generate the the traversable map based on layer differences. The safe traversable map serves as the input for efficient autonomous navigation. Furthermore, we employ a variable step size Rapidly Exploring Random Trees that could adjust the step size automatically, eliminating tuning step sizes according to environments. We conduct extensive experiments in simulation and real-world, demonstrating the efficacy and efficiency of our system. (Supplemented video link: https://youtu.be/6XJWcsH1fk0).
Chaoqun Wang 0009, Lili Meng, Sizhen She, Ian M. Mitchell, Teng Li 0005, Frederick Tung, Weiwei Wan, Max Q.-H. Meng, Clarence W. de Silva
IROS2
2016 Exploiting Random RGB and Sparse Features for Camera Pose Estimation
Lili Meng, Frederick Tung, James J. Little, Clarence W. de Silva
BMVC1
2014 Multiple Description Coding With Randomly and Uniformly Offset Quantizers
abstract
In this paper, two multiple description coding schemes are developed, based on prediction-induced randomly offset quantizers and unequal-deadzone-induced near-uniformly offset quantizers, respectively. In both schemes, each description encodes one source subset with a small quantization stepsize, and other subsets are predictively coded with a large quantization stepsize. In the first method, due to predictive coding, the quantization bins that a coefficient belongs to in different descriptions are randomly overlapped. The optimal reconstruction is obtained by finding the intersection of all received bins. In the second method, joint dequantization is also used, but near-uniform offsets are created among different low-rate quantizers by quantizing the predictions and by employing unequal deadzones. By generalizing the recently developed random quantization theory, the closed-form expression of the expected distortion is obtained for the first method, and a lower bound is obtained for the second method. The schemes are then applied to lapped transform-based multiple description image coding. The closed-form expressions enable the optimization of the lapped transform. An iterative algorithm is also developed to facilitate the optimization. Theoretical analyzes and image coding results show that both schemes achieve better performance than other methods in this category.
Lili Meng, Jie Liang 0001, Upul Samarawickrama, Yao Zhao 0001, Huihui Bai 0001, André Kaup
IEEE Trans. Image Process.1
2013 M-channel multiple description coding based on uniformly offset quantizers with optimal deadzone
abstract
This paper proposes an improved source-splitting-based two-rate M-channel multiple description coding scheme, where the source is split into M subsets. In each description, one subset is coded at a high rate, and others are predictively coded at a low rate. Uniform offsets among low-rate quantizers of different descriptions are achieved by employing unequal deadzones and by quantizing the predictions. When several descriptions are received, the optimal reconstruction of each subset is achieved by finding the intersection of all received quantization bins. The closed-form expression of the expected distortion is obtained. The proposed scheme is applied to lapped transform-based multiple description image coding and achieves improved performance. The optimal deadzone selection and its impact are also given in this paper.
Lili Meng, Jie Liang 0001, Yao Zhao 0001, Huihui Bai 0001, Chunyu Lin, André Kaup
ICASSP1
2013 Multiple description coding with randomly offset quantizers
abstract
A multiple description coding scheme based on prediction-induced randomly offset quantizers is proposed, where each description encodes one source subset with a small quantization stepsize, and other subsets are predictively coded with a large quantization stepsize. Due to the prediction, the quantization bins that a coefficient belongs to in different descriptions are randomly overlapped with each others. The optimal reconstruction is obtained by finding the intersection of all received quantization bins. Using the recently developed random quantization theory, the closed-form expression of the expected distortion is obtained. The proposed scheme is then applied to lapped transform-based multiple-description image coding, and an iterative optimization scheme is developed to find the optimal lapped transform. Experimental results show that the proposed scheme achieves better performance than other methods in this category.
Lili Meng, Jie Liang 0001, Upul Samarawickrama, Yao Zhao 0001, Huihui Bai 0001, André Kaup
ISCAS1
2010 GOP-Flexible Distributed Multiview Video Coding with Adaptive Side Information
Lili Meng, Yao Zhao 0001, Jeng-Shyang Pan 0001, Huihui Bai 0001, Anhong Wang
ICCCI (3)1