EDBT 2026 Demo / reviewers in the wild / expert
Mu Li 0005
dblp:36/4526-5
· DBLP profile ↗
33ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0002-7327-3304ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 16 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DynaQuant: Dynamic Mixed-Precision Quantization for Learned Image CompressionabstractPrevailing quantization techniques in Learned Image Compression (LIC) typically employ a static, uniform bit-width across all layers, failing to adapt to the highly diverse data distributions and sensitivity characteristics inherent in LIC models. This leads to a suboptimal trade-off between performance and efficiency. In this paper, we introduce DynaQuant, a novel framework for dynamic mixed-precision quantization that operates on two complementary levels. First, we propose content-aware quantization, where learnable scaling and offset parameters dynamically adapt to the statistical variations of latent features. This fine-grained adaptation is trained end-to-end using a novel Distance-aware Gradient Modulator (DGM), which provides a more informative learning signal than the standard Straight-Through Estimator. Second, we introduce a data-driven, dynamic bit-width selector that learns to assign an optimal bit precision to each layer, dynamically reconfiguring the network's precision profile based on the input data. Our fully dynamic approach offers substantial flexibility in balancing rate-distortion (R-D) performance and computational cost. Experiments demonstrate that DynaQuant achieves R-D performance comparable to full-precision models while significantly reducing computational and storage requirements, thereby enabling the practical deployment of advanced LIC on diverse hardware platforms. Youneng Bao, Yulong Cheng, Mu Li 0005, Yongsheng Liang 0001 |
AAAI | 6 |
| 2026 | MPR-net: Medicinal plant recognition network with dual-branch attention fusion
Zhanyan Tang, Yusen Fu, Mu Li 0005, Huiling Liang, Yibing Tang, Jie Wen 0001 |
Pattern Recognit. | 3 |
| 2026 | Gradient descent-driven sampling for multimodal long-term scanpath prediction in panoramic videos
Tianming Zhou, Yulong Cheng, Kanglong Fan, Youneng Bao, Mu Li 0005 |
Pattern Recognit. | 5 |
| 2026 | Pseudocylindrical Convolutions for Learned Omnidirectional Image CompressionabstractEquirectangular projection (ERP) is a convenient form to store omnidirectional images, but it is neither equal-area nor conformal, creating challenges for subsequent visual communication. When used for image compression, ERP amplifies sampling density and deforms objects near the poles, hindering perceptually optimal bit allocation. Here, we present one of the earliest endeavors to apply deep neural networks to omnidirectional image compression. We first propose parametric pseudocylindrical representations that generalize common pseudocylindrical map projections. A tractable greedy algorithm is introduced to identify (sub-)optimal representation configurations, guided by a proxy objective for rate-distortion performance. We then develop pseudocylindrical convolutions, which can be efficiently implemented by standard convolutions with “pseudocylindrical padding.” To demonstrate the utility of the proposed pseudocylindrical representations and convolutions, we implement an end-to-end omnidirectional image compression method, consisting of an analysis transform, a uniform quantizer, a synthesis transform, and an entropy model. Experiments show that our optimized method achieves consistently better rate-distortion performance compared to the state-of-the-art. Mu Li 0005, Kede Ma, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | DeCenter: Density-Center Guided Perception Enhancement for UAV Object DetectionabstractUnmanned aerial vehicle (UAV) object detection is essential for applications such as surveillance, agriculture, and disaster response. However, UAV imagery often contains small, dense, and occluded objects, posing challenges for existing methods. To address these challenges, we propose DeCenter, a novel Density-Center Guided Perception Enhancement framework for UAV object detection. DeCenter is composed of two key modules that jointly enhance the perception of small and crowded objects. First, the Density-Guided Object Center Heatmap Generator (DOCHG) adaptively generates Gaussian kernel-based heatmaps according to local density information, guiding the model to emphasize central neighborhoods of objects in crowded regions. This mechanism reduces overlaps between adjacent instances and alleviates missed detections under occlusion. Second, the Density-Center Feature Enhancement module (DCFE) integrates complementary cues from density features and object centers, adaptively balancing region-level object distribution with fine-grained localization. By fusing these signals, DCFE enhances the quality of feature representations, making them more discriminative for dense small objects while suppressing background noise. Experimental results on VisDrone and UAVDT datasets show that DeCenter achieves competitive overall accuracy with clear improvements in detecting dense small objects, offering an effective solution for UAV object detection. The code will be available at https://github.com/bluuzzz/decenter. Zhiqing Shi, Zhihao Wu 0002, Jie Wen 0001, Mu Li 0005, Xiaopeng Fan 0001, Yaowei Wang 0001, LinLin Shen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Adaptive Fine-Grained Fusion Network for Multimodal UAV Object DetectionabstractMultimodal perception and fusion play a vital role in uncrewed aerial vehicle (UAV) object detection. Existing methods typically adopt global fusion strategies across modalities. However, due to illumination variation, the effectiveness of RGB and infrared modalities may differ across local regions within the same image, particularly in UAV perspectives where occlusions and dense small objects are prevalent, leading to suboptimal performance of global fusion methods. To address this issue, we propose an adaptive fine-grained fusion network for multimodal UAV object detection. First, we design a local feature consistency-based modality fusion module, which adaptively assigns local fusion weights according to the structural consistency of high-response regions across modalities, thereby enabling more effective aggregation of object-relevant features. Second, we introduce a mutual information-guided feature contrastive loss to encourage the preservation of modality-specific information during the early training phase. Experimental results demonstrate that the proposed method effectively addresses the issue of object occlusion in UAV perspectives, achieving state-of-the-art performance on multimodal UAV object detection benchmarks. Code will be available at https://github.com/lingf5877/AFFNet. Zhanyan Tang, Zhihao Wu 0002, Mu Li 0005, Jie Wen 0001, Bob Zhang 0001, Yong Xu 0001, Jianqiang Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | DiffusionREC: Diffusion Model with Adaptive Condition for Referring Expression ComprehensionabstractThe objective of referring expression comprehension (REC) is to accurately identify the object in an image described by a given expression. Existing REC methods, including transformer-based and graph-based approaches among others, have shown robust performance in REC tasks. In this study, we present a groundbreaking framework named DiffusionREC for REC task. This framework reimagines the REC task as a text guided bounding box denoising diffusion process, through which noisy bounding boxes are refined and distilled to pinpoint the target box. Throughout the training process, the bounding box of the target object diffuses from its ground-truth position towards a random distribution. Simultaneously, a filtering-based object decoder is introduced to reverse this diffusion of noise, conditional on the provided expression, the result from previous denoised step and the interaction between the expression and the image. At the inference stage, we begin by randomly generating a collection of boxes. Subsequently, the filtering-based object decoder is iteratively employed to refine and prune these bounding boxes, taking into account the conditions on the given expression, the results from the previous denoised step, and the interaction between the expression and the image. Extensive experiments conducted on six datasets demonstrate that DiffusionREC outperforms previous REC methods, yielding superior performances. Jingcheng Ke, Wai Keung Wong, Jia Wang 0020, Mu Li 0005, Lunke Fei, Jie Wen 0001 |
AAAI | 4 |
| 2025 | Learned Image Compression with Dictionary-based Entropy ModelabstractLearned image compression methods have attracted great research interest and exhibited superior rate-distortion performance to the best classical image compression standards of the present. The entropy model plays a key role in learned image compression, which estimates the probability distribution of the latent representation for further entropy coding. Most existing methods employed hyper-prior and auto-regressive architectures to form their entropy models. However, they only aimed to explore the internal dependencies of latent representation while neglecting the importance of extracting prior from training data. In this work, we propose a novel entropy model named Dictionary-Based Cross Attention Entropy model, which introduces a learnable dictionary to summarize the typical structures occurring in the training dataset to enhance the entropy model. Extensive experimental results have demonstrated that the proposed model strikes a better balance between performance and latency, achieving state-of-the-art results on various benchmark datasets. Jingbo Lu, Leheng Zhang, Mu Li 0005, Wen Li 0001, Shuhang Gu |
CVPR | 4 |
| 2025 | Dataset Distillation as Data Compression: A Rate-Utility PerspectiveabstractDriven by the ``scale-is-everything'' paradigm, modern machine learning increasingly demands ever-larger datasets and models, yielding prohibitive computational and storage requirements. Dataset distillation mitigates this by compressing an original dataset into a small set of synthetic samples, while preserving its full utility. Yet, existing methods either maximize performance under fixed storage budgets or pursue suitable synthetic data representations for redundancy removal, without jointly optimizing both objectives. In this work, we propose a joint rate-utility optimization method for dataset distillation. We parameterize synthetic samples as optimizable latent codes decoded by extremely lightweight networks. We estimate the Shannon entropy of quantized latents as the rate measure and plug any existing distillation loss as the utility measure, trading them off via a Lagrange multiplier. To enable fair, cross-method comparisons, we introduce bits per class (bpc), a precise storage metric that accounts for sample, label, and decoder parameter costs. On CIFAR-10, CIFAR-100, and ImageNet-128, our method achieves up to $170\times$ greater compression than standard distillation at comparable accuracy. Across diverse bpc budgets, distillation losses, and backbone architectures, our approach consistently establishes better rate-utility trade-offs. Youneng Bao, Yongsheng Liang 0001, Mu Li 0005, Kede Ma |
ICCV | 5 |
| 2025 | Learning Compact Semantic Information for Incomplete Multi-View Missing Multi-Label ClassificationabstractMulti-view data involves various data forms, such as multi-feature, multi-sequence and multimodal data, providing rich semantic information for downstream tasks. The inherent challenge of incomplete multi-view missing multi-label learning lies in how to effectively utilize limited supervision and insufficient data to learn discriminative representation. Starting from the sufficiency of multi-view shared information for downstream tasks, we argue that the existing contrastive learning paradigms on missing multi-view data show limited consistency representation learning ability, leading to the bottleneck in extracting multi-view shared information. In response, we propose to minimize task-independent redundant information by pursuing the maximization of cross-view mutual information. Additionally, to alleviate the hindrance caused by missing labels, we develop a dual-branch soft pseudo-label cross-imputation strategy to improve classification performance. Extensive experiments on multiple benchmarks validate our advantages and demonstrate strong compatibility with both missing and complete data. Jie Wen 0001, Zhanyan Tang, Yuting He 0001, Mu Li 0005, Chengliang Liu 0003 |
ICML | 6 |
| 2025 | Dual structure-aware consensus graph learning for incomplete multi-view clustering
Lilei Sun, Wai Keung Wong, Yusen Fu, Jie Wen 0001, Mu Li 0005, Yuwu Lu, Lunke Fei |
Pattern Recognit. | 5 |
| 2025 | Stable successive Neural Image Compression via coherent demodulation-based transformation
Youneng Bao, Wen Tan 0001, Mu Li 0005, Fanyang Meng, Yongsheng Liang 0001 |
Signal Process. | 3 |
| 2025 | ShiftLIC: Lightweight Learned Image Compression With Spatial-Channel Shift OperationsabstractLearned Image Compression (LIC) has attracted considerable attention due to their outstanding rate-distortion (R-D) performance and flexibility. However, the substantial computational cost poses challenges for practical deployment. The issue of feature redundancy in LIC is rarely addressed. Our findings indicate that many features within the LIC backbone network exhibit similarities. This paper introduces ShiftLIC, a novel and efficient LIC framework that employs parameter-free shift operations to replace large-kernel convolutions, significantly reducing the model’s computational burden and parameter count. Specifically, we propose the Spatial Shift Block (SSB), which combines shift operations with small-kernel convolutions to replace large-kernel. This approach maintains feature extraction efficiency while reducing both computational complexity and model size. To further enhance the representation capability in the channel dimension, we propose a channel attention module based on recursive feature fusion. This module enhances feature interaction while minimizing computational overhead. Additionally, we introduce an improved entropy model integrated with the SSB module, making the entropy estimation process more lightweight and thereby comprehensively reducing computational costs. Experimental results demonstrate that ShiftLIC outperforms leading compression methods, such as VVC Intra and GMM, in terms of computational cost, parameter count, and decoding latency. Additionally, ShiftLIC sets a new SOTA benchmark with a BD-rate gain per MACs/pixel of −102.6%, showcasing its potential for practical deployment in resource-constrained environments. The code is released athttps://github.com/baoyu2020/ShiftLIC. Youneng Bao, Wen Tan 0001, Chuanmin Jia, Mu Li 0005, Yongsheng Liang 0001, Yonghong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | 3DFACENet: 3D Facial Attractiveness Computation and Enhancement NetworkabstractThe development of facial editing, virtual makeup, AR/VR technologies and 3D games applications underscore the need for advanced 3D facial attractiveness research. However, due to the lack of 3D beauty face data and the complexity of handling 3D face data, 3D facial aesthetics research remains largely unexplored. To fill this gap, we propose 3DFACENet, an innovative system designed for the computation and enhancement of 3D facial attractiveness. Our approach employs a 3D facial reconstruction encoder to generate encoded vectors from images and a render module to obtain 3D face models. To minimize computational load, we innovatively propose an attractiveness computation module which leverages 3D shape and texture coefficients rather than 3D mesh models to access facial attractiveness, achieving state-of-the-art results. To balance aesthetic enhancement and identity preservation, we design a controllable beautification decoder. For the first time, we introduce the concept of "attractive centers", demonstrating that an individual's distance to these centers is significantly negatively correlated with their beauty scores. Our beautification decoder edits 3D facial coefficients towards these centers, achieving a significant and controllable enhancement in facial attractiveness. Extensive experiments on the SCUT-FBP5500 and MEBeauty dataset validate the effectiveness and feasibility of 3DFACENet. Tianhao Peng 0003, Mu Li 0005, Baoyuan Wu, David Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Learned Scanpaths Aid Blind Panoramic Video Quality AssessmentabstractPanoramic videos have the advantage of providing an immersive and interactive viewing experience. Nevertheless, their spherical nature gives rise to various and uncertain user viewing behaviors, which poses significant challenges for panoramic video quality assessment (PVQA). In this work, we propose an end-to-end optimized, blind PVQA method with explicit modeling of user viewing patterns through visual scanpaths. Our method consists of two modules: a scanpath generator and a quality assessor. The scanpath generator is initially trained to predict future scanpaths by minimizing their expected code length and then jointly optimized with the quality assessor for quality prediction. Our blind PVQA method enables direct quality assessment of panoramic images by treating them as videos composed of identical frames. Experiments on three public panoramic image and video quality datasets, encompassing both synthetic and authentic distortions, validate the superiority of our blind PVQA model over existing methods. Kanglong Fan, Wen Wen 0007, Mu Li 0005, Yifan Peng 0001, Kede Ma |
CVPR | 3 |
| 2024 | Modular Blind Video Quality AssessmentabstractBlind video quality assessment (BVQA) plays a pivotal role in evaluating and improving the viewing experience of end-users across a wide range of video-based platforms and services. Contemporary deep learning-based models primarily analyze video content in its aggressively subsampled format, while being blind to the impact of the actual spatial resolution and frame rate on video quality. In this paper, we propose a modular BVQA model and a method of training it to improve its modularity. Our model comprises a base quality predictor, a spatial rectifier, and a temporal rectifier, responding to the visual content and distortion, spatial resolution, and frame rate changes on video quality, respectively. During training, spatial and temporal rectifiers are dropped out with some probabilities to render the base quality predictor a standalone BVQA model, which should work better with the rectifiers. Extensive experiments on both professionally-generated content and user-generated content video databases show that our quality model achieves superior or comparable performance to current methods. Additionally, the modularity of our model offers an opportunity to analyze existing video quality databases in terms of their spatial and temporal complexity. Wen Wen 0007, Mu Li 0005, Yabin Zhang 0002, Yiting Liao, Kede Ma |
CVPR | 2 |
| 2024 | ISFB-GAN: Interpretable semantic face beautification with generative adversarial network
Tianhao Peng 0003, Mu Li 0005, Fangmei Chen, Yong Xu 0001, Yahan Sun, David Zhang 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Deep dual incomplete multi-view multi-label classification via label semantic-guided contrastive learning
Jinrong Cui, Yazi Xie, Chengliang Liu 0003, Qiong Huang 0001, Mu Li 0005, Jie Wen 0001 |
Neural Networks | 5 |
| 2024 | Omnidirectional image super-resolution via position attention network
Xin Wang 0160, Shiqi Wang 0001, Jinxing Li 0003, Mu Li 0005, Yong Xu 0001 |
Neural Networks | 4 |
| 2024 | Perceptual Quality Assessment of Virtual Reality Videos in the WildabstractInvestigating how people perceive virtual reality (VR) videos in the wild (i.e., those captured by everyday users) is a crucial and challenging task in VR-related applications due to complexauthenticdistortionslocalizedinspaceandtime.Existingpanoramic video databases only consider synthetic distortions, assume fixed viewing conditions, and are limited in size. To overcome these shortcomings, we construct the VR Video Quality in the Wild (VRVQW) database, containing 502 user-generated videos with diverse content and distortion characteristics. Based on VRVQW, we conduct a formal psychophysical experiment to record the scanpaths and perceived quality scores from 139 participants under two different viewing conditions. We provide a thorough statistical analysis of the recordeddata, observing significantimpact of viewing conditions on both human scanpaths and perceived quality. Moreover, we develop an objective quality assessment model for VR videos based on pseudocylindrical representation and convolution. Results on the proposed VRVQW show that our method is superior to existing video quality assessment models.We have made the database and code available at https://github.com/ limuhit/VR-Video-Quality-in-the-Wild. Wen Wen 0007, Mu Li 0005, Yiru Yao, Xiangjie Sui, Yabin Zhang 0002, Long Lan, Yuming Fang 0001, Kede Ma |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Learning Content-Weighted Pseudocylindrical Representation for 360° Image CompressionabstractLearned 360° image compression methods using equirectangular projection (ERP) often confront a non-uniform sampling issue, inherent to sphere-to-rectangle projection. While uniformly or nearly uniformly sampling representations, along with their corresponding convolution operations, have been proposed to mitigate this issue, these methods often concentrate solely on uniform sampling rates, thus neglecting the content of the image. In this paper, we urge that different contents within 360° images have varying significance and advocate for the adoption of a content-adaptive parametric representation in 360° image compression, which takes into account both the content and sampling rate. We first introduce the parametric pseudocylindrical representation and corresponding convolution operation, upon which we build a learned 360° image codec. Then, we model the hyperparameter of the representation as the output of a network, derived from the image's content and its spherical coordinates. We treat the optimization of hyperparameters for different 360° images as distinct compression tasks and propose a meta-learning algorithm to jointly optimize the codec and the metaknowledge, i.e., the hyperparameter estimation network. A significant challenge is the lack of a direct derivative from the compression loss to the hyperparameter network. To address this, we present a novel method to relax the rate-distortion loss as a function of the hyperparameters, enabling gradient-based optimization of the metaknowledge. Experimental results on omnidirectional images demonstrate that our method achieves state-of-the-art performance and superior visual quality. Mu Li 0005, Youneng Bao, Xiaohang Sui, Jinxing Li 0003, Guangming Lu 0002, Yong Xu 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Learning efficient facial landmark model for human attractiveness analysis
Tianhao Peng 0003, Mu Li 0005, Fangmei Chen, Yong Xu 0001, David Zhang 0001 |
Pattern Recognit. | 2 |
| 2023 | From Global to Local: Multi-Patch and Multi-Scale Contrastive Similarity Learning for Unsupervised Defocus Blur DetectionabstractDefocus blur detection (DBD), which aims to detect out-of-focus or in-focus pixels from a single image, has been widely applied to many vision tasks. To remove the limitation on the abundant pixel-level manual annotations, unsupervised DBD has attracted much attention in recent years. In this paper, a novel deep network named Multi-patch and Multi-scale Contrastive Similarity (M2CS) learning is proposed for unsupervised DBD. Specifically, the predicted DBD mask from a generator is first exploited to re-generate two composite images by transporting the estimated clear and unclear areas from the source image to realistic full-clear and full-blurred images, respectively. To encourage these two composite images to be completely in-focus or out-of-focus, a global similarity discriminator is exploited to measure the similarity of each pair in a contrastive way, through which each two positive samples (two clear images or two blurred images) are enforced to be close while each two negative samples (a clear image and a blurred image) are inversely far. Since the global similarity discriminator only focuses on the blur-level of a whole image and there do exist some fail-detected pixels which only cover a small part of areas, a set of local similarity discriminators are further designed to measure the similarity of image patches in multiple scales. Thanks to this joint global and local strategy, as well as the contrastive similarity learning, the two composite images are more efficiently moved to be all-clear or all-blurred. Experimental results on real-world datasets substantiate the superiority of our proposed method both in quantification and visualization. The source code is released at: https://github.com/jerysaw/M2CS. Jinxing Li 0003, Beicheng Liang, Xiangwei Lu, Mu Li 0005, Guangming Lu 0002, Yong Xu 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Learning Context-Based Nonlocal Entropy Modeling for Image CompressionabstractThe entropy of the codes usually serves as the rate loss in the recent learned lossy image compression methods. Precise estimation of the probabilistic distribution of the codes plays a vital role in reducing the entropy and boosting the joint rate-distortion performance. However, existing deep learning based entropy models generally assume the latent codes are statistically independent or depend on some side information or local context, which fails to take the global similarity within the context into account and thus hinders the accurate entropy estimation. To address this issue, we propose a special nonlocal operation for context modeling by employing the global similarity within the context. Specifically, due to the constraint of context, nonlocal operation is incalculable in context modeling. We exploit the relationship between the code maps produced by deep neural networks and introduce the proxy similarity functions as a workaround. Then, we combine the local and the global context via a nonlocal attention block and employ it in masked convolutional networks for entropy modeling. Taking the consideration that the width of the transforms is essential in training low distortion models, we finally produce a U-net block in the transforms to increase the width with manageable memory consumption and time complexity. Experiments on Kodak and Tecnick datasets demonstrate the priority of the proposed context-based nonlocal attention block in entropy modeling and the U-net block in low distortion situations. On the whole, our model performs favorably against the existing image compression standards and recent deep image compression models. Mu Li 0005, Kai Zhang 0008, Jinxing Li 0003, Wangmeng Zuo, Radu Timofte, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | End-to-End Optimized 360° Image CompressionabstractThe 360° image that offers a 360-degree scenario of the world is widely used in virtual reality and has drawn increasing attention. In 360° image compression, the spherical image is first transformed into a planar image with a projection such as equirectangular projection (ERP) and then saved with the existing codecs. The ERP images that represent different circles of latitude with the same number of pixels suffer from the unbalance sampling problem, resulting in inefficiency using planar compression methods, especially for the deep neural network (DNN) based codecs. To tackle this problem, we introduce a latitude adaptive coding scheme for DNNs by allocating variant numbers of codes for different regions according to the latitude on the sphere. Specifically, taking both the number of allocated codes for each region and their entropy into consideration, we introduce a flexible regional adaptive rate loss for region-wise rate controlling. Latitude adaptive constraints are then introduced to prevent spending too many codes on the over-sampling regions. Furthermore, we introduce viewport-based distortion loss by calculating the average distortion on a set of viewports. We optimize and test our model on a large 360° dataset containing 19,790 images collected from the Internet. The experiment results demonstrate the superiority of the proposed latitude adaptive coding scheme. On the whole, our model outperforms the existing image compression standards, including JPEG, JPEG2000, HEVC Intra Coding, and VVC Intra Coding, and helps to save around 15% bits compared to the baseline learned image compression model for planar images. Mu Li 0005, Jinxing Li 0003, Shuhang Gu, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Learning Content-Weighted Deep Image CompressionabstractLearning-based lossy image compression usually involves the joint optimization of rate-distortion performance, and requires to cope with the spatial variation of image content and contextual dependence among learned codes. Traditional entropy models can spatially adapt the local bit rate based on the image content, but usually are limited in exploiting context in code space. On the other hand, most deep context models are computationally very expensive and cannot efficiently perform decoding over the symbols in parallel. In this paper, we present a content-weighted encoder-decoder model, where the channel-wise multi-valued quantization is deployed for the discretization of the encoder features, and an importance map subnet is introduced to generate the importance masks for spatially varying code pruning. Consequently, the summation of importance masks can serve as an upper bound of the length of bitstream. Furthermore, the quantized representations of the learned code and importance map are still spatially dependent, which can be losslessly compressed using arithmetic coding. To compress the codes effectively and efficiently, we propose an upper-triangular masked convolutional network (triuMCN) for large context modeling. Experiments show that the proposed method can produce visually much better results, and performs favorably against deep and traditional lossy image compression approaches. Mu Li 0005, Wangmeng Zuo, Shuhang Gu, Jane You, David Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Online Multi-view Subspace Learning with Mixed NoiseabstractMulti-view learning reveals the latent correlation between different input modalities and has achieved outstanding performances in many fields. Recent approaches aim to find a low-dimensional subspace to reconstruct each view, in which the gross residual or noise follows either Gaussian or Laplacian distribution. However, the noise distribution is often more complex in practical applications, and a deterministic distribution assumption is incapable of modeling it. Additionally, referring to time-changed data, e.g., videos, the noise is temporal smooth, preventing us from processing the data with the whole input, as have generally been done in many existing multi-view learning methods. To tackle these problems, a novel online multi-view subspace learning is proposed in this paper. Particularly, our proposed method not only estimates a transformation for each view to extract the correlation among various views, but also introduces a Mixture of Gausssians (MoG) model into the multi-view data, successfully exploiting numbers of Gaussian Distributions to adaptively fit a wider range of the complex noise. Furthermore, we further design a novel online Expectation Maximization (EM) algorithm, being capable of efficiently processing the dynamic data. Experimental results substantiate the effectiveness and superiority of our approach. Jinxing Li 0003, Hongwei Yong, Feng Wu 0001, Mu Li 0005 |
ACM Multimedia | 4 |
| 2020 | Similarity and diversity induced paired projection for cross-modal retrieval
Jinxing Li 0003, Mu Li 0005, Guangming Lu 0002, Bob Zhang 0001, Hongpeng Yin, David Zhang 0001 |
Inf. Sci. | 2 |
| 2020 | Efficient and Effective Context-Based Convolutional Entropy Modeling for Image CompressionabstractPrecise estimation of the probabilistic structure of natural images plays an essential role in image compression. Despite the recent remarkable success of end-to-end optimized image compression, the latent codes are usually assumed to be fully statistically factorized in order to simplify entropy modeling. However, this assumption generally does not hold true and may hinder compression performance. Here we present contextbased convolutional networks (CCNs) for efficient and effective entropy modeling. In particular, a 3D zigzag scanning order and a 3D code dividing technique are introduced to define proper coding contexts for parallel entropy decoding, both of which boil down to place translation-invariant binary masks on convolution filters of CCNs. We demonstrate the promise of CCNs for entropy modeling in both lossless and lossy image compression. For the former, we directly apply a CCN to the binarized representation of an image to compute the Bernoulli distribution of each code for entropy estimation. For the latter, the categorical distribution of each code is represented by a discretized mixture of Gaussian distributions, whose parameters are estimated by three CCNs. We then jointly optimize the CCNbased entropy model along with analysis and synthesis transforms for rate-distortion performance. Experiments on the Kodak and Tecnick datasets show that our methods powered by the proposed CCNs generally achieve comparable compression performance to the state-of-the-art while being much faster. Mu Li 0005, Kede Ma, Jane You, David Zhang 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 1 |
| 2018 | A Probabilistic Hierarchical Model for Multi-View and Multi-Feature ClassificationabstractSome recent works in classification show that the data obtained from various views with different sensors for an object contributes to achieving a remarkable performance. Actually, in many real-world applications, each view often contains multiple features, which means that this type of data has a hierarchical structure, while most of existing works do not take these features with multi-layer structure into consideration simultaneously. In this paper, a probabilistic hierarchical model is proposed to address this issue and applied for classification. In our model, a latent variable is first learned to fuse the multiple features obtained from a same view, sensor or modality. Particularly, mapping matrices corresponding to a certain view are estimated to project the latent variable from a shared space to the multiple observations. Since this method is designed for the supervised purpose, we assume that the latent variables associated with different views are influenced by their ground-truth label. In order to effectively solve the proposed method, the Expectation-Maximization (EM) algorithm is applied to estimate the parameters and latent variables. Experimental results on the extensive synthetic and two real-world datasets substantiate the effectiveness and superiority of our approach as compared with state-of-the-art. Jinxing Li 0003, Hongwei Yong, Bob Zhang 0001, Mu Li 0005, Lei Zhang 0006, David Zhang 0001 |
AAAI | 4 |
| 2018 | Learning Convolutional Networks for Content-Weighted Image CompressionabstractLossy image compression is generally formulated as a joint rate-distortion optimization problem to learn encoder, quantizer, and decoder. Due to the non-differentiable quantizer and discrete entropy estimation, it is very challenging to develop a convolutional network (CNN)-based image compression system. In this paper, motivated by that the local information content is spatially variant in an image, we suggest that: (i) the bit rate of the different parts of the image is adapted to local content, and (ii) the content-aware bit rate is allocated under the guidance of a content-weighted importance map. The sum of the importance map can thus serve as a continuous alternative of discrete entropy estimation to control compression rate. The binarizer is adopted to quantize the output of encoder and a proxy function is introduced for approximating binary operation in backward propagation to make it differentiable. The encoder, decoder, binarizer and importance map can be jointly optimized in an end-to-end manner. And a convolutional entropy encoder is further presented for lossless compression of importance map and binary codes. In low bit rate image compression, experiments show that our system significantly outperforms JPEG and JPEG 2000 by structural similarity (SSIM) index, and can produce the much better visual result with sharp edges, rich textures, and fewer artifacts. Mu Li 0005, Wangmeng Zuo, Shuhang Gu, Debin Zhao, David Zhang 0001 |
CVPR | 1 |
| 2018 | Shift-Net: Image Inpainting via Deep Feature Rearrangement
Zhaoyi Yan, Xiaoming Li 0002, Mu Li 0005, Wangmeng Zuo, Shiguang Shan |
ECCV (14) | 3 |
| 2017 | Joint distance and similarity measure learning based on triplet-based constraints
Mu Li 0005, Qilong Wang 0001, David Zhang 0001, Peihua Li, Wangmeng Zuo |
Inf. Sci. | 1 |