Meiqin Liu 0002

dblp:89/2278-2 · DBLP profile ↗
← Back
40ranked-venue papers
6as first author
32since 2021 · last 2026
0000-0001-8428-5098ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 2 first-author · 22 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Distortion Perception Enhanced Image Quality Assessment for Color Block Artifacts and Color Saturation Distortions
Wenya Yang, Meiqin Liu 0002, Yao Zhao 0001
QoMEX3
2026 TCSAF: Learnable RefBank for reference-based video super-resolution
Meiqin Liu 0002, Houguo Ji, Yao Zhao 0001
Expert Syst. Appl.1
2026 Point cloud accumulation via multi-dimensional pseudo label and progressive instance association
Chunyu Lin, Lang Nie, Meiqin Liu 0002, Yao Zhao 0001
J. Vis. Commun. Image Represent.5
2026 Perceiving degradation prompts in residual diffusion model for real image denoising
Meiqin Liu 0002, Xuan Long, Yao Zhao 0001
Pattern Recognit.1
2026 TransVFC: A transformable video feature compression framework for machines
Yuxiao Sun, Yao Zhao 0001, Meiqin Liu 0002, Huihui Bai 0001, Chunyu Lin, Weisi Lin
Pattern Recognit.3
2026 Customizable ROI-Based Deep Image Compression
abstract
Region of Interest (ROI)-based image compression optimizes bit allocation by prioritizing ROI for higher-quality reconstruction. However, as the users (including human clients and downstream machine tasks) become more diverse, ROI-based image compression needs to be customizable to support various preferences. For example, different users may define distinct ROI or require different quality trade-offs between ROI and non-ROI. Existing ROI-based image compression schemes predefine the ROI, making it unchangeable, and lack effective mechanisms to balance reconstruction quality between ROI and non-ROI. This work proposes a paradigm for customizable ROI-based deep image compression. First, we develop a Text-controlled Mask Acquisition (TMA) module, which allows users to easily customize their ROI for compression by just inputting the corresponding semantictext. It makes the encoder controlled bytext. Second, we design a Customizable Value Assign (CVA) mechanism, which masks the non-ROI with a changeable extent decided by users instead of a constant one to manage the reconstruction quality trade-off between ROI and non-ROI. Finally, we present a Latent Mask Attention (LMA) module, where the latent spatial prior of the mask and the latent Rate-Distortion Optimization (RDO) prior of the image are extracted and fused in the latent space, and further used to optimize the latent representation of the source image. Experimental results demonstrate that our proposed customizable ROI-based deep image compression paradigm effectively addresses the needs of customization for ROI definition and mask acquisition as well as the reconstruction quality trade-off management between the ROI and non-ROI. Additionally, even by using the uniform mask as input, our method still outperforms the anchor methods in image reconstruction and machine vision tasks (such as object detection and instance segmentation). Our source code will be available at: https://github.com/hccavgcyv/Customizable-ROI-Based-Deep-Image-Compression.
Fanxin Xia, Feng Ding 0007, Xinfeng Zhang 0001, Meiqin Liu 0002, Yao Zhao 0001, Weisi Lin, Lili Meng
IEEE Trans. Circuits Syst. Video Technol.5
2026 Symmetric Entropy-Constrained Video Coding for Machines
abstract
As video transmission increasingly serves machine vision systems (MVS) instead of human vision systems (HVS), video coding for machines (VCM) has become a critical research topic. Existing VCM methods often bind codecs to specific downstream models, requiring retraining or supervised data, thus limiting generalization in multi-task scenarios. Recently, unified VCM frameworks have employed visual backbones (VB) and visual foundation models (VFM) to support multiple video understanding tasks with a single codec. They mainly utilize VB/VFM to maintain semantic consistency or suppress non-semantic information, but seldom explore how to directly link video coding with understanding under VB/VFM guidance. Hence, we propose a Symmetric Entropy-Constrained Video Coding framework for Machines (SEC-VCM). It establishes a symmetric alignment between the video codec and VB, allowing the codec to leverage VB's representation capabilities to preserve semantics and discard MVS-irrelevant information. Specifically, a bi-directional entropy-constraint (BiEC) mechanism ensures symmetry between the process of video decoding and VB encoding by suppressing conditional entropy. This helps the codec to explicitly handle semantic information beneficial to MVS while squeezing useless information. Furthermore, a semantic-pixel dual-path fusion (SPDF) module injects pixel-level priors into the final reconstruction. Through semantic-pixel fusion, it suppresses artifacts harmful to MVS and improves machine-oriented reconstruction quality. Experimental results on classical video understanding tasks and MLLM-based tasks show state-of-the-art (SOTA) rate-task performance. It achieves significant bitrate savings over H.266/VVC reference software VTM on video instance segmentation (37.4%), video object segmentation (29.8%), object detection (46.2%), multiple object tracking (44.9%), and MLLM-based video grounding (97.6%). The code is at https://github.com/Ws-Syx/SEC-VCM.
Yuxiao Sun, Meiqin Liu 0002, Weisi Lin, Frédéric Dufaux, Yao Zhao 0001
IEEE Trans. Image Process.2
2025 DATA-VSR: Dynamic Trajectory Attention and Texture Adaptive Rooter for Video Super-Resolution
abstract
Video Super-Resolution (VSR) is essential for reconstructing high-definition sequences from correlated video frames. While Transformer-based VSR methods have improved reconstruction quality, they require substantial computational resources, limiting deployment on resource-constrained devices. To tackle this issue, we propose a novel framework named Dynamic Trajectory Attention and Texture Adaptive Rooter for Video Super-Resolution (DATA-VSR). There are two key innovations: the Temporal Redundancy-aware Alignment Network (TRAN) and the Spatial Redundancy-aware Refinement Network (SRRN). Specifically, features are aligned by focusing on dynamic temporal trajectories instead of static redundancies in TRAN, and then features are adaptively refined based on the texture complexity of different regions in SRRN. Additionally, the Dual-Domain Enhancement Block (DDEB) is incorporated to effectively capture global dependencies in the frequency domain and enhance the representation of local features in the spatial domain. The experimental results on standard VSR benchmarks show that DATA-VSR achieves competitive performance with fewer parameters, lower FLOPs, and a specific reduction of 17%.
Linfeng He, Meiqin Liu 0002, Yao Zhao 0001
ICASSP2
2025 ALIC: Adaptive Fusion Entropy Model for Learned Image Compression
abstract
Recently, learned image compression algorithms have achieved significant performance. The entropy model is crucial for improving the rate-distortion performance by estimating the probability distribution of latent representation. In this paper, we propose an adaptive fusion entropy model for learned image compression (ALIC). To explore the correlation between channel and global spatial features, an adaptive fusion entropy model (AFEM) is designed. AFEM first slices the latent representation along the channels and leverages the adaptive channel fusion context module (ACFC) to capture correlations between the decoded and current slices. Subsequently, AFEM uses the adaptive spatial fusion context module (ASFC) to further divide the current slice into encoding pixels and reference pixels, thus improving the accuracy of probability estimation. The attention map and modulation parameter are introduced in ACFC and ASFC to interact with channel and spatial features. In addition, the variable-rate residual transformer (VResFormer) is proposed to control dynamic bit-rate by selectively modulating the high-frequency component according to coefficient weight and bias. Experimental results indicate that our ALIC outperforms other learned image compression algorithms. Our ALIC saves 5.89% bit-rate compared with VVC (4:4:4) on Kodak dataset.
Lingxue Li, Meiqin Liu 0002, Yao Zhao 0001
ICASSP2
2025 Embedding Compression Distortion in Video Coding for Machines
abstract
Currently, video transmission serves not only the Human Visual System (HVS) for viewing but also machine perception for analysis. However, existing codecs are primarily optimized for pixel-domain and HVS-perception metrics rather than the needs of machine vision tasks. To address this issue, we propose a Compression Distortion Representation Embedding (CDRE) framework, which extracts machine-perception-related distortion representation and embeds it into downstream models, addressing the information lost during compression and improving task performance. Specifically, to better analyze the machine-perception-related distortion, we design a compression-sensitive extractor that identifies compression degradation in the feature domain. For efficient transmission, a lightweight distortion codec is introduced to compress the distortion information into a compact representation. Subsequently, the representation is progressively embedded into the downstream model, enabling it to be better informed about compression degradation and enhancing performance. Experiments across various codecs and downstream tasks demonstrate that our framework can effectively boost the rate-task performance of existing codecs with minimal overhead in terms of bitrate, execution time, and number of parameters. Our codes and supplementary materials are released in https://github.com/Ws-Syx/CDRE/.
Yuxiao Sun, Yao Zhao 0001, Meiqin Liu 0002, Weisi Lin
ICME3
2025 Incremental few-shot instance segmentation via feature enhancement and prototype calibration
Weixiang Gao, Caijuan Shi, Rui Wang 0128, Ao Cai, Changyu Duan, Meiqin Liu 0002
Comput. Vis. Image Underst.6
2025 Context perturbation: A Consistent alignment approach for Domain Adaptive Semantic Segmentation
Meiqin Liu 0002, Yao Zhao 0001, Wei Wang 0108, Yunchao Wei
Comput. Vis. Image Underst.1
2024 Semantic Lens: Instance-Centric Semantic Alignment for Video Super-resolution
abstract
As a critical clue of video super-resolution (VSR), inter-frame alignment significantly impacts overall performance. However, accurate pixel-level alignment is a challenging task due to the intricate motion interweaving in the video. In response to this issue, we introduce a novel paradigm for VSR named Semantic Lens, predicated on semantic priors drawn from degraded videos. Specifically, video is modeled as instances, events, and scenes via a Semantic Extractor. Those semantics assist the Pixel Enhancer in understanding the recovered contents and generating more realistic visual results. The distilled global semantics embody the scene information of each frame, while the instance-specific semantics assemble the spatial-temporal contexts related to each instance. Furthermore, we devise a Semantics-Powered Attention Cross-Embedding (SPACE) block to bridge the pixel-level features with semantic knowledge, composed of a Global Perspective Shifter (GPS) and an Instance-Specific Semantic Embedding Encoder (ISEE). Concretely, the GPS module generates pairs of affine transformation parameters for pixel-level feature modulation conditioned on global semantics. After that the ISEE module harnesses the attention mechanism to align the adjacent frames in the instance-centric semantic space. In addition, we incorporate a simple yet effective pre-alignment module to alleviate the difficulty of model training. Extensive experiments demonstrate the superiority of our model over existing state-of-the-art VSR methods.
Yao Zhao 0001, Meiqin Liu 0002
AAAI3
2024 Noisy-Residual Continuous Diffusion Models for Real Image Denoising
abstract
The generation paradigm of diffusion model (DM) inspires numerous works to approach the image denoising problem iteratively. However, DM-based image denoising methods typically require long serial sampling chains, resulting in substantial sampling time and computation. To address this issue, we propose a Noisy-Residual Continuous Diffusion Model (RCDM). It constructs a path between clean and noisy images by shifting their noisy residual during forward process, which significantly shortens diffusion distance. To approximate the path, Noisy Residual Tracer Network (NRTNet) is adopted to estimate the derivative of each point along the path. For further acceleration, clean images are iteratively sampled from noisy images in the reverse process, where the sampling intervals are learnable and skippable. Moreover, we devise a two-stage training strategy to minimize the curvature of the learned path. Experimental results demonstrate that the proposed method achieves superior performance with fewer sampling steps in real image denoising.
Xuan Long, Meiqin Liu 0002, Yao Zhao 0001
ICME2
2024 TLVC: Temporal Bit-rate Allocation for Learned Video Compression
abstract
Most of the existing neural video compression methods adopt the hybrid coding framework, which only focus on the motion and residual coding of adjacent frames and ignore the long-term inter-frame dependency and bit-rate allocation. To address the shortcoming, we propose a temporal bit-rate allocation strategy for learned video compression (TLVC). Specifically, Motion-driven Temporal Gate (MTG) is designed to yield temporal bit-rate coefficient by considering the impact of residual coding on subsequent motion estimation. Sequentially, Texture-conditioned Spatial Gate (TSG) is proposed to take the generated coefficient to guide the residual compression with different bit-rate. Experimental results demonstrate that TLVC can achieve effective bit-rate allocation compared with the traditional codec H.266/VVC (VTM-13.2) of low delay p-frame (LDP) configuration.
Meiqin Liu 0002, Yao Zhao 0001
ICME2
2024 SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-Resolution
abstract
Diffusion-based Video Super-Resolution (VSR) is renowned for generating perceptually realistic videos, yet it grapples with maintaining detail consistency across frames due to stochastic fluctuations. The traditional approach of pixel-level alignment is ineffective for diffusion-processed frames because of iterative disruptions. To overcome this, we introduce SeeClear--a novel VSR framework leveraging conditional video generation, orchestrated by instance-centric and channel-wise semantic controls. This framework integrates a Semantic Distiller and a Pixel Condenser, which synergize to extract and upscale semantic details from low-resolution frames. The Instance-Centric Alignment Module (InCAM) utilizes video-clip-wise tokens to dynamically relate pixels within and across frames, enhancing coherency. Additionally, the Channel-wise Texture Aggregation Memory (CaTeGory) infuses extrinsic knowledge, capitalizing on long-standing semantic textures. Our method also innovates the blurring diffusion process with the ResShift mechanism, finely balancing between sharpness and diffusion effects. Comprehensive experiments confirm our framework's advantage over state-of-the-art diffusion-based VSR techniques.
Yao Zhao 0001, Meiqin Liu 0002
NeurIPS3
2024 Camouflaged object segmentation with prior via two-stage training
Rui Wang 0128, Caijuan Shi, Changyu Duan, Weixiang Gao, Hongli Zhu, Yunchao Wei, Meiqin Liu 0002
Comput. Vis. Image Underst.7
2024 Multimodal spatiotemporal aggregation for point cloud accumulation
Chunyu Lin, Lang Nie, Meiqin Liu 0002, Yao Zhao 0001
J. Vis. Commun. Image Represent.4
2024 IBVC: Interpolation-driven B-frame video compression
Meiqin Liu 0002, Weisi Lin, Yao Zhao 0001
Pattern Recognit.2
2024 Implicit-Explicit Motion Learning for Video Camouflaged Object Detection
abstract
Video camouflaged object detection aims to identify objects that are visually concealed within the surroundings in a video. Most of the existing methods fall into analyzing the implicit inter-frame motion to capture the camouflaged object. However, due to a lack of exploring the prior explicit motion of the camouflaged object, these works generally encounter difficulty in capturing the complete camouflaged object. To address this issue, we propose to integrate implicit and explicit motion learning into a unified framework, namelyImplicit-Explicit Motion Learning network (IMEX), for video camouflaged object detection. Specifically, to promote the identifiability of the camouflaged object, a cross-scale representation fusion was proposed for global inter-frame alignment. By establishing cross-scale temporal-spatial association and aggregating the temporal-spatial attentive representations, it also achieves an elimination of the implicit motion of inter-frame to some extent. Moreover, to further improve the discriminability of boundary regions of the detected object, an explicit motion-induced consistency preserving of camouflaged objects is proposed, in which the prior boundary-aware explicit motion field is leveraged to supervise the consistency of camouflaged objects in consecutive frames. Extensive experiments show that our proposed IMEX achieves substantial performance improvements by a large margin.
Wenjun Hui, Zhenfeng Zhu, Guanghua Gu, Meiqin Liu 0002, Yao Zhao 0001
IEEE Trans. Multim.4
2023 3D-Aware Multi-Class Image-to-Image Translation with NeRFs
abstract
Recent advances in 3D-aware generative models (3D-aware GANs) combined with Neural Radiance Fields (NeRF) have achieved impressive results. However no prior works investigate 3D-aware GANs for 3D consistent multiclass image-to-image (3D-aware 121) translation. Naively using 2D-121 translation methods suffers from unrealistic shape/identity change. To perform 3D-aware multiclass 121 translation, we decouple this learning process into a multiclass 3D-aware GAN step and a 3D-aware 121 translation step. In the first step, we propose two novel techniques: a new conditional architecture and an effective training strategy. In the second step, based on the well-trained multiclass 3D-aware GAN architecture, that preserves view-consistency, we construct a 3D-aware 121 translation system. To further reduce the view-consistency problems, we propose several new techniques, including a U-net-like adaptor network design, a hierarchical representation constrain and a relative regularization loss. In exten-sive experiments on two datasets, quantitative and qualitative results demonstrate that we successfully perform 3D-aware 121 translation with multi-view consistency. Code is available in 3DI2I.
Senmao Li, Joost van de Weijer 0001, Yaxing Wang, Fahad Shahbaz Khan, Meiqin Liu 0002, Jian Yang 0003
CVPR5
2023 SIGVIC: Spatial Importance Guided Variable-Rate Image Compression
abstract
Variable-rate mechanism has improved the flexibility and efficiency of learning-based image compression that trains multiple models for different rate-distortion tradeoffs. One of the most common approaches for variable-rate is to channel- wisely or spatial-uniformly scale the internal features. However, the diversity of spatial importance is instructive for bit allocation of image compression. In this paper, we introduce a Spatial Importance Guided Variable-rate Image Compression (SigVIC), in which a spatial gating unit (SGU) is designed for adaptively learning a spatial importance mask. Then, a spatial scaling network (SSN) takes the spatial importance mask to guide the feature scaling and bit allocation for variablerate. Moreover, to improve the quality of decoded image, Top-K shallow features are selected to refine the decoded features through a shallow feature fusion module (SFFM). Experiments show that our method outperforms other learning- based methods (whether variable-rate or not) and traditional codecs, with storage saving and high flexibility.
Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001
ICASSP2
2023 Kernel Dimension Matters: To Activate Available Kernels for Real-time Video Super-Resolution
abstract
Real-time video super-resolution requires low latency with high-quality reconstruction. Existing methods mostly use pruning schemes or neglect complicated modules to reduce the calculation complexity. However, the video contains large amounts of temporal redundancies due to the inter-frame correlation, which is rarely investigated in existing methods. The static and dynamic information lies in feature maps and represents the redundant complements and temporal offsets respectively. It is crucial to split channels with dynamic and static information for efficient processing. Thus, this paper proposes a kernel-split strategy to activate available kernels for real-time inference. This strategy focuses on the dimensions of convolutional kernels, including the channel and depth dimensions. Available kernel dimensions are activated according to the split of high-value and low-value channels. Specifically, a multi-channel selection unit is designed to discriminate the importance of channels and filter the high-value channels hierarchically. At each hierarchy, low-dimensional convolutional kernels are activated to reuse the low-value channel and re-parameterized convolutional kernels are employed on the high-value channel to merge the depth dimension. In addition, we design a multiple flow deformable alignment module for a sufficient temporal representation with affordable calculation cost. Experimental results demonstrate that our method outperforms other state-of-the-art (SOTA) ones in terms of reconstruction quality and runtime. Codes will be available at https://github.com/Kimsure/KSNet.
Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001
ACM Multimedia2
2023 MSPNet: Multi-stage progressive network for image denoising
Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001
Neurocomputing2
2023 Temporal Consistency Learning of Inter-Frames for Video Super-Resolution
abstract
Video super-resolution (VSR) is a task that aims to reconstruct high-resolution (HR) frames from the low-resolution (LR) reference frame and multiple neighboring frames. The vital operation is to utilize the relative misaligned frames for the current frame reconstruction and preserve the consistency of the results. Existing methods generally explore information propagation and frame alignment to improve the performance of VSR. However, few studies focus on the temporal consistency of inter-frames. In this paper, we propose a Temporal Consistency learning Network (TCNet) for VSR in an end-to-end manner, to enhance the consistency of the reconstructed videos. A spatio-temporal stability module is designed to learn the self-alignment from inter-frames. Especially, the correlative matching is employed to exploit the spatial dependency from each frame to maintain structural stability. Moreover, a self-attention mechanism is utilized to learn the temporal correspondence to implement an adaptive warping operation for temporal consistency among multi-frames. Besides, a hybrid recurrent architecture is designed to leverage short-term and long-term information. We further present a progressive fusion module to perform a multistage fusion of spatio-temporal features. And the final reconstructed frames are refined by these fused features. Objective and subjective results of various experiments demonstrate that TCNet has superior performance on different benchmark datasets, compared to several state-of-the-art methods.
Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 JNMR: Joint Non-Linear Motion Regression for Video Frame Interpolation
abstract
Video frame interpolation (VFI) aims to generate predictive frames by motion-warping from bidirectional references. Most examples of VFI utilize spatiotemporal semantic information to realize motion estimation and interpolation. However, due to variable acceleration, irregular movement trajectories, and camera movement in real-world cases, they can not be sufficient to deal with non-linear middle frame estimation. In this paper, we present a reformulation of the VFI as a joint non-linear motion regression (JNMR) strategy to model the complicated inter-frame motions. Specifically, the motion trajectory between the target frame and multiple reference frames is regressed by a temporal concatenation of multi-stage quadratic models. Then, a comprehensive joint distribution is constructed to connect all temporal motions. Moreover, to reserve more contextual details for joint regression, the feature learning network is devised to explore clarified feature expressions with dense skip-connection. Later, a coarse-to-fine synthesis enhancement module is utilized to learn visual dynamics at different resolutions with multi-scale textures. The experimental VFI results show the effectiveness and significant improvement of joint motion regression over the state-of-the-art methods. The code is available at https://github.com/ruhig6/JNMR.
Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001
IEEE Trans. Image Process.1
2021 Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline
abstract
Depth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR) is a practical and valuable task, which up-scales the depth map into high-resolution (HR) space. However, limited by the lack of real-world paired low-resolution (LR) and HR depth maps, most existing methods use down-sampling to obtain paired training samples. To this end, we first construct a large-scale dataset named "RGB-D-D", which can greatly promote the study of depth map SR and even more depth-related real-world tasks. The "D-D" in our dataset represents the paired LR and HR depth maps captured from mobile phone and Lucid Helios respectively ranging from indoor scenes to challenging outdoor scenes. Besides, we provide a fast depth map super-resolution (FDSR) baseline, in which the high-frequency component adaptively decomposed from RGB image to guide the depth map SR. Extensive experiments on existing public datasets demonstrate the effectiveness and efficiency of our network compared with the state-of-the-art methods. Moreover, for the real-world LR depth maps, our algorithm can produce more accurate HR depth maps with clearer boundaries and to some extent correct the depth value errors.
Lingzhi He, Hongguang Zhu, Feng Li 0037, Huihui Bai 0001, Runmin Cong, Chunjie Zhang 0001, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001
CVPR8
2021 Depth Super-Resolution by Texture-Depth Transformer
abstract
Depth maps have been still suffering from some non-negligible effects, resulting from the consumer-level sensors. The limited resolution of the acquired depth maps is one of these annoying issues. Many prominent researchers have recently made a lot of efforts, such as traditional filters, as well as the deep learning paradigms. However, depth super-resolution is still an open challenge. In this paper, we design a texture-depth transformer for depth super-resolution task, which can learn the corresponding structural information of the high-resolution texture images and the corresponding interpolated depth maps. Moreover, a multi-scale feature fusion strategy is exploited to further enhance the fusion feature. Complementary to a quantitative evaluation, we demonstrate the effectiveness of the proposed approach.
Shuaiyong Zhang, Mengyao Yang, Meiqin Liu 0002, Junpeng Qi
ICME4
2021 ODE-Inspired Image Denoiser: An End-to-End Dynamical Denoising Network
Meiqin Liu 0002, Chunyu Lin, Yao Zhao 0001
PRCV (3)2
2021 An End-to-End Mutual Enhancement Network Toward Image Compression and Semantic Segmentation
Meiqin Liu 0002, Yao Zhao 0001
PRCV (2)3
2021 Image Outpainting with Depth Assistance
Lei Zhang 0116, Kang Liao, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001
PRCV (3)4
2021 Joint distortion rectification and super-resolution for self-driving scene perception
Keyao Zhao, Kang Liao, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001
Neurocomputing4
2020 IET Image Processing
abstract
360 video is very popular due to its 360 views of a scene. Although 360 videos are also compressed by a hybrid coding framework like 2D video, its high resolution and serious shape deformation affect coding efficiency. In equirectangular projection (ERP) format of 360 videos, if an object moves from equator regions to pole regions or vice versa, large deformation will be introduced and motion estimation cannot find the best‐matched part. To solve the above problem, the authors propose to generate a better reference frame for the current to be encoded frame. First, they project the frame prior to the current one from ERP to the sphere and rotate it at an appropriate angle depending on motion vectors. Subsequently, they insert this generated frame to the rear of the reference queue and let the encoder work as usual. The advantage is that the inserted frame has a more similar shape deformation as the current frame, which greatly helps motion estimation and makes full use of 360 video characters. Their method is simple and friendly compatible with the existing compression standard. Experiments prove that their method achieves 1.57% Bjøntegaard Delta (BD)‐gain compared with standard high efficiency video coding.
Chunyu Lin, Yao Zhao 0001, Meiqin Liu 0002, Xue Zhang 0008
IET Image Process.4
2020 A view-free image stitching network based on global homography
Lang Nie, Chunyu Lin, Kang Liao, Meiqin Liu 0002, Yao Zhao 0001
J. Vis. Commun. Image Represent.4
2020 Unsupervised fisheye image correction through bidirectional loss with geometric prior
Shangrong Yang, Chunyu Lin, Kang Liao, Yao Zhao 0001, Meiqin Liu 0002
J. Vis. Commun. Image Represent.5
2019 Improving Cube-to-ERP Conversion Performance with Geometry Features of 360 Video Structure
abstract
360 videos provide an omnidirectional view of the scene with extremely large data. Therefore, representing 360 videos with less data has become more and more important. Cube format is such a popular representation of 360 videos. However, we have to convert cube to Equirectangula(ERP) for displaying convenience. In this paper, we enhance Cube-to-ERP conversion performance by joint using Convolutional Neural Network(CNN) and classical interpolation method. The optimal threshold of boundary is derived according to geometry features of the cube-to-ERP format. This threshold is the guidance of how to combine CNN and classical interpolation method. Our experiment results prove that the derived threshold has a certain degree of guiding significance. Furthermore, we propose a new evaluation criterion with the help of Marsaglia model. It is much easier and more accurate to evaluate geometry conversion process.
Chunyu Lin, Huihui Bai 0001, Meiqin Liu 0002, Yao Zhao 0001
DCC4
2019 Block Partitioning Decision Based on Content Complexity for Future Video Coding
Yanhong Zhang, Yao Zhao 0001, Chunyu Lin, Meiqin Liu 0002
ICIG (3)4
2017 Depth map up-sampling with fractal dimension and texture-depth boundary consistencies
Meiqin Liu 0002, Yao Zhao 0001, Jie Liang 0001, Chunyu Lin, Huihui Bai 0001
Neurocomputing1
2014 Two-Stage Multiview Image Compression Using Interview SIFT Matching
abstract
In this paper, a novel scheme of two-stage multiview image compression is proposed to create two-level reconstructed quality. Differently from the conventional multiview image compression algorithms, SIFT (Scale-Invariant Feature Transform) features matching from interview images are exploited to remove the correlations between multiple views. In the first stage coding, SIFT and RANSAC (RANdom SAmple Consensus) algorithms are combined to calculate the correlation matrix of interview, which then can be developed to obtain the coarse reconstruction of the current view. In the second stage coding, the reconstructed quality can be improved further by using the residual information. The experimental results have shown that at higher compression ratio, the proposed scheme can obtain better rate-distortion performance than intra coding in MVC (Multiview Video Coding). Furthermore, with the change of the compression ratio, the proposed scheme can achieve more stable reconstructed quality.
Huihui Bai 0001, Mengmeng Zhang 0008, Meiqin Liu 0002, Anhong Wang, Yao Zhao 0001
DCC3
2012 Multiple Description Video Coding Using Macro Block Level Correlation of Inter-/Intra-Descriptions
abstract
Multiple description coding (MDC) is a promising technology for robust transmission over error-prone channels, which has attracted a lot research interests. The basic idea of MDC is to how to utilize redundant information of the descriptions for robust transmission. In view of practical applications, many MDC approaches have been proposed compatible with a certain standard codec, especially H.264/AVC. In this paper, we attempt to develop a novel MD video codec with generalized compatibility, which aims to the effective redundancy allocation from inter-/intra-descriptions. In [1], the redundancy allocation may be not enough effective due to frame level. As a result, in this paper, the redundant information will be taken into account at MB level.
Huihui Bai 0001, Mengmeng Zhang 0008, Meiqin Liu 0002, Anhong Wang, Yao Zhao 0001
DCC3