VLDB 2026 Research / reviewers in the wild / expert
Huihui Bai 0001
dblp:75/5070-1
· DBLP profile ↗
97ranked-venue papers
10as first author
48since 2021 · last 2026
0000-0002-3879-8957ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 8 first-author · 30 since 2021Artificial intelligence and machine learning · 29 · 1 first-author · 22 since 2021Databases, data management, data science and information retrieval · 13 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BlindDiff: empowering degradation modeling in diffusion models for blind image super-resolution
Feng Li 0037, Zichao Liang, Runmin Cong, Huihui Bai 0001, Yao Zhao 0001, Meng Wang 0001 |
Sci. China Inf. Sci. | 5 |
| 2026 | Learning a joint mutual-guidance enhancement network for degraded low-light color image and low-resolution depth map
Bintao Chen, Lijun Zhao 0002, Jinjing Zhang, Anhong Wang, Huihui Bai 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | You can mask more for extremely low-bitrate image compression
Feng Li 0037, Jiaxin Han, Runmin Cong, Yunchao Wei, Weisi Lin, Yao Zhao 0001, Huihui Bai 0001 |
Pattern Recognit. | 8 |
| 2026 | TransVFC: A transformable video feature compression framework for machines
Yuxiao Sun, Yao Zhao 0001, Meiqin Liu 0002, Huihui Bai 0001, Chunyu Lin, Weisi Lin |
Pattern Recognit. | 5 |
| 2026 | An interpretable depth map super-resolution method via unrolling dual-boundary consistency constrained optimization
Lijun Zhao 0002, Jinjing Zhang, Huihui Bai 0001, Anhong Wang |
Pattern Recognit. | 4 |
| 2026 | A survey on image compressive sensing: From classical theory to the latest explicable deep learning
Lijun Zhao 0002, Xinlu Wang, Jinjing Zhang, Huihui Bai 0001, Anhong Wang |
Pattern Recognit. | 5 |
| 2026 | LGVF: A Low-Bitrate Generative Video Compression Framework With Spatio-Temporal Diffusion ModelabstractLow-bitrate video compression remains a fundamental challenge in video coding. Recent advances in generative technologies show great potential for addressing this problem, yet existing generative models often suffer from uncontrollability caused by error accumulation and unpredictable semantic mutation. To tackle these issues, we propose a novel low-bitrate generative video compression framework based on a spatio-temporal diffusion model. A quality comparator is designed to select the optimal feature during generation, which effectively suppresses error accumulation. In addition, we design a Semantic Mutation Feature Extraction Network (SMFN) to accurately predict frame sequences, thereby effectively handling semantically mutated features. Extensive experiments demonstrate that our framework achieves superior performance on FVD and LPIPS metrics, significantly reducing bitrate while preserving both visual fidelity and semantic consistency. Mengmeng Zhang 0008, Hongyun Lu, Huihui Bai 0001, Hongyuan Jing, Zhi Liu 0008 |
IEEE Signal Process. Lett. | 6 |
| 2026 | DiffLLFace: Learning Alternate Illumination-Diffusion Adaptation for Low-Light Face Super-Resolution and BeyondabstractFacial image acquisition under constrained illumination and with limited-resolution imaging devices often results in coupled photometric and geometric degradations, manifesting as low-light and low-resolution (LLR) conditions. Prevailing research predominantly follows fragmented optimization paradigms that address low-light image enhancement (LLIE) and face super-resolution (FSR) as isolated tasks. This approach overlooks the compound nature of the degradations, thereby significantly limiting their applicability in practical scenarios. To bridge this gap, we present DiffLLFace, a unified framework that harnesses diffusive generative capabilities with illumination-aware trajectories to achieve robust FSR from LLR observations. The core of our method lies in its alternate illumination-diffusion adaptation, which operates throughout the generation process. This mechanism not only captures degradation patterns in both brightness and structure to harmonize latent representations but also dynamically calibrates the illumination prior with the generative knowledge inherent to diffusion models. As such, DiffLLFace attains precise control over conditional adaptation and illumination rectification. We further devise a simple yet effective non-parametric Fourier enhancement strategy, which provides structural appearance clues that work in concert with the alternate adaptation to ensure texture and color consistency. Extensive experiments demonstrate the superiority of DiffLLFace over existing methods and remarkable generalizability on complex natural scenes. Code is available at https://github.com/KaishengPang/DiffLLFace. Runmin Cong, Kaisheng Pang, Feng Li 0037, Hua Li 0012, Huihui Bai 0001, Sam Kwong, Wei Zhang 0021 |
IEEE Trans. Image Process. | 5 |
| 2026 | Harnessing Group-Oriented Consistency Constraints for Semi-Supervised Semantic Segmentation in CdZnTe Semiconductors
Man Liu 0003, Huihui Bai 0001, Anhong Wang, Yunchao Wei, Yao Zhao 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Bridging Component Learning With Degradation Modelling for Blind Image Super-ResolutionabstractConvolutional Neural Network (CNN)-based image super-resolution (SR) has exhibited impressive success on known degraded low-resolution (LR) images. However, this type of approach is hard to hold its performance in practical scenarios when the degradation process (i.e.blur and downsampling) is unknown. Despite existing blind SR methods proposed to solve this problem using blur kernel estimation, the perceptual quality and reconstruction accuracy are still unsatisfactory. In this paper, we analyze the degradation of a high-resolution (HR) image from image intrinsic components according to a degradation-based formulation model. We propose a components decomposition and co-optimization network (CDCN) for blind SR. Firstly, CDCN decomposes the input LR image into structure and detail components in feature space. Then, the mutual collaboration block (MCB) is presented to exploit the relationship between both two components. In this way, the detail component can provide informative features to enrich the structural context and the structure component can carry structural context for better detail revealing via a mutual complementary manner. After that, we present a degradation-driven learning strategy to jointly supervise the HR image detail and structure restoration process. Finally, a multi-scale fusion module followed by an upsampling layer is designed to fuse the structure and detail features and perform SR reconstruction. Empowered by such degradation-based components decomposition, collaboration, and mutual optimization, we can bridge the correlation between component learning and degradation modelling for blind SR, thereby producing SR results with more accurate textures. Extensive experiments on both synthetic SR datasets and real-world images show that the proposed method achieves the state-of-the-art performance compared to existing methods. Feng Li 0037, Huihui Bai 0001, Weisi Lin, Runmin Cong, Yao Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Attend and Enrich: Enhanced Visual Prompt for Zero-Shot LearningabstractZero-shot learning (ZSL) endeavors to transfer knowledge from the seen categories to recognize unseen categories, which mostly relies on the semantic-visual interactions between image and attribute tokens. Recently, the prompt learning has emerged in ZSL and demonstrated significant potential as it allows the zero-shot transfer of diverse visual concepts to downstream tasks. However, current methods explore the fixed adaptation of the learnable prompt on the seen domains, which make them over-emphasize the primary visual features observed during training, limiting their generalization capabilities to the unseen domains. In this work, we propose AENet, which endows semantic information into the visual prompt to distill semantic-enhanced prompt for visual representation enrichment, enabling effective knowledge transfer for ZSL. AENet comprises two key steps: 1) exploring the concept-harmonized tokens for the visual and attribute modalities, grounded on the modal-sharing token that represents consistent visual-semantic concepts; and 2) yielding the semantic-enhanced prompt via the visual residual refinement unit with attribute consistency supervision. It is further integrated with primary visual features to attend to semantic-related information for visual enhancement, thus strengthening transferable ability. Experimental results on three benchmarks show that our AENet outperforms existing state-of-the-art ZSL methods. Man Liu 0003, Huihui Bai 0001, Feng Li 0037, Chunjie Zhang 0001, Yunchao Wei, Tat-Seng Chua, Yao Zhao 0001 |
AAAI | 2 |
| 2025 | CNC: Cross-modal Normality Constraint for Unsupervised Multi-class Anomaly DetectionabstractExisting unsupervised distillation-based methods rely on the differences between encoded and decoded features to locate abnormal regions in test images. However, the decoder trained only on normal samples still reconstructs abnormal patch features well, degrading performance. This issue is particularly pronounced in unsupervised multi-class anomaly detection tasks. We attribute this behavior to ‘over-generalization’ (OG) of decoder: the significantly increasing diversity of patch patterns in multi-class training enhances the model generalization on normal patches, but also inadvertently broadens its generalization to abnormal patches. To mitigate ‘OG’, we propose a novel approach that leverages class-agnostic learnable prompts to capture common textual normality across various visual patterns, and then apply them to guide the decoded features towards a ‘normal’ textual representation, suppressing ‘over-generalization’ of the decoder on abnormal patterns. To further improve performance, we also introduce a gated mixture-of-experts module to specialize in handling diverse patch patterns and reduce mutual interference between them in multi-class training. Our method achieves competitive performance on the MVTec AD and VisA datasets, demonstrating its effectiveness. Xiaoyang Wang 0007, Huihui Bai 0001, Eng Gee Lim, Jimin Xiao |
AAAI | 3 |
| 2025 | EvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with EventsabstractContinuous space-time video super-resolution (C-STVSR) endeavors to upscale videos simultaneously at arbitrary spatial and temporal scales, which has recently garnered increasing interest. However, prevailing methods struggle to yield satisfactory videos at out-of-distribution spatial and temporal scales. On the other hand, event streams characterized by high temporal resolution and high dynamic range, exhibit compelling promise in vision tasks. This paper presents EvEnhancer, an innovative approach that marries the unique advantages of event streams to elevate effectiveness, efficiency, and generalizability for C-STVSR. Our approach hinges on two pivotal components: 1) Event-adapted synthesis capitalizes on the spatiotemporal correlations between frames and events to discern and learn long-term motion trajectories, enabling the adaptive interpolation and fusion of informative spatiotemporal features; 2) Local implicit video transformer integrates local implicit video neural function with cross-scale spatiotemporal attention to learn continuous video representations utilized to generate plausible videos at arbitrary resolutions and frame rates. Experiments show that EvEnhancer achieves superiority on synthetic and real-world datasets and preferable generalizability on out-of-distribution scales against state-of-the-art methods. Code is available at https://github.com/W-Shuoyan/EvEnhancer. Shuoyan Wei, Feng Li 0037, Shengeng Tang, Yao Zhao 0001, Huihui Bai 0001 |
CVPR | 5 |
| 2025 | DecAD: Decoupling Anomalies in Latent Space for Multi-Class Unsupervised Anomaly Detection
Xiaoyang Wang 0007, Huihui Bai 0001, Eng Gee Lim, Jimin Xiao |
ICCV | 3 |
| 2025 | Dual-Edge Consistency Constrained Unfolding Network for Depth Map Super-Resolution
Lijun Zhao 0002, Jinjing Zhang, Huihui Bai 0001, Anhong Wang |
ICIG (1) | 4 |
| 2025 | Once-for-All: Controllable Generative Image Compression with Dynamic Granularity AdaptationabstractAlthough recent generative image compression methods have demonstrated impressive potential in optimizing the rate-distortion-perception trade-off, they still face the critical challenge of flexible rate adaptation to diverse compression necessities and scenarios. To overcome this challenge, this paper proposes a $\textbf{Control}$lable $\textbf{G}$enerative $\textbf{I}$mage $\textbf{C}$ompression framework, $\textbf{Control-GIC}$, the first capable of fine-grained bitrate adaptation across a broad spectrum while ensuring high-fidelity and generality compression. We base Control-GIC on a VQGAN framework representing an image as a sequence of variable-length codes ($\textit{i.e.}$ VQ-indices), which can be losslessly compressed and exhibits a direct positive correlation with bitrates. Drawing inspiration from the classical coding principle, we correlate the information density of local image patches with their granular representations. Hence, we can flexibly determine a proper allocation of granularity for the patches to achieve dynamic adjustment for VQ-indices, resulting in desirable compression rates. We further develop a probabilistic conditional decoder capable of retrieving historic encoded multi-granularity representations according to transmitted codes, and then reconstruct hierarchical granular features in the formalization of conditional probability, enabling more informative aggregation to improve reconstruction realism. Our experiments show that Control-GIC allows highly flexible and controllable bitrate adaptation where the results demonstrate its superior performance over recent state-of-the-art methods. Feng Li 0037, Yuxi Liu 0020, Runmin Cong, Yao Zhao 0001, Huihui Bai 0001 |
ICLR | 6 |
| 2025 | Learning A Deep Second-Order Unfolding Model for Arbitrary-Scale Depth Map Super-Resolution
Lijun Zhao 0002, Jinjing Zhang, Huihui Bai 0001, Anhong Wang |
ICXR | 4 |
| 2025 | MM-Prompt: Multi-modality and Multi-granularity Prompts for Few-Shot SegmentationabstractDespite the effectiveness of Segment Anything Model (SAM) based methods in Few-Shot Segmentation (FSS) tasks, our closer examination of their prompt encoding mechanism reveals that these methods rely solely on visual information to generate a single type of prompt. Consequently, they suffer from semantic granularity representation bias and a loss of spatial information. To address these limitations, this paper introduces an innovative multi-modal prompt encoder, enabling SAM to leverage both annotated reference images and textual descriptions of class names as segmentation prompts. This approach generates text prompts, dense visual prompts, and sparse visual prompts, spanning multiple modalities and granularities. These prompts provide enhanced representations of the target class, capturing both abstract semantics and specific details, while ensuring granularity appropriateness. When our multi-modal prompt encoder is integrated with SAM's image encoder and mask decoder, the overall model is referred to as MM-Prompt. To validate its effectiveness, we conducted extensive empirical studies on the PASCAL-5^i and COCO-20^i datasets. The experimental results demonstrate that MM-Prompt achieves state-of-the-art performance in FSS tasks, highlighting its substantial potential and value in this domain. Runmin Cong, Jinpeng Chen 0003, Chen Zhang 0013, Feng Li 0037, Huihui Bai 0001, Sam Kwong |
ACM Multimedia | 6 |
| 2025 | Unifying Reconstruction and Density Estimation via Invertible Contraction Mapping in One-Class ClassificationabstractDue to the difficulty in collecting all unexpected abnormal patterns, One-Class Classification (OCC) has become the most popular approach to anomaly detection (AD). Reconstruction-based AD method relies on the discrepancy between inputs and reconstructed results to identify unobserved anomalies. However, recent methods trained only on normal samples may generalize to certain abnormal inputs, leading to well-reconstructed anomalies and degraded performance. To address this, we constrain reconstructions to remain on the normal manifold using a novel AD framework based on contraction mapping. This mapping guarantees that any input converges to a fixed point through iterations of this mapping. Based on this property, training the contraction mapping using only normal data ensures that its fixed point lies within the normal manifold. As a result, abnormal inputs are iteratively transformed toward the normal manifold, increasing the reconstruction error. In addition, the inherent invertibility of contraction mapping enables flow-based density estimation, where a prior distribution learned from the previous reconstruction is used to estimate the input likelihood for anomaly detection, further improving the performance. Using both mechanisms, we propose a bidirectional structure with forward reconstruction and backward density estimation. Extensive experiments on tabular data, natural image, and industrial image data demonstrate the effectiveness of our method. The code is available at URD. Tianhong Dai, Huihui Bai 0001, Yao Zhao 0001, Jimin Xiao |
NeurIPS | 3 |
| 2025 | Ultra-Low Bitrate Multimodal Generative Face Video Coding Framework
Zhi Liu 0008, Hongyun Lu, Huihui Bai 0001, Hongyuan Jing, Mengmeng Zhang 0008 |
PCS | 4 |
| 2025 | SRConvNet: A Transformer-Style ConvNet for Lightweight Image Super-Resolution
Feng Li 0037, Runmin Cong, Jingjing Wu 0001, Huihui Bai 0001, Meng Wang 0001, Yao Zhao 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | PSVMA+: Exploring Multi-Granularity Semantic-Visual Adaption for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) endeavors to identify the unseen categories using knowledge from the seen domain, necessitating the intrinsic interactions between the visual features and attribute semantic features. However, GZSL suffers from insufficient visual-semantic correspondences due to the attribute diversity and instance diversity. Attribute diversity refers to varying semantic granularity in attribute descriptions, ranging from low-level (specific, directly observable) to high-level (abstract, highly generic) characteristics. This diversity challenges the collection of adequate visual cues for attributes under a uni-granularity. Additionally, diverse visual instances corresponding to the same sharing attributes introduce semantic ambiguity, leading to vague visual patterns. To tackle these problems, we propose a multi-granularity progressive semantic-visual mutual adaption (PSVMA+) network, where sufficient visual elements across granularity levels can be gathered to remedy the granularity inconsistency. PSVMA+ explores semantic-visual interactions at different granularity levels, enabling awareness of multi-granularity in both visual and semantic elements. At each granularity level, the dual semantic-visual transformer module (DSVTM) recasts the sharing attributes into instance-centric attributes and aggregates the semantic-related visual regions, thereby learning unambiguous visual features to accommodate various instances. Given the diverse contributions of different granularities, PSVMA+ employs selective cross-granularity learning to leverage knowledge from reliable granularities and adaptively fuses multi-granularity features for comprehensive representations. Experimental results demonstrate that PSVMA+ consistently outperforms state-of-the-art methods. Man Liu 0003, Huihui Bai 0001, Feng Li 0037, Chunjie Zhang 0001, Yunchao Wei, Meng Wang 0001, Tat-Seng Chua, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Reference-Based Iterative Interaction With P2-Matching for Stereo Image Super-ResolutionabstractStereo Image Super-Resolution (SSR) holds great promise in improving the quality of stereo images by exploiting the complementary information between left and right views. Most SSR methods primarily focus on the inter-view correspondences in low-resolution (LR) space. The potential of referencing a high-quality SR image of one view benefits the SR for the other is often overlooked, while those with abundant textures contribute to accurate correspondences. Therefore, we propose Reference-based Iterative Interaction (RIISSR), which utilizes reference-based iterative pixel-wise and patch-wise matching, dubbed $P^{2}$ -Matching, to establish cross-view and cross-resolution correspondences for SSR. Specifically, we first design the information perception block (IPB) cascaded in parallel to extract hierarchical contextualized features for different views. Pixel-wise matching is embedded between two parallel IPBs to exploit cross-view interaction in LR space. Iterative patch-wise matching is then executed by utilizing the SR stereo pair as another mutual reference, capitalizing on the cross-scale patch recurrence property to learn high-resolution (HR) correspondences for SSR performance. Moreover, we introduce the supervised side-out modulator (SSOM) to re-weight local intra-view features and produce intermediate SR images, which seamlessly bridge two matching mechanisms. Experimental results demonstrate the superiority of RIISSR against existing state-of-the-art methods. Runmin Cong, Rongxin Liao, Feng Li 0037, Ronghui Sheng, Huihui Bai 0001, Renjie Wan, Sam Kwong, Wei Zhang 0021 |
IEEE Trans. Image Process. | 5 |
| 2025 | Prompt to Restore, Restore to Prompt: Cyclic Prompting for Universal Adverse Weather RemovalabstractUniversal adverse weather removal (UAWR) seeks to address various weather degradations within a unified framework. Recent methods are inspired by prompt learning using pre-trained vision-language models (e.g., CLIP), leveraging degradation-aware prompts to facilitate weather-free image restoration, yielding significant improvements. In this work, we propose CyclicPrompt, an innovative cyclic prompt approach designed to enhance the effectiveness, adaptability, and generalizability of UAWR. CyclicPrompt comprises two key components: 1) a composite context prompt that integrates weather-related information and context-aware representations into the network to guide restoration. This prompt differs from previous methods by marrying learnable input-conditional vectors with weather-specific knowledge, thereby improving adaptability across various degradations and 2) the erase-and-paste mechanism, after the initial guided restoration, substitutes weather-specific knowledge with constrained restoration priors, inducing high-quality weather-free concepts into the composite prompt to further fine-tune the restoration process. Therefore, we can form a cyclic "Prompt-Restore-Prompt" pipeline that adeptly harnesses weather-specific knowledge, textual contexts, and reliable textures. Extensive experiments on synthetic and real-world datasets validate the superior performance of CyclicPrompt. The code is available at: https://github.com/RongxinL/CyclicPrompt. Rongxin Liao, Feng Li 0037, Yanyan Wei, Zenglin Shi, Le Zhang 0001, Huihui Bai 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | Enhancing Light Field Salient Object Detection With Variance-Maximized Key Focal Slice SelectionabstractLight field saliency object detection (LF SOD) methods have made significant progress recently. Most of them explore abundant multi-modal information from the all-focus image and the focal stacks at all focal planes to enrich scene details and depth perception. However, in light-field images, the spatial and depth information varies slightly across different slices, raising redundancy within focal stacks. Besides, the noise can appear repeatedly in multiple images of the focal stacks, which brings interference. To address these issues, in this work, we propose VMKNet, an effective approach that leverages innovative variance-maximized key slice selection and interacts with the all-focus image, to improve LF SOD. Specifically, we measure consistency differences between the all-focus image and each focal slice in the salient region as saliency scores. Then, we randomly assemble sets of them, where each score corresponds to a certain slice. The one exhibiting the highest variance is singled out to determine key focal slices as they reveal the diversity of salient objects. Then, the bidirectional guidance module (BGM) is presented to learn attentive features of all-focus and selected key slices in a mutual guidance manner, thus producing enhanced and holistic features. With hierarchical BGMs, our model can progressively aggregate common salient semantics and meaningful contextual details, generating more discriminative representations. Moreover, we introduce the edge enhancement module in conjunction with BGM to improve the sharpness of saliency maps. Extensive experiments on common light field datasets demonstrate that our method, termed VMKNet, outperforms recent state-of-the-art LF, RGB-D, and RGB methods. Our code is available athttps://github.com/Han-jiaxin/VMKNet. Jiaxin Han, Feng Li 0037, Mengmeng Zhang 0008, Huihui Bai 0001, Jimin Xiao, Yao Zhao 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Joint Deep-Unfolding Optimization Learning for Depth Map Arbitrary-Scale Super-Resolution
Lijun Zhao 0002, Jinjing Zhang, Anhong Wang, Huihui Bai 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | MPA-Det: multi-path aggregation-based object detection framework for aerial visual computing
Yaolin Lei, Zifeng Qiu, Huihui Bai 0001, Wenzhi He |
Vis. Comput. | 5 |
| 2024 | Towards the Uncharted: Density-Descending Feature Perturbation for Semi-supervised Semantic SegmentationabstractSemi-supervised semantic segmentation allows model to mine effective supervision from unlabeled data to complement label-guided training. Recent research has primarily focused on consistency regularization techniques, exploring perturbation-invariant training at both the image and feature levels. In this work, we proposed a novel feature-level consistency learning framework named Density-Descending Feature Perturbation (DDFP). Inspired by the low-density separation assumption in semi-supervised learning, our key insight is that feature density can shed a light on the most promising direction for the segmentation classifier to explore, which is the regions with lower density. We propose to shift features with confident predictions towards lower-density regions by perturbation injection. The perturbed features are then super-vised by the predictions on the original features, thereby compelling the classifier to explore less dense regions to effectively regularize the decision boundary. Central to our method is the estimation of feature density. To this end, we introduce a lightweight density estimator based on normalizing flow, allowing for efficient capture of the feature density distribution in an online manner. By extracting gradients from the density estimator, we can determine the direction towards less dense regions for each feature. The proposed DDFP outperforms other designs on feature-level perturbations and shows state of the art performances on both Pascal VOC and Cityscapes dataset under various partition protocols. The project is available at https://github.com/Gavinwxy/DDFP. Xiaoyang Wang 0007, Huihui Bai 0001, Limin Yu, Yao Zhao 0001, Jimin Xiao |
CVPR | 2 |
| 2024 | Deep Video Compression Based on Key ObjectivesabstractCurrent deep learning-based video encoding frameworks primarily focus on uniform global encoding, lacking specific optimizations for popular applications in recent years, such as video chat and conferencing. These applications typically feature key objectives that are more important than the background. Addressing this, this paper proposes a Depth Video Compression Algorithm based on Key Objectives, aimed at enhancing the encoding quality of these key elements while conservatively using the overall bitrate. Tianyu Gong, Ye Zhong, Huihui Bai 0001 |
DCC | 3 |
| 2024 | Region-Adaptive Transform with Segmentation Prior for Image Compression
Yuxi Liu 0020, Wenhan Yang, Huihui Bai 0001, Yunchao Wei, Yao Zhao 0001 |
ECCV (46) | 3 |
| 2024 | Exploring Resolution Fields for Scalable Image Compression With Uncertainty GuidanceabstractRecently, there are significant advancements in learning-based image compression methods surpassing traditional coding standards. Most of them prioritize achieving the best rate-distortion performance for a particular compression rate, which limits their flexibility and adaptability in various applications with complex and varying constraints. In this work, we explore the potential of resolution fields in scalable image compression and propose the reciprocal pyramid network (RPN) that fulfills the need for more adaptable and versatile compression. Specifically, RPN first builds a compression pyramid and generates the resolution fields at different levels in a top-down manner. The key design lies in the cross-resolution context mining module between adjacent levels, which performs feature enriching and distillation to mine meaningful contextualized information and remove unnecessary redundancy, producing informative resolution fields as residual priors. The scalability is achieved by progressive bitstream reusing and resolution field incorporation varying at different levels. Furthermore, between adjacent compression levels, we explicitly quantify the aleatoric uncertainty from the bottom decoded representations and develop an uncertainty-guided loss to update the upper-level compression parameters, forming a reverse pyramid process that enforces the network to focus on the textured pixels with high variance for more reliable and accurate reconstruction. Combining resolution field exploration and uncertainty guidance in a pyramid manner, RPN can effectively achieve spatial and quality scalable image compression. Experiments show the superiority of RPN against existing classical and deep learning-based scalable codecs. Code will be available athttps://github.com/JGIroro/RPNSIC. Dongyi Zhang, Feng Li 0037, Man Liu 0003, Runmin Cong, Huihui Bai 0001, Meng Wang 0001, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Perception-Oriented UAV Image Dehazing Based on Super-Pixel Scene PriorabstractCurrent unmanned aerial vehicle (UAV) image defogging techniques often result in images with chromatic aberrations, color distortions, and increased noise due to flight variables and environmental factors, impacting downstream mission objectives. This article presents a novel UAV image dehazing framework designed to enhance perceptual tasks in foggy conditions. Moving beyond traditional pixel-level dehazing, our approach utilizes a super-pixel scene prior (SPSP) method, improving the UAV defogging process. By shifting the dehazing operation from RGB to Lab color space using SPSP, we minimize chromatic confusion and pinpoint reliable defogging areas, especially in the L channel. To address light inconsistency challenges during defogging, we introduce a new guided filtering algorithm that leverages simple linear iterative clustering (SLIC). This algorithm utilizes super-pixel clusters instead of large guided windows, preserving crucial information and boosting efficiency. SPSP also guides the SLIC algorithm in Lab space, facilitating faster defogging. Our framework incorporates a quantitative analysis of super-pixel segmentation and target detection, utilizing a feedback loop with the alternating direction multiplier method (ADMM) to optimize perception and defogging concurrently, thus enhancing UAV visual capabilities in fog. Our UAV image dehazing technique outperforms existing methods, as evidenced by quantitative and qualitative assessments, effectively eliminating haze from UAV images and significantly improving perceptual processing in foggy conditions. Zifeng Qiu, Tianyu Gong, Zichao Liang, Taoyi Chen, Runmin Cong, Huihui Bai 0001, Yao Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Progressive Complementary Knowledge Aggregation for CdZnTe Defect SegmentationabstractAutomatic quality inspection of industrial products is an indispensable part of modern manufacturing. Cadmium zinc telluride (CdZnTe) crystal is an important industrial raw material, but the special photosensitive properties of CdZnTe make it show different defect boundaries under different lighting angles, which poses challenges for quality inspection. In this article, we propose progressive complementary knowledge aggregation (PCKA) for CdZnTe defect segmentation, which is model-agnostic. First, the 12 images of CdZnTe crystal with different lighting angles are fed into the preliminary aggregation net to aggregate unique pixel-level clues. Second, we use a latent aggregation net to acquire the feature-level complementary clues under the guidance of the pixel-level clues within latent space. Such a learning paradigm is an effective solution for the special photosensitive properties of CdZnTe crystal. Extensive experiments on self-collected dataset demonstrate the effectiveness and efficiency of our PCKA compared with other solutions. Feng Li 0037, Man Liu 0003, Huihui Bai 0001, Yunchao Wei, Anhong Wang, Shijie Ma, Yao Zhao 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Part-Object Progressive Refinement Network for Zero-Shot LearningabstractZero-shot learning (ZSL) recognizes unseen images by sharing semantic knowledge transferred from seen images, encouraging the investigation of associations between semantic and visual information. Prior works have been devoted to the alignment of global visual features with semantic information, i.e., attribute vectors, or further mining the local part regions related to each attribute and then simply concatenating them for category decisions. Although effective, these works ignore intrinsic interactions between local parts and the whole object, which enables a more discriminative and representative knowledge transfer for ZSL. In this paper, we propose a Part-Object Progressive Refinement Network (POPRNet), where discriminative and transferable semantics are progressively refined by the cooperation between parts and the whole object. Specifically, POPRNet incorporates discriminative part semantics and object-centric semantics guided by semantic intensity to improve cross-domain transferability. To achieve part-object learning, a semantic-augment transformer (SaT) is proposed to model the part-object relation at the part-level via an encoder and at the object-level via a decoder, generating a comprehensive semantic representation to boost discriminability and transferability. By introducing the prototype updating module embedded with the prototype selection layers, the discriminative ability of the updated category prototype is enhanced to further improve the recognition performance of ZSL. Extensive experiments are conducted to demonstrate the superiority and competitiveness of our proposed POPRNet method on three public benchmark datasets. The code is available at https://github.com/ManLiuCoder/POPRNet. Man Liu 0003, Chunjie Zhang 0001, Huihui Bai 0001, Yao Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Enhanced Video Super-Resolution Network towards Compressed DataabstractVideo super-resolution (VSR) algorithms aim at recovering a temporally consistent high-resolution (HR) video from its corresponding low-resolution (LR) video sequence. Due to the limited bandwidth during video transmission, most available videos on the internet are compressed. Nevertheless, few existing algorithms consider the compression factor in practical applications. In this paper, we propose an enhanced VSR model towards compressed videos, termed as ECVSR, to simultaneously achieve compression artifacts reduction and SR reconstruction end-to-end. ECVSR contains a motion-excited temporal adaption network (METAN) and a multi-frame SR network (SRNet). The METAN takes decoded LR video frames as input and models inter-frame correlations via bidirectional deformable alignment and motion-excited temporal adaption, where temporal differences are calculated as motion prior to excite the motion-sensitive regions of temporal features. In SRNet, cascaded recurrent multi-scale blocks (RMSB) are employed to learn deep spatio-temporal representations from adapted multi-frame features. Then, we build a reconstruction module for spatio-temporal information integration and HR frame reconstruction, which is followed by a detail refinement module for texture and visual quality enhancement. Extensive experimental results on compressed videos demonstrate the superiority of our method for compressed VSR. Code will be available at https://github.com/lifengcs/ECVSR . Feng Li 0037, Huihui Bai 0001, Runmin Cong, Yao Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Progressive Semantic-Visual Mutual Adaption for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) identifies unseen categories by knowledge transferred from the seen domain, relying on the intrinsic interactions between visual and semantic information. Prior works mainly localize regions corresponding to the sharing attributes. When various visual appearances correspond to the same attribute, the sharing attributes inevitably introduce semantic ambiguity, hampering the exploration of accurate semantic-visual interactions. In this paper, we deploy the dual semantic-visual transformer module (DSVTM) to progressively model the correspondences between attribute prototypes and visual features, constituting a progressive semantic-visual mutual adaption (PSVMA) network for semantic disambiguation and knowledge transferability improvement. Specifically, DSVTM devises an instance-motivated semantic encoder that learns instance-centric prototypes to adapt to different images, enabling the recast of the unmatched semantic-visual pair into the matched one. Then, a semantic-motivated instance decoder strengthens accurate cross-domain interactions between the matched pair for semantic-related instance adaption, en-couraging the generation of unambiguous visual representations. Moreover, to mitigate the bias towards seen classes in GZSL, a debiasing loss is proposed to pursue response consistency between seen and unseen predictions. The PSVMA consistently yields superior performances against other state-of-the-art methods. Code will be available at: https://github.com/ManLiuCoder/PSVMA. Man Liu 0003, Feng Li 0037, Chunjie Zhang 0001, Yunchao Wei, Huihui Bai 0001, Yao Zhao 0001 |
CVPR | 5 |
| 2023 | Learning deep texture-structure decomposition for low-light image restoration and enhancement
Lijun Zhao 0002, Jinjing Zhang, Anhong Wang, Huihui Bai 0001 |
Neurocomputing | 5 |
| 2023 | Boundary-constrained interpretable image reconstruction network for deep compressive sensing
Lijun Zhao 0002, Xinlu Wang, Jinjing Zhang, Anhong Wang, Huihui Bai 0001 |
Knowl. Based Syst. | 5 |
| 2023 | Joint depth map super-resolution method via deep hybrid-cross guidance filter
Lijun Zhao 0002, Jinjing Zhang, Anhong Wang, Huihui Bai 0001 |
Pattern Recognit. | 6 |
| 2023 | Plausible Proxy Mining With Credibility for Unsupervised Person Re-IdentificationabstractOne effective way to address unsupervised person re-identification is to use a clustering-based contrastive learning approach. Existing state-of-the-art methods adopt clustering algorithms (e.g., DBSCAN) and camera ID information to divide all person images into several camera-aware proxies. Then, for each person image, the extracted feature representation is pulled closer to the centroids of its pseudo-positive proxies (the proxies that share the same pseudo-identity label with this image) and pushed away from the centroids of other pseudo-negative proxies (the proxies that share the different pseudo-identity label with this image). However, the quality of the proxy centroid is significantly affected by the proxy impurity issue and thus deteriorates the learned feature representations. On the premise that we cannot introduce superior supervision signals by thoroughly solving the proxy impurity issue, for a person image, identifying its plausible proxies: the pseudo-negative proxies which potentially include its wrongly-clustered instances (the instances with the same ground-truth identity with this image), and further fixing the resulted incorrect supervision signals become an urgent and challenging problem. This paper proposes a simple yet effective approach to address this problem. With a given image, our method can effectively locate its plausible proxies. Then we introduce credibility to measure how much we should treat the centroid of each mined plausible proxy as a positive supervision signal rather than entirely negative. Extensive experiments on three widely-used person re-ID datasets validate the effectiveness of our proposed approach. Codes will be available at:https://github.com/Dingyuan-Zheng/PPCL. Dingyuan Zheng, Jimin Xiao, Mingjie Sun, Huihui Bai 0001, Junhui Hou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Learning Detail-Structure Alternative Optimization for Blind Super-ResolutionabstractExisting convolutional neural networks (CNN) based image super-resolution (SR) methods have achieved impressive performance on bicubic kernel, which is not valid to handle unknown degradations in real-world applications. Recent blind SR methods suggest to reconstruct SR images relying on blur kernel estimation. However, their results still remain visible artifacts and detail distortion due to the estimation errors. To alleviate these problems, in this paper, we propose an effective and kernel-free network, namely DSSR, which enables recurrent detail-structure alternative optimization without blur kernel prior incorporation for blind SR. Specifically, in our DSSR, a detail-structure modulation module (DSMM) is built to exploit the interaction and collaboration of image details and structures. The DSMM consists of two components: a detail restoration unit (DRU) and a structure modulation unit (SMU). The former aims at regressing the intermediate HR detail reconstruction from LR structural contexts, and the latter performs structural contexts modulation conditioned on the learned detail maps at both HR and LR spaces. Besides, we use the output of DSMM as the hidden state and design our DSSR architecture from a recurrent convolutional neural network (RCNN) view. In this way, the network can alternatively optimize the image details and structural contexts, achieving co-optimization across time. Moreover, equipped with the recurrent connection, our DSSR allows low- and high-level feature representations complementary by observing previous HR details and contexts at every unrolling time. Extensive experiments on synthetic datasets and real-world images demonstrate that our method achieves the state-of-the-art against existing methods. Feng Li 0037, Huihui Bai 0001, Weisi Lin, Runmin Cong, Yao Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Fine-Grained Image Classification by Class and Image-Specific Decomposition With Multiple ViewsabstractFine-grained image classification attempts to accurately classify images that are similar to each other. Multiview information is often used to improve the classification accuracy. Although great progress has been made, fine-grained image classification methods still have two drawbacks. On the one hand, they often treat each image independently without considering image correlations within the same class along with the distinctive characters of each image. On the other hand, multiview correlations are often used during classifier training, leaving the correlations of different views unconsidered. To solve these two problems, in this paper, we propose a novel fine-grained image classification method by class and image-specific decomposition with multiviews (CISD-MV). For each view, we treat images of the same class jointly by decomposing the class and image-specific information. Since images of different classes are similar and correlated, we linearly model class correlations of images using decomposed low-rank parts. In addition, for each image, the representations of different views are correlated, and we use linear transformation to model view correlations. We jointly optimise for the class and image-specific components along with the class correlation and view correlation transformation matrixes. A testing image is assigned to the class that has the minimum summed reconstruction error. We conduct fine-grained image classification experiments on several public fine-grained image datasets. Experimental results and analysis show the effectiveness of the proposed method. Chunjie Zhang 0001, Huihui Bai 0001, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Adaptive feature fusion network based on boosted attention mechanism for single image dehazing
Feng Li 0037, Runmin Cong, Huihui Bai 0001, Yao Zhao 0001 |
Multim. Tools Appl. | 4 |
| 2022 | LMDC: Learning a multiple description codec for deep learning-based image compression
Lijun Zhao 0002, Jinjing Zhang, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001 |
Multim. Tools Appl. | 3 |
| 2022 | Multi-level augmented inpainting network using spatial similarity
Huihui Bai 0001, Yao Zhao 0001 |
Pattern Recognit. | 2 |
| 2022 | Cross-Part Learning for Fine-Grained Image ClassificationabstractRecent techniques have achieved remarkable improvements depended on mining subtle yet distinctive features for fine-grained visual classification (FGVC). While prior works directly combine discriminative features extracted from different parts, we argue that the potential interactions between different parts and their abilities to category predictions should be taken into consideration, which enables significant parts to contribute more to the decision of the sub-category. To this end, we present a Cross-Part Convolutional Neural Network (CP-CNN) in a weakly supervised manner to explore cross-learning among multi-regional features. Specifically, the context transformer is implemented to encourage joint feature learning across different parts under the guidance of a navigator. The part with the highest confidence is regarded as a navigator to deliver distinguishing characteristics to the others with lower confidence while the complementary information is retained. To locate discriminative but subtle parts precisely, a part proposal generator (PPG) is designed with the feature enhancement blocks, through which complex scale variations caused by the viewpoint diversity can be effectively alleviated. Extensive experiments on three benchmark datasets demonstrate that our proposed method consistently outperforms existing state-of-the-art methods. Man Liu 0003, Chunjie Zhang 0001, Huihui Bai 0001, Riquan Zhang, Yao Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and BaselineabstractDepth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR) is a practical and valuable task, which up-scales the depth map into high-resolution (HR) space. However, limited by the lack of real-world paired low-resolution (LR) and HR depth maps, most existing methods use down-sampling to obtain paired training samples. To this end, we first construct a large-scale dataset named "RGB-D-D", which can greatly promote the study of depth map SR and even more depth-related real-world tasks. The "D-D" in our dataset represents the paired LR and HR depth maps captured from mobile phone and Lucid Helios respectively ranging from indoor scenes to challenging outdoor scenes. Besides, we provide a fast depth map super-resolution (FDSR) baseline, in which the high-frequency component adaptively decomposed from RGB image to guide the depth map SR. Extensive experiments on existing public datasets demonstrate the effectiveness and efficiency of our network compared with the state-of-the-art methods. Moreover, for the real-world LR depth maps, our algorithm can produce more accurate HR depth maps with clearer boundaries and to some extent correct the depth value errors. Lingzhi He, Hongguang Zhu, Feng Li 0037, Huihui Bai 0001, Runmin Cong, Chunjie Zhang 0001, Chunyu Lin, Meiqin Liu 0002, Yao Zhao 0001 |
CVPR | 4 |
| 2021 | Multi-scale attention network for image inpainting
Huihui Bai 0001, Yao Zhao 0001 |
Comput. Vis. Image Underst. | 2 |
| 2020 | Deep Interleaved Network for Single Image Super-Resolution with Asymmetric Co-AttentionabstractRecently, Convolutional Neural Networks (CNN) based image super-resolution (SR) have shown significant success in the literature. However, these methods are implemented as single-path stream to enrich feature maps from the input for the final prediction, which fail to fully incorporate former low-level features into later high-level features. In this paper, to tackle this problem, we propose a deep interleaved network (DIN) to learn how information at different states should be combined for image SR where shallow information guides deep representative features prediction. Our DIN follows a multi-branch pattern allowing multiple interconnected branches to interleave and fuse at different states. Besides, the asymmetric co-attention (AsyCA) is proposed and attacked to the interleaved nodes to adaptively emphasize informative features from different states and improve the discriminative ability of networks. Extensive experiments demonstrate the superiority of our proposed DIN in comparison with the state-of-the-art SR methods. Feng Li 0037, Runming Cong, Huihui Bai 0001 |
IJCAI | 3 |
| 2020 | Face inpainting network for large missing regions based on weighted facial similarity
Huihui Bai 0001, Yao Zhao 0001 |
Neurocomputing | 2 |
| 2020 | FilterNet: Adaptive Information Filtering Network for Accurate and Fast Image Super-ResolutionabstractDeep convolutional neural network (CNN) approaches have achieved impressive performance for image super-resolution (SR). The main issue of image SR is to effectively recover the high-frequency detail of low-resolution (LR) input. However, existing CNN methods often inevitably exhibit a large amount of memory consumption and computational cost. In addition, in most SR networks, the low-frequency and high-frequency components of the LR features are treated equally in the training process, which can ignore the local detailed information and hinder the representational capacity of networks. To solve these issues, in this paper, we propose a deep adaptive information filtering network (FilterNet) for accurate and fast image SR. In contrast to the existing methods that adopt fully CNN methods to directly predict the HR images, the proposed FilterNet concentrates on more useful features and adaptively filters the redundant low-frequency information. In general, we present the dilated residual group (DRG), which consists of multiple dilated residual units. The DRGs can directly expand the receptive field of the network to efficiently exploit the contextual information of the LR input. In the dilated residual unit, a gated selective mechanism is proposed to adaptively learn more high-frequency information and filter the low-frequency information. Besides, we introduce a novel adaptive information fusion structure, which builds long scaling skip connections among the DRGs to rescale the hierarchical features and fuse more detailed information. The scaling weights can be deemed as the part parameters of our network and trained adaptively. The Extensive evaluations on benchmark datasets demonstrate that our FilterNet achieves superior performance both on accuracy and speed compared with recent state-of-the-art methods. Feng Li 0037, Huihui Bai 0001, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Learning a Deep Dual Attention Network for Video Super-ResolutionabstractRecently, deep learning based video super-resolution (SR) methods combine the convolutional neural networks (CNN) with motion compensation to estimate a high-resolution (HR) video from its low-resolution (LR) counterpart. However, most previous methods conduct downscaling motion estimation to handle large motions, which can lead to detrimental effects on the accuracy of motion estimation due to the reduction of spatial resolution. Besides, these methods usually treat different types of intermediate features equally, which lack flexibility to emphasize meaningful information for revealing the high-frequency details. In this paper, to solve above issues, we propose a deep dual attention network (DDAN), including a motion compensation network (MCNet) and a SR reconstruction network (ReconNet), to fully exploit the spatio-temporal informative features for accurate video SR. The MCNet progressively learns the optical flow representations to synthesize the motion information across adjacent frames in a pyramid fashion. To decrease the mis-registration errors caused by the optical flow based motion compensation, we extract the detail components of original LR neighboring frames as complementary information for accurate feature extraction. In the ReconNet, we implement dual attention mechanisms on a residual unit and form a residual attention unit to focus on the intermediate informative features for high-frequency details recovery. Extensive experimental results on numerous datasets demonstrate the proposed method can effectively achieve superior performance in terms of quantitative and qualitative assessments compared with state-of-the-art methods. Feng Li 0037, Huihui Bai 0001, Yao Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Improving Cube-to-ERP Conversion Performance with Geometry Features of 360 Video Structureabstract360 videos provide an omnidirectional view of the scene with extremely large data. Therefore, representing 360 videos with less data has become more and more important. Cube format is such a popular representation of 360 videos. However, we have to convert cube to Equirectangula(ERP) for displaying convenience. In this paper, we enhance Cube-to-ERP conversion performance by joint using Convolutional Neural Network(CNN) and classical interpolation method. The optimal threshold of boundary is derived according to geometry features of the cube-to-ERP format. This threshold is the guidance of how to combine CNN and classical interpolation method. Our experiment results prove that the derived threshold has a certain degree of guiding significance. Furthermore, we propose a new evaluation criterion with the help of Marsaglia model. It is much easier and more accurate to evaluate geometry conversion process. Chunyu Lin, Huihui Bai 0001, Meiqin Liu 0002, Yao Zhao 0001 |
DCC | 3 |
| 2019 | Rate Control Algorithm in HEVC Based on Scene-Change DetectionabstractIn HEVC, bit-allocation model is based on the hierarchical control, which can divide video sequence into three levels: Group of Picture (GOP), frame and Coding Tree Unit (CTU). However, the fixed size of GOP fails to consider the influence of scene change in the video coding process, which may decrease the compression efficiency and reconstructed quality. In this paper, the main idea of the proposed algorithm is to detect the scene change efficiently, and then apply it in the rate control algorithm of HEVC to decrease the BD-rate and save the coding time. Huihui Bai 0001, Yao Zhao 0001 |
DCC | 2 |
| 2019 | Deep Multiple Description Coding by Learning Scalar QuantizationabstractIn this paper, we propose a deep multiple description coding framework, whose quantizers are adaptively learned via the minimization of multiple description compressive loss. Firstly, our framework is built upon auto-encoder networks, which have multiple description multi-scale dilated encoder network and multiple description decoder networks. Secondly, two entropy estimation networks are learned to estimate the informative amounts of the quantized tensors, which can further supervise the learning of multiple description encoder network to represent the input image delicately. Thirdly, a pair of scalar quantizers accompanied by two importance-indicator maps is automatically learned in an end-to-end self-supervised way. Finally, multiple description structural dissimilarity distance loss is imposed on multiple description decoded images in pixel domain for diversified multiple description generations rather than on feature tensors in feature domain, in addition to multiple description reconstruction loss. Through testing on two commonly used datasets, it is verified that our method is beyond several state-of-the-art multiple description coding approaches in terms of coding efficiency. Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001 |
DCC | 2 |
| 2019 | Semantic Map Based Image Compression via Conditional Generative Adversarial Network
Zhensong Wei, Zeyi Liao, Huihui Bai 0001, Yao Zhao 0001 |
ICIG (3) | 3 |
| 2019 | Detail-preserving image super-resolution via recursively dilated residual network
Feng Li 0037, Huihui Bai 0001, Yao Zhao 0001 |
Neurocomputing | 2 |
| 2019 | Learning a virtual codec based on deep convolutional neural network to compress image
Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Iterative range-domain weighted filter for structural preserving image smoothing and de-noising
Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001 |
Multim. Tools Appl. | 2 |
| 2019 | Simultaneous color-depth super-resolution with conditional generative adversarial networks
Lijun Zhao 0002, Huihui Bai 0001, Jie Liang 0001, Bing Zeng 0001, Anhong Wang, Yao Zhao 0001 |
Pattern Recognit. | 2 |
| 2019 | Local activity-driven structural-preserving filtering for noise removal and image smoothing
Lijun Zhao 0002, Huihui Bai 0001, Jie Liang 0001, Anhong Wang, Bing Zeng 0001, Yao Zhao 0001 |
Signal Process. | 2 |
| 2019 | Multiple Description Convolutional Neural Networks for Image CompressionabstractMultiple description coding (MDC) is able to stably transmit signal in un-reliable and non-prioritized networks, which has been broadly studied for several decades. However, traditional MDC does not well leverage image's context features to generate multiple descriptions. In this paper, we propose a novel standard-compliant convolutional neural network-based MDC framework, which efficiently leverages image's context information to compress the image. First, multiple description generator network (MDGN) is designed to produce appearance-similar yet feature-different multiple descriptions automatically according to image's content, which are compressed by a standard codec. Second, we present multiple description reconstruction network (MDRN) including side reconstruction networks (SRNs) and central reconstruction network (CRN). When any one of two lossy descriptions is received at decoder, SRN network is used to improve the quality of this decoded lossy description by simultaneously removing compression artifact and up-sampling. Meanwhile, we utilize CRN network with two decoded descriptions as inputs for better reconstruction, if both of lossy descriptions are available. Third, multiple description virtual codec network is proposed to bridge the gap between MDGN network and MDRN network in order to train an end-to-end MDC framework. Here, two learning algorithms are provided to train our whole framework. In addition to structural dis-similarity loss function, the produced descriptions are used as opposing labels with multiple description distance loss function to regularize the training of MDGN network. These losses guarantee that the generated descriptions are structurally similar yet finely diverse. Experimental results show a great deal of objective and subjective quality measurements to validate the effectiveness of our framework. Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Warping and Blending Enhancement for 3D View Synthesis Based on Grid Deformation
Ningning Hu, Yao Zhao 0001, Huihui Bai 0001 |
ICIG (3) | 3 |
| 2017 | Convolutional neural network-based depth image artifact removalabstractIn 3D video coding and depth-based image rendering, the distortion of the compressed depth image often leads to wrong 3D warpping. In this paper, by generalizing the recent work of convolutional neural network (CNN)-based depth image up-sampling, we propose a CNN-based depth image artifact removal scheme, where both the compressed depth and color images are used to enhance the depth accuracy. The proposed CNN has two sub-networks: joint depth-color sub-network and joint depth sub-network. During the depth and color feature extraction, the gradient of the depth image is used as the input to color image, while the gradient of color image is used as the input of depth feature extraction. Such an exchange of gradient information improves the learned features. Experimental results in terms of both objective and subjective quality of the depth and color images verify the efficiency of the proposed method. Lijun Zhao 0002, Jie Liang 0001, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001 |
ICIP | 3 |
| 2017 | Single depth image super-resolution with multiple residual dictionary learning and refinementabstractLearning-based image super-resolution methods often use large datasets to learn texture features. When these methods are applied to depth images, emphasis should be given on learning the geometrical structures at object boundaries, since depth images do not have much texture information. In this paper, we develop a scheme to learn multiple residual dictionaries from only one external image. After depth image super-resolution, some artifacts may appear. An adaptive depth map refinement method is then proposed to remove these artifacts along the depth edges, based on the shape-adaptive weighted median filtering method. Experimental results demonstrate the advantage of the proposed method over many other methods. Lijun Zhao 0002, Huihui Bai 0001, Jie Liang 0001, Anhong Wang, Yao Zhao 0001 |
ICME | 2 |
| 2017 | Depth map up-sampling with fractal dimension and texture-depth boundary consistencies
Meiqin Liu 0002, Yao Zhao 0001, Jie Liang 0001, Chunyu Lin, Huihui Bai 0001 |
Neurocomputing | 5 |
| 2017 | Two-stage filtering of compressed depth images with Markov Random Field
Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001, Bing Zeng 0001 |
Signal Process. Image Commun. | 2 |
| 2016 | Just Noticeable Difference Based Fast Coding Unit Partition in 3D-HEVC Intra CodingabstractSummary form only given. This paper mainly studies currently developing 3D video coding based on HEVC. HEVC-based 3D video coding mainly focuses on 3DTV and auto-stereoscopic video compression system. A variety of new encoding tools, such as inter-view motion prediction and depth modeling modes, have been added in 3D-HEVC. Although 3D-HEVC provides greater bit rate saving, it also brings the enormous encoding complexity increase. The coding time is increased correspondingly. It is necessary to reduce the encoding time. In this paper, a fast CU-sized partition algorithm is proposed for 3D-HEVC intra coding. The key point of this algorithm is to find the relationship between the texture characteristic and the sub-partition in each CU. It needs to determine whether the LCU can be subdivided to smaller CU according to the relationship. In order to reduce the redundancy of the human eye, just noticeable difference (JND) is a high efficiency model in the base of psychology and physiology. Instead of the time-consuming rate distortion optimization for coding mode decision, the variance of JND in each CU can be exploited to partition the coding unit according to human visual system characteristics. In other words, the larger blocks with higher JND variance will be subdivided to smaller blocks with lower JND variance. Consequently, the rules of CU preliminary partition are decided as follows: (a) For a 64×64 CU, if the variance of JND is larger than 0.25, the CU will be sub-divided into four 32×32 sub-blocks. (b) For a 32×32 CU, if the variance of JND is larger than 0.15, the CU will be sub-divided into four 16×16 sub-blocks. (c) For a 16×16 CU, if the variance of JND is larger than 0.10, the CU will be sub-divided into four 8×8 sub-blocks. The proposed algorithm is implemented based on HTM-13.1 reference software. The experiment condition is set up as "All Intra-Main" (AI-Main) configuration [1]. The quantization parameter (QP) values of texture are set to 25, 30, 35 and 40, respectively and the corresponding QPs of depth can be set to 34,39,42,45. The experimental results show that the fast intra mode decision algorithm provides over 29.25% encoding time saving on average with comparable rate distortion performance. Hai Ren, Huihui Bai 0001, Chunyu Lin, Mengmeng Zhang 0008, Yao Zhao 0001 |
DCC | 2 |
| 2016 | Joint iterative guidance filtering for compressed depth imagesabstractIn the general 3D scene, the correlation of depth image and corresponding color image exists, so many filtering methods have been proposed to improve the quality of depth images according to this correlation. Unlike the conventional methods, in this paper both depth and color information can be jointly employed to improve the quality of compressed depth image by the way of iterative guidance. Firstly, due to noises and blurring in the compressed image, a depth pre-filtering method is essential to remove artifact noises. Considering that the received geometry structure in the distorted depth image is more reliable than its color image, the color information is merged with depth image to get depth-merged color image. Then the depth image and its corresponding depth-merged color image can be used to refine the quality of the distorted depth image using joint iterative guidance filtering method. Therefore, the efficient depth structural information included in the distorted depth images are preserved relying on depth itself, while the corresponding color structural information are employed to improve the quality of depth image. We demonstrate the efficiency of the proposed filtering method by comparing objective and visual quality of the synthesized image with many existing depth filtering methods. Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001 |
VCIP | 2 |
| 2016 | Depth Map Down-Sampling and Coding Based on Synthesized View DistortionabstractIn this paper, we propose a depth map down-sampling and coding scheme that minimizes the view synthesis distortion. Moreover, a solution for the optimal depth map down-sampling problem that minimizes the depth-caused distortion in the virtual view by exploiting the depth map and the associated texture information along with the up-sampling method to be used in the decoder side is derived. Furthermore, to enhance compression performance, the synthesized view distortion, which is evaluated by emulating the interpolation and the virtual view synthesis process, is used in the optimization objective function for coding mode selection in the video encoder. Experimental results show that both the proposed depth map down-sampling and encoding methods lead to good performance, and the average bit rate reduction is 2.62% compared with 3D-AVC. Jimin Xiao, Tammam Tillo, Yao Zhao 0001, Chunyu Lin, Huihui Bai 0001 |
IEEE Trans. Multim. | 6 |
| 2015 | Intra-/inter-View Correlation Based Multiple Description Coding for Multiview TransmissionabstractWith the development of 3D video technology, many studies have paid attention to compression efficiency and rate distortion performance. When 3D videos are transmitted over error-prone channels, they may suffer significant quality degradation. In this paper, we combine multiview video coding (MVC) with multiple description coding (MDC) for robust transmission. The proposed scheme can give full consideration of both intra-view and inter-view correlation for better estimation. Furthermore, an adaptive mode decision is designed to generate a label as redundant information. The experiments show that the redundant information occupies just a few bits while the PSNR values of the reconstructed videos demonstrate a significant improvement. Jiansheng Guo, Huihui Bai 0001, Chunyu Lin, Mengmeng Zhang 0008, Yao Zhao 0001 |
DCC | 2 |
| 2015 | Texture Characteristics Based Fast Coding Unit Partition in HEVC Intra CodingabstractHigh efficiency video coding (HEVC) is an emerging video compression standard, developed by the Joint Collaborative Team on Video Coding (JCT-VC). The aim of HEVC standardization effort is to save about 50% bit rate for equal perceptual video quality relative to H.264/AVC. Although HEVC provides greater bit rate saving, it also brings the enormous encoding complexity increase. In this paper, we propose a fast intra CU decision algorithm based on the texture characteristics of video. Furthermore, we also consider the coding bits of each CU as auxiliary information to refine the partition results. Experimental results show that the fast intra mode decision algorithm provides over 33% complexity reduction in terms of encoding time with negligible quality loss, compared with the original HEVC test model version HM-12.0+RExt-4.0rc2. Huihui Bai 0001, Chunyu Lin, Mengmeng Zhang 0008, Yao Zhao 0001 |
DCC | 2 |
| 2015 | Quantized dictionary for sparse representationabstractDictionary learning for sparse representation has drawn considerable attention in recent years. In particular, the K-SVD algorithm is an efficient approach, and various modifications of the K-SVD have been developed for applications such as face recognition. However, the efficient storage of the dictionary has not been studied. Currently, the dictionary is simply normalized and saved as floating-point numbers, which could be quite large and lead to excessive cost and delay if the dictionary needs to be transmitted, e.g., to mobile users. In this paper, we develop a quantized K-SVD (Q-KSVD) to reduce the storage of the dictionary. We compress each basis image in the dictionary by the conventional image coding method. Moreover, we integrate the image compression step into various modified K-SVD optimization schemes, and develop an algorithm to find the optimal dictionary when there is a constraint on the total bits of the compressed dictionary. Our algorithm selects dictionary bases by ranking the contribution-rate slopes of all bases. This method also serves as an efficient approach to find the optimal number of bases of the dictionary at each rate constraint. Face recognition experiments using four K-SVD-based methods show that our method can achieve different tradeoffs between the dictionary storage space and the recognition accuracy. It can achieve comparable performance with as little as 3% of the original storage space. It can even yield higher accuracy than the uncompressed dictionary in some cases. Jie Liang 0001, Yao Zhao 0001, Chunyu Lin, Huihui Bai 0001 |
MMSP | 5 |
| 2014 | Two-Stage Multiview Image Compression Using Interview SIFT MatchingabstractIn this paper, a novel scheme of two-stage multiview image compression is proposed to create two-level reconstructed quality. Differently from the conventional multiview image compression algorithms, SIFT (Scale-Invariant Feature Transform) features matching from interview images are exploited to remove the correlations between multiple views. In the first stage coding, SIFT and RANSAC (RANdom SAmple Consensus) algorithms are combined to calculate the correlation matrix of interview, which then can be developed to obtain the coarse reconstruction of the current view. In the second stage coding, the reconstructed quality can be improved further by using the residual information. The experimental results have shown that at higher compression ratio, the proposed scheme can obtain better rate-distortion performance than intra coding in MVC (Multiview Video Coding). Furthermore, with the change of the compression ratio, the proposed scheme can achieve more stable reconstructed quality. Huihui Bai 0001, Mengmeng Zhang 0008, Meiqin Liu 0002, Anhong Wang, Yao Zhao 0001 |
DCC | 1 |
| 2014 | SNR Scalable Extension for 3D-HEVCabstractIn this paper we present a SNR scalable extension design on three-dimensional video compression using High Efficiency Video Coding (3D-HEVC). A multi-loop decoder solution is integrated into the proposed scalable coding serves as the whole framework for the SNR scalable 3D-HEVC. To effectively improve the coding performance, an inter-layer texture prediction is extended into the proposed scalable scenario for texture views and depth maps. To further reduce complexity and bitrate, a novel inter-layer distortion less prediction method is added because of the smooth texture characteristic in depth maps. Mengmeng Zhang 0008, Hongyun Lu, Huihui Bai 0001 |
DCC | 3 |
| 2014 | Fast Intra Prediction Based BCIM for Depth-Map in 3D-HEVCabstractThis paper presents a novel compression algorithm to replace Depth Modeling Mode for coding the depth-map. The result demonstrates that the execution time is reduced on an average 54.7% while the BD-rate of virtual views increase only 1.47%. Mengmeng Zhang 0008, Shenghui Qiu, Huihui Bai 0001 |
DCC | 3 |
| 2014 | Fast intra partition algorithm for HEVC screen content codingabstractSince the publication of the High Efficiency Video Coding standard as the newest video coding standard, several extensions have been made. Among these, the use of the screen content coding in many fields is one of the important extensions. In terms of coding tree unit (CTU) partitioning, rate distortion optimization is still used in screen content coding. The complexity of the process has resulted in problems in relation to real-time application. Thus, this paper proposes a fast-deciding CTU partition mode algorithm based on entropy and coding bits. Experimental results show that the proposed algorithm can save 32% of encoding time on average compared with the default algorithm in HM-12.1+RExt-5.1 with only 0.8% bit rate increment in coding performance. Mengmeng Zhang 0008, Yuhui Guo, Huihui Bai 0001 |
VCIP | 3 |
| 2014 | Multiple description video coding using correlation optimized temporal sampling
Huihui Bai 0001, Mengmeng Zhang 0008, Anhong Wang, Yao Zhao 0001 |
Sci. China Inf. Sci. | 1 |
| 2014 | Multiple Description Video Coding Based on Human Visual System CharacteristicsabstractIn this paper, a novel multiple description video coding scheme is proposed based on the characteristics of the human visual system (HVS). Due to the underlying spatial-temporal masking properties, human eyes cannot sense any changes below the just noticeable difference (JND) threshold. Therefore, at an encoder, only the visual information that cannot be predicted well within the JND tolerance needs to be encoded as redundant information, which leads to more effective redundancy allocation according to the HVS characteristics. Compared with the relevant existing schemes, the experimental results exhibit better performance of the proposed scheme at same bit rates, in terms of perceptual evaluation and subjective viewing. Huihui Bai 0001, Weisi Lin, Mengmeng Zhang 0008, Anhong Wang, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Multiple Description Coding With Randomly and Uniformly Offset QuantizersabstractIn this paper, two multiple description coding schemes are developed, based on prediction-induced randomly offset quantizers and unequal-deadzone-induced near-uniformly offset quantizers, respectively. In both schemes, each description encodes one source subset with a small quantization stepsize, and other subsets are predictively coded with a large quantization stepsize. In the first method, due to predictive coding, the quantization bins that a coefficient belongs to in different descriptions are randomly overlapped. The optimal reconstruction is obtained by finding the intersection of all received bins. In the second method, joint dequantization is also used, but near-uniform offsets are created among different low-rate quantizers by quantizing the predictions and by employing unequal deadzones. By generalizing the recently developed random quantization theory, the closed-form expression of the expected distortion is obtained for the first method, and a lower bound is obtained for the second method. The schemes are then applied to lapped transform-based multiple description image coding. The closed-form expressions enable the optimization of the lapped transform. An iterative algorithm is also developed to facilitate the optimization. Theoretical analyzes and image coding results show that both schemes achieve better performance than other methods in this category. Lili Meng, Jie Liang 0001, Upul Samarawickrama, Yao Zhao 0001, Huihui Bai 0001, André Kaup |
IEEE Trans. Image Process. | 5 |
| 2013 | M-channel multiple description coding based on uniformly offset quantizers with optimal deadzoneabstractThis paper proposes an improved source-splitting-based two-rate M-channel multiple description coding scheme, where the source is split into M subsets. In each description, one subset is coded at a high rate, and others are predictively coded at a low rate. Uniform offsets among low-rate quantizers of different descriptions are achieved by employing unequal deadzones and by quantizing the predictions. When several descriptions are received, the optimal reconstruction of each subset is achieved by finding the intersection of all received quantization bins. The closed-form expression of the expected distortion is obtained. The proposed scheme is applied to lapped transform-based multiple description image coding and achieves improved performance. The optimal deadzone selection and its impact are also given in this paper. Lili Meng, Jie Liang 0001, Yao Zhao 0001, Huihui Bai 0001, Chunyu Lin, André Kaup |
ICASSP | 4 |
| 2013 | Multiple description coding with randomly offset quantizersabstractA multiple description coding scheme based on prediction-induced randomly offset quantizers is proposed, where each description encodes one source subset with a small quantization stepsize, and other subsets are predictively coded with a large quantization stepsize. Due to the prediction, the quantization bins that a coefficient belongs to in different descriptions are randomly overlapped with each others. The optimal reconstruction is obtained by finding the intersection of all received quantization bins. Using the recently developed random quantization theory, the closed-form expression of the expected distortion is obtained. The proposed scheme is then applied to lapped transform-based multiple-description image coding, and an iterative optimization scheme is developed to find the optimal lapped transform. Experimental results show that the proposed scheme achieves better performance than other methods in this category. Lili Meng, Jie Liang 0001, Upul Samarawickrama, Yao Zhao 0001, Huihui Bai 0001, André Kaup |
ISCAS | 5 |
| 2013 | A fast depth-map wedgelet partitioning scheme for intra prediction in 3D video codingabstractBy employing 35 intra prediction modes, HEVC standard performs better to remove spatial redundancy between the current block and its neighbors. Although 3D video coding has adopted HEVC intra prediction and Depth Modeling Modes (DMM) technology to improve the performance, the Explicit Wedgelet Partition Mode within DMM brings unaffordable complexity to compress depth map. Based on the way of obtaining the picture texture from the mode with Sum of Absolute Transform Difference in rough mode decision, we propose a fast scheme to determine Wedgelet Partition in intra prediction, which leads to significant computational saving with marginal BD-rate increase after decoder-side view synthesis. Mengmeng Zhang 0008, Jizheng Xu, Huihui Bai 0001 |
ISCAS | 4 |
| 2013 | Fast bottom-up pruning for HEVC intraframe codingabstractIn intraframe coding of the High Efficiency Video Coding (HEVC) standard, up to 35 modes are defined for intra prediction and the quadtree structure is used for adaptive block partition. While such flexibility leads to more efficient compression, it also dramatically increases the encoder complexity. In this paper, a simple yet effective fast bottom-up pruning algorithm is proposed to reduce the computational cost. Mode decision at a large coding unit (CU) is selectively skipped based on the block structures of its sub-CUs. Our experimental results show that the proposed scheme can effectively reduce the encoder complexity without compromising the compression efficiency. Han Huang 0001, Yao Zhao 0001, Chunyu Lin, Huihui Bai 0001 |
VCIP | 4 |
| 2013 | Control-Point Representation and Differential Coding Affine-Motion CompensationabstractThe affine-motion model is able to capture rotation, zooming, and the deformation of moving objects, thereby providing a better motion-compensated prediction. However, it is not widely used due to difficulty in both estimation and efficient coding of its motion parameters. To alleviate this problem, a new control-point representation that favors differential coding is proposed for efficient compression of affine parameters. By exploiting the spatial correlation between adjacent coding blocks, motion vectors at control points can be predicted and thus efficiently coded, leading to overall improved performance. To evaluate the proposed method, four new affine prediction modes are designed and embedded into the high-efficiency video coding test model HM1.0. The encoder adaptively chooses whether to use the new affine mode in an operational rate-distortion optimization. Bitrate savings up to 33.82% in low-delay and 23.90% in random-access test conditions are obtained for low-complexity encoder settings. For high-efficiency settings, bitrate savings up to 14.26% and 4.89% for these two modes are observed. Han Huang 0001, John W. Woods, Yao Zhao 0001, Huihui Bai 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | Generalized Gradient Vector Flow for Snakes: New Observations, Analysis, and ImprovementabstractSnakes, or active contours, have been widely used in image processing applications. An external force for snakes called gradient vector flow (GVF) attempts to address traditional snake problems of initialization sensitivity and poor convergence to concavities, while generalized GVF (GGVF) aims to improve GVF snake convergence to long and thin indentations (LTIs). In this paper, we find and show that both GVF and GGVF snakes essentially yield the same performance in capturing LTIs of odd widths, and generally neither can converge to even-width LTIs. Based on a thorough investigation of the GVF and GGVF fields within the LTI during their iterative processes, we identify the crux of the convergence problem, and accordingly propose a novel external force termed as component-normalized GGVF (CN-GGVF) to eliminate the problem. CN-GGVF is obtained by normalizing each component of initial GGVF vectors with respect to its own magnitude. Experimental results and comparisons against GGVF snakes show that the proposed CN-GGVF snakes can capture LTIs regardless of odd or even widths with a remarkably faster convergence speed, while preserving other desirable properties of GGVF snakes with lower computational complexity in vector normalization. Lunming Qin, Ce Zhu, Yao Zhao 0001, Huihui Bai 0001, Huawei Tian |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | Multiple Description Video Coding Using Macro Block Level Correlation of Inter-/Intra-DescriptionsabstractMultiple description coding (MDC) is a promising technology for robust transmission over error-prone channels, which has attracted a lot research interests. The basic idea of MDC is to how to utilize redundant information of the descriptions for robust transmission. In view of practical applications, many MDC approaches have been proposed compatible with a certain standard codec, especially H.264/AVC. In this paper, we attempt to develop a novel MD video codec with generalized compatibility, which aims to the effective redundancy allocation from inter-/intra-descriptions. In [1], the redundancy allocation may be not enough effective due to frame level. As a result, in this paper, the redundant information will be taken into account at MB level. Huihui Bai 0001, Mengmeng Zhang 0008, Meiqin Liu 0002, Anhong Wang, Yao Zhao 0001 |
DCC | 1 |
| 2012 | Temporal Sampling Based Multiple Description Video Coding for Scenes SwitchingabstractDue to network congestion and delay sensibility, it is always a great challenge for video transmission over lossy network. Multiple description coding (MDC) is an attractive approach to solve this problem. It can efficiently combat packet loss without any retransmission thus satisfying the demand of real time services and relieving the network congestion. In view of perfect compatibility with the standard source and channel codec, temporal sampling based MDC has become a better choice for practical applications. However, for the frames switching from one scene to another temporal correlation may be destroyed by sampled in temporal domain, which may result in the false estimation when the related frames are lost at the side decoder. To address this problem, in this paper an improved MD coding based on temporal sampling is proposed to make sure the decoder can work correctly when scenes changing. Mengmeng Zhang 0008, Huihui Bai 0001 |
DCC | 2 |
| 2012 | Entropy analysis on multiple description video coding based on pre- and post-processingabstractMultiple description (MD) video coding is a promising method to solve real-time video transmission over unreliable network. In the conventional MD video coding, the original video sequence can be split directly into two subsequences by odd and even means. Then the two sub-sequences can be compressed as two descriptions by the standard video encoder. The conventional MD scheme is simple to realize but it may lead to worse reconstructed quality when one description is lost. To solve this problem, the MD scheme based on pre- and post-processing is proposed in this paper. Before odd and even splitting, the original video sequence can be pre-processed by effective redundancy allocation, which is helpful for the estimation of the lost description. Furthermore, the entropy of the descriptions is used to analyse the rate-distortion performance of the two MD schemes. Lastly, the experimental results have shown the proposed MD scheme has better reconstructed quality when information lost has happened, while the conventional MD scheme has better compression efficiency when information can be transmitted accurately. It can be found that the experimental results can be consistent with the entropy analysis. Huihui Bai 0001, Anhong Wang, Ajith Abraham |
HIS | 1 |
| 2012 | Stereo video coding using distributed compressive sensing with joint dictionaryabstractFor many practical applications, stereo-paired video is an important special case of multiview video coding (MVC). This paper presents a novel framework of stereo video coding based on distributed compressive sensing, which also can be easily extended to MVC. According to distributed video coding (DVC) at the encoder the video sequences from each view can be compressed independently without any communications between the cameras while at the decoder the inter-view correlation can be exploited for quality enhancement. Furthermore, due to compressive sensing (CS) principles, low complexity at the encoder side can result in low power consumption in the cameras, which may be promising in wireless camera sensor network. Here, the joint dictionary is applied in compressive sensing, which can make good use of inter-view and temporal correlation for better reconstruction quality of convex optimization. The experimental results validate the effectiveness of the proposed scheme with better performance than other compared schemes. Huihui Bai 0001, Mengmeng Zhang 0008, Anhong Wang, Yao Zhao 0001 |
ICIP | 1 |
| 2012 | Affine SKIP and DIRECT modes for efficient video codingabstractHigher-order motion models were introduced in video coding a couple of decades ago, but have not been widely used due to both difficulty in parameters estimation and their requirement of more side information. Recently, researchers have put them back into consideration. In this paper, the affine motion model is employed in SKIP and DIRECT modes to produce a better prediction. In affine SKIP/DIRECT, candidate predictors of the motion parameters are derived from the motions of neighboring coded blocks, with the best predictor determined by rate-distortion tradeoff. Extensive experiments have shown the efficiency of these new affine modes. No additional motion estimation is needed, so the proposed method is also quite practical. Han Huang 0001, John W. Woods, Yao Zhao 0001, Huihui Bai 0001 |
VCIP | 4 |
| 2010 | GOP-Flexible Distributed Multiview Video Coding with Adaptive Side Information
Lili Meng, Yao Zhao 0001, Jeng-Shyang Pan 0001, Huihui Bai 0001, Anhong Wang |
ICCCI (3) | 4 |
| 2009 | Robust multiple description distributed video coding using optimized zero-padding
Anhong Wang, Yao Zhao 0001, Huihui Bai 0001 |
Sci. China Ser. F Inf. Sci. | 3 |
| 2008 | Priority Encoding Transmission Based Multiple Description Video Coding over Packet Loss NetworkabstractIn this paper, we attempt to overcome the limitation of specific scalable video codec and apply FEC-MDC to a common video coder, such as the standard H.264. The proposed scheme is explained as follows. Firstly, according to motion vector changes, an original video sequence is divided into several sub-sequences as messages, so in each message better temporal correlation can be maintained for better estimation when information losses occur. Secondly, the standard H.264 encoder is used to encode the messages. Thirdly, based on priority encoding transmission, unequal protections are assigned in each message. Lastly, at the decoder, the segments whose priorities are not higher than the fraction of packets received can be recover totally. Huihui Bai 0001, Yao Zhao 0001, Ce Zhu |
DCC | 1 |
| 2007 | Multiple Description Video Coding using Adaptive Temporal Sub-SamplingabstractMultiple description coding (MDC) is a promising alternative for robust transmission of information over non-prioritized and unpredictable networks. Especially, MDC has emerged as an attractive approach for video applications where retransmission is unacceptable or infeasible. In this paper, an effective MD video codec is designed based on pre-and post-processing of video sequences, without any modification to the source or channel codec. Considering different motion information of inter-frame, adaptive temporal sub-sampling is employed to the original video data. As a result, adaptive redundancy can make a better tradeoff between the reconstruction quality and compression efficiency. The experimental results exhibit better performance of the proposed scheme than other schemes. Huihui Bai 0001, Yao Zhao 0001, Ce Zhu |
ICME | 1 |
| 2007 | Optimized Multiple Description Lattice Vector Quantization for Wavelet Image CodingabstractMultiple description (MD) coding is a promising alternative for robust transmission of information over non-prioritized and unpredictable networks. In this paper, an effective MD image coding scheme is introduced based on the MD lattice vector quantization (MDLVQ) for the wavelet transformed images. In view of the characteristics of wavelet coefficients in different frequency subbands, MDLVQ is applied in an optimized way, including an appropriate construction of wavelet coefficient vectors, the optimization of MDLVQ encoding parameters such as the choice of sublattice index values and the quantization accuracy for different subbands. More importantly, optimized side decoding is employed to predict lost information based on inter-vector correlation and an alternative transmission way for further reducing side distortion. Experimental results validate the effectiveness of the proposed scheme with better performance than some other tested MD image codecs including that based on optimized MD scalar quantization. Huihui Bai 0001, Ce Zhu, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | Multiple Description Shifted Lattice Vector Quantization for Progressive Wavelet Image CodingabstractMultiple description (MD) coding is a promising alternative for robust transmission of information over non-prioritized and unpredictable networks. Furthermore, practical variable-bandwidth channels also require fine grain scalability of the descriptions (bit streams). In this paper, according to the geometrical structure and the special relationship of lattice vector quantizers, a MD quantizer called shifted lattice vector quantization (SLVQ) is employed in MD image coding to realize progressive transmission over unreliable channels. In view of the characteristics of wavelet coefficients in different frequency subbands, besides an appropriate construction of wavelet coefficient vectors, the algorithm of modified zerotree coding is also applied to improve compression performance. Experimental results validate the effectiveness of the proposed scheme with better performance than the other schemes based on MD scalar quantization for progressive transmission. Huihui Bai 0001, Yao Zhao 0001, Ce Zhu |
ICIP | 1 |