Bo Yan 0001

dblp:63/6796-1 · DBLP profile ↗
← Back
97ranked-venue papers
20as first author
45since 2021 · last 2025
0000-0003-0256-9682ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 85 · 17 first-author · 40 since 2021Artificial intelligence and machine learning · 21 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Efficient Online Training for Zero-Shot Time-Lapse Microscopy Denoising and Super-Resolution
abstract
In time-lapse microscopy, inherent noise significantly limits imaging sensitivity and increases measurement uncertainty. Due to the scarcity of clean data, zero-shot approaches have emerged as highly data-efficient solutions for microscopy denoising. However, existing methods typically process video frames independently, resulting in long training times and issues such as temporal noise and over-smoothing. In this paper, we introduce MDSR-Zero, a zero-shot online learning method designed for plug-and-play noise suppression and super-resolution of microscopy videos. Our approach leverages an efficient online training strategy that reuses denoising models from previous frames. By treating the video as a continuous stream, our model significantly reduces training time and ensures temporally consistent denoising. Additionally, we propose a novel loss function tailored for denoising in the context of super-resolution, which enhances the detail in the denoised results. Extensive experiments on both synthetic and real-world noise demonstrate that our method achieves state-of-the-art performance among zero-shot denoising approaches and is competitive with self-supervised methods. Notably, our method can reduce training time by up to 10x compared to the previous SOTA method.
Ruian He, Ri Cheng, Xinkai Lyu, Weimin Tan, Bo Yan 0001
AAAI5
2025 Efficient Trajectory Space-Time Super-Resolution for Fast Live-cell Imaging
abstract
Live-cell imaging is a powerful tool for studying dynamic subcellular processes by capturing the spatiotemporal organization of the biological microenvironment. However, limitations due to phototoxicity and photobleaching prevent microscopes from achieving high frame rates and high-quality images. Although current deep learning methods can enhance both frame rates and image resolution without compromising cell health, they often overlook the continuity of subcellular trajectories, which leads to discontinuous temporal modeling. It also incurs prohibitive computational costs due to exhaustive correlation computation that hinder real-time applications. To address these issues with high efficiency, we propose Trajectory Space-Time Super-Resolution (T-STSR), a method designed to boost frame rates and resolution in fast subcellular imaging while significantly reducing computational overhead. Our approach incorporates Spatial-Temporal Trajectory Modeling (STTM), which learns a state-space model over spatiotemporal slices to reconstruct particle trajectories at low cost. In addition, our novel Trajectory-Aware Loss randomly subsamples trajectory data during training, promoting continuous trajectory representation and mitigating noise with minimal additional computation. We validated T-STSR on both synthesized and real-world datasets with various particle types and noise conditions, demonstrating that our method achieves superior restoration results while saving 75% inference time compared to the previous SOTA model.
Ruian He, Zixian Zhang, Ri Cheng, Weimin Tan, Bo Yan 0001
ACM Multimedia5
2025 Scaling Laws for Data-Efficient Visual Transfer Learning
Wenxuan Yang, Qingqv Wei, Chenxi Ma, Weimin Tan, Bo Yan 0001
ACM Multimedia5
2025 MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks
abstract
Medical vision-language models (VLMs) have shown promise as clinical assistants across various medical fields. However, specialized dermatology VLM capable of delivering professional and detailed diagnostic analysis remains underdeveloped, primarily due to less specialized text descriptions in current dermatology multimodal datasets. To address this issue, we propose MM-Skin, the first large-scale multimodal dermatology dataset that encompasses 3 imaging modalities, including clinical, dermoscopic, and pathological and nearly 10k high-quality image-text pairs collected from professional textbooks. In addition, we generate over 27k diverse, instruction-following vision question answering (VQA) samples (9× the size of current largest dermatology VQA dataset). Leveraging public datasets and MM-Skin, we developed SkinVL, a dermatology-specific VLM designed for precise and nuanced skin disease interpretation. Comprehensive benchmark evaluations of SkinVL on VQA, supervised fine-tuning (SFT) and zero-shot classification tasks across 8 datasets, reveal its exceptional performance for skin diseases in comparison to both general and medical VLM models. The introduction of MM-Skin and SkinVL offers a meaningful contribution to advancing the development of clinical dermatology VLM assistants. Code and dataset are available at https://github.com/ZwQ803/MM-Skin.
Chenxi Ma, Weimin Tan, Bo Yan 0001
ACM Multimedia5
2025 TabiMed: Tabularizing Medical Images for Few-Shot In-Context Diagnosis
abstract
Achieving accurate predictions with limited samples is a key challenge in biomedical image artificial intelligence. Previous methods rely on pre-trained image foundation models with supervised fine-tuning (SFT) or zero-shot inference to enhance small-data performance. However, SFT is time-consuming and prone to overfitting, whereas zero-shot inference fails to fully exploit available data. Inspired by recent tabular foundation models, which show superior performance on small-sample tasks with in-context learning (ICL), we propose TabiMed, a novel framework that transforms visual representations into structured tabular data, leveraging pre-trained tabular models for fast and accurate analysis on small data. TabiMed consists of three key components: dynamic modality-aware representation engine, tabularization adapter and in-context inference module. Experiments on 10 datasets from different fields demonstrate three major advantages of TabiMed: 1) excellent performance on small datasets, with an average AUC of 14.1% higher than zero-shot; 2) high efficiency, with a training time 250x faster than SFT; 3) scalability to larger datasets through our tabularization adapter. TabiMed proposes a novel pathway to address the challenges of analyzing biomedical images with few samples.
Wanying Zhou, Yu Ling, Chenxi Ma, Weimin Tan, Bo Yan 0001
ACM Multimedia7
2025 Spatiotemporal-Aware Self-Supervised Fluorescence Microscopy Image Denoising
abstract
Fluorescence microscopy has been an indispensable tool in many scientific disciplines. However, the expensive imaging cost and the photo-toxicity problem make it difficult to obtain high-quality images. The independent shot noise in fluorescence microscopy images always overwhelms signals and limits the imaging resolution, hindering progress in related research. Recently, self-supervised image denoising has received wide attention for its ability to train a denoiser without paired low Signal-to-Noise Ratio (SNR) and high SNR images. Existing self-supervised fluorescence microscopy image denoising works either suffer from high imaging/computational cost or large training difficulty. Here, we propose a Single-image based Self-Supervised Denoising approach (TriS-D) by utilizing the spatiotemporal redundancy of the fluorescence microscopy imaging data, which facilitates the low-cost and convenient training. The TriS-D can generate the training data from a raw image, not only releasing the demand for multiple low SNR time-lapse imaging data but also enabling the building of a 2D convolution-based model. Comprehensive experiments across different imaging modalities and biological samples verify the effectiveness of the TriS-D.
Chenxi Ma, Weimin Tan, Zhaohui Zhou, Bo Yan 0001
IEEE Signal Process. Lett.4
2024 Context-Aware Iteration Policy Network for Efficient Optical Flow Estimation
abstract
Existing recurrent optical flow estimation networks are computationally expensive since they use a fixed large number of iterations to update the flow field for each sample. An efficient network should skip iterations when the flow improvement is limited. In this paper, we develop a Context-Aware Iteration Policy Network for efficient optical flow estimation, which determines the optimal number of iterations per sample. The policy network achieves this by learning contextual information to realize whether flow improvement is bottlenecked or minimal. On the one hand, we use iteration embedding and historical hidden cell, which include previous iterations information, to convey how flow has changed from previous iterations. On the other hand, we use the incremental loss to make the policy network implicitly perceive the magnitude of optical flow improvement in the subsequent iteration. Furthermore, the computational complexity in our dynamic network is controllable, allowing us to satisfy various resource preferences with a single trained model. Our policy network can be easily integrated into state-of-the-art optical flow networks. Extensive experiments show that our method maintains performance while reducing FLOPs by about 40%/20% for the Sintel/KITTI datasets.
Ri Cheng, Ruian He, Xuhao Jiang, Shili Zhou, Weimin Tan, Bo Yan 0001
AAAI6
2024 Low-Latency Space-Time Supersampling for Real-Time Rendering
abstract
With the rise of real-time rendering and the evolution of display devices, there is a growing demand for post-processing methods that offer high-resolution content in a high frame rate. Existing techniques often suffer from quality and latency issues due to the disjointed treatment of frame supersampling and extrapolation. In this paper, we recognize the shared context and mechanisms between frame supersampling and extrapolation, and present a novel framework, Space-time Supersampling (STSS). By integrating them into a unified framework, STSS can improve the overall quality with lower latency. To implement an efficient architecture, we treat the aliasing and warping holes unified as reshading regions and put forth two key components to compensate the regions, namely Random Reshading Masking (RRM) and Efficient Reshading Module (ERM). Extensive experiments demonstrate that our approach achieves superior visual fidelity compared to state-of-the-art (SOTA) methods. Notably, the performance is achieved within only 4ms, saving up to 75\% of time against the conventional two-stage pipeline that necessitates 17ms.
Ruian He, Shili Zhou, Ri Cheng, Weimin Tan, Bo Yan 0001
AAAI6
2024 MGQFormer: Mask-Guided Query-Based Transformer for Image Manipulation Localization
abstract
Deep learning-based models have made great progress in image tampering localization, which aims to distinguish between manipulated and authentic regions. However, these models suffer from inefficient training. This is because they use ground-truth mask labels mainly through the cross-entropy loss, which prioritizes per-pixel precision but disregards the spatial location and shape details of manipulated regions. To address this problem, we propose a Mask-Guided Query-based Transformer Framework (MGQFormer), which uses ground-truth masks to guide the learnable query token (LQT) in identifying the forged regions. Specifically, we extract feature embeddings of ground-truth masks as the guiding query token (GQT) and feed GQT and LQT into MGQFormer to estimate fake regions, respectively. Then we make MGQFormer learn the position and shape information in ground-truth mask labels by proposing a mask-guided loss to reduce the feature distance between GQT and LQT. We also observe that such mask-guided training strategy has a significant impact on the convergence speed of MGQFormer training. Extensive experiments on multiple benchmarks show that our method significantly improves over state-of-the-art methods.
Kunlun Zeng, Ri Cheng, Weimin Tan, Bo Yan 0001
AAAI4
2024 SAMFlow: Eliminating Any Fragmentation in Optical Flow with Segment Anything Model
abstract
Optical Flow Estimation aims to find the 2D dense motion field between two frames. Due to the limitation of model structures and training datasets, existing methods often rely too much on local clues and ignore the integrity of objects, resulting in fragmented motion estimation. Through theoretical analysis, we find the pre-trained large vision models are helpful in optical flow estimation, and we notice that the recently famous Segment Anything Model (SAM) demonstrates a strong ability to segment complete objects, which is suitable for solving the fragmentation problem. We thus propose a solution to embed the frozen SAM image encoder into FlowFormer to enhance object perception. To address the challenge of in-depth utilizing SAM in non-segmentation tasks like optical flow estimation, we propose an Optical Flow Task-Specific Adaption scheme, including a Context Fusion Module to fuse the SAM encoder with the optical flow context encoder, and a Context Adaption Module to adapt the SAM features for optical flow task with Learned Task-Specific Embedding. Our proposed SAMFlow model reaches 0.86/2.10 clean/final EPE and 3.55/12.32 EPE/F1-all on Sintel and KITTI-15 training set, surpassing Flowformer by 8.5%/9.9% and 13.2%/16.3%. Furthermore, our model achieves state-of-the-art performance on the Sintel and KITTI-15 benchmarks, ranking #1 among all two-frame methods on Sintel clean pass.
Shili Zhou, Ruian He, Weimin Tan, Bo Yan 0001
AAAI4
2024 Bridging The Domain Gap Arising from Text Description Differences for Stable Text-To-Image Generation
abstract
Generating high-quality images that conform to the semantics of captions has numerous potential applications. However, text-to-image generation is a challenging task due to its cross-modality nature. Current generative models are typically unstable, meaning that complex sentences can result in poor image quality. In this paper, we propose a novel model to bridge the domain gap arising from sentence complexity to achieve stable text-to-image generation. Our model includes two key modules, the attribute extraction module and the attribute fusion module. These modules can extract attributes from the captions and fuse them with image features to encourage the model to accurately understand the semantics. Our modules are plug-and-play and extensive experiments demonstrate that our approach outperforms the state-of-the-art GAN model. Our code and trained model are available at https://github.com/tantian21/stable-t2i-generation.
Tian Tan 0016, Weimin Tan, Xuhao Jiang, Yueming Jiang, Bo Yan 0001
ICASSP5
2024 Facial Micro-Motion-Aware Mixup for Micro-Expression Recognition
abstract
Data-driven learning models have demonstrated strong benefits in capturing subtle facial movements for micro-expression recognition (MER), but are limited by the available data. Generative models can generate a variety of new data, but are typically computationally prohibitive compared to efficient Mixup-like methods. In this paper, we propose a novel Facial Micro-Motion-Aware Mixup approach for MER, namely MEMix. Our MEMix constructs a micro-motion-aware mask to select the most salient facial motions and generate a new sample with a mixed motion feature. This mixed motion feature can effectively expand the data distribution, leading to smoother decision boundaries for MER models. To demonstrate the good generality of MEMix, we integrate it with three advanced vision transformer-based models. The results show that the three integrated models consistently achieve performance improvements ranging from 4.07% to 7.32% in accuracy and from 6.54% to 9.18% in F1-score. Besides, to further explore the ability of MEMix, we propose a two-stream network called MixMeFormer, which unlocks the potential of the transformer by simply integrating mixed motion features with facial semantics for MER. Extensive experiments demonstrate that our MixMeFormer outperforms other state-of-the-art methods on three well-known micro-expression datasets.
Zhuoyao Gu, Miao Pang, Weimin Tan, Xuhao Jiang, Bo Yan 0001
ICASSP6
2024 Addressing Imbalance for Class Incremental Learning in Medical Image Classification
abstract
Deep convolutional neural networks have made significant breakthroughs in medical image classification, under the assumption that training samples from all classes are simultaneously available. However, in real-world medical scenarios, there's a common need to continuously learn about new diseases, leading to the emerging field of class incremental learning (CIL) in the medical domain. Typically, CIL suffers from catastrophic forgetting when trained on new classes. This phenomenon is mainly caused by the imbalance between old and new classes, and it becomes even more challenging with imbalanced medical datasets. In this work, we introduce two simple yet effective plug-in methods to mitigate the adverse effects of the imbalance. First, we propose a CIL-balanced classification loss to mitigate the classifier bias toward majority classes via logit adjustment. Second, we propose a distribution margin loss that not only alleviates the inter-class overlap in embedding space but also enforces the intra-class compactness. We evaluate the effectiveness of our method with extensive experiments on three benchmark datasets (CCH5000, HAM10000, and EyePACS). The results demonstrate that our approach outperforms state-of-the-art methods.
Xuze Hao, Wenqian Ni, Xuhao Jiang, Weimin Tan, Bo Yan 0001
ACM Multimedia5
2024 FacialFlowNet: Advancing Facial Optical Flow Estimation with a Diverse Dataset and a Decomposed Model
abstract
Facial movements play a crucial role in conveying altitude and intentions, and facial optical flow provides a dynamic and detailed representation of it. However, the scarcity of datasets and a modern baseline hinders the progress in facial optical flow research. This paper proposes FacialFlowNet (FFN), a novel large-scale facial optical flow dataset, and the Decomposed Facial Flow Model (DecFlow), the first method capable of decomposing facial flow. FFN comprises 9,635 identities and 105,970 image pairs, offering unprecedented diversity for detailed facial and head motion analysis. DecFlow features a facial semantic-aware encoder and a decomposed flow decoder, excelling in accurately estimating and decomposing facial flow into head and expression components. Comprehensive experiments demonstrate that FFN significantly enhances the accuracy of facial flow estimation across various optical flow methods, achieving up to an 11% reduction in Endpoint Error (EPE) (from 3.91 to 3.48). Moreover, DecFlow, when coupled with FFN, outperforms existing methods in both synthetic and real-world scenarios, enhancing facial expression analysis. The decomposed expression flow achieves a substantial accuracy improvement of 18% (from 69.1% to 82.1%) in micro-expressions recognition. These contributions represent a significant advancement in facial motion analysis and optical flow estimation. Codes and datasets can be found.
Jianzhi Lu, Ruian He, Shili Zhou, Weimin Tan, Bo Yan 0001
ACM Multimedia5
2024 Learning Cross-Spectral Prior for Image Super-Resolution
abstract
With the rising interest in multi-camera cross-spectral systems, cross-spectral images have been widely used in computer vision and image processing. Therefore, an effective super-resolution (SR) method provides high-resolution (HR) cross-spectral images for different research and applications. However, existing SR methods rarely consider utilizing cross-spectral information to assist the SR of visible images. They cannot handle complex degradation (noise, high brightness, low light) and misalignment problems in low-resolution (LR) cross-spectral images. Here, we first explore the potential of using near-infrared (NIR) image guidance for better SR, based on the observation that NIR images can preserve valuable information for recovering adequate image details. To take full advantage of the cross-spectral prior, we propose a novel Cross-Spectral Prior guided image SR approach (CSPSR). The cross-view matching (CVM) module and the dynamic multi-modal fusion (DMF) module can enhance the spatial correlation between cross-spectral images and bridge the multi-modal feature gap, respectively. Extensive experiments demonstrate the effectiveness of our CSPSR.
Chenxi Ma, Weimin Tan, Shili Zhou, Bo Yan 0001
ACM Multimedia4
2024 Audio-Driven Identity Manipulation for Face Inpainting
abstract
Recent advances in multimodal artificial intelligence have greatly improved the integration of vision-language-audio cues to enrich the content creation process. Inspired by these developments, in this paper, we first integrate audio into the face inpainting task to facilitate identity manipulation. Our main insight is that a person's voice carries distinct identity markers, such as age and gender, which provide an essential supplement for identity-aware face inpainting. By extracting identity information from audio as guidance, our method can naturally support tasks of identity preservation and identity swapping in face inpainting. Specifically, we introduce a dual-stream network architecture comprising a face branch and an audio branch. The face branch is tasked with extracting deterministic information from the visible parts of the input masked face, while the audio branch is designed to capture heuristic identity priors from the speaker's voice. The identity codes from two streams are integrated using a multi-layer perceptron (MLP) to create a virtual unified identity embedding that represennts comprehensive identity features. In addition, to explicitly exploit the information from audio, we introduce an audio-face generator to generate an 'fake' audio face directly from audio and fuse the multi-scale intermediate features from the audio-face generator into face inpainting network through an audio-visual feature fusion (AVFF) module. Extensive experiments demonstrate the positive impact of extracting identity information from audio on face inpainting task, especially in identity preservation.
Weimin Tan, Bo Yan 0001
ACM Multimedia4
2024 A Medical Data-Effective Learning Benchmark for Highly Efficient Pre-training of Foundation Models
Wenxuan Yang, Weimin Tan, Bo Yan 0001
ACM Multimedia4
2024 Prompt-Guided Semantic-Aware Distillation for Weakly Supervised Incremental Semantic Segmentation
abstract
Weakly Supervised Incremental Semantic Segmentation (WISS) aims to enable deep neural networks to incrementally learn new classes using only image-level labels without catastrophic forgetting. Despite WISS eliminating the usage of costly and time-consuming pixel-by-pixel annotations, the image-level labels can not provide details about the location of new classes, resulting in inferior performance. To address these issues, we take inspiration from zero-shot learning to model the inter-class semantic relation utilizing class names as text prompts, thereby facilitating knowledge transfer between classes. However, some class names of the segmentation datasets are polysemous. Thus, we design a new prompt template to better capture the semantic relation by appending synonyms and definitions of the corresponding classes. Guided by this semantic relation, we propose semantic relation weighted distillation to transfer the knowledge from old to new classes, significantly improving plasticity while reducing forgetting. Additionally, we introduce a novel superclass-level distillation aimed at preserving shared global knowledge within the superclass, further alleviating catastrophic forgetting. We extensively evaluate our method by integrating it into state-of-the-art WISS approaches on Pascal VOC and COCO datasets. We observe consistent gains in performance across diverse experimental scenarios. Code is available athttps://github.com/Magic-Nova77/PGSD.
Xuze Hao, Xuhao Jiang, Wenqian Ni, Weimin Tan, Bo Yan 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 A Motion Distillation Framework for Video Frame Interpolation
abstract
In recent years, we have seen the success of deep video enhancement models. However, the performance improvement of new methods has gradually entered a bottleneck period. Optimizing model structures or increasing training data brings less and less improvement. We argue that existing models with advanced structures have not fully demonstrated their performance and demand further exploration. In this study, we statistically analyze the relationship between motion estimation accuracy and video interpolation quality of existing video frame interpolation methods, and find that only supervising the final output leads to inaccurate motion and further affects the interpolation performance. Based on this important observation, we propose a general motion distillation framework that can be widely applied to flow-based and kernel-based video frame interpolation methods. Specifically, we begin by training a teacher model, which uses the ground-truth target frame and adjacent frames to estimate motion. These motion estimates then guide the training of a student model for video frame interpolation. Our experimental results demonstrate the effectiveness of this approach in enhancing performance across diverse advanced video interpolation model structures. For example, after applying our motion distillation framework, the CtxSyn model achieves a PSNR gain of 3.047 dB.
Shili Zhou, Weimin Tan, Bo Yan 0001
IEEE Trans. Multim.3
2024 Lesion-Decoupling-Based Segmentation With Large-Scale Colon and Esophageal Datasets for Early Cancer Diagnosis
abstract
Lesions of early cancers often show flat, small, and isochromatic characteristics in medical endoscopy images, which are difficult to be captured. By analyzing the differences between the internal and external features of the lesion area, we propose a lesion-decoupling-based segmentation (LDS) network for assisting early cancer diagnosis. We introduce a plug-and-play module called self-sampling similar feature disentangling module (FDM) to obtain accurate lesion boundaries. Then, we propose a feature separation loss (FSL) function to separate pathological features from normal ones. Moreover, since physicians make diagnoses with multimodal data, we propose a multimodal cooperative segmentation network with two different modal images as input: white-light images (WLIs) and narrowband images (NBIs). Our FDM and FSL show a good performance for both single-modal and multimodal segmentations. Extensive experiments on five backbones prove that our FDM and FSL can be easily applied to different backbones for a significant lesion segmentation accuracy improvement, and the maximum increase of mean Intersection over Union (mIoU) is 4.58. For colonoscopy, we can achieve up to mIoU of 91.49 on our Dataset A and 84.41 on the three public datasets. For esophagoscopy, mIoU of 64.32 is best achieved on the WLI dataset and 66.31 on the NBI dataset.
Weimin Tan, Shilun Cai, Bo Yan 0001, Yunshi Zhong
IEEE Trans. Neural Networks Learn. Syst.4
2023 Multi-Modality Deep Network for Extreme Learned Image Compression
abstract
Image-based single-modality compression learning approaches have demonstrated exceptionally powerful encoding and decoding capabilities in the past few years , but suffer from blur and severe semantics loss at extremely low bitrates. To address this issue, we propose a multimodal machine learning method for text-guided image compression, in which the semantic information of text is used as prior information to guide image compression for better compression performance. We fully study the role of text description in different components of the codec, and demonstrate its effectiveness. In addition, we adopt the image-text attention module and image-request complement module to better fuse image and text features, and propose an improved multimodal semantic-consistent loss to produce semantically complete reconstructions. Extensive experiments, including a user study, prove that our method can obtain visually pleasing results at extremely low bitrates, and achieves a comparable or even better performance than state-of-the-art methods, even though these methods are at 2x to 4x bitrates of ours.
Xuhao Jiang, Weimin Tan, Tian Tan 0016, Bo Yan 0001, Liquan Shen
AAAI4
2023 Fine-Grained Blind Face Inpainting with 3D Face Component Disentanglement
abstract
Inpainting is a task to restore occlusion or other corruption on images. However, previous works require mask of the occluded area to restore the occluded image, which is inconvenient for application. Blind face inpainting aims to automatically restore the occluded face without position information of the corrupt region. In this paper, we propose a novel fine-grained blind face inpainting framework, combining 3D face components disentanglement with generative network. Canonical face texture and shape disentangled by unsupervised 3D face model is restored separately to get occlusion-free rendered result. Finally, the pixel-to-pixel generative module utilize the occlusion image and the coarse de-occlusion face to get refined inpainted result. We also build up a new dataset called CelebO-3D which consists of occluded face images synthesized with 3D occlusion and rendered by 3D face model. Extensive experiments show that the proposed method is effective and robust in face blind inpainting both in synthesized and real images. Extensive evaluations and comparison with previous methods also show our superior effectiveness and lightweight architecture.
Ruian He, Weimin Tan, Bo Yan 0001, Yangle Lin
ICASSP4
2023 Uncer2Natural: Uncertainty-Aware Unsupervised Image Denoising
abstract
Recently, unsupervised image denoising methods learning from paired noisy samples have received increasing attention. These methods build on the idea that the mean of multiple noisy images of the same scene is the ideal clean image. However, these methods ignore the effect of Aleatoric uncertainty in the noisy image (e.g., pixels deviating from the expected distribution). The presence of Aleatoric uncertainty causes degradation of the reconstructed target pixels, resulting in high uncertainty for these pixels (i.e., low confidence), which in turn leads to sub-optimal denoising results. To address this problem, we propose a novel uncertainty-aware unsupervised image denoising method named Uncer2Natural (U2N). It dynamically predicts the Aleatoric uncertainty for each noisy sample and produces satisfactory denoising results by reducing the effect of Aleatoric uncertainty. Extensive experimental results show that U2N outperforms state-of-the- art unsupervised image denoising methods in terms of both quantitative metrics and qualitative visual quality.
Weimin Tan, Jiaxing Shi, Bo Yan 0001
ICASSP5
2023 Multi-Modality Deep Network for JPEG Artifacts Reduction
abstract
In recent years, many convolutional neural network-based models are designed for JPEG artifacts reduction, and have achieved notable progress. However, few methods are suitable for extreme low-bitrate image compression artifacts reduction. The main challenge is that the highly compressed image loses too much information, resulting in reconstructing high-quality image difficultly. To address this issue, we propose a multimodal fusion learning method for text-guided JPEG artifacts reduction, in which the corresponding text description not only provides the potential prior information of the highly compressed image, but also serves as supplementary information to assist in image deblocking. We fuse image features and text semantic features from the global and local perspectives respectively, and design a contrastive loss built upon contrastive learning to produce visually pleasing results. Extensive experiments, including a user study, prove that our method can obtain better deblocking results compared to the state-of-the-art methods.
Xuhao Jiang, Weimin Tan, Chenxi Ma, Bo Yan 0001, Liquan Shen
IJCAI5
2023 Learning Survival Distribution with Implicit Survival Function
abstract
Survival analysis aims at modeling the relationship between covariates and event occurrence with some untracked (censored) samples. In implementation, existing methods model the survival distribution with strong assumptions or in a discrete time space for likelihood estimation with censorship, which leads to weak generalization. In this paper, we propose Implicit Survival Function (ISF) based on Implicit Neural Representation for survival distribution estimation without strong assumptions, and employ numerical integration to approximate the cumulative distribution function for prediction and optimization. Experimental results show that ISF outperforms the state-of-the-art methods in three public datasets and has robustness to the hyperparameter controlling estimation precision.
Yu Ling, Weimin Tan, Bo Yan 0001
IJCAI3
2023 Uncertainty-Guided Spatial Pruning Architecture for Efficient Frame Interpolation
abstract
The video frame interpolation (VFI) model applies the convolution operation to all locations, leading to redundant computations in regions with easy motion. We can use dynamic spatial pruning method to skip redundant computation, but this method cannot properly identify easy regions in VFI tasks without supervision. In this paper, we develop an Uncertainty-Guided Spatial Pruning (UGSP) architecture to skip redundant computation for efficient frame interpolation dynamically. Specifically, pixels with low uncertainty indicate easy regions, where the calculation can be reduced without bringing undesirable visual results. Therefore, we utilize uncertainty-generated mask labels to guide our UGSP in properly locating the easy region. Furthermore, we propose a self-contrast training strategy that leverages an auxiliary non-pruning branch to improve the performance of our UGSP. Extensive experiments show that UGSP maintains performance but reduces FLOPs by 34%/52%/30% compared to baseline without pruning on Vimeo90K/UCF101/MiddleBury datasets. In addition, our method achieves state-of-the-art performance with lower FLOPs on multiple benchmarks.
Ri Cheng, Xuhao Jiang, Ruian He, Shili Zhou, Weimin Tan, Bo Yan 0001
ACM Multimedia6
2023 MVFlow: Deep Optical Flow Estimation of Compressed Videos with Motion Vector Prior
abstract
In recent years, many deep learning-based methods have been proposed to tackle the problem of optical flow estimation and achieved promising results. However, they hardly consider that most videos are compressed and thus ignore the pre-computed information in compressed video streams. Motion vectors, one of the compression information, record the motion of the video frames. They can be directly extracted from the compression code stream without computational cost and serve as a solid prior for optical flow estimation. Therefore, we propose an optical flow model, MVFlow, which uses motion vectors to improve the speed and accuracy of optical flow estimation for compressed videos. In detail, MVFlow includes a key Motion-Vector Converting Module, which ensures that the motion vectors can be transformed into the same domain of optical flow and then be utilized fully by the flow estimation module. Meanwhile, we construct four optical flow datasets for compressed videos containing frames and motion vectors in pairs. The experimental results demonstrate the superiority of our proposed MVFlow, which can reduce the AEPE by 1.09 compared to existing models or save 52% time to achieve similar accuracy to existing models.
Shili Zhou, Xuhao Jiang, Weimin Tan, Ruian He, Bo Yan 0001
ACM Multimedia5
2023 Abnormal Event Detection via Hypergraph Contrastive Learning
abstract
Abnormal event detection, which refers to mining unusual interactions among involved entities, plays an important role in many real applications. Previous works mostly oversimplify this task as detecting abnormal pair-wise interactions. However, real-world events may contain multi-typed attributed entities and complex interactions among them, which forms an Attributed Heterogeneous Information Network (AHIN). With the boom of social networks, abnormal event detection in AHIN has become an important, but seldom explored task. In this paper, we firstly study the unsupervised abnormal event detection problem in AHIN. The events are considered as star-schema instances of AHIN and are further modeled by hypergraphs. A novel hypergraph contrastive learning method, named AEHCL, is proposed to fully capture abnormal event patterns. AEHCL designs the intra-event and inter-event contrastive modules to exploit self-supervised AHIN information. The intra-event contrastive module captures the pair-wise and multivariate interaction anomalies within an event, and the inter-event module captures the contextual anomalies among events. These two modules collaboratively boost the performance of each other and improve the detection results. During the testing phase, a contrastive learning-based abnormal event score function is further proposed to measure the abnormality degree of events. Extensive experiments on three datasets in different scenarios demonstrate the effectiveness of AEHCL, and the results improve state-of-the-art baselines up to 12.0% in Average Precision (AP) and 4.6% in Area Under Curve (AUC) respectively.
Bo Yan 0001, Cheng Yang 0002, Chuan Shi 0001, Jiawei Liu 0006
SDM1
2023 Self-Supervised Digital Histopathology Image Disentanglement for Arbitrary Domain Stain Transfer
abstract
Diagnosis of cancerous diseases relies on digital histopathology images from stained slides. However, the staining varies among medical centers, which leads to a domain gap of staining. Existing generative adversarial network (GAN) based stain transfer methods highly rely on distinct domains of source and target, and cannot handle unseen domains. To overcome these obstacles, we propose a self-supervised disentanglement network (SDN) for domain-independent optimization and arbitrary domain stain transfer. SDN decomposes an image into features of content and stain. By exchanging the stain features, the staining style of an image is transferred to the target domain. For optimization, we propose a novel self-supervised learning policy based on the consistency of stain and content among augmentations from one instance. Therefore, the process of training SDN is independent on the domain of training data, and thus SDN is able to tackle unseen domains. Exhaustive experiments demonstrate that SDN achieves the top performance in intra-dataset and cross-dataset stain transfer compared with the state-of-the-art stain transfer models, while the number of parameters in SDN is three orders of magnitude smaller parameters than that of compared models. Through stain transfer, SDN improves AUC of downstream classification model on unseen data without fine-tuning. Therefore, the proposed disentanglement framework and self-supervised learning policy have significant advantages in eliminating the stain gap among multi-center histopathology images.
Yu Ling, Weimin Tan, Bo Yan 0001
IEEE Trans. Medical Imaging3
2023 Rethinking and Improving Few-Shot Segmentation From a Contour-Aware Perspective
abstract
Existing few-shot segmentation approaches basically adopt the idea of comparing the semantic prototype vector of the query image and support images, and then obtaining the segmentation result. However, recent studies have shown that a single feature vector in feature map cannot accurately represent pixel-level categories, thus leading to poor segmentation of object boundary and semantic ambiguity. To address this common problem, we propose a novel contour-aware network (CTANet) for few-shot segmentation in this paper. Unlike the usual practice of classifying each pixel separately, CTANet regards all pixels within the same contour as a whole, which can take advantage of the internal consistency of objects to obtain a more accurate representation of category information. To obtain the accurate object contour, our network consists of a contour generation module and a contour refinement module, where the former exploits multiple levels of features to generate a primary contour map and the latter learns to refine the primary contour map. Furthermore, a novel contour-aware mixed loss is proposed to fuse the common BCE loss and our contour-aware loss to supervise the training process on two levels, pixel-level and contour-level. Extensive experiments demonstrate that our CTANet achieves a new state-of-the-art performance on$ \text{PASCAL-5}^{i}$and$ \text{COCO-20}^{i}$. Hopefully, our new perspective could provide more clues for future research on few-shot segmentation. Our code is freely available at:https://github.com/hardtogetA/CTANet.
Weimin Tan, Ganghui Ru, Yueming Jiang, Bo Yan 0001
IEEE Trans. Multim.5
2022 Promoting Single-Modal Optical Flow Network for Diverse Cross-Modal Flow Estimation
abstract
In recent years, optical flow methods develop rapidly, achieving unprecedented high performance. Most of the methods only consider single-modal optical flow under the well-known brightness-constancy assumption. However, in many application systems, images of different modalities need to be aligned, which demands to estimate cross-modal flow between the cross-modal image pairs. A lot of cross-modal matching methods are designed for some specific cross-modal scenarios. We argue that the prior knowledge of the advanced optical flow models can be transferred to the cross-modal flow estimation, which may be a simple but unified solution for diverse cross-modal matching tasks. To verify our hypothesis, we design a self-supervised framework to promote the single-modal optical flow networks for diverse corss-modal flow estimation. Moreover, we add a Cross-Modal-Adapter block as a plugin to the state-of-the-art optical flow model RAFT for better performance in cross-modal scenarios. Our proposed Modality Promotion Framework and Cross-Modal Adapter have multiple advantages compared to the existing methods. The experiments demonstrate that our method is effective on multiple datasets of different cross-modal scenarios.
Shili Zhou, Weimin Tan, Bo Yan 0001
AAAI3
2022 Learning Robust Image-Based Rendering on Sparse Scene Geometry via Depth Completion
abstract
Recent image-based rendering (IBR) methods usually adopt plenty of views to reconstruct dense scene geometry. However, the number of available views is limited in prac-tice. When only few views are provided, the performance of these methods drops off significantly, as the scene geometry becomes sparse as well. Therefore, in this paper, we propose Sparse-IBRNet (SIBRNet) to perform robust IBR on sparse scene geometry by depth completion. The SIBR-Net has two stages, geometry recovery (GR) stage and light blending (LB) stage. Specifically, GR stage takes sparse depth map and RGB as input to predict dense depth map by exploiting the correlation between two modals. As in-accuracy of the complete depth map may cause projection biases in the warping process, LB stage first uses a bias-corrected module (BCM) to rectify deviations, and then ag-gregates modified features from different views to render a novel view. Extensive experimental results demonstrate that our method performs best on sparse scene geometry than re-cent IBR methods, and it can generate better or comparable results as well when the geometric information is dense.1
Shili Zhou, Ri Cheng, Weimin Tan, Bo Yan 0001, Lang Fu
CVPR5
2022 Geometry-Aware Reference Synthesis for Multi-View Image Super-Resolution
abstract
Recent multi-view multimedia applications struggle between high-resolution (HR) visual experience and storage or bandwidth constraints. Therefore, this paper proposes a Multi-View Image Super-Resolution (MVISR) task. It aims to increase the resolution of multi-view images captured from the same scene. One solution is to apply image or video super-resolution (SR) methods to reconstruct HR results from the low-resolution (LR) input view. However, these methods cannot handle large-angle transformations between views and leverage information in all multi-view images. To address these problems, we propose the MVSRnet, which uses geometry information to extract sharp details from all LR multi-view to support the SR of the LR input view. Specifically, the proposed Geometry-Aware Reference Synthesis module in MVSRnet uses geometry information and all multi-view LR images to synthesize pixel-aligned HR reference images. Then, the proposed Dynamic High-Frequency Search network fully exploits the high-frequency textural details in reference images for SR. Extensive experiments on several benchmarks show that our method significantly improves over the state-of-the-art approaches.
Ri Cheng, Bo Yan 0001, Weimin Tan, Chenxi Ma
ACM Multimedia3
2022 Learning Parallax Transformer Network for Stereo Image JPEG Artifacts Removal
abstract
Under stereo settings, the performance of image JPEG artifacts removal can be further improved by exploiting the additional information provided by a second view. However, incorporating this information for stereo image JPEG artifacts removal is a huge challenge, since the existing compression artifacts make pixel-level view alignment difficult. In this paper, we propose a novel parallax transformer network (PTNet) to integrate the information from stereo image pairs for stereo image JPEG artifacts removal. Specifically, a well-designed symmetric bi-directional parallax transformer module is proposed to match features with similar textures between different views instead of pixel-level view alignment. Due to the issues of occlusions and boundaries, a confidence-based cross-view fusion module is proposed to achieve better feature fusion for both views, where the cross-view features are weighted with confidence maps. Especially, we adopt a coarse-to-fine design for the cross-view interaction, leading to better performance. Comprehensive experimental results demonstrate that our PTNet can effectively remove compression artifacts and achieves superior performance than other testing state-of-the-art methods.
Xuhao Jiang, Weimin Tan, Ri Cheng, Shili Zhou, Bo Yan 0001
ACM Multimedia5
2022 Rethinking Super-Resolution as Text-Guided Details Generation
abstract
Deep neural networks have greatly promoted the performance of single image super-resolution (SISR). Conventional methods still resort to restoring the single high-resolution (HR) solution only based on the input of image modality. However, the image-level information is insufficient to predict adequate details and photo-realistic visual quality facing large upscaling factors (×8, ×16). In this paper, we propose a new perspective that regards the SISR as a semantic image detail enhancement problem to generate semantically reasonable HR image that are faithful to the ground truth. To enhance the semantic accuracy and the visual quality of the reconstructed image, we explore the multi-modal fusion learning in SISR by proposing a Text-Guided Super-Resolution (TGSR) framework, which can effectively utilize the information from the text and image modalities. Different from existing methods, the proposed TGSR could generate HR image details that match the text descriptions through a coarse-to-fine process. Extensive experiments and ablation studies demonstrate the effect of the TGSR, which exploits the text reference to recover realistic images.
Chenxi Ma, Bo Yan 0001, Weimin Tan, Siming Chen 0001
ACM Multimedia2
2022 Co-Completion for Occluded Facial Expression Recognition
abstract
The existence of occlusions brings in semantically irrelevant visual patterns and leads to the content loss of occluded regions. Although previous works have made improvement on occluded facial expression recognition, they do not explicitly handle the interference factors aforementioned. In this paper, we propose an intuitive and simplified workflow, Co-Completion, which combines occlusion discarding and feature completion together to reduce the impact of occlusions on facial expression recognition. To protect key features from being contaminated and reduce the dependency of feature completion on occlusion discarding, guidance from discriminative regions is also introduced for joint feature completion. Moreover, we release the COO-RW database for occlusion simulation and refine the occlusion generation protocol for fair comparison in this filed. Experiments on synthetic and realistic databases demonstrate the superiority of our method. The COO-RW database can be downloaded from https://github.com/loveSmallOrange/COO-RW.
Weimin Tan, Ruian He, Yangle Lin, Bo Yan 0001
ACM Multimedia5
2022 Prior embedding multi-degradations super resolution network
Chenxi Ma, Weimin Tan, Bo Yan 0001, Shili Zhou
Neurocomputing3
2021 Perceptual Variousness Motion Deblurring with Light Global Context Refinement
abstract
Deep learning algorithms have made significant progress in dynamic scene deblurring. However, several challenges are still unsettled: 1) The degree and scale of blur in different regions of a blurred image can have a considerable variation in a large range. However, the traditional input pyramid or downscaling-upscaling, is designed to have limited and inflexible perceptual variousness to cope with large blur scale variation. 2) The nonlocal block is proved to be effective in the image enhancement tasks, but it requires high computation and memory cost. In this paper, we are the first to propose a light-weight globally-analyzing module into the image deblurring field, named Light Global Context Refinement (LGCR) module. With exponentially lower cost, it achieves even better performance than the nonlocal unit. Moreover, we propose the Perceptual Variousness Block (PVB) and PVB-piling strategy. By placing PVB repeatedly, the whole method possesses abundant reception field spectrum to be aware of the blur with various degrees and scales. Comprehensive experimental results from the different benchmarks and assessment metrics show that our method achieves excellent performance to set a new state-of-the-art in motion deblurring.1
Weimin Tan, Bo Yan 0001
ICCV3
2021 Organ-Branched CNN for Robust Face Super-Resolution
abstract
In this paper, we present a novel organ-branched CNN method for face super-resolution, named OBC-FSR. It is the first work focusing on facial-part-specific face SR, which consists of a local (facial part) network and a global network. Specifically, local network enhances the five key regions of human faces separately by Wasserstein generative adversarial networks (WGAN). Simultaneously, it also predicts five key regions’ masks, namely, eyes, eyebrows, mouth, nose, and other parts. The output of the local network is obtained by merging super-resolved five key regions. In order to alleviate boundary effects and distortions in the result of local network, our proposed network also includes a global network, which learns the direct mapping between LR and HR human faces. The final HR result of our FSR method is a fusion of the out-puts of local and global networks. Experimental results verify the superior performance of our method compared to the state-of-the-art.
Bahetiyaer Bare, Shili Zhou, Bo Yan 0001, Ke Li 0010
ICME4
2021 Multimodal Asymmetric Dual Learning for Unsupervised Eyeglasses Removal
abstract
Glasses removal is a challenging task due to the diversity of glasses species and the difficulty of obtaining paired datasets. Most existing methods need to build different models for different glasses or expensive paired datasets for supervised training, which lacks universality. In this paper, we propose a multimodal asymmetric dual learning method for unsupervised glasses removal. This method uses large-scale face images with and without glasses for dual feature learning, which does not require intensive manual marking of the glasses. Given a face image with glasses, we aim to generate a glasses-free image preserving the person identity. Thus, in order to make up for the lack of semantic features in the glasses region, we introduce the text description of the target image into the task, and propose a text-guided multimodal feature fusion method. We adaptively select the glasses-free image closest to the target one for better dual feature learning. We also propose a exchange residual loss to generate more precise mask of glasses. Extensive experiments prove that our method can generate real glasses-free images, and better retain the person identity, which can be useful for face recognition.
Bo Yan 0001, Weimin Tan
ACM Multimedia2
2021 Perception-Oriented Stereo Image Super-Resolution
abstract
Recent studies of deep learning based stereo image super-resolution (StereoSR) have promoted the development of StereoSR. However, existing StereoSR models mainly concentrate on improving quantitative evaluation metrics and neglect the visual quality of super-resolved stereo images. To improve the perceptual performance, this paper proposes the first perception-oriented stereo image super-resolution approach by exploiting the feedback, provided by the evaluation on the perceptual quality of StereoSR results. To provide accurate guidance for the StereoSR model, we develop the first special stereo image super-resolution quality assessment (StereoSRQA) model, and further construct a StereoSRQA database. Extensive experiments demonstrate that our StereoSR approach significantly improves the perceptual quality and enhances the reliability of stereo images for disparity estimation.
Chenxi Ma, Bo Yan 0001, Weimin Tan, Xuhao Jiang
ACM Multimedia2
2021 Space-Angle Super-Resolution for Multi-View Images
abstract
The limited spatial and angular resolutions in multi-view multimedia applications restrict their visual experience in practical use. In this paper, we first argue the space-angle super-resolution (SASR) problem for irregular arranged multi-view images. It aims to increase the spatial resolution of source views and synthesize arbitrary virtual high resolution (HR) views between them jointly. One feasible solution is to perform super-resolution (SR) and view synthesis (VS) methods separately. However, it cannot fully exploit the intra-relationship between SR and VS tasks. Intuitively, multi-view images can provide more angular references, and higher resolution can provide more high-frequency details. Therefore, we propose a one-stage space-angle super-resolution network called SASRnet, which simultaneously synthesizes real and virtual HR views. Extensive experiments on several benchmarks demonstrate that our proposed method outperforms two-stage methods, meanwhile prove that SR and VS can promote each other. To our knowledge, this work is the first to address the SASR problem for unstructured multi-view images in an end-to-end learning-based manner.
Ri Cheng, Bo Yan 0001, Shili Zhou
ACM Multimedia3
2021 JROTM: Jointly reinforced object tracking with temporal content reference and motion guidance
Bo Yan 0001, Chuming Lin, Weimin Tan
Neurocomputing2
2021 A Distortion-Aware Multi-Task Learning Framework for Fractional Interpolation in Video Coding
abstract
Motion-compensated prediction adopts fractional-pixel interpolation to obtain the best motion vector. Traditional fixed interpolation filters cannot handle various content and structures well, and existing convolutional neural network based methods cannot fully exploit the distortion characteristics for fractional interpolation. Therefore, this paper proposes a distortion-aware multi-task learning framework (DA-MLF) to perform fractional interpolation. First, a multi-task training framework is proposed to provide the distortion characteristics as complementary information for improving the performance of subsequent interpolation. Then, a uniform interpolation sub-network is proposed to accomplish fractional interpolation, which utilizes the feature fusion module to fuse abundant local features, and the distortion awareness module to capture the multi-scale information of compression artifacts. Furthermore, DA-MLF is integrated into High Efficiency Video Coding (HEVC) test model, and multiple experiments are performed to evaluate the effectiveness of our method. On HEVC testing sequences, DA-MLF achieves 5.0%, 4.0% and 1.7% BD-rate reduction on average compared to the HEVC baseline, under low-delay P, low-delay B and random-access configurations, respectively. The experimental results validate that our framework not only achieves the best interpolation performance but also has the lowest computational complexity compared with state-of-the-art methods.
Liangwei Yu, Liquan Shen, Hao Yang 0008, Xuhao Jiang, Bo Yan 0001
IEEE Trans. Circuits Syst. Video Technol.5
2021 Motion Blur Removal With Quality Assessment Guidance
abstract
Non-uniform blind motion deblurring is a challenging yet fundamental task in the computer vision field, which aims to restore the latent sharp image from the blurry input. Recently, deep-learning-based methods have made significant improvement and progress, on the metric of PSNR. They achieve good results mainly because they adopt Mean Squared Error (MSE) as the optimization objective, in addition to their good model design. However, simple adoption of the PSNR metric and the MSE loss, has non-ignorable disadvantages. PSNR cannot always succeed in assessing the deblurred quality in accordance with the human visual system (HVS), and MSE guides the network to generate over-smoothed images. To address these problems, we are the first to propose the deep-learning-based multi-scale non-reference quality assessment network (Deep DEBLUR-IQA) for assessing the quality of deblurred results. Moreover, a deblurring network of high efficiency is presented. It is more than 50 times faster than other SOTA multi-scale Convolution Neural Network (CNN) methods, with the newly propose Residual Dilated Block (RDB) and Light ResBlock (LRB). The deblurring network's performance can be further boosted with Multiple Dilation Block (MDB), with an acceptable speed decrease. Finally, and most importantly, we are the first to let Deep DEBLUR-IQA guide the deblurring network's optimization. This IQA-guided enhancement paradigm can significantly improve the deblurring results’ subjective quality while achieving excellent PSNR. Experimental results demonstrate that the proposed method performs favorably against state-of-the-art methods quantitatively and qualitatively.
Bo Yan 0001, Chenxi Ma
IEEE Trans. Multim.2
2020 Assessing Eye Aesthetics for Automatic Multi-Reference Eye In-Painting
abstract
With the wide use of artistic images, aesthetic quality assessment has been widely concerned. How to integrate aesthetics into image editing is still a problem worthy of discussion. In this paper, aesthetic assessment is introduced into eye in-painting task for the first time. We construct an eye aesthetic dataset, and train the eye aesthetic assessment network on this basis. Then we propose a novel eye aesthetic and face semantic guided multi-reference eye inpainting GAN approach (AesGAN), which automatically selects the best reference under the guidance of eye aesthetics. A new aesthetic loss has also been introduced into the network to learn the eye aesthetic features and generate highquality eyes. We prove the effectiveness of eye aesthetic assessment in our experiments, which may inspire more applications of aesthetics assessment. Both qualitative and quantitative experimental results show that the proposed AesGAN can produce more natural and visually attractive eyes compared with state-of-the-art methods.
Bo Yan 0001, Weimin Tan, Shili Zhou
CVPR1
2020 Disparity-Aware Domain Adaptation in Stereo Image Restoration
abstract
Under stereo settings, the problems of disparity estimation, stereo magnification and stereo-view synthesis have gathered wide attention. However, the limited image quality brings non-negligible difficulties in developing related applications and becomes the main bottleneck of stereo images. To the best of our knowledge, stereo image restoration is rarely studied. Towards this end, this paper analyses how to effectively explore disparity information, and proposes a unified stereo image restoration framework. The proposed framework explicitly learn the inherent pixel correspondence between stereo views and restores stereo image with the cross-view information at image and feature level. A Feature Modulation Dense Block (FMDB) is introduced to insert disparity prior throughout the whole network. The experiments in terms of efficiency, objective and perceptual quality, and the accuracy of depth estimation demonstrates the superiority of the proposed framework on various stereo image restoration tasks.
Bo Yan 0001, Chenxi Ma, Bahetiyaer Bare, Weimin Tan, Steven C. H. Hoi
CVPR1
2020 Deep Image Quality Assessment Driven Single Image Deblurring
abstract
Motion deblurring is a challenging task in computer vision, which aims to recover the sharp image from the blurry one. Recently, deep learning based methods have made a significant improvement in the metric of PSNR due to the optimazation of the Mean Squared Error (MSE) loss function between deblurring results and sharp images. However, PSNR prefers smooth images and fails to evaluate the sharpness of deblurred images. To solve this problem, we firstly propose the deep learning based multi-scale non-reference quality assessment network (Deep DEBLUR-IQA) for assessing deblurred results. Moreover, we propose an efficient deblurring network, which is over 50 times faster than SOTA multi-scale networks. Finally, the combination of our Deep DEBLUR-IQA network and novel single image deblurring network can significantly increase subjective quality while maintaining satisfactory PSNR. Experiments show that the proposed method outperforms state-of-the-art methods, both qualitatively and quantitatively.
Chenxi Ma, Bo Yan 0001
ICME5
2020 MMFL: Multimodal Fusion Learning for Text-Guided Image Inpainting
abstract
Painters can successfully recover severely damaged objects, yet current inpainting algorithms still can not achieve this ability. Generally, painters will have a conjecture about the seriously missing image before restoring it, which can be expressed in a text description. This paper imitates the process of painters' conjecture, and proposes to introduce the text description into the image inpainting task for the first time, which provides abundant guidance information for image restoration through the fusion of multimodal features. We propose a multimodal fusion learning method for image inpainting (MMFL). To make better use of text features, we construct an image-adaptive word demand module to reasonably filter the effective text features. We introduce a text guided attention loss and a text-image matching loss to make the network pay more attention to the entities in the text description. Extensive experiments prove that our method can better predict the semantics of objects in the missing regions and generate fine grained textures.
Bo Yan 0001, Weimin Tan
ACM Multimedia2
2020 Flow-guided feature enhancement network for video-based person re-identification
Weichao Gong, Bo Yan 0001, Chuming Lin
Neurocomputing2
2020 Effective image restoration for semantic segmentation
Xuejing Niu, Bo Yan 0001, Weimin Tan
Neurocomputing2
2020 Cycle-IR: Deep Cyclic Image Retargeting
abstract
Supervised deep learning techniques have achieved great success in various fields due to getting rid of the limitation of handcrafted representations. However, most previous image retargeting algorithms still employ fixed design principles such as using gradient map or handcrafted features to compute saliency map, which inevitably restricts its generality. Deep learning techniques may help to address this issue, but the challenging problem is that we need to build a large-scale image retargeting dataset for the training of deep retargeting models. However, building such a dataset requires huge human efforts. In this paper, we propose a novel deep cyclic image retargeting approach, called Cycle-IR, to firstly implement image retargeting with a single deep model, without relying on any explicit user annotations. Our idea is built on the reverse mapping from the retargeted images to the given images. If the retargeted image has serious distortion or excessive loss of important visual information, the reverse mapping is unlikely to restore the input image well. We constrain this forward-reverse consistency by introducing a cyclic perception coherence loss. In addition, we propose a simple yet effective image retargeting network (IRNet) to implement the image retargeting process. Our IRNet contains a spatial and channel attention layer, which is able to discriminate visually important regions of input images effectively, especially in cluttered images. Given arbitrary sizes of input images and desired aspect ratios, our Cycle-IR can produce visually pleasing target images directly. Extensive experiments on the standard RetargetMe dataset show the superiority of our Cycle-IR.
Weimin Tan, Bo Yan 0001, Chuming Lin, Xuejing Niu
IEEE Trans. Multim.2
2020 Semantic Segmentation Guided Pixel Fusion for Image Retargeting
abstract
Image retargeting aims to obtain high visual quality of target images for human vision. Through semantic segmentation and understanding of input images, we can better preserve the important semantic regions, so as to effectively improve the performance of image retargeting. Benefit from the successful application of deep neural network in the field of semantic segmentation, in this paper, we propose a novel image retargeting approach using semantic segmentation and pixel fusion. Compared with existing image retargeting methods, our approach can effectively reduce geometric distortion during image retargeting by finely reallocating scaling factors for each region based on the semantic segmentation results. Experimental results demonstrate that the proposed approach can well preserve important semantic regions while leaving less unnatural geometric distortion. Our approach also shows the important role of semantic segmentation and understanding of scenes in image retargeting in detail.
Bo Yan 0001, Xuejing Niu, Bahetiyaer Bare, Weimin Tan
IEEE Trans. Multim.1
2019 Frame and Feature-Context Video Super-Resolution
abstract
For video super-resolution, current state-of-the-art approaches either process multiple low-resolution (LR) frames to produce each output high-resolution (HR) frame separately in a sliding window fashion or recurrently exploit the previously estimated HR frames to super-resolve the following frame. The main weaknesses of these approaches are: 1) separately generating each output frame may obtain high-quality HR estimates while resulting in unsatisfactory flickering artifacts, and 2) combining previously generated HR frames can produce temporally consistent results in the case of short information flow, but it will cause significant jitter and jagged artifacts because the previous super-resolving errors are constantly accumulated to the subsequent frames.In this paper, we propose a fully end-to-end trainable frame and feature-context video super-resolution (FFCVSR) network that consists of two key sub-networks: local network and context network, where the first one explicitly utilizes a sequence of consecutive LR frames to generate local feature and local SR frame, and the other combines the outputs of local network and the previously estimated HR frames and features to super-resolve the subsequent frame. Our approach takes full advantage of the inter-frame information from multiple LR frames and the context information from previously predicted HR frames, producing temporally consistent highquality results while maintaining real-time speed by directly reusing previous features and frames. Extensive evaluations and comparisons demonstrate that our approach produces state-of-the-art results on a standard benchmark dataset, with advantages in terms of accuracy, efficiency, and visual quality over the existing approaches.
Bo Yan 0001, Chuming Lin, Weimin Tan
AAAI1
2019 Scale-Aware Deep Network with Hole Convolution for Blind Motion Deblurring
abstract
The removal of non-uniform motion blur, which is caused by the jittering of the camera and the rapid movement of the target objects, is quite a challenging task. A majority of deblurring works focus on the traditional two-step method: firstly, the motion kernel estimation, and then the energy function minimization. In this paper, we address the problem of single image non-uniform motion blur with deep convolutional neural network (CNN) in a decent end-to-end manner. We proposed a brand-new convolution architecture named as 'hole convolution', of which the kernel takes a rectangular ring of the neighbors to the center pixel into computation and thus the reception field is greatly expanded. Moreover, we present a scale-aware convolutional neural network to recover the latent sharp image. With the carefully designed units, the hole convolution layers as well as the exquisite network structure, our proposed deep CNN can eliminate blur from various scales without upsampling and downsampling. Experimental results show that our method can effectively restore the latent sharp image, which outperforms the state-of-the-art method by a large margin.
Ke Li 0010, Bo Yan 0001
ICME3
2019 A Multi-level Aggregated Network for Image Restoration
abstract
Recently, significant progress has been witnessed in image restoration benefited from the development of deep convolutional neural networks (CNN). However, we note that many state-of-the-art image restoration networks can be unfolded as a one-level architecture, which is constructed by stacking multiple convolution layers or blocks. As the depth of network grows, the information flow is weakened. And, existing skip connection has restricted ability to pass the previous state to the latter layers in the network. Based on above observations, we explore a multi-level aggregated network (MLAN) to fully exploit features of deeper layers. The proposed MLAN can extract more features by merging layers at different levels progressively, and can better aggregate the features by the augmented connection manner. Experimental results demonstrate a satisfactory performance of the proposed model on different image restoration tasks.
Chenxi Ma, Weimin Tan, Bahetiyaer Bare, Bo Yan 0001
ICME4
2019 RDGAN: Retinex Decomposition Based Adversarial Learning for Low-Light Enhancement
abstract
Pictures taken under the low-light condition often suffer from low contrast and loss of image details, thus an approach that can effectively improve low-light images is demanded. Traditional Retinex-based methods assume that the reflectance components of low-light images keep unchanged, which neglect the color distortion and lost details. In this paper, we propose an end-to-end learning-based framework that first decomposes the low-light image and then learns to fuse the decomposed results to obtain the high quality enhanced result. Our framework can be divided into a RDNet (Retinex Decomposition Network) for decomposition and a FENet (Fusion Enhancement Network) for fusion. Specific multi-term losses are respectively designed for the two networks. We also present a new RDGAN (Retinex Decomposition based Generative Adversarial Network) loss, which is computed on the decomposed reflectance components of the enhanced and the reference images. Experiments demonstrate that our approach is good at color and detail restoration, which outperforms other state-of-the-art methods.
Weimin Tan, Xuejing Niu, Bo Yan 0001
ICME4
2019 Real-time video super-resolution via motion convolution kernel estimation
Bahetiyaer Bare, Bo Yan 0001, Chenxi Ma, Ke Li 0010
Neurocomputing2
2019 Deep Objective Quality Assessment Driven Single Image Super-Resolution
abstract
Single-image super-resolution (SISR) is a classic problem in the image processing community, which aims at generating a high-resolution image from a low-resolution one. In recent years, deep learning based SISR methods emerged and achieved a performance leap than previous methods. However, because the evaluation metrics of SISR methods is peak signal-to-noise ratio (PSNR), previous methods usually choose L2-norm as the loss function. This leads to a significant improvement in the final PSNR value but little improvement in perceptual quality. In this paper, in order to achieve better results in both perceptual quality and PSNR values, we propose an objective quality assessment driven SISR method. First, we propose a novel full-reference image quality assessment approach for SISR and employ it as a loss function, namely super-resolution image quality assessment (SR-IQA) loss. Then, we combine SR-IQA loss with L2-norm to guide our proposed SISR method to achieve better results. Besides that, our proposed SISR method consists of several proposed highway units. Furthermore, in order to verify the generalization ability of our new kind of loss function, we integrate SR-IQA loss to generative adversarial networks based SR method and achieve better perceptual quality. Experimental results prove that our proposed SISR method achieves better performance than other methods both qualitatively and quantitatively in most of the cases.
Bo Yan 0001, Bahetiyaer Bare, Chenxi Ma, Ke Li 0010, Weimin Tan
IEEE Trans. Multim.1
2019 Naturalness-Aware Deep No-Reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) is a non-trivial task, because it is hard to find a pristine counterpart for an image in real applications, such as image selection, high quality image recommendation, etc. In recent years, deep learning-based NR-IQA methods emerged and achieved better performance than previous methods. In this paper, we present a novel deep neural networks-based multi-task learning approach for NR-IQA. Our proposed network is designed by a multi-task learning manner that consists of two tasks, namely, natural scene statistics (NSS) features prediction task and the quality score prediction task. NSS features prediction is an auxiliary task, which helps the quality score prediction task to learn better mapping between the input image and its quality score. The main contribution of this work is to integrate the NSS features prediction task to the deep learning-based image quality prediction task to improve the representation ability and generalization ability. To the best of our knowledge, it is the first attempt. We conduct the same database validation and cross database validation experiments on LIVE1, TID20132, CSIQ3, LIVE multiply distorted image quality database (LIVE MD)4, CID20135, and LIVE in the wild image quality challenge (LIVE challenge)6databases to verify the superiority and generalization ability of the proposed method. Experimental results confirm the superior performance of our method on the same database validation; our method especially achieves 0.984 and 0.986 on the LIVE image quality assessment database in terms of the Pearson linear correlation coefficient (PLCC) and Spearman rank-order correlation coefficient (SROCC), respectively. Also, experimental results from cross database validation verify the strong generalization ability of our method. Specifically, our method gains significant improvement up to 21.8% on unseen distortion types.
Bo Yan 0001, Bahetiyaer Bare, Weimin Tan
IEEE Trans. Multim.1
2018 Feature Super-Resolution: Make Machine See More Clearly
abstract
Identifying small size images or small objects is a notoriously challenging problem, as discriminative representations are difficult to learn from the limited information contained in them with poor-quality appearance and unclear object structure. Existing research works usually increase the resolution of low-resolution image in the pixel space in order to provide better visual quality for human viewing. However, the improved performance of such methods is usually limited or even trivial in the case of very small image size (we will show it in this paper explicitly). In this paper, different from image super-resolution (ISR), we propose a novel super-resolution technique called feature super-resolution (FSR), which aims at enhancing the discriminatory power of small size image in order to provide high recognition precision for machine. To achieve this goal, we propose a new Feature Super-Resolution Generative Adversarial Network (FSR-GAN) model that transforms the raw poor features of small size images to highly discriminative ones by performing super-resolution in the feature space. Our FSR-GAN consists of two subnetworks: a feature generator network G and a feature discriminator network D. By training the G and the D networks in an alternative manner, we encourage the G network to discover the latent distribution correlations between small size and large size images and then use G to improve the representations of small images. Extensive experiment results on Oxford5K, Paris, Holidays, and Flick100k datasets demonstrate that the proposed FSR approach can effectively enhance the discriminatory ability of features. Even when the resolution of query images is reduced greatly, e.g., 1/64 original size, the query feature enhanced by our FSR approach achieves surprisingly high retrieval performance at different image resolutions and increases the retrieval precision by 25% compared to the raw query feature.
Weimin Tan, Bo Yan 0001, Bahetiyaer Bare
CVPR2
2018 A Deep Learning Based No-Reference Image Quality Assessment Model for Single-Image Super-Resolution
abstract
Single-image super-resolution (SISR) is a very important and classic problem of the computer vision community. Although a lot of SISR methods have been proposed, few studies have been conducted to address the quality assessment of SISR methods. In this paper, we proposed a deep learning based no-reference image quality assessment (NR-IQA) model for SISR. We took small patches from images to form our training set and labeled them with different scores. With the aid of well-designed architecture and training strategy, our method achieved a performance leap than state-of-the-art methods. Experimental results proved the generalizability and the effectiveness of the proposed model.
Bahetiyaer Bare, Ke Li 0010, Bo Yan 0001, Bailan Feng, Chunfeng Yao
ICASSP3
2018 Face Hallucination Based on Key Parts Enhancement
abstract
Face hallucination aims to generate a high resolution face from a low resolution one. Generic super resolution methods can not solve this problem well, because human face has a strong structure. With the rapid development of the deep learning technique, some convolutional neural networks (CNNs) models for face hallucination emerged and achieved state-of-the-art performance. In this paper, we proposed a five-branch network based on five key parts of human face. Each branch of this network aims to generate a high resolution key part. The final high resolution face is the combination of the five branches' output. In addition, we designed a gated enhance unit (GEU) and cascade it to form our network architecture. Experimental results confirm that our method can generate pleasing high resolution faces.
Ke Li 0010, Bahetiyaer Bare, Bo Yan 0001, Bailan Feng, Chunfeng Yao
ICASSP3
2018 HNSR: Highway Networks Based Deep Convolutional Neural Networks Model for Single Image Super-Resolution
abstract
Convolutional neural networks (CNNs) have been widely used in computer vision community. Single image super-resolution (SISR) is a classic computer vision problem, which aims to output a high-resolution image from a low-resolution one. In recent years, CNNs-based SISR methods emerged and achieved a performance leap. In this paper, we present a highly accurate deep CNNs model for SISR. Inspired by the ideas in highway networks, we propose a highway unit and cascade highway units to ensemble our model. Furthermore, we employ structural similarity index (SSIM) as a part of loss function to enhance the accuracy of trained deep CNNs model. Experimental results show that our proposed model outperforms other state-of-the-art methods.
Ke Li 0010, Bahetiyaer Bare, Bo Yan 0001, Bailan Feng, Chunfeng Yao
ICASSP3
2018 Deep Residual Network for Enhancing Quality of the Decoded Intra Frames of Hevc
abstract
High Efficiency Video Coding (HEVC) has been widely used to encode video sequences and output streams for its outstanding performance on compression. However, the lossy compression process still results in blocking, blurring and ring effects, which are quite significant at low bit-rates. In this paper, we propose a post-processing method based on a residual convolutional neural network (CNN) to effectively improve the visual quality of the decoded intra frames of HEVC. By learning the spatial mapping between high and low quality image patches, our approach can effectively predict lost image details. Since the proposed approach only works on decoded frames, it does not require to modify the original HEVC, i.e., the enhancement can be done in parallel with HEVC baseline, which makes it flexible to be included into a decoding system. Experimental results show that the proposed approach outperforms the state-of-the-arts by a large margin and makes a 5.5% bit-rate reduction on average compared to HEVC baseline. Furthermore, the PSNR measures of intra frames increase about 0.395 dB when encoded at high quality parameters (QP).
Weimin Tan, Bo Yan 0001
ICIP3
2018 Foreground Detection in Surveillance Video with Fully Convolutional Semantic Network
abstract
Foreground detection is an important part of surveillance video analysis, and also has challenges. For example, the classical methods are difficult to distinguish the foreground, which is similar to the background. In recent years, Convolutional Neural Networks (CNNs) have been widely used in image processing and achieved better performance. In this paper, we proposed an efficient deep Fully Convolutional Semantic Networks (FCSN) model for foreground detection in surveillance video. Our model aimed at learning the global differences between the video frame and the background image, and the semantic information by utilizing the pre-trained weights on semantic segmentation. In the experiment, unlike other related work, we proposed a reasonable method, which is able to avoid overfitting results to construct training data with 20 videos and test data with 6 videos on the dataset of 2014 ChangeDetection.net (CDnet 2014). Experimental results verified that our model outperforms the state-of-the-art methods in the foreground detection of surveillance video.
Chuming Lin, Bo Yan 0001, Weimin Tan
ICIP2
2018 Beyond Visual Retargeting: A Feature Retargeting Approach for Visual Recognition and Its Applications
abstract
The popularity of mobile applications has greatly enriched and facilitated our lives. However, the rapid increase of digital images and the problem of narrow bandwidth of the wireless network call for an appropriate approach to reduce the amount of data transmitted over the wireless network (i.e., low bit-rate transmission) while ensuring high recognition accuracy at the cloud. We propose a simple and effective feature retargeting (FR) approach for retargeting an image while preserving the representative local features (e.g., SIFT, SURF, and BRIEF) in the image. Our feature retargeting approach aims at low bit-rate visual recognition instead of high-quality visual perception that visual retargeting methods dedicate to. Our algorithm consists of two key novelties: estimating feature saliency and retargeting image: Estimating feature saliency focuses on predicting the relative importance of different features in an image by analyzing uniqueness in a specific context; Retargeting image aims at finding the optimal resolution for the retargeted image to maximize feature-saliency energy. We evaluate the proposed approach for two different applications in three large data sets and observe that our FR approach consistently outperforms state-of-the-art retargeting algorithms, resulting in both higher precision and lower bit-rates. We also demonstrate that even when the resolution of source image is reduced greatly, e.g., 1/7 original size, our algorithm produces superior results as compared with other approaches.
Weimin Tan, Bo Yan 0001, Chuming Lin
IEEE Trans. Circuits Syst. Video Technol.2
2017 Salient Object Detection via Google Image Retrieval
Weimin Tan, Bo Yan 0001
ICIG (1)2
2017 An accurate deep convolutional neural networks model for no-reference image quality assessment
abstract
The goal of image quality assessment (IQA) is to use computational models to measure the consistency between image quality and subjective evaluations. In recent years, convolutional neural networks (CNNs) have been widely used in image processing community and have achieved performance leaps than non CNNs-based methods. In this work, we describe an accurate deep CNNs model for no-reference IQA. Taking image patches as input, our deep CNNs model achieves an end-to-end method without any handcrafted features and pre-processing procedures that are employed by previous no-reference IQA methods. The proposed model consists of six convolutional layers, two fully connected layers, one max pooling layer and two sum layers. The experimental results verify that our model outperforms the state-of-the-art no-reference IQA methods and most of the full-reference IQA metrics.
Bahetiyaer Bare, Ke Li 0010, Bo Yan 0001
ICME3
2017 An efficient deep convolutional neural networks model for compressed image deblocking
abstract
Convolutional neural networks (CNNs) have been widely used in image processing community. Image deblocking is a post-processing strategy, which aims to reduce the visually annoying blocking artifacts that are caused by block-based transform coding at low bit rates. In recent years, CNNs based methods have been proposed to solve this classic image processing problem. In this paper, we present an efficient deep C-NNs model for image deblocking. Our model can well alleviate the conflict between bit reduction and quality preservation by taking local small patches into consideration. Our trained model can be used to deblock lossy compressed images with different quality factors. The proposed model can be easily integrated into the existing codecs as a post-processing procedure without changing the codec. Experimental results verify that our proposed model outperforms the state-of-the-art methods in both the objective quality and the perceptual quality.
Ke Li 0010, Bahetiyaer Bare, Bo Yan 0001
ICME3
2017 Salient object detection via multiple saliency weights
Weimin Tan, Bo Yan 0001
Multim. Tools Appl.2
2017 Learning quality assessment of retargeted images
Bo Yan 0001, Bahetiyaer Bare, Ke Li 0010, Alan C. Bovik
Signal Process. Image Commun.1
2017 Codebook Guided Feature-Preserving for Recognition-Oriented Image Retargeting
abstract
Traditional image resizing methods, such as uniform scaling and content-aware image retargeting, are designed to preserve the visually salient contents of an image while resizing it. In this paper, we propose a novel image resizing approach called recognition-oriented image retargeting. Its goal is to preserve the distinctive local features for recognition instead of the traditional visual saliency during resizing. Moreover, we also apply our approach to image matching and image retrieval applications to verify its performance. Meanwhile, using our approach to these applications is able to solve some of the challenging problems in their fields. In image matching application, we find that our approach shows promising preservation of local feature descriptors. In image retrieval task, extensive experiments on Oxford5K, Holidays, Paris, and Flickr100k data sets demonstrate that our approach consistently outperforms other image retargeting methods by large margins in the aspects of retrieval precision and query bits.
Bo Yan 0001, Weimin Tan, Ke Li 0010, Qi Tian 0001
IEEE Trans. Image Process.1
2016 A survey on high coherence visual media retargeting: recent advances and applications
Weimin Tan, Bo Yan 0001
Frontiers Comput. Sci.2
2016 An Effective Video Synopsis Approach with Seam Carving
abstract
With a growth of surveillance cameras, the amount of captured videos expands. Manually analyzing and retrieving surveillance video is labor intensive and expensive. It would be much more convenient to generate a video digest, with which we can view the video in a fast and motion-preserving way. In this paper, we propose a novel video synopsis approach to generate condensed video, which uses an object tracking method for extracting important objects. This method will generate video tubes and a seam carving method to condense the original video. Experimental results demonstrate that our proposed method can achieve a high condensation rate while preserving all the important objects of interest. Therefore, this approach can enable users to view the surveillance video with great efficiency.
Ke Li 0010, Bo Yan 0001, Hamid Gharavi
IEEE Signal Process. Lett.2
2016 Image Retargeting for Preserving Robust Local Feature: Application to Mobile Visual Search
abstract
With the sharp increasing of mobile devices, conducting search on mobile devices becomes pervasive, and one of the most popular applications is mobile visual search. To achieve low bit-rate visual search, most of the existing works focus on addressing local descriptor coding and BoW histogram compression . In this paper, we extend the concept of image retargeting and propose a new image resizing approach that is devoted to preserving the robust local features in the query image while resizing it. Based on the extended concept, we introduce a novel mobile-visual-search scheme that conducts the proposed approach to reduce the size of the query image for achieving low bit-rate visual search. Extensive experiments on Oxford 5 K and Flickr 100k datasets show that our approach obtains superior retrieval performance than state-of-the-art image resizing approaches at the similar query size; meanwhile, it is cost effective in terms of processing time.
Weimin Tan, Bo Yan 0001, Ke Li 0010, Qi Tian 0001
IEEE Trans. Multim.2
2015 Pixel fusion based stereo image retargeting
abstract
Image retargeting attempts to adapt images to different devices while preserving the salient contents. Most existing methods address retargeting of a single image. In this paper, we propose a novel image retargeting method for resizing a pair of stereo images. Naively retargeting each image independently will distort the geometric structure and will impair the perception of the 3-D structure of the scene. We introduce an extension to the 2-D image retargeting method that works on a pair of stereo images. We demonstrate the performance of our method on a number of challenging indoor and outdoor stereo images. Experimental results show that our method is able to provide visually comfortable resized images when the resizing ratio is relatively high.
Bahetiyaer Bare, Ke Li 0010, Bo Yan 0001, Xiaoyu Qi, Hamid Gharavi
ICME3
2015 Seam Searching-Based Pixel Fusion for Image Retargeting
abstract
This paper presents a new image retargeting method based on seam searching and pixel fusion. First, the importance map of the original image, which is represented by the saliency of every pixel, is generated. Second, all the pixels in the original image are divided into W (width of the original image) groups with the help of seam searching. Then, it performs inter-row coherence filtering on importance map to maintain the spatial coherence, and directly utilizes the filtered importance map to generate the scaling map. Finally, after pixel fusion with the scaling map, the image can be retargeted effectively. Experimental results show that our method is able to provide visually pleasing resized images.
Bo Yan 0001, Ke Li 0010, Xiaochu Yang, Tingxiao Hu
IEEE Trans. Circuits Syst. Video Technol.1
2014 Efficient DCT-based image retargeting in compressed domain
abstract
With the increasing requirement of efficient image retargeting, many algorithms have been proposed for adapting images contents to various display settings. However, most of these algorithms work in spatial domain of raw images. Since images are mostly stored in DCT-based compressed format such as JPEG, it will be very attractive to realize the image retargeting in compressed domain. In this paper, we propose a new low complexity DCT-based image retargeting method, which is completely performed in compressed domain. This proposed algorithm is based on the construction of block level importance map and calculation of block level forward energy. Experimental results prove that our proposed algorithm is able to significantly reduce image decoding complexity as well as image retargeting complexity, while providing the satisfying image retargeting quality compared with other methods.
Ke Li 0010, Bo Yan 0001, Liu Liu 0006, Kairan Sun
ICME2
2014 Learning to Assess Image Retargeting
abstract
Content-aware image retargeting enables images to fit different devices with various aspect ratios while preserving salient contents. Meanwhile, assessing the quality of image retargeting and unifying both subjective and objective evaluation have become a prominent challenge. In this paper, we propose an image quality assessment based on Radial Basis Function (RBF) neural network. We propose a new feature of image retargeting evaluation, which adapts structural similarity (SSIM) and saliency. By also including other existing features, we build a neural network to assess the quality of the retargeted image. The neural network is trained to combine the above-mentioned features. The accuracy of our proposed assessment is verified by simulations and it possesses huge practical significance.
Bahetiyaer Bare, Ke Li 0010, Bo Yan 0001
ACM Multimedia4
2014 Efficient Image Retargeting via Adaptive Pixel Fusion
abstract
This paper presents a new image retargeting method, which is able to provide both high efficiency and quality. Firstly, our method calculates the image resizing factors in both horizontal and vertical directions. Then, it resizes the image in these both directions sequentially in order to achieve the target aspect ratio with the help of our proposed mapping functions. Finally, the reconstructed image can be uniformly scaled to the target resolution. Experimental results show that our method is not only cost effective in terms of computational resource, but also provides visually pleasing resized images.
Bo Yan 0001, Xiaochu Yang, Ke Li 0010
ACM Multimedia1
2013 Efficient seam carving for object removal
abstract
This paper introduces a new object removal approach for images based on discontinuous seam carving. Existing seam carving based object removal methods generally are time-consuming and oftentimes cause image distortion or cutting off many more seams than is necessary. In order to solve these limitations, our proposed method only considers the energy of all the pixels outside the target region. Firstly, based on the discontinuous seam carving, we calculate the energy map of both directions, up-down and bottom-up. Then we carve the seam respectively from the upper bound of the target region to up, from the lower bound of the target region to down and the middle part within the region. Experimental results prove that our proposed method outperform others in terms of efficiency and image quality significantly.
Bo Yan 0001, Yiqi Gao, Kairan Sun
ICIP1
2013 Effective retargeting for image coding
abstract
An effective retargeting scheme is presented in this paper in order to improve the efficiency of image coding. Firstly, we propose to calculate the information loss in image retargeting based on learning. Then before encoding an image, we divide the image into several grids and retarget each grid according to our learning results. The retargeted image can be encoded by the traditional image codec. In the receiver side, with the help of the side information, the compressed image can be decoded and reconstructed. Simulation results show that the proposed scheme is able to provide much higher compression performance compared with the JPEG method.
Tingxiao Hu, Bo Yan 0001
ISCAS2
2013 An efficient framework for image/video inpainting
Miaohui Wang, Bo Yan 0001, King Ngi Ngan
Signal Process. Image Commun.2
2013 Matching-Area-Based Seam Carving for Video Retargeting
abstract
This paper presents a video retargeting method considering both spatial and temporal coherence for resizing videos. Our algorithm is based on a novel matching-area-based temporal energy adjustment that allows per-frame seam carving to remove the optimal pixels to achieve spatially and temporally continuous resized videos. The temporal energy adjustment allows the seam to track the object it previously carved, and avoid carving the seam on different objects in two consecutive frames to achieve both spatial and temporal coherence. Our method outperforms other state-of-the-art retargeting systems, as demonstrated in the results and widely supported by the conducted user study.
Bo Yan 0001, Kairan Sun, Liu Liu 0006
IEEE Trans. Circuits Syst. Video Technol.1
2012 Lowcomplexity content-aware image retargeting
abstract
Image retargeting plays a more and more significant role recently, thanks to the escalating diversity of display devices. In this paper, we present a novel low complexity content-aware image retargeting method, which can provide both high efficiency and quality. Specifically, the importance map is used as the basis of our method. The important regions tend to preserve its original size. Our method performs inter-row coherence filtering on importance maps in order to maintain the spatial coherence, and then directly utilizes the filtered importance maps to generate the scaling map. Experimental results show the proposed algorithm improves the efficiency significantly compared with other existing methods. At the same time, the resized image quality of our algorithm is as good as, if not better than, that of the other methods. As a result, our method possesses huge practical significance.
Kairan Sun, Bo Yan 0001, Yiqi Gao
ICIP2
2012 Joint Complexity Estimation of I-Frame and P-Frame for H.264/AVC Rate Control
abstract
Rate control plays a significant role for high quality video coding. This paper presents a rate control method for H.264/AVC, using a new bit allocation scheme for both I-frame and P-frame. This scheme is based on our proposed frame complexity measurement and estimation model. The measurement model considers not only the absolute complexity of I-frame and P-frame, respectively, but also the complexity relationship between I-frame and P-frame. Our proposed bit allocation scheme is able to allocate bit rate more efficiently for both I-frame and P-frame in order to achieve a relatively steady visual quality. Experimental results show that our proposed method is capable of providing more stable video quality than other existing methods.
Bo Yan 0001, Kairan Sun
IEEE Trans. Circuits Syst. Video Technol.1
2012 Efficient Frame Concealment for Depth Image-Based 3-D Video Transmission
abstract
In depth image-based 3-D video transmission, the compressed video stream is very likely to be corrupted by channel errors. Due to the high compression ratio of H.264/AVC, it is often common that an entire coded picture is packetized into one packet. Thus the loss of a packet may result in the loss of the whole video frame. Currently, most of the frame concealment methods are mainly for 2-D video transmission. In this paper, we have proposed an efficient frame concealment algorithm for depth image-based 3-D video transmission, which is able to provide accurate estimation for the motion vectors of the lost frame with the help of the depth information. Simulation results show that it is highly effective and significantly outperforms other existing frame recovery methods by up to 2.91 dB.
Bo Yan 0001
IEEE Trans. Multim.1
2011 Efficient P-frame complexity estimation for frame layer rate control of H.264/AVC
abstract
Rate control plays a significant role for high quality video coding. This paper presents a rate control method for H.264/AVC using a new bit allocation scheme for P-frame. This scheme is based on our proposed P-frame complexity measurement and estimation model. The measurement model considers not only the absolute complexity of the picture via average gradient, but also the relative histogram difference between two frames. Our proposed bit allocation scheme is able to allocate bitrate more efficiently for P-frame, which means allocating more bits to high-complexity frames and less bits to low-complexity ones. Experimental results show that our proposed method is able to achieve better quality than the original JVT-G012 and other existing methods.
Kairan Sun, Bo Yan 0001
ICIP2
2011 Selective pixel interpolation for spatial error concealment
abstract
This paper proposes an effective algorithm for spatial error concealment with accurate edge detection and partitioning interpolation. Firstly, a new method is used for detecting possible edge pixels and their matching pixels around the lost block. Then, the true edge lines can be determined, with which the lost block is partitioned. Finally, based on the partition result, each lost pixel can be interpolated with correct reference pixels, which are in the same region with the lost pixel. Experimental results show that the proposed spatial error concealment method is obviously superior to the previous methods for different sequences by up to 4.04 dB.
Yi Ge, Bo Yan 0001, Kairan Sun, Hamid Gharavi
MMSP2
2010 Pyramid model based Down-sampling for image inpainting
abstract
Image inpainting is a useful and powerful technique for automatically restoring or removing objects in films and damaged pictures. In the last ten years, many excellent inpainting algorithms have been proposed after Bertalmio et al. [1]. However, no paper systematically and theoretically analyzes the factors, which limit the performances of existing algorithms. Based on extensive experiments, we firstly construct a universal framework for image inpainting, which contains three crucial factors—Area, Shape, and Perimeter (ASP). Then we propose a Pyramid model based Down-sampling Inpainting (PDI) model according to the ASP principles. Experimental results show that the performances of existing methods can be tremendously improved after incorporating the PDI model.
Miaohui Wang, Bo Yan 0001, Hamid Gharavi
ICIP2
2010 A Hybrid Frame Concealment Algorithm for H.264/AVC
abstract
In packet-based video transmissions, packets loss due to channel errors may result in the loss of the whole video frame. Recently, many error concealment algorithms have been proposed in order to combat channel errors; however, most of the existing algorithms can only deal with the loss of macroblocks and are not able to conceal the whole missing frame. In order to resolve this problem, in this paper, we have proposed a new hybrid motion vector extrapolation (HMVE) algorithm to recover the whole missing frame, and it is able to provide more accurate estimation for the motion vectors of the missing frame than other conventional methods. Simulation results show that it is highly effective and significantly outperforms other existing frame recovery methods.
Bo Yan 0001, Hamid Gharavi
IEEE Trans. Image Process.1
2009 Lagrangian Multiplier Based Joint Three-Layer Rate Control for H.264/AVC
abstract
Lagrangian multiplier (LM) based mode decision is one of the most important technologies in standard H.264/AVC encoder. Based on LM theory, this paper presents a joint three-layer (JTL) model for H.264/AVC rate control. At macroblock (MB) level, we dynamically revise LM for each of MBs by its estimated complexities, which is able to select a better coding mode than the current scheme with the constant LM adopted in H.264/AVC. At frame level, a more flexible and effective quantization parameter (QP) adjustment scheme is designed for I-frame to avoid buffer overflow or underflow. In addition, we also present a new target bits allocation scheme in group of picture (GOP) level. Experimental results show that our JTL model can not only significantly improve the video quality with the average PSNR gain up to 0.97 dB, but also provide a more stable buffer occupancy with respect to other existing rate control methods.
Miaohui Wang, Bo Yan 0001
IEEE Signal Process. Lett.2
2009 Adaptive Distortion-Based Intra-Rate Estimation for H.264/AVC Rate Control
abstract
Rate distortion of Intra-frame (I-frame) plays a significant role in controlling video quality for H.264/AVC. This letter presents an adaptive distortion-based Intra-rate estimation (ADIE) algorithm for H.264/AVC rate control. In this algorithm, a new rate control model is established based on the distortion by taking image complexity, buffer status, and scene change into consideration. After adaptively updating this model, our proposed algorithm is capable of providing more accurate estimation for the quantization step (Qstep) of the I-frame than other conventional methods. Experimental results show that the proposed method can significantly improve video quality (up to 1.62 dB average PSNR performance), while achieving an accurate output bit rate. More importantly, it provides a more stable visual quality which is a result of keeping the buffer occupancy steadier than other existing rate control algorithms.
Bo Yan 0001, Miaohui Wang
IEEE Signal Process. Lett.1
2008 Efficient error concealment for the whole-frame loss based on H.264/AVC
abstract
For low bitrate video communications, each video frame usually fills the payload of a single network packet. In this situation, the loss of a packet may result in loosing the entire video frame. Currently, most existing error concealment algorithms can only deal with the loss of macroblocks and are not able to conceal the whole missing frame. In this paper, we have proposed a new hybrid motion vector extrapolation (HMVE) algorithm to recover the whole missing frame. The proposed algorithm is capable of estimating the missing motion vectors with much greater accuracy than other existing methods. Experimental results show that it is highly effective and significantly outperforms other existing frame recovery methods.
Bo Yan 0001, Hamid Gharavi
ICIP1
2006 Multi-Path Multi-Channel Routing Protocol
abstract
In this paper we present a DSR-based multi-path routing protocol, which has been developed for transmission of multiple description coded (MDC) packets in wireless ad-hoc network environments. The protocol is designed to eliminate co-channel interference between multiple routes from source to destination by assigning a different frequency band to each route. In the route discovery process we use three metrics to select the best multiple routes. These are hop count, power budget, and the number of joint nodes between the different routes. For continuous media communications we show that in order to effectively benefit from the advantages associated with multipath diversity routing, it is important to use a multi-channel protocol such as the one developed here
Bo Yan 0001, Hamid Gharavi
NCA1
2006 Power Control in Multihop CSMA
abstract
This paper aims at improving the power efficiency of the CSMA/CA protocol for transmission of multimedia information over multihop wireless channels. Using a distance dependent propagation model, we present a power control scheme, which is based on the receiver sensitivity adjustment mechanism. The receiver sensitivity approach aims at exploiting a tradeoff between the interference and the contention in accessing the shared medium. For real-time traffic, we show that by controlling the receiver sensitivity threshold we can significantly improve the multihop link performance.
Bo Yan 0001, Hamid Gharavi
WOWMOM1