Huanjing Yue

dblp:119/0275 · DBLP profile ↗
← Back
71ranked-venue papers
23as first author
46since 2021 · last 2026
0000-0003-2517-9783ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 20 first-author · 33 since 2021Artificial intelligence and machine learning · 22 · 7 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
abstract
Abstract Facial expression recognition (FER) has emerged as an important research topic in recent years. However, current FER paradigms face challenges in generalization, lack semantic information aligned with natural language, and struggle to process both images and videos within a unified framework. Multimodal Large Language Models (MLLMs) have recently achieved success, offering advantages in addressing these issues and potentially overcoming the limitations of current FER paradigms. Nonetheless, directly applying pre-trained MLLMs to FER remains challenging due to insufficient instruction datasets and the inability of vision encoders to extract fine-grained facial information. Our zero-shot evaluations of existing open-source MLLMs on FER reveal a significant performance gap compared to GPT-4V and state-of-the-art supervised methods. In this paper, we aim to enhance MLLMs’ capabilities in understanding facial expressions. We first introduce a facial expression recognition instruction dataset ( FERID ), which has 376k category instructions and 339k conversational instructions. We then propose a novel MLLM, named EMO-LLaMA , which incorporates facial priors from a pretrained facial analysis network to enhance its understanding of human facial information. Specifically, we design a Face Info Mining module to extract both global and local facial information. Furthermore, we utilize a handcrafted prompt to introduce age-gender-race attributes, considering the emotional differences across diverse human groups. Extensive experiments show that EMO-LLaMA achieves results comparable to or competitive with SOTA on both static and dynamic FER datasets. The instruction dataset and code will be available at https://github.com/xxtars/EMO-LLaMA .
Bohao Xing, Zitong Yu, Xin Liu 0012, Kaishen Yuan, Qilang Ye, Weicheng Xie 0001, Huanjing Yue, Heikki Kälviäinen
Int. J. Comput. Vis.7
2026 DIC-DDA: Learned Asymmetric Distributed Image Compression via Dual Domain Alignment
abstract
Multi-view or stereo image compression is an essential technology in 3D related applications. Due to the overlap between different views, exploring their correlations can help improve the compression rate. However, the computing complexity of joint encoding at the encoding side is a heavy burden for terminal encoders. To solve this problem, the learned Distributed Image Coding (DIC), which only uses the correlated view (namely the side image, SI) in the decoder side, has gained much attention in recent years. In this work, we explore asymmetric DIC where one view is selected as the SI and is losslessly compressed. The key problem in learned asymmetric DIC is alignment between the transmitted low-quality target image and high-quality SI. Previous methods usually adopt patch-level alignment with the offset index obtained from degraded (via re-encoded and decoded) SI and the decoded target image, which hinders the alignment accuracy. In this work, we propose a dual domain alignment strategy, which includes degraded domain and fused domain pixel-wise offset estimation. For the degraded domain alignment, we estimate the offset between the degraded SI feature and the degraded target image feature, which eliminates the difficulties in cross-domain matching. For the fused-domain alignment, we observe that the fusion result of degraded target feature and aligned side image feature implicitly contains fine-scale disparity information. Therefore, we estimate the fine-scale offset from the fusion result, which helps refine the degraded domain offsets. We further propose a selective enhancement module to repair the mismatched region in the aligned feature. Extensive experiments on three datasets demonstrate the superiority of our proposed method, outperforming the second-best method by 16% in terms of average BD-rate reduction on the KITTI Stereo dataset. Our code is available at https://github.com/lixianghuitju/DIC-DDA.
Huanjing Yue, Gaosheng Liu, Xin Liu 0012, Jing-Yu Yang 0002
IEEE Trans. Image Process.1
2025 Zero-shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model
abstract
Diffusion-based zero-shot image restoration and enhancement models have achieved great success in various tasks of image restoration and enhancement. However, directly applying them to video restoration and enhancement results in severe temporal flickering artifacts. In this paper, we propose the first framework for zero-shot video restoration and enhancement based on the pre-trained image diffusion model. By replacing the spatial self-attention layer with the proposed short-long-range (SLR) temporal attention layer, the pre-trained image diffusion model can take advantage of the temporal correlation between frames. We further propose temporal consistency guidance, spatial-temporal noise sharing, and an early stopping sampling strategy to improve temporally consistent sampling. Our method is a plug-and-play module that can be inserted into any diffusion-based image restoration or enhancement methods to further improve their performance. Experimental results demonstrate the superiority of our proposed method.
Cong Cao 0005, Huanjing Yue, Xin Liu 0012, Jing-Yu Yang 0002
AAAI2
2025 Incomplete Multi-modal Brain Tumor Segmentation via Learnable Sorting State Space Model
abstract
Brain tumor segmentation plays a crucial role in clinical diagnosis, yet the frequent unavailability of certain MRI modalities poses a significant challenge. In this paper, we introduce the Learnable Sorting State Space Model (LS3M), a novel framework designed to maximize the utilization of available modalities for brain tumor segmentation. LS3M excels at efficiently modeling long-range dependencies based on the Mamba design, while incorporating differentiable permutation matrices that reorder input sequences based on modality-specific characteristics. This dynamic reordering ensures that critical spatial inductive biases and long-range semantic correlations inherent in 3D brain MRI are preserved, which is crucial for imcomplete multi-modal brain tumor segmentation. Once the input sequences are reordered using the generated permutation matrix, the Series State Space Model (S3M) block models the relationships between them, capturing both local and long-range dependencies. This enables effective representation of intra-modal and inter-modal relationships, significantly improving segmentation accuracy. Extensive experiments on the BraTS2018 and BraTS2020 datasets demonstrate that LS3M outperforms existing methods, offering a robust solution for brain tumor segmentation, particularly in scenarios with missing modalities.
Zheyu Zhang 0002, Yayuan Lu, Feipeng Ma, Yueyi Zhang 0001, Huanjing Yue, Xiaoyan Sun 0001
CVPR5
2025 Learning Adaptive Lighting via Channel-Aware Guidance
abstract
Learning lighting adaptation is a crucial step in achieving good visual perception and supporting downstream vision tasks. Current research often addresses individual light-related challenges, such as high dynamic range imaging and exposure correction, in isolation. However, we identify shared fundamental properties across these tasks: i) different color channels have different light properties, and ii) the channel differences reflected in the spatial and frequency domains are different. Leveraging these insights, we introduce the channel-aware Learning Adaptive Lighting Network (LALNet), a multi-task framework designed to handle multiple light-related tasks efficiently. Specifically, LALNet incorporates color-separated features that highlight the unique light properties of each color channel, integrated with traditional color-mixed features by Light Guided Attention (LGA). The LGA utilizes color-separated features to guide color-mixed features focusing on channel differences and ensuring visual consistency across all channels. Additionally, LALNet employs dual domain channel modulation for generating color-separated features and a mixed channel modulation and light state space module for producing color-mixed features. Extensive experiments on four representative light-related tasks demonstrate that LALNet significantly outperforms state-of-the-art methods on benchmark tests and requires fewer computational resources. We provide an anonymous online demo at LALNet.
Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0026, Huanjing Yue, Jing-Yu Yang 0002
ICML6
2025 FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
Zhuozhao Hu, Kaishen Yuan, Xin Liu 0012, Zitong Yu, Yuan Zong, Jingang Shi, Huanjing Yue, Jing-Yu Yang 0002
ACM Multimedia7
2025 DSDNet: Raw Domain Demoiréing via Dual Color-Space Synergy
abstract
With the rapid advancement of mobile imaging, capturing screens using smartphones has become a prevalent practice in distance learning and conference recording. However, moiré artifacts, caused by frequency aliasing between display screens and camera sensors, are further amplified by the image signal processing pipeline, leading to severe visual degradation. Existing sRGB domain demoiréing methods struggle with irreversible information loss, while recent two-stage raw domain approaches suffer from information bottlenecks and inference inefficiency. To address these limitations, we propose a single-stage raw domain demoiréing framework, Dual-Stream Demoiréing Network (DSDNet), which leverages the synergy of raw and YCbCr images to remove moiré while preserving luminance and color fidelity. Specifically, to guide luminance correction and moiré removal, we design a raw-to-YCbCr mapping pipeline and introduce the Synergic Attention with Dynamic Modulation (SADM) module. This module enriches the raw-to-sRGB conversion with cross-domain contextual features. Furthermore, to better guide color fidelity, we develop a Luminance-Chrominance Adaptive Transformer (LCAT), which decouples luminance and chrominance representations. Extensive experiments demonstrate that DSDNet outperforms state-of-the-art methods in both visual quality and quantitative evaluation and achieves an inference speed 2.4x faster than the second-best method, highlighting its practical advantages. We provide an anonymous online demo at https://dsdnet.github.io/DSDNet/.
Fangpu Zhang, Yeying Jin, Qihua Cheng, Peng-Tao Jiang, Huanjing Yue, Jing-Yu Yang 0002
ACM Multimedia6
2025 Learning Differential Pyramid Representation for Tone Mapping
abstract
Existing tone mapping methods operate on downsampled inputs and rely on handcrafted pyramids to recover high-frequency details. Existing tone mapping methods operate on downsampled inputs and rely on handcrafted pyramids to recover high-frequency details. These designs typically fail to preserve fine textures and structural fidelity in complex HDR scenes. Furthermore, most methods lack an effective mechanism to jointly model global tone consistency and local contrast enhancement, leading to globally flat or locally inconsistent outputs such as halo artifacts. We present the Differential Pyramid Representation Network (DPRNet), an end-to-end framework for high-fidelity tone mapping. At its core is a learnable differential pyramid that generalizes traditional Laplacian and Difference-of-Gaussian pyramids through content-aware differencing operations across scales. This allows DPRNet to adaptively capture high-frequency variations under diverse luminance and contrast conditions. To enforce perceptual consistency, DPRNet incorporates global tone perception and local tone tuning modules operating on downsampled inputs, enabling efficient yet expressive tone adaptation. Finally, an iterative detail enhancement module progressively restores the full-resolution output in a coarse-to-fine manner, reinforcing structure and sharpness. Experiments show that DPRNet achieves state-of-the-art results, improving PSNR by **2.39 dB** on the 4K HDR+ dataset and **3.01 dB** on the 4K HDRI Haven dataset, while producing perceptually coherent, detail-preserving results. Demo available at [DPRNet](https://xxxxxxdprnet.github.io/DPRNet/).
Yinbo Li, Yihao Liu 0001, Peng-Tao Jiang, Fangpu Zhang, Qihua Cheng, Huanjing Yue, Jing-Yu Yang 0002
NeurIPS7
2025 Multi-Scale Promoted Self-Adjusting Correlation Learning for Facial Action Unit Detection
abstract
Facial Action Unit (AU) detection is a crucial task in affective computing and social robotics as it helps to identify emotions expressed through facial expressions. Anatomically, there are innumerable correlations between AUs, which contain rich information and are vital for AU detection. Previous methods used fixed AU correlations based on expert experience or statistical rules on specific benchmarks, but it is challenging to comprehensively reflect complex correlations between AUs via hand-crafted settings. There are alternative methods that employ a fully connected graph to learn these dependencies exhaustively. However, these approaches can result in a computational explosion and high dependency with a large dataset. To address these challenges, this paper proposes a novel self-adjusting AU-correlation learning (SACL) method with less computation for AU detection. This method adaptively learns and updates AU correlation graphs by efficiently leveraging the characteristics of different levels of AU motion and emotion representation information extracted in different stages of the network. Moreover, this paper explores the role of multi-scale learning in correlation information extraction, and design a simple yet effective multi-scale feature learning (MSFL) method to promote better performance in AU detection. By integrating AU correlation information with multi-scale features, the proposed method obtains a more robust feature representation for the final AU detection. Extensive experiments show that the proposed method outperforms the state-of-the-art methods on widely used AU detection benchmark datasets, with only 28.7% and 12.0% of the parameters and FLOPs of the best method, respectively.
Xin Liu 0012, Kaishen Yuan, Xuesong Niu, Jingang Shi, Zitong Yu, Huanjing Yue, Jing-Yu Yang 0002
IEEE Trans. Affect. Comput.6
2025 RViDeformer: Efficient Raw Video Denoising Transformer With a Larger Benchmark Dataset
abstract
In recent years, raw video denoising has garnered increased attention due to the consistency with the imaging process and well-studied noise modeling in the raw domain. However, two problems still hinder the denoising performance. Firstly, there is no large dataset with realistic motions for supervised raw video denoising, as capturing noisy and clean frames for real dynamic scenes is difficult. To address this, we propose recapturing existing high-resolution videos displayed on a 4K screen with high-low ISO settings to construct noisy-clean paired frames. In this way, we construct a video denoising dataset (named as ReCRVD) with 120 groups of noisy-clean videos, whose ISO values ranging from 1600 to 25600. Secondly, while non-local temporal-spatial attention is beneficial for denoising, it often leads to heavy computation costs. We propose an efficient raw video denoising transformer network (RViDeformer) that explores both short and long-distance correlations. Specifically, we propose multi-branch spatial and temporal attention modules, which explore the patch correlations from local window, local low-resolution window, global downsampled window, and neighbor-involved window, and then they are fused together. We employ reparameterization to reduce computation costs. Our network is trained in both supervised and unsupervised manners, achieving the best performance compared with state-of-the-art methods. Additionally, the model trained with our proposed dataset (ReCRVD) outperforms the model trained with previous benchmark dataset (CRVD) when evaluated on the real-world outdoor noisy videos.Our code and dataset will be released.
Huanjing Yue, Cong Cao 0005, Jing-Yu Yang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2025 TransDiff: Unsupervised Non-Line-of-Sight Imaging With Aperture-Limited Relay Surfaces
abstract
Non-line-of-sight (NLOS) imaging aims to reconstruct scenes hidden from direct view and has broad applications in robotic vision, rescue operations, autonomous driving, and remote sensing. However, most existing methods rely on densely sampled transients from large, continuous relay surfaces, which limits their practicality in real-world scenarios with aperture constraints. To address this limitation, we propose an unsupervised zero-shot framework tailored for confocal NLOS imaging with aperture-limited relay surfaces. Our method leverages latent diffusion models to recover fully-sampled transients from undersampled versions by enforcing measurement consistency during the sampling process. To further improve recovered transient quality, we introduce a progressive recovery strategy that incrementally recovers missing transient values, effectively mitigating the impact of severe aperture limitations. In addition, to suppress error propagation during recovery, we develop a backpropagation-based error correction reconstruction algorithm that refines intermediate recovered transients by enforcing sparsity regularization in the voxel domain, enabling high-fidelity final reconstructions. Extensive experiments on both simulated and real-world datasets validate the robustness and generalization capability of our method across diverse aperture-limited relay surfaces. Notably, our method follows a zero-shot paradigm, requiring only a single pretraining stage without paired data or pattern-specific retraining, which makes it a more practical and generalizable framework for NLOS imaging.
Xingyu Cui, Huanjing Yue, Shida Sun, Yusen Hou, Zhiwei Xiong, Jing-Yu Yang 0002
IEEE Trans. Image Process.2
2025 A Simple Yet Effective Network Based on Vision Transformer for Camouflaged Object and Salient Object Detection
abstract
Camouflaged object detection (COD) and salient object detection (SOD) are two distinct yet closely-related computer vision tasks widely studied during the past decades. Though sharing the same purpose of segmenting an image into binary foreground and background regions, their distinction lies in the fact that COD focuses on concealed objects hidden in the image, while SOD concentrates on the most prominent objects in the image. Building universal segmentation models is currently a hot topic in the community. Previous works achieved good performance on certain task by stacking various hand-designed modules and multi-scale features. However, these careful task-specific designs also make them lose their potential as general-purpose architectures. Therefore, we hope to build general architectures that can be applied to both tasks. In this work, we propose a simple yet effective network (SENet) based on vision Transformer (ViT), by employing a simple design of an asymmetric ViT-based encoder-decoder structure, we yield competitive results on both tasks, exhibiting greater versatility than meticulously crafted ones. To enhance the performance of universal architectures on both tasks, we propose some general methods targeting some common difficulties of the two tasks. First, we use image reconstruction as an auxiliary task during training to increase the difficulty of training, forcing the network to have a better perception of the image as a whole to help with segmentation tasks. In addition, we propose a local information capture module (LICM) to make up for the limitations of the patch-level attention mechanism in pixel-level COD and SOD tasks and a dynamic weighted loss (DW loss) to solve the problem that small target samples are more difficult to locate and segment in both tasks. Finally, we also conduct a preliminary exploration of joint training, trying to use one model to complete two tasks simultaneously. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of our method. The code is available at https://github.com/linuxsino/SENet.
Chao Hao, Zitong Yu, Xin Liu 0012, Jun Xu 0019, Huanjing Yue, Jing-Yu Yang 0002
IEEE Trans. Image Process.5
2025 Learning to See Low-Light Images via Feature Domain Adaptation
abstract
Raw low-light image enhancement (LLIE) has achieved much better performance than the sRGB domain enhancement methods due to the merits of raw data. However, the ambiguity between noisy to clean and raw to sRGB mappings may mislead the single-stage enhancement networks. The two-stage networks avoid ambiguity by step-by-step or decoupling the two mappings but usually have large computing complexity. To solve this problem, we propose a single-stage network empowered by Feature Domain Adaptation (FDA) to decouple the denoising and color mapping tasks in raw LLIE. The denoising encoder is supervised by the clean raw image, and then the denoised features are adapted for the color mapping task by an FDA module. We propose a Lineformer to serve as the FDA, which can well explore the global and local correlations with fewer line buffers (friendly to the line-based imaging process). During inference, the raw supervision branch is removed. In this way, our network combines the advantage of a two-stage enhancement process with the efficiency of single-stage inference. Experiments on four benchmark datasets demonstrate that our method achieves state-of-the-art performance with fewer computing costs (60% FLOPs of the two-stage method DNF). Our codes will be released after the acceptance of this work.
Qihua Cheng, Huanjing Yue, Yihao Liu 0001, Jing-Yu Yang 0002
IEEE Trans. Image Process.3
2025 Learned Focused Plenoptic Image Compression With Local-Global Correlation Learning
abstract
The dense light field sampling of focused plenoptic images (FPIs) yields substantial amounts of redundant data, necessitating efficient compression in practical applications. However, the presence of discontinuous structures and long-distance properties in FPIs poses a challenge. In this paper, we propose a novel end-to-end approach for learned focused plenoptic image compression (LFPIC). Specifically, we introduce a local-global correlation learning strategy to build the nonlinear transforms. This strategy can effectively handle the discontinuous structures and leverage long-distance correlations in FPI for high compression efficiency. Additionally, we propose a spatial-wise context model tailored for LFPIC to help emphasize the most related symbols during coding and further enhance the rate-distortion performance. Experimental results demonstrate the effectiveness of our proposed method, achieving a 22.16% BD-rate reduction (measured in PSNR) on the public dataset compared to the recent state-of-the-art LFPIC method. This improvement holds significant promise for benefiting the applications of focused plenoptic cameras.
Gaosheng Liu, Huanjing Yue, Bihan Wen, Jing-Yu Yang 0002
IEEE Trans. Multim.2
2024 TMFormer: Token Merging Transformer for Brain Tumor Segmentation with Missing Modalities
abstract
Numerous techniques excel in brain tumor segmentation using multi-modal magnetic resonance imaging (MRI) sequences, delivering exceptional results. However, the prevalent absence of modalities in clinical scenarios hampers performance. Current approaches frequently resort to zero maps as substitutes for missing modalities, inadvertently introducing feature bias and redundant computations. To address these issues, we present the Token Merging transFormer (TMFormer) for robust brain tumor segmentation with missing modalities. TMFormer tackles these challenges by extracting and merging accessible modalities into more compact token sequences. The architecture comprises two core components: the Uni-modal Token Merging Block (UMB) and the Multi-modal Token Merging Block (MMB). The UMB enhances individual modality representation by adaptively consolidating spatially redundant tokens within and outside tumor-related regions, thereby refining token sequences for augmented representational capacity. Meanwhile, the MMB mitigates multi-modal feature fusion bias, exclusively leveraging tokens from present modalities and merging them into a unified multi-modal representation to accommodate varying modality combinations. Extensive experimental results on the BraTS 2018 and 2020 datasets demonstrate the superiority and efficacy of TMFormer compared to state-of-the-art methods when dealing with missing modalities.
Zheyu Zhang 0002, Yueyi Zhang 0001, Huanjing Yue, Aiping Liu, Yunwei Ou, Xiaoyan Sun 0001
AAAI4
2024 KeDuSR: Real-World Dual-Lens Super-Resolution via Kernel-Free Matching
abstract
Dual-lens super-resolution (SR) is a practical scenario for reference (Ref) based SR by utilizing the telephoto image (Ref) to assist the super-resolution of the low-resolution wide-angle image (LR input). Different from general RefSR, the Ref in dual-lens SR only covers the overlapped field of view (FoV) area. However, current dual-lens SR methods rarely utilize these specific characteristics and directly perform dense matching between the LR input and Ref. Due to the resolution gap between LR and Ref, the matching may miss the best-matched candidate and destroy the consistent structures in the overlapped FoV area. Different from them, we propose to first align the Ref with the center region (namely the overlapped FoV area) of the LR input by combining global warping and local warping to make the aligned Ref be sharp and consistent. Then, we formulate the aligned Ref and LR center as value-key pairs, and the corner region of the LR is formulated as queries. In this way, we propose a kernel-free matching strategy by matching between the LR-corner (query) and LR-center (key) regions, and the corresponding aligned Ref (value) can be warped to the corner region of the target. Our kernel-free matching strategy avoids the resolution gap between LR and Ref, which makes our network have better generalization ability. In addition, we construct a DuSR-Real dataset with (LR, Ref, HR) triples, where the LR and HR are well aligned. Experiments on three datasets demonstrate that our method outperforms the second-best method by a large margin. Our code and dataset are available at https://github.com/ZifanCui/KeDuSR.
Huanjing Yue, Zifan Cui, Kun Li 0001, Jing-Yu Yang 0002
AAAI1
2024 Anatomical Consistency Distillation and Inconsistency Synthesis for Brain Tumor Segmentation with Missing Modalities
abstract
Multi-modal Magnetic Resonance Imaging (MRI) is imperative for accurate brain tumor segmentation, offering indispensable complementary information. Nonetheless, the absence of modalities poses significant challenges in achieving precise segmentation. Recognizing the shared anatomical structures between mono-modal and multi-modal representations, it is noteworthy that mono-modal images typically exhibit limited features in specific regions and tissues. In response to this, we present Anatomical Consistency Distillation and Inconsistency Synthesis (ACDIS), a novel framework designed to transfer anatomical structures from multi-modal to mono-modal representations and synthesize modality-specific features. ACDIS consists of two main components: Anatomical Consistency Distillation (ACD) and Modality Feature Synthesis Block (MFSB). ACD incorporates the Anatomical Feature Enhancement Block (AFEB), meticulously mining anatomical information. Simultaneously, Anatomical Consistency ConsTraints (ACCT) are employed to facilitate the consistent knowledge transfer, i.e., the richness of information and the similarity in anatomical structure, ensuring precise alignment of structural features across mono-modality and multi-modality. Complementarily, MFSB produces modality-specific features to rectify anatomical inconsistencies, thereby compensating for missing information in the segmented features. Through validation on the BraTS2018 and BraTS2020 datasets, ACDIS substantiates its efficacy in the segmentation of brain tumors with missing MRI modalities.
Zheyu Zhang 0002, Xinzhao Liu, Yueyi Zhang 0001, Huanjing Yue, Yunwei Ou, Xiaoyan Sun 0001
ECAI5
2024 AUFormer: Vision Transformers Are Parameter-Efficient Facial Action Unit Detectors
Kaishen Yuan, Zitong Yu, Xin Liu 0012, Weicheng Xie 0001, Huanjing Yue, Jing-Yu Yang 0002
ECCV (50)5
2024 Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality Reweighter
abstract
Deep neural networks (DNNs) have been applied in many computer vision tasks and achieved state-of-the-art (SOTA) performance. However, misclassification will occur when DNNs predict adversarial examples which are created by adding human-imperceptible adversarial noise to natural examples. This limits the application of DNN in security-critical fields. In order to enhance the robustness of models, previous research has primarily focused on the unimodal domain, such as image recognition and video understanding. Although multi-modal learning has achieved advanced performance in various tasks, such as action recognition, research on the robustness of RGB-skeleton action recognition models is scarce. In this paper, we systematically investigate how to improve the robustness of RGB-skeleton action recognition models. We initially conducted empirical analysis on the robustness of different modalities and observed that the skeleton modality is more robust than the RGB modality. Motivated by this observation, we propose the Attention-based Modality Reweighter (AMR), which utilizes an attention layer to re-weight the two modalities, enabling the model to learn more robust features. Our AMR is plug-and-play, allowing easy integration with multimodal models. To demonstrate the effectiveness of AMR, we conducted extensive experiments on various datasets. For example, compared to the SOTA methods, AMR exhibits a 43.77% improvement against PGD20 attacks on the NTURGB+D 60 dataset. Furthermore, it effectively balances the differences in robustness between different modalities.
Xin Liu 0012, Zitong Yu, Yonghong Hou, Huanjing Yue, Jing-Yu Yang 0002
IJCB5
2024 Efficient Screen Content Image Compression via Superpixel-based Content Aggregation and Dynamic Feature Fusion
Sheng Shen 0010, Huanjing Yue, Jing-Yu Yang 0002
IJCAI2
2024 Virtual Scanning: Unsupervised Non-line-of-sight Imaging from Irregularly Undersampled Transients
abstract
Non-line-of-sight (NLOS) imaging allows for seeing hidden scenes around corners through active sensing. Most previous algorithms for NLOS reconstruction require dense transients acquired through regular scans over a large relay surface, which limits their applicability in realistic scenarios with irregular relay surfaces. In this paper, we propose an unsupervised learning-based framework for NLOS imaging from irregularly undersampled transients~(IUT). Our method learns implicit priors from noisy irregularly undersampled transients without requiring paired data, which is difficult and expensive to acquire and align. To overcome the ambiguity of the measurement consistency constraint in inferring the albedo volume, we design a virtual scanning process that enables the network to learn within both range and null spaces for high-quality reconstruction. We devise a physics-guided SURE-based denoiser to enhance robustness to ubiquitous noise in low-photon imaging conditions. Extensive experiments on both simulated and real-world data validate the performance and generalization of our method. Compared with the state-of-the-art (SOTA) method, our method achieves higher fidelity, greater robustness, and remarkably faster inference times by orders of magnitude. The code and model are available at https://github.com/XingyuCuii/Virtual-Scanning-NLOS.
Xingyu Cui, Huanjing Yue, Xiangjun Yin, Yusen Hou, Yun Meng, Jing-Yu Yang 0002
NeurIPS2
2024 Unsupervised HDR Image and Video Tone Mapping via Contrastive Learning
abstract
Capturing high dynamic range (HDR) images (videos) is attractive because it can reveal the details in both dark and bright regions. Since the mainstream screens only support low dynamic range (LDR) content, tone mapping algorithm is required to compress the dynamic range of HDR images (videos). Although image tone mapping has been widely explored, video tone mapping is lagging behind, especially for the deep-learning-based methods, due to the lack of HDR-LDR video pairs. In this work, we propose a unified framework (IVTMNet) for unsupervised image and video tone mapping. To improve unsupervised training, we propose domain and instance based contrastive learning loss. Instead of using a universal feature extractor, such as VGG to extract the features for similarity measurement, we propose a novel latent code, which is an aggregation of the brightness and contrast of extracted features, to measure the similarity of different pairs. We totally construct two negative pairs and three positive pairs to constrain the latent codes of tone mapped results. For the network structure, we propose a spatial-feature-enhanced (SFE) module to enable information exchange and transformation of nonlocal regions. For video tone mapping, we propose a temporal-feature-replaced (TFR) module to efficiently utilize the temporal correlation and improve the temporal consistency of video tone-mapped results. We construct a large-scale unpaired HDR-LDR video dataset to facilitate the unsupervised training process for video tone mapping. Experimental results demonstrate that our method outperforms state-of-the-art image and video tone mapping methods. Our code and dataset are available athttps://github.com/cao-cong/UnCLTMO.
Cong Cao 0005, Huanjing Yue, Xin Liu 0012, Jing-Yu Yang 0002
IEEE Trans. Circuits Syst. Video Technol.2
2024 rPPG-MAE: Self-Supervised Pretraining With Masked Autoencoders for Remote Physiological Measurements
abstract
Remote photoplethysmography (rPPG) is an important technique for detecting human vital signs and has received extensive attention. For a long time, researchers have focused attention on supervised methods that rely on large amounts of labeled data. These methods are limited by their need for large amounts of data and the difficulty of acquiring ground truth physiological signals. To address these issues, several self-supervised methods based on contrastive learning have been proposed. However, they focus on contrastive learning between samples, which neglects inherent self-similar priors in physiological signals and seems to have a limited ability to cope with noise. In this paper, a linear self-supervised reconstruction task was designed for extracting the inherent self-similar priors in physiological signals. In addition, a specific noise-insensitive strategy was explored for reducing the interference of motion and illumination. The framework proposed in this paper, rPPG-MAE, demonstrates excellent performance even on the challenging VIPL-HR dataset. We also evaluate the proposed method on two public datasets, namely, PURE and UBFC-rPPG. The results show that our method not only outperforms existing self-supervised methods but also outperforms state-of-the-art (SOTA) supervised methods. One important observation is that the quality of the dataset appears to be more important than the size of the dataset used in self-supervised pretraining of the rPPG. The source code is available athttps://github.com/linuxsino/rPPG-MAE.
Xin Liu 0012, Yuting Zhang 0008, Zitong Yu, Hao Lu 0009, Huanjing Yue, Jing-Yu Yang 0002
IEEE Trans. Multim.5
2024 From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation
abstract
The action anticipation task refers to predicting what action will happen based on observed videos, which requires the model to have a strong ability to summarize the present and then reason about the future. Experience and common sense suggest that there is a significant correlation between different actions, which provides valuable prior knowledge for the action anticipation task. However, previous methods have not effectively modeled this underlying statistical relationship. To address this issue, we propose a novel end-to-end video modeling architecture that utilizes attention mechanisms, named Anticipation via Recognition and Reasoning (ARR). ARR decomposes the action anticipation task into action recognition and sequence reasoning tasks and effectively learns the statistical relationship between actions by next action prediction (NAP). In comparison to existing temporal aggregation strategies, ARR is able to extract more effective features from observable videos to make more reasonable predictions. In addition, to address the challenge of relationship modeling that requires extensive training data, we propose an innovative approach for the unsupervised pre-training of the decoder, which leverages the inherent temporal dynamics of video to enhance the reasoning capabilities of the network. Extensive experiments on the Epic-kitchen-100, EGTEA Gaze+, and 50salads datasets demonstrate the efficacy of the proposed methods. The code is available at https://github.com/linuxsino/ARR .
Xin Liu 0012, Chao Hao, Zitong Yu, Huanjing Yue, Jing-Yu Yang 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Dec-Adapter: Exploring Efficient Decoder-Side Adapter for Bridging Screen Content and Natural Image Compression
abstract
Natural image compression has been greatly improved in the deep learning era. However, the compression performance will be heavily degraded if the pretrained encoder is directly applied on screen content image compression. Meanwhile, we observe that parameter-efficient transfer learning (PETL) methods have shown great adaptation ability in high-level vision tasks. Therefore, we propose a Dec-Adapter, a pioneering entropy-efficient transfer learning module for the decoder to bridge natural image and screen content compression. The adapter’s parameters are learned during encoding and transmitted to the decoder for image-adaptive decoding. Our Dec-Adapter is lightweight, domain-transferable, and architecture-agnostic with generalized performance in bridging the two domains. Experiments demonstrate that our method outperforms all existing methods by a large margin in terms of BD-rate performance on screen content image compression. Specifically, our method achieves over 2 dB gain compared with the baseline when transferred to screen content image compression.
Sheng Shen 0010, Huanjing Yue, Jing-Yu Yang 0002
ICCV2
2023 Adversarial Dual-Student With Differentiable Spatial Warping for Semi-Supervised Semantic Segmentation
abstract
A common challenge posed to robust semantic segmentation is the expensive data annotation cost. Existing semi-supervised solutions show great potential for solving this problem. Their key idea is constructing consistency regularization with unsupervised data augmentation from unlabeled data for model training. The perturbations for unlabeled data enable the consistency training loss, which benefits semi-supervised semantic segmentation. However, these perturbations destroy image context and introduce unnatural boundaries, which is harmful for semantic segmentation. Besides, the widely adopted semi-supervised learning framework, i.e. mean-teacher, suffers performance limitation since the student model finally converges to the teacher model. In this paper, first of all, we propose a context friendly differentiable geometric warping to conduct unsupervised data augmentation; secondly, a novel adversarial dual-student framework is proposed to improve the Mean-Teacher from the following two aspects: (1) dual student models are learned independently except for a stabilization constraint to encourage exploiting model diversities; (2) adversarial training scheme is applied to both students and the discriminators are resorted to distinguish reliable pseudo-label of unlabeled data for self-training. Effectiveness is validated via extensive experiments on PASCAL VOC2012 and Cityscapes. Our solution significantly improves the performance and state-of-the-art results are achieved on both datasets. Remarkably, compared with fully supervision, our solution achieves comparable mIoU of 73.4% using only 12.5% annotated data on PASCAL VOC2012. Our codes and models are available athttps://github.com/cao-cong/ADS-SemiSeg.
Cong Cao 0005, Dongliang He, Fu Li 0003, Huanjing Yue, Jing-Yu Yang 0002, Errui Ding
IEEE Trans. Circuits Syst. Video Technol.5
2023 Intra-Inter View Interaction Network for Light Field Image Super-Resolution
abstract
Light field (LF) cameras, which can record real-word scenes from multiple viewpoints in a single shot, are widely used in 3D reconstruction, re-focusing, and virtual realityetc. However, the inherent trade-off between spatial resolution and angular resolution of LF images hinders their applications for scenarios requiring high resolutions. In this paper, we propose a novel intra-inter view interaction network for LF image super-resolution, termed as LF-IINet, to exploit the correlations among all views and simultaneously preserve the parallax structure of LF views. The proposed LF-IINet consists of two parallel branches. Specifically, the top branch extracts global inter-view information, and the bottom branch first independently maps each view to deep representations and then models the correlations among all intra-view features via proposed multi-view context block (MCB). The two branches interact with each other by proposed inter-assist-intra feature updating module (IntraFUM, where the intra feature are updated with the assistance of the inter feature) and intra-assist-inter feature updating module (InterFUM, where the inter feature are updated with the assistance of the intra feature). In this way, our LF-IINet incorporates rich angular and spatial information for LF image super-resolution. Extensive comparison with state-of-the-art methods demonstrates that our method achieves superior performance visually and quantitatively. Furthermore, quantitative results also show that our method is effective for LF images with either small or large disparities. Our code is shared inhttps://github.com/GaoshengLiu/LF-IINet.
Gaosheng Liu, Huanjing Yue, Jing-Yu Yang 0002
IEEE Trans. Multim.2
2023 Efficient Light Field Angular Super-Resolution With Sub-Aperture Feature Learning and Macro-Pixel Upsampling
abstract
The acquisition of densely-sampled light field (LF) images is costly, which hampers the applications of LF imaging technology in 3D reconstruction, digital refocusing, virtual reality,etc. To mitigate the obstacle, various approaches have been proposed to reconstruct densely-sampled LF images from sparsely-sampled ones. However, most existing methods still suffer from the non-Lambertian effect and large disparity issue. In this paper, we embrace the challenges by introducing a new paradigm for LF angular super-resolution (SR), which first explores the multi-scale spatial-angular correlations on the sparse sub-aperture images (SAIs) and then performs angular SR on macro-pixel features. In this way, we propose an efficient LF angular SR network, termed as EASR, with simple 3D (2D) CNNs and reshaping operations. The proposed EASR can extract effective feature representations on SAIs and can handle large disparities well by performing angular SR on macro-pixel features. Extensive comparisons with state-of-the-art methods demonstrate that our method achieves superior performance visually and quantitatively. Furthermore, our method achieves efficient angular SR by providing an excellent tradeoff between reconstruction performance and inference time.
Gaosheng Liu, Huanjing Yue, Jing-Yu Yang 0002
IEEE Trans. Multim.2
2023 Recaptured Screen Image Demoiréing in Raw Domain
abstract
Capturing screen content by smart-phone cameras has become a daily routine to record or share instant information from display screens for convenience. However, these recaptured screen images are often degraded by moiré patterns and usually present color cast against the original screen source. We observe that performing demoiréing in raw domain before feeding into the image signal processor (ISP) is more effective than demoiréing in the sRGB domain as done in recent demoiréing works. In this paper, we investigate the demoiréing of raw images through a class-specific learning approach. To this end, we build the first well-aligned raw moiré image dataset by pixel-wise alignment between the recaptured images and source ones. Noting that document images occupy a large portion of screen contents and have different properties from generic images, we propose a class-specific learning strategy for textual images and natural color images. In addition, to deal with moiré patterns with various scales, a multi-scale encoder with multi-level feature fusion is proposed. The shared encoder enables us to extract rich representations for the two kinds of contents and the class-specific decoders benefit the specific content reconstruction by focusing on targeted representations. Experiment results demonstrate that our method achieves state-of-the-art demoiréing performance. We have released the code and dataset inhttps://github.com/tju-chengyijia/RDNet
Huanjing Yue, Yijia Cheng, Cong Cao 0005, Jing-Yu Yang 0002
IEEE Trans. Multim.1
2022 Reference-Based Speech Enhancement via Feature Alignment and Fusion Network
abstract
Speech enhancement aims at recovering a clean speech from a noisy input, which can be classified into single speech enhancement and personalized speech enhancement. Personalized speech enhancement usually utilizes the speaker identity extracted from the noisy speech itself (or a clean reference speech) as a global embedding to guide the enhancement process. Different from them, we observe that the speeches of the same speaker are correlated in terms of frame-level short-time Fourier Transform (STFT) spectrogram. Therefore, we propose reference-based speech enhancement via a feature alignment and fusion network (FAF-Net). Given a noisy speech and a clean reference speech spoken by the same speaker, we first propose a feature level alignment strategy to warp the clean reference with the noisy speech in frame level. Then, we fuse the reference feature with the noisy feature via a similarity-based fusion strategy. Finally, the fused features are skipped connected to the decoder, which generates the enhanced results. Experimental results demonstrate that the performance of the proposed FAF-Net is close to state-of-the-art speech enhancement methods on both DNS and Voice Bank+DEMAND datasets. Our code is available at https://github.com/HieDean/FAF-Net.
Huanjing Yue, Wenxin Duo, Xiulian Peng, Jing-Yu Yang 0002
AAAI1
2022 Real-RawVSR: Real-World Raw Video Super-Resolution with a Benchmark Dataset
Huanjing Yue, Jing-Yu Yang 0002
ECCV (6)1
2022 CdCLR: Clip-Driven Contrastive Learning for Skeleton-Based Action Recognition
abstract
In this study, we propose a Clip-Driven Contrastive Learning for Skeleton-Based Action Recognition (CdCLR). In-stead of considering sequences as instances, CdCLR extracts clips from the sequences as new instances. Aim to implement inherent supervision-guided contrastive learning through joint optimal training of sequences discrimination, clips discrimination, and order verification. Mining abundant positive/negative pairs inside sequence while learning inter-and intra-sequence semantic repre-sentations. Extensive experiments on the NTU RGB+D 60, UCLA and iMiGUE datasets present that CdCLR exhibits superior performance under various evaluation protocols and reaches state-of-the-art. Our code is available at https://github.com/Erich-G/CdCLRI.
Rong Gao 0005, Xin Liu 0012, Jing-Yu Yang 0002, Huanjing Yue
VCIP4
2022 Cloud Detection From Remote Sensing Imagery Based on Domain Translation Network
abstract
Cloud detection in optical imagery has drawn remarkable attention in the era of big Earth observation data analytic. While multiple supervised learning models have been developed for such purpose, large volumes of paired training samples annotated at the pixel level are essential to ensure the model’s generalization capacity. However, constructing a comprehensive cloud detection training database is a tedious and time-consuming process. To tackle this dilemma, we simply regard cloud-contaminated remote sensing (RS) imagery as the combination of cloud and background domains and propose a cloud detection framework based on image-to-image domain translation network (DTNet) to separate cloud-contaminated RS imagery into two target domains of cloud and background object images without using any paired and pixel-level annotation training data. The framework was evaluated with multispectral images from two types of sensors, Landsat-8 Operational Land Imager (OLI) (30 m) and GaoFen-1 (16 m), and demonstrated superior or comparable performance compared with several state-of-the-art cloud detection models.
Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Yang Chen 0015, Chunping Hou, Kun Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 Unsupervised Domain Adaptation for Cloud Detection Based on Grouped Features Alignment and Entropy Minimization
abstract
Most convolutional neural network (CNN)-based cloud detection methods are built upon the supervised learning framework that requires a large number of pixel-level labels. However, it is expensive and time-consuming to manually annotate pixelwise labels for massive remote sensing images. To reduce the labeling cost, we propose an unsupervised domain adaptation (UDA) approach to generalize the model trained on labeled images of source satellite to unlabeled images of the target satellite. To effectively address the domain shift problem on cross-satellite images, we develop a novel UDA method based on grouped features alignment (GFA) and entropy minimization (EM) to extract domain-invariant representations to improve the cloud detection accuracy of cross-satellite images. The proposed UDA method is evaluated on “Landsat-$8~\rightarrow $ZY-3” and “GF-$1\rightarrow $ZY-3” domain adaptation tasks. Experimental results demonstrate the effectiveness of our method against existing state-of-the-art UDA approaches. The code of this paper has been made available online (https://github.com/nkszjx/grouped-features-alignment).
Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Kun Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Unsupervised Domain-Invariant Feature Learning for Cloud Detection of Remote Sensing Images
abstract
The detection of clouds in remote sensing (RS) images is an important task, and convolutional neural networks (CNNs) have been used to perform it. However, supervised cloud detection CNNs rely heavily on a large number of samples annotated at pixel level to tune their parameter. Annotating RS images is a labor-intensive procedure and requires expert-level human knowledge. To reduce the labeling cost, we propose an unsupervised domain adaptation (UDA) approach to enable the model trained on labeled source satellite images to generalize to unlabeled target satellite images. Specifically, we propose a fine-grained feature alignment (FGFA) domain adaptation strategy that encourages a cloud detection network to extract domain-invariant representations, which improves the accuracy of cloud detection in unlabeled target satellite images. The proposed FGFA strategy consists of two steps: 1) fine-grained class-relevant feature selection based on an attention-guided mechanism and 2) class-relevant feature alignment (FA) based on a proposed grouped FA approach. Experimental results on the “Landsat-$8~\rightarrow $ZY-3” and “GF-$1\rightarrow $ZY-3” domain adaptation tasks demonstrate the effectiveness of our method and its superiority to existing state-of-the-art UDA approaches.
Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Xin Liu 0012, Kun Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 Spatio-temporal Contrastive Domain Adaptation for Action Recognition
abstract
Compared with image-based UDA, video-based UDA is comprehensive to bridge the domain shift on both spatial representation and temporal dynamics. Most previous works focus on short-term modeling and alignment with frame-level or clip-level features, which is not discriminative sufficiently for video-based UDA tasks. To address these problems, in this paper we propose to establish the cross-modal domain alignment via self-supervised contrastive framework, i.e., spatio-temporal contrastive domain adaptation (STCDA), to learn the joint clip-level and video-level representation alignment. Since the effective representation is modeled from unlabeled data by self-supervised learning (SSL), spatio-temporal contrastive learning (STCL) is proposed to explore the useful long-term feature representation for classification, using self-supervision setting trained from the contrastive clip/video pairs with positive or negative properties. Besides, we involve a novel domain metric scheme, i.e., video-based contrastive alignment (VCA), to optimize the category-aware video-level alignment and generalization between source and target. The proposed STCDA achieves stat-of-the-art results on several UDA benchmarks for action recognition.
Sicheng Zhao, Jing-Yu Yang 0002, Huanjing Yue, Pengfei Xu 0013, Runbo Hu
CVPR4
2021 Implicit Transformer Network for Screen Content Image Continuous Super-Resolution
abstract
Nowadays, there is an explosive growth of screen contents due to the wide application of screen sharing, remote cooperation, and online education. To match the limited terminal bandwidth, high-resolution (HR) screen contents may be downsampled and compressed. At the receiver side, the super-resolution (SR)of low-resolution (LR) screen content images (SCIs) is highly demanded by the HR display or by the users to zoom in for detail observation. However, image SR methods mostly designed for natural images do not generalize well for SCIs due to the very different image characteristics as well as the requirement of SCI browsing at arbitrary scales. To this end, we propose a novel Implicit Transformer Super-Resolution Network (ITSRN) for SCISR. For high-quality continuous SR at arbitrary ratios, pixel values at query coordinates are inferred from image features at key coordinates by the proposed implicit transformer and an implicit position encoding scheme is proposed to aggregate similar neighboring pixel values to the query one. We construct benchmark SCI1K and SCI1K-compression datasets withLR and HR SCI pairs. Extensive experiments show that the proposed ITSRN significantly outperforms several competitive continuous and discrete SR methods for both compressed and uncompressed SCIs.
Jing-Yu Yang 0002, Sheng Shen 0010, Huanjing Yue, Kun Li 0001
NeurIPS3
2021 Low-light image enhancement based on Retinex decomposition and adaptive gamma correction
abstract
Abstract Low‐light images suffer from poor visibility and noise. In this paper, a low‐light image enhancement method based on Retinex decomposition is proposed. A pyramid network is first utilized to extract multi‐scale features to improve the quality of Retinex decomposition. Then the decomposed illumination is refined via an adaptive Gamma correction network to handle non‐uniform illumination, while the decomposed reflectance is refined with a lightweight network. Finally, the enhanced image is obtained by element‐wise multiplication between the refined illumination and reflectance components. Quantitative and qualitative experiments demonstrate the superiority of our method over state‐of‐the‐art image enhancement methods.
Jing-Yu Yang 0002, Huanjing Yue, Zhongyu Jiang, Kun Li 0001
IET Image Process.3
2021 Unsupervised moiré pattern removal for recaptured screen images
Huanjing Yue, Yijia Cheng, Fanglong Liu, Jing-Yu Yang 0002
Neurocomputing1
2021 Deep edge map guided depth super resolution
Zhongyu Jiang, Huanjing Yue, Yukun Lai, Jing-Yu Yang 0002, Yonghong Hou, Chunping Hou
Signal Process. Image Commun.2
2021 Deep noise estimation and removal for real-world noisy images
Huanjing Yue, Zhongyu Jiang, Shengdi Zhou, Jing-Yu Yang 0002, Yonghong Hou, Chunping Hou
Signal Process. Image Commun.1
2021 Reference guided image super-resolution via efficient dense warping and adaptive fusion
Huanjing Yue, Zhongyu Jiang, Jing-Yu Yang 0002, Chunping Hou
Signal Process. Image Commun.1
2021 Recaptured Screen Image Demoiréing
abstract
In many situations, such as transferring data between devices and recording precious moments, we would like to capture the contents on screens using digital cameras for convenience. These recaptured screen images and videos suffer from a special type of degradation called “moiré pattern”, which is caused by the aliasing between the grid of display screen and the array of camera sensor. However, few works are proposed to tackle this problem. Considering the great success of convolutional neural networks (CNNs) in image restoration, we propose a CNN-based moiré removal method for recaptured screen images. There are mainly two contributions in this paper. First, for the generation of training data, we propose an image registration algorithm via global homography transform and local patch matching to compensate the significant viewpoint disparity between the recaptured screen image and the moiré-free image obtained via screenshot. We construct a moiré removal and brightness improvement (MRBI) database with aligned moiré-free and moiré images. Second, we propose a convolutional neural Network with Additive and Multiplicative modules (termed as AMNet) to transfer the low light moiré image to the bright moiré-free image. The proposed network is trained with pixel-wise loss, perceptual loss, and adversarial loss. Extensive experiments on 340 test images demonstrate that the proposed method outperforms state-of-the-art moiré removal methods.
Huanjing Yue, Lipu Liang, Hongteng Xu, Chunping Hou, Jing-Yu Yang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2021 Landsat-8 OLI Multispectral Image Dehazing Based on Optimized Atmospheric Scattering Model
abstract
Optical satellite images are often affected by haze atmospheric conditions, which degrades the quality of remote sensing (RS) data and reduces the accuracy of interpretation and classification. Hence, haze removal becomes a necessary preprocessing step for most of the applications of RS image. In this article, we propose a novel haze removal method for Landsat-8 OLI multispectral image based on an optimized atmospheric scattering model. We focus on adaptively estimating the haze transmission map of each band by taking into account the effect of both wavelength and haze atmospheric conditions (haze particle size and haze particle concentration) thus improving dehazing performance. The experimental results on Landsat-8 OLI multispectral images show that the proposed dehazing model is able to remove haze successfully and significantly improve the image visibility as well as correct the spectral bias to some degree. Moreover, this method is simple and feasible, and has good practical value.
Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 CDnetV2: CNN-Based Cloud Detection for Remote Sensing Imagery With Cloud-Snow Coexistence
abstract
Cloud detection is a crucial preprocessing step for optical satellite remote sensing (RS) images. This article focuses on the cloud detection for RS imagery with cloud-snow coexistence and the utilization of the satellite thumbnails that lose considerable amount of high resolution and spectrum information of original RS images to extract cloud mask efficiently. To tackle this problem, we propose a novel cloud detection neural network with an encoder-decoder structure, named CDnetV2, as a series work on cloud detection. Compared with our previous CDnetV1, CDnetV2 contains two novel modules, that is, adaptive feature fusing model (AFFM) and high-level semantic information guidance flows (HSIGFs). AFFM is used to fuse multilevel feature maps by three submodules: channel attention fusion model (CAFM), spatial attention fusion model (SAFM), and channel attention refinement model (CARM). HSIGFs are designed to make feature layers at decoder of CDnetV2 be aware of the locations of the cloud objects. The high-level semantic information of HSIGFs is extracted by a proposed high-level feature fusing model (HFFM). By being equipped with these two proposed key modules, AFFM and HSIGFs, CDnetV2 is able to fully utilize features extracted from encoder layers and yield accurate cloud detection results. Experimental results on the ZY-3 satellite thumbnail data set demonstrate that the proposed CDnetV2 achieves accurate detection accuracy and outperforms several state-of-the-art methods.
Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 RSDehazeNet: Dehazing Network With Channel Refinement for Multispectral Remote Sensing Images
abstract
Multispectral remote sensing (RS) images are often contaminated by the haze that degrades the quality of RS data and reduces the accuracy of interpretation and classification. Recently, the emerging deep convolutional neural networks (CNNs) provide us new approaches for RS image dehazing. Unfortunately, the power of CNNs is limited by the lack of sufficient hazy-clean pairs of RS imagery, which makes supervised learning impractical. To meet the data hunger of supervised CNNs, we propose a novel haze synthesis method to generate realistic hazy multispectral images by modeling the wavelength-dependent and spatial-varying characteristics of haze in RS images. The proposed haze synthesis method not only alleviates the lack of realistic training pairs in multispectral RS image dehazing but also provides a benchmark data set for quantitative evaluation. Furthermore, we propose an end-to-end RSDehazeNet for haze removal. We utilize both local and global residual learning strategies in RSDehazeNet for fast convergence with superior performance. Channel attention modules are incorporated to exploit strong channel correlation in multispectral RS images. Experimental results show that the proposed network outperforms the state-of-the-art methods for synthetic data and real Landsat-8 OLI multispectral RS images.
Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2020 Supervised Raw Video Denoising With a Benchmark Dataset on Dynamic Scenes
abstract
In recent years, the supervised learning strategy for real noisy image denoising has been emerging and has achieved promising results. In contrast, realistic noise removal for raw noisy videos is rarely studied due to the lack of noisy-clean pairs for dynamic scenes. Clean video frames for dynamic scenes cannot be captured with a long-exposure shutter or averaging multi-shots as was done for static images. In this paper, we solve this problem by creating motions for controllable objects, such as toys, and capturing each static moment for multiple times to generate clean video frames. In this way, we construct a dataset with 55 groups of noisy-clean videos with ISO values ranging from 1600 to 25600. To our knowledge, this is the first dynamic video dataset with noisy-clean pairs. Correspondingly, we propose a raw video denoising network (RViDeNet) by exploring the temporal, spatial, and channel correlations of video frames. Since the raw video has Bayer patterns, we pack it into four sub-sequences, i.e RGBG sequences, which are denoised by the proposed RViDeNet separately and finally fused into a clean video. In addition, our network not only outputs a raw denoising result, but also the sRGB result by going through an image signal processing (ISP) module, which enables users to generate the sRGB result with their favourite ISPs. Experimental results demonstrate that our method outperforms state-of-the-art video and raw image denoising algorithms on both indoor and outdoor videos.
Huanjing Yue, Cong Cao 0005, Ronghe Chu, Jing-Yu Yang 0002
CVPR1
2020 Reference Image Guided Super-Resolution via Progressive Channel Attention Networks
Huanjing Yue, Sheng Shen 0010, Jing-Yu Yang 0002, Haofeng Hu, Yan-Fang Chen
J. Comput. Sci. Technol.1
2020 Spatiotemporally scalable matrix recovery for background modeling and moving object detection
Jing-Yu Yang 0002, Huanjing Yue, Kun Li 0001, Chunping Hou
Signal Process.3
2020 IENet: Internal and External Patch Matching ConvNet for Web Image Guided Denoising
abstract
From the non-local self-similarity (NSS)-based image denoising to the convolutional-network (ConvNet)-based image denoising, the denoising performance has been greatly improved. However, it is still not clear how to utilize similar web images to guide image denoising using ConvNet. This paper proposes a novel ConvNet for image denoising to explore both internal (NSS) and external correlations when external similar images are available. Since external similar images may be taken with different viewpoints, focal lengths, and may contain different objects, it is difficult to directly explore external correlations at image level using ConvNet. Therefore, we propose an internal and external patch matching ConvNet (IENet), whose inputs are similar patch cubes extracted from the noisy input and its external similar images. We design three different network structures, namely early-fusion, middle-fusion, and late-fusion of the internal and external cubes to fully combine the strengths of internal and external correlations. The experimental results demonstrate that the proposed method achieves the best denoising results compared with the seven state-of-the-art denoising methods. In specific, the proposed method outperforms the state-of-the-art web image guided denoising method by more than 1 dB on average, which further demonstrates the superiority of the proposed IENet-based filtering over the hand-crafted filtering methods.
Huanjing Yue, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Truong Q. Nguyen, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Learning to Reconstruct and Understand Indoor Scenes From Sparse Views
abstract
This paper proposes a new method for simultaneous 3D reconstruction and semantic segmentation for indoor scenes. Unlike existing methods that require recording a video using a color camera and/or a depth camera, our method only needs a small number of (e.g., 3~5) color images from uncalibrated sparse views, which significantly simplifies data acquisition and broadens applicable scenarios. To achieve promising 3D reconstruction from sparse views with limited overlap, our method first recovers the depth map and semantic information for each view, and then fuses the depth maps into a 3D scene. To this end, we design an iterative deep architecture, named IterNet, to estimate the depth map and semantic segmentation alternately. To obtain accurate alignment between views with limited overlap, we further propose a joint global and local registration method to reconstruct a 3D scene with semantic information. We also make available a new indoor synthetic dataset, containing photorealistic high-resolution RGB images, accurate depth maps and pixel-level semantic labels for thousands of complex layouts. Experimental results on public datasets and our dataset demonstrate that our method achieves more accurate depth estimation, smaller semantic segmentation errors, and better 3D reconstruction results over state-of-the-art methods.
Jing-Yu Yang 0002, Kun Li 0001, Yukun Lai, Huanjing Yue, Jianzhi Lu, Hao Wu 0042, Yebin Liu
IEEE Trans. Image Process.5
2019 Single Image De-Raining via Generative Adversarial Nets
abstract
In this paper, we propose a Generative Adversarial Network for Single Image De-raining(GAN-SID). We observe that batch normalization has side effects in the de-raining task. Therefore, we introduce instance normalization to replace the traditional batch normalization layers in both generator and discriminator. Motivated by the Squeeze-and-Excitation (SE) network that can learn the importance of channels, we introduce SE module in the generator to give different weights to the learned features. To preserve image details while removing rain streaks, we propose to utilize pixel-wise loss, perceptual loss, and adversarial loss to train the proposed network. Experiments on two synthetic datasets and real world images demonstrate that the proposed method outperforms state-of-the-art de-raining works in both objective and subjective measurements.
Shichao Li 0006, Yonghong Hou, Huanjing Yue, Zihui Guo
ICME3
2019 CDnet: CNN-Based Cloud Detection for Remote Sensing Imagery
abstract
Cloud detection is one of the important tasks for remote sensing image (RSI) preprocessing. In this paper, we utilize the thumbnail (i.e., preview image) of RSI, which contains the information of original multispectral or panchromatic imagery, to extract cloud mask efficiently. Compared with detection cloud mask from original RSI, it is more challenging to detect cloud mask using thumbnails due to the loss of resolution and spectrum information. To tackle this problem, we propose a cloud detection neural network (CDnet) with an encoder-decoder structure, a feature pyramid module (FPM), and a boundary refinement (BR) block. The FPM extracts the multiscale contextual information without the loss of resolution and coverage; the BR block refines object boundaries; and the encoder-decoder structure gradually recovers segmentation results with the same size as input image. Experimental results on the ZY-3 satellite thumbnails cloud cover validation data set and two other validation data sets (GF-1 WFV Cloud and Cloud Shadow Cover Validation Data and Landsat-8 Cloud Cover Assessment Validation Data) demonstrate that the proposed method achieves accurate detection accuracy and outperforms several state-of-the-art methods.
Jing-Yu Yang 0002, Jianhua Guo 0002, Huanjing Yue, Haofeng Hu, Kun Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 High ISO JPEG Image Denoising by Deep Fusion of Collaborative and Convolutional Filtering
abstract
Capturing images at high ISO modes will introduce much realistic noise, which is difficult to be removed by traditional denoising methods. In this paper, we propose a novel denoising method for high ISO JPEG images via deep fusion of collaborative and convolutional filtering. Collaborative filtering explores the non-local similarity of natural images, while convolutional filtering takes advantage of the large capacity of convolutional neural networks (CNNs) to infer noise from noisy images. We observe that the noise variance map of a high ISO JPEG image is spatial-dependent and has a Bayer-like pattern. Therefore, we introduce the Bayer pattern prior in our noise estimation and collaborative filtering stages. Since collaborative filtering is good at recovering repeatable structures and convolutional filtering is good at recovering irregular patterns and removing noise in flat regions, we propose to fuse the strengths of the two methods via deep CNN. The experimental results demonstrate that our method outperforms the state-of-the-art realistic noise removal methods for a wide variety of testing images in both subjective and objective measurements. In addition, we construct a dataset with noisy and clean image pairs for high ISO JPEG images to facilitate research on this topic.
Huanjing Yue, Jing-Yu Yang 0002, Truong Q. Nguyen, Feng Wu 0001
IEEE Trans. Image Process.1
2018 Image Alignment via Multi-Model Geometric Fitting and Hierarchical Homography Estimation
abstract
It is challenging to achieve accurate alignment for building images containing multiple planes. We propose a multi-model geometric fitting and hierarchical homography estimation method to improve the alignment performance for building images. We first extract scale-invariant feature transform (SIFT) features of the images, and then adopt the multi-homography fitting algorithm to classify the feature points into different deformation models. According to the deduced deformation models, we partition the source image into base and transition regions. For the base regions, we adopt the moving direct linear transformation (Moving DLT) to estimate homographies. For the transition regions, we propose a hierarchical homography estimation method to select appropriate homographies. Experimental results show that our method achieves more accurate alignment results compared with state-of-the-art alignment methods for building images.
Jing-Yu Yang 0002, Huanjing Yue, Kun Li 0001, Chunping Hou
ICASSP3
2018 Deep Joint Noise Estimation and Removal for High ISO JPEG Images
abstract
Capturing images under high ISO mode introduces much noise. The statistics of high ISO noise is quite different from that of Gaussian noise. Therefore, this kind of noise is difficult to be removed by traditional Gaussian noise removal methods. This paper proposes a convolutional neural network (CNN) based method to jointly estimate and remove high ISO noise. There are two contributions in this paper. First, we propose a CNN based noise estimation method to estimate the pixel-wise noise level. Due to the Bayer down-sampling process in imaging, the noise variance map is characterized by Bayer patterns. Therefore, we propose packing 2 × 2 blocks in a noisy image into 4D vectors, which makes the pixels with similar noise levels be neighbors. Second, the noise variance map is correlated with the image content. Thus, we propose concatenating the estimated noise variance map with the noisy image, and feed the fused data to the denoising network. The two networks are trained together in an end-to-end fashion. Experimental results demonstrate that the proposed method outperforms state-of-the-art noise estimation and removal methods.
Huanjing Yue, Shengdi Zhou, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Chunping Hou
ICPR1
2018 Depth Super-Resolution From RGB-D Pairs With Transform and Spatial Domain Regularization
abstract
This paper proposes a depth super-resolution method with both transform and spatial domain regularization. In the transform domain regularization, nonlocal correlations are exploited via an auto-regressive model, where each patch is further sparsified with a locally-trained transform to consider intra-patch correlations. In the spatial domain regularization, we propose a multi-directional total variation (MTV) prior to characterize the geometrical structures spatially orientated at arbitrary directions in depth maps. To achieve adaptive regularization, the MTV is weighted for each directional finite difference considering local characteristics of RGB-D data. We develop an accelerated proximal gradient algorithm to solve the proposed model. Quantitative and qualitative evaluations compared with state-of-the-art methods demonstrate that the proposed method achieves superior depth super-resolution performance for various configurations of magnification factors and datasets.
Zhongyu Jiang, Yonghong Hou, Huanjing Yue, Jing-Yu Yang 0002, Chunping Hou
IEEE Trans. Image Process.3
2017 Underwater image enhancement based on structure-texture decomposition
abstract
Underwater images generally suffer from low contrast, serious noise and color distortion. The main challenges of underwater image enhancement are to preserve details in dark regions while avoiding oversaturetion in bright regions. This paper proposes a novel underwater image enhancement method based on image decomposition. By decomposing the high-frequency texture and noise into the texture layer, the transmission map is estimated from the noise-free structure layer to avoid the noise amplification problem in underwater image enhancement. Both the structure layer and texture layer are descattered with the estimated transmission map. After denoising by gradient residual minimizition, the texture layer is enhanced and added back into the structure layer to recover the final enhanced image. Experimental results verify that the proposed approach can recover the high-quality images with fine details and edges while improving contrast and color naturalness, especially for images taken in the high turbidity environment.
Jing-Yu Yang 0002, Huanjing Yue, Xiaomei Fu, Chunping Hou
ICIP3
2017 Image noise estimation and removal considering the bayer pattern of noise variance
abstract
Traditional image denoising methods are designed for Gaussian or Poisson noise, which are not suitable for realistic noise introduced in the complicated imaging pipeline. We observe that, due to the demosaicing process in imaging, the noise variance maps of captured JPEG images are characterized by Bayer patterns. In this paper, we propose a novel noise estimation and removal method based on the Bayer pattern of noise variance maps. There are two key contributions in the proposed method. First, to the best of our knowledge, we are the first to consider the Bayer patterns of noise variance maps in noise estimation and denoising. Second, we extend the state-of-the-art denoising method CBM3D to deal with realistic noise by integrating the estimated noise variance map and Bayer-pattern down-sampling into the denoising process. Experimental results show that the proposed method achieves the best noise estimation performance compared with two state-of-the-art methods. In addition, the denoising performance of CBM3D for realistic noise is significantly improved using the proposed approach and outperforms state-of-the-art blind denoising methods.
Huanjing Yue, Jing-Yu Yang 0002, Truong Q. Nguyen, Chunping Hou
ICIP1
2017 Depth Map Super-Resolution Considering View Synthesis Quality
abstract
Accurate and high-quality depth maps are required in lots of 3D applications, such as multi-view rendering, 3D reconstruction and 3DTV. However, the resolution of captured depth image is much lower than that of its corresponding color image, which affects its application performance. In this paper, we propose a novel depth map super-resolution (SR) method by taking view synthesis quality into account. The proposed approach mainly includes two technical contributions. First, since the captured low-resolution (LR) depth map may be corrupted by noise and occlusion, we propose a credibility based multi-view depth maps fusion strategy, which considers the view synthesis quality and interview correlation, to refine the LR depth map. Second, we propose a view synthesis quality based trilateral depth-map up-sampling method, which considers depth smoothness, texture similarity and view synthesis quality in the up-sampling filter. Experimental results demonstrate that the proposed method outperforms state-of-the-art depth SR methods for both super-resolved depth maps and synthesized views. Furthermore, the proposed method is robust to noise and achieves promising results under noise-corruption conditions.
Jianjun Lei 0001, Lele Li, Huanjing Yue, Feng Wu 0001, Nam Ling, Chunping Hou
IEEE Trans. Image Process.3
2017 Textured Image Demoiréing via Signal Decomposition and Guided Filtering
abstract
Moiré artifacts are generally caused by the interference between the overlap of the sensor's sampling grid and high-frequency (nearly) periodic textures, and heavily affect the image quality. However, it is difficult to effectively remove moiré artifacts from textured images as the structure of moiré patterns is similar to that of textures in some sense. In this paper, we propose a novel textured image demoiréing method by signal decomposition and guided filtering. Given a textured image with moiré artifacts, we first remove moiré artifacts in the green (G) channel using the proposed low-rank and sparse matrix decomposition model. This model regularizes the texture layer by the low-rank prior in spatial domain and the moiré layer by sparse representation in frequency domain. An alternating direction method under the augmented Lagrangian multiplier framework is used to solve the matrix decomposition model. Then, since the red (R) and blue (B) channels are more heavily polluted by moiré artifacts than the G channel, we propose to remove moiré artifacts in its R and B channels via guided filtering by the obtained texture layer of the G channel. Experimental results demonstrate that our method outperforms the state-of-the-art methods for both synthetic and real images.
Jing-Yu Yang 0002, Fanglei Liu, Huanjing Yue, Xiaomei Fu, Chunping Hou, Feng Wu 0001
IEEE Trans. Image Process.3
2017 Contrast Enhancement Based on Intrinsic Image Decomposition
abstract
In this paper, we propose to introduce intrinsic image decomposition priors into decomposition models for contrast enhancement. Since image decomposition is a highly illposed problem, we introduce constraints on both reflectance and illumination layers to yield a highly reliable solution. We regularize the reflectance layer to be piecewise constant by introducing a weighted ℓ1norm constraint on neighboring pixels according to the color similarity, so that the decomposed reflectance would not be affected much by the illumination information. The illumination layer is regularized by a piecewise smoothness constraint. The proposed model is effectively solved by the Split Bregman algorithm. Then, by adjusting the illumination layer, we obtain the enhancement result. To avoid potential color artifacts introduced by illumination adjusting and reduce computing complexity, the proposed decomposition model is performed on the value channel in HSV space. Experiment results demonstrate that the proposed method performs well for a wide variety of images, and achieves better or comparable subjective and objective quality compared with the state-of-the-art methods.
Huanjing Yue, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Feng Wu 0001, Chunping Hou
IEEE Trans. Image Process.1
2016 Background recovery from video sequences via online motion-assisted RPCA
abstract
Background modeling is an important technique for video analysis. Robust principal component analysis (RPCA) assisted with motion information has shown improved background recovery performance, but still suffers from the deficiency in handling steaming video due to the batch-mode formulation and implementation. This paper proposes an online motion-assisted robust principal component analysis (OMA-RPCA) model for background recovery from video sequences. The inherent batch-mode nuclear norm for low-rank approximation is replaced with an explicitly low-rank matrix factorization. Motion information extracted by an optical flow method is incorporated into the data term to facilitate the separation of moving objects from the background. The proposed model is effectively solved by an alternating optimization scheme in an online mode. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods with lower memory cost and scalability to online applications.
Jiaoru Yang, Jing-Yu Yang 0002, Xuemeng Yang, Huanjing Yue
VCIP4
2015 Moiré pattern removal from texture images via low-rank and sparse matrix decomposition
abstract
Moiré patterns, an artifact of aliasing interference between details in the subject matter and the grid of the sensor, heavily disturb the qualitative and quantitative analysis of images. It is hard to effectively remove moiré patterns since they are similar to image textures. We propose a novel low-rank and sparse matrix decomposition model for moiré pattern removal. This method is grounded on the observation: textures are locally well-patterned while moiré patterns are dissimilar, and the energy distribution of moiré patterns in the frequency domain is concentrated and almost no mixed with that of textures. For each patch, texture component is regularized by a low-rank prior and moiré component is regularized by a sparse prior in the discrete cosine transform (DCT) domain. This model is effectively solved by the alternating direction method under the augmented Lagrangian multiplier (ALM-ADM) algorithm. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods.
Fanglei Liu, Jing-Yu Yang 0002, Huanjing Yue
VCIP3
2015 Image Denoising by Exploring External and Internal Correlations
abstract
Single image denoising suffers from limited data collection within a noisy image. In this paper, we propose a novel image denoising scheme, which explores both internal and external correlations with the help of web images. For each noisy patch, we build internal and external data cubes by finding similar patches from the noisy and web images, respectively. We then propose reducing noise by a two-stage strategy using different filtering approaches. In the first stage, since the noisy patch may lead to inaccurate patch selection, we propose a graph based optimization method to improve patch matching accuracy in external denoising. The internal denoising is frequency truncation on internal cubes. By combining the internal and external denoising patches, we obtain a preliminary denoising result. In the second stage, we propose reducing noise by filtering of external and internal cubes, respectively, on transform domain. In this stage, the preliminary denoising result not only enhances the patch matching accuracy but also provides reliable estimates of filtering parameters. The final denoising image is obtained by fusing the external and internal filtering results. Experimental results show that our method constantly outperforms state-of-the-art denoising schemes in both subjective and objective quality measurements, e.g., it achieves >2 dB gain compared with BM3D at a wide range of noise levels.
Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001
IEEE Trans. Image Process.1
2014 CID: Combined Image Denoising in Spatial and Frequency Domains Using Web Images
abstract
In this paper, we propose a novel two-step scheme to filter heavy noise from images with the assistance of retrieved Web images. There are two key technical contributions in our scheme. First, for every noisy image block, we build two three dimensional (3D) data cubes by using similar blocks in retrieved Web images and similar nonlocal blocks within the noisy image, respectively. To better use their correlations, we propose different denoising strategies. The denoising in the 3D cube built upon the retrieved images is performed as median filtering in the spatial domain, whereas the denoising in the other 3D cube is performed in the frequency domain. These two denoising results are then combined in the frequency domain to produce a denoising image. Second, to handle heavy noise, we further propose using the denoising image to improve image registration of the retrieved Web images, 3D cube building, and the estimation of filtering parameters in the frequency domain. Afterwards, the proposed denoising is performed on the noisy image again to generate the final denoising result. Our experimental results show that when the noise is high, the proposed scheme is better than BM3D by more than 2 dB in PSNR and the visual quality improvement is clear to see.
Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001
CVPR1
2013 SIFT-based image super-resolution
abstract
This paper presents a new exemplar-based image super-resolution (SR) method in which we propose making use of scale invariant image features for high frequency (HF) approximation. We introduce the scale invariant feature transform (SIFT) descriptors in both building an exemplar dataset adaptively and producing the HF details with respect to the features of an input low resolution image. Given a large image database, we propose using the highly correlated images retrieved by SIFT descriptors for exemplar training rather than using a general set of images to increase the matching accuracy. Through building the training set of high resolution/low resolution exemplar pairs, the HF details for SR are retrieved from the training set by matching the SIFT features in a dense way. The flexibility as well as effectiveness of our SR approach is demonstrated at different magnification factors, e.g. 3 and 4. Experimental results show that our SIFT-based SR approach achieves enhanced high resolution images in terms of both objective and subjective qualities in comparison with the state-of-the-art exemplar-based methods.
Huanjing Yue, Jing-Yu Yang 0002, Xiaoyan Sun 0001, Feng Wu 0001
ISCAS1
2013 Landmark Image Super-Resolution by Retrieving Web Images
abstract
This paper proposes a new super-resolution (SR) scheme for landmark images by retrieving correlated web images. Using correlated web images significantly improves the exemplar-based SR. Given a low-resolution (LR) image, we extract local descriptors from its up-sampled version and bundle the descriptors according to their spatial relationship to retrieve correlated high-resolution (HR) images from the web. Though similar in content, the retrieved images are usually taken with different illumination, focal lengths, and shot perspectives, resulting in uncertainty for the HR detail approximation. To solve this problem, we first propose aligning these images to the up-sampled LR image through a global registration, which identifies the corresponding regions in these images and reduces the mismatching. Second, we propose a structure-aware matching criterion and adaptive block sizes to improve the mapping accuracy between LR and HR patches. Finally, these matched HR patches are blended together by solving an energy minimization problem to recover the desired HR image. Experimental results demonstrate that our SR scheme achieves significant improvement compared with four state-of-the-art schemes in terms of both subjective and objective qualities.
Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001
IEEE Trans. Image Process.1
2013 Cloud-Based Image Coding for Mobile Devices - Toward Thousands to One Compression
abstract
Current image coding schemes make it hard to utilize external images for compression even if highly correlated images can be found in the cloud. To solve this problem, we propose a method of cloud-based image coding that is different from current image coding even on the ground. It no longer compresses images pixel by pixel and instead tries to describe images and reconstruct them from a large-scale image database via the descriptions. First, we describe an input image based on its down-sampled version and local feature descriptors. The descriptors are used to retrieve highly correlated images in the cloud and identify corresponding patches. The down-sampled image serves as a target to stitch retrieved image patches together. Second, the down-sampled image is compressed using current image coding. The feature vectors of local descriptors are predicted by the corresponding vectors extracted in the decoded down-sampled image. The predicted residual vectors are compressed by transform, quantization, and entropy coding. The experimental results show that the visual quality of reconstructed images is significantly better than that of intra-frame coding in HEVC and JPEG at thousands to one compression .
Huanjing Yue, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Feng Wu 0001
IEEE Trans. Multim.1
2012 SIFT-Based Image Compression
abstract
This paper proposes a novel image compression scheme based on the local feature descriptor - Scale Invariant Feature Transform (SIFT). The SIFT descriptor characterizes an image region invariantly to scale and rotation. It is used widely in image retrieval. By using SIFT descriptors, our compression scheme is able to make use of external image contents to reduce visual redundancy among images. The proposed encoder compresses an input image by SIFT descriptors rather than pixel values. It separates the SIFT descriptors of the image into two groups, a visual description which is a significantly sub sampled image with key SIFT descriptors embedded and a set of differential SIFT descriptors, to reduce the coding bits. The corresponding decoder generates the SIFT descriptors from the visual description and the differential set. The SIFT descriptors are used in our SIFT-based matching to retrieve the candidate predictive patches from a large image dataset. These candidate patches are then integrated into the visual description, presenting the final reconstructed images. Our preliminary but promising results demonstrate the effectiveness of our proposed image coding scheme towards perceptual quality. Our proposed image compression scheme provides a feasible approach to make use of the visual correlation among images.
Huanjing Yue, Xiaoyan Sun 0001, Feng Wu 0001, Jing-Yu Yang 0002
ICME1
2012 IMShare: instantly sharing your mobile landmark images by search-based reconstruction
abstract
Instantly sharing captured landmark images is becoming fashionable, much like when you write a blog or chat with friends by mobile phone. However, real-time transmission of high-resolution images poses a significant challenge to contemporary mobile networks. Either long delays in transmission or largely reduced image resolution can lead to bad user experience. In this paper, we propose a novel mobile-cloud scheme IMShare to enable instant sharing of high-resolution images. On the mobile side, high-resolution images are described by their thumbnails and SIFT (Scale-Invariant Feature Transform) descriptors. After compression, data sent by mobile phones can be reduced to an average of 2.6 kilobytes (KB) per mega pixel. On the cloud side, high-resolution images are reproduced from a large-scale image database by retrieving partial duplicate images by SIFT descriptors and stitching corresponding image patches together under the guidance of the thumbnails. IMShare is the first scheme to demonstrate that not only visually pleasant images can be reconstructed using this mobile-cloud method but also the reconstruction can be done in seconds using parallel computing. Our user study of a half million images in a database shows that the proposed IMShare significantly outperforms the current method on subjective quality.
Lican Dai, Huanjing Yue, Xiaoyan Sun 0001, Feng Wu 0001
ACM Multimedia2