VLDB 2026 Research / reviewers in the wild / expert
Yuming Fang 0001
dblp:31/9004-1
· DBLP profile ↗
271ranked-venue papers
48as first author
160since 2021 · last 2026
0000-0002-6946-3586ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 204 · 39 first-author · 115 since 2021Artificial intelligence and machine learning · 51 · 6 first-author · 41 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 2 since 2021Security and privacy · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Computer networks · 6 · 6 since 2021Systems, architecture and hardware · 5 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-CancerabstractSingle-cell RNA sequencing (scRNA-seq) is essential for decoding tumor heterogeneity. However, pan-cancer research still faces two key challenges: learning discriminative and efficient single-cell representations, and establishing a comprehensive evaluation benchmark. In this paper, we introduce \algoname, a lightweight hybrid neural network that combines the strengths of Transformers and state-space models to achieve a balance between performance and efficiency. \algoname consists of a front-end local-context encoder with shared self-attention layers to capture complex, order-independent gene interactions; and a back-end global sequential feature decoder that efficiently integrates global context using a linear-time state-space model. This modular design preserves the expressive power of Transformers while leveraging the scalability of Mamba to enable transcriptome modeling, effectively capturing both local and global regulatory signals. To enable robust evaluation, we also construct a large-scale pan-cancer single-cell benchmark, \algoname Bench, containing over 3.5 million high-quality cells across 33 cancer subtypes, curated through a rigorous preprocessing pipeline. Experimental results show that \algoname outperforms state-of-the-art models on our pan-cancer benchmark (+4.0\%) and across multiple public tasks, including cell type annotation (+7.4\%), batch integration (+4.0\%) and multi-omics integration (+3.1\%). Xiaoshui Huang, Tianlin Zhu, Yifan Zuo 0001, Xue Xia 0005, Zonghan Wu, Jiebin Yan, Dingli Hua, Zongyi Xu, Yuming Fang 0001, Jian Zhang 0002 |
AAAI | 9 |
| 2026 | Event-Guided Super-Resolving Blurry Image via Asymmetric Integral Driven ConsistencyabstractSuper-Resolution from a Blurry low-resolution image (SRB) constitutes a severely ill-posed inverse problem. Current learning-based SRB approaches primarily rely on synthetic, well-labeled paired datasets to regularize solution spaces, yet they exhibit limited generalizability in practical applications due to significant domain discrepancies between simulated degradations and real-world imaging conditions. To bridge this synthetic-to-real gap, we propose a novel Self-supervised Event-based SRB (SE-SRB) framework that leverages neuromorphic event streams as physical priors and adopts a lightweight neural architecture tailored for effective domain adaptation. Specifically, the proposed SE-SRB introduces a self-supervised learning paradigm based on asymmetric integral driven consistency, which enforces temporal coherence between predictions derived from RGB and asynchronous event streams at different time points. Extensive experiments validate that SE-SRB consistently outperforms state-of-the-art methods on both synthetic and real-world datasets. Built upon a lightweight parallel two-stream architecture, SE-SRB achieves high computational efficiency, featuring reduced parameter count, lower FLOPs, and real-time inference capability (40 FPS). Chi Zhang 0027, Xiang Zhang 0022, Lei Yu 0006, Gui-Song Xia, Yuming Fang 0001, Wenhan Yang |
AAAI | 5 |
| 2026 | Evidential Robust Feature Learning for Generalized Few-Shot Segmentation
Weide Liu, Xiaoyang Zhong, Lu Wang 0001, Chunbo Lang, Yuming Fang 0001, Jun Cheng 0003, Xulei Yang, Gong Cheng 0003 |
Int. J. Comput. Vis. | 5 |
| 2026 | Blind Omnidirectional Image Quality Assessment: Embracing the Magic Power of Multimodal Large Language Models
Jiebin Yan, Junjie Chen 0008, Pengfei Chen 0003, Xuelin Liu, Ziwen Tan, Yuming Fang 0001 |
Int. J. Comput. Vis. | 7 |
| 2026 | SinMDGan: A Hybrid Deep Learning Framework for Single Motion Synthesis Using Diffusion-GAN ModelsabstractABSTRACT Generating diverse and realistic movements has long been a central challenge in computer graphics. Generative Adversarial Networks (GANs) remain a compelling solution due to their ability to perform well even with limited training data. However, traditional GANs generate samples directly, which can lead to the omission of certain data patterns. To address this limitation, we introduce SinMDGan , a hybrid deep learning framework for single‐motion synthesis that leverages a Diffusion‐GAN model. Our approach integrates the strengths of GANs, which capture global motion characteristics, with diffusion techniques, which refine local details, ensuring both authenticity and diversity in generated movements. Unlike conventional cascaded GANs, our framework employs a single generator‐discriminator pair, utilizing different diffusion time steps to synthesize novel and diverse motions from a single short sequence. Experimental evaluations demonstrate the effectiveness of our model in achieving stable data distribution coverage and enhancing output diversity. Additionally, we showcase various applications, including motion composition and long‐sequence generation, highlighting the versatility of our approach. Binsong Zuo, Tingsong Lu, Yuming Fang 0001, Xiaolu Mu, Xiaogang Jin 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2026 | Inter-modality feature prediction through multimodal fusion for 3D shape defect detection
Mujtaba Asad, Waqar Azeem, Hafiz Tayyab Mustafa, Yuming Fang 0001, Jie Yang 0002, Yifan Zuo 0001, Wei Liu 0044 |
Neural Networks | 4 |
| 2026 | IC-Bench: Benchmarking robustness of large multimodal models to common corruptions on image captioning
Xuelin Liu, Xinpeng Fang, Jiebin Yan, Chengyang Fang, Yuming Fang 0001 |
Pattern Recognit. | 5 |
| 2026 | Audio-visual saliency prediction based on joint adversarial learning and Co-Attention mechanism
Amin Mao, Jiebin Yan, Yuming Fang 0001 |
Pattern Recognit. | 3 |
| 2026 | RGB-D salient object detection via cross-modal adaptive correlation learning network
Guanqun Ding, Yuming Fang 0001, Jiebin Yan |
Signal Process. Image Commun. | 2 |
| 2026 | DSRAS: Dual-Stage Reasoning and Answer Selection for Video-Text Visual Question AnsweringabstractVideo text-based visual question answering (Video TextVQA) aims to answer questions by spatio-temporal joint reasoning over textual and visual information in a video. Existing methods have achieved remarkable progress using uniform sampling strategies. However, uniform frame sampling may introduce noisy frames and miss keyframes. Meanwhile, current methods treat all video frames equally, which is suboptimal, especially since video text question answering tasks primarily target questions that include text within the video. To address the aforementioned issues, we introduce a novel Dual-Stage Reasoning and Answer Selection (DSRAS) model, which not only adaptively focuses on and extracts keyframes, but also significantly enhances attention to video frames containing text through an answer selection mechanism. Specifically, we propose a Dual-Stage Reasoning (DSR) module to achieve adaptive frame selection. Then, we introduce an Answer Selection Module (ASM) to guide our model to focus on keyframes containing textual information. Extensive experiments demonstrate that our model outperforms existing approaches on the RoadTextVQA and M4-ViteVQA datasets. Chengyang Fang, Xiankun Wan, Wenhui Jiang 0001, Yuming Fang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2026 | Text-Conditional Visual-Language Alignment for Video CaptioningabstractVideo captioning remains a challenging task due to the diverse video content and the complex relationships between visual and textual elements. Recent efforts predominantly focus on multimodal architecture designs trained with paired video-caption data. Nonetheless, the learning paradigm suffers from the “one-to-many” corresponding problem, since one source video is mapped to multiple caption annotations. The difficulty of video captioning is further exacerbated by the poor-written captions, which mislead the captioner with irrelevant information. Essentially, the problem stems from the inadequate alignment between video and caption. In this work, we propose a Text-Conditional Alignment Transformer, which fully exploits the rich information provided by diverse labeled captions, and avoids the impacts of label ambiguity and noise. To alleviate the challenge of the “one-to-many” correspondence, we introduce Text-conditioned Video Encoding, which diversifies the video representation by emphasizing the spatial-temporal visual areas relevant to the given descriptions while filtering out redundant visual information. The refined video representation is well-aligned to match the corresponding text description, and naturally converts the “one-to-many” mapping to “one-to-one” mapping. To deal with the noisy annotations, we propose Quality-aware Caption Decoding. We first dynamically measure the qualities of different captions corresponding to the same video in a reference-free manner. Then the estimated qualities are further utilized as auxiliary signals, guiding the model to perform quality-aligned learning from noisy captions. We conduct extensive experiments on MSR-VTT, MSVD, VATEX and ActivityNet-Entities datasets, and demonstrate their consistent performance improvements compared to state-of-the-arts. Wenhui Jiang 0001, Wenbin Guan, Zhizhen Li, Yuming Fang 0001, Yuxin Peng 0001, Yang Liu 0293 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | CMF-IoU: Multi-Stage Cross-Modal Fusion 3D Object Detection With IoU Joint PredictionabstractMulti-modal methods based on camera and LiDAR sensors have garnered significant attention in the field of 3D detection. However, many prevalent works focus on single or partial stage fusion, leading to insufficient feature extraction and suboptimal performance. In this paper, we introduce a multi-stage cross-modal fusion 3D detection framework, termed CMF-IOU, to effectively address the challenge of aligning 3D spatial and 2D semantic information. Specifically, we first project the pixel information into 3D space via a depth completion network to get the pseudo points, which unifies the representation of the LiDAR and camera information. Then, a bilateral cross-view enhancement 3D backbone is designed to encode LiDAR points and pseudo points. The first sparse-to-distant (S2D) branch utilizes an encoder-decoder structure to reinforce the representation of sparse LiDAR points. The second residual view consistency (ResVC) branch is proposed to mitigate the influence of inaccurate pseudo points via both the 3D and 2D convolution processes. Subsequently, we introduce an iterative voxel-point aware fine grained pooling module, which captures the spatial information from LiDAR points and textural information from pseudo points in the proposal refinement stage. To achieve more precise refinement during iteration, an intersection over union (IoU) joint prediction branch integrated with a novel proposals generation technique is designed to preserve the bounding boxes with both high IoU and classification scores. Extensive experiments show the superior performance of our method on the KITTI, nuScenes and Waymo datasets. The code is available at https://github.com/pami-zwning/CMF-IOU. Zhiwei Ning, Zhaojiang Liu, Xuanang Gao, Yifan Zuo 0001, Jie Yang 0002, Yuming Fang 0001, Wei Liu 0044 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Objective Quality Assessment of AI-Generated Content Videos With Transformation Consistency FocusabstractUnnatural motion artifacts—such as implausible object dynamics or discontinuous scene transitions—characterize a critical challenge in AI-generated content (AIGC) videos. Assessing these temporal inconsistencies is essential for benchmarking the performance of text-to-video (T2V) models. Current assessment methods derive motion features from either action recognition models or optical flow estimators. However, these motion features cannot faithfully reflect the human-aligned interpretability of how objects or scenes transform between frames. For instance, a model might detect a “running” action but fail to penalize implausible leg movements. To address this gap, we propose Transformation Consistency-based Video Quality Assessment (TCVQA), a novel framework that quantifies transformation consistency by measuring the recognizability of semantic transformations across frames. The core module of TCVQA is the TC-branch, which includes three core components: A Feature Extractor to capture high-level, fine-grained, and low-level motion features. A Flow-Driven Transformation module Warps extracted features from the source frame to the target frame using predicted optical flow. A Differential Perceiver computes discrepancies between warped source features and actual target features, yielding a consistency score that reflects deviations from natural motion patterns. Besides, the TCVQA also integrates three other branches, the TV-branch, the V-branch, and the F-branch, to perceive multiple aspects of distortions. Extensive experiments on AIGC-VQA benchmarks—including T2VQA-DB, LGVQ, FETV, and MQT demonstrate TCVQA’s superiority, achieving consistent improvement in correlation with human judgments over state-of-the-art methods. Our work establishes transformation consistency as a pivotal axis, enabling more reliable evaluation of AIGC video quality assessment. Jiebin Yan, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Viewport-Unaware Full-Reference Omnidirectional Image Quality Assessment With Inter-Patch and Sequence SimilarityabstractFull-reference (FR) image quality assessment (IQA) (FR-IQA) has been extensively explored in the past two decades and is one of the most basic and hot topics in the image processing community, due to its indispensable role in quantitatively describing image quality degradation and guiding algorithm and system optimization. However, FR omnidirectional image quality assessment (OIQA) (FR-OIQA) has achieved less success, due to the natural gap between 2D images and omnidirectional images (OIs). To this end, we present a novel FR-OIQA model with Inter-Patch and Sequence Similarity (IPSS). Specifically, to avoid the extra computational load of viewport generation/prediction methods, IPSS processes OIs in aviewport-unawaremanner,i.e., directly extracting a patch sequence from an OI in the format of Equirectangular Projection (ERP) with retaining regions of interest. Furthermore, since the patches from ERP image contain inborn geometry deformation, thedeformation-awareconvolution is plugged into feature extraction and used to distill quality-aware features from theintrinsic pseudo-degradation, which are then utilized to measure inter-patch similarity. Finally, a distortion-aware interaction module is used to aggregate patch-wise quality-aware features, whose output is used to calculate patch-sequence similarity,i.e., the global quality of OI. Through comprehensive experiments on a large-scale OIQA database, we demonstrate the superiority of the proposed IPSS and the effectiveness of each module. Jiebin Yan, Junjie Chen 0008, Pengfei Chen 0003, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | I've Got Proof! Dataset-Specific Watermarking for Detecting Excessive Dataset Usage in Text-to-Image Diffusion Model Fine-TuningabstractCurrently, text-to-image diffusion models, which exhibit remarkable proficiency in image generation, have prompted the emergence of diverse fine-tuning methodologies due to the considerable resource demands entailed in their training. Simultaneously, apprehensions have arisen regarding the excessive utilization of open-source datasets for fine-tuning the models. Therefore, intellectual property protection for such datasets is important and necessary. In this regard, watermarking is a commonly employed method. However, the existing watermarking methods fail to establish an accurate correspondence between the watermarked dataset and the malicious model, which may lead to erroneous accusations against the model. In this paper, we propose a dataset-specific watermarking method, DSW, to detect excessive usage of protected datasets by suspect diffusion models fine-tuned using LoRA. DSW can establish a one-to-one correspondence between the watermarked dataset and the malicious model, thereby enhancing the accuracy of intellectual property verification. We simultaneously train an encoder and a decoder based on a bi-level optimization strategy. The encoder embeds a watermark image, which symbolizes the dataset owner's identity, into the dataset samples to yield a watermarked dataset. For a text-to-image diffusion model fine-tuned on this watermarked dataset, the decoder extracts the watermark from the generated image. Comprehensive experiments validate the availability, effectiveness, and stealthiness of DSW. The code is available athttps://github.com/hzhwis/DSW. Xiangli Xiao, Yushu Zhang 0001, Yuming Fang 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | Medical Archive in an Image: Generating a High-Capacity Customizable Cover Image for Medical Privacy ProtectionabstractThe rapid development of medical information technology promotes the popularization of telemedicine. With remote transmission of patient archives, medical institutions can provide more efficient diagnostic service. Because of the highly sensitive nature of medical data, medical archives typically require the support of encryption to prevent leakage of patient privacy. However, the encrypted archives visually appear as snowflake-like noise, which is easily perceived by an attacker and thus appealing to the crack. In this paper, we propose an imperceptible medical archive construction scheme, which can hide personal medical data in a high-capacity customizable cover image, thus making it impossible for attackers to perceive its presence. Specifically, we first conduct a fusion of multimodal medical image data, integrating image information from various imaging techniques into a unified image, thereby enhancing diagnostic efficacy. Secondly, a high-capacity customizable cover image is generated based on the patient facial mask. Lastly, the unified image with patient demographic information is hidden in the cover image to construct the imperceptible medical archive. To enhance the hiding capacity of the cover image, we propose a Collaborative Steganography Generation Model (CSGM), which increases the low-frequency components during image synthesis, enabling more data to be embedded while maintaining high visual quality. CSGM also supports identity-aware customization using facial structure and semantic text. Experiments show that CSGM improves the PSNR of stego images by 6.95 dB and the PSNR of recovered secret images by 2.98 dB compared to state-of-the-art methods, confirming its effectiveness in imperceptibility and data recovery. Wenying Wen, Zhouxin Wu, Yushu Zhang 0001, Tao Wang 0084, Xiangli Xiao, Yuming Fang 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | DeVIL: A Dual Verification Framework for Integrity and Ownership of Light Field Images With Customizable WatermarkingabstractLight field (LF) images capture both spatial and angular data, providing enhanced detail in multi-view and 3D scenes, which makes them highly applicable across various domains. Thus, compared to traditional images, LF images not only contain more complex data and are more susceptible to unauthorized tampering or malicious uses during transmission due to their unique structure, thereby increasing the need for verification. Existing dual watermarking techniques are designed for single-view images and lack adaptability to the multi-view structure and geometric consistency of LF images, making it difficult to simultaneously address integrity verification and ownership verification for LF images. Given the multi-view characteristics and geometric consistency of LF images, there is an urgent need for a highly robust and adaptable dual watermarking method. Therefore, we propose a dual verification framework with customizable watermarking, called DeVIL, which specifically designed for LF images and enables both ownership and integrity verification. By embedding reversible ownership watermarks into the sub-aperture images (SAIs), affiliation can be verified after transmission. Among them, the embedding technique can flexibly choose either a reversible robust watermarking technique or a robust zero-watermarking technique based on different application scenarios. Subsequently, we perform Arnold transformation on key SAIs and embed them into non-key SAIs to create stego images, effectively reducing the risk of information leakage. Furthermore, integrity watermarks are embedded in the stego images, forming stego images with integrity watermarks to detect subtle tampering during transmission. Experimental results show that DeVIL can successfully recover SAIs and demonstrates strong robustness and adaptability against various attacks. Compared to state-of-the-art methods, DeVIL reduces the average bit error rate by 9.32% under different attacks, with average normalized cross-correlation improvement of 0.17, significantly enhancing the robustness of ownership protection in different datasets. Wenying Wen, Xiangli Xiao, Yuming Fang 0001, Yushu Zhang 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | A Dual-Protection Method for 3D Object Security and Copyright: Watermark Embedding During DecryptionabstractWith advancements in the computer industry, 3D objects are now widely used in various applications, including game development, animation production, and industrial design. This growing adoption has increased the need for effective content security and copyright protection for 3D objects. However, existing encryption and watermarking techniques often operate independently, leading to low efficiency and weak coupling between security and copyright protection. To address these gaps, this paper presents a novel method that integrates watermark embedding into the 3D object decryption process, simultaneously ensuring content security and copyright protection. Specifically, a Look-Up Table (LUT)-based encryption method is employed to secure 3D object data, while a Spread Transform Dither Modulation (ST-DM)-based watermarking method is used to embed user-specific identity information during decryption. Unlike conventional approaches that apply encryption and watermarking separately, the proposed method enables efficient 3D object sharing, as the owner only needs to encrypt the object once, regardless of the number of authorized users. The encrypted model can then be distributed securely via multicast and caching. Decryption with personalized keys produces distinct watermarked 3D objects, allowing for reliable traceability of unauthorized redistribution. Experimental results and theoretical evaluations demonstrate that the proposed method delivers satisfactory visual quality, efficiency, robustness, and security. Xiangli Xiao, Yushu Zhang 0001, Zhongyun Hua, Wenying Wen, Yuming Fang 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | Detecting Malicious Concepts Without Image Generation in AI-Generated Content (AIGC)abstractThe task of text-to-image generation has achieved tremendous success in practice, with emerging concept generation models capable of producing highly personalized and customized content. Fervor for concept generation is increasing rapidly among users, and platforms for concept sharing have sprung up. The concept owners may upload malicious concepts and disguise them with non-malicious text descriptions and example images to deceive users into downloading and generating malicious content. The platform needs a quick method to determine whether a concept is malicious to prevent the spread of malicious concepts. However, simply relying on concept image generation to judge whether a concept is malicious requires time and computational resources. Especially, as the number of concepts uploaded and downloaded on the platform continues to increase, this approach becomes impractical and poses a risk of generating malicious content. In this paper, we propose Concept QuickLook, the first systematic work to incorporate malicious concept detection into research, which performs detection based solely on concept files without generating any images. We define malicious concepts and design two operational modes for detection: concept matching and fuzzy detection. Extensive experiments demonstrate that the proposed Concept QuickLook can detect malicious concepts and demonstrate practicality in concept sharing platforms. We also design robustness experiments to further validate the effectiveness of the solution. We hope this work can initiate malicious concept detection tasks and provide some inspiration. Kun Xu 0019, Wenying Wen, Tao Wang 0084, Yushu Zhang 0001, Yuming Fang 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | Tracing the Use of Open-Source Training Datasets for Neural Radiance Field Models
Yushu Zhang 0001, Xiangli Xiao, Zhongyun Hua, Yuming Fang 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | Traffic-Aware Asynchronous Trajectory Planning and Scheduling in UAV-Assisted Wireless Networks With Heterogeneous Traffic Demands
Che Chen, Bo Gu 0003, Bin Lyu, Shimin Gong, Zhi Liu 0002, Yuming Fang 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | Distributed Two-Tier Cache Optimization in Metaverse Scenarios Combining MADDPG and GCNabstractThe rapid emergence of the Metaverse requires higher network throughput and lower latency to deliver immersive and responsive virtual experiences. Traditional centralized data processing approaches are constrained by limited computational and bandwidth resources when handling large-scale user data. A Cloud-Edge-End transmission architecture is proposed in this study, tailored for Metaverse scenarios to optimize resource allocation, minimize latency, and enhance rendering efficiency. A real-time trajectory segment prediction scheme (FDK) was developed, which combines FastDTW with K-means by leveraging user behavior trajectories to determine subscene popularity and store them on GPU servers, thereby reducing user wait time. A two-tier cache optimization scheme (MAE2C) is also proposed, incorporating GCN for subscene feature identification. GPU servers employ the MADDPG strategy to cache popular subscenes, while edge servers utilize DDPG to cache missed scenes. This approach effectively reduces cloud access and cache replacement frequency. Simulation results demonstrate that the subscene cache hit rate of the MAE2C scheme significantly outperforms existing methods across various cache capacities, with a 6.9% reduction in cache replacement frequency. This research provides effective technical support for Metaverse scene rendering and offers insights into the development of generative Metaverse systems. Shenglu Zhao, Xuelin Liu, Yifeng Tan, Yuming Fang 0001 |
IEEE Trans. Multim. | 6 |
| 2026 | Data-Driven Control of Insect Flapping Flight via Deep Reinforcement LearningabstractModeling and simulating realistic insect flight pose unique challenges due to the complex interaction between multi-degree-of-freedom wing kinematics and highly precise aerodynamic forces. To solve this challenge, this article presents a bidirectional kinematics-aerodynamics coupled simulation framework for miniature insect flight. Our approach first models the kinematics of flying insects by parameterizing natural wingbeat cycles based on available real-world datasets. Subsequently, we compute aerodynamic forces utilizing an improved semi-empirical model, which extends from quasi-steady formulation by incorporating critical unsteady force components. To achieve closed-loop control for both kinematics and aerodynamics, we employ deep reinforcement learning to train a virtual insect to adaptively adjust flapping strategies in response to dynamic flight states. Finally, an integrated controller enables the simulated insect to autonomously regulate the wing motion and perform complex tasks such as visual obstacle avoidance. Extensive experiments and comparisons demonstrate that our framework can effectively generate physically plausible and autonomous insect flight across a variety of scenarios. Tingsong Lu, Yuming Fang 0001, Camille Le Roy, Xiaogang Jin 0001, Zhigang Deng 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | FreDeGS: frequency-guided 3D gaussian splatting for scene deblurring
Xiaofeng Quan, Junzhe Wan, Amin Mao, Yuming Fang 0001 |
Vis. Comput. | 6 |
| 2025 | PSReg: Prior-guided Sparse Mixture of Experts for Point Cloud RegistrationabstractThe discriminative feature is crucial for point cloud registration. Recent methods improve the feature discriminative by distinguishing between non-overlapping and overlapping region points. However, they still face challenges in distinguishing the ambiguous structures in the overlapping regions. Therefore, the ambiguous features they extracted resulted in a significant number of outlier matches from overlapping regions. To solve this problem, we propose a prior-guided SMoE-based registration method to improve the feature distinctiveness by dispatching the potential correspondences to the same experts. Specifically, we propose a prior-guided SMoE module by fusing prior overlap and potential correspondence embeddings for routing, assigning tokens to the most suitable experts for processing. In addition, we propose a registration framework by a specific combination of Transformer layer and prior-guided SMoE module. The proposed method not only pays attention to the importance of locating the overlapping areas of point clouds, but also commits to finding more accurate correspondences in overlapping areas. Our extensive experiments demonstrate the effectiveness of our method, achieving state-of-the-art registration recall (95.7%/79.3%) on the 3DMatch/3DLoMatch benchmark. Moreover, we also test the performance on ModelNet40 and demonstrate excellent performance. Xiaoshui Huang, Zhou Huang 0006, Yifan Zuo 0001, Yongshun Gong, Chengdong Zhang, Deyang Liu, Yuming Fang 0001 |
AAAI | 7 |
| 2025 | Enhancing Low-Light Images: A Synthetic Data Perspective on Practical and Generalizable SolutionsabstractRecently, deep neural networks (DNNs) have emerged as the leading approach for low-light image enhancement (LLIE). However, training these models generally requires large-scale paired datasets, which are challenging to obtain due to the labor-intensive and time-consuming nature of real-world data collection. To alleviate this issue, synthetic data are often combined with real-captured data for training. However, most existing low-light image synthesis methods are simply performed in the sRGB domain using Gamma correction or manual adjustments via Lightroom, which fail to incorporate the physical imaging prior through the image signal processing (ISP) pipeline and thus result in limited dataset size and degradation space. Consequently, LLIE methods trained on such data often exhibit some drawbacks in the results, such as inaccurate white balance and abnormal enhancement artifacts, which limit their practicality and generalizability. In this paper, we propose a practical low-light image synthesis pipeline capable of generating unlimited paired training data. Our pipeline starts with a reverse ISP model that converts sRGB images back to the unprocessed RAW domain, where we then simulate low-light degradation, noise degradation, and white balance adjustments. Finally, the degraded RAW images are processed through a forward ISP model to produce low-light sRGB images. The pipeline further employs multiple tone mapping curves and color correction matrices (CCMs) to expand the degradation space. Hence, trained with our proposed synthetic data, existing state-of-the-art (SOTA) LLIE deep models are expected to improve their performance. Extensive experiments across various datasets demonstrate that our synthetic data can indeed effectively enhance existing LLIE deep models, improving both their practicality and generalizability. Qinghua Lin, Zhihua Wang 0002, Jianguo Zhang 0001, Yuming Fang 0001 |
AAAI | 6 |
| 2025 | Decoding Emotions: How Eye Features Influence Perception of Emotional Intensity in Virtual Characters
Mingfang Mao, Tingsong Lu, Shikun Zhou, Yuming Fang 0001 |
CGI (1) | 5 |
| 2025 | GCN and MADDPG-Based Two-Tier Distributed Cache Optimization for Metaverse ScenariosabstractWith the rapid development of the Metaverse, the demand for high transmission rates and low latency is increasing. Traditional centralized data processing architectures, however, are unable to meet these demands due to resource and bandwidth bottlenecks. This paper proposes a cloud-edge-end collaborative transmission architecture to optimize resource allocation and improve rendering efficiency. A real-time trajectory segmentation prediction scheme (FDK) integrates the FastDTW algorithm with KMeans clustering to predict sub-scene popularity, enabling GPU cache allocation and reducing user wait times. Additionally, a two-tier cache optimization scheme (MAE2C) uses GCN to analyze sub-scene features, employing MADDPG to cache popular scenes on GPU servers and DDPG to cache missed scenes on edge servers. Simulation results show that the MAE2C scheme significantly improves cache hit rates, reducing cache replacement frequency by45.14%. This study provides efficient support for Metaverse scene rendering and insights for the development of generative Metaverse technologies. Shenglu Zhao, Xuelin Liu, Yifeng Tan, Yuming Fang 0001 |
HPCC | 6 |
| 2025 | PointGAC: Geometric-Aware Codebook for Masked Point Cloud ModelingabstractMost masked point cloud modeling (MPM) methods follow a regression paradigm to reconstruct the coordinate or feature of masked regions. However, they tend to over-constrain the model to learn the details of the masked region, resulting in failure to capture generalized features. To address this limitation, we propose \textbf{\textit{PointGAC}}, a novel clustering-based MPM method that aims to align the feature distribution of masked regions. Specially, it features an online codebook-guided teacher-student framework. Firstly, it presents a geometry-aware partitioning strategy to extract initial patches. Then, the teacher model updates a codebook via online k-means based on features extracted from the complete patches. This procedure facilitates codebook vectors to become cluster centers. Afterward, we assigns the unmasked features to their corresponding cluster centers, and the student model aligns the assignment for the reconstructed masked features. This strategy focuses on identifying the cluster centers to which the masked features belong, enabling the model to learn more generalized feature representations. Benefiting from a proposed codebook maintenance mechanism, codebook vectors are actively updated, which further increases the efficiency of semantic feature learning. Experiments validate the effectiveness of the proposed method on various downstream tasks. Code is available at https://github.com/LAB123-tech/PointGAC Abiao Li, Chenlei Lv, Yuming Fang 0001, Yifan Zuo 0001, Jian Zhang 0002, Guofeng Mei |
ICCV | 3 |
| 2025 | Dynamic 3D Gaussian Reconstruction with Specular Reflectionabstract3D Gaussian Splatting (3DGS) has shown remarkable potential in novel view synthesis. However, it still encounters significant challenges in reconstructing dynamic scenes, particularly when dealing with reflective surfaces. To address this issue, we propose a novel 3DGS-based method for dynamic scene reconstruction with explicit reflection modeling. Our approach integrates deferred shading with a dual-environment map that combines static and dynamic components, enabling effective modeling of specular reflections. This allows our method to capture both steady and temporally varying lighting, resulting in more realistic renderings. We evaluate the proposed method on the benchmark dynamic reflection dataset, NERF-DS, and compare it with state-of-the-art approaches. Experimental results show that our method achieves superior or comparable performance in terms of PSNR, SSIM, and LPIPS metrics compared to competing approaches. Mingyang Zhao 0001, Yuanzhi Xu, Yifan Zuo 0001, Xiaoshui Huang, Yuming Fang 0001 |
ICIP | 6 |
| 2025 | Cross-Structure and Semantic Enhancement for Diabetic Retinopathy GradingabstractChallenges such as highly variable lesion appearances and complex structural distributions hinder model performance in diabetic retinopathy (DR) grading tasks. To address these issues, we focus on guiding the network toward discriminative feature representation by prioritizing diagnostically relevant information and modeling intricate dependencies between features and DR grades. Specifically, we propose a DR grading network that integrates the Kolmogorov-Arnold Network (KAN) into a Convolution-Vision Transformer (CNN-ViT) cooperative framework, where convolutions and Transformer encoders capture spatial patterns, hierarchical structures, and context, while KAN enhances non-linear semantic dependency modeling. Additionally, we introduce a Cross-Structure (CS) attention module to emphasize relevant features. The proposed modules form the Convolution-Cross-Structure-KAN (CCSK) block, which serves as the backbone of our network, CCSKFormer, enabling more accurate DR grading. The proposed model achieves outstanding performance on two public datasets, with comparisons and ablation studies further validating the effectiveness of the individual modules (https://github.com/xia-xx-cv/CCSKformer). Xue Xia 0005, Zipeng Lin, Jingying Zhu, Jiebin Yan, Yuming Fang 0001 |
ICME | 5 |
| 2025 | MERD-360VR: A Multimodal Emotional Response Dataset from 360° VR Videos Across Different Age Groups
Shikun Zhou, Yuming Fang 0001, Tingsong Lu |
ICMI | 3 |
| 2025 | What Happens in the Surroundings: A Benchmark for 360° image CaptioningabstractImage captioning has been widely studied by the computer vision and natural language processing communities. However, conventional image captioning models are mainly built upon 2D images with narrow field-of-views. To comprehensively analyze what happens in real scenes, using 360° cameras to capture 360° images has been a popular research trend. However, 360° image captioning is rarely studied due to the lack of related datasets. To bridge the research gap, we introduce a novel dataset for 360° image captioning, namely 360IC (360° Image Captioning). It contains 1250 360° images from rich scenes, and each image is manually labeled with at least three detailed descriptions, which will greatly promote the research of 360° image captioning. We also propose a Multi-View Transformer Network (MVTransNet) for 360° image captioning based on multi-view analysis and fusion, which deals with the characteristics of large resolution, wide field-of-view and complex visual content of 360° images. Specifically, it builds a hierarchical architecture to model the spatial dependency of images with larger content, thus forming rich visual features of 360° images and making the generated descriptions more accurate. Extensive experiments on 360IC show that the proposed network outperforms other competing methods considerably. Our dataset will be released soon. Wenhui Jiang 0001, Tiancong Xu, Zichen Li, Yuming Fang 0001 |
IJCNN | 5 |
| 2025 | Scale Margin Loss for Object Detection
Yuxuan Cheng, Yanjun Zhang 0002, Leo Yu Zhang, Donald Donglong Chen, Yuming Fang 0001 |
KSEM (2) | 5 |
| 2025 | Weak-shot Keypoint Estimation via Keyness and Correspondence TransferabstractKeypoint estimation is a fundamental task in computer vision, but generally requires large-scale annotated data for training. Few-shot and unsupervised keypoint estimation are prevalent economical paradigms, but the former still requires annotations for extensive novel classes while the latter only supports for single class. In this paper, we focus on the task of weak-shot keypoint estimation, where multiple novel classes are learned from unlabeled images with the help of labeled base classes. The key problem is what to transfer from base classes to novel classes, and we propose to transfer keyness and correspondence, which essentially belong to comparing entities and thus are class-agnostic and class-wise transferable. The keyness compares which pixel in the local region is more key, which can guide the keypoints of novel classes to move towards the local maximum (i.e., obtaining keypoints). The correspondence compares whether the two pixels belongs to the same semantic part, which can activate the keypoints of novel classes by reinforcing the consistency between corresponding points on two paired images. By transferring keyness and correspondence, our framework achieves favourable performance for weak-shot keypoint estimation. Extensive experiments and analyses on large-scale benchmark MP-100 demonstrate our effectiveness. Junjie Chen 0008, Zeyu Luo, Zezheng Liu, Wenhui Jiang 0001, Li Niu 0002, Yuming Fang 0001 |
NeurIPS | 6 |
| 2025 | Atkscopes: Multiresolution Adversarial Perturbation as a Unified Attack on Perceptual Hashing and Beyond
Yushu Zhang 0001, Zhongyun Hua, Wenying Wen, Yuming Fang 0001 |
USENIX Security Symposium | 6 |
| 2025 | Frequency-Aware Native Resolution Assessment of 8K Omnidirectional ImagesabstractOmnidirectional images (ODIs) serve as fundamental visual medium for presenting virtual reality (VR) contents, supporting fully immersive experiences through 360-degree scene representation. Typically, a high pixel density is essential for visual quality in VR environments, which in turn requires sufficiently high-resolution imagery to achieve. However, capturing native high-resolution ODIs requires expensive omnidirectional cameras with large sensors (e.g., Insta360 TITAN). An alternative approach is to use low-resolution cameras to acquire original images and then enhance their resolution via super-resolution algorithms. In this work, we explore whether super-resolution ODIs can be easily distinguished from native high-resolution ODIs at 8K scale. To this end, we firstly construct the Native Resolution Assessment of 8K Omnidirectional Images (NRA- 8KODI) dataset, whose native 8K ODIs are collected with an Insta360 TITAN camera and 8K super-resolution images are generated from SOTA open-sourced algorithms. Recognizing high-frequency signals are essential for differentiating non-native 8K ODIs, a frequency-aware model is designed to capture high-frequency details. Specially, to maintain high-frequency details kept in high-resolutions while reduce computational costs brought by high-resolutions, we propose a frequency-aware compressor module to suppress feature channels dominated by low-frequency details. Finally, our model achieves 97.2% accuracy in detecting non-native 8K ODIs, implying that super-resolution for ODIs can still be improved for visual experience in VR applications. Jingwen Hou, Zengliang Li, Jiebin Yan, Weide Liu, Yuming Fang 0001, Wei Zhou 0021 |
VCIP | 5 |
| 2025 | Texture-Aware Network for Enhancing Inner Smoke Representation in Visual Smoke Density EstimationabstractABSTRACT Smoke often appears before visible flames in the early stages of fire disasters, making accurate pixel‐wise detection essential for fire alarms. Although existing segmentation models effectively identify smoke pixels, they generally treat all pixels within a smoke region as having the same prior probability. This assumption of rigidity, common in natural object segmentation, fails to account for the inherent variability within smoke. We argue that pixels within smoke exhibit a probabilistic relationship with both smoke and background, necessitating density estimation to enhance the representation of internal structures within the smoke. To this end, we propose enhancements across the entire network. First, we improve the backbone by adaptively integrating scene information into texture features through separate paths, enabling smoke‐tailored feature representation for further exploit. Second, we introduce a texture‐aware head with long convolutional kernels to integrate both global and orientation‐specific information, enhancing representation for intricate smoke structure. Third, we develop a dual‐task decoder for simultaneous density and location recovery, with the frequency‐domain alignment in the final stage to preserve internal smoke details. Extensive experiments on synthetic and real smoke datasets demonstrate the effectiveness of our approach. Specifically, comparisons with 17 models show the superiority of our method, with mean IoU improvements of 4.88%, 2.63%, and 3.17% on three test sets. (The code will be available on https://github.com/xia‐xx‐cv/TANet_smoke ). Xue Xia 0005, Yajing Peng, Zichen Li, Jinting Shi, Yuming Fang 0001 |
IET Comput. Vis. | 5 |
| 2025 | Improving multi-modal brain tumor segmentation via pre-training and knowledge distillation based post-training
Weide Liu, Jingwen Hou, Xiaoyang Zhong, Huijing Zhan, Jun Cheng 0003, Yuming Fang 0001, Guanghui Yue 0001 |
Neurocomputing | 6 |
| 2025 | Integrating large foundation models into multimodal named entity recognition with evidential fusion
Weide Liu, Xiaoyang Zhong, Jingwen Hou, Haozhe Huang, Wei Zhou 0021, Yuming Fang 0001 |
Neurocomputing | 7 |
| 2025 | Opinion-unaware blind quality assessment of AI-generated omnidirectional images based on deep feature statistics
Xuelin Liu, Jiebin Yan, Yuming Fang 0001, Jingwen Hou |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | Hierarchical boundary feature alignment network for video salient object detectionabstractThe deep learning based video salient object detection (VSOD) models have achieved great success in the past few years, however, these VSOD models still suffer from the following two problems: i) struggle in accurately predicting those pixels surrounding salient objects; ii) unaligned features of different scales lead to deviations in feature fusion . To tackle these problems, we propose a hierarchical boundary feature alignment network (HBFA). Specifically, the proposed HBFA consists of a temporal–spatial fusion module (TSM) and three decoding branches. TSM captures multi-scale spatiotemporal information. The two boundary feature branches are used to guide the whole network to pay more attention to the boundary of salient objects, while the feature alignment branch is capable of fusing the features from the internal and external branches while aligning features across different scales. Our extensive experiments show that the proposed method reaches a new state-of-the-art performance. Amin Mao, Jiebin Yan, Yuming Fang 0001, Hantao Liu |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | Opinion-unaware blind stereoscopic image quality assessment: A comprehensive study
Jiebin Yan, Yuming Fang 0001, Xuelin Liu, Wenhui Jiang 0001, Yang Liu 0293 |
Pattern Recognit. | 2 |
| 2025 | Max360IQ: Blind omnidirectional image quality assessment with multi-axis attention
Jiebin Yan, Ziwen Tan, Yuming Fang 0001, Jiale Rao, Yifan Zuo 0001 |
Pattern Recognit. | 3 |
| 2025 | Learning Stage-wise Fusion Transformer for light field saliency detection
Wenhui Jiang 0001, Qi Shu, Hongwei Cheng, Yuming Fang 0001, Yifan Zuo 0001 |
Pattern Recognit. Lett. | 4 |
| 2025 | Viewport-Independent Blind Quality Assessment of AI-Generated Omnidirectional Images via Vision-Language CorrespondenceabstractThe advancement of deep generation technology has significantly enhanced the growth of artificial intelligencegenerated content (AIGC). Among these, AI-generated omnidirectional images (AGOIs), hold considerable promise for applications in virtual reality (VR). However, the quality of AGOIs varies widely, and there has been limited research focused on their quality assessment. In this letter, inspired by the characteristics of the human visual system, we propose a novel viewportindependent blind quality assessment method for AGOIs, termed VI-AGOIQA, which leverages vision-language correspondence. Specifically, to minimize the computational burden associated with viewport-based prediction methods for omnidirectional image quality assessment, a set of image patches are first extracted from AGOIs in Equirectangular Projection (ERP) format. Then, the correspondence between visual and textual inputs is effectively learned by utilizing the pre-trained image and text encoders of the Contrastive Language-Image Pre-training (CLIP) model. Finally, a multimodal feature fusion module is applied to predict human visual preferences based on the learned knowledge of visual-language consistency. Extensive experiments conducted on publicly available database demonstrate the promising performance of the proposed method. The source code will be made available at https://github.com/LXLHXL123/VI-AGOIQA. Xuelin Liu, Jiebin Yan, Chenyi Lai, Yuming Fang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Towards Scalable and Efficient Full-Reference Omnidirectional Image Quality AssessmentabstractFull-Reference (FR) image quality assessment (IQA) (FR-IQA) has achieved notable success due to its irreplaceable role in algorithm and system optimization; however, it has less been investigated in omnidirectional image quality assessment (OIQA). In this paper, we make an attempt to FR-OIQA considering the constraint of the computation budget, in which this issue is formulated as “quality perception from patch to sequence”,i.e.,Intra-PatchSequence degradation modeling andInter-PatchSequence similarity calculation (denoted by IPS$^{2}$). Specifically, IPS$^{2}$directly accepts local patches from the omnidirectional image (OI) in the format of Equirectangular Projection as input, avoiding other preprocessing operations, such as scan-path prediction and projection transformation. Subsequently, IPS$^{2}$uses a deep feature extractor to capture patch quality and then sends the patch- wise quality maps to the cross-patch similarity (CPS) module, which explicitly models intra-patch sequence degradation and inter-patch sequence similarity via self-attention. Finally, a quality regressor is used to aggregate these features of the CPS module and predict the global quality of the OI. The experimental results on a large-scale OIQA database show that the proposed IPS$^{2}$outperforms most state-of-the-art methods in quality prediction accuracy while offering substantial reductions in computational cost and model size. Jiebin Yan, Zhihua Wang 0002, Yuming Fang 0001, Hantao Liu |
IEEE Signal Process. Lett. | 4 |
| 2025 | Multiscale Feature-Guided Adversarial Examples Quality Assessment via Hierarchical Perception of Human Visual SystemabstractDeep neural networks (DNNs) reveal significant robustness deficiencies due to their susceptibility to being misled by small and imperceptible adversarial examples, thus it is crucial to improve the robustness of DNNs against such harmful perturbations. The current$L_{p}$specification ignores differences in human visual perception when measuring similarity, and most existing image quality assessment (IQA) methods and adversarial example datasets lack subjective scores for evaluation. In this paper, we construct a new database of adversarial examples, called the AED, which contains 35 original images, 1050 adversarial examples, and the corresponding subjective scores of adversarial examples. Then, a novel full-reference IQA model for the quality evaluation of the adversarial examples is proposed by taking into full consideration the hierarchical perception of human visual system (HVS) and the outstanding capabilities of the multi-scale feature extraction network in feature extraction. Specifically, a feature encoding network that uses continuous convolution layers to pre-extract features and expand the receptive field of the image is employed. To simulate the HVS hierarchical perception, the features of different scales are further obtained by designing a multi-scale feature extraction network. The structural similarity scores of the feature maps at different scales are calculated for jointly arriving at the final IQA score of the adversarial examples. Experimental results have demonstrated that our proposed model is closer to the perception of HVS in small imperceptible distortions evaluation of adversarial examples compared with other classical and state-of-the-art models. Wenying Wen, Minghui Huang, Li Dong 0006, Yushu Zhang 0001, Yuming Fang 0001 |
IEEE Trans. Big Data | 5 |
| 2025 | MLVPP: Multilevel Visual Privacy Protection via Thumbnail Preservation and Key SharingabstractNowadays, shared social images often contain multiple privacy subjects, and the disclosure of these privacy subjects increases the risk of privacy violations; thus, the protection of visual privacy is particularly important. However, existing means of visual privacy protection render images unavailable and are typically protected only for images with a single privacy subject. Combining compressed sensing (CS) and 2DCS, this article proposes a multilevel visual privacy protection scheme (MLVPP) via thumbnail preservation (TPE) and key sharing (KS), which includes two stages, that is, CS-TPE multilevel encryption and KS multilevel decryption. In the first stage, we utilize 2DCS to enable the compressed sampled observations to preserve the structural similarity of the nonsensitive part and leverage CS to encrypt multiple sensitive parts of the image. The CS-processed image is then made to strike a good balance between privacy and availability through TPE. In the second stage, with the help of KS mechanism, the CS encryption keys and TPE key are shared in different combinations with the users related to the privacy subjects for meeting the MLVPP requirements. The quality of decrypted images varies for different classes of related users, e.g., unrelated users, semirelated users, and full-related users. The proposed approach preserves the visual information in the nonsensitive part of the image while providing security for the sensitive parts. Compared with related works, MLVPP reveals that the PSNR of the fully decrypted image can reach 34.04, and the SSIM is greater than 0.96 at a compression rate of 0.25. Experimental results demonstrate the superior performance of MLVPP. Wenying Wen, Haigang Huang, Yushu Zhang 0001, Yuming Fang 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Separate, Locate, and Align: Determine Context Relation of Scene Text From Multiple Perspectives in TextVQAabstractText-based Visual Question Answering (TextVQA) focuses on answering questions about the scene text in images. Most works in this field uses transformer based models to modeling the interaction of question and scene texts which means the scene texts will be treated as a natural language sentence and concatenated in reading order as a part of input. However, they ignore the fact that different from words in natural language sentence which have inherent context relation, the context relation of scene texts in images need to be determined. To tackle this problem, we propose a novel method named Separate, Locate and Align (SLA) that discriminate the context relation of scene texts from semantic, visual and spatial aspects. Specifically, based on scene texts with similar visual information (e.g. background color, font color, font style, etc.) having semantic contextual relations, we propose a Text Semantic Separate (TSS) module to discriminate the semantic relation between different scene texts according to their visual contextual information. Then, we introduce a Spatial Circle Position (SCP) module that helps the model discriminate the spatial relation between different scene texts. Last, we design a Visual Alignment (VA) module to help the model distinguish the visual relationships between different scene texts according to the color distribution differences. Extensive experiments show that our method outperforms existing alternatives on TextVQA and ST-VQA datasets without pre-training tasks. Chengyang Fang, Wenhui Jiang 0001, Yuming Fang 0001, Yuxin Peng 0001, Yang Liu 0293 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Perceptual Transform Fusion of Infrared and Visible ImagesabstractInfrared and visible image fusion aims to generate fused images with rich textures and clear target representations. Existing methods generally assume high-quality input images, thus overlooking issues such as reduced contrast and loss of details in visible images under low-light conditions. The naive enhance-then-fuse strategy cannot perform fuse-oriented image enhancement, which always reaches a sub-optimal result. To address this challenge, we propose a perceptual transform fusion of infrared and visible images, which simultaneously optimizes low-light enhancement and image fusion. Specifically, to improve computational efficiency and optimize key feature representations while suppressing noise interactions caused by lighting variations, we introduce a lightweight adaptive sparse Transformer block (ASTBlock). This model adaptively integrates sparse and dense attention mechanisms to enhance feature representations and employs a feed-forward network to eliminate redundant information, thereby ensuring the quality of image fusion. Subsequently, to retain significant details while reducing the impact of noise introduced by low-light enhancement, we incorporate discrete wavelet transform (DWT) for feature decomposition and fusion, further enhancing the representation capability and feature preservation of fused images. Meanwhile, to tackle the issues of insufficient contrast and hidden details in low-light conditions, we design an illumination perception module and an illumination consistency loss to improve the contrast and clarity of fused images. Experimental results on multiple public benchmark datasets for quality assessment and downstream tasks, e.g., pedestrian detection, demonstrate that our method significantly outperforms the state-of-the-art (SOTA) methods. The code is available at https://github.com/hinmouc/PIVFusion. Dingli Hua, Qingmao Chen, Zhiliang Wu, Yifan Zuo 0001, Wenying Wen, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Learning Comprehensive Visual Grounding for Video CaptioningabstractThe grounding accuracy of existing video captioners is still behind the expectation. The majority of existing methods perform grounded video captioning on sparse entity annotations. However, grounded captioning models rely on deliberate grounding annotations as supervision, which are relatively hard to obtain. Moreover, the captioning accuracy often suffers from degenerated object appearances on the annotated area such as motion blur and video defocus, and these models seldom consider the complex interactions among entities. In this paper, we propose a comprehensive visual grounding network to improve video captioning, by using inexpensive pseudo annotation while avoiding the need to collect large amounts of manual annotations. Specifically, the network consists of spatial-temporal entity grounding and action grounding. The proposed entity grounding encourages the attention mechanism to focus on informative spatial areas across video frames. The action grounding dynamically associates the verbs to related subjects and the corresponding context, which keeps fine-grained spatial and temporal details for action prediction. Both entity grounding and action grounding are formulated as a unified task guided by a soft grounding supervision. More importantly, the grounding objective is supervised by pseudo annotations automatically produced by a grounding annotation generation module, thus our model can be easily applied to the challenging dataset without any grounding annotation provided. We conduct extensive experiments on three benchmark datasets and demonstrate significant performance improvements of +2.4 CIDEr on MSR-VTT, +4.7 CIDEr on MSVD, and +5.1 CIDEr on ActivityNet-Entities compared to state-of-the-arts. Wenhui Jiang 0001, Linxin Liu, Yuming Fang 0001, Yibo Cheng, Yuxin Peng 0001, Yang Liu 0293 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | UNI-IQA: A Unified Approach for Mutual Promotion of Natural and Screen Content Image Quality AssessmentabstractTo date, the image quality assessment (IQA) research field has mainly focused on natural images (NIs)-based IQA and screen content images (SCIs)-based IQA. Usually, these two research branches are quite independent due to the large differences between NIs and SCIs, where NIs, captured by cameras directly, contain pictorial information solely, yet, SCIs, synthesized or GPU-rendered, have pictures and textures. Moreover, the distortion types are also different, and subjective scores of different datasets assigned by participants are usually not well aligned. So, due to the above-mentioned “domain shifts” and “dataset misalignments”, our research community has widely believed that it could be very difficult to achieve joint mutual promotions between NIs- and SCIs-based IQA. In this paper, we argue that despite the “differences”, there still are some “common characteristics” — our human visual system perceives the “pictures” in both SCIs and NIs almost the same way. Thus, we can still achieve mutual performance promotion if we can appropriately use the “common characteristics” between SCIs and NIs. Our key idea is to devise a “content-aware” data switch, which, from the perspective of input’s contents (i.e., pictures or textures), aims at letting the model automatically enhance the commonness and compress the discrepancies between the two tasks. Notice that none of the existing fusion schemes can reach this goal since they are actually content-unaware, degenerating the “mutual interactions” into “mutual interferences”. This paper is the first attempt to achieve full end-to-end “mutual interactions” between NIs- and SCIs-based IQA. Using the proposed switch, we are also the first to achieve solid mutual promotions for the two tasks, reaching new SOTA results. Mengke Song, Chenglizhao Chen, Wenfeng Song, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Spatial Quality Oriented Rate Control for Volumetric Video Streaming via Deep Reinforcement LearningabstractVolumetric videos offer an incredibly immersive viewing experience but encounters challenges in maintaining quality of experience (QoE) due to its ultra-high bandwidth requirements. One significant challenge stems from user’s spatial interactions, potentially leading to discrepancies between transmission bitrates and the actual quality of rendered viewports. In this study, we conduct comprehensive measurement experiments to investigate the impact of six degrees of freedom information on received video quality. Our results indicate that the correlation between spatial quality and transmission bitrates is influenced by the user’s viewing distance, exhibiting variability among users. To address this, we propose a spatial quality oriented rate control system, namely sparkle, that aims to satisfy spatial quality requirements while maximizing long-term QoE for volumetric video streaming services. Leveraging richer user interaction information, we devise a tailored learning-based algorithm to enhance long-term QoE. To address the complexity brought by richer state input and precise allocation, we integrate pre-constraints derived from three-dimensional displays to intervene action selection, efficiently reducing the action space and speeding up convergence. Extensive experimental results illustrate that sparkle significantly enhances the averaged QoE by up to 29% under practical network and user tracking scenarios. Xi Wang 0050, Wei Liu 0004, Shimin Gong, Zhi Liu 0002, Jing Xu 0005, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Multitask Auxiliary Network for Perceptual Quality Assessment of Non-Uniformly Distorted Omnidirectional ImagesabstractOmnidirectional image quality assessment (OIQA) has been widely investigated in the past few years and achieved much success. However, most of existing studies are dedicated to solve the uniform distortion problem in OIQA, which has a natural gap with the non-uniform distortion problem, and their ability in capturing non-uniform distortion is far from satisfactory. To narrow this gap, in this paper, we propose a multitask auxiliary network for non-uniformly distorted omnidirectional images, where the parameters are optimized by jointly training the main task and other auxiliary tasks. The proposed network mainly consists of three parts: a backbone for extracting multiscale features from the viewport sequence, a multitask feature selection module for dynamically allocating specific features to different tasks, and auxiliary sub-networks for guiding the proposed model to capture local distortion and global quality change. Extensive experiments conducted on two large-scale OIQA databases demonstrate that the proposed model outperforms other state-of-the-art OIQA metrics, and these auxiliary sub-networks contribute to improve the performance of the proposed model. The source code is available athttps://github.com/RJL2000/MTAOIQA. Jiebin Yan, Jiale Rao, Junjie Chen 0008, Ziwen Tan, Weide Liu, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Beyond Privacy: Generating Privacy-Preserving Faces Supporting Robust Image AuthenticationabstractThe prevalence of face capturing along with the advancement of face recognition poses a potential threat to individual privacy. To protect privacy, plenty of methods have been proposed to change identity in the face, thus blocking malicious face recognition. However, these methods fail to satisfy authentication requirements for special application scenarios, e.g., face authentication in surveillance capture. In this paper, we propose a novel face privacy protection model, which supports robust image authentication via information-conditional identity transformation. Specifically, we first introduce a basic face manipulation model (FMM), which can preserve identity-irrelevant attributes when manipulating identity. Based on FMM, we further design a lightweight protector called AIDPro, outputting a transformed identity which is different from the original one and is embedded a message presenting authentication information. Benefiting from the semantic robustness, our model does not require noise layers to achieve accurate message extraction after various image distortions. In addition, the message can be the condition to guide the identity transformation for privacy protection, which avoids extra resource consumption from supporting image authentication. Extensive experimental results demonstrate our model has comparable privacy protection performance, superior attribute preservation performance, and robust authentication performance especially in JPEG compression and screen shooting. Our code is available athttps://github.com/daizigege/AIDPro. Tao Wang 0084, Wenying Wen, Xiangli Xiao, Zhongyun Hua, Yushu Zhang 0001, Yuming Fang 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Webly Supervised Fine-Grained Classification by Integrally Tackling Noises and Subtle DifferencesabstractWebly-supervised fine-grained visual classification (WSL-FGVC) aims to learn similar sub-classes from cheap web images, which suffers from two major issues: label noises in web images and subtle differences among fine-grained classes. However, existing methods for WSL-FGVC only focus on suppressing noise at image-level, but neglect to mine cues at pixel-level to distinguish the subtle differences among fine-grained classes. In this paper, we propose a bag-level top-down attention framework, which could tackle label noises and mine subtle cues simultaneously and integrally. Specifically, our method first extracts high-level semantic information from a bag of images belonging to the same class, and then uses the bag-level information to mine discriminative regions in various scales of each image. Besides, we propose to derive attention weights from attention maps to weight the bag-level fusion for a robust supervision. We also propose an attention loss on self-bag attention and cross-bag attention to facilitate the learning of valid attention. Extensive experiments on four WSL-FGVC datasets, i.e., Web-Aircraft, Web-Bird, Web-Car, and WebiNat-5089, demonstrate the effectiveness of our method against the state-of-the-art methods. Junjie Chen 0008, Jiebin Yan, Yuming Fang 0001, Li Niu 0002 |
IEEE Trans. Image Process. | 3 |
| 2025 | SRENet: Saliency-Based Lighting Enhancement NetworkabstractLighting enhancement is a classical topic in low-level image processing. Existing studies mainly focus on global illumination optimization while overlooking local semantic objects, and this limits the performance of exposure compensation. In this paper, we introduce SRENet, a novel lighting enhancement network guided by saliency information. It adopts a two-step strategy of foreground-background separation optimization to achieve a balance between global and local illumination. In the first step, we extract salient regions and implement the local illumination enhancement that ensures the exposure quality of salient objects. Next, we utilize a fusion module to process global lighting optimization based on local enhanced results. With the two-step strategy, the proposed SRENet yield better lighting enhancement for local illumination while preserving the globally optimal results. Experimental results demonstrate that our method obtains more effective enhancement results for various tasks of exposure correction and lighting quality improvement. The source code and pre-trained models are available at https://github.com/PlanktonQAQ/SRENet. Yuming Fang 0001, Chenlei Lv, Weisi Lin |
IEEE Trans. Image Process. | 1 |
| 2025 | Toward Transparent Deep Image Aesthetics Assessment With Tag-Based Content DescriptorsabstractDeep learning approaches for Image Aesthetics Assessment (IAA) have shown promising results in recent years, but the internal mechanisms of these models remain unclear. Previous studies have demonstrated that image aesthetics can be predicted using semantic features, such as pre-trained object classification features. However, these semantic features are learned implicitly, and therefore, previous works have not elucidated what the semantic features are representing. In this work, we aim to create a more transparent deep learning framework for IAA by introducing explainable semantic features. To achieve this, we propose Tag-based Content Descriptors (TCDs), where each value in a TCD describes the relevance of an image to a human-readable tag that refers to a specific type of image content. This allows us to build IAA models from explicit descriptions of image contents. We first propose the explicit matching process to produce TCDs that adopt predefined tags to describe image contents. We show that a simple MLP-based IAA model with TCDs only based on predefined tags can achieve an SRCC of 0.767, which is comparable to most state-of-the-art methods. However, predefined tags may not be sufficient to describe all possible image contents that the model may encounter. Therefore, we further propose the implicit matching process to describe image contents that cannot be described by predefined tags. By integrating components obtained from the implicit matching process into TCDs, the IAA model further achieves an SRCC of 0.817, which significantly outperforms existing IAA methods. Both the explicit matching process and the implicit matching process are realized by the proposed TCD generator. To evaluate the performance of the proposed TCD generator in matching images with predefined tags, we also labeled 5101 images with photography-related tags to form a validation set. And experimental results show that the proposed TCD generator can meaningfully assign photography-related tags to images. Jingwen Hou, Weisi Lin, Yuming Fang 0001, Haoning Wu 0001, Chaofeng Chen, Weide Liu |
IEEE Trans. Image Process. | 3 |
| 2025 | Diffusion-Based Facial Aesthetics Enhancement With 3D Structure GuidanceabstractFacial Aesthetics Enhancement (FAE) aims to improve facial attractiveness by adjusting the structure and appearance of a facial image while preserving its identity as much as possible. Most existing methods adopted deep feature-based or score-based guidance for generation models to conduct FAE. Although these methods achieved promising results, they potentially produced excessively beautified results with lower identity consistency or insufficiently improved facial attractiveness. To enhance facial aesthetics with less loss of identity, we propose the Nearest Neighbor Structure Guidance based on Diffusion (NNSG-Diffusion), a diffusion-based FAE method that beautifies a 2D facial image with 3D structure guidance. Specifically, we propose to extract FAE guidance from a nearest neighbor reference face. To allow for less change of facial structures in the FAE process, a 3D face model is recovered by referring to both the matched 2D reference face and the 2D input face, so that the depth and contour guidance can be extracted from the 3D face model. Then the depth and contour clues can provide effective guidance to Stable Diffusion with ControlNet for FAE. Extensive experiments demonstrate that our method is superior to previous relevant methods in enhancing facial aesthetics while preserving facial identity. Lisha Li, Jingwen Hou, Weide Liu, Yuming Fang 0001, Jiebin Yan |
IEEE Trans. Image Process. | 4 |
| 2025 | Perceptual Quality Assessment of 360° Images Based on Generative Scanpath RepresentationabstractDespite substantial efforts dedicated to the design of heuristic models for omnidirectional (i.e., 360°) image quality assessment (OIQA), a conspicuous gap remains due to the lack of consideration for the diversity of viewing behaviors that leads to the varying perceptual quality of 360° images. Two critical aspects underline this oversight: the neglect of viewing conditions that significantly sway user gaze patterns and the overreliance on a single viewport sequence from the 360° image for quality inference. To address these issues, we introduce a unique generative scanpath representation (GSR) for effective quality inference of 360° images, which aggregates varied perceptual experiences of multi-hypothesis users under a predefined viewing condition. More specifically, given a viewing condition characterized by the starting point of viewing and exploration time, a set of scanpaths consisting of dynamic visual fixations can be produced using an apt scanpath generator. Following this vein, we use the scanpaths to convert the 360° image into the unique GSR, which provides a global overview of gazed-focused contents derived from scanpaths. As such, the quality inference of the 360° image is swiftly transformed to that of GSR. We then propose an efficient OIQA computational framework by learning the quality maps of GSR. Comprehensive experimental results validate that the predictions of the proposed framework are highly consistent with human perception in the spatiotemporal domain, especially in the challenging context of locally distorted 360° images under varied viewing conditions. The code will be released at https://github.com/xiangjieSui/GSR. Xiangjie Sui, Hanwei Zhu, Xuelin Liu, Yuming Fang 0001, Shiqi Wang 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Visual Quality Assessment of Composite Images: A Compression-Oriented Database and MeasurementabstractComposite images (CIs) have experienced unprecedented growth, especially with the prosperity of a large number of generative AI technologies. They are usually created by combining multiple visual elements from different sources to form a single cohesive composition, which have an increasing impact on a variety of vision applications. However, transmission of CIs can degrade their visual quality, especially undergoing lossy compression to reduce bandwidth and storage. To facilitate the development of objective measurements for CIs and investigate the influence of compression distortions on their perception, we establish a compression-oriented image quality assessment (CIQA) database for CIs (called ciCIQA) with 30 typical encoding distortions. Compressed with six representative codecs, we have carried out a large-scale subjective experiment that delivered 3,000 encoded CIs with labeled quality scores, making ciCIQA one of the earliest CI databases with the most compression types. ciCIQA enables us to explore the encoding effects on visual quality from the first five just noticeable difference (JND) points, offering insights for perceptual CI compression and related tasks. Moreover, we have proposed a new multi-masked no-reference CIQA method(called mmCIQA), including a multi-masked quality representation module, a self-supervised quality alignment module, and a multi-masked attentive fusion module. Experimental results demonstrate the outstanding performance of our mmCIQA in assessing the quality of CIs, outperforming 17 competitive approaches. The proposed method and database as well as the collected objective metrics are made publicly available on https://charwill.github.io/mmciqa.html. Miaohui Wang, Zhuowei Xu, Yuming Fang 0001, Weisi Lin |
IEEE Trans. Image Process. | 4 |
| 2025 | Omnidirectional Image Quality Captioning: A Large-Scale Database and a New ModelabstractThe fast growing application of omnidirectional images calls for effective approaches for omnidirectional image quality assessment (OIQA). Existing OIQA methods have been developed and tested on homogeneously distorted omnidirectional images, but it is hard to transfer their success directly to the heterogeneously distorted omnidirectional images. In this paper, we conduct the largest study so far on OIQA, where we establish a large-scale database called OIQ-10K containing 10,000 omnidirectional images with both homogeneous and heterogeneous distortions. A comprehensive psychophysical study is elaborated to collect human opinions for each omnidirectional image, together with the spatial distributions (within local regions or globally) of distortions, and the head and eye movements of the subjects. Furthermore, we propose a novel multitask-derived adaptive feature-tailoring OIQA model named IQCaption360, which is capable of generating a quality caption for an omnidirectional image in a manner of textual template. Extensive experiments demonstrate the effectiveness of IQCaption360, which outperforms state-of-the-art methods by a significant margin on the proposed OIQ-10K database. The OIQ-10K database and the related source codes are available at https://github.com/WenJuing/IQCaption360. Jiebin Yan, Ziwen Tan, Yuming Fang 0001, Junjie Chen 0008, Wenhui Jiang 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Learning Guided Implicit Depth Function With Scale-Aware Feature FusionabstractRecently, the single image super-resolution based on implicit image function is a hot topic, which learns a universal model for arbitrary upsampling scales. By contrast, color-guided depth map super-resolution is less explored based on implicit function learning. The related research faces three questions. First, is it also necessary and applicable to fuse the depth feature and the color feature in the encoder with continuous upsampling scales? Second, is the scale information in the encoder as important as that in the decoder? Third, how to efficiently and effectively model the affinity of location distance and content similarity within cross domains in the decoder? This paper proposes a transformer-based network to answer the above questions, which includes a depth super-resolution branch and a guidance extraction branch. Specifically, in the encoder, the effective implicit cross transformer is designed to fuse the guidance from the color feature with continuous coordinate mapping. In addition, the unrelated guidance is filtered out by correlation evaluation in the high-dimension feature space. Unlike the scale only introduced in the decoder, this paper additionally embeds the scale into the position encoding and the feed-forward network in the encoder to learn the scale-aware feature representation. In the decoder, the high-resolution depth feature is reconstructed by using the internal prior and the external guidance. The internal prior is implemented by implicit self-attention in the depth super-resolution branch, and the external guidance is exploited via implicit cross-attention between both branches. Finally, the above decoded features are complementary to generate the high-resolution depth map. The sufficient experiments on the synthetic and real datasets for in-distribution and out-of-distribution upsampling scales validate the improved performance. The code and the models are public via https://github.com/NaNRan13/GIDF. Yifan Zuo 0001, Yuming Fang 0001, Jiebin Yan, Wenhui Jiang 0001, Yuxin Peng 0001, Yan Huang 0023 |
IEEE Trans. Image Process. | 5 |
| 2025 | Subjective and Objective Quality Assessment of Non-Uniformly Distorted Omnidirectional ImagesabstractOmnidirectional image quality assessment (OIQA) has been one of the hot topics in IQA with the continuous development of VR techniques, and achieved much success in the past few years. However, most studies devote themselves to the uniform distortion issue, i.e., all regions of an omnidirectional image are perturbed by the “same amount” of noise, while ignoring the non-uniform distortion issue, i.e., partial regions undergo “different amount” of perturbation with the other regions in the same omnidirectional image. Additionally, nearly all OIQA models are verified on the platforms containing a limited number of samples, which largely increases the over-fitting risk and therefore impedes the development of OIQA. To alleviate these issues, we elaborately explore this topic from both subjective and objective perspectives. Specifically, we construct a large OIQA database containing 10,320 non-uniformly distorted omnidirectional images, each of which is generated by considering quality impairments on one or two camera len(s). Then we meticulously conduct psychophysical experiments and delve into the influence of both holistic and individual factors (i.e., distortion range and viewing condition) on omnidirectional image quality. Furthermore, we propose a perception-guided OIQA model for non-uniform distortion by adaptively simulating users' viewing behavior. Experimental results demonstrate that the proposed model outperforms state-of-the-art methods. Jiebin Yan, Jiale Rao, Xuelin Liu, Yuming Fang 0001, Yifan Zuo 0001, Weide Liu |
IEEE Trans. Multim. | 4 |
| 2025 | Computational Analysis of Degradation Modeling in Blind Panoramic Image Quality AssessmentabstractBlind panoramic image quality assessment (BPIQA) has recently brought a new challenge to the visual quality community, due to the complex interaction between immersive content and human behavior. Although many efforts have been made to advance BPIQA from both conducting psychophysical experiments and designing performance-driven objective algorithms, limited content and few samples in those closed sets inevitably would result in shaky conclusions, thereby hindering the development of BPIQA; we refer to it as the easy-database issue. In this article, we present a sufficient computational analysis of degradation modeling in BPIQA to thoroughly explore the easy-database issue , where we carefully design three types of experiments via investigating the gap between BPIQA and blind image quality assessment (BIQA), the necessity of specific design in BPIQA models, and the generalization ability of BPIQA models. From extensive experiments, we find that easy databases narrow the gap between the performance of BPIQA and BIQA models, which is unconducive to the development of BPIQA. And the easy databases make the BPIQA models be closed to saturation; therefore, the effectiveness of the associated specific designs cannot be well verified. Besides, the BPIQA models trained on our recently proposed databases with complicated degradation show better generalization ability. Thus, we believe that much more efforts are highly desired to put into BPIQA from both subjective viewpoint and objective viewpoint. Jiebin Yan, Ziwen Tan, Jiale Rao, Yifan Zuo 0001, Yuming Fang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Viewport-Unaware Blind Omnidirectional Image Quality Assessment: A Flexible and Effective ParadigmabstractMost of the existing blind omnidirectional image quality assessment (BOIQA) models rely on viewport generation by modeling user viewing behavior or transforming omnidirectional images (OIs) into varying formats; however, these methods are either computationally expensive or less scalable. To solve these issues, in this article, we present a flexible and effective paradigm, which is viewport-unaware and can be easily adapted to 2D plane image quality assessment (2D-IQA). Specifically, the proposed BOIQA model includes an adaptive prior-equator sampling module for extracting a patch sequence from the equirectangular projection (ERP) image in a resolution-agnostic manner, a progressive deformation-unaware feature fusion module which is able to capture patch-wise quality degradation in a deformation-immune way, and a local-to-global quality aggregation module to adaptively map local perception to global quality. Extensive experiments across four OIQA databases (including uniformly distorted OIs and non-uniformly distorted OIs) demonstrate that the proposed model achieves competitive performance with low complexity against other state-of-the-art models, and we also verify its adaptive capacity to 2D-IQA. The source code is available at https://github.com/KangchengWu/OIQA . Jiebin Yan, Kangcheng Wu, Junjie Chen 0008, Ziwen Tan, Yuming Fang 0001, Weide Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | A Perceptually Optimized and Self-Calibrated Tone Mapping OperatorabstractWith the increasing popularity and accessibility of high dynamic range (HDR) photography, tone mapping operators (TMOs) for dynamic range compression are practically demanding. In this paper, we develop a two-stage neural network-based TMO that is self-calibrated and perceptually optimized. In Stage one, motivated by the physiology of the early stages of the human visual system, we first decompose an HDR image into a normalized Laplacian pyramid. We then use two lightweight deep neural networks, taking the normalized representation as input and estimating the Laplacian pyramid of the corresponding LDR image. We optimize the tone mapping network by minimizing the normalized Laplacian pyramid distance, a perceptual metric aligning with human judgments of tone-mapped image quality. In Stage two, we input the same HDR image-self-calibrated to different maximum luminance levels-into the learned tone mapping network, and generate a pseudo-multi-exposure image stack with varying detail visibility and color saturation. We then train another fusion network to merge the LDR image stack into a desired LDR image by maximizing a variant of the structural similarity index for multi-exposure image fusion, proven perceptually relevant to fused image quality. Extensive experiments show that our method produces images with consistently better visual quality while ranking among the fastest local TMOs. Peibei Cao, Chenyang Le, Yuming Fang 0001, Kede Ma |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | A Bio-Inspired Model for Bee SimulationsabstractAs eusocial creatures, bees display unique macro collective behavior and local body dynamics that hold potential applications in various fields, such as computer animation, robotics, and social behavior. Unlike birds and fish, bees fly in a low-aligned zigzag pattern. Additionally, bees rely on visual signals for foraging and predator avoidance, exhibiting distinctive local body oscillations, such as body lifting, thrusting, and swaying. These inherent features pose significant challenges to realistic bee simulations in practical animation applications. In this article, we present a bio-inspired model for bee simulations capable of replicating both macro collective behavior and local body dynamics of bees. Our approach utilizes a visually-driven system to simulate a bee's local body dynamics, incorporating obstacle perception and body rolling control for effective collision avoidance. Moreover, we develop an oscillation rule that captures the dynamics of the bee's local bodies, drawing on insights from biological research. Our model extends beyond simulating individual bees' dynamics; it can also represent bee swarms by integrating a fluid-based field with the bees' innate noise and zigzag motions. To fine-tune our model, we utilize pre-collected honeybee flight data. Through extensive simulations and comparative experiments, we demonstrate that our model can efficiently generate realistic low-aligned and inherently noisy bee swarms. Wenxiu Guo, Yuming Fang 0001, Yang Tong, Tingsong Lu, Xiaogang Jin 0001, Zhigang Deng 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Comprehensive Visual Grounding for Video DescriptionabstractThe grounding accuracy of existing video captioners is still behind the expectation. The majority of existing methods perform grounded video captioning on sparse entity annotations, whereas the captioning accuracy often suffers from degenerated object appearances on the annotated area such as motion blur and video defocus. Moreover, these methods seldom consider the complex interactions among entities. In this paper, we propose a comprehensive visual grounding network to improve video captioning, by explicitly linking the entities and actions to the visual clues across the video frames. Specifically, the network consists of spatial-temporal entity grounding and action grounding. The proposed entity grounding encourages the attention mechanism to focus on informative spatial areas across video frames, albeit the entity is annotated in only one frame of a video. The action grounding dynamically associates the verbs to related subjects and the corresponding context, which keeps fine-grained spatial and temporal details for action prediction. Both entity grounding and action grounding are formulated as a unified task guided by a soft grounding supervision, which brings architecture simplification and improves training efficiency as well. We conduct extensive experiments on two challenging datasets, and demonstrate significant performance improvements of +2.3 CIDEr on ActivityNet-Entities and +2.2 CIDEr on MSR-VTT compared to state-of-the-arts. Wenhui Jiang 0001, Yibo Cheng, Linxin Liu, Yuming Fang 0001, Yuxin Peng 0001, Yang Liu 0293 |
AAAI | 4 |
| 2024 | Meta-Point Learning and Refining for Category-Agnostic Pose EstimationabstractCategory-agnostic pose estimation (CAPE) aims to predict keypoints for arbitrary classes given a few support images annotated with keypoints. Existing methods only rely on the features extracted at support keypoints to predict or refine the keypoints on query image, but a few support feature vectors are local and inadequate for CAPE. Considering that human can quickly perceive potential keypoints of arbitrary objects, we propose a novel framework for CAPE based on such potential keypoints (named as meta-points). Specifically, we maintain learnable embeddings to capture inherent information of various keypoints, which interact with image feature maps to produce meta-points without any support. The produced meta-points could serve as meaningful potential keypoints for CAPE. Due to the inevitable gap between inherency and annotation, we finally utilize the identities and details offered by support key-points to assign and refine meta-points to desired keypoints in query image. In addition, we propose a progressive deformable point decoder and a slacked regression loss for better prediction and supervision. Our novel framework not only reveals the inherency of key points but also outperforms existing methods of CAPE. Comprehensive experiments and in-depth studies on large-scale MP-100 dataset demon-strate the effectiveness of our framework. Code is avaiable at https://github.com/chenbys/MetaPoint Junjie Chen 0008, Jiebin Yan, Yuming Fang 0001, Li Niu 0002 |
CVPR | 3 |
| 2024 | Multiscale Sliced Wasserstein Distances as Perceptual Color Difference Measures
Zhihua Wang 0002, Leon Wang, Tsein-I Liu, Yuming Fang 0001, Qilin Sun 0001, Kede Ma |
ECCV (53) | 5 |
| 2024 | Arbitrary-Scale Video Super-Resolution with Structural and Textural Priors
Wei Shang 0001, Dongwei Ren, Yuming Fang 0001, Wangmeng Zuo, Kede Ma |
ECCV (57) | 4 |
| 2024 | Quality of Experience of Viewport Adaptive Omnidirectional Video StreamingabstractWith the explosive growth of multimedia streaming services and virtual reality devices, omnidirectional video (ODV) is becoming increasingly popular in practical applications. However, streaming the entire ODV with high definition and high frame rate induces a waste of bandwidth. The tile-based viewport adaptive streaming provides a solution to overcome volatile network conditions, while the scheme would lead to quality adaptation when the network changes dynamically. In this paper, we focus on investigating how the human visual quality of experience (QoE) changes with time-varying ODV quality. Specifically, we construct a new quality of experience database for viewport adaptive ODV streaming named JUFEOVQoE, which includes twelve original ODVs with diverse content, and corresponding 378 viewport videos generated by compressing the raw viewport videos using a variety of combinations of quantization parameter (QP), spatial (S), and temporal resolutions (T). We conduct a series of subjective experiments to collect the mean opinion scores of the viewport video sequences and the viewing direction data of the subjects. Furthermore, we test several state-of-the-art objective QoE models on the proposed database. Experimental results demonstrate that existing mainstream QoE methods cannot predict the QoE of the viewport adaptive streaming ODVs accurately. The database will be released to facilitate further research. Xuelin Liu, Haoyun Zhang, Jiebin Yan, Yuming Fang 0001, Shiqi Wang 0001 |
ICIP | 5 |
| 2024 | GSTran: Joint Geometric and Semantic Coherence for Point Cloud Segmentation
Abiao Li, Chenlei Lv, Guofeng Mei, Yifan Zuo 0001, Jian Zhang 0002, Yuming Fang 0001 |
ICPR (18) | 6 |
| 2024 | Blind Quality Assessment of Panoramic Images Based on Multiple Viewport SequencesabstractWith the development of virtual reality (VR) technology, panoramic image (PI), which is an important digital form of immersive multimedia, has drawn much attention from researchers. However, distortions are inevitably introduced in the process of processing, encoding and compression, which damages their quality and affects the user’s experience. Therefore, assessing the quality of panoramic images is urgent. In this paper, with the consideration of viewing behavior, we propose a novel blind panoramic image quality assessment model, which consists of three parts, viewport generation, feature extraction and quality prediction. Specifically, inspired by the viewing process of PI, we first generate multiple viewport sequences according to the real viewing trajectory and then extract multilevel features with a pre-trained backbone. The concatenated features are taken as the input of a recurrent neural network to evaluate the perceptual quality of PI. To validate the effectiveness of the proposed method, objective experiments are conducted on the public subjective panoramic image quality database. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods. Xuelin Liu, Jiebin Yan, Yuming Fang 0001, Hantao Liu |
ISCAS | 4 |
| 2024 | Adaptive Image Quality Assessment via Teaching Large Multimodal Model to CompareabstractWhile recent advancements in large multimodal models (LMMs) have significantly improved their abilities in image quality assessment (IQA) relying on absolute quality rating, how to transfer reliable relative quality comparison outputs to continuous perceptual quality scores remains largely unexplored. To address this gap, we introduce an all-around LMM-based NR-IQA model, which is capable of producing qualitatively comparative responses and effectively translating these discrete comparison outcomes into a continuous quality score. Specifically, during training, we present to generate scaled-up comparative instructions by comparing images from the same IQA dataset, allowing for more flexible integration of diverse IQA datasets. Utilizing the established large-scale training corpus, we develop a human-like visual quality comparator. During inference, moving beyond binary choices, we propose a soft comparison method that calculates the likelihood of the test image being preferred over multiple predefined anchor images. The quality score is further optimized by maximum a posteriori estimation with the resulting probability matrix. Extensive experiments on nine IQA datasets validate that the Compare2Score effectively bridges text-defined comparative levels during training with converted single image quality scores for inference, surpassing state-of-the-art IQA models across diverse scenarios. Moreover, we verify that the probability-matrix-based inference conversion not only improves the rating accuracy of Compare2Score but also zero-shot general-purpose LMMs, suggesting its intrinsic effectiveness. Hanwei Zhu, Haoning Wu 0001, Baoliang Chen, Lingyu Zhu 0006, Yuming Fang 0001, Guangtao Zhai, Weisi Lin, Shiqi Wang 0001 |
NeurIPS | 7 |
| 2024 | PosCap: Boosting Video Captioning with Part-of-Speech Guidance
Jingfu Xiao, Wenhui Jiang 0001, Yuming Fang 0001 |
PRCV (10) | 4 |
| 2024 | Harmonizing Base and Novel Classes: A Class-Contrastive Approach for Generalized Few-Shot Segmentation
Weide Liu, Yuming Fang 0001, Chuan-Sheng Foo, Jun Cheng 0003, Guosheng Lin |
Int. J. Comput. Vis. | 4 |
| 2024 | CD-iNet: Deep Invertible Network for Perceptual Image Color Difference Measurement
Zhihua Wang 0002, Keshuo Xu, Keyan Ding, Qiuping Jiang, Yifan Zuo 0001, Zhangkai Ni, Yuming Fang 0001 |
Int. J. Comput. Vis. | 7 |
| 2024 | CFNet: Conditional filter learning with dynamic noise estimation for real image denoising
Yifan Zuo 0001, Wenhao Yao, Yifeng Zeng, Yuming Fang 0001, Yan Huang 0023, Wenhui Jiang 0001 |
Knowl. Based Syst. | 5 |
| 2024 | Benchmarking deep models on retinal fundus disease diagnosis and a large-scale datasetabstractRetinal fundus imaging contributes to monitoring the vision of patients by providing views of the interior surface of the eyes. Machine learning models greatly aided ophthalmologists in detecting retinal disorders from color fundus images. Hence, the quality of the data is pivotal for enhancing diagnosis algorithms, which ultimately benefits vision care and maintenance. To facilitate further research in this domain, we introduce the Eye Disease Diagnosis and Fundus Synthesis (EDDFS) dataset, comprising 28,877 fundus images. These include 15,000 healthy samples and a diverse range of images depicting various disorders such as diabetic retinopathy, age-related macular degeneration, glaucoma, pathological myopia, hypertension retinopathy, retinal vein occlusion, and Laser photocoagulation. In addition to providing the dataset, we propose a Transformer-joint convolution network for automated eye disease screening. Firstly, a co-attention structure is integrated to capture long-range attention information along with local features. Secondly, a cross-stage feature fusion module is designed to extract multi-level and disease-related information. By leveraging the dataset and our proposed network, we establish benchmarks for disease screening and grading tasks. Our experimental results underscore the network’s proficiency in both multi-label and single-label disease diagnosis, while also showcasing the dataset’s capability in supporting fundus synthesis. (The dataset and code will be available on https://github.com/xia-xx-cv/EDDFS_dataset). Xue Xia 0005, Guobei Xiao, Kun Zhan, Jinhua Yan, Yuming Fang 0001, Guofu Huang |
Signal Process. Image Commun. | 7 |
| 2024 | Learning content-aware feature fusion for guided depth map super-resolution
Yifan Zuo 0001, Xiaoshui Huang, Xue Xia 0005, Yuming Fang 0001 |
Signal Process. Image Commun. | 7 |
| 2024 | TPE-DF: Thumbnail Preserving Encryption via Dual-2DCS FusionabstractThumbnail preserving encryption (TPE) images protect privacy while still maintaining visual usability when stored in the cloud. However, existing thumbnail preserving encryption techniques are insufficient in defending against statistical attack. Drawing on the ideas of 2D compressed sensing (2DCS), this letter proposes thumbnail preserving encryption via dual-2DCS fusion, called TPE-DF, which can resist statistical attack. The original image is processed by 2DCS with deterministic binary block diagonal (DBBD) sampling matrix to generate the sampled image. Then, the residual matrix (RM) is then constructed by subtracting the predictive reconstructed values of the sampled image from the original image. The sampled image is processed by dual-2DCS fusion with arbitrary scaling sensing degradation (ASSD) matrix to obtain a carrier image. Then, the bitstream of RM is embedded in the carrier image to finally generate the TPE image. Experimental results show that a reversible TPE scheme with resistance to statistical attack is achieved. Wenying Wen, Qiyu Jiang, Haigang Huang, Yushu Zhang 0001, Yuming Fang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Perceptual Quality Assessment of Virtual Reality Videos in the WildabstractInvestigating how people perceive virtual reality (VR) videos in the wild (i.e., those captured by everyday users) is a crucial and challenging task in VR-related applications due to complexauthenticdistortionslocalizedinspaceandtime.Existingpanoramic video databases only consider synthetic distortions, assume fixed viewing conditions, and are limited in size. To overcome these shortcomings, we construct the VR Video Quality in the Wild (VRVQW) database, containing 502 user-generated videos with diverse content and distortion characteristics. Based on VRVQW, we conduct a formal psychophysical experiment to record the scanpaths and perceived quality scores from 139 participants under two different viewing conditions. We provide a thorough statistical analysis of the recordeddata, observing significantimpact of viewing conditions on both human scanpaths and perceived quality. Moreover, we develop an objective quality assessment model for VR videos based on pseudocylindrical representation and convolution. Results on the proposed VRVQW show that our method is superior to existing video quality assessment models.We have made the database and code available at https://github.com/ limuhit/VR-Video-Quality-in-the-Wild. Wen Wen 0007, Mu Li 0005, Yiru Yao, Xiangjie Sui, Yabin Zhang 0002, Long Lan, Yuming Fang 0001, Kede Ma |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Video Quality Assessment for Online Processing: From Spatial to Temporal SamplingabstractWith the rapid development of multimedia processing and deep learning technologies, especially in the field of video understanding, video quality assessment (VQA) has achieved significant progress. Although researchers have moved from designing efficient video quality mapping models to various research directions, in-depth exploration of the effectiveness-efficiency trade-offs of spatio-temporal modeling in VQA models is still less sufficient. Considering the fact that videos have highly redundant information, this paper investigates this problem from the perspective of joint spatial and temporal sampling, aiming to seek the answer to how little information we should keep at least when feeding videos into the VQA models while with acceptable performance sacrifice. To this end, we drastically sample the video’s information from both spatial and temporal dimensions, and the heavily squeezed video is then fed into a stable VQA model. Comprehensive experiments regarding joint spatial and temporal sampling are conducted on six public video quality databases, and the results demonstrate the acceptable performance of the VQA model when throwing away most of the video information. Furthermore, with the proposed joint spatial and temporal sampling strategy, we make an initial attempt to design an online VQA model, which is instantiated by as simple as possible a spatial feature extractor, a temporal feature fusion module, and a global quality regression module. Through quantitative and qualitative experiments, we verify the feasibility of online VQA model by simplifying itself and reducing input. Jiebin Yan, Yuming Fang 0001, Xuelin Liu, Xue Xia 0005, Weide Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | 2AFC Prompting of Large Multimodal Models for Image Quality AssessmentabstractWhile abundant research has been conducted on improving high-level visual understanding and reasoning capabilities of large multimodal models (LMMs), their image quality assessment (IQA) ability has been relatively under-explored. Here we take initial steps towards this goal by employing the two-alternative forced choice (2AFC) prompting, as 2AFC is widely regarded as the most reliable way of collecting human opinions of visual quality. Subsequently, the global quality score of each image estimated by a particular LMM can be efficiently aggregated using the maximum a posteriori estimation. Meanwhile, we introduce three evaluation criteria: consistency, accuracy, and correlation, to provide comprehensive quantifications and deeper insights into the IQA capability of five LMMs. Extensive experiments show that existing LMMs exhibit remarkable IQA ability on coarse-grained quality comparison, but there is room for improvement on fine-grained quality discrimination. The proposed dataset sheds light on the future development of IQA models based on LMMs. The codes will be made publicly available athttps://github.com/h4nwei/2AFC-LMMs. Hanwei Zhu, Xiangjie Sui, Baoliang Chen, Xuelin Liu, Peilin Chen 0001, Yuming Fang 0001, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | A2 GSTran: Depth Map Super-Resolution via Asymmetric Attention With Guidance SelectionabstractCurrently, Convolutional Neural Network (CNN) has dominated guided depth map super-resolution (SR). However, the inefficient receptive field growing and input-independent convolution limit the generalization of CNN. Motivated by vision transformer, this paper proposes an efficient transformer-based backbone A2GSTran for guided depth map SR, which resolves the above intrinsic defect of CNN. In addition, state-of-the-art (SOTA) models only refine depth features with the guidance which is implicitly selected without supervision. So, there is no explicit guarantee to mitigate the artifacts of texture copying and edge blurring. Accordingly, the proposed A2GSTran simultaneously solves two sub-problems,i.e., guided monocular depth estimation and guided depth SR, in separate branches. Specifically, the explicit supervision upon monocular depth estimation lifts the efficiency of guidance selection. The feature fusion between branches is designed via bi-directional cross attention. Moreover, since guidance domain is defined in high resolution (HR), we propose asymmetric cross attention to maintain the guidance information via pixel unshuffle instead of pooling which has unequal channel number to depth features. Based on the supervisions to depth reconstruction and guidance selection, the final depth features are refined by fusing the output features of the corresponding branches via channel attention to generate the HR depth map. Sufficient experimental results on synthetic and real datasets for multiple scales validate our contributions compared with SOTA models. The code and models are public via https://github.com/alex-cate/Depth_Map_Super-resolution_via_Asymmetric_Attention_with_Guidance_Selection Yifan Zuo 0001, Yifeng Zeng, Yuming Fang 0001, Xiaoshui Huang, Jiebin Yan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | FedCov: Enhanced Trustworthy Federated Learning for Machine RUL Prediction With Continuous-to-Discrete ConversionabstractNumerous approaches have been proposed for predicting machine remaining useful life (RUL), which helps prevent unnecessary downtime and reduces the maintenance cost in industrial systems. Most existing RUL methods rely on centralized learning and require large-scale datasets with manual labels, which are infeasible to collect. As a decentralized learning paradigm, federated learning (FL) has recently been integrated into these approaches, which aims to utilize the distributed data from local users for model training, while preserving their data privacy. However, data heterogeneity in the industry poses a critical challenge for FL, leading to model drifting issue and degraded global model performance. A straightforward method to tackle this problem is to estimate the data distribution of clients. However, it is difficult to apply this method to a regression task, since the prediction space is continuous, which increases the difficulty in estimating the data distribution. Motivated by the digital-to-analog converter in electronics, we propose a novel approach called FedCov, which involves a converter module that transforms continuous RUL values into discrete categories. Subsequently, a generator is trained to aggregate user information based on the discrete label distribution, and it is broadcasted to users as a data enhancement tool to address the data heterogeneity problem. Furthermore, to improve the performance and reliability of our FedCov, an uncertainty estimation module is proposed, which utilizes the confidence level of model predictions to adjust the training direction. Extensive experiments are conducted on theC-MAPSSbenchmark, which demonstrates that our proposed FedCov effectively solves the model drift issues and improves the performances on the RUL task, achieving state-of-the-arts performances. Yuming Fang 0001, Weide Liu, Ruibing Jin, Jun Cheng 0003, Zhenghua Chen |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Saliency Guided Deep Neural Network for Color Transfer With Light OptimizationabstractColor transfer aims to change the color information of the target image according to the reference one. Many studies propose color transfer methods by analysis of color distribution and semantic relevance, which do not take the perceptual characteristics for visual quality into consideration. In this study, we propose a novel color transfer method based on the saliency information with brightness optimization. First, a saliency detection module is designed to separate the foreground regions from the background regions for images. Then a dual-branch module is introduced to implement color transfer for images. Finally, a brightness optimization operation is designed during the fusion of foreground and background regions for color transfer. Experimental results show that the proposed method can implement the color transfer for images while keeping the color consistency well. Compared with other existing studies, the proposed method can obtain significant performance improvement. The source code and pre-trained models are available at https://github.com/PlanktonQAQ/SCTNet. Yuming Fang 0001, Pengwei Yuan, Chenlei Lv, Jiebin Yan, Weisi Lin |
IEEE Trans. Image Process. | 1 |
| 2024 | Smoke-Aware Global-Interactive Non-Local Network for Smoke Semantic SegmentationabstractCompared with other objects, smoke semantic segmentation (SSS) is more difficult and challenging due to some special characteristics of smoke, such as non-rigid, translucency, variable mode and so on. To achieve accurate positioning of smoke in real complex scenes and promote the development of intelligent fire detection, we propose a Smoke-Aware Global-Interactive Non-local Network (SAGINN) for SSS, which harness the power of both convolution and transformer to capture local and global information simultaneously. Non-local is a powerful means for modeling long-range context dependencies, however, friendliness to single-scale low-resolution features limits its potential to produce high-quality representations. Therefore, we propose a Global-Interactive Non-local (GINL) module, leveraging global interaction between multi-scale key information to improve the robustness of feature representations. To solve the interference of smoke-like objects, a Pyramid High-level Semantic Aggregation (PHSA) module is designed, where the learned high-level category semantics from classification aids model by providing additional guidance to correct the wrong information in segmentation representations at the image level and alleviate the inter-class similarity problem. Besides, we further propose a novel loss function, termed Smoke-aware loss (SAL), by assigning different weights to different objects contingent on their importance. We evaluate our SAGINN on extensive synthetic and real data to verify its generalization ability. Experimental results show that SAGINN achieves 83% average mIoU on the three testing datasets (83.33%, 82.72% and 82.94%) of SYN70K with an accuracy improvement of about 0.5%, 0.002 mMse and 0.805Fβ on SMOKE5K, which can obtain more accurate location and finer boundaries of smoke, achieving satisfactory results on smoke-like objects. Lin Zhang 0061, Feiniu Yuan, Yuming Fang 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Image Super-Resolution via Efficient Transformer Embedding Frequency Decomposition With RestartabstractRecently, transformer-based backbones show superior performance over the convolutional counterparts in computer vision. Due to quadratic complexity with respect to the token number in global attention, local attention is always adopted in low-level image processing with linear complexity. However, the limited receptive field is harmful to the performance. In this paper, motivated by Octave convolution, we propose a transformer-based single image super-resolution (SISR) model, which explicitly embeds dynamic frequency decomposition into the standard local transformer. All the frequency components are continuously updated and re-assigned via intra-scale attention and inter-scale interaction, respectively. Specifically, the attention in low resolution is enough for low-frequency features, which not only increases the receptive field, but also decreases the complexity. Compared with the standard local transformer, the proposed FDRTran layer simultaneously decreases FLOPs and parameters. By contrast, Octave convolution only decreases FLOPs of the standard convolution, but keeps the parameter number unchanged. In addition, the restart mechanism is proposed for every a few frequency updates, which first fuses the low and high frequency, then decomposes the features again. In this way, the features can be decomposed in multiple viewpoints by learnable parameters, which avoids the risk of early saturation for frequency representation. Furthermore, based on the FDRTran layer with restart mechanism, the proposed FDRNet is the first transformer backbone for SISR which discusses the Octave design. Sufficient experiments show our model reaches state-of-the-art performance on 6 synthetic and real datasets. The code and the models are available at https://github.com/catnip1029/FDRNet. Yifan Zuo 0001, Wenhao Yao, Yuming Fang 0001, Wei Liu 0044, Yuxin Peng 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Bayesian Uncertainty Calibration for Federated Time Series AnalysisabstractDeep learning models for time series analysis often require large-scale labeled datasets for training. However, acquiring such datasets is cost-intensive and challenging, particularly for individual institutions. To overcome this challenge and concern about data confidentiality among different institutions, federated learning (FL) servers as a viable solution to this dilemma by offering a decentralized learning framework. However, the datasets collected by each institution often suffer from imbalance and may not adhere to uniform protocols, leading to diverse data distributions. To address this problem, we design a global model to approximate the global data distribution of all participant clients, then transfer it to local clients as an induction in the training phase. While discrepancies between the approximate distribution and the actual distribution result in uncertainty in the predicted results. Moreover, the diverse data distributions among various clients within the FL framework, combined with the inherent lack of reliability and interpretability in deep learning models, further amplify the uncertainty of the prediction results. To address these issues, we propose an uncertainty calibration method based on Bayesian deep learning techniques, which captures uncertainty by learning a fidelity transformation to reconstruct the output of time series regression and classification tasks, utilizing deterministic pre-trained models. Extensive experiments on the regression dataset (C-MAPSS) and classification datasets (ESR, Sleep-EDF, HAR, and FD) in the Independent and Identically Distributed (IID) and non-IID settings show that our approach effectively calibrates uncertainty within the FL framework and facilitates better generalization performance in both the regression and classification tasks, achieving state-of-the-art performance. Weide Liu, Xue Xia 0005, Zhenghua Chen, Yuming Fang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Adaptive Structure and Texture Similarity Metric for Image Quality Assessment and OptimizationabstractObjective Image Quality Assessment (IQA) aims to design computational models that can automatically predict the perceived quality of images. The state-of-the-art full-reference IQA metric – Deep Image Structure and Texture Similarity (DISTS), neglects the fact that natural images often consist of local structure and texture, and requires supervised training on the annotated dataset. In this article, we introduce multiple adaptive strategies to improve DISTS, resulting in an opinion-unaware IQA metric, named A-DISTS. Specifically, A-DISTS first uses the dispersion index as a statistical feature to adaptively localize structure and texture regions at different scales. Second, it adaptively assigns the spatial weights between local structure and texture similarity measurements according to the estimated structure or texture probability maps. Finally, it calculates the entropy of image representation to adaptively weigh the importance of each feature map. As a result, A-DISTS is adapted to local image content and does not require any training. The experimental results demonstrated that the proposed metric correlates well with human rating in the standard and algorithm-dependent IQA databases, and exhibits competitive performance in the optimization tasks of single image super-resolution, motion deblurring, and multi-distortion removal. Keyan Ding, Rijin Zhong, Zhihua Wang 0002, Yang Yu 0014, Yuming Fang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | PPM-SEM: A Privacy-Preserving Mechanism for Sharing Electronic Patient Records and Medical Images in TelemedicineabstractDespite the various privacy protection methods that are available through medical services platforms, it is still challenging for patients to achieve a desirable level of privacy protection during image sharing. Therefore, this paper proposes a privacy protection mechanism, called PPM-SEM, for the secure sharing of electronic patient records (EPRs) and medical images in telemedicine; it includes two stages: privacy preparation and privacy protection and reconstruction. In the first stage, a dual watermark (i.e., an image watermark) is generated by combining the patient's EPRs with an image, which can be utilized to ensure the security of patient identity data (i.e., a text watermark).Inthe second stage, a modal transformation network is constructed by training the dual watermark together as an additional channel. This network is called watermark-CycleGAN (W-CycleGAN), which can address the privacy and security issues concerning medical images and provide a double protection mechanism for EPRs. Experimental results demonstrate that only the recovery network with the correct key can restore high-quality medical images. In addition, the patient's EPRs can be fully extracted; i.e., 100% accuracy can be maintained. It is noted that the nonpaired recovery network can also recover visually meaningful medical images, thereby realizing privacy protection for patients in telemedicine scenarios. Wenying Wen, Ziye Yuan, Yushu Zhang 0001, Yuming Fang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Perceptual Quality Assessment of Omnidirectional Images: A Benchmark and Computational ModelabstractCompared with traditional 2D images, omnidirectional images (also referred to as 360 ∘ images) have more complicated perceptual characteristics due to the particularities of imaging and display. How humans perceive omnidirectional images in an immersive environment and form the immersive quality of experience are important problems. Thus, it is crucial to measure the quality of omnidirectional images under different viewing conditions, which suffer from realistic distortions. In this article, we build a large-scale subjective assessment database for omnidirectional images and carry out a comprehensive psychophysical experiment to study the relationships between different factors (viewing conditions and viewing behaviors) and the perceptual quality of omnidirectional images. In addition, we collect both subjective ratings and head movement data. A thorough analysis of the collected subjective data is also provided, where we make several interesting findings. Moreover, with the proposed database, we propose a novel transformer-based omnidirectional image quality assessment model. To be consistent with the human viewing process, viewing conditions and behaviors are naturally incorporated into the proposed model. Specifically, the proposed model mainly consists of three parts: viewport sequence generation, multi-scale feature extraction, and perceptual quality prediction. Extensive experimental results conducted on the proposed database demonstrate the effectiveness of the proposed method over existing image quality assessment methods. Xuelin Liu, Jiebin Yan, Yuming Fang 0001, Yang Liu 0293 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Improving Image Aesthetic Assessment via Multiple Image Joint LearningabstractImage Aesthetic Assessment (IAA) is an emerging paradigm that predicts aesthetic score as the popular aesthetic taste for an image. Previous IAA approaches take a single image as input to predict the aesthetic score of the image. However, we discover that most existing IAA methods fail dramatically to predict the images with a large variance of aesthetic voting distribution. Motivated by the practice that people consider similar experiences to improve the consistence of the voting result, we present a novel Multiple Image joint Learning Network (MILNet) to mimic this natural process. Our novelty is mainly three-fold: (a) Semantic-based retrieval method that constructs aesthetic similarity (the similarity of aesthetic attribution) to select reference images; (b) Graph network reasoning that initializes and updates the weight of intrinsic relationships among multiple images; (c) Adaptive Earth Mover’s Distance (AdaEMD) loss function that adjusts weight for easy and hard instances to mitigate unbalanced distribution of aesthetic datasets. Our evaluation with the benchmark AVA and TAD datasets demonstrates that the proposed MILNet outperforms state-of-the-art IAA methods. The code is available at https://github.com/flyingbird93/MILNet . Tengfei Shi, Chenglizhao Chen, Aimin Hao, Yuming Fang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Visual Security Index Combining CNN and Filter for Perceptually Encrypted Light Field ImagesabstractVisual security index (VSI) represents a quantitative index for the visual security evaluation of perceptually encrypted images. Recently, the research on visual security of encrypted light field (LF) images faces two challenges. One is that the existing perceptually encrypted image databases are often too small, which is easy to cause overfitting in convolutional neural network (CNN). The other is that existing VSI models did not take a full account the intrinsic characteristics of the LF images and highly relied on handcrafted feature extraction. In this article, we construct a new database of perceptually encrypted LF images, called the PE-SLF, which is 2.6 times as big as the existing largest perceptual encrypted image database. Moreover, a novel visual security index (VSI) model is proposed by taking into full consideration the intrinsic spatial-angular characteristics of the LF images and the outstanding capabilities of CNN in feature extraction. First, we exploit CNN to detect the texture and structure features of encrypted sub-aperture images in the spatial domain. Second, we apply the Gabor filter to detect the Gabor feature over the epi-polar plane images in angular domain. Last, the spatial and angular similarity measurements are subsequently calculated for jointly yielding the final visual security score. Experimental results on the constructed PE-SLF demonstrate that the proposed VSI model is closer to the perception of HVS in visual security evaluation of encrypted LF images compared to other classical and state-of-the-art models. Wenying Wen, Minghui Huang, Yushu Zhang 0001, Yuming Fang 0001, Yifan Zuo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | ScanDMM: A Deep Markov Model of Scanpath Prediction for 360° ImagesabstractScanpath prediction for 360° images aims to produce dynamic gaze behaviors based on the human visual perception mechanism. Most existing scanpath prediction methods for 360° images do not give a complete treatment of the time-dependency when predicting human scanpath, resulting in inferior performance and poor generalizability. In this paper, we present a scanpath prediction method for 360° images by designing a novel Deep Markov Model (DMM) architecture, namely ScanDMM. We propose a semantics-guided transition function to learn the nonlinear dynamics of time-dependent attentional landscape. Moreover, a state initialization strategy is proposed by considering the starting point of viewing, enabling the model to learn the dynamics with the correct “launcher”. We further demonstrate that our model achieves state-of-the-art performance on four 360° image databases, and exhibit its generalizability by presenting two applications of applying scanpath prediction models to other visual tasks - saliency detection and image quality assessment, expecting to provide profound insights into these fields. Xiangjie Sui, Yuming Fang 0001, Hanwei Zhu, Shiqi Wang 0001, Zhou Wang 0001 |
CVPR | 2 |
| 2023 | Laptran: Transformer Embedding Graph Laplacian for Point Cloud Part SegmentationabstractSince the feature representations of the points located at the junction regions of various parts are ambiguous, it is still challenging to exploit the fine-grained semantic features of point clouds on part segmentation tasks. To resolve the issue, we design a modified transformer module, named Laplacian transformer, to investigate the local differences between each point and its corresponding neighbors based on graph Laplacian theory. This module constructs a more accurate local geometric representation of the point cloud. It concentrates on the points located at the junction areas of various parts while boosting the recognition effect of these points. Encapsulated with the Laplacian module, we propose a Unet-like transformer framework to perform part segmentation for point clouds. Experimental results demonstrate that the proposed framework achieves more accurate results on public benchmark datasets. Abiao Li, Chenlei Lv, Yuming Fang 0001, Yifan Zuo 0001 |
ICIP | 3 |
| 2023 | Learning a Multilevel Cooperative View Reconstruction Network for Light Field Angular Super-ResolutionabstractRecently, many methods have been proposed to improve the angular resolution of sparsely-sampled Light Field (LF). However, the synthesized dense LF inevitably exhibits blurry edges and artifacts. This paper intents to model the global relations of LF views and quality degradation model by learning a multilevel cooperative view reconstruction network to further enhance LF angular Super-Resolution (SR) performance. The proposed LF angular SR network consists of three sub-networks including the Cooperative Angular Transformer Network (CATNet), the Deblurring Network (DBNet), and the Texture Repair Network (TRNet). The CATNet simultaneously captures global features of all LF views and local features within each view, which benefits in characterizing the inherent LF structure. The DBNet models a quality degradation model by estimating blur kernels to reduce the blurry edges and artifacts. The TRNet focuses on restoring fine-scale texture details. Experimental results over various LF datasets including large baseline LF images demonstrate the significant superiority of our method when compared with state-of-the-art ones. Deyang Liu, Xiaofei Zhou 0003, Ping An 0001, Yuming Fang 0001 |
ICME | 5 |
| 2023 | Feature Adaptive YOLO for Remote Sensing Detection in Adverse Weather ConditionsabstractTarget detection in remote sensing has been one of the most challenging tasks in the past few decades. However, the detection performance in adverse weather conditions still needs to be satisfactory, mainly caused by the low-quality image features and the fuzzy boundary information. This work proposes a novel framework called Feature Adaptive YOLO (FA-YOLO). Specifically, we present a Hierarchical Feature Enhancement Module (HFEM), which adaptively performs feature-level enhancement to tackle the adverse impacts of different weather conditions. Then, we propose an Adaptive receptive Field enhancement Module (AFM) that dynamically adjusts the receptive field of the features and thus can enrich the context information for feature augmentation. In addition, we introduce Deformable Gated Head (DG-Head) which reduces the clutter caused by adverse weather. Experimental results on RTTS and two synthetic datasets demonstrate that our proposed FA-YOLO significantly outperforms other state-of-the-art target detection models. Chaojun Ni, Wenhui Jiang 0001, Qishou Zhu, Yuming Fang 0001 |
VCIP | 5 |
| 2023 | Geometry-assisted multi-representation view reconstruction network for Light Field image angular super-resolution
Deyang Liu, Zaidong Tong, Yan Huang 0023, Yifan Zuo 0001, Yuming Fang 0001 |
Knowl. Based Syst. | 6 |
| 2023 | Authenticable medical image-sharing scheme based on embedded small shadow QR code and blockchain framework
Wenying Wen, Yunpeng Jian, Yuming Fang 0001, Yushu Zhang 0001, Baolin Qiu |
Multim. Syst. | 3 |
| 2023 | A multi-level approach with visual information for encrypted H.265/HEVC videos
Wenying Wen, Rongxin Tu, Yushu Zhang 0001, Yuming Fang 0001, Yong Yang 0001 |
Multim. Syst. | 4 |
| 2023 | Measuring Perceptual Color Differences of Smartphone PhotographsabstractMeasuring perceptual color differences (CDs) is of great importance in modern smartphone photography. Despite the long history, most CD measures have been constrained by psychophysical data of homogeneous color patches or a limited number of simplistic natural photographic images. It is thus questionable whether existing CD measures generalize in the age of smartphone photography characterized by greater content complexities and learning-based image signal processors. In this article, we put together so far the largest image dataset for perceptual CD assessment, in which the photographic images are 1) captured by six flagship smartphones, 2) altered by Photoshop, 3) post-processed by built-in filters of the smartphones, and 4) reproduced with incorrect color profiles. We then conduct a large-scale psychophysical experiment to gather perceptual CDs of 30,000 image pairs in a carefully controlled laboratory environment. Based on the newly established dataset, we make one of the first attempts to construct an end-to-end learnable CD formula based on a lightweight neural network, as a generalization of several previous metrics. Extensive experiments demonstrate that the optimized formula outperforms 33 existing CD measures by a large margin, offers reasonable local CD maps without the use of dense supervision, generalizes well to homogeneous color patch data, and empirically behaves as a proper metric in the mathematical sense. Our dataset and code are publicly available at https://github.com/hellooks/CDNet. Zhihua Wang 0002, Keshuo Xu, Yang Yang 0201, Jianlei Dong, Shuhang Gu, Lihao Xu, Yuming Fang 0001, Kede Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | UDNet: Uncertainty-aware deep network for salient object detection
Yuming Fang 0001, Jiebin Yan, Wenhui Jiang 0001, Yang Liu 0293 |
Pattern Recognit. | 1 |
| 2023 | Toward a blind image quality evaluator in the wild by learning beyond human opinion scoresabstractNowadays, most existing blind image quality assessment (BIQA) models i n t h e w i l d heavily rely on human ratings, which are extraordinarily labor-expensive to collect. Here, we propose an o p i n i o n − f r e e BIQA method that learns from multiple annotators to assess the perceptual quality of images captured in the wild. Specifically, we first synthesize distorted images based on the pristine counterparts. We then randomly assemble a set of image pairs from the synthetic images, and use a group of IQA models to assign pseudo-binary labels for each pair indicating which image has higher quality as the supervisory signal. Based on the newly established pseudo-labeled dataset, we train a deep neural network (DNN)-based BIQA model to rank the perceptual quality, optimized for consistency with the binary rank labels. Since there exists domain shift, e.g., distortion shift and content shift, between the synthetic and in-the-wild images, we leverage two ways to alleviate this issue. First, the simulated distortions should be similar to authentic distortions as much as possible. Second, an unsupervised domain adaptation (UDA) module is further applied to encourage learning domain-invariant features between two domains. Extensive experiments demonstrate the effectiveness of our proposed o p i n i o n − f r e e BIQA model, yielding SOTA performance in terms of correlation with human opinion scores, as well as gMAD competition. Our code is available at: https://github.com/wangzhihua520/OF_BIQA . Zhihua Wang 0002, Jianguo Zhang 0001, Yuming Fang 0001 |
Pattern Recognit. | 4 |
| 2023 | Corrigendum to 'Toward a Blind Image Quality Evaluator in the Wild by Learning beyond Human Opinion Scores' Pattern Recognition. Volume 137 (2023) 109296
Zhihua Wang 0002, Jianguo Zhang 0001, Yuming Fang 0001 |
Pattern Recognit. | 4 |
| 2023 | Optical flow-assisted multi-level fusion network for Light Field image angular reconstruction
Deyang Liu, Yan Huang 0023, Yuanzhi Wang, Yuming Fang 0001 |
Signal Process. Image Commun. | 6 |
| 2023 | APCAS: Autonomous Privacy Control and Authentication Sharing in Social NetworksabstractRecently, the increasing development of social networks has brought about many privacy-related issues. Because users (participants) have different privacy sharing behaviors when multiple parties participate. The way of privacy sharing depends on the wishes of publishers, and users usually cannot control privacy sharing independently. To address this issue, this article proposes an autonomous privacy control and authentication sharing (APCAS) scheme based on quick response (QR) codes in social networks. Using the error correction of QR codes with high-quality images can prevent photographs that are uploaded to the social network from being subjected to lossy operations, which meets the participants’ needs of availability and privacy simultaneously. Combining the superiority of the polynomial-based secret image sharing and the visual secret image sharing can provide single-share certification for both dealer participation and dealer nonparticipation during certification. By comparing advanced image sharing algorithms, the proposed APCAS can restore the secret images one by one according to users’ wishes and can thus effectively mitigate the problem of autonomous privacy control in online sharing. Compared with related methods in recent years in terms of peak signal-to-noise ratio (PSNR), authentication operation, and authentication complexity, the proposed APCAS algorithm provides lossless decryption, lower authentication computation complexity, and no pixel scalability. Wenying Wen, Jiacong Fan, Yushu Zhang 0001, Yuming Fang 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | Gradient-Guided Single Image Super-Resolution Based on Joint Trilateral Feature FilteringabstractThe state of the arts (SOTAs) of single image super-resolution always exploit guidance from gradient prior. The fusion of gradient guidance is implemented by channel-wise concatenation followed by a convolutional layer. However, the kernels sharing in spatial positions cannot adaptively tune the effect of gradient guidance for all feature positions. To resolve this problem, a novel network module is proposed to simulate the traditional Joint Trilateral Filter (JTF) by extending the definition domain from pixels to features. Moreover, to improve the efficiency and flexibility, the functions of JTF kernel generation for image features and gradient features are explicitly learned instead of individual kernel weights, e.g., the exponential functions in the traditional JTF. Based on the proposed JTF modules, this paper follows the gradient-guided framework which simultaneously infers high-resolution (HR) image features and HR gradient features within two parallel branches, respectively. Specifically, by treating image features and gradient features as cross guidance to each other, the proposed JTF modules adaptively adjust the fusion patterns for local features via a bi-directional way. By doing so, the quality of image features and gradient features is alternatively enhanced. Compared with SOTAs, the proposed JTF-SISR shows improvement which is evaluated for multiple upsampling scales and degradation modes on 5 synthetic datasets, i.e., Set5, Set14, B100, Urban100 and Manga109, and 1 real dataset, i.e., RealSRSet. The code is public inhttps://github.com/a239xjc/JTF-SISR. Yifan Zuo 0001, Yuming Fang 0001, Deyang Liu, Wenying Wen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Study of Spatio-Temporal Modeling in Video Quality AssessmentabstractVideo quality assessment (VQA) has received remarkable attention recently. Most of the popular VQA models employ recurrent neural networks (RNNs) to capture the temporal quality variation of videos. However, each long-term video sequence is commonly labeled with a single quality score, with which RNNs might not be able to learn long-term quality variation well: What's the real role of RNNs in learning the visual quality of videos? Does it learn spatio-temporal representation as expected or just aggregating spatial features redundantly? In this study, we conduct a comprehensive study by training a family of VQA models with carefully designed frame sampling strategies and spatio-temporal fusion methods. Our extensive experiments on four publicly available in- the-wild video quality datasets lead to two main findings. First, the plausible spatio-temporal modeling module (i. e., RNNs) does not facilitate quality-aware spatio-temporal feature learning. Second, sparsely sampled video frames are capable of obtaining the competitive performance against using all video frames as the input. In other words, spatial features play a vital role in capturing video quality variation for VQA. To our best knowledge, this is the first work to explore the issue of spatio-temporal modeling in VQA. Yuming Fang 0001, Zhaoqian Li, Jiebin Yan, Xiangjie Sui, Hantao Liu |
IEEE Trans. Image Process. | 1 |
| 2023 | Multi-Stream Dense View Reconstruction Network for Light Field Image CompressionabstractRecently, many view synthesis-based methods are proposed for high-efficiency light field (LF) image compression. However, most existing methods fail to recover more texture details on occlusion regions, which reduces the compression efficiency. In this paper, we propose a multi-stream dense view reconstruction network to further improve LF image compression performance. In our method, only sparsely-sampled LF views are transmitted and the rest of the views are reconstructed at the decoder side. During the reconstruction process, we firstly constitute a multi-disparity geometry (MDG) structure based on the decoded sparse LF views, which can reflect abundant disparity characteristics. Subsequently, a multi-stream view reconstruction network (MSVRNet) is put forward to reconstruct a high-quality dense LF image, which consists of a multi-scale feature fusion sub-network, a fusion reconstruction sub-network, and a detail refinement sub-network. The multi-scale feature fusion sub-network can implicitly lean abundant multiscale geometric structure features from the constituted MDG structure. The fusion reconstruction sub-network and the detail refinement sub-network are respectively utilized to fuse the learned multiscale geometric features and restore more texture details, especially for occlusion regions. Moreover, 3D convolutional operations are adopted in the whole reconstruction process, which allow information propagation among the learned multiscale geometric features. Comprehensive experimental results demonstrate the effectiveness of the proposed method. The perceptual quality of reconstructed views and application on depth estimation also demonstrate that the proposed method can keep structural consistency of the reconstructed LF image and recover more texture details. Deyang Liu, Yan Huang 0023, Yuming Fang 0001, Yifan Zuo 0001, Ping An 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Perceptual Quality Assessment of Omnidirectional ImagesabstractOmnidirectional images, also called 360◦images, have attracted extensive attention in recent years, due to the rapid development of virtual reality (VR) technologies. During omnidirectional image processing including capture, transmission, consumption, and so on, measuring the perceptual quality of omnidirectional images is highly desired, since it plays a great role in guaranteeing the immersive quality of experience (IQoE). In this paper, we conduct a comprehensive study on the perceptual quality of omnidirectional images from both subjective and objective perspectives. Specifically, we construct the largest so far subjective omnidirectional image quality database, where we consider several key influential elements, i.e., realistic non-uniform distortion, viewing condition, and viewing behavior, from the user view. In addition to subjective quality scores, we also record head and eye movement data. Besides, we make the first attempt by using the proposed database to train a convolutional neural network (CNN) for blind omnidirectional image quality assessment. To be consistent with the human viewing behavior in the VR device, we extract viewports from each omnidirectional image and incorporate the user viewing conditions naturally in the proposed model. The proposed model is composed of two parts, including a multi-scale CNN-based feature extraction module and a perceptual quality prediction module. The feature extraction module is used to incorporate the multi-scale features, and the perceptual quality prediction module is designed to regress them to perceived quality scores. The experimental results on our database verify that the proposed model achieves the competing performance compared with the state-of-the-art methods. Yuming Fang 0001, Jiebin Yan, Xuelin Liu, Yang Liu 0293 |
AAAI | 1 |
| 2022 | Informative Attention Supervision for Grounded Video DescriptionabstractAttention supervision encourages grounded video description models (GVDMs) to focus on the related visual content when generating words. Thus, it improves the description performance of GVDMs. However, existing GVDMs often fail to focus on small but informative regions because these regions are considered as negative by using the intersection-over-union (IoU) based attention groundtruth sampling method. Moreover, the prevailing attention loss functions enforce the GVDMs to focus equally on all sampled regions when the GVDMs generate words, which may make it difficult for the model to attend to informative regions and thus degrade the quality of the generated sentences. To alleviate the above problems, we propose an informative attention supervision method including a novel attention groundtruth sampling method and a group-based weak grounding supervision. Specifically, our attention groundtruth sample method captures small proposal regions that overlap with the entity boxes. The proposed grounding supervision allows the GVDMs to dynamically focus on some of the most informative attention regions instead of all of them. Our approach yields competitive results on the ActivityNet Entities dataset without bells and whistles, surpassing previous methods without increasing inference costs. Boyang Wan, Wenhui Jiang 0001, Yuming Fang 0001 |
ICASSP | 3 |
| 2022 | A Database of Visual Color Differences of Modern Smartphone PhotographyabstractMeasures for visual color differences (CDs) are pivotal in hardware and software upgrading of modern smartphone photography. Towards this goal, we construct currently the largest database for visual CDs of smartphone photography. Our database consists of 15, 335 natural images 1) captured by six latest flagship smartphones, 2) altered by Photoshop®, 3) post-processed by built-in filters of smartphones, and 4) reproduced with incorrect color profiles. Moreover, we conduct a large-scale psychophysical experiment to gather visual CDs of 30, 000 image pairs from 20 human subjects in a well-designed laboratory environment. Last, we apply our human-rated database to compare a total of 27 classical and recent CD metrics. We show that existing metrics are limited in assessing CDs of smartphone photography, and point out promising future directions of learning-based CD metrics. Keshuo Xu, Zhihua Wang 0002, Yang Yang 0201, Jianlei Dong, Lihao Xu, Yuming Fang 0001, Kede Ma |
ICIP | 6 |
| 2022 | A Practical Method for Butterfly Motion CaptureabstractSimulating realistic butterfly motion has been a widely-known challenging problem in computer animation. Arguably, one of its main reasons is the difficulty of acquiring accurate flight motion of butterflies. In this paper we propose a practical yet effective, optical marker-based approach to capture and process the detailed motion of a flying butterfly. Specifically, we first capture the trajectories of the wings and thorax of a flying butterfly using optical marker-based motion tracking. After that, our method automatically fills the positions of missing markers by exploiting the continuity and relevance of neighboring frames, and improves the quality of the captured motion via noise filtering with optimized parameter settings. Through comparisons with existing motion processing methods, we demonstrate the effectiveness of our approach to obtain accurate flight motions of butterflies. Furthermore, we created and will release a first-of-its-kind butterfly motion capture dataset to research community. Tingsong Lu, Yang Tong, Yuming Fang 0001, Zhigang Deng 0001 |
MIG | 4 |
| 2022 | Benchmarking 360° Saliency Models by General-Purpose MetricsabstractHow to effectively evaluate a model's capability to predict the visual attention of observers in 360° scenes gains interest along with the advancement of saliency prediction modeling of omnidirectional images (ODIs). So far, many general-purpose metrics from 2D saliency literature have been adopted to evaluate the 360° saliency models. However, whether they are still effective when being adopted to evaluate the 360° saliency models has not been explored. In this paper, we testify several standard saliency evaluation metrics on the 360° saliency models and comprehensively analyze their behaviors in the omnidirectional scenario. We find that 1) most metrics under-penalize false positives; 2) existing 360° datasets involve severe equator bias that few metrics can effectively penalize. We hope this case study can provide a guideline for benchmarking 360° image/video saliency models. Xiangjie Sui, Jiebin Yan, Yuming Fang 0001 |
MMSP | 4 |
| 2022 | Bilinear CNNs for Blind Quality Assessment of Fine-Grained ImagesabstractMost of the existing image quality assessment (IQA) studies focus on discriminable images, whose relative visual quality could be easily determined by human beings (we also call this issue coarse-grained (CG)-IQA). The effective models designed for CG-IQA struggle for quality assessment of the images with subtle differences (often exist in many real applications), which is also called fine-grained (FG) IQA problem. Thus, we make the first, to the best of our knowledge, attempt to build a novel blind IQA (BIQA) model for the images with FG distortion, aiming to fill the gap between objective IQA model and real applications. Specifically, the proposed model mainly consists of a feature extraction module (a sequence of convolution layers), a squeeze-and-excitation module, and a bilinear pooling module, whose objectives are extracting quality-aware features, enhancing features' representation ability, and discriminability. We conduct extensive experiments on a public FG-IQA database, and demonstrate the superiority of the proposed method and the effectiveness of each module. Jiebin Yan, Yuming Fang 0001, Wenhui Jiang 0001 |
MMSP | 3 |
| 2022 | Eye Disease Diagnosis and Fundus Synthesis: A Large-Scale Dataset and BenchmarkabstractAs one of the most common imaging modalities, retinal fundus imaging offers images of interior surface of eyes for initial examination of disorders. Data-driven machine learning methods, especially deep learning models in recent years, provide automatic ophthalmological disease diagnosis techniques from color fundus images. Data with high quality, diversity and balanced distribution supports deep model-based eye disease diagnosis. However, many existing datasets focus on a specific kind of eye disease, and some suffer from label noise or quality degeneration, which hinders automatic screening algorithms from dealing with multiple eye diseases. To solve this, we propose a high-quality dataset containing 28877 color fundus images for deep learning-based diagnosis. Except for 15000 healthy samples, the dataset consists of 8 eye disorders including diabetic retinopathy, agerelated macular degeneration, glaucoma, pathological myopia, hypertension, retinal vein occlusion, LASIK spot and others. Based on this, we propose a co-attention network for disease diagnosis, establish benchmark on screening and grading tasks, and demonstrate that the proposed dataset supports generative adversarial network-based image synthesis. The dataset will be made publicly available. Xue Xia 0005, Kun Zhan, Guobei Xiao, Jinhua Yan, Zhuxiang Huang, Guofu Huang, Yuming Fang 0001 |
MMSP | 8 |
| 2022 | FundusGAN: A One-Stage Single Input GAN for Fundus Synthesis
Xue Xia 0005, Yuming Fang 0001 |
PRCV (2) | 3 |
| 2022 | Dual-stream Self-attention Network for Image CaptioningabstractSelf-attention based encoder-decoder models achieve dominant performance in image captioning. However, most existing image captioning models (ICMs) only focus on modeling the relation between spatial tokens, while channel-wise attention is neglected for getting visual representation. Considering that different channels of visual representation usually denote different visual objects, it may lead to poor performance in terms of object and attribute words in the captioning sentences generated by the ICMs. In this paper, we propose a novel dual-stream self-attention module (DSM) to alleviate the above issue. Specifically, we propose a parallel self-attention based module that simultaneously encodes visual information from the spatial and channel dimensions. Besides, to obtain channel-wise visual features effectively and efficiently, we introduce a group self-attention block with linear computational complexity. To validate the effectiveness of our model, we conduct extensive experiments on the standard IC benchmarks including MSCOCO and Flickr30k. Without bells and whistles, the proposed model performs new SOTAs containing 135.4 CIDEr score on MSCOCO and 70.8 CIDEr score on Flickr30k. Boyang Wan, Wenhui Jiang 0001, Yuming Fang 0001, Wenying Wen, Hantao Liu |
VCIP | 3 |
| 2022 | Texture-aware Network for Smoke Density EstimationabstractSmoke density estimation, also termed as soft segmentation, was developed from pixel-wise smoke (hard) segmen-tation and it aims at providing transparency and segmentation confidence for each pixel. The key difference between them lies in that segmentation focuses on classifying pixels into smoke and non-smoke ones, while density estimation obtains inner transparency of smoke component rather than treat all smoke pixels as an equal value. Based on this, we propose a texture-aware network being able to capture inner transparency of smoke components rather than merely focus on general smoke distribution for pixel-wise smoke density estimation. Besides, we adapt the Squeeze-and-Excitation (SE) layer for smoke feature extraction by involving max values for robustness. In order to represent inhomogeneous smoke pixels, we proposed a simple yet efficient attention-based texture-aware module that involves both gradient and semantic information. Experimental results show that our method outperforms others in both single image density estimation or segmentation and video smoke detection. Xue Xia 0005, Kun Zhan, Yajing Peng, Yuming Fang 0001 |
VCIP | 4 |
| 2022 | Revisiting image captioning via maximum discrepancy competition
Boyang Wan, Wenhui Jiang 0001, Yuming Fang 0001, Minwei Zhu, Yang Liu 0293 |
Pattern Recognit. | 3 |
| 2022 | A Novel Video Salient Object Detection Method via Semisupervised Motion Quality PerceptionabstractPrevious video salient object detection (VSOD) approaches have mainly focused on the perspective of network design for achieving performance improvements. However, with the recent slowdown in the development of deep learning techniques, it might become increasingly difficult to anticipate another breakthrough solely via complex networks. Therefore, this paper proposes a universal learning scheme to obtain a further 3% performance improvement for all state-of-the-art (SOTA) VSOD models. The major highlight of our method is that we propose the ‘motion quality’, a new concept for mining video frames from the ‘buffered’ testing video stream for constructing a fine-tuning set. By using our approach, all frames in this set can all well-detect their salient object by the ‘target SOTA model’ — the one we want to improve. Thus, the VSOD results of the mined set, which were previously derived by the target SOTA model, can be directly applied as pseudolearning objectives to fine-tune a completely new spatial model that has been pretrained on the widely used DAVIS-TR set. Since some spatial scenes in the buffered testing video stream are shown, the fine-tuned spatial model can perform very well for the remaining unseen testing frames, outperforming the target SOTA model significantly. Although offline model fine tuning requires additional time costs, the performance gain can still benefit scenarios without speed requirements. Moreover, its semisupervised methodology might have considerable potential to inspire the VSOD community in the future. Chenglizhao Chen, Chong Peng 0001, Guodong Wang 0001, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | A Novel Long-Term Iterative Mining Scheme for Video Salient Object DetectionabstractThe existing state-of-the-art (SOTA) video salient object detection (VSOD) models have widely followed short-term methodology, which dynamically determines the balance between spatial and temporal saliency fusion by solely considering the current consecutive limited frames. However, the short-term methodology has one critical limitation, which conflicts with the real mechanism of our visual system — a typical long-term methodology. As a result, failure cases keep showing up in the results of the current SOTA models, and the short-term methodology becomes the major technical bottleneck. To solve this problem, this paper proposes a novel VSOD approach, which performs VSOD in a complete long-term way. Our approach converts the sequential VSOD, a sequential task, to a data mining problem, i.e., decomposing the input video sequence to object proposals in advance and then mining salient object proposals as much as possible in an easy-to-hard way. Since all object proposals are simultaneously available, the proposed approach is a complete long-term approach, which can alleviate some difficulties rooted in conventional short-term approaches. In addition, we devised an online updating scheme that can grasp the most representative and trustworthy pattern profile of the salient objects, outputting framewise saliency maps with rich details and smoothing both spatially and temporally. The proposed approach outperforms almost all SOTA models on five widely used benchmark datasets. Chenglizhao Chen, Hengsen Wang, Yuming Fang 0001, Chong Peng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Distilling Knowledge From Object Classification to Aesthetics AssessmentabstractIn this work, we point out that the major dilemma of image aesthetics assessment (IAA) comes from the abstract nature of aesthetic labels. That is, a vast variety of distinct contents can correspond to the same aesthetic label. On the one hand, during inference, the IAA model is required to relate various distinct contents to the same aesthetic label. On the other hand, when training, it would be hard for the IAA model to learn to distinguish different contents merely with the supervision from aesthetic labels, since aesthetic labels are not directly related to any specific content. To deal with this dilemma, we propose to distill knowledge on semantic patterns for a vast variety of image contents from multiple pre-trained object classification (POC) models to an IAA model. Expecting the combination of multiple POC models can provide sufficient knowledge on various image contents, the IAA model can easier learn to relate various distinct contents to a limited number of aesthetic labels. By supervising an end-to-end single-backbone IAA model with the distilled knowledge, the performance of the IAA model is significantly improved by 4.8% in SRCC compared to the version trained only with ground-truth aesthetic labels. On specific categories of images, the SRCC improvement brought by the proposed method can achieve up to 7.2%. Peer comparison also shows that our method outperforms 10 previous IAA methods. Jingwen Hou, Henghui Ding, Weisi Lin, Weide Liu, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | DACNN: Blind Image Quality Assessment via a Distortion-Aware Convolutional Neural NetworkabstractDeep neural networks have achieved great performance on blind Image Quality Assessment (IQA), but it is still challenging for using one network to accurately predict the quality of images with different distortions. In this paper, a Distortion-Aware Convolutional Neural Network (DACNN) is proposed for blind IQA, which works effectively for not only synthetically distorted images but also authentically distorted images. The proposed DACNN consists of a distortion aware module, a distortion fusion module, and a quality prediction module. In the distortion aware module, a Siamese network-based pretraining strategy is proposed to design a synthetic distortion-aware network for full learning the synthetic distortions, and an authentic distortion-aware network is used for extracting the authentic distortions. To efficiently fuse the learned distortion features, and make the network pay more attention to the essential features, a weight-adaptive fusion network is proposed to adaptively adjust the weight of each distortion. Finally, the quality prediction module is adopted to map the fused features to a quality score. Extensive experiments on four authentic IQA databases and four synthetic IQA databases have proved the effectiveness of the proposed DACNN. Zhaoqing Pan, Jianjun Lei 0001, Yuming Fang 0001, Xiao Shao, Nam Ling, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Depth Estimation Using a Self-Supervised Network Based on Cross-Layer Feature Fusion and the Quadtree ConstraintabstractDepth estimation from a camera is an important task for 3D perception. Recently, without using the labeled ground truth of depth map, a self-supervised deep learning network can use relative pose to synthesize the target image from the reference image, and the photometric error between synthesized reference image and real one is used as self-supervisory signal. In this paper, we propose a novel self-supervised depth estimation network, which takes advantage of the quadtree constraint to optimize the depth estimation network. Based on the quadtree constraint, the photometric loss and depth loss of quadtree are proposed. In order to solve the problem that multiple depth values in repeated structures and uniform texture regions can cause relatively low photometric loss, we use quadtree-based photometric loss, which calculates the averaged photometric loss in quadtree blocks instead of the pixel-wise loss. For the problem of imbalanced depth distribution, we use quadtree depth loss, which constrains the depth inconsistency within quadtree blocks. The depth estimation network is composed of deep fusion module and cross-layer feature fusion module, which can better extract the feature information of RGB image and sparse keypoints depths, and makes full use of the detail information of the shallow feature map and the semantic information of the deep feature map to enrich the feature information extraction. Experimental results demonstrate that our method outperforms the state-of-the-art approaches of depth estimation. Yongbin Gao, Zhijun Fang 0001, Yuming Fang 0001, Hamido Fujita, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Omnidirectional Image Quality Assessment by Distortion Discrimination Assisted Multi-Stream NetworkabstractOmnidirectional image (OI) quality assessment is crucial to facilitate the development of virtual reality (VR) related technology. In this work, a distortion discrimination assisted multi-stream network is proposed for OI quality assessment. The multi-stream architecture is constructed by generating the viewport images received by the retina at one point to simulate the characteristics of humans perceiving VR contents. Additionally, the strategy of generating several viewport image sets from one OI is proposed for data augmentation. Furthermore, the facts that the human brain has the ability for both quality assessment and distortion type distinguishment, and the process of human brain handling two tasks exists information interaction inspire us to employ an auxiliary distortion discrimination task to facilitate the quality assessment task learning. Extensive experiments conducted on two public OI databases demonstrate the superiority of the proposed method to both traditional 2D quality metrics and existing metrics specific for OIs. Moreover, utilizing the assistant task is proven to be more effective than the single task learning for OI quality evaluation. Better generalization performance is also verified to be another valuable trait of the proposed method. Yu Zhou 0009, Yanjing Sun, Leida Li, Ke Gu 0001, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Visual Cluster Grounding for Image CaptioningabstractAttention mechanisms have been extensively adopted in vision and language tasks such as image captioning. It encourages a captioning model to dynamically ground appropriate image regions when generating words or phrases, and it is critical to alleviate the problems of object hallucinations and language bias. However, current studies show that the grounding accuracy of existing captioners is still far from satisfactory. Recently, much effort is devoted to improving the grounding accuracy by linking the words to the full content of objects in images. However, due to the noisy grounding annotations and large variations of object appearance, such strict word-object alignment regularization may not be optimal for improving captioning performance. In this paper, to improve the performance of both grounding and captioning, we propose a novel grounding model which implicitly links the words to the evidence in the image. The proposed model encourages the captioner to dynamically focus on informative regions of the objects, which could be either discriminative parts or full object content. With slacked constraints, the proposed captioning model can capture correct linguistic characteristics and visual relevance, and then generate more grounded image captions. In addition, we propose a novel quantitative metric for evaluating the correctness of the soft attention mechanism by considering the overall contribution of all object proposals when generating certain words. The proposed grounding model can be seamlessly plugged into most attention-based architectures without introducing inference complexity. We conduct extensive experiments on Flickr30k (Young et al., 2014) and MS COCO datasets (Lin et al., 2014), demonstrating that the proposed method consistently improves image captioning in both grounding and captioning. Besides, the proposed attention evaluation metric shows better consistency with the captioning performance. Wenhui Jiang 0001, Minwei Zhu, Yuming Fang 0001, Guangming Shi, Yang Liu 0293 |
IEEE Trans. Image Process. | 3 |
| 2022 | VCRNet: Visual Compensation Restoration Network for No-Reference Image Quality AssessmentabstractGuided by the free-energy principle, generative adversarial networks (GAN)-based no-reference image quality assessment (NR-IQA) methods have improved the image quality prediction accuracy. However, the GAN cannot well handle the restoration task for the free-energy principle-guided NR-IQA methods, especially for the severely destroyed images, which results in that the quality reconstruction relationship between the distorted image and its restored image cannot be accurately built. To address this problem, a visual compensation restoration network (VCRNet)-based NR-IQA method is proposed, which uses a non-adversarial model to efficiently handle the distorted image restoration task. The proposed VCRNet consists of a visual restoration network and a quality estimation network. To accurately build the quality reconstruction relationship between the distorted image and its restored image, a visual compensation module, an optimized asymmetric residual block, and an error map-based mixed loss function, are proposed for increasing the restoration capability of the visual restoration network. For further addressing the NR-IQA problem of severely destroyed images, the multi-level restoration features which are obtained from the visual restoration network are used for the image quality estimation. To prove the effectiveness of the proposed VCRNet, seven representative IQA databases are used, and experimental results show that the proposed VCRNet achieves the state-of-the-art image quality prediction accuracy. The implementation of the proposed VCRNet has been released at https://github.com/NUIST-Videocoding/VCRNet. Zhaoqing Pan, Jianjun Lei 0001, Yuming Fang 0001, Xiao Shao, Sam Kwong |
IEEE Trans. Image Process. | 4 |
| 2022 | Subjective and Objective Quality of Experience of Free Viewpoint VideosabstractFree viewpoint videos (FVVs) provide immersive experiences for end-users, and they have been applied in many applications, such as movies, sports, and TV shows. However, the development of quantifying the quality of experience (QoE) of FVVs is still relatively slow due to the high costs of data collection and limited public databases. In this paper, we conduct a comprehensive study on FVV QoE. First, we construct the largest, to the best of our knowledge, FVV QoE database called Youku-FVV from two complex real scenarios, i. e., entertainment and sports. Specifically, Youku-FVV originates from the videos captured by dozens of real cameras arranged annularly. We use these videos to generate virtual viewpoints, which make up FVVs together with real views. In constructing the FVV QoE database, we consider both internal and external influencing factors of QoE, which correspond to FVV generation and playback, respectively. Besides, we make an initial attempt to train an efficient no reference FVV QoE prediction model using this database, where several sparse frame sampling strategies are validated. And we demonstrate the feasibility of striving for the balance between effectiveness and efficiency of FVV QoE prediction. The proposed FVV QoE database and source codes are publicly available at https://github.com/QTJiebin/FVV_QoE. Jiebin Yan, Jing Li 0026, Yuming Fang 0001, Zhaohui Che, Xue Xia 0005, Yang Liu 0293 |
IEEE Trans. Image Process. | 3 |
| 2022 | Identity-Aware Facial Expression Recognition Via Deep Metric Learning Based on Synthesized ImagesabstractPerson-dependent facial expression recognition has received considerable research attention in recent years. Unfortunately, different identities can adversely influence recognition accuracy, and the recognition task becomes challenging. Other adverse factors, including limited training data and improper measures of facial expressions, can further contribute to the above dilemma. To solve these problems, a novel identity-aware method is proposed in this study. Furthermore, this study also represents the first attempt to fulfill the challenging person-dependent facial expression recognition task based on deep metric learning and facial image synthesis techniques. Technically, a StarGAN is incorporated to synthesize facial images depicting different but complete basic emotions for each identity to augment the training data. Then, a deep-convolutional-neural-network-based network is employed to automatically extract latent features from both real facial images and all synthesized facial images. Next, a Mahalanobis metric network trained based on extracted latent features outputs a learned metric that measures facial expression differences between images, and the recognition task can thus be realized. Extensive experiments based on several well-known publicly available datasets are carried out in this study for performance evaluations. Person-dependent datasets, including CK+, Oulu (all 6 subdatasets), MMI, ISAFE, ISED, etc., are all incorporated. After comparing the new method with several popular or state-of-the-art facial expression recognition methods, its superiority in person-dependent facial expression recognition can be proposed from a statistical point of view. Wei Huang 0013, Peng Zhang 0005, Yufei Zha, Yuming Fang 0001, Yanning Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | MIG-Net: Multi-Scale Network Alternatively Guided by Intensity and Gradient Features for Depth Map Super-ResolutionabstractThe studies of previous decades have shown that the quality of depth maps can be significantly lifted by introducing the guidance from intensity images describing the same scenes. With the rising of deep convolutional neural network, the performance of guided depth map super-resolution is further improved. The variants always consider deep structure, optimized gradient flow and feature reusing. Nevertheless, it is difficult to obtain sufficient and appropriate guidance from intensity features without any prior. In fact, features in the gradient domain, e.g., edges, present strong correlations between the intensity image and the corresponding depth map. Therefore, the guidance in the gradient domain can be more efficiently explored. In this paper, the depth features are iteratively upsampled by 2×. In each upsampling stage, the low-quality depth features and the corresponding gradient features are iteratively refined by the guidance from the intensity features via two parallel streams. Then, to make full use of depth features in the image and gradient domains, the depth features and gradient features are alternatively complemented with each other. Compared with state-of-the-art counterparts, the sufficient experimental results show improvements according to the objective and subjective assessments. The code is available athttps://github.com/Yifan-Zuo/MIG-net-gradient_guided_depth_enhancement. Yifan Zuo 0001, Yuming Fang 0001, Xiaoshui Huang, Xiwu Shang, Qiang Wu 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Perceptual Quality Assessment of Omnidirectional Images as Moving Camera VideosabstractOmnidirectional images (also referred to as static 360$^{\circ }$panoramas) impose viewing conditions much different from those of regular 2D images. How do humans perceive image distortions in immersive virtual reality (VR) environments is an important problem which receives less attention. We argue that, apart from the distorted panorama itself, two types of VR viewing conditions are crucial in determining the viewing behaviors of users and the perceived quality of the panorama: the starting point and the exploration time. We first carry out a psychophysical experiment to investigate the interplay among the VR viewing conditions, the user viewing behaviors, and the perceived quality of 360$^{\circ }$images. Then, we provide a thorough analysis of the collected human data, leading to several interesting findings. Moreover, we propose a computational framework for objective quality assessment of 360$^{\circ }$images, embodying viewing conditions and behaviors in a delightful way. Specifically, we first transform an omnidirectional image to several video representations using different user viewing behaviors under different viewing conditions. We then leverage advanced 2D full-reference video quality models to compute the perceived quality. We construct a set of specific quality measures within the proposed framework, and demonstrate their promises on three VR quality databases. Xiangjie Sui, Kede Ma, Yiru Yao, Yuming Fang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | Image Quality Assessment in the Modern AgeabstractThis tutorial provides the audience with the basic theories, methodologies, and current progresses of image quality assessment (IQA). From an actionable perspective, we will first revisit several subjective quality assessment methodologies, with emphasis on how to properly select visual stimuli. We will then present in detail the design principles of objective quality assessment models, supplemented by an in-depth analysis of their advantages and disadvantages. Both hand-engineered and (deep) learning-based methods will be covered. Moreover, the limitations with the conventional model comparison methodology for objective quality models will be pointed out, and novel comparison methodologies such as those based on the theory of "analysis by synthesis" will be introduced. We will last discuss the real-world multimedia applications of IQA, and give a list of open challenging problems, in the hope of encouraging more and more talented researchers and engineers devoting to this exciting and rewarding research field. Kede Ma, Yuming Fang 0001 |
ACM Multimedia | 2 |
| 2021 | Anomaly detection in video sequences: A benchmark and computational modelabstractAbstract Anomaly detection has attracted considerable search attention. However, existing anomaly detection databases encounter two major problems. Firstly, they are limited in scale. Secondly, training sets contain only video‐level labels indicating the existence of an abnormal event during the full video while lacking annotations of precise time durations. To tackle these problems, we contribute a new L arge‐scale A nomaly D etection ( LAD ) database as the benchmark for anomaly detection in video sequences, which is featured in two aspects. 1) It contains 2000 video sequences including normal and abnormal video clips with 14 anomaly categories including crash, fire, violence etc . with large scene varieties, making it the largest anomaly analysis database to date. 2) It provides the annotation data, including video‐level labels (abnormal/normal video, anomaly type) and frame‐level labels (abnormal/normal video frame) to facilitate anomaly detection. Leveraging the above benefits from the LAD database, we further formulate anomaly detection as a fully supervised learning problem and propose a multi‐task deep neural network to solve it. We firstly obtain the local spatiotemporal contextual feature by using an Inflated 3D convolutional (I3D) network. Then we construct a recurrent convolutional neural network fed the local spatiotemporal contextual feature to extract the spatiotemporal contextual feature. With the global spatiotemporal contextual feature, the anomaly type and score can be computed simultaneously by a multi‐task neural network. Experimental results show that the proposed method outperforms the state‐of‐the‐art anomaly detection methods on our database and other public databases of anomaly detection. Supplementary materials are available at http://sim.jxufe.cn/JDMKL/ymfang/anomaly‐detection.html . Boyang Wan, Wenhui Jiang 0001, Yuming Fang 0001, Zhiyuan Luo 0002, Guanqun Ding |
IET Image Process. | 3 |
| 2021 | Exposing Semantic Segmentation Failures via Maximum Discrepancy Competition
Jiebin Yan, Yuming Fang 0001, Zhangyang Wang, Kede Ma |
Int. J. Comput. Vis. | 3 |
| 2021 | Unsupervised stereoscopic image retargeting via view synthesis and stereo cycle consistency losses
Xiaoting Fan, Jianjun Lei 0001, Jie Liang 0001, Yuming Fang 0001, Xiaochun Cao, Nam Ling |
Neurocomputing | 4 |
| 2021 | Dynamic proposal sampling for weakly supervised object detection
Wenhui Jiang 0001, Zhicheng Zhao 0001, Yuming Fang 0001 |
Neurocomputing | 4 |
| 2021 | An efficient and high-quality pansharpening model based on conditional random fields
Yong Yang 0001, Hangyuan Lu, Shuying Huang, Yuming Fang 0001, Wei Tu 0002 |
Inf. Sci. | 4 |
| 2021 | Single Image Deraining: From Model-Based to Data-Driven and BeyondabstractThe goal of single-image deraining is to restore the rain-free background scenes of an image degraded by rain streaks and rain accumulation. The early single-image deraining methods employ a cost function, where various priors are developed to represent the properties of rain and background layers. Since 2017, single-image deraining methods step into a deep-learning era, and exploit various types of networks, i.e., convolutional neural networks, recurrent neural networks, generative adversarial networks, etc., demonstrating impressive performance. Given the current rapid development, in this paper, we provide a comprehensive survey of deraining methods over the last decade. We summarize the rain appearance models, and discuss two categories of deraining approaches: model-based and data-driven approaches. For the former, we organize the literature based on their basic models and priors. For the latter, we discuss the developed ideas related to architectures, constraints, loss functions, and training datasets. We present milestones of single-image deraining methods, review a broad selection of previous works in different categories, and provide insights on the historical development route from the model-based to data-driven methods. We also summarize performance comparisons quantitatively and qualitatively. Beyond discussing the technicality of deraining methods, we also discuss the future possible directions. Wenhan Yang, Robby T. Tan, Shiqi Wang 0001, Yuming Fang 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Asymmetrically distorted 3D video quality assessment: From the motion variation to perceived quality
Yuming Fang 0001, Xiangjie Sui, Jiebin Yan, Yifan Zuo 0001, Jiheng Wang, Zhaoqian Li |
Signal Process. | 1 |
| 2021 | Light field image encryption based on spatial-angular characteristic
Kangkang Wei, Wenying Wen, Yuming Fang 0001 |
Signal Process. | 3 |
| 2021 | Visual attention prediction for Autism Spectrum Disorder with hierarchical semantic fusion
Yuming Fang 0001, Yifan Zuo 0001, Wenhui Jiang 0001, Hanqin Huang, Jiebin Yan |
Signal Process. Image Commun. | 1 |
| 2021 | Objective quality assessment of synthesized images by local variation measurement
Xiangjie Sui, Mengna Ding, Jiebin Yan, Yuming Fang 0001, Yifan Zuo 0001, Zuowen Tan |
Signal Process. Image Commun. | 4 |
| 2021 | Stereoscopic Image Retargeting Based on Deep Convolutional Neural NetworkabstractStereoscopic image retargeting aims at converting stereoscopic images to the target resolution adaptively. Different from 2D image retargeting, stereoscopic image retargeting needs to preserve both the shape structure of salient objects and depth consistency of 3D scenes. In this paper, we present a stereoscopic image retargeting method based on deep convolutional neural network to obtain high-quality retargeted images with both object shape preservation and scene depth preservation. First, a cross-attention extraction mechanism is constructed to generate attention map, which contains the valuable attention features of the left and right images and the common attention features between them. Second, since the disparity map can provide accurate depth information of objects in 3D scenes, a disparity-assisted 3D significance map generation module is utilized to further preserve the valuable depth information of stereoscopic images. Finally, in order to predict the retargeted stereoscopic images accurately, an image consistency loss is developed to preserve the geometric structure of salient objects, and a disparity consistency loss is introduced to eliminate depth distortions. Experimental results demonstrate that the proposed deep convolutional neural network can provide favorable stereoscopic image retargeting results. Xiaoting Fan, Jianjun Lei 0001, Jie Liang 0001, Yuming Fang 0001, Nam Ling, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Perceptual Quality Assessment for Asymmetrically Distorted Stereoscopic Video by Temporal Binocular RivalryabstractIn this paper, we propose a two-stage weighting based perceptual quality assessment framework for asymmetrically distorted stereoscopic video (SV) sequences by temporal binocular rivalry. Firstly, a traditional 2D image quality assessment (IQA) method is employed to measure spatial distortion, and the temporal distortion is evaluated by the magnitude differences between motion vectors of distorted and reference video frames. Secondly, the structural strength (SS) computed by gradient map and the motion energy (ME) computed by frame difference map are used to estimate the intensity of visual stimulus in spatial and temporal domain respectively. Then, SS and ME are considered as the importance indexes to combine the quality scores of spatial and temporal distortion to estimate perceived distortion of single-view video sequences, which is denoted as the first-stage weighting. Finally, considering that the difference of intensity of visual stimulus between two eyes results in binocular rivalry, a novel temporal binocular rivalry inspired weighting method is designed to integrate the quality scores of left- and right-views for the final visual quality prediction of SV sequences, which is denoted as the second-stage weighting. Experimental results on Waterloo-IVC SV quality databases show that several specific examples of 2D-IQA methods within the proposed framework can obtain highly competitive performance over other existing ones. Yuming Fang 0001, Xiangjie Sui, Jiheng Wang, Jiebin Yan, Jianjun Lei 0001, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Predicting the Quality of View Synthesis With Color-Depth Image FusionabstractWith the increasing prevalence of free-viewpoint video applications, virtual view synthesis has attracted extensive attention. In view synthesis, a new viewpoint is generated from the input color and depth images with a depth-image-based rendering (DIBR) algorithm. Current quality evaluation models for view synthesis typically operate on the synthesized images, i.e. after the DIBR process, which is computationally expensive. So a natural question is that can we infer the quality of DIBR-based synthesized images using the input color and depth images directly without performing the intricate DIBR operation. With this motivation, this paper presents a no-reference image quality prediction model for view synthesis via COlor-Depth Image Fusion, dubbed CODIF, where the actual DIBR is not needed. First, object boundary regions are detected from the color image, and a Wavelet-based image fusion method is proposed to imitate the interaction between color and depth images during the DIBR process. Then statistical features of the interactional regions and natural regions are extracted from the fused color-depth image to portray the influences of distortions in color/depth images on the quality of synthesized views. Finally, all statistical features are utilized to learn the quality prediction model for view synthesis. Extensive experiments on public view synthesis databases demonstrate the advantages of the proposed metric in predicting the quality of view synthesis, and it even suppresses the state-of-the-art post-DIBR view synthesis quality metrics. Leida Li, Yipo Huang, Jinjian Wu, Ke Gu 0001, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Visual Quality Assessment for Perceptually Encrypted Light Field ImagesabstractPerceptual encryption has received widespread attention as a technology for protecting multimedia visual information. In multimedia applications, perceptual encryption only protects a portion of the content without previewing all the information. At present, perceptual encryption is mostly oriented to conventional plain images while few methods aim at visual security measures for perceptually encrypted images. Existing solutions usually adopt well-known quality assessment metrics to measure the visual quality of encrypted images. However, they often exhibit undesired behavior on perceptually encrypted images with low quality. As a typical representation of three-dimensional scenes, light field images record the intensity and direction of light during propagation which is distinct from 2-D images. In this paper, we construct a perceptually encrypted light field image database (PE-LFID) for quality assessment based on 14 reference light field plaintext images. For each scene, we employ four encryption methods, each of which has six levels. Additionally, a novel visual security evaluation method based on PE-LFID is proposed by taking into account the local and global features of the light field images. First, we use a multi-threshold edge detection method to obtain the edge similarity in the spatial domain of the light field image. Afterwards, the epipolar plane image (EPI) generated from the angular domain of the light field image is used to calculate the gradient magnitude similarity, which is expressed as a global feature. Furthermore, the final quality prediction score is calculated by adaptively weighting between local and global features. We conduct extensive experiments on the proposed PE-LFID to assess the performance of classical and state-of-the-art IQA models. The experimental results demonstrate the effectiveness of the proposed method for visual security evaluation of perceptually encrypted light field images, as well as the scalability of PE-LFID. Wenying Wen, Kangkang Wei, Yuming Fang 0001, Yushu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Exploring Rich and Efficient Spatial Temporal Interactions for Real-Time Video Salient Object DetectionabstractWe have witnessed a growing interest in video salient object detection (VSOD) techniques in today's computer vision applications. In contrast with temporal information (which is still considered a rather unstable source thus far), the spatial information is more stable and ubiquitous, thus it could influence our vision system more. As a result, the current main-stream VSOD approaches have inferred and obtained their saliency primarily from the spatial perspective, still treating temporal information as subordinate. Although the aforementioned methodology of focusing on the spatial aspect is effective in achieving a numeric performance gain, it still has two critical limitations. First, to ensure the dominance by the spatial information, its temporal counterpart remains inadequately used, though in some complex video scenes, the temporal information may represent the only reliable data source, which is critical to derive the correct VSOD. Second, both spatial and temporal saliency cues are often computed independently in advance and then integrated later on, while the interactions between them are omitted completely, resulting in saliency cues with limited quality. To combat these challenges, this paper advocates a novel spatiotemporal network, where the key innovation is the design of its temporal unit. Compared with other existing competitors (e.g., convLSTM), the proposed temporal unit exhibits an extremely lightweight design that does not degrade its strong ability to sense temporal information. Furthermore, it fully enables the computation of temporal saliency cues that interact with their spatial counterparts, ultimately boosting the overall VSOD performance and realizing its full potential towards mutual performance improvement for each. The proposed method is easy to implement yet still effective, achieving high-quality VSOD at 50 FPS in real-time applications. Chenglizhao Chen, Guotao Wang 0004, Chong Peng 0001, Yuming Fang 0001, Dingwen Zhang, Hong Qin 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Superpixel-Based Quality Assessment of Multi-Exposure Image Fusion for Both Static and Dynamic ScenesabstractMulti-exposure image fusion (MEF) algorithms have been used to merge a stack of low dynamic range images with various exposure levels into a well-perceived image. However, little work has been dedicated to predicting the visual quality of fused images. In this work, we propose a novel and efficient objective image quality assessment (IQA) model for MEF images of both static and dynamic scenes based on superpixels and an information theory adaptive pooling strategy. First, with the help of superpixels, we divide fused images into large- and small-changed regions using the structural inconsistency map between each exposure and fused images. Then, we compute the quality maps based on the Laplacian pyramid for large- and small-changed regions separately. Finally, an information theory induced adaptive pooling strategy is proposed to compute the perceptual quality of the fused image. Experimental results on three public databases of MEF images demonstrate the proposed model achieves promising performance and yields a relatively low computational complexity. Additionally, we also demonstrate the potential application for parameter tuning of MEF algorithms. Yuming Fang 0001, Yan Zeng 0001, Wenhui Jiang 0001, Hanwei Zhu, Jiebin Yan |
IEEE Trans. Image Process. | 1 |
| 2021 | Data-Level Recombination and Lightweight Fusion Scheme for RGB-D Salient Object DetectionabstractExisting RGB-D salient object detection methods treat depth information as an independent component to complement RGB and widely follow the bistream parallel network architecture. To selectively fuse the CNN features extracted from both RGB and depth as a final result, the state-of-the-art (SOTA) bistream networks usually consist of two independent subbranches: one subbranch is used for RGB saliency, and the other aims for depth saliency. However, depth saliency is persistently inferior to the RGB saliency because the RGB component is intrinsically more informative than the depth component. The bistream architecture easily biases its subsequent fusion procedure to the RGB subbranch, leading to a performance bottleneck. In this paper, we propose a novel data-level recombination strategy to fuse RGB with D (depth) before deep feature extraction, where we cyclically convert the original 4-dimensional RGB-D into DGB, RDB and RGD. Then, a newly lightweight designed triple-stream network is applied over these novel formulated data to achieve an optimal channel-wise complementary fusion status between the RGB and D, achieving a new SOTA performance. Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Yuming Fang 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Band Representation-Based Semi-Supervised Low-Light Image Enhancement: Bridging the Gap Between Signal Fidelity and Perceptual QualityabstractIt has been widely acknowledged that under-exposure causes a variety of visual quality degradation because of intensive noise, decreased visibility, biased color, etc. To alleviate these issues, a novel semi-supervised learning approach is proposed in this paper for low-light image enhancement. More specifically, we propose a deep recursive band network (DRBN) to recover a linear band representation of an enhanced normal-light image based on the guidance of the paired low/normal-light images. Such design philosophy enables the principled network to generate a quality improved one by reconstructing the given bands based upon another learnable linear transformation which is perceptually driven by an image quality assessment neural network. On one hand, the proposed network is delicately developed to obtain a variety of coarse-to-fine band representations, of which the estimations benefit each other in a recursive process mutually. On the other hand, the extracted band representation of the enhanced image in the recursive band learning stage of DRBN is capable of bridging the gap between the restoration knowledge of paired data and the perceptual quality preference to high-quality images. Subsequently, the band recomposition learns to recompose the band representation towards fitting perceptual regularization of high-quality images with the perceptual guidance. The proposed architecture can be flexibly trained with both paired and unpaired data. Extensive experiments demonstrate that our method produces better enhanced results with visually pleasing contrast and color distributions, as well as well-restored structural details. Wenhan Yang, Shiqi Wang 0001, Yuming Fang 0001, Yue Wang 0032, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Dense Attention Fluid Network for Salient Object Detection in Optical Remote Sensing ImagesabstractDespite the remarkable advances in visual saliency analysis for natural scene images (NSIs), salient object detection (SOD) for optical remote sensing images (RSIs) still remains an open and challenging problem. In this paper, we propose an end-to-end Dense Attention Fluid Network (DAFNet) for SOD in optical RSIs. A Global Context-aware Attention (GCA) module is proposed to adaptively capture long-range semantic context relationships, and is further embedded in a Dense Attention Fluid (DAF) structure that enables shallow attention cues flow into deep layers to guide the generation of high-level feature attention maps. Specifically, the GCA module is composed of two key components, where the global feature aggregation module achieves mutual reinforcement of salient feature embeddings from any two spatial locations, and the cascaded pyramid attention module tackles the scale variation issue by building up a cascaded pyramid framework to progressively refine the attention map in a coarse-to-fine manner. In addition, we construct a new and challenging optical RSI dataset for SOD that contains 2,000 images with pixel-wise saliency annotations, which is currently the largest publicly available benchmark. Extensive experiments demonstrate that our proposed DAFNet significantly outperforms the existing state-of-the-art SOD competitors. https://github.com/rmcong/DAFNet_TIP20. Qijian Zhang, Runmin Cong, Chongyi Li, Ming-Ming Cheng, Yuming Fang 0001, Xiaochun Cao, Yao Zhao 0001, Sam Kwong |
IEEE Trans. Image Process. | 5 |
| 2021 | Blind Quality Assessment for Tone-Mapped Images by Analysis of Gradient and Chromatic StatisticsabstractA tone-mapped image (TMI) obtained from the corresponding high dynamic range (HDR) image induces artifacts and distortion, which might result in the loss of structure information and impaired color. By analyzing the visual characteristics of TMIs, this work proposes a robust blind visual quality evaluation method for TMIs by using gradient and chromatic statistics (VQGC). First, motivated by the perceptual mechanism that the human visual system (HVS) is sensitive to image structure variation, we employ the gradient features to measure structure degradation in TMIs. To predict structure distortion accurately, we compute the gradient magnitude and orientation to measure image structure variation, and the relative gradient magnitude and orientation are also computed to capture microstructure change. Second, the color invariance descriptors are utilized to capture the visual degradation of colorfulness by local binary pattern (LBP) on four chromatic feature maps. Finally, the gradient and chromatic features are combined together as the final quality-aware feature vector, which is applied to assess the perceptual quality of TMIs by support vector regression (SVR). Comparison experiments show that the performance of the proposed method is better than other existing blind quality assessment methods on public databases. Yuming Fang 0001, Jiebin Yan, Rengang Du, Yifan Zuo 0001, Wenying Wen, Yan Zeng 0001, Leida Li |
IEEE Trans. Multim. | 1 |
| 2021 | Quality Evaluation for Image Retargeting With Instance SemanticsabstractTo meet the ever-increasing demand for devices with diversified displays, image retargeting has become a prevalent technique for adaptive image resizing. In practice, the retargeting operation inevitably causes impairments in the images; thus, image retargeting quality assessment (IRQA) is urgently needed and, can be used to guide algorithm optimization, selection and design. Unlike traditional image quality assessment, image retargeting introduces geometric distortions, which typically affect high-level image semantics. With this motivation, this paper presents a quality evaluation model for image retargeting based on INstance SEMantics (INSEM). Considering that the human visual system (HVS) perceives images highly dependent on apprehensible areas and that impairments in image retargeting mainly degrade the salient instances, an image instance is utilized as the basic semantic unit, and a top-down method is devised to extract instance-level semantic features for IRQA. In addition, taking into account the influence of semantic categories on the perception of retargeting quality, we further propose Semantic-based self-adaptive pooling (SSAP) to integrate instance-based semantic features. Finally, global features are incorporated to generate quality scores that are more consistent with people's perceptions. Extensive experiments and comparisons of three public databases, in terms of both intradatabase and cross-database settings, demonstrate the superiority of the proposed metric over state-of-the-art methods. Leida Li, Jinjian Wu, Lin Ma 0002, Yuming Fang 0001 |
IEEE Trans. Multim. | 5 |
| 2021 | Frequency-Dependent Depth Map Enhancement via Iterative Depth-Guided Affine Transformation and Intensity-Guided RefinementabstractRecently, deep convolutional neural network sho-ws significant improvement for intensity-guided depth map enhancement. The most networks focus on either increasing depth or easing features propagation via residual learning and dense connection. However, it has not been explicitly considered yet to mitigate the artifacts caused by the differences of the distributions between the depth map and the corresponding color image, e.g., edge misalignment. In this paper, a novel depth-guided affine transformation is used to filter out the unrelated intensity features, which is further used to refine the depth features. Since the quality of initial depth features is low, the depth-guided intensity features filtering and the intensity-guided depth features refinement are iteratively performed, which progressively promotes effects of such tasks. To make full use of the iterations, all the refined depth features are dense connected followed by a 1 × 1 convolution layer. In addition, to improve the performance in the case of large upsampling factors (e.g., 16×), the depth features are enhanced from coarse to fine. In each frequency-dependent refinement of the depth features, the above iterative subnetwork as well as the residual learning are introduced. The proposed method is tested for the noise-free and noisy cases which compares against 16 state-of-the-art methods. Our experimental results show the improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Yuming Fang 0001, Ping An 0001, Xiwu Shang, Junnan Yang |
IEEE Trans. Multim. | 2 |
| 2020 | Perceptual Quality Assessment of Smartphone PhotographyabstractAs smartphones become people's primary cameras to take photos, the quality of their cameras and the associated computational photography modules has become a de facto standard in evaluating and ranking smartphones in the consumer market. We conduct so far the most comprehensive study of perceptual quality assessment of smartphone photography. We introduce the Smartphone Photography Attribute and Quality (SPAQ) database, consisting of 11,125 pictures taken by 66 smartphones, where each image is attached with so far the richest annotations. Specifically, we collect a series of human opinions for each image, including image quality, image attributes (brightness, colorfulness, contrast, noisiness, and sharpness), and scene category labels (animal, cityscape, human, indoor scene, landscape, night scene, plant, still life, and others) in a well-controlled laboratory environment. The exchangeable image file format (EXIF) data for all images are also recorded to aid deeper analysis. We also make the first attempts using the database to train blind image quality assessment (BIQA) models constructed by baseline and multi-task deep neural networks. The results provide useful insights on how EXIF data, image attributes and high-level semantics interact with image quality, how next-generation BIQA models can be designed, and how better computational photography systems can be optimized on mobile devices. The database along with the proposed BIQA models are available at https://github.com/h4nwei/SPAQ. Yuming Fang 0001, Hanwei Zhu, Yan Zeng 0001, Kede Ma, Zhou Wang 0001 |
CVPR | 1 |
| 2020 | From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image EnhancementabstractUnder-exposure introduces a series of visual degradation, i.e. decreased visibility, intensive noise, and biased color, etc. To address these problems, we propose a novel semi-supervised learning approach for low-light image enhancement. A deep recursive band network (DRBN) is proposed to recover a linear band representation of an enhanced normal-light image with paired low/normal-light images, and then obtain an improved one by recomposing the given bands via another learnable linear transformation based on a perceptual quality-driven adversarial learning with unpaired data. The architecture is powerful and flexible to have the merit of training with both paired and unpaired data. On one hand, the proposed network is well designed to extract a series of coarse-to-fine band representations, whose estimations are mutually beneficial in a recursive process. On the other hand, the extracted band representation of the enhanced image in the first stage of DRBN (recursive band learning) bridges the gap between the restoration knowledge of paired data and the perceptual quality preference to real high-quality images. Its second stage (band recomposition) learns to recompose the band representation towards fitting perceptual properties of high-quality images via adversarial learning. With the help of this two-stage design, our approach generates enhanced results with well-reconstructed details and visually promising contrast and color distributions. Qualitative and quantitative evaluations demonstrate the superiority of our DRBN. Wenhan Yang, Shiqi Wang 0001, Yuming Fang 0001, Yue Wang 0032, Jiaying Liu 0001 |
CVPR | 3 |
| 2020 | Blind Stereoscopic Image Quality Assessment By Deep Neural Network Of Multi-Level Feature FusionabstractIn this paper, we propose an effective blind image quality assessment (BIQA) method for stereoscopic images by deep neural network (DNN) of multi-level feature fusion (MLFF) inspired by the multi-scale characteristics and binocular properties of the human visual system (HVS). Specifically, we firstly feed the left- and right-view images into a weight sharing convolutional neural network (CNN) for jointly feature extraction. To aggregate multi-level features, we concatenate the low-, middle-, and high-level feature maps of stereoscopic images to simulate the complicated visual interaction processing in the HVS. Two fully connected layers are used to build the nonlinear mapping from the highly abstract features to the quality scores of stereoscopic images. The experiments conducted on two public databases prove the validity of the proposed MLFF method. Jiebin Yan, Yuming Fang 0001, Xiongkuo Min, Yiru Yao, Guangtao Zhai |
ICME | 2 |
| 2020 | Image dehazing based on a transmission fusion strategy by automatic image matting
Feiniu Yuan, Yu Zhou 0009, Xue Xia 0005, Jinting Shi, Yuming Fang 0001, Xueming Qian |
Comput. Vis. Image Underst. | 5 |
| 2020 | Blind quality assessment for tone-mapped images based on local and global features
Xuelin Liu, Yuming Fang 0001, Rengang Du, Yifan Zuo 0001, Wenying Wen |
Inf. Sci. | 2 |
| 2020 | DevsNet: Deep Video Saliency Network using Short-term and Long-term Cues
Yuming Fang 0001, Chi Zhang 0027, Xiongkuo Min, Hanqin Huang, Yugen Yi, Guangtao Zhai, Chia-Wen Lin |
Pattern Recognit. | 1 |
| 2020 | A visually secure image encryption scheme based on semi-tensor product compressed sensing
Wenying Wen, Yukun Hong, Yuming Fang 0001, Ming Li 0029 |
Signal Process. | 3 |
| 2020 | Perceptual objective quality assessment of stereoscopic stitched images
Weiqing Yan, Guanghui Yue 0001, Yuming Fang 0001, Hua Chen 0004, Chang Tang, Gangyi Jiang |
Signal Process. | 3 |
| 2020 | Perceptual Quality Assessment for Screen Content Images by Spatial ContinuityabstractIn this paper, we propose an effective blind quality assessment method for screen content images (SCIs), called perceptual quality measure by spatial continuity (PQSC). With the center-surround mechanism in the human visual system (HVS), the proposed method extracts the statistical features on chromatic and textural variations in SCIs to measure the visual distortion. First, by considering the chromatic continuity between spatially adjacent pixels, photo-metric invariant chromatic descriptors are extracted as zero-order and first-order features. Second, motivated by the perceptual mechanism that the HVS is sensitive to image texture variation, we employ local ternary pattern operator to effectively depict the spatial continuity of texture. With these extracted chromatic and textural features, we further adopt histogram to compute the statistical chromatic and textural features. Support vector regression (SVR) is used to train the quality prediction model from visual features to human ratings. Experimental results on three public benchmark databases demonstrate that the performance of our method is superior to the current blind image quality assessment methods, even better than some full reference image quality assessment counterparts. Yuming Fang 0001, Rengang Du, Yifan Zuo 0001, Wenying Wen, Leida Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Blind Realistic Blur Assessment Based on Discrepancy LearningabstractBlur is one of the most common distortions that degrade natural images. This stimulates the blossom of sharpness assessment metrics. Existing sharpness metrics possess good performance for evaluating simulated blur, but are limited for the more common realistic blur that are introduced during image capture and processing in real life. To this end, we propose an effective Realistic Blur Assessment method (RBA) based on discrepancy learning. First, motivated by the fact that the distortion-free reference images are usually unavailable in practice, but the Human Visual System (HVS) can still accurately perceive image sharpness by quantifying the perceptual discrepancy between the distorted image and the hallucinated reference image in mind, we propose to train a discrepancy generation model to automatically generate the discrepancy map from the distorted image analogous to the HVS. This is achieved by using a deep neural network with rich training images. With the discrepancy map, two sharpness-aware features, i.e. sparse representation based entropy of primitive and content-guided variation of power, are then extracted to severally quantify spatial visual information amount and spectral power. Finally, the two features are integrated to produce the overall sharpness score. Extensive experiments demonstrate the superiority of the proposed method over the state-of-the-arts. Leida Li, Yu Zhou 0009, Ke Gu 0001, Yuzhe Yang 0001, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Multi-Exposure Decomposition-Fusion Model for High Dynamic Range Image Saliency DetectionabstractHigh dynamic range (HDR) imaging techniques have witnessed a great improvement in the past few decades. However, saliency detection task on HDR content is still far from well explored. In this paper, we introduce a multi-exposure decomposition-fusion model for HDR image saliency detection inspired by the brightness adaption mechanism. The proposed model is composed of three modules. Firstly, a decomposition module converts the input raw HDR image into a stack of LDR images by uniformly sampling the exposure time range. Secondly, a saliency region proposal network is employed to generate the candidate saliency maps for each LDR image in the exposure stack. Finally, an uncertainty weighting based fusion algorithm is applied to generate the overall saliency map for the input HDR image by merging the obtained LDR saliency maps. Extensive experiments show that our proposed model achieves superior performance compared with the state-of-the-art methods on the existing HDR eye fixation databases. The source code of the proposed model are made publicly available at https://github.com/sunnycia/DFHSal. Xu Wang 0006, Zhenhao Sun, Qiudan Zhang, Yuming Fang 0001, Lin Ma 0002, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Depth Map Enhancement by Revisiting Multi-Scale Intensity Guidance Within Coarse-to-Fine StagesabstractBeing different from the most methods of guided depth map enhancement based on deep convolutional neural network which focus on increasing the depth of networks, this paper is to improve the effectiveness of intensity guidance when the network goes deep. Overall, the proposed network upsamples the low-resolution depth maps from coarse to fine. Within each refinement stage of certain-scale depth features, the current-scale and all coarse-scales of the guidance features are revisited by dense connection. Therefore, the multi-scale guidance is efficiently maintained as the propagation of features. Furthermore, the proposed network maintains the intensity features in the high-resolution domain from which the multi-scale guidance is directly extracted. This design further improves the quality of intensity guidance. In addition, the shallow depth features upsampled via transposed convolution layer are directly transferred to the final depth features for reconstruction, which is called global residual learning in feature domain. Similarly, the global residual learning in pixel domain learns the difference between the depth ground truth and the coarsely upsampled depth map. Also, the local residual learning is to maintain the low frequency within each refinement stage and progressively recover the high frequency. The proposed method is tested for noise-free and noisy cases which compares against 16 state-of-the-art methods. Our experimental results show the improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Yuming Fang 0001, Yong Yang 0001, Xiwu Shang, Qiang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Multi-Scale Frequency Reconstruction for Guided Depth Map Super-Resolution via Deep Residual NetworkabstractThe depth maps obtained by the consumer-level sensors are always noisy in the low-resolution (LR) domain. Existing methods for the guided depth super-resolution, which are based on the pre-defined local and global models, perform well in general cases (e.g., joint bilateral filter and Markov random field). However, such model-based methods may fail to describe the potential relationship between RGB-D image pairs. To solve this problem, this paper proposes a data-driven approach based on the deep convolutional neural network with global and local residual learning. It progressively upsamples the LR depth map guided by the high-resolution intensity image in multiple scales. A global residual learning is adopted to learn the difference between the ground truth and the coarsely upsampled depth map, and the local residual learning is introduced in each scale-dependent reconstruction sub-network. This scheme can restore the depth structure from coarse to fine via multi-scale frequency synthesis. In addition, batch normalization layers are used to improve the performance of depth map denoising. Our method is evaluated in noise-free and noisy cases. A comprehensive comparison against 17 state-of-the-art methods is carried out. The experimental results show that the proposed method has faster convergence speed as well as improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Qiang Wu 0001, Yuming Fang 0001, Ping An 0001, Liqin Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Perceptual Evaluation for Multi-Exposure Image Fusion of Dynamic ScenesabstractA common approach to high dynamic range (HDR) imaging is to capture multiple images of different exposures followed by multi-exposure image fusion (MEF) in either radiance or intensity domain. A predominant problem of this approach is the introduction of the ghosting artifacts in dynamic scenes with camera and object motion. While many MEF methods (often referred to as deghosting algorithms) have been proposed for reduced ghosting artifacts and improved visual quality, little work has been dedicated to perceptual evaluation of their deghosting results. Here we first construct a database that contains 20 multiexposure sequences of dynamic scenes and their corresponding fused images by nine MEF algorithms. We then carry out a subjective experiment to evaluate fused image quality, and find that none of existing objective quality models for MEF provides accurate quality predictions. Motivated by this, we develop an objective quality model for MEF of dynamic scenes. Specifically, we divide the test image into static and dynamic regions, measure structural similarity between the image and the corresponding sequence in the two regions separately, and combine quality measurements of the two regions into an overall quality score. Experimental results show that the proposed method significantly outperforms the state-of-the-art. In addition, we demonstrate the promise of the proposed model in parameter tuning of MEF methods.1. Yuming Fang 0001, Hanwei Zhu, Kede Ma, Zhou Wang 0001, Shutao Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Deep Guided Learning for Fast Multi-Exposure Image FusionabstractWe propose a fast multi-exposure image fusion (MEF) method, namely MEF-Net, for static image sequences of arbitrary spatial resolution and exposure number. We first feed a low-resolution version of the input sequence to a fully convolutional network for weight map prediction. We then jointly upsample the weight maps using a guided filter. The final image is computed by a weighted fusion. Unlike conventional MEF methods, MEF-Net is trained end-to-end by optimizing the perceptually calibrated MEF structural similarity (MEF-SSIM) index over a database of training sequences at full resolution. Across an independent set of test sequences, we find that the optimized MEF-Net achieves consistent improvement in visual quality for most sequences, and runs 10 to 1000 times faster than state-of-the-art methods. The code is made publicly available at. Kede Ma, Zhengfang Duanmu, Hanwei Zhu, Yuming Fang 0001, Zhou Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | No Reference Quality Assessment for 3D Synthesized Views by Local Structure Variation and Global Naturalness ChangeabstractDepth image based rendering (DIBR) has been widely used to generate different virtual viewpoints of the same scene from the new perspective. However, DIBR tends to introduce annoying artifacts including blurring, discontinuity, blocking, and stretching, etc.. Thus, to improve DIBR performance, it is important to accurately measure the visual quality of synthesized views. In this paper, we propose a novel and effective no reference (NR) quality assessment method for 3D synthesized views by local variation and global change (LVGC). More specifically, we firstly compute the Gaussian derivatives for the input image to extract structure and chromatic features. Then, we use the local binary pattern (LBP) operator to encode the structure and chromatic feature maps, which are used to calculate quality-aware features to measure the local structural and chromatic distortion. Besides, we extract luminance features by global change to evaluate the naturalness of 3D synthesized views. With these extracted features, we utilize random forest regression (RFR) to train the quality prediction model from visual features to human ratings. Experimental results on three public benchmark databases demonstrate the effectiveness of our method on estimating visual quality of 3D synthesized views. Jiebin Yan, Yuming Fang 0001, Rengang Du, Yan Zeng 0001, Yifan Zuo 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Stereoscopic Image Stitching via Disparity-Constrained Warping and BlendingabstractAs a significant branch of virtual reality, stereoscopic image stitching aims to generating wide perspectives and natural-looking scenes. Existing 2D image stitching methods cannot be successfully applied to the stereoscopic images without considering the disparity consistency of stereoscopic images. To address this issue, this paper presents a stereoscopic image stitching method based on disparity-constrained warping and blending, which could avoid visual distortion and preserve disparity consistency. First, a point-line-driven homography based disparity minimization method is designed to pre-align the left and right images and reduce vertical disparity. Afterwards, a multi-constraint warping is proposed to further align the left and right images, where the initial disparity map is introduced to control the consistency of disparities. Finally, a disparity consistency seam-cutting and blending method is presented to determine the optimal seam and conduct stereoscopic image stitching. Experimental results demonstrate that the proposed method achieves competitive performance compared with other state-of-the-art methods. Xiaoting Fan, Jianjun Lei 0001, Yuming Fang 0001, Qingming Huang, Nam Ling, Chunping Hou |
IEEE Trans. Multim. | 3 |
| 2019 | Blind Image Quality Assessment by Learning from Multiple AnnotatorsabstractModels for image quality assessment (IQA) are generally optimized and tested by comparing to human ratings, which are expensive to obtain. Here, we develop a blind IQA (BIQA) model, and a method of training it without human ratings. We first generate a large number of corrupted image pairs, and use a set of existing IQA models to identify which image of each pair has higher quality. We then train a convolutional neural network to estimate perceived image quality along with the uncertainty, optimizing for consistency with the binary labels. The reliability of each IQA annotator is also estimated during training. Experiments demonstrate that our model outperforms state-of-the-art BIQA models in terms of correlation with human ratings in existing databases, as well in group maximum differentiation (gMAD) competition. Kede Ma, Xuelin Liu, Yuming Fang 0001, Eero P. Simoncelli |
ICIP | 3 |
| 2019 | Image Quality Assessment of Multi-exposure Image Fusion for Both Static and Dynamic ScenesabstractOver the past decade, many multi-exposure image fusion (MEF) methods have been proposed to obtain perceptually appealing results for both static and dynamic scenes. However, little work has been dedicated to evaluate perceptual quality of fused images. In this work, we propose a novel objective image quality assessment (IQA) model for MEF images of both static and dynamic scenes based on a pyramid subband contrast preservation scheme and an information theory adaptive pooling strategy. Firstly, we decompose the images using a Laplacian pyramid, and each pyramid subband is used to extract gradient and contrast features. Secondly, we binarize the structure inconsistency map between each exposure and fused image to obtain large-changed and small-changed regions. Finally, an information theory adaptive pooling strategy is used to combine these two quality scores from the individual regions. Experimental results on two large scale MEF databases of static and dynamic sequences show that the proposed model can obtain superior performance than state-of-the-art models designed for fused images. Yuming Fang 0001, Yan Zeng 0001, Hanwei Zhu, Guangtao Zhai |
ICME | 1 |
| 2019 | A Spatial-Temporal Weighted Method for Asymmetrically Distorted Stereo Video Quality AssessmentabstractWe propose a 2D-TO-3D video quality prediction model for assessing the perceptual quality of asymmetrically compressed stereoscopic 3D video. In the case that the distortions between left- and right-views are significantly different, directly averaging the qualities of single-view videos to estimate 3D video perceptual quality may lead a strong prediction bias. In order to eliminate or reduce the prediction bias, we design a two-stage approach for 3D video quality prediction. Firstly, we evaluate the perceptual quality of single-view videos with the state-of-the-art 2D image/video quality assessment approaches. Secondly, we design a binocular rivalry inspired model considering both spatial and temporal perceptual information to integrate 2D video perceptual quality of both views into the assessment of 3D video quality. We validate the highly competitive performance of the proposed approach on the Waterloo-IVC 3D video quality database. Yuming Fang 0001, Xiangjie Sui, Jiheng Wang |
ISCAS | 1 |
| 2019 | Blind Quality Assessment for DIBR-Synthesized Images Based on Chromatic and Disoccluded Information
Mengna Ding, Yuming Fang 0001, Yifan Zuo 0001, Zuowen Tan |
PRCV (2) | 2 |
| 2019 | Blind image quality assessment based on joint log-contrast statistics
Qiaohong Li, Weisi Lin, Ke Gu 0001, Yabin Zhang 0002, Yuming Fang 0001 |
Neurocomputing | 5 |
| 2019 | Residual dense network for intensity-guided depth map enhancement
Yifan Zuo 0001, Yuming Fang 0001, Yong Yang 0001, Xiwu Shang |
Inf. Sci. | 2 |
| 2019 | Reduced-reference quality assessment of image super-resolution by energy change and texture variation
Yuming Fang 0001, Jiaying Liu 0001, Yabin Zhang 0002, Weisi Lin, Zongming Guo |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Stereoscopic image quality assessment by deep convolutional neural network
Yuming Fang 0001, Jiebin Yan, Xuelin Liu, Jiheng Wang |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Video saliency detection by gestalt theory
Yuming Fang 0001, Xiaoqiang Zhang 0007, Feiniu Yuan, Nevrez Imamoglu, Haiwen Liu |
Pattern Recognit. | 1 |
| 2019 | Deep3DSaliency: Deep Stereoscopic Video Saliency Detection Model by 3D Convolutional NetworksabstractStereoscopic saliency detection plays an important role in various stereoscopic video processing applications. However, conventional stereoscopic video saliency detection methods mainly use independent low-level features instead of extracting them automatically, and thus, they ignore the intrinsic relationship between the spatial and temporal information. In this paper, we propose a novel stereoscopic video saliency detection method based on 3D convolutional neural networks, namely Deep 3D Video Saliency (Deep3DSaliency). The proposed network consists of two sub-models: Spatiotemporal Saliency Model (STSM), and Stereoscopic Saliency Aware Model (SSAM). STSM directly takes three consecutive video frames as the input to extract visual spatiotemporal features, while SSAM attempts to further infer the depth and semantic features from the left and right video frames by shared parameters from STSM. The visual spatiotemporal features from STSM, and the depth and semantic features from SSAM are learned by an alternating optimization scheme. Finally, all these saliency-related features are combined together for the final stereoscopic saliency detection via 3D deconvolution. Experimental results show the superior performance of the proposed model over other existing ones in saliency estimation for 3D video sequences. Yuming Fang 0001, Guanqun Ding, Jia Li 0003, Zhijun Fang 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Visual Attention Prediction for Stereoscopic Video by Multi-Module Fully Convolutional NetworkabstractVisual attention is an important mechanism in the human visual system (HVS) and there have been numerous saliency detection algorithms designed for 2D images/video recently. However, the research for fixation detection of stereoscopic video is still limited and challenging due to the complicated depth and motion information. In this paper, we design a novel multi-module fully convolutional network (MM-FCN) for fixation detection of stereoscopic video. Specifically, we design a fully convolutional network for spatial saliency prediction (S-FCN), where the initial spatial saliency map of stereoscopic video is learned by image database of object detection. Furthermore, the fully convolutional network for temporal saliency prediction (T-FCN) is constructed by combining saliency results from S-FCN and motion information from video frames. Finally, the fully convolutional network for depth fixation prediction (D-FCN) is designed to compute the final fixation map of stereoscopic video by learning depth features with spatiotemporal features from T-FCN. The experimental results show that the proposed MM-FCN can predict fixation results for stereoscopic video more effectively and efficiently than other related fixation prediction methods. Yuming Fang 0001, Chi Zhang 0027, Hanqin Huang, Jianjun Lei 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | No-Reference Quality Assessment for View Synthesis Using DoG-Based Edge Statistics and Texture NaturalnessabstractView synthesis is a key technique in free-viewpoint video, which renders virtual views based on texture and depth images. The distortions in synthesized views come from two stages, i.e., the stage of the acquisition and processing of texture and depth images, and the rendering stage using depth-image-based-rendering (DIBR) algorithms. The existing view synthesis quality metrics are designed for the distortions caused by a single stage, which cannot accurately evaluate the quality of the entire view synthesis process. With the considerations that the distortions introduced by two stages both cause edge degradation and texture unnaturalness, and the Difference-of-Gaussian (DoG) representation is powerful in capturing image edge and texture characteristics by simulating the center-surrounding receptive fields of retinal ganglion cells of human eyes, this paper presents a no-reference quality index for Synthesized views using DoG-based Edge statistics and Texture naturalness (SET). To mimic the multi-scale property of the Human Visual System (HVS), DoG images are first calculated at multiple scales. Then the orientation selective statistics features and the texture naturalness features are calculated on the DoG images and the coarsest scale image, producing two groups of quality-aware features. Finally, the quality model is learnt from these features using the random forest regression model. Experimental results on two view synthesis image databases demonstrate that the proposed metric is advantageous over the relevant state-of-the-arts in dealing with the distortions in the whole view synthesis process. Yu Zhou 0009, Leida Li, Shiqi Wang 0001, Jinjian Wu, Yuming Fang 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Multimodal Medical Image Fusion Based on Fuzzy Discrimination With Structural Patch DecompositionabstractMultimodal medical image fusion, emerging as a hot topic, aims to fuse images with complementary multi-source information. In this paper, we propose a novel multimodal medical image fusion method based on structural patch decomposition (SPD) and fuzzy logic technology. First, the SPD method is employed to extract two salient features for fusion discrimination. Next, two novel fusion decision maps called an incomplete fusion map and supplemental fusion map are constructed from salient features. In this step, the supplemental map is constructed by our defined two different fuzzy logic systems. The supplemental and incomplete maps are then combined to construct an initial fusion map. The final fusion map is obtained by processing the initial fusion map with a Gaussian filter. Finally, a weighted average approach is adopted to create the final fused image. Additionally, an effective color medical image fusion scheme that can effectively prevent color distortion and obtain superior diagnostic effects is also proposed to enhance fused images. Experimental results clearly demonstrate that the proposed method outperforms state-of-the-art methods in terms of subjective visual and quantitative evaluations. Yong Yang 0001, Jiahua Wu 0004, Shuying Huang, Yuming Fang 0001, Pan Lin, Yue Que 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2018 | Hyperspectral Image Dataset for Benchmarking on Salient Object DetectionabstractMany works have been done on salient object detection using supervised or unsupervised approaches on colour images. Recently, a few studies demonstrated that efficient salient object detection can also be implemented by using spectral features in visible spectrum of hyperspectral images from natural scenes. However, these models on hyperspectral salient object detection were tested with a very few number of data selected from various online public dataset, which are not specifically created for object detection purposes. Therefore, here, we aim to contribute to the field by releasing a hyperspectral salient object detection dataset with a collection of 60 hyperspectral images with their respective ground-truth binary images and representative rendered colour images (sRGB). We took several aspects in consideration during the data collection such as variation in object size, number of objects, foreground-background contrast, object position on the image, and etc. Then, we prepared ground truth binary images for each hyperspectral data, where salient objects are labelled on the images. Finally, we did performance evaluation using Area Under Curve (AUC) metric on some existing hyperspectral saliency detection models in literature. Nevrez Imamoglu, Yu Oishi, Xiaoqiang Zhang 0007, Guanqun Ding, Yuming Fang 0001, Toru Kouyama, Ryosuke Nakamura |
QoMEX | 5 |
| 2018 | Blind visual quality assessment for image super-resolution by convolutional neural network
Yuming Fang 0001, Chi Zhang 0027, Wenhan Yang, Jiaying Liu 0001, Zongming Guo |
Multim. Tools Appl. | 1 |
| 2018 | Image salient regions encryption for generating visually meaningful ciphertext image
Wenying Wen, Yushu Zhang 0001, Yuming Fang 0001, Zhijun Fang 0001 |
Neural Comput. Appl. | 3 |
| 2018 | A novel superpixel-based saliency detection model for 360-degree images
Yuming Fang 0001, Xiaoqiang Zhang 0007, Nevrez Imamoglu |
Signal Process. Image Commun. | 1 |
| 2018 | Optimal Region Selection for Stereoscopic Video Subtitle InsertionabstractStereoscopic subtitle insertion is a fundamental and essential element in stereoscopic film and TV industry. However, little work has been dedicated to the optimal region selection for stereoscopic subtitle insertion. In addition, there is no public database reported for the performance evaluation of it. In this paper, we build the first large-scale video database (TJU3D) for stereoscopic video subtitle insertion, which includes 50 video sequences with rich screen scenes. Compared with 2D subtitle region selection, there are several problems we have to consider in stereoscopic subtitle region selection: 1) the subtitle should avoid depth cue collision and occlusion from objects in stereoscopic video sequences; 2) the disparity value of the subtitle must be minimized to reduce visual discomfort; and 3) the temporal coherence constraint must be considered during region selection for subtitles in video sequences. By considering these constraints, we propose an optimal region selection algorithm for stereoscopic subtitle insertion. First, we compute the disparity map of each video frame in video sequences. For each frame, the optimal position and disparity value of the subtitle are determined by a subtitle region selection algorithm, which contains two parts (i.e., the coarse selection and fine selection). After that, by considering the temporal consistency between adjacent frames, the position and disparity value of each frame are further classified and processed in order to avoid the subtitle jitter. We evaluate the proposed method on TJU3D video database through two visual discomfort prediction metrics and one subjective experiment. To further verify the effectiveness of the proposed method, we also validate the performance of the proposed method on video comfort assessment database, i.e., IEEE-SA Stereo Database. Experimental results demonstrate that the visual discomfort is greatly reduced when using the proposed method compared with the basic method. Guanghui Yue 0001, Chunping Hou, Jianjun Lei 0001, Yuming Fang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Joint-Feature Guided Depth Map Super-Resolution With Face PriorsabstractIn this paper, we present a novel method to super-resolve and recover the facial depth map nicely. The key idea is to exploit the exemplar-based method to obtain the reliable face priors from high-quality facial depth map to improve the depth image. Specifically, a new neighbor embedding (NE) framework is designed for face prior learning and depth map reconstruction. First, face components are decomposed to form specialized dictionaries and then reconstructed, respectively. Joint features, i.e., low-level depth, intensity cues and high-level position cues, are put forward for robust patch similarity measurement. The NE results are used to obtain the face priors of facial structures and smooth maps, which are then combined in an uniform optimization framework to recover high-quality facial depth maps. Finally, an edge enhancement process is implemented to estimate the final high resolution depth map. Experimental results demonstrate the superiority of our method compared to state-of-the-art depth map super-resolution techniques on both synthetic data and real-world data from Kinect. Shuai Yang 0001, Jiaying Liu 0001, Yuming Fang 0001, Zongming Guo |
IEEE Trans. Cybern. | 3 |
| 2018 | No Reference Quality Assessment for Screen Content Images With Both Local and Global Feature RepresentationabstractIn this paper, we propose a novel no reference quality assessment method by incorporating statistical luminance and texture features (NRLT) for screen content images (SCIs) with both local and global feature representation. The proposed method is designed inspired by the perceptual property of the human visual system (HVS) that the HVS is sensitive to luminance change and texture information for image perception. In the proposed method, we first calculate the luminance map through the local normalization, which is further used to extract the statistical luminance features in global scope. Second, inspired by existing studies from neuroscience that high-order derivatives can capture image texture, we adopt four filters with different directions to compute gradient maps from the luminance map. These gradient maps are then used to extract the second-order derivatives by local binary pattern. We further extract the texture feature by the histogram of high-order derivatives in global scope. Finally, support vector regression is applied to train the mapping function from quality-aware features to subjective ratings. Experimental results on the public large-scale SCI database show that the proposed NRLT can achieve better performance in predicting the visual quality of SCIs than relevant existing methods, even including some full reference visual quality assessment methods. Yuming Fang 0001, Jiebin Yan, Leida Li, Jinjian Wu, Weisi Lin |
IEEE Trans. Image Process. | 1 |
| 2018 | Structure-Guided Image Inpainting Using Homography TransformationabstractIn this paper, we present a novel structure-guided framework for exemplar-based image inpainting to maintain the neighborhood consistence and structure coherence of an inpainted region. The proposed method consists of a data term for pixel validity and boundary continuity, a smoothness term to depict the compatibility of neighboring pixels for contextual continuity, and a coherence term to investigate image inherent regularities to ensure image self-similarity. To better reconstruct image structures, the method utilizes image regularity statistics to extract dominant linear structures of the target image. Guided by these structures, homography transformations are estimated and combined to globally repair the missing region using the Markov random field model. To reduce computational complexity, a hierarchical process is implemented to utilize the regularity effectively. The experimental results demonstrate that our method yields better results for various real-world scenes than existing state-of-the-art image inpainting techniques. Jiaying Liu 0001, Shuai Yang 0001, Yuming Fang 0001, Zongming Guo |
IEEE Trans. Multim. | 3 |
| 2017 | Perceptual quality assessment of HDR deghosting algorithmsabstractHigh dynamic range (HDR) imaging techniques aim to extend the dynamic range of images that cannot be well captured using conventional camera sensors. A common practice is to take a stack of pictures with different exposure levels and fuse them to produce a final image with more details. However, a small displacement between images caused by either camera or scene motion would void the benefits and cause the so-called ghosting artifacts. Over the past decade, many HDR deghosting algorithms have been proposed, but little work has been dedicated to evaluate HDR deghosting results either subjectively or objectively. In this work, we present a comprehensive subjective study for HDR deghosting. Specifically, we create a database that contains 20 dynamic image sequences and their corresponding deghosting results by 9 deghosting algorithms. A subjective user study is then carried out to evaluate the perceptual quality of deghosted images. The experimental results demonstrate the performance and limitations of existing HDR deghosting algorithm as well as no-reference image quality assessment models. In the future, we will make the database available to the public. Yuming Fang 0001, Hanwei Zhu, Kede Ma, Zhou Wang 0001 |
ICIP | 1 |
| 2017 | Saliency detection by forward and backward cues in deep-CNNabstractAs prior knowledge of objects or object features helps us make relations for similar objects on attentional tasks, pre-trained deep convolutional neural networks (CNNs) can be used to detect salient objects on images regardless of the object class is in the network knowledge or not. In this paper, we propose a top-down saliency model using CNN, a weakly supervised CNN model trained for 1000 object labelling task from RGB images. The model detects attentive regions based on their objectness scores predicted by selected features from CNNs. To estimate the salient objects effectively, we combine both forward and backward features, while demonstrating that partially-guided backpropagation will provide sufficient information for selecting the features from forward run of CNN model. Finally, these top-down cues are enhanced with a state-of-the-art bottom-up model as complementing the overall saliency. As the proposed model is an effective integration of forward and backward cues through objectness without any supervision or regression to ground truth data, it gives promising results compared to state-of-the-art models in two different datasets. Nevrez Imamoglu, Chi Zhang 0027, Wataru Shimoda, Yuming Fang 0001, Boxin Shi |
ICIP | 4 |
| 2017 | No reference quality assessment for stereoscopic images by statistical featuresabstractIn this paper, we propose a novel no reference (NR) quality assessment metric for stereoscopic images by statistical features. First, we calculate the luminance map through the local normalization, which is further used to extract the statistic luminance features. Second, we predict the disparity map of the stereoscopic image, which is further combined with the corresponding left and right views to extract the statistical structure and depth features for the stereoscopic image. The support vector regression (SVR) is employed as the mapping function from the quality-aware features to subjective quality scores. Experimental results on four publicly available large-scale stereoscopic image databases show that the proposed metric can obtain high-accuracy performance and is competitive with the state-of-the-art methods designed for visual quality prediction of stereoscopic images. Yuming Fang 0001, Jiebin Yan, Jiheng Wang |
QoMEX | 1 |
| 2017 | Learning visual saliency from human fixations for stereoscopic images
Yuming Fang 0001, Jianjun Lei 0001, Jia Li 0003, Long Xu 0001, Weisi Lin, Patrick Le Callet |
Neurocomputing | 1 |
| 2017 | BSD: Blind image quality assessment based on structural degradation
Qiaohong Li, Weisi Lin, Yuming Fang 0001 |
Neurocomputing | 3 |
| 2017 | Multi-Task Rank Learning for Image Quality AssessmentabstractIn practice, images are distorted by more than one distortion. For image quality assessment (IQA), existing machine learning (ML)-based methods generally establish a unified model for all the distortion types, or each model is trained independently for each distortion type, which is therefore distortion aware. In distortion-aware methods, the common features among different distortions are not exploited. In addition, there are fewer training samples for each model training task, which may result in overfitting. To address these problems, we propose a multi-task learning framework to train multiple IQA models together, where each model is for each distortion type; however, all the training samples are associated with each model training task. Thus, the common features among different distortion types and the said underlying relatedness among all the learning tasks are exploited, which would benefit the generalization ability of trained models and prevent overfitting possibly. In addition, pairwise image quality ranking instead of image quality rating is optimized in our learning task, which is fundamentally departed from traditional ML-based IQA methods toward better performance. The experimental results confirm that the proposed multi-task rank-learning-based IQA metric is prominent against all state-of-the-art nonreference IQA approaches. Long Xu 0001, Jia Li 0003, Weisi Lin, Yongbing Zhang 0002, Lin Ma 0002, Yuming Fang 0001, Yihua Yan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2017 | Stereoscopic Image Stitching Based on a Hybrid Warping ModelabstractTraditional image editing techniques cannot be directly used to process stereoscopic media, as extra constraints are required to ensure consistent changes between the left and right images. In this paper, we propose a hybrid warping model for stereoscopic image stitching by combining projective and content-preserving warping. First, a uniform homography algorithm is proposed to prewarp the left and right images, and thus ensure consistent changes. Second, a content-preserving warping is introduced to locally refine alignment and reduce vertical disparities. Finally, a seam-cutting-based algorithm is used to find a blending seam, and the multiband blending algorithm is used to produce the final stitched image. Experimental results show that the proposed method can effectively stitch stereoscopic images, which not only avoids local distortions, but also reduces vertical disparities reasonably. Weiqing Yan, Chunping Hou, Jianjun Lei 0001, Yuming Fang 0001, Zhouye Gu, Nam Ling |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Objective Quality Assessment of Screen Content Images by Uncertainty WeightingabstractIn this paper, we propose a novel full-reference objective quality assessment metric for screen content images (SCIs) by structure features and uncertainty weighting (SFUW). The input SCI is first divided into textual and pictorial regions. The visual quality of textual regions is estimated based on perceptual structural similarity, where the gradient information is adopted as the structural feature. To predict the visual quality of pictorial regions in SCIs, we extract the structural features and luminance features for similarity computation between the reference and distorted pictorial patches. To obtain the final visual quality of SCI, we design an uncertainty weighting method by perceptual theories to fuse the visual quality of textual and pictorial regions effectively. Experimental results show that the proposed SFUW can obtain better performance of visual quality prediction for SCIs than other existing ones. Yuming Fang 0001, Jiebin Yan, Jiaying Liu 0001, Shiqi Wang 0001, Qiaohong Li, Zongming Guo |
IEEE Trans. Image Process. | 1 |
| 2017 | Visual Attention Modeling for Stereoscopic Video: A Benchmark and Computational ModelabstractIn this paper, we investigate the visual attention modeling for stereoscopic video from the following two aspects. First, we build one large-scale eye tracking database as the benchmark of visual attention modeling for stereoscopic video. The database includes 47 video sequences and their corresponding eye fixation data. Second, we propose a novel computational model of visual attention for stereoscopic video based on Gestalt theory. In the proposed model, we extract the low-level features, including luminance, color, texture, and depth, from discrete cosine transform coefficients, which are used to calculate feature contrast for the spatial saliency computation. The temporal saliency is calculated by the motion contrast from the planar and depth motion features in the stereoscopic video sequences. The final saliency is estimated by fusing the spatial and temporal saliency with uncertainty weighting, which is estimated by the laws of proximity, continuity, and common fate in Gestalt theory. Experimental results show that the proposed method outperforms the state-of-the-art stereoscopic video saliency detection models on our built large-scale eye tracking database and one other database (DML-ITRACK-3D). Yuming Fang 0001, Chi Zhang 0027, Jing Li 0026, Jianjun Lei 0001, Matthieu Perreira Da Silva, Patrick Le Callet |
IEEE Trans. Image Process. | 1 |
| 2017 | No-Reference and Robust Image Sharpness Evaluation Based on Multiscale Spatial and Spectral FeaturesabstractThe human visual system exhibits multiscale characteristic when perceiving visual scenes. The hierarchical structures of an image are contained in its scale space representation, in which the image can be portrayed by a series of increasingly smoothed images. Inspired by this, this paper presents a no-reference and robust image sharpness evaluation (RISE) method by learning multiscale features extracted in both the spatial and spectral domains. For an image, the scale space is first built. Then sharpness-aware features are extracted in gradient domain and singular value decomposition domain, respectively. In order to take into account the impact of viewing distance on image quality, the input image is also down-sampled by several times, and the DCT-domain entropies are calculated as quality features. Finally, all features are utilized to learn a support vector regression model for sharpness prediction. Extensive experiments are conducted on four synthetically and two real blurred image databases. The experimental results demonstrate that the proposed RISE metric is superior to the relevant state-of-the-art methods for evaluating both synthetic and real blurring. Furthermore, the proposed metric is robust, which means that it has very good generalization ability. Leida Li, Wenhan Xia, Weisi Lin, Yuming Fang 0001, Shiqi Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | Optimized Multioperator Image Retargeting Based on Perceptual Similarity MeasureabstractWith various emerging mobile devices, the visual content have be to resized into different sizes or aspect ratios for good viewing experiences. In this paper, we propose a new multioperator retargeting algorithm by using four retargeting operators of seam carving, cropping, warping, and scaling iteratively. To determine which retargeting operator should be used at each iteration, we adopt structural similarity (SSIM) to evaluate the similarity between the original and retargeted images. The retargeting operator sequence is constructed based on the four types of retargeting operators by an optimization process. Since the sizes of original and retargeted images are different, scale-invariant feature transform flow is used for dense correspondence between the original and retargeted images for similarity evaluation. Additionally, visual saliency is used to weight SSIM results based on the characteristics of the human visual system. Experimental results on a public image retargeting database have shown the promising performance of the proposed multioperator retargeting algorithm. Yuming Fang 0001, Zhijun Fang 0001, Feiniu Yuan, Yong Yang 0001, Shouyuan Yang, Naixue Xiong |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2016 | A benchmark for robustness analysis of visual tracking algorithmsabstractIn this study, we investigate the robustness of existing visual tracking algorithms with quality-degraded video. A video database including the reference video sequences and their distorted versions is created as the benchmark for robustness analysis of visual tracking algorithms. Ten existing visual tracking algorithms are used to conduct the experiments for robustness analysis based on the benchmark. Our initial investigation demonstrates that all the existing visual tracking algorithms cannot obtain the robust visual tracking results for quality-degraded video sequences. The experimental results in this study show that there is still much room for the design of robust visual tracking algorithms. Yuming Fang 0001, Yuan Yuan 0029, Long Xu 0001, Weisi Lin |
ICASSP | 1 |
| 2016 | Aspect Ratio Similarity (ARS) for image retargeting quality assessmentabstractDuring the past few years, there have been various kinds of content-aware image retargeting methods proposed for image resizing. However, the lack of effective objective retargeting quality metric limits the further development of image retargeting. Different from the traditional image quality assessment, the quality degradation of the retargeted images is mainly caused by the geometric changes due to retargeting. In this paper, we propose a practical approach to reveal the geometric changes during image retargeting, and design an Aspect Ratio Similarity (ARS) metric to predict the visual quality of the retargeted image. The experimental results on the widely used dataset show that the proposed metric outperforms the state of the arts. Yabin Zhang 0002, Weisi Lin, Xinfeng Zhang 0001, Yuming Fang 0001, Leida Li |
ICASSP | 4 |
| 2016 | Quality assessment for image super-resolution based on energy change and texture variationabstractIn this paper, we propose a novel reduced-reference quality assessment metric for image super-resolution (RRIQA-SR) based on the low-resolution (LR) image information. First, we use the Markov Random Field (MRF) to model the pixel correspondence between LR and high-resolution (HR) images. Based on the pixel correspondence, we predict the perceptual similarity between image patches of LR and HR images by two components: the energy change and texture variation. The overall quality of HR images is estimated by the perceptual similarity between local image patches of LR and HR images. Experimental results demonstrate that the proposed method can obtain better performance of quality prediction for HR images than other existing ones, even including some full-reference (FR) metrics. Yuming Fang 0001, Jiaying Liu 0001, Yabin Zhang 0002, Weisi Lin, Zongming Guo |
ICIP | 1 |
| 2016 | Robust and automatic video colorization via multiframe reordering refinementabstractIn this paper, we propose a robust video colorization method automatically through limited color references in a video sequence. The proposed method first estimates motion vectors between a monochrome frame and colored reference frames for initial matching by optical flow. Then it transfers color information to matched points in the monochrome frame and further propagates color information of matched points to other parts of the monochrome frame. Furthermore, we design a multiframe reordering refinement to colorize video sequences robustly. Experimental results demonstrate that the proposed method achieves much better performance in video colorization than state-of-the-art methods. Sifeng Xia, Jiaying Liu 0001, Yuming Fang 0001, Wenhan Yang, Zongming Guo |
ICIP | 3 |
| 2016 | Quality assessment of 3D synthesized images via disoccluded region discoveryabstractDepth-Image-Based-Rendering (DIBR) is fundamental in free-viewpoint 3D video, which has been widely used to generate synthesized views from multi-view images. The majority of DIBR algorithms cause disoccluded regions, which are the areas invisible in original views but emerge in synthesized views. The quality of synthesized images is mainly contaminated by distortions in these disoccluded regions. Unfortunately, traditional image quality metrics are not effective for these synthesized images because they are sensitive to geometric distortions. To solve the problem, this paper proposes an objective quality evaluation method for 3D Synthesized images via Disoccluded Region Discovery (SDRD). A self-adaptive scale transform model is first adopted to preprocess the images on account of the impacts of view distance. Then disoccluded regions are detected by comparing the absolute difference between the preprocessed synthesized image and the warped image of preprocessed reference image. Furthermore, the disoccluded regions are weighted by a weighting function proposed to account for the varying sensitivities of human eyes to the size of disoccluded regions. Experiments conducted on IRCCyN/IVC DIBR image database demonstrate that the proposed SDRD method remarkably outperforms traditional 2D and existing DIBR-related quality metrics. Yu Zhou 0009, Leida Li, Ke Gu 0001, Yuming Fang 0001, Weisi Lin |
ICIP | 4 |
| 2016 | No-reference image quality assessment based on high order derivativesabstractResearch in human visual perception has found that the sense of natural scences cannot be conveyed only through lines and edges. It also needs the knowledge of texture regions within the image, which can be obtained through the analysis of higher derivatives. Inspired by the research from neuroscience that high order derivatives can capture the details of image structure, we propose a novel simple yet effective blind image quality assessment (IQA) metric based on high order derivatives (BHOD). In the proposed metric, we extract multi-scale structural features up to fourth order image derivatives, to obtain the image structural features. Support vector regression (SVR) is used to learn the mapping between feature space and subjective opinion scores. The proposed method is extensively evaluated on three image databases and shows highly competitive performance to state-of-the-art NR-IQA methods. Qiaohong Li, Weisi Lin, Yuming Fang 0001 |
ICME | 3 |
| 2016 | Blind quality assessment of compressed images via pseudo structural similarityabstractBlock-based compression causes severe pseudo structures. We find that the pseudo structures of images compressed by different levels show some degree of similarity. So we propose to evaluate the quality of compressed images via the similarity between pseudo structures of two images. To obtain a “reference” image, we introduce the most distorted image (MDI), which is derived from the distorted image and suffers from the highest degree of compression. The proposed pseudo structural similarity (PSS) model calculates the similarity between pseudo structures of the distorted image and MDI. Pseudo structures of the distorted image become similar to the MDI's under the condition of severe compression. Via comparative tests, the proposed PSS model, on one hand, is shown to be comparable to state-of-the-art competitors, and on the other hand, it is not only good at assessing natural scene images but also performs the best in the hotly-researched screen content image (SCI) database. It deserves to mention that PSS is able to boost the performance of mainstream general-purpose no-reference (NR) quality measures. Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Yuming Fang 0001, Xiaokang Yang 0001, Xiaolin Wu 0001, Jiantao Zhou 0001, Xianming Liu 0005 |
ICME | 4 |
| 2016 | No-reference Image Quality Assessment Based on Structural and Luminance Information
Qiaohong Li, Weisi Lin, Jingtao Xu, Yuming Fang 0001, Daniel Thalmann |
MMM (1) | 4 |
| 2016 | Perceptual evaluation of Compressive Sensing Image RecoveryabstractCompressive sensing (CS) has been attracting tremendous attention in recent years. Extensive CS recovery algorithms have been proposed for effective image reconstruction. However, little work has been dedicated to the perceptual evaluation of CS image recovery algorithms and the corresponding recovered images. In this paper, we first build a Compressive Sensing Recovered Image Database (CSRID), which contains images generated by ten popular CS image recovery algorithms at different sensing rates. We then carry out a subjective experiment using the single-stimulus method to obtain the subjective qualities of the images. The subjective scores are then used to evaluate the performances of the CS image recovery algorithms. Finally, the performances of general-purpose no-reference (NR) quality metrics and image blur metrics are investigated on the CSRID database. Experimental results show that the state-of-the-art quality metrics are very limited in predicting the quality of CS recovered images. Bo Hu 0008, Leida Li, Jiansheng Qian, Yuming Fang 0001 |
QoMEX | 4 |
| 2016 | No-reference image quality assessment based on local region statisticsabstractIn this paper, we propose an effective no-reference image quality assessment (IQA) method based on local region statistics (NRLRS). The proposed method is built on the hypothesis that image distortions may alter the local region statistics which can be well characterized by the inter-pixel relationship. Hence, by extracting perceptual features that describe the inter-pixel patterns of a distorted image, we can effectively quantify the impact of image degradation. For this purpose, the perceptual gray-level differences between neighboring pixels are extracted and a Gaussian Mixture Model (GMM) codebook is constructed as the generative model of extracted features. The Fisher vector representation is then derived to describe image as their derivations from the GMM model. Finally, partial least square regression is used to map the Fisher encodings to quality scores. Experimental results indicate that the proposed method achieves better performance in quality prediction as compared to relevant full-reference and no-reference IQA methods. Qiaohong Li, Weisi Lin, Yuming Fang 0001, Xinfeng Zhang 0001, Yabin Zhang 0002 |
VCIP | 3 |
| 2016 | A novel selective image encryption method based on saliency detectionabstractSalient regions usually carry important information in images. Existing feature encryption algorithms aim at extracting edge features as significant information rather than salient regions for encryption purpose. Moreover, most of them protect significant information by transforming the input image into texture-like or noise-like encrypted image which is obviously a visual sign of encrypted image, and thus can be easily attacked. In this paper, we propose a salient regions encryption scheme to generate visually meaningful ciphertext. First, salient regions are efficiently extracted by a saliency detection model in the compressed domain. Then we pre-encrypt these salient regions by a chaos-based encryption algorithm. With optical encryption theory, the pre-encrypted salient regions are finally transformed into a visually meaningful ciphertext. To the best of our knowledge, it is the first time to use salient regions as important visual information for encryption to obtain cipertext in images. The experimental results demonstrate that the salient regions can be largely hidden with the proposed method. Wenying Wen, Yushu Zhang 0001, Yuming Fang 0001, Zhijun Fang 0001 |
VCIP | 3 |
| 2016 | A general effective rate control system based on matching measurement and inter-quantizer
Zhijun Fang 0001, Yongbin Gao, Naixue Xiong, Athanasios V. Vasilakos, Yuming Fang 0001 |
Inf. Sci. | 5 |
| 2016 | Saliency-based stereoscopic image retargeting
Yuming Fang 0001, Junle Wang, Yuan Yuan 0029, Jianjun Lei 0001, Weisi Lin, Patrick Le Callet |
Inf. Sci. | 1 |
| 2016 | Orientation selectivity based visual pattern for reduced-reference image quality assessment
Jinjian Wu, Weisi Lin, Guangming Shi, Leida Li, Yuming Fang 0001 |
Inf. Sci. | 5 |
| 2016 | High-order local ternary patterns with locality preserving projection for smoke detection and image classification
Feiniu Yuan, Jinting Shi, Xue Xia 0005, Yuming Fang 0001, Zhijun Fang 0001, Tao Mei 0001 |
Inf. Sci. | 4 |
| 2016 | A Novel Spatial Pooling Strategy for Image Quality Assessment
Qiaohong Li, Yuming Fang 0001, Jingtao Xu |
J. Comput. Sci. Technol. | 2 |
| 2016 | Color image quality assessment based on sparse representation and reconstruction residual
Leida Li, Wenhan Xia, Yuming Fang 0001, Ke Gu 0001, Jinjian Wu, Weisi Lin, Jiansheng Qian |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Abnormal event detection in crowded scenes based on deep learning
Zhijun Fang 0001, Fengchang Fei, Yuming Fang 0001, Changhoon Lee, Naixue Xiong, Lei Shu 0001 |
Multim. Tools Appl. | 3 |
| 2016 | Perceptual quality evaluation for image defocus deblurring
Leida Li, Ya Yan, Yuming Fang 0001, Shiqi Wang 0001, Lu Tang 0001, Jiansheng Qian |
Signal Process. Image Commun. | 3 |
| 2016 | Visual structural degradation based reduced-reference image quality assessment
Jinjian Wu, Weisi Lin, Yuming Fang 0001, Leida Li, Guangming Shi, S. Issac Niwas |
Signal Process. Image Commun. | 3 |
| 2016 | No-Reference Quality Assessment for Multiply-Distorted Images in Gradient DomainabstractIn practice, images available to consumers usually undergo several stages of processing including acquisition, compression, transmission, and presentation, and each stage may introduce certain type of distortion. It is common that images are simultaneously distorted by multiple types of distortions. Most existing objective image quality assessment (IQA) methods have been designed to estimate perceived quality of images corrupted by a single image processing stage. In this letter, we propose a no-reference (NR) IQA method to predict the visual quality of multiply-distorted images based on structural degradation. In the proposed method, a novel structural feature is extracted as the gradient-weighted histogram of local binary pattern (LBP) calculated on the gradient map (GWH-GLBP), which is effective to describe the complex degradation pattern introduced by multiple distortions. Extensive experiments conducted on two public multiply-distorted image databases have demonstrated that the proposed GWH-GLBP metric compares favorably with existing full-reference and NR IQA methods in terms of high accordance with human subjective ratings. Qiaohong Li, Weisi Lin, Yuming Fang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2016 | Just Noticeable Difference Estimation for Screen Content ImagesabstractWe propose a novel just noticeable difference (JND) model for a screen content image (SCI). The distinct properties of the SCI result in different behaviors of the human visual system when viewing the textual content, which motivate us to employ a local parametric edge model with an adaptive representation of the edge profile in JND modeling. In particular, we decompose each edge profile into its luminance, contrast, and structure, and then evaluate the visibility threshold in different ways. The edge luminance adaptation, contrast masking, and structural distortion sensitivity are studied in subjective experiments, and the final JND model is established based on the edge profile reconstruction with tolerable variations. Extensive experiments are conducted to verify the proposed JND model, which confirm that it is accurate in predicting the JND profile, and outperforms the state-of-the-art schemes in terms of the distortion masking ability. Furthermore, we explore the applicability of the proposed JND model in the scenario of perceptually lossless SCI compression, and experimental results show that the proposed scheme can outperform the conventional JND guided compression schemes by providing better visual quality at the same coding bits. Shiqi Wang 0001, Lin Ma 0002, Yuming Fang 0001, Weisi Lin, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Backward Registration-Based Aspect Ratio Similarity for Image Retargeting Quality AssessmentabstractDuring the past few years, there have been various kinds of content-aware image retargeting operators proposed for image resizing. However, the lack of effective objective retargeting quality assessment metrics limits the further development of image retargeting techniques. Different from traditional image quality assessment (IQA) metrics, the quality degradation during image retargeting is caused by artificial retargeting modifications, and the difficulty for image retargeting quality assessment (IRQA) lies in the alternation of the image resolution and content, which makes it impossible to directly evaluate the quality degradation like traditional IQA. In this paper, we interpret the image retargeting in a unified framework of resampling grid generation and forward resampling. We show that the geometric change estimation is an efficient way to clarify the relationship between the images. We formulate the geometric change estimation as a backward registration problem with Markov random field and provide an effective solution. The geometric change aims to provide the evidence about how the original image is resized into the target image. Under the guidance of the geometric change, we develop a novel aspect ratio similarity (ARS) metric to evaluate the visual quality of retargeted images by exploiting the local block changes with a visual importance pooling strategy. Experimental results on the publicly available MIT RetargetMe and CUHK data sets demonstrate that the proposed ARS can predict more accurate visual quality of retargeted images compared with the state-of-the-art IRQA metrics. Yabin Zhang 0002, Yuming Fang 0001, Weisi Lin, Xinfeng Zhang 0001, Leida Li |
IEEE Trans. Image Process. | 2 |
| 2016 | A Universal Framework for Salient Object DetectionabstractIn this paper, we propose a novel universal framework for salient object detection, which aims to enhance the performance of any existing saliency detection method. First, rough salient regions are extracted from any existing saliency detection model with distance weighting, adaptive binarization, and morphological closing. With the superpixel segmentation, a Bayesian decision model is adopted to refine the rough saliency map to obtain a more accurate saliency map. An iterative optimization method is designed to obtain better saliency results by exploiting the characteristics of the output saliency map each time. Through the iterative optimization process, the rough saliency map is updated step by step with better and better performance until an optimal saliency map is obtained. Experimental results on the public salient object detection datasets with ground truth demonstrate the promising performance of the proposed universal framework subjectively and objectively. Jianjun Lei 0001, Bingren Wang, Yuming Fang 0001, Weisi Lin, Patrick Le Callet, Nam Ling, Chunping Hou |
IEEE Trans. Multim. | 3 |
| 2016 | Blind Image Quality Assessment Using Statistical Structural and Luminance FeaturesabstractBlind image quality assessment (BIQA) aims to develop quantitative measures to automatically and accurately estimate perceptual image quality without any prior information about the reference image. In this paper, we introduce a novel BIQA metric by structural and luminance information, based on the characteristics of human visual perception for distorted image. We extract the perceptual structural features of distorted image by the local binary pattern distribution. Besides, the distribution of normalized luminance magnitudes is extracted to represent the luminance changes in distorted image. After extracting the features for structures and luminance, support vector regression is adopted to model the complex nonlinear relationship from feature space to quality measure. The proposed BIQA model is called no-reference quality assessment using statistical structural and luminance features (NRSL). Extensive experiments conducted on four synthetically distorted image databases and three naturally distorted image databases have demonstrated that the proposed NRSL metric compares favorably with the relevant state-of-the-art BIQA models in terms of high correlation with human subjective ratings. The MATLAB source code and validation results of NRSL are publicly online at http://www.ntu.edu.sg/home/wslin/Publications.htm. Qiaohong Li, Weisi Lin, Jingtao Xu, Yuming Fang 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | Free-Energy Principle Inspired Video Quality Metric and Its Use in Video CodingabstractIn this paper, we extend the free-energy principle to video quality assessment (VQA) by incorporating with the recent psychophysical study on human visual speed perception (HVSP). A novel video quality metric, namely the free-energy principle inspired video quality metric (FePVQ), is therefore developed and applied to perceptual video coding optimization. The free-energy principle suggests that the human visual system (HVS) can actively predict “orderly” information and avoid “disorderly” information for image perception. Basically, “orderly” is associated with the skeletons and edges of objects, and “disorderly” mostly concerns textures in images. Based on this principle, an image is separated into orderly and disorderly regions, and processed differently in image quality assessment. For videos, visual attention, or fixation, is associated with the objects with significant motion according to HVSP, resulting in a motion strength factor in the FePVQ so that the free-energy principle is extended into spatio-temporal domain for VQA. In addition, we investigate the application of the FePVQ in perceptual rate distortion optimization (RDO). For this purpose, the FePVQ is realized with low computational cost by using the relative total variation model and the block-wise motion vectors of video coding to simulate the free-energy principle and the HVSP, respectively. The experimental results indicate that the proposed FePVQ is highly consistent with the HVS perception. The linear correlation coefficient and Spearman's rank-order correlation coefficient are up to 0.8324 and 0.8281 on the LIVE video database. Better perceptual quality of encoded video sequences is achieved by FePVQ-motivated RDO in video coding. Long Xu 0001, Weisi Lin, Lin Ma 0002, Yongbing Zhang 0002, Yuming Fang 0001, King Ngi Ngan, Songnan Li, Yihua Yan |
IEEE Trans. Multim. | 5 |
| 2015 | Multi-task rank learning for image quality assessmentabstractIn practice, multiple types of distortions are associated with an image quality degradation process. The existing machine learning (ML) based image quality assessment (IQA) approaches generally established a unified model for all distortion types, or each model is trained independently for each distortion type by using single-task learning, which lead to the poor generalization ability of the models as applied to practical image processing. There are often the underlying cross relatedness amongst these single-task learnings in IQA, which is ignored by the previous approaches. To solve this problem, we propose a multi-task learning framework to train IQA models simultaneously across individual tasks each of which concerns one distortion type. These relatedness can be therefore exploited to improve the generalization ability of IQA models from single-task learning. In addition, pairwise image quality rank instead of image quality rating is optimized in learning task. By mapping image quality rank to image quality rating, a novel no-reference (NR) IQA approach can be derived. The experimental results confirm that the proposed Multi-task Rank Learning based IQA (MRLIQ) approach is prominent among all state-of-the-art NR-IQA approaches. Long Xu 0001, Jia Li 0003, Weisi Lin, Yongbing Zhang 0002, Lin Ma 0002, Yuming Fang 0001, Yun Zhang 0002, Yihua Yan |
ICASSP | 6 |
| 2015 | Is pedestrian detection robust for surveillance?abstractIn surveillance systems, pedestrian detection is a fundamental task. To improve the detection accuracy, various approaches have been proposed to address severe occlusion, pose variation, etc. However, apart from the detection accuracy, a robust surveillance system also requires stable detection performance even when the video quality degrades due to the bandwidth limitation and environment variation. To study the robustness of detection algorithms, we introduce the Distorted Surveillance Video Database (DSurVD) which includes four types of common distortions in surveillance video; we benchmark several state-of-the-art pedestrian detection algorithms on this database; miss rate index (MRI) is proposed to evaluate the performance stability of the detectors on distorted videos. Performance-Quality curves of these algorithms regarding to different types of distortion are provided. We also provide discussion on how the quality affects the detection performance. Yuan Yuan 0029, Weisi Lin, Yuming Fang 0001 |
ICIP | 3 |
| 2015 | Gradient-weighted structural similarity for image quality assessmentsabstractThe goal of Image Quality Assessment (IQA) is to design computational models that can automatically predict the perceived image quality consistent with human subjective ratings. In this paper, we propose a full reference IQA metric gradient weighted structural similarity (GW-SSIM) by incorporating the gradient information to the well-known IQA metric SSIM. Experimental results demonstrate that GW-SSIM can greatly improve the quality prediction accuracy and achieve the best performance among the SSIM-based methods by addressing SSIM's shortcomings. Additionally, incorporating the proposed gradient weighting (GW) map into peak-signal-to-noise ratio (PSNR) also makes it quite competitive to state-of-the-art IQA models, and this is meaningful since PSNR is still a widely adopted metric. Qiaohong Li, Yuming Fang 0001, Weisi Lin, Daniel Thalmann |
ISCAS | 2 |
| 2015 | Moving Object Tracking with Structure Complexity Coefficients
Yuan Yuan 0029, Yuming Fang 0001, Weisi Lin |
MMM (1) | 2 |
| 2015 | Real-time image smoke detection using staircase searching-based dual threshold AdaBoost and dynamic analysisabstractIt is very challenging to accurately detect smoke from images because of large variances of smoke colour, textures, shapes and occlusions. To improve performance, the authors combine dual threshold AdaBoost with staircase searching technique to propose and implement an image smoke detection method. First, extended Haar‐like features and statistical features are efficiently extracted from integral images from both intensity and saturation components of RGB images. Then, a dual threshold AdaBoost algorithm with a staircase searching technique is proposed to classify the features of smoke for smoke detection. The staircase searching technique aims at keeping consistency of training and classifying as far as possible. Finally, dynamic analysis is proposed to further validate the existence of smoke. Experimental results demonstrate that the proposed system has a good robustness in terms of early smoke detection and low false alarm rate, and it can detect smoke from videos with size of 320 × 240 in real time. Feiniu Yuan, Zhijun Fang 0001, Shiqian Wu, Yong Yang 0001, Yuming Fang 0001 |
IET Image Process. | 5 |
| 2015 | Visual acuity inspired saliency detection by using sparse features
Yuming Fang 0001, Weisi Lin, Zhijun Fang 0001, Zhenzhong Chen 0001, Chia-Wen Lin, Chenwei Deng |
Inf. Sci. | 1 |
| 2015 | Subjective quality evaluation of compressed digital compound images
Huan Yang 0001, Yuming Fang 0001, Yuan Yuan 0029, Weisi Lin |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Exploiting entropy masking in perceptual graphic rendering
Lu Dong 0001, Yuming Fang 0001, Weisi Lin, Chenwei Deng, Ce Zhu, Seah Hock Soon |
Signal Process. Image Commun. | 2 |
| 2015 | No-Reference Quality Assessment of Contrast-Distorted Images Based on Natural Scene StatisticsabstractContrast distortion is often a determining factor in human perception of image quality, but little investigation has been dedicated to quality assessment of contrast-distorted images without assuming the availability of a perfect-quality reference image. In this letter, we propose a simple but effective method for no-reference quality assessment of contrast distorted images based on the principle of natural scene statistics (NSS). A large scale image database is employed to build NSS models based on moment and entropy features. The quality of a contrast-distorted image is then evaluated based on its unnaturalness characterized by the degree of deviation from the NSS models. Support vector regression (SVR) is employed to predict human mean opinion score (MOS) from multiple NSS features as the input. Experiments based on three publicly available databases demonstrate the promising performance of the proposed method. Yuming Fang 0001, Kede Ma, Zhou Wang 0001, Weisi Lin, Zhijun Fang 0001, Guangtao Zhai |
IEEE Signal Process. Lett. | 1 |
| 2015 | Perceptual Quality Assessment of Screen Content ImagesabstractResearch on screen content images (SCIs) becomes important as they are increasingly used in multi-device communication applications. In this paper, we present a study on perceptual quality assessment of distorted SCIs subjectively and objectively. We construct a large-scale screen image quality assessment database (SIQAD) consisting of 20 source and 980 distorted SCIs. In order to get the subjective quality scores and investigate, which part (text or picture) contributes more to the overall visual quality, the single stimulus methodology with 11 point numerical scale is employed to obtain three kinds of subjective scores corresponding to the entire, textual, and pictorial regions, respectively. According to the analysis of subjective data, we propose a weighting strategy to account for the correlation among these three kinds of subjective scores. Furthermore, we design an objective metric to measure the visual quality of distorted SCIs by considering the visual difference of textual and pictorial regions. The experimental results demonstrate that the proposed SCI perceptual quality assessment scheme, consisting of the objective metric and the weighting strategy, can achieve better performance than 11 state-of-the-art IQA methods. To the best of our knowledge, the SIQAD is the first large-scale database published for quality evaluation of SCIs, and this research is the first attempt to explore the perceptual quality assessment of distorted SCIs. Huan Yang 0001, Yuming Fang 0001, Weisi Lin |
IEEE Trans. Image Process. | 2 |
| 2015 | Perceptual Quality Assessment for 3D Triangle Mesh Based on CurvatureabstractTriangle meshes are used in representation of 3D geometric models, and they are subject to various visual distortions during geometrical processing and transmission. In this study, we propose a novel objective quality assessment method for 3D meshes based on curvature information; according to characteristics of the human visual system (HVS), two new components including visual masking and saturation effect are designed for the proposed method. Besides, inspired by the fact that the HVS is sensitive to structural information, we compute the structure distortion of 3D meshes. We test the performance of the proposed method on three publicly available databases of 3D mesh quality evaluation. We rotate among these databases for parameter determination to demonstrate the robustness of the proposed scheme. Experimental results demonstrate that the proposed method can predict consistent results in terms of correlation to the subjective scores across the databases. Lu Dong 0001, Yuming Fang 0001, Weisi Lin, Seah Hock Soon |
IEEE Trans. Multim. | 2 |
| 2015 | Depth Sensation Enhancement for Multiple Virtual View RenderingabstractDepth information is an indispensable element in depth image-based rendering (DIBR) for three-dimensional (3-D) display. In this paper, we propose a novel depth sensation enhancement method to address the problems in multiple virtual view rendering. First, as the depth sensation is decreased when rendering intermediate multiple virtual views, the basic principle of depth sensation enhancement is derived according to the number of rendering views. Second, with the increase of the scene complexity, it is difficult to ensure the depth sensation of all neighboring objects. The saliency analysis is adopted to give preferred guarantee to the depth sensation between the salient object and its neighbors. Then, the depth sensation enhancement for multiple virtual view rendering is performed based on a defined energy function built by the number of rendering views and the saliency analysis. Finally, considering the temporal consistency between adjacent frames, the depth sensation enhancement is extended to video applications with a newly designed energy function with energy term of temporal consistency preservation. Experimental results on a public database demonstrate that the proposed method can obtain promising performance in depth sensation. Jianjun Lei 0001, Cuicui Zhang, Yuming Fang 0001, Zhouye Gu, Nam Ling, Chunping Hou |
IEEE Trans. Multim. | 3 |
| 2015 | Visual Object Tracking by Structure Complexity CoefficientsabstractAppearance change of moving targets is a challenging problem in visual tracking. In this paper, we present a novel visual object tracking algorithm based on the observation dependent hidden Markov model (OD-HMM) framework. The observation dependency is computed by structure complexity coefficients (SCC) which is defined to predict the target appearance change. Unlike conventional methods addressing the appearance change problem by investigating different online appearance models, we handle this problem by addressing the fundamental reason of motion -related appearance change during visual tracking. Based on the analysis of motion-related appearance change, we investigate the relationship between the structure of the object surface and the appearance stability. The appearance of complex structural regions is easier to change compared with that of smooth structural regions with object moving. Based on this, we define SCC to predict the appearance stability of moving objects. Different from the standard HMM-based tracking algorithms where observations between different frames are assumed to be independent, we consider the observation dependency between consecutive frames with the information provided by SCC. Moreover , we present a novel outlier removing method in appearance model updating which helps to avoid error accumulation. Experimental results on challenging video sequences demonstrate that the proposed visual tracking algorithm with OD-HMM and SCC achieves better performance than existing related tracking algorithms. Yuan Yuan 0029, Huan Yang 0001, Yuming Fang 0001, Weisi Lin |
IEEE Trans. Multim. | 3 |
| 2014 | Stereoscopic image retargeting based on 3D saliency detectionabstractIn this paper, we propose a novel stereoscopic image retargeting algorithm based on 3D visual saliency detection. A new 3D visual attention model is designed based on 2D visual feature detection, depth feature detection and the modeling of various viewing bias in stereo vision. A geometrically consistent seam carving technique is adopted for retargeting stereo image pair. Experimental results demonstrated that both the proposed visual attention model and the proposed retargeting method outperform the state-of-the-art studies. Junle Wang, Yuming Fang 0001, Manish Narwaria, Weisi Lin, Patrick Le Callet |
ICASSP | 2 |
| 2014 | Rank learning on training set selection and image quality assessmentabstractMachine learning (ML) techniques are widely used in recent no-reference visual quality assessment (NR-VQA) metrics by training on subjective image quality databases. In these metrics, the optimization function is constructed based on L2norm of the distance between subjective image quality and predicted image quality. There are two problems in these L2norm based methods: (1) human's opinion on subjective image quality rating is not reliable at fine-scale level. A small difference between subjective image qualities represented by mean opinion scores (MOSs) of two images may not truly reflect the real quality difference between these two images, but acts as noise. The optimization process should avoid such noise. (2) Generally, human's opinion on pairwise comparison (PC) for image quality is more reliable and believable than MOS. The importance of PC is ignored during the optimization process of existing ML-based studies, which are designed based on the numerical rating system. In this paper, we introduce image quality ranking concept to establish a new optimization objective instead of L2norm optimization, and then a novel NR-VQA is constructed based on ranking learning. The proposed metric firstly suggests a reasonable training set for ML, which is ignored by existing ML-based NR-VQA. The ranking theory is adopted to build optimization function, which reflects the properties of PC over the numerical ranting system used by traditional NR-VQA. By ignoring the small difference between MOSs from two images during the optimization process, the proposed ranking-based NR-VQA can also well address the first problem from the existing related metrics. Experimental results show that the proposed ranking-based NR-VQA can obtain better performance over the state-of-the-art NR-VQA approaches. Long Xu 0001, Weisi Lin, Jia Li 0003, Xu Wang 0006, Yihua Yan, Yuming Fang 0001 |
ICME | 6 |
| 2014 | Saliency detection in computer rendered images based on object-level contrast
Lu Dong 0001, Weisi Lin, Yuming Fang 0001, Shiqian Wu, Seah Hock Soon |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | A Video Saliency Detection Model in Compressed DomainabstractSaliency detection is widely used to extract regions of interest in images for various image processing applications. Recently, many saliency detection models have been proposed for video in uncompressed (pixel) domain. However, video over Internet is always stored in compressed domains, such as MPEG2, H.264, and MPEG4 Visual. In this paper, we propose a novel video saliency detection model based on feature contrast in compressed domain. Four types of features including luminance, color, texture, and motion are extracted from the discrete cosine transform coefficients and motion vectors in video bitstream. The static saliency map of unpredicted frames (I frames) is calculated on the basis of luminance, color, and texture features, while the motion saliency map of predicted frames (P and B frames) is computed by motion feature. A new fusion method is designed to combine the static saliency and motion saliency maps to get the final saliency map for each video frame. Due to the directly derived features in compressed domain, the proposed model can predict the salient regions efficiently for video frames. Experimental results on a public database show superior performance of the proposed video saliency detection model in compressed domain. Yuming Fang 0001, Weisi Lin, Zhenzhong Chen 0001, Chia-Ming Tsai, Chia-Wen Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Visual Object Tracking Based on Backward Model ValidationabstractAppearance model updating is a challenging task in visual object tracking with occlusion and appearance variation. To avoid error accumulation in model updating, validation of updating is generally performed in tracking algorithms. These algorithms use the existing appearance model to validate incoming data. However, the existing appearance model may not be able to distinguish the valid training data (resulting from large appearance variation) from the invalid ones (resulting from occlusion), since both appearance variation and occlusion would lead to a good deal of appearance change of the estimated tracking result. The root of the problem is: the existing (outdated) model with information from frame 1 to n-1 may not be able to predict large appearance variations in frame n and, as a result, the appearance variations may be excluded from model updating. This defeats the purpose of model updating, which is to include new changes in appearance variations to the model, because the existing methods do not have the provision to include such changes in model updating by validating changes with the outdated model. To address this problem, we propose a backward model validation-based visual tracking (BVT) algorithm, which performs model updating first in frame n and then uses the information from the incoming frame (frame n + 1) to backward-check whether the updating is valid (occurrence of appearance variation) or invalid (occurrence of occlusion). In this way, the uncertainty of validating unpredictable features with the existing appearance models can be avoided. Moreover, an adaptive feature fusion method is designed to properly integrate the color-based feature with texture-based feature. The proposed feature extraction method provides a robust representation of the target with both rotation and shape deformation. Experimental results demonstrate that the proposed BVT algorithm outperforms the relevant existing algorithms on both publicly available and proprietary databases. Yuan Yuan 0029, Sabu Emmanuel, Yuming Fang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Saliency-Based Defect Detection in Industrial Images by Using Phase SpectrumabstractFor computer vision-based inspection of electronic chips or dies in semiconductor production lines, we propose a new method to effectively and efficiently detect defects in images. Different from the traditional methods that compare the image of each test chip or die with the template image one by one, which are sensitive to misalignment between the test and template images, a collection of multiple test images are used as the input image for processing simultaneously in our method with two steps. The first step is to obtain salient regions of the whole collection of test images, and the second step is to evaluate local discrepancy between salient regions in test images and the corresponding regions in the defect-free template image. To be more specific, in the first step of our method, phase-only Fourier transform (POFT), which is computationally efficient for online applications in industry, is used for saliency detection. We provide the theoretical justification for POFT to be effective to attenuate the normal regions and amplify the defects in multiple test images, which are usually arranged in a matrix format in industrial practice. By comparing with four other popular methods, the proposed algorithm can efficiently accommodate small variations (inevitable in practice) in test chips or dies, such as the spatial misalignments and product variations. Experimental results on a large-scale database including 1073 images, 94 of which are defective, show that our method performs much better than the other methods in terms of precision, recall, and F-measure. Xiaolong Bai, Yuming Fang 0001, Weisi Lin, Lipo Wang 0001, Bing-Feng Ju |
IEEE Trans. Ind. Informatics | 2 |
| 2014 | Video Saliency Incorporating Spatiotemporal Cues and Uncertainty WeightingabstractWe propose a novel algorithm to detect visual saliency from video signals by combining both spatial and temporal information and statistical uncertainty measures. The main novelty of the proposed method is twofold. First, separate spatial and temporal saliency maps are generated, where the computation of temporal saliency incorporates a recent psychological study of human visual speed perception. Second, the spatial and temporal saliency maps are merged into one using a spatiotemporally adaptive entropy-based uncertainty weighting approach. The spatial uncertainty weighing incorporates the characteristics of proximity and continuity of spatial saliency, while the temporal uncertainty weighting takes into account the variations of background motion and local contrast. Experimental results show that the proposed spatiotemporal uncertainty weighting algorithm significantly outperforms state-of-the-art video saliency detection models. Yuming Fang 0001, Zhou Wang 0001, Weisi Lin, Zhijun Fang 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Saliency Detection for Stereoscopic ImagesabstractMany saliency detection models for 2D images have been proposed for various multimedia processing applications during the past decades. Currently, the emerging applications of stereoscopic display require new saliency detection models for salient region extraction. Different from saliency detection for 2D images, the depth feature has to be taken into account in saliency detection for stereoscopic images. In this paper, we propose a novel stereoscopic saliency detection framework based on the feature contrast of color, luminance, texture, and depth. Four types of features, namely color, luminance, texture, and depth, are extracted from discrete cosine transform coefficients for feature contrast calculation. A Gaussian model of the spatial distance between image patches is adopted for consideration of local and global contrast calculation. Then, a new fusion method is designed to combine the feature maps to obtain the final saliency map for stereoscopic images. In addition, we adopt the center bias factor and human visual acuity, the important characteristics of the human visual system, to enhance the final saliency map for stereoscopic images. Experimental results on eye tracking databases show the superior performance of the proposed model over other existing methods. Yuming Fang 0001, Junle Wang, Manish Narwaria, Patrick Le Callet, Weisi Lin |
IEEE Trans. Image Process. | 1 |
| 2014 | Salient Region Detection by Fusing Bottom-Up and Top-Down Features Extracted From a Single ImageabstractRecently, some global contrast-based salient region detection models have been proposed based on only the low-level feature of color. It is necessary to consider both color and orientation features to overcome their limitations, and thus improve the performance of salient region detection for images with low-contrast in color and high-contrast in orientation. In addition, the existing fusion methods for different feature maps, like the simple averaging method and the selective method, are not effective sufficiently. To overcome these limitations of existing salient region detection models, we propose a novel salient region model based on the bottom-up and top-down mechanisms: the color contrast and orientation contrast are adopted to calculate the bottom-up feature maps, while the top-down cue of depth-from-focus from the same single image is used to guide the generation of final salient regions, since depth-from-focus reflects the photographer's preference and knowledge of the task. A more general and effective fusion method is designed to combine the bottom-up feature maps. According to the degree-of-scattering and eccentricities of feature maps, the proposed fusion method can assign adaptive weights to different feature maps to reflect the confidence level of each feature map. The depth-from-focus of the image as a significant top-down feature for visual attention in the image is used to guide the salient regions during the fusion process; with its aid, the proposed fusion method can filter out the background and highlight salient regions for the image. Experimental results show that the proposed model outperforms the state-of-the-art models on three public available data sets. Huawei Tian, Yuming Fang 0001, Yao Zhao 0001, Weisi Lin, Zhenfeng Zhu |
IEEE Trans. Image Process. | 2 |
| 2013 | 2D mel-cepstrum based saliency detectionabstractStudies on the Human Visual System (HVS) have demonstrated that human eyes are more attentive to spatial or spectral components with irregularities on the scene. This fact was modeled differently in many computational methods such as saliency residual (SR) approach, which tries to find the irregularity of frequency components by subtracting average filtered and original amplitude spectra. However, studies showed that the high frequency components have more effect on the HVS perception. In this paper, we propose a 2D mel-cepstrum based spectral residual saliency detection model (MCSR) to provide perceptually meaningful and more informative saliency map with less redundancy and without down-sampling as in saliency residual approach. Experimental results demonstrate that proposed MCSR model can yield promising results compared to the relevant state of the art models. Nevrez Imamoglu, Yuming Fang 0001, Wenwei Yu, Weisi Lin |
ICIP | 2 |
| 2013 | Video saliency incorporating spatiotemporal cues and uncertainty weightingabstractWe propose a method to detect visual saliency from video signals by combing both spatial and temporal information and statistical uncertainty measures. The main novelty of the proposed method is twofold. First, separate spatial and temporal saliency maps are generated, where the computation of temporal saliency incorporates a recent psychological study of human visual speed perception, where the perceptual prior probability distribution of the speed of motion is measured through a series of psychovisual experiments. Second, the spatial and temporal saliency maps are merged into one using a spatiotemporally adaptive entropy-based uncertainty weighting approach. Experimental results show that the proposed method significantly outperforms state-of-the-art video saliency detection models. Yuming Fang 0001, Zhou Wang 0001, Weisi Lin |
ICME | 1 |
| 2013 | A saliency detection model based on sparse features and visual acuityabstractIn this paper, we propose a novel computational model of visual attention based on the relevant characteristics of the Human Visual System (HVS). The input image is firstly divided into small image patches. Then the sparse features for each image patch are extracted based on the learned sparse coding basis. The human visual acuity is adopted in the calculation of the center-surround feature differences for saliency detection. In addition, the neighboring image patches for computing the saliency value of each center image patch are selected based on the characteristics of HVS. Experimental results show that the proposed saliency detection algorithm outperforms other existing schemes tested with a large public image database. Yuming Fang 0001, Weisi Lin, Zhenzhong Chen 0001, Chia-Wen Lin, Zhijun Fang 0001, Chenwei Deng |
ISCAS | 1 |
| 2013 | Detection of salient objects in computer synthesized images based on object-level contrastabstractIn this work, we propose a method to detect visually salient objects in computer synthesized images from 3D meshes. Different from existing detection methods on graphic saliency which compute saliency based on pixel-level contrast, the proposed method computes saliency by measuring object-level contrast of each object to the other objects in a rendered image. Given a synthesized image, the proposed method first extracts dominant colors from each object, and represents each object with the dominant color descriptor (DCD). Saliency is measured as the contrast between the DCD of the object and the DCDs of its surrounding objects. We evaluate the proposed method on a data set of computer rendered images, and the results show that the proposed method obtains much better performance compared with existing related methods. Lu Dong 0001, Weisi Lin, Yuming Fang 0001, Shiqian Wu, Seah Hock Soon |
VCIP | 3 |
| 2013 | Saliency detection for stereoscopic imagesabstractSaliency detection techniques have been widely used in various 2D multimedia processing applications. Currently, the emerging applications of stereoscopic display require new saliency detection models for stereoscopic images. Different from saliency detection for 2D images, depth features have to be taken into account in saliency detection for stereoscopic images. In this paper, we propose a new stereoscopic saliency detection framework based on the feature contrast of color, intensity, texture, and depth. Four types of features including color, luminance, texture, and depth are extracted from DC-T coefficients to represent the energy for image patches. A Gaussian model of the spatial distance between image patches is adopted for the consideration of local and global contrast calculation. A new fusion method is designed to combine the feature maps for computing the final saliency map for stereoscopic images. Experimental results on a recent eye tracking database show the superior performance of the proposed method over other existing ones in saliency estimation for 3D images. Yuming Fang 0001, Junle Wang, Manish Narwaria, Patrick Le Callet, Weisi Lin |
VCIP | 1 |
| 2013 | Objective quality assessment for image retargeting based on perceptual distortion and information lossabstractImage retargeting techniques aim to obtain retargeted images with different sizes or aspect ratios for various display screens. Various content-aware image retargeting algorithms have been proposed recently. However, there is still no accurate objective metric for visual quality assessment of retargeted images. In this paper, we propose a novel objective metric for assessing visual quality of retargeted images based on perceptual geometric distortion and information loss. The proposed metric measures the geometric distortion of retargeted images by SIFT flow variation. Furthermore, a visual saliency map is derived to characterize human perception of the geometric distortion. On the other hand, the information loss in a retargeted image, which is calculated based on the saliency map, is integrated into the proposed metric. A user study is conducted to evaluate the performance of the proposed metric. Experimental results show the consistency between the objective assessments from the proposed metric and subjective assessments. Chih-Chung Hsu, Chia-Wen Lin, Yuming Fang 0001, Weisi Lin |
VCIP | 3 |
| 2013 | A Saliency Detection Model Using Low-Level Features Based on Wavelet TransformabstractResearchers have been taking advantage of visual attention in various image processing applications such as image retargeting, video coding, etc. Recently, many saliency detection algorithms have been proposed by extracting features in spatial or transform domains. In this paper, a novel saliency detection model is introduced by utilizing low-level features obtained from the wavelet transform domain. Firstly, wavelet transform is employed to create the multi-scale feature maps which can represent different features from edge to texture. Then, we propose a computational model for the saliency map from these features. The proposed model aims to modulate local contrast at a location with its global saliency computed based on the likelihood of the features, and the proposed model considers local center-surround differences and global contrast in the final saliency map. Experimental evaluation depicts the promising results from the proposed model by outperforming the relevant state of the art saliency detection models. Nevrez Imamoglu, Weisi Lin, Yuming Fang 0001 |
IEEE Trans. Multim. | 3 |
| 2012 | Video saliency detection in the compressed domainabstractSaliency detection is widely used to extract the regions of interest in images. Many saliency detection models have been proposed for videos in the uncompressed domain. However, videos are always stored in the compressed domain such as MPEG2, H.264, MPEG4 Visual, etc. In this study, we propose a video saliency detection model based on feature contrast in the compressed domain. Four features of luminance, color, texture and motion are extracted from DCT coefficients and motion vectors in the video bitstream. The static saliency map of video frames is calculated based on the luminance, color and texture features, while the motion saliency map for video frames is computed by motion feature. The final saliency map for video frames is obtained through combining the static saliency map and motion saliency map. Experimental results show good performance of the proposed video saliency detection model in the compressed domain. Yuming Fang 0001, Weisi Lin, Zhenzhong Chen 0001, Chia-Ming Tsai, Chia-Wen Lin |
ACM Multimedia | 1 |
| 2012 | Saliency Detection in the Compressed Domain for Adaptive Image RetargetingabstractSaliency detection plays important roles in many image processing applications, such as regions of interest extraction and image resizing. Existing saliency detection models are built in the uncompressed domain. Since most images over Internet are typically stored in the compressed domain such as joint photographic experts group (JPEG), we propose a novel saliency detection model in the compressed domain in this paper. The intensity, color, and texture features of the image are extracted from discrete cosine transform (DCT) coefficients in the JPEG bit-stream. Saliency value of each DCT block is obtained based on the Hausdorff distance calculation and feature map fusion. Based on the proposed saliency detection model, we further design an adaptive image retargeting algorithm in the compressed domain. The proposed image retargeting algorithm utilizes multioperator operation comprised of the block-based seam carving and the image scaling to resize images. A new definition of texture homogeneity is given to determine the amount of removal block-based seams. Thanks to the directly derived accurate saliency information from the compressed domain, the proposed image retargeting algorithm effectively preserves the visually important regions for images, efficiently removes the less crucial regions, and therefore significantly outperforms the relevant state-of-the-art algorithms, as demonstrated with the in-depth analysis in the extensive experiments. Yuming Fang 0001, Zhenzhong Chen 0001, Weisi Lin, Chia-Wen Lin |
IEEE Trans. Image Process. | 1 |
| 2012 | Bottom-Up Saliency Detection Model Based on Human Visual Sensitivity and Amplitude SpectrumabstractWith the wide applications of saliency information in visual signal processing, many saliency detection methods have been proposed. However, some key characteristics of the human visual system (HVS) are still neglected in building these saliency detection models. In this paper, we propose a new saliency detection model based on the human visual sensitivity and the amplitude spectrum of quaternion Fourier transform (QFT). We use the amplitude spectrum of QFT to represent the color, intensity, and orientation distributions for image patches. The saliency value for each image patch is calculated by not only the differences between the QFT amplitude spectrum of this patch and other patches in the whole image, but also the visual impacts for these differences determined by the human visual sensitivity. The experiment results show that the proposed saliency detection model outperforms the state-of-the-art detection models. In addition, we apply our proposed model in the application of image retargeting and achieve better performance over the conventional algorithms. Yuming Fang 0001, Weisi Lin, Bu-Sung Lee, Chiew Tong Lau, Zhenzhong Chen 0001, Chia-Wen Lin |
IEEE Trans. Multim. | 1 |
| 2011 | A visual attention model combining top-down and bottom-up mechanisms for salient object detectionabstractSelective attention in the human visual system is performed as the way that humans focus on the most important parts when observing a visual scene. Many bottom-up computational models of visual attention have been devised to get the saliency map for an image, which are data-driven or task-independent. However, studies show that the task-driven or top-down mechanism also plays an important role during the formation of visual attention, especially with the cases of object detection and location. In this paper, we proposed a new computational visual attention model by combining bottom-up and top-down mechanisms for man-made object detection in scenes. This model shows that the statistical characteristics of orientation features can be used as top-down clues to help for determining the location for salient objects in natural scenes. Experiments confirm the effectiveness of this visual attention model. Yuming Fang 0001, Weisi Lin, Chiew Tong Lau, Bu-Sung Lee |
ICASSP | 1 |
| 2011 | Saliency-based image retargeting in the compressed domainabstractIn this paper, we propose a novel image retargeting algorithm to resize images based on the extracted saliency information from the compressed domain. Firstly, we utilize DCT coefficients in JPEG bitstream to perform saliency detection with the consideration of the human visual sensitivity. The obtained saliency information is used to determine the relative visual importance of each 8 x 8 block for the image. Furthermore, we propose a new adaptive block-level seam removal operation for connected blocks to resize the image. Thanks to the directly derived saliency information from the compressed domain, the proposed image retargeting algorithm effectively preserves the objects of attention, efficiently removes the less crucial regions, and therefore significantly outperforms the relevant state-of-the-art algorithms, as demonstrated with the careful analysis and in the extensive experiments. Yuming Fang 0001, Zhenzhong Chen 0001, Weisi Lin, Chia-Wen Lin |
ACM Multimedia | 1 |
| 2011 | Bottom-Up Saliency Detection Model Based on Amplitude Spectrum
Yuming Fang 0001, Weisi Lin, Bu-Sung Lee, Chiew Tong Lau, Chia-Wen Lin |
MMM (1) | 1 |
| 2011 | Image retargeting based on the sensitivity-tuned visual significance mapabstractIn this paper, we propose a novel image retargeting algorithm based on the sensitivity-tuned visual significance map which is composed of a saliency map and a gradient map. We develop a new saliency detection model based on the human visual sensitivity and amplitude spectrum of image patches. We use a coherent normalization based fusion method to combine the saliency map and the gradient map to generate the visual significance map. The seam carving technique is adopted for image retargeting, based on the sensitivity-tuned visual significance map. Experiment results show that the proposed algorithm outperforms the relevant state-of-the-arts image retargeting algorithms significantly. Yuming Fang 0001, Zhenzhong Chen 0001, Weisi Lin, Chia-Wen Lin, Chia-Ming Tsai |
VCIP | 1 |