Yuzhen Niu

dblp:28/7442 · DBLP profile ↗
← Back
92ranked-venue papers
34as first author
46since 2021 · last 2026
0000-0002-9874-9719ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 62 · 27 first-author · 34 since 2021Artificial intelligence and machine learning · 27 · 8 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2Computer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Flare detection and detail compensation for nighttime flare removal
Yuzhen Niu, Yuezhou Li, Jingyuan Zheng
Eng. Appl. Artif. Intell.1
2026 Unpaired underwater image enhancement via Cross-Modal Dual Contrastive Learning
Yuzhen Niu, Shiyun Yang, Yaojie Huang
J. Vis. Commun. Image Represent.1
2026 Integrating perceptual cues with mixture-of-experts for low-light image restoration
Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Rui Xu 0028, Hui Da, Wenxi Liu, Lifang Wei
Neural Networks2
2026 ProMoT: Progressive Prompting of Modality and Temporal Dynamics for RGB-T Tracking
abstract
RGB-T tracking benefits from the complementary nature of RGB and TIR modalities, yet their relative reliability for target localization often shifts over time. Most existing trackers fail to adapt to such modality and temporal dynamics in a unified and effective manner, resulting in target representations that are neither discriminative nor temporally consistent. In this paper, we propose ProMoT, a novel tracking framework that jointly integrates cross-modal and temporal cues into a progressive prompting process, enabling continuous retrieval of target-aware representations. Specifically, we design an adaptive target query generator (QueryGen), which selectively aggregates informative spatio-temporal cues from diverse ghost representations through the dynamic sparse ghost fusion mechanism, thereby enabling the generation of target-aware queries. To further preserve fine-grained, temporally consistent target cues, we introduce a high-order contextual prompt updater (PromptUpdater), which encodes high-order cross-modal representations from current and previous frames. These prompts establish the compact and discriminative inter-frame context to not only refine the current frame’s features but also guide target localization in future frames. All components are built upon a parameter-shared backbone for RGB and TIR inputs, forming our complete ProMoT framework. Extensive experiments on both complete and missing modality RGB-T tracking benchmarks show that ProMoT consistently achieves state-of-the-art performance while balancing efficiency.
Rui Xu 0028, Si Chen 0002, Yuzhen Niu, Yan Yan 0001, Dahan Wang
IEEE Trans. Circuits Syst. Video Technol.4
2026 Multimodal Image Representation Learning With Limited Visual-Tactile Data
abstract
Previous multimodal visual-tactile image representation learning (VTL) methods have achieved significant success in object understanding through large-scale training data. However, obtaining sufficient training data is often infeasible, and the above methods struggle to effectively focus on discriminative visual and tactile features with limited data, resulting in degraded performance. To solve the above issue, we introduce a new task called visual-tactile image representation learning with limited data (VTL-L), which better facilitates real-world applications. To address the challenges of limited data and modality discrepancy in the VTL-L task, we propose a novel multi-order feature enhancement-based, alignment-free fusion network (MOA-Net). First, we introduce a multi-order feature enhancement (MFE) module to hierarchically strengthen the detailed and structural representation by aggregating the low- and high-order topological information. This approach can effectively reduce the attention noise and obtain discriminative features with limited data. Then, we propose the alignment-free visual-tactile fusion (AVTF) module to achieve representative spatial and channel features and perform the cross-modality fusion without alignment, which efficiently mitigates the modality discrepancy. Finally, we develop a dual counterfactual intervention (DCI) loss to jointly optimize fused visual-tactile feature and probability distributions, thereby improving the performance of the MOA-Net in the VTL-L task. Extensive experiments demonstrate the superiority of the proposed method across three types of tasks on four datasets under diverse limited-data settings (source code available at: https://github.com/liuxiangqiu007/MOA-Net).
Liuxiang Qiu, Hui Da, Wenxi Liu, Yuzhen Niu, Hanli Wang, Tiesong Zhao
IEEE Trans. Image Process.4
2026 Illumination-Guided Grouped Attention and Masked Progressive Denoising for Low-Light Image Enhancement
abstract
Existing Transformer- or Mamba-based low-light image enhancement (LLIE) methods can capture long-range dependencies in images and achieve global degradation restoration, but suffer from insufficient detail recovery. Besides, these methods do not consider the different illumination levels in various regions or enhance the regional illumination accordingly. Furthermore, existing denoising methods often overfit to single noise distribution and type, showing poor generalization to the noises in low-light images. To address these issues, we propose an Illumination-guided Grouped Attention (IGA) module and a Masked Progressive Denoising (MPD) module for low-light image enhancement. The IGA module first improves the local feature representation through detail recovery operations, and then iteratively groups the image representation based on the illumination levels and conducts grouped attention to achieve both fine detail recovery and adaptive regional illumination optimization. The MPD module explores correlations within and among regions to enhance the texture and structure representations from a local to global perspective, thus achieving both local and global denoising and well denoising generalization capability. Extensive quantitative and qualitative experiments on eight datasets demonstrate that our proposed method outperforms the existing state-of-the-art low-light image enhancement methods. Furthermore, plugging our proposed IGA and MPD modules into existing Transformer- or Mamba-based LLIE methods can significantly improve their performance, further demonstrating the modules' ability to address the common issues in such methods.
Hui Da, Yuzhen Niu, Liuxiang Qiu, Tiesong Zhao, Yuzhong Chen 0001
IEEE Trans. Multim.2
2025 URWKV: Unified RWKV Model with Multi-state Perspective for Low-light Image Restoration
abstract
Existing low-light image enhancement (LLIE) and joint LLIE and deblurring (LLIE-deblur) models have made strides in addressing predefined degradations, yet they are often constrained by dynamically coupled degradations. To address these challenges, we introduce a Unified Receptance Weighted Key Value (URWKV) model with multi-state perspective, enabling flexible and effective degradation restoration for low-light images. Specifically, we customize the core URWKV block to perceive and analyze complex degradations by leveraging multiple intra- and inter-stage states. First, inspired by the pupil mechanism in the human visual system, we propose Luminance-adaptive Normalization (LAN) that adjusts normalization parameters based on rich inter-stage states, allowing for adaptive, scene-aware luminance modulation. Second, we aggregate multiple intra-stage states through exponential moving average approach, effectively capturing subtle variations while mitigating information loss inherent in the single-state mechanism. To reduce the degradation effects commonly associated with conventional skip connections, we propose the State-aware Selective Fusion (SSF) module, which dynamically aligns and integrates multi-state features across encoder stages, selectively fusing contextual information. In comparison to state-of-the-art models, our URWKV model achieves superior performance on various benchmarks, while requiring significantly fewer parameters and computational resources. Code is available at: https://github.com/FZU-N/URWKV.
Rui Xu 0028, Yuzhen Niu, Yuezhou Li, Huangbiao Xu, Wenxi Liu, Yuzhong Chen 0001
CVPR2
2025 Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition (PAR) seeks to predict multiple semantic attributes associated with a specific pedestrian. There are two types of approaches for PAR: unimodal framework and bimodal framework. The former one is to seek a robust visual feature. However, the lack of exploiting semantic feature of linguistic modality is the main concern. The latter one utilizes prompt learning techniques to integrate linguistic data. However, static prompt templates and simple bimodal concatenation cannot to capture the extensive intra-class attribute variability and support active modalities collaboration. In this paper, we propose an Enhanced Visual-Semantic Interaction with Tailored Prompts (EVSITP) framework for PAR. We present an Image-Conditional Dual-Prompt Initialization Module (IDIM) to adaptively generate context-sensitive prompts from visual inputs. Subsequently, a Prompt Enhanced and Regularization Module (PERM) is proposed to strengthen linguistic information from IDIM. We further design a Bimodal Mutual Interaction Module (BMIM) to ensure bidirectional modalities communication. In addition, existing PAR datasets are collected over a short period in limited scenarios, which do not align with real-world scenarios. Therefore, we annotate a long-term person re-identification dataset to create a new PAR dataset, Celeb-PAR. Experiments on several challenging PAR datasets show that our method outperforms state-of-the-art approaches.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001
CVPR4
2025 Projection, Interaction and Fusion: A Progressive Difference Fusion Network for Salient Object Detection
abstract
In recent years, deep learning-based Salient Object Detection (SOD) methods have made tremendous progress; however, their performance in complex scenarios has reached a bottleneck. In this paper, we propose a novel Progressive Difference Fusion Network (PDFNet) based on fine-grained feature fusion. First, to address the scale variability of salient objects, we introduce a Self-Guided Module (SGM) with dynamic receptive fields. Second, to tackle the shape variability of salient objects, we design a Feature Aggregation Module (FAM) incorporating cross convolutions and a feedback loop. Finally, to alleviate the issue of confusion between global and detail information during multi-scale feature fusion in existing models, we develop a Progressive Difference Fusion Unit (PDFU) to project multi-scale features into fine-grained nodes and enhance them through node interaction based on difference features. Additionally, we propose a Conditional Random Field Based on Patch (CRFbp), which focuses on handling discrete points, further improving the model’s performance. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) performance on five benchmark datasets. Code is available at: https://github.com/pdfnet2025/PDFNet.git.
Xiao Ke, Yuzhen Niu
IJCAI3
2025 IPCMoE: Integrating Perceptual Cues with Mixture-of-Experts for Joint Low-Light Image Enhancement and Deblurring
abstract
Visual perception of nighttime images is often compromised by co-existing low-light and blur degradations. While recent methods have made progress in jointly solving these degradations, the diversity of patterns and intensities in degradation has not been properly considered, leading to inconsistent illumination and unintended artifacts. In response, we propose to integrate perceptual cues with mixture-of-experts (IPCMoE) to achieve flexible processing for low-light blurry images. By exploiting the perceptual cues, we strategically combine dedicated experts with the selective collaboration approach for feature enlightening and texture restoration. To this end, we develop perceptual-integrated MoEs by designing customized routers and task-depended experts. Specifically, the texture memorial MoE is developed to preserve valuable features to restore high-fidelity details, and the enhancement MoE that adaptively integrates enlightening cues and texture cues is designed to formulate the relationship between feature enlightening and texture restoration, thereby achieving dynamic image processing. Extensive experiments show that our method achieves state-of-the-art performance on LOL-Blur and Real-LOL-Blur datasets.
Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Hui Da, Rui Xu 0028, Wenxi Liu
ACM Multimedia2
2025 CoFiVLA: Synergistic Coarse-Fine Vision-Language Alignment for Image Aesthetic Assessment
Yuzhen Niu, Siling Chen 0002, Yuzhong Chen 0001, Rui Xu 0028, Hui Da
ACM Multimedia1
2025 Parallax-aware dual-view feature enhancement and adaptive detail compensation for dual-pixel defocus deblurring
Yuzhen Niu, Rui Xu 0028, Yuezhou Li, Yuzhong Chen 0001
Eng. Appl. Artif. Intell.1
2025 Rethinking attention mechanism for enhanced pedestrian attribute recognition
abstract
Pedestrian Attribute Recognition (PAR) plays a crucial role in various computer vision applications, demanding precise and reliable identification of attributes from pedestrian images. Traditional PAR methods, though effective in leveraging attention mechanisms, often suffer from the lack of direct supervision on attention, leading to potential overfitting and misallocation. This paper introduces a novel and model-agnostic approach, Attention-Aware Regularization (AAR), which rethinks the attention mechanism by integrating causal reasoning to provide direct supervision of attention maps. AAR employs perturbation techniques and a unique optimization objective to assess and refine attention quality, encouraging the model to prioritize attribute-specific regions. Our method demonstrates significant improvement in PAR performance by mitigating the effects of incorrect attention and fostering a more effective attention mechanism. Experiments on standard datasets showcase the superiority of our approach over existing methods, setting a new benchmark for attention-driven PAR models.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001
Neurocomputing4
2025 Progressive fusion of local and global image features for cross-modal image aesthetic assessment
Yuzhen Niu, Siling Chen 0002
Multim. Syst.1
2025 High-order diversity feature learning for pedestrian attribute recognition
abstract
Pedestrian attribute recognition (PAR) involves accurately identifying multiple attributes present in pedestrian images. There are two main approaches for PAR: part-based method and attention-based method. The former relies on existing segmentation or region detection methods to localize body parts and learn corresponding attribute-specific feature from the corresponding regions, where the performance heavily depends on the accuracy of body region localization. The latter adopts the embedded attention modules or transformer attention to exploit detailed feature. However, it can focus on certain body regions but often provide coarse attention, failing to capture fine-grained details, the learned feature may also be interfered with by irrelevant information. Meanwhile, these methods overlook the global contextual information. This work argues for replacing coarse attention with detailed attention and integrating it with global contextual feature from ViT to jointly represent attribute-specific regions. To tackle this issue, we propose a High-order Diversity Feature Learning (HDFL) method for PAR based on ViT. We utilize a polynomial predictor to design an Attribute-specific Detailed Feature Exploration (ADFE) module, which can construct the high-order statistics and gain more fine-grained feature. Our ADFE module is a parameter-friendly method that provides flexibility in deciding its utilization during the inference phase. A Soft-redundancy Perception Loss (SPLoss) is proposed to adaptively measure the redundancy between feature of different orders, which can promote diverse characterization of features. Experiments on several PAR datasets show that our method achieves a new state-of-the-art (SOTA) performance. On the most challenging PA100K dataset, our method outperforms previous SOTA by 1.69% and achieves the highest mA of 84.92%.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001
Neural Networks4
2025 Collaboratively enhanced and integrated detail-context information for low-light image enhancement
Yuzhen Niu, Huangbiao Xu, Rui Xu 0028, Yuzhong Chen 0001
Pattern Recognit.1
2025 Learning Comprehensive Representation via Selective Activation and Dual-Level Orthogonality for Pedestrian Attribute Recognition
abstract
Multi-label Pedestrian Attribute Recognition (PAR) involves identifying a series of semantic attributes in person images. Existing PAR solutions typically rely on CNN as the backbone network to extract pedestrian features. Unfortunately, CNNs process only one adjacent region at a time, resulting in the disappearance of long-range relations between different attribute-specific regions. To address this limitation, we adopt the Vision Transformer (ViT) instead of CNN as the backbone for PAR, aiming to build long-range relations and extract more robust features. However, PAR suffers from an inherent attribute imbalance issue, causing ViT to naturally focus more on attributes that appear frequently in the training set and ignore some pedestrian attributes that appear less. The native features extracted by ViT are not able to tolerate the imbalance attribute distribution issue. To tackle this issue, we propose a novel component and a dual-level loss: the Selective Feature Activation Method (SFAM), the Orthogonal Feature Activation Loss (OFALoss), and Orthogonal Weight Regularization Loss (OWRLoss). SFAM smartly suppresses the more informative attribute-specific features, thus compelling the PAR model to pay greater attention to attribute-specific regions that are often overlooked. The proposed OFALoss enforces an orthogonal constraint on the original feature extracted by ViT and the suppressed features from SFAM, promoting the comprehensiveness of feature representation in each attribute-specific region. Furthermore, OWRLoss is employed for decreasing correlations among entries of the last shared classification layer, which can alleviate the highly correlated of weight vectors caused by non-uniform distribution. This can prevent excessive mutual interference among different attributes during attribute recognition. Our model-agnostic approach is plug-and-play, requiring no additional training parameters in the training process. We conduct experiments on several benchmark PAR datasets, including PETA, PA100K, RAPv1, and RAPv2, demonstrating the effectiveness of our method. Specifically, our method outperforms existing state-of-the-art approaches.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001, Jianqiang Zhao
IEEE Trans. Circuits Syst. Video Technol.4
2025 HUGS-Net: A Lightweight and Unified Network for Adverse Weather Image Denoising
abstract
Image denoising under adverse weather conditions aims to eliminate multiple weather-related noises and restore bright and clear images. Until now, most methods are task-specific while all-in-one algorithms often require a large number of parameters, limiting their model efficiency. Our theoretical analysis and statistical experiments reveal that adverse weather images in Hue channel contain rich contextual information for further processing. With this observation, we propose a novel lightweight HUe-Guided Synergistic Network (HUGS-Net) with multi-scale detail refinement. First, we design a Fourier interaction and evolution module to capture global information from Hue channel without introducing excessive network parameters. Second, we develop a lightweight residue group convolution block to model local texture features, incorporating them with global information to guide noises removal. Third, we introduce a multi-scale fusion module to enhance high-frequency details at a small feature resolution in RGB color space. With the above design, HUGS-Net further supervises and supplements refined background information. Comprehensive experiments showcase the superiority of HUGS-Net across various adverse weather datasets (e.g., image deraining, desnowing, dehazing) with the least parameter size and fast running speed.The source code will be made public after peer review process.
Runjie Wang, Yuzhen Niu, Tiesong Zhao
IEEE Trans. Multim.3
2025 Skeleton-Boundary-Guided Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) aims to resolve the tough issue of accurately segmenting objects hidden in the surroundings. However, the existing methods suffer from two major problems: the incomplete interior and the inaccurate boundary of the object. To address these difficulties, we propose a three-stage skeleton-boundary–guided network (SBGNet) for the COD task. Specifically, we design a novel skeleton-boundary label to be complementary to the typical pixel-wise mask annotation, emphasizing the interior skeleton and the boundary of the camouflaged object. Furthermore, the proposed feature guidance module (FGM) leverages the skeleton-boundary feature to guide the model to focus on both the interior and the boundary of the camouflaged object. Besides, we design a bidirectional feature flow path with the information interaction module (IIM) to propagate and integrate the semantic and texture information. Finally, we propose the dual feature distillation module (DFDM) to progressively refine the segmentation results in a fine-grained manner. Comprehensive experiments demonstrate that our SBGNet outperforms 20 state-of-the-art methods on three benchmarks in both qualitative and quantitative comparisons.
Yuzhen Niu, Yeyuan Xu, Yuezhou Li, Jiabang Zhang, Yuzhong Chen 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Selective and Orthogonal Feature Activation for Pedestrian Attribute Recognition
abstract
Pedestrian Attribute Recognition (PAR) involves identifying the attributes of individuals in person images. Existing PAR methods typically rely on CNNs as the backbone network to extract pedestrian features. However, CNNs process only one adjacent region at a time, leading to the loss of long-range inter-relations between different attribute-specific regions. To address this limitation, we leverage the Vision Transformer (ViT) instead of CNNs as the backbone for PAR, aiming to model long-range relations and extract more robust features. However, PAR suffers from an inherent attribute imbalance issue, causing ViT to naturally focus more on attributes that appear frequently in the training set and ignore some pedestrian attributes that appear less. The native features extracted by ViT are not able to tolerate the imbalance attribute distribution issue. To tackle this issue, we propose two novel components: the Selective Feature Activation Method (SFAM) and the Orthogonal Feature Activation Loss. SFAM smartly suppresses the more informative attribute-specific features, compelling the PAR model to capture discriminative features from regions that are easily overlooked. The proposed loss enforces an orthogonal constraint on the original feature extracted by ViT and the suppressed features from SFAM, promoting the complementarity of features in space. We conduct experiments on several benchmark PAR datasets, including PETA, PA100K, RAPv1, and RAPv2, demonstrating the effectiveness of our method. Specifically, our method outperforms existing state-of-the-art approaches by GRL, IAA-Caps, ALM, and SSC in terms of mA on the four datasets, respectively.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Jianqiang Zhao
AAAI4
2024 MiNet: Weakly-Supervised Camouflaged Object Detection through Mutual Interaction between Region and Edge Cues
abstract
Existing weakly-supervised camouflaged object detection (WSCOD) methods have much difficulty in detecting accurate object boundaries due to insufficient and imprecise boundary supervision in scribble annotations. Drawing inspiration from human perception that discerns camouflaged objects by incorporating both object region and boundary information, we propose a novel Mutual Interaction Network (MiNet) for scribble-based WSCOD to alleviate the detection difficulty caused by insufficient scribbles. The proposed MiNet facilitates mutual reinforcement between region and edge cues, thereby integrating more robust priors to enhance detection accuracy. In this paper, we first construct an edge cue refinement net, featuring a core region-aware guidance module (RGM) aimed at leveraging the extracted region feature as a prior to generate the discriminative edge map. By considering both object semantic and positional relationships between edge feature and region feature, RGM highlights the areas associated with the object in the edge feature. Subsequently, to tackle the inherent similarity between camouflaged objects and the surroundings, we devise a region-boundary refinement net. This net incorporates a core edge-aware guidance module (EGM), which uses the enhanced edge map from the edge cue refinement net as guidance to refine the object boundaries in an iterative and multi-level manner. Experiments on CAMO, CHAMELEON, COD10K, and NC4K datasets demonstrate that the proposed MiNet outperforms the state-of-the-art methods.
Yuzhen Niu, Lifen Yang, Rui Xu 0028, Yuezhou Li, Yuzhong Chen 0001
ACM Multimedia1
2024 Zero-Referenced Enlightening and Restoration for UAV Nighttime Vision
abstract
Unmanned aerial vehicle (UAV) based visual systems suffer from poor perception at nighttime. There are three challenges for enlightening nighttime vision for UAVs: Firstly, the UAV nighttime images differ from underexposed images in the statistical characteristic, limiting the performance of general low-light image enhancement (LLIE) methods. Secondly, when enlightening nighttime images, the artifacts tend to be amplified, distracting the visual perception of UAVs. Thirdly, due to the inherent scarcity of paired data in the real world, it is difficult for UAV nighttime vision to benefit from supervised learning. To meet these challenges, we propose a zero-referenced enlightening and restoration network (ZERNet) for improving the perception of UAV vision at nighttime. Specifically, by estimating the nighttime enlightening map (NE-map), a pixel-to-pixel transformation is then conducted to enlighten the dark pixels while suppressing overbright pixels. Furthermore, we propose the self-regularized restoration to preserve the semantic contents and restrict the artifacts in the final result. Finally, our method is derived from zero-referenced learning, which is free from paired training data. Comprehensive experiments show that the proposed ZERNet effectively improves the nighttime visual perception of UAVs on quantitative metrics, qualitative comparisons, and application-based analysis.
Yuezhou Li, Yuzhen Niu, Rui Xu 0028
IEEE Geosci. Remote. Sens. Lett.2
2024 Two-stream network with viewport selection for blind omnidirectional video quality assessment
Yuzhen Niu
Multim. Tools Appl.2
2024 Blind consumer video quality assessment with spatial-temporal perception and fusion
Yuzhen Niu, Yuming Zheng, Zhenlong Wang, Mengzhen Zhong, Tiesong Zhao
Multim. Tools Appl.1
2024 Multi-View Graph Embedding Learning for Image Co-Segmentation and Co-Localization
abstract
Image co-segmentation and co-localization exploit inter-image information to identify and extract foreground objects with a batch mode. However, they remain challenging when confronted with large object variations or complex backgrounds. This paper proposes a multi-view graph embedding (MV-Gem) learning scheme which integrates diversity, robustness and discernibility of object features to alleviate this phenomenon. To encourage the diversity, the deep co-information containing both low-layer general representations and high-layer semantic information is generated to form a multi-view feature pool for comprehensive co-object description. To enhance the robustness, a multi-view adaptive weighted learning is formulated to fuse the deep co-information for feature complementation. To ensure the discernibility, the graph embedding and sparse constraint are embedded into the fusion formulation for feature selection. The former aims to inherit important structures from multiple views, and the latter further selects important features to restrain irrelevant backgrounds. With these techniques, MV-Gem gradually recovers all co-objects through optimization iterations. Extensive experimental results on real-world datasets demonstrate that MV-Gem is capable of locating and delineating co-objects in an image group.
Aiping Huang, Lijian Li 0004, Le Zhang 0001, Yuzhen Niu, Tiesong Zhao, Chia-Wen Lin
IEEE Trans. Circuits Syst. Video Technol.4
2024 STD-Net: Spatio-Temporal Decomposition Network for Video Demoiréing With Sparse Transformers
abstract
The problem of video demoiréing is a new challenge in video restoration. Unlike image demoiréing, which involves removing static and uniform patterns, video demoiréing requires tackling dynamic and varied moiré patterns while maintaining video details, colors, and temporal consistency. It is particularly challenging to model moiré patterns for videos with camera or object motions, where separating moiré from the original video content across frames is extremely difficult. Nonetheless, we observe that the spatial distribution of moiré patterns is often sparse on each frame, and their long-range temporal correlation is not significant. To fully leverage this phenomenon, a sparsity-constrained spatial self-attention scheme is proposed to concentrate on removing sparse moiré efficiently for each frame without being distracted by dynamic video content. The frame-wise spatial features are then correlated and aggregated via the local temporal cross-frame-attention module to produce temporal-consistent high-quality moiré-free videos. The above decoupled spatial and temporal transformers constitute the Spatio-Temporal Decomposition Network, dubbed STD-Net. For evaluation, we present a large-scale video demoiréing benchmark featuring various real-life scenes, camera motions, and object motions. We demonstrate that our proposed model can effectively and efficiently achieve superior performance on video demoiréing and single image demoiréing tasks.The proposed dataset will be released after the paper is accepted.
Yuzhen Niu, Rui Xu 0028, Zhihua Lin, Wenxi Liu
IEEE Trans. Circuits Syst. Video Technol.1
2024 Perception-Driven Similarity-Clarity Tradeoff for Image Super-Resolution Quality Assessment
abstract
Super-Resolution (SR) algorithms aim to enhance the resolutions of images. Massive deep-learning-based SR techniques have emerged in recent years. In such case, a visually appealing output may contain additional details compared with its reference image. Accordingly, fully referenced Image Quality Assessment (IQA) cannot work well; however, reference information remains essential for evaluating the qualities of SR images. This poses a challenge to SR-IQA: How to balance the referenced and no-reference scores for user perception? In this paper, we propose a Perception-driven Similarity-Clarity Tradeoff (PSCT) model for SR-IQA. Specifically, we investigate this problem from both referenced and no-reference perspectives, and design two deep-learning-based modules to obtain referenced and no-reference scores. We present a theoretical analysis based on Human Visual System (HVS) properties on their tradeoff and also calculate adaptive weights for them. Experimental results indicate that our PSCT model is superior to the state-of-the-arts on SR-IQA. In addition, the proposed PSCT model is also capable of evaluating quality scores in other image enhancement scenarios, such as deraining, dehazing and underwater image enhancement. The source code is available at https://github.com/kekezhang112/PSCT.
Tiesong Zhao, Yuzhen Niu, Jinsong Hu 0001, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.4
2024 Perceptual Decoupling With Heterogeneous Auxiliary Tasks for Joint Low-Light Image Enhancement and Deblurring
abstract
Capturing images at night are susceptible to inadequate illumination conditions and motion blurring. Given the typical coupling of these two forms of degradation, a pioneer work takes a compact approach of brightening followed by deblurring. However, this sequential approach may compromise informative features and elevate the likelihood of generating unintended artifacts. In this paper, we observe that the co-existing low light and blurs intuitively impair multiple perceptions, making it difficult to produce visually appealing results. To meet these challenges, we propose perceptual decoupling with heterogeneous auxiliary tasks (PDHAT) for joint low-light image enhancement and deblurring. Based on the crucial perceptual properties of the two degradations, we construct two individual auxiliary tasks: coarse preview prediction (CPP) and high-frequency reconstruction (HFR), so that the perception of color, brightness, edges, and details are decoupled into heterogeneous auxiliary tasks to obtain task-specific representations for parallel assisting the main task: joint low-light enhancement and deblurring (LLE-Deblur). Furthermore, we develop dedicated modules to build the network blocks in each branch based on the exclusive properties of each task. Comprehensive experiments are conducted on LOL-Blur and Real-LOL-Blur datasets, showing that our method outperforms existing methods on quantitative metrics and qualitative results.
Yuezhou Li, Rui Xu 0028, Yuzhen Niu, Wenzhong Guo, Tiesong Zhao
IEEE Trans. Multim.3
2024 Bilateral Interaction for Local-Global Collaborative Perception in Low-Light Image Enhancement
abstract
Low-light image enhancement is a challenging task due to the limited visibility in dark environments. While recent advances have shown progress in integrating CNNs and Transformers, the inadequate local-global perceptual interactions still impedes their application in complex degradation scenarios. To tackle this issue, we propose BiFormer, a lightweight framework that facilitates local-global collaborative perception via bilateral interaction. Specifically, our framework introduces a core CNN-Transformer collaborative perception block (CPB) that combines local-aware convolutional attention (LCA) and global-aware recursive Transformer (GRT) to simultaneously preserve local details and ensure global consistency. To promote perceptual interaction, we adopt bilateral interaction strategy for both local and global perception, which involves local-to-global second-order interaction (SoI) in the dual-domain, as well as a mixed-channel fusion (MCF) module for global-to-local interaction. The MCF is also a highly efficient feature fusion module tailored for degraded features. Extensive experiments conducted on low-level and high-level tasks demonstrate that BiFormer achieves state-of-the-art performance. Furthermore, it exhibits a significant reduction in model parameters and computational cost compared to existing Transformer-based low-light image enhancement methods.
Rui Xu 0028, Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Yuzhong Chen 0001, Tiesong Zhao
IEEE Trans. Multim.3
2023 HCSD-Net: Single Image Desnowing with Color Space Transformation
abstract
Single-image desnowing aims at depressing snowflake noises while preserving a clean background. Existing methods usually mask the locations of noises and remove them in RGB color space. In this paper, we rethink this problem by investigating the impacts of color space selection. Theoretical analysis and experiments reveal that the feature of snowflake noises exhibit different distributions in different color spaces. In particular, these noises are barely seen in Hue channel, which inspires us to recover global structure and texture information of the clean background from Hue channel. More low-frequency information is also found in the Hue channel. With these observations, we propose a novel Hybrid-Color-Space-based Desnowing Network (HCSD-Net). The proposed HCSD-Net extracts low-frequency and high-frequency features in Hue channel and RGB color space, respectively. After that, it utilizes a multi-scale fusion module to enhance high-frequency details at a small feature resolution. These details are further used to supervise and supplement the background information. Extensive experiments demonstrate that our proposed HCSD-Net outperforms state-of-the-art methods on various synthetic and real-world desnowing datasets. Codes are available at https://github.com/ttz-rainbow/HCSD-Net.
Nanfeng Jiang, Hongxin Wu, Yuzhen Niu, Tiesong Zhao
ACM Multimedia5
2023 Zero-referenced low-light image enhancement with adaptive filter network
Yuezhou Li, Yuzhen Niu, Rui Xu 0028, Yuzhong Chen 0001
Eng. Appl. Artif. Intell.2
2023 Joint Shared-and-Specific Information for Deep Multi-View Clustering
abstract
Multi-view data describes an image sample with different modalities of features, thus provides a more comprehensive description of data. Its three basic characteristics, i.e., consensus, complementary and redundancy, determine its performances in computer vision tasks. In this paper, we effectively exploit the above three characteristics to propose a deep learning scheme with joint shared-and-specific information (JSSI) for multi-view clustering. Aiming at facilitating the consensus, JSSI extracts shared information of multi-view data via an adversarial similarity constraint, which is realized by classification and discrimination interactions. Aiming at reducing the redundancy, JSSI separate out view-specific features and prevent them from interfering with the shared features via a difference constraint. Aiming at ensuring the complementary, JSSI aligns the shared features and then concatenates them with the specific features. We examine the effectiveness of JSSI with multi-view clustering on real-world datasets, such as faces and indoor scenes. Extensive experiments and comparisons show that JSSI outperforms other state-of-the-art methods in most of these datasets.
Aiping Huang, Wei Gao 0003, Yuzhen Niu, Tiesong Zhao
IEEE Trans. Circuits Syst. Video Technol.4
2023 Comment-Guided Semantics-Aware Image Aesthetics Assessment
abstract
Existing image aesthetics assessment methods mainly rely on the visual features of images but ignore their rich semantics. Nowadays, with the widespread application of social media, the comments corresponding to images in the form of texts can be easily accessed and provide rich semantic information, which can be utilized to effectively complement image features. This paper proposes a comment-guided semantics-aware image aesthetics assessment method, which is built upon a multi-task learning framework for image aesthetics prediction and comment-guided semantics classification. To assist image aesthetics assessment, we first model the semantics of an image as the topic features of its corresponding comments using Latent Dirichlet Allocation. We then propose a two-stream multitask learning framework for both topic feature prediction and aesthetic score distribution prediction. Topic feature prediction task enables to infer the semantics from images, since the comments are usually unavailable during inference and comment-guided semantics can only serve as supervision during training. We further propose to deeply fuse aesthetics and semantic features using a layerwise feature fusion method. Experimental results demonstrate that the proposed method outperforms state-of-the-art image aesthetics assessment methods.
Yuzhen Niu, Bingrui Song, Wenxi Liu
IEEE Trans. Circuits Syst. Video Technol.1
2023 Progressive Moire Removal and Texture Complementation for Image Demoireing
abstract
Taking photos of digital screens often produces color-distorting moire patterns caused by inconsistency between the color filter array of cameras and the sub-pixel layout of screens, which severely degrades the quality of photos. Most existing demoireing methods employ multi-stream network architecture to simultaneously process the same moire image with different resolutions, but they neglect the complementarity among different resolutions. In this paper, we propose a novel moire removal model to address this issue. Unlike the existing multi-stream based approaches, in this model, we present a progressive texture complementation block to exploit the complementary information from different resolutions in order to progressively remove moire textures and restore image content. Additionally, we propose a residual moire removal block, in which the depthwise separable convolution is utilized to remove moire from image while reducing computation overhead. This block also includes a local color correction structure, which is used to correct color shifts presented in the moire images. Experimental results on two public datasets show that our method outperforms state-of-the-art methods. Besides, the quantity of parameters and FLOPs of our model are tens of times fewer than the off-the-shelf models. Furthermore, our network framework can adapt well to another low-level vision task, rain removal, in which our model also achieves state-of-the-art performance.
Yuzhen Niu, Zhihua Lin, Wenxi Liu, Wenzhong Guo
IEEE Trans. Circuits Syst. Video Technol.1
2023 High Efficiency Vibrotactile Codec Based on Gate Recurrent Network
abstract
The multimedia has achieved dominant positions in both local storage and internet bandwidth, which inevitably promotes the compression of audio, image and video information. Nowadays, the emerging haptic technology, which enhances the immersion in virtual reality and remote control, has also brought new challenges in its codec design. It is thus imperative to develop haptic codecs, including kinesthetic and vibrotactile codecs, with high efficiency and low delay. In this paper, we exploit statistical features of vibrotactile data to develop a Recurrent-Network-based Vibrotactile Codec (RNVC) with high compression efficiency and low coding delay. The proposed encoder consists of vibrotactile estimation by Gate Recurrent Unit (GRU), non-uniform quantization/compensation of residuals and an entropy encoder. In particular, the GRU-based recurrent network is utilized for its high efficiency to predict signals and low complexity to converge. The decoder consists of all counterparts of encoder. Experimental results show the proposed RNVC significantly reduces of original bitrates with negligible encoding delay, which achieves the state-of-the-art coding performance of vibrotactile signal.
Tiesong Zhao, Qian Liu 0001, Yuzhen Niu
IEEE Trans. Multim.5
2022 Perceptual-Aware and Restorable Real-Time Image Downscaling
abstract
Image downscaling has been a classical problem and has recently been linked to super-resolution (SR). In this paper, we aim to propose a learning-based image downscaling model, FastDownscaler, which can efficiently produce low-resolution (LR) images that not only preserve the rich details of the original high-resolution images but also be highly restorable for existing SR models. We first present two separate lightweight networks with different upsampling losses, the bilinear loss and the bicubic loss, which are better for SR restoration and LR downscaling, respectively. To produce versatile LR images, we then propose to distill bilinear loss guided network with bicubic loss guided one. To our best knowledge, we establish the first image downscaling quality assessment dataset to evaluate the downscaling performance. Experimental results demonstrate the superior performance of the proposed model on image downscaling and SR. Furthermore, our model can achieve over 600 FPS for downscaling a$1920\times 1280$image.
Yuzhen Niu, Luwei Zheng, Jianbin Wu, Wenxi Liu
ICME1
2022 Learning-Based Video Coding with Joint Deep Compression and Enhancement
abstract
End-to-end learning-based video coding has attracted substantial attentions by compressing video signals as stacked visual features. This paper proposes an end-to-end deep video codec with jointly optimized compression and enhancement modules (JCEVC). First, we propose a dual-path generative adversarial network (DPEG) to reconstruct video details after compression. An α-path and a β-path concurrently reconstruct the structure information and local textures. Second, we reuse the DPEG network in both motion compensation and quality enhancement modules, which are further combined with other necessary modules to formulate our JCEVC framework. Third, we employ a joint training of deep video compression and enhancement that further improves the rate-distortion (RD) performance of compression. Compared with x265 LDP very fast mode, our JCEVC reduces the average bit-per-pixel (bpp) by 39.39%/54.92% at the same PSNR/MS-SSIM, which outperforms the state-of-the-art deep video codecs by a considerable margin. Sourcecode is available at: https://github.com/fwz1021/JCEVC.
Tiesong Zhao, Weize Feng, Hongji Zeng, Yuzhen Niu, Jiaying Liu 0001
ACM Multimedia5
2022 Continuous Transformation Superposition for Visual Comfort Enhancement of Casual Stereoscopic Photography
abstract
Casual stereoscopic photography allows ordinary users to create a stereoscopic photo using two photos taken casually by a monocular camera. The visual comfort of a casual stereoscopic photo can greatly affect its visual experience. In this paper, we present a novel visual comfort enhancement method for casual stereoscopic photography via reinforcement learning based on continuous transformation superposition. We consider the transformation, in a continuous transformation space, to transform each view as superpositions of several basic continuous transformations, enabling more subtle and flexible image transformation operations to approach better solutions. To achieve the continuous transformation superposition, we prepare a collection of continuous transformation models for translation, rotation, and perspective transformations. Then we train a policy model to determine an optimal transformation chain to recurrently handle both the geometric constraints and disparity adjustment, and thereby enhance the visual comfort of casual stereoscopic images. We further propose an attention-based stereo feature fusion module that enhances and integrates the binocular information between the left and right views. Experimental results on three datasets demonstrate that our proposed method achieves superior performance to state-of-the-art methods.
Yuzhong Chen 0001, Qijin Shen, Yuzhen Niu, Wenxi Liu
VR3
2022 Accelerating the discovery of anticancer peptides targeting lung and breast cancers with the Wasserstein autoencoder model and PSO algorithm
abstract
In the development of targeted drugs, anticancer peptides (ACPs) have attracted great attention because of their high selectivity, low toxicity and minimal non-specificity. In this work, we report a framework of ACPs generation, which combines Wasserstein autoencoder (WAE) generative model and Particle Swarm Optimization (PSO) forward search algorithm guided by attribute predictive model to generate ACPs with desired properties. It is well known that generative models based on Variational AutoEncoder (VAE) and Generative Adversarial Networks (GAN) are difficult to be used for de novo design due to the problems of posterior collapse and difficult convergence of training. Our WAE-based generative model trains more successfully (lower perplexity and reconstruction loss) than both VAE and GAN-based generative models, and the semantic connections in the latent space of WAE accelerate the process of forward controlled generation of PSO, while VAE fails to capture this feature. Finally, we validated our pipeline on breast cancer targets (HIF-1) and lung cancer targets (VEGR, ErbB2), respectively. By peptide-protein docking, we found candidate compounds with the same binding sites as the peptides carried in the crystal structure but with higher binding affinity and novel structures, which may be potent antagonists that interfere with these target-mediated signaling.
Guanghui Yang, Zhi-Tong Bing, Liang Huang 0005, Yuzhen Niu, Lei Yang 0021
Briefings Bioinform.6
2022 Second-order information bottleneck based spiking neural networks for sEMG recognition
Anguo Zhang, Yuzhen Niu, Yueming Gao, Junyi Wu 0001
Inf. Sci.2
2022 Efficient Encoder-Decoder Network With Estimated Direction for SAR Ship Detection
abstract
Synthetic aperture radar (SAR) image ship detection has important applications in marine surveillance. There are two limitations when applying advanced detection methods naively for SAR ship detection. First, most detectors construct the model as an encoder and rely on the feature pyramid network (FPN) head for accurate prediction, which may lead to high computational costs. Second, the background noises in the ground truth (annotated as rectangular bounding boxes) of angular ships bring difficulties for model training. To meet these challenges, we propose an efficient encoder–decoder network with estimated direction for ship detection in SAR images. First, we present an anchor-free encoder–decoder model that can efficiently extract multiple-level features. Second, we formulate ship detection as a multitask learning problem, including a bounding box prediction and a ship direction regression. The estimated ship direction can weakly supervise and benefit ship detection. Furthermore, we develop a center-weighted labeling method for overlapped annotations. Comprehensive experiments on SAR-Ship-Detection and SSDD datasets show that our method achieves state-of-the-art performance with a high running speed.
Yuzhen Niu, Yuezhou Li, Jiangyi Huang, Yuzhong Chen 0001
IEEE Geosci. Remote. Sens. Lett.1
2022 Event-Driven Intrinsic Plasticity for Spiking Convolutional Neural Networks
abstract
The biologically discovered intrinsic plasticity (IP) learning rule, which changes the intrinsic excitability of an individual neuron by adaptively turning the firing threshold, has been shown to be crucial for efficient information processing. However, this learning rule needs extra time for updating operations at each step, causing extra energy consumption and reducing the computational efficiency. The event-driven or spike-based coding strategy of spiking neural networks (SNNs), i.e., neurons will only be active if driven by continuous spiking trains, employs all-or-none pulses (spikes) to transmit information, contributing to sparseness in neuron activations. In this article, we propose two event-driven IP learning rules, namely, input-driven and self-driven IP, based on basic IP learning. Input-driven means that IP updating occurs only when the neuron receives spiking inputs from its presynaptic neurons, whereas self-driven means that IP updating only occurs when the neuron generates a spike. A spiking convolutional neural network (SCNN) is developed based on the ANN2SNN conversion method, i.e., converting a well-trained rate-based artificial neural network to an SNN via directly mapping the connection weights. By comparing the computational performance of SCNNs with different IP rules on the recognition of MNIST, FashionMNIST, Cifar10, and SVHN datasets, we demonstrate that the two event-based IP rules can remarkably reduce IP updating operations, contributing to sparse computations and accelerating the recognition process. This work may give insights into the modeling of brain-inspired SNNs for low-power applications.
Anguo Zhang, Xiumin Li, Yueming Gao, Yuzhen Niu
IEEE Trans. Neural Networks Learn. Syst.4
2021 Coarse-To-Fine Person Re-Identification With Auxiliary-Domain Classification and Second-Order Information Bottleneck
abstract
Person re-identification (Re-ID) is to retrieve a particular person captured by different cameras, which is of great significance for security surveillance and pedestrian behavior analysis. However, due to the large intra-class variation of a person across cameras, e.g., occlusions, illuminations, viewpoints, and poses, Re-ID is still a challenging task in the field of computer vision. In this paper, to attack the issues concerning with intra-class variation, we propose a coarse-to-fine Re-ID framework with the incorporation of auxiliary-domain classification (ADC) and second-order information bottleneck (2O-IB). In particular, as an auxiliary task, ADC is introduced to extract the coarse-grained essential features to distinguish a person from miscellaneous backgrounds, which leads to the effective coarse- and fine-grained feature representations for Re-ID. On the other hand, to cope with the redundancy, irrelevance, and noise contained in the Re-ID features caused by intra-class variations, we integrate 2O-IB into the network to compress and optimize the features, without increasing additional computation overhead during inference. Experimental results demonstrate that our proposed method significantly reduces the neural network output variance of intra-class person images and achieves the superior performance to state-of-the-art methods.
Anguo Zhang, Yueming Gao, Yuzhen Niu, Wenxi Liu, Yongcheng Zhou
CVPR3
2021 Game Theory-driven Rate Control for 360-Degree Video Coding
abstract
The 360-degree video (omnidirectional video) has become popular recently due to its capability of providing immersive experience, which is generally achieved via spherical moving pictures with freedom of viewpoint changing. Nevertheless, the support of full-view visual contents has inevitably reshaped its perceptual quality metric and dramatically increased its bitrate output after video coding. Therefore in 360-degree video coding, the Rate Control (RC) problem, which aims to maximize the resulted perceptual quality under bitrate constraint, has become a challenging task yet to be addressed. In this paper, we observe a latitude-based bitrate discrepancy in equirectangular-projected 360-degree video coding and further utilize this feature in bitrate allocation under panoramic vision. We introduce game theory to find optimal inter/intra-frame bit allocations that maximize the overall RC performance in terms of utility function. Finally, an overall framework is proposed that is capable of providing both an improved bitrate accuracy and an enhanced perceptual quality. Experimental results demonstrate the efficiency of proposed method, with promising RC performances for 4K and 8K 360-degree videos.
Tiesong Zhao, Jielian Lin, Xu Wang 0006, Yuzhen Niu
ACM Multimedia5
2021 Image Retargeting Quality Assessment Based on Registration Confidence Measure and Noticeability-Based Pooling
abstract
Nowadays, image retargeting approaches have been widely applied to adapt images of various resolutions to heterogenous display devices. To assess the quality of the retargeted images, image retargeting quality assessment (IRQA) has emerged as a critical problem in image quality assessment. In this paper, we address the IRQA problem with a newly proposed framework based on registration confidence measurement (RCM) and noticeability-based pooling (NBP). First, we define the RCM to evaluate the accuracy of image registration, which aligns scenes between the original and retargeted images. We then integrate the proposed RCM with the computed local fidelity of each image block to alleviate the negative influence of inaccurate registration on fidelity measurements. Meanwhile, we present a visual attention fusion (VAF) framework to enhance faces and lines in the saliency map, which are observed to be highly sensitive in the human visual system (HVS). Finally, we propose the NBP strategy, which aggregates the local fidelity of each image block into the overall quality of the retargeted image. Specifically, the NBP strategy sets larger quality ranges for the regions where the visual distortions are more accessible to HVS to reflect the easy noticeability of these regions. Experimental results on the MIT RetargetMe and CUHK datasets demonstrate that the proposed IRQA metric based on RCM and NBP outperforms the state-of-the-art IRQA metrics.
Yuzhen Niu, Zhishan Wu, Tiesong Zhao
IEEE Trans. Circuits Syst. Video Technol.1
2021 HDR-GAN: HDR Image Reconstruction From Multi-Exposed LDR Images With Large Motions
abstract
Synthesizing high dynamic range (HDR) images from multiple low-dynamic range (LDR) exposures in dynamic scenes is challenging. There are two major problems caused by the large motions of foreground objects. One is the severe misalignment among the LDR images. The other is the missing content due to the over-/under-saturated regions caused by the moving objects, which may not be easily compensated for by the multiple LDR exposures. Thus, it requires the HDR generation model to be able to properly fuse the LDR images and restore the missing details without introducing artifacts. To address these two problems, we propose in this paper a novel GAN-based model, HDR-GAN, for synthesizing HDR images from multi-exposed LDR images. To our best knowledge, this work is the first GAN-based approach for fusing multi-exposed LDR images for HDR reconstruction. By incorporating adversarial learning, our method is able to produce faithful information in the regions with missing content. In addition, we also propose a novel generator network, with a reference-based residual merging block for aligning large object motions in the feature domain, and a deep HDR supervision scheme for eliminating artifacts of the reconstructed HDR images. Experimental results demonstrate that our model achieves state-of-the-art reconstruction performance over the prior HDR methods on diverse scenes.
Yuzhen Niu, Jianbin Wu, Wenxi Liu, Wenzhong Guo, Rynson W. H. Lau
IEEE Trans. Image Process.1
2020 Recurrent Enhancement of Visual Comfort for Casual Stereoscopic Photography
abstract
Creating stereoscopic 3D media content has wide applications in virtual reality. In this paper, we are interested in a challenging application, casual stereoscopic photography, that allows ordinary users to create a stereoscopic photo using two images captured by a hand-held monocular camera. To handle the geometric constraints and disparity adjustment for casually captured left and right images, we present a coarse-to-fine framework. In the coarse stage, we propose a unified reinforcement learning-based method, in which the produced stereo image is iteratively adjusted and evaluated in the term of visual comfort. In addition, to further enhance the visual comfort of the stereoscopic image produced in the coarse stage, we introduce another independent recurrent network to fine-tune its disparity range. Lastly, we perform comprehensive experiments to evaluate our method and demonstrate the applicability of our model for real images.
Yuzhen Niu, Qingyang Zheng, Wenxi Liu, Wenzhong Guo
VR1
2020 No-reference stereoscopic image quality assessment using a multi-task CNN and registered distortion representation
Yiqing Shi, Wenzhong Guo, Yuzhen Niu, Jiamei Zhan
Pattern Recognit.3
2020 Single image reflection removal based on structure-texture layering
Nanfeng Jiang, Yuzhen Niu, Liqun Lin, Nadir Mustafa, Tiesong Zhao
Signal Process. Image Commun.2
2020 Matting-Based Residual Optimization for Structurally Consistent Image Color Correction
abstract
Image color correction aims to eliminate color differences between images, especially for color consistency in panoramic or stereoscopic images. Nowadays, the global color correction methods cannot correct local color differences, while the local color correction methods usually lead to structural inconsistency between local regions and image clarity reduction. To address these problems, we propose a matting-based residual optimization for structurally consistent image color correction (MROC). Inspired by the residual image, we formulate the image color correction problem as the optimization of a residual image between the input target image and the resulting image. The residual image is initialized and improved by the soft matting method with a closed-form solution. Besides, a data term is introduced to identify those pixels with higher color and structural consistencies and preserve these pixels during optimization. The whole computational infrastructure operates at the pixel level to correct local color differences while maintaining image clarity. Experimental results demonstrate that the performance of the proposed MROC method is superior to the state-of-the-art image color correction methods. Furthermore, the proposed matting-based residual optimization can also be incorporated in a variety of color correction methods, with enhanced outcomes justified by a group of image quality assessment metrics.
Yuzhen Niu, Tiesong Zhao
IEEE Trans. Circuits Syst. Video Technol.1
2020 Visually Consistent Color Correction for Stereoscopic Images and Videos
abstract
In stereoscopic 3D (S3D) color correction, visual inconsistency is a common problem that leads to perceptual quality degradations. In this paper, we propose an S3D image/video color correction strategy that resolves global, local, and temporal color discrepancies simultaneously. We achieve the image-based S3D color correction by three steps: a coarse-grain color correction for global color matching, a fine-grain color correction to further improve both global and local color consistencies, and a guided filtering process to guarantee the structural consistency before and after color correction. In addition, we extend the above strategy to S3D and multiview video color correction. To achieve temporal consistency between successive video frames, we develop an improved histogram matching within a sliding window on time axis. In our method, the mapping functions for each color channel change gradually following the video stream to avoid abrupt temporal changes in colors. The experimental results demonstrate that the proposed strategy outperforms the state-of-the-art color correction algorithms for images and videos.
Yuzhen Niu, Xiaohua Zheng, Tiesong Zhao
IEEE Trans. Circuits Syst. Video Technol.1
2020 Boundary-Aware RGBD Salient Object Detection With Cross-Modal Feature Sampling
abstract
Mobile devices usually mount a depth sensor to resolve ill-posed problems, like salient object detection on cluttered background. The main barrier of exploring RGBD data is to handle the information from two different modalities. To cope with this problem, in this paper, we propose a boundary-aware cross-modal fusion network for RGBD salient object detection. In particular, to enhance the fusion of color and depth features, we present a cross-modal feature sampling module to balance the contribution of the RGB and depth features based on the statistics of their channel values. In addition, in our multi-scale dense fusion network architecture, we not only incorporate edge-sensitive losses to preserve the boundary of the detected salient region, but also refine its structure by merging the estimated saliency maps of different scales. We accomplish the multi-scale saliency map merging using two alternative methods which produce refined saliency maps via per-pixel weighted combination and an encoder-decoder network. Extensive experimental evaluations demonstrate that our proposed framework can achieve the state-of-the-art performance on several public RGBD-based datasets.
Yuzhen Niu, Guanchao Long, Wenxi Liu, Wenzhong Guo, Shengfeng He
IEEE Trans. Image Process.1
2019 End-to-End Automatic Image Annotation Based on Deep CNN and Multi-Label Data Augmentation
abstract
Automatic image annotation is a key step in image retrieval and image understanding. In this paper, we present an end-to-end automatic image annotation method based on a deep convolutional neural network (CNN) and multi-label data augmentation. Different from traditional annotation models that usually perform feature extraction and annotation as two independent tasks, we propose an end-to-end automatic image annotation model based on deep CNN (E2E-DCNN). E2E-DCNN transforms the image annotation problem into a multi-label learning problem. It uses a deep CNN structure to carry out the adaptive feature learning before constructing the end-to-end annotation structure using multiple cross-entropy loss functions for training. It is difficult to train a deep CNN model using small-scale datasets or scale up multi-label datasets using traditional data augmentation methods; hence, we propose a multi-label data augmentation method based on Wasserstein generative adversarial networks (ML-WGAN). The ML-WGAN generator can approximate the data distribution of a single multi-label image. The images generated by ML-WGAN can assist in the reduction of the over-fitting problem of training a deep CNN model and enhance the generalization ability of the trained CNN model. We optimize the network structure by using deformable convolution and spatial pyramid pooling. We experiment the proposed E2E-DCNN model with data augmentation by the proposed ML-WGAN on several public datasets. The experimental results demonstrate that the proposed model outperforms the state-of-the-art automatic image annotation models.
Xiao Ke, Jiawei Zou, Yuzhen Niu
IEEE Trans. Multim.3
2018 Stereoscopic Image Quality Assessment Based on both Distortion and Disparity
abstract
Understanding the characteristics of high-quality stereoscopic 3D (S3D) images has great significance for S3D image classification, quality assessment, and quality enhancement. Existing works assess the quality of an S3D image from a single perspective and use databases with subjective opinion scores obtained by conducting subjective experiments in labs. So the performance of these works in real applications is unclear and questionable. In this paper, we propose an S3D image quality assessment index based on two important factors, namely DisTortion and DisParity (DTDP). Distortion reflects the extent of distortion of each view of an S3D image, respectively. Disparity is a distinguishing factor for an S3D image as compared with a monocular image. We design some features to represent these two factors and use a random forest regression (RFR) to learn the mapping between the features and subjective opinion scores. The database we choose is the NVIDIA 3D VISION LIVE Highest Rated database which is from real application. The images and the subjective opinion scores are directly come from the website viewers all over the world. Experimental results demonstrate a superior performance of the proposed DTDP index as compared to the existing methods.
Yuzhen Niu, Yini Zhong, Xiao Ke, Yiqing Shi
VCIP1
2018 CF-based optimisation for saliency detection
abstract
In view of the observation that saliency maps generated by saliency detection algorithms usually show similarity imperfection against the ground truth, the authors propose an optimisation algorithm based on clustering and fitting (CF) for saliency detection. The algorithm uses a fitting model to represent the quantitative relationship between ground truth and algorithm‐generated saliency maps. The authors use the K ‐means method to cluster the images into k clusters according to the similarities among images. Image similarity is measured in terms of scene and colour by using the GIST and colour histogram features, after which the fitting model for each cluster is calculated. The saliency map of a new image is optimised by using one of the fitting models which correspond to the cluster to which the image belongs. Experimental results show that their CF‐based optimisation algorithm improves the performance of various single image saliency detection algorithms. Moreover, the improvement achieved by their algorithm when using both CF strategies is greater than the improvement achieved by the same algorithm when not using the clustering strategy. In addition, their proposed optimisation algorithm can also effectively optimise co‐saliency detection algorithms which already consider multiple similar images simultaneously to improve saliency of single images.
Yuzhen Niu, Wenqi Lin, Xiao Ke
IET Comput. Vis.1
2018 Meta-metric for saliency detection evaluation metrics based on application preference
Yuzhen Niu, Jianer Chen, Wenzhong Guo
Multim. Tools Appl.1
2018 Region-Aware Image Denoising by Exploring Parameter Preference
abstract
The goal of an image denoising algorithm is to preserve the details of clean images while reducing the noise in noisy images. Some existing image denoising algorithms preserve the details by using external information. However, external information needs to be obtained from external images or regions similar to the noisy images or regions. In this letter, we propose a region-aware image denoising algorithm (RAID) by exploring parameter preference. The proposed RAID algorithm is based on the observation that different regions of noisy images prefer denoised results that differ due to being obtained with different denoising parameters. The RAID algorithm, first, measures the extent of preferences of denoising parameters for different image regions. Then, it combines the denoised results obtained by using the various denoising parameters according to the extent of preferences determined in the previous step to get the final denoised image. Experimental results show that the proposed RAID algorithm can be combined with some existing image denoising algorithms to improve their denoising performance. Our denoised results are better at preserving the details while reducing the noise than existing algorithms which use the same parameter for the whole image.
Yuzhen Niu, Yang Yang 0026, Wenzhong Guo, Lening Lin
IEEE Trans. Circuits Syst. Video Technol.1
2018 Image Quality Assessment for Color Correction Based on Color Contrast Similarity and Color Value Difference
abstract
Color correction plays an important role in the image processing field. But substantial research on the assessment of color correction is still insufficient. In this paper, we present an image quality assessment metric for color correction. It assesses the color consistency between the reference and target/result images of color correction according to their color contrast similarity and color value difference. Both the average difference and difference span are considered during the assessment. To compensate for the scene difference between the reference and target/result images of color correction, we propose to use an image registration algorithm to build their matching relationship, upon which a matching image is built. The matching image has the same scene as the target image and the same color feature as the reference image, and thus the matching image is regarded as the real reference image of our color correction assessment. Furthermore, we combine a confidence map of the matching image and a saliency map of the target/result image as a weighting map for assessment, which helps to improve the consistency between the objective and subjective assessment results. The experimental results show that our color correction assessment metric has better correlation, accuracy, and monotonicity with users' subjective scores than 19 state-of-the-art metrics.
Yuzhen Niu, Haifeng Zhang 0012, Wenzhong Guo, Rongrong Ji
IEEE Trans. Circuits Syst. Video Technol.1
2017 Fitting-based optimisation for image visual salient object detection
abstract
To overcome some major problems with traditional saliency evaluation metrics, full‐reference image quality assessment (IQA) metrics, which have similar but stricter objectives, are used. Inspired by the root mean absolute error, the authors propose a fitting‐based optimisation method for salient object detection algorithms. Their algorithm analyses the quantitative relationship between saliency and ground truth values, and uses the derived relationship to fit the saliency values to the original saliency maps. This ensures that the resulting images, which are composed of fitted values, are closer to the ground truth. The proposed algorithm first computes the statistics of the ground truth and saliency maps computed by each salient object detection algorithm. These statistics are used to compute the parameters of four fitting models, which generally agree with the characteristics of the statistical data. For a new saliency map, they use the fitting model with the computed parameters to obtain the fitted saliency values, which are confined to the range [0, 255]. Finally, they evaluate their saliency optimisation algorithm using traditional evaluation metrics, IQA metrics, and a content‐based image retrieval application. The results show that the proposed approach improves the quality of the optimised saliency maps.
Yuzhen Niu, Wenqi Lin, Xiao Ke, Lingling Ke
IET Comput. Vis.1
2017 Machine learning-based framework for saliency detection in distorted images
Yuzhen Niu, Lening Lin, Yuzhong Chen 0001, Lingling Ke
Multim. Tools Appl.1
2017 Data equilibrium based automatic image annotation by fusing deep model and semantic propagation
Xiao Ke, Ming-Ke Zhou, Yuzhen Niu, Wenzhong Guo
Pattern Recognit.3
2016 Accumulative Energy-Based Seam Carving for Image Resizing
abstract
With the diversified development of the digital devices, such as computer, mobile phone, pad and television, how to resize an image or video to adapt to different display screens has been attracting more and more peoples' attention. Seam carving has been an important method for image resizing. If multiple removed or inserted seams are located within a certain region, it can lead to discontinuity image content. Besides, the salient objects tend to be destroyed if the energy function only contains the gradient information. Therefore, we propose an accumulative energy-based seam carving method for image resizing. When removing a certain seam, we distribute the energy of each pixel on the seam to its adjacent 8-connected pixels in order to avoid the extreme concentration of seams, especially within a texture region. In addition, we add the image saliency and the edge information into the energy function to reduce the distortion. Since the computational complexity of seam carving method is very high, we use parallel computing environment to achieve efficient computation. Experimental results show that compared with the existing methods, our method can both avoid the discontinuity of image content and distortions as well as better maintain the shape of the salient objects.
Yuzhen Niu, Jiawen Lin, Haifeng Zhang 0012
PDCAT2
2016 Evaluation of visual saliency analysis algorithms in noisy images
Yuzhen Niu, Lingling Ke, Wenzhong Guo
Mach. Vis. Appl.1
2016 Erratum to: Geometry-shader-based real-time voxelization and applications
Shu-Huai Chang, Yu-Chi Lai, Chih-Yuan Yao, Kai-Lung Hua, Yuzhen Niu, Feng Liu 0015
Vis. Comput.5
2015 XGRouter: high-quality global router in X-architecture with particle swarm optimization
Genggeng Liu, Wenzhong Guo, Rongrong Li, Yuzhen Niu
Frontiers Comput. Sci.4
2015 Multi-scale saliency detection via inter-regional shortest colour path
abstract
Saliency detection has attracted considerable attention, and numerous approaches aimed at locating meaningful regions in images have been presented. Nevertheless, accurate saliency detection algorithms remain in urgent demand. Many algorithms work well when dealing with simple images, but work poorly with complex images that contain small‐scale and high‐contrast structures. Moreover, most existing local and global regional saliency detection methods measure image saliency through region contrast. Such measurement is achieved by directly computing the difference between non‐adjacent regions. In this study, the authors introduce a new perspective for evaluating region contrast. We propose a novel multi‐scale saliency region detection method by optimising the shortest path of two non‐adjacent regions in the colour space and by measuring the region contrast from different scales. The final saliency maps indicate that the proposed method can work well with images containing small patches, but with high contrast. The proposed approach can also make the foreground significantly more uniform. Experimental results on three public benchmark datasets show that the proposed method achieves better precision–recall curve than some state‐of‐the‐art methods.
Wenzhong Guo, Xiaolong Sun, Yuzhen Niu
IET Comput. Vis.3
2015 A PSO-based timing-driven Octilinear Steiner tree algorithm for VLSI routing considering bend reduction
Genggeng Liu, Wenzhong Guo, Yuzhen Niu, Xing Huang 0001
Soft Comput.3
2015 Multilayer Obstacle-Avoiding X-Architecture Steiner Minimal Tree Construction Based on Particle Swarm Optimization
abstract
As the basic model for very large scale integration routing, the Steiner minimal tree (SMT) can be used in various practical problems, such as wire length optimization, congestion, and time delay estimation. In this paper, an effective algorithm based on particle swarm optimization is presented to construct a multilayer obstacle-avoiding X-architecture SMT (ML-OAXSMT). First, a pretreatment strategy is presented to reduce the total number of judgments for the routing conditions around obstacles and vias. Second, an edge transformation strategy is employed to make the particles have the ability to bypass the obstacles while the union-find partition is used to prevent invalid solutions. Third, according to the feature of ML-OAXSMT problem, we design an edge-vertex encoding strategy, which has the advantage of simple and effective. Moreover, a penalty mechanism is proposed to help the particle bypass the obstacles, and reduce the generation of via at the same time. Experimental results show that our algorithm from a global perspective of multilayer structure can achieve the best solution quality among the existing algorithms. Finally, to our best knowledge, we redefine the edge cost and then construct the obstacle-avoiding preferred direction X-architecture Steiner tree, which is the first work to address this problem and can offer the theory supports for chip design based on non-Manhattan architecture.
Genggeng Liu, Xing Huang 0001, Wenzhong Guo, Yuzhen Niu
IEEE Trans. Cybern.4
2015 Obstacle-Avoiding Algorithm in X-Architecture Based on Discrete Particle Swarm Optimization for VLSI Design
abstract
Obstacle-avoiding Steiner minimal tree (OASMT) construction has become a focus problem in the physical design of modern very large-scale integration (VLSI) chips. In this article, an effective algorithm is presented to construct an OASMT based on X-architecturex for a given set of pins and obstacles. First, a kind of special particle swarm optimization (PSO) algorithm is proposed that successfully combines the classic genetic algorithm (GA), and greatly improves its own search capability. Second, a pretreatment strategy is put forward to deal with obstacles and pins, which can provide a fast information inquiry for the whole algorithm by generating a precomputed lookup table. Third, we present an efficient adjustment method, which enables particles to avoid all the obstacles by introducing some corner points of obstacles. Finally, an excellent refinement method is discussed to further enhance the quality of the final routing tree, which can improve the quality of the solution by 7.93% on average. To our best knowledge, this is the first time to specially solve the single-layer obstacle-avoiding problem in X-architecture. Experimental results show that the proposed algorithm can further shorten wirelength in the presence of obstacles. And it achieves the best solution quality in a reasonable runtime among the existing algorithms.
Xing Huang 0001, Genggeng Liu, Wenzhong Guo, Yuzhen Niu
ACM Trans. Design Autom. Electr. Syst.4
2015 A PSO-Optimized Real-Time Fault-Tolerant Task Allocation Algorithm in Wireless Sensor Networks
abstract
One of challenging issues for task allocation problem in wireless sensor networks (WSNs) is distributing sensing tasks rationally among sensor nodes to reduce overall power consumption and ensure these tasks finished before deadlines. In this paper, we propose a soft real-time fault-tolerant task allocation algorithm (FTAOA) for WSNs in using primary/backup (P/B) technique to support fault tolerance mechanism. In the proposed algorithm, the construction process of discrete particle swarm optimization (DPSO) is achieved through adopting a binary matrix encoding form, minimizing tasks execution time, saving node energy cost, balancing network load, and defining a fitness function for improving scheduling effectiveness and system reliability. Furthermore, FTAOA employs passive backup copies overlapping technology and is capable to determinate the mode of backup copies adaptively through scheduling primary copies as early as possible and backup copies as late as possible. To improve resource utilization, we allocate tasks to the nodes with high performance in terms of load, energy consumption, and failure ratio. Analysis and simulation results show the feasibility and effectiveness of FTAOA. FTAOA can strike a good balance between local solution and global exploration and achieve a satisfactory result within a short period of time.
Wenzhong Guo, Jie Li 0002, Yuzhen Niu, Chengyu Chen
IEEE Trans. Parallel Distributed Syst.4
2014 Direct manipulation video navigation on touch screens
abstract
Direct Manipulation Video Navigation (DMVN) systems allow a user to directly drag an object of interest along its motion trajectory and have been shown effective for space-centric video browsing tasks. This paper designs touch-based interface techniques to support DMVN on touchscreen devices. While touch screens can suit DMVN systems naturally and enhance the directness during video navigation, the fat finger problems, such as precise selection and occlusion handling, must be properly addressed. In this paper, we discuss the effect of the fat finger problems on DMVN and develop three touch-based object dragging techniques for DMVN on touch screens, namely Offset Drag, Window Drag, and Drag Anywhere. We conduct user studies to evaluate our techniques as well as two baseline solutions on a smartphone and a desktop touch screen. Our studies show that two of our techniques can support DMVN on touch screen devices well and perform better than the baseline solutions.
Cuong Nguyen 0003, Yuzhen Niu, Feng Liu 0015
Mobile HCI2
2014 Interpolation-tuned salient region detection
Yang Liu 0009, Lei Wang 0184, Yuzhen Niu
Sci. China Inf. Sci.4
2014 Oscillation analysis for salient object detection
Yang Liu 0009, Lei Wang 0184, Yuzhen Niu, Feng Liu 0015
Multim. Tools Appl.4
2014 Fast Gaussian kernel learning for classification tasks based on specially structured global optimization
Shangping Zhong, Tianshun Chen, Fengying He, Yuzhen Niu
Neural Networks4
2014 Geometry-shader-based real-time voxelization and applications
Hsu-Huai Chang, Yu-Chi Lai, Chin-Yuan Yao, Kai-Lung Hua, Yuzhen Niu, Feng Liu 0015
Vis. Comput.5
2013 Direct manipulation video navigation in 3D
abstract
Direct Manipulation Video Navigation (DMVN) systems allow a user to navigate a video by dragging an object along its motion trajectory. These systems have been shown effective for space-centric video browsing. Their performance, however, is often limited by temporal ambiguities in a video with complex motion, such as recurring motion, self-intersecting motion, and pauses. The ambiguities come from reducing the 3D spatial-temporal motion (x, y, t) to the 2D spatial motion (x, y) in visualizing the motion and dragging the object. In this paper, we present a 3D DMVN system that maps the spatial-temporal motion (x, y, t) to 3D space (x, y, z) by mapping time t to depth z, visualizes the motion and video frame in 3D, and allows to navigate the video by spatial-temporally manipulating the object in 3D. We show that since our 3D DMVN system preserves all the motion information, it resolves the temporal ambiguities and supports intuitive navigation on challenging videos with complex motion.
Cuong Nguyen 0003, Yuzhen Niu, Feng Liu 0015
CHI2
2013 Saliency Aggregation: A Data-Driven Approach
abstract
A variety of methods have been developed for visual saliency analysis. These methods often complement each other. This paper addresses the problem of aggregating various saliency analysis methods such that the aggregation result outperforms each individual one. We have two major observations. First, different methods perform differently in saliency analysis. Second, the performance of a saliency analysis method varies with individual images. Our idea is to use data-driven approaches to saliency aggregation that appropriately consider the performance gaps among individual methods and the performance dependence of each method on individual images. This paper discusses various data-driven approaches and finds that the image-dependent aggregation method works best. Specifically, our method uses a Conditional Random Field (CRF) framework for saliency aggregation that not only models the contribution from individual saliency map but also the interaction between neighboring pixels. To account for the dependence of aggregation on an individual image, our approach selects a subset of images similar to the input image from a training data set and trains the CRF aggregation model only using this subset instead of the whole training set. Our experiments on public saliency benchmarks show that our aggregation method outperforms each individual saliency method and is robust with the selection of aggregated methods.
Long Mai, Yuzhen Niu, Feng Liu 0015
CVPR2
2013 Joint Subspace Stabilization for Stereoscopic Video
abstract
Shaky stereoscopic video is not only unpleasant to watch but may also cause 3D fatigue. Stabilizing the left and right view of a stereoscopic video separately using a monocular stabilization method tends to both introduce undesirable vertical disparities and damage horizontal disparities, which may destroy the stereoscopic viewing experience. In this paper, we present a joint subspace stabilization method for stereoscopic video. We prove that the low-rank subspace constraint for monocular video [10] also holds for stereoscopic video. Particularly, the feature trajectories from the left and right video share the same subspace. Based on this proof, we develop a stereo subspace stabilization method that jointly computes a common subspace from the left and right video and uses it to stabilize the two videos simultaneously. Our method meets the stereoscopic constraints without 3D reconstruction or explicit left-right correspondence. We test our method on a variety of stereoscopic videos with different scene content and camera motion. The experiments show that our method achieves high-quality stabilization for stereoscopic video in a robust and efficient way.
Feng Liu 0015, Yuzhen Niu, Hailin Jin
ICCV2
2013 Making stereo photo cropping easy
abstract
The increasing popularity of stereoscopic 3D brings the demand for tools for editing and authoring stereoscopic images and videos. This paper shows that even a simple task like cropping is difficult for amateur users with little stereoscopic photography knowledge. Unlike regular monocular (2D) images, cropping a stereoscopic image needs to be carefully executed to avoid stereoscopic violations, which otherwise cause an unpleasant stereoscopic viewing experience. In this paper, we present a system that assists in stereoscopic photo cropping by automatically measuring the stereoscopic photography violations and alerting users with the potential violations. Our study shows that compared to a popular stereoscopic photo editing system, our system makes stereoscopic photo cropping easier even for amateur users with little stereoscopic photography knowledge and provides a good user experience.
Yuzhen Niu, Feng Liu 0015
ICME2
2013 Casual Stereoscopic Photo Authoring
abstract
Stereoscopic 3D displays become more and more popular these years. However, authoring high-quality stereoscopic 3D content remains challenging. In this paper, we present a method for easy stereoscopic photo authoring with a regular (monocular) camera. Our method takes two images or video frames using a monocular camera as input and transforms them into a stereoscopic image pair that provides a pleasant viewing experience. The key technique of our method is a perceptual-plausible image rectification algorithm that warps the input image pairs to meet the stereoscopic geometric constraint while avoiding noticeable visual distortion. Our method uses spatially-varying mesh-based image warps. Our warping method encodes a variety of constraints to best meet the stereoscopic geometric constraint and minimize visual distortion. Since each energy term is quadratic, our method eventually formulates the warping problem as a quadratic energy minimization which is solved efficiently using a sparse linear solver. Our method also allows both local and global adjustments of the disparities, an important property for adapting resulting stereoscopic images to different viewing conditions. Our experiments demonstrate that our spatially-varying warping technique can better support casual stereoscopic photo authoring than existing methods and our results and user study show that our method can effectively use casually-taken photos to create high-quality stereoscopic photos that deliver a pleasant 3D viewing experience.
Feng Liu 0015, Yuzhen Niu, Hailin Jin
IEEE Trans. Multim.2
2012 Video summagator: an interface for video summarization and navigation
abstract
This paper presents Video Summagator (VS), a volume-based interface for video summarization and navigation. VS models a video as a space-time cube and visualizes the video cube using real-time volume rendering techniques. VS empowers a user to interactively manipulate the video cube. We show that VS can quickly summarize both the static and dynamic video content by visualizing the space-time information in 3D. We demonstrate that VS enables a user to quickly look into the video cube, understand the content, and navigate to the content of interest.
Cuong Nguyen 0003, Yuzhen Niu, Feng Liu 0015
CHI2
2012 Leveraging stereopsis for saliency analysis
abstract
Stereopsis provides an additional depth cue and plays an important role in the human vision system. This paper explores stereopsis for saliency analysis and presents two approaches to stereo saliency detection from stereoscopic images. The first approach computes stereo saliency based on the global disparity contrast in the input image. The second approach leverages domain knowledge in stereoscopic photography. A good stereoscopic image takes care of its disparity distribution to avoid 3D fatigue. Particularly, salient content tends to be positioned in the stereoscopic comfort zone to alleviate the vergence-accommodation conflict. Accordingly, our method computes stereo saliency of an image region based on the distance between its perceived location and the comfort zone. Moreover, we consider objects popping out from the screen salient as these objects tend to catch a viewer's attention. We build a stereo saliency analysis benchmark dataset that contains 1000 stereoscopic images with salient object masks. Our experiments on this dataset show that stereo saliency provides a useful complement to existing visual saliency analysis and our method can successfully detect salient content from images that are difficult for monocular saliency analysis methods.
Yuzhen Niu, Yujie Geng, Feng Liu 0015
CVPR1
2012 Detecting rule of simplicity from photos
abstract
Simplicity refers to one of the most important photography composition rules. Simplicity states that simplifying the image background can draw viewers' attention to the subject of interest in a photograph and help them better comprehend and appreciate it. Understanding whether a photo respects photography rules or not facilitates photo quality assessment. In this paper, we present a method to automatically detect whether a photo is composed according to the rule of simplicity. We design features according to the definition, implementation and effect of the rule. First, we make use of saliency analysis to infer the subject of interest in a photo and measure its compactness. Second, we segment an image into background and foreground and measure the homogeneity within the background as another feature. Third, when looking at an image created with the rule of simplicity, different viewers tend to agree on what the subject of interest is in this photo. We accordingly measure the consistency among various saliency detection results as a feature. We experiment with these features in a range of machine learning methods. Our experiments show that our methods, together with these features, provide an encouraging result in detecting the rule of simplicity in a photo.
Long Mai, Hoang Le, Yuzhen Niu, Yu-Chi Lai, Feng Liu 0015
ACM Multimedia3
2012 Image resizing via non-homogeneous warping
Yuzhen Niu, Feng Liu 0015, Michael Gleicher
Multim. Tools Appl.1
2012 What Makes a Professional Video? A Computational Aesthetics Approach
abstract
Understanding the characteristics of high-quality professional videos is important for video classification, video quality measurement, and video enhancement. A professional video is good not only for its interesting story but also for its high visual quality. In this paper, we study what makes a professional video from the perspective of aesthetics. We discuss how a professional video is created and correspondingly design a variety of features that distinguish professional videos from amateur ones. We study general aesthetics features that are applied to still photos and extend them to videos. We design a variety of features that are particularly relevant to videos. We examined the performance of these features in the problem of professional and amateur video classification. Our experiments show that with these features, 97.3% professional and amateur shot classification accuracy rate is achieved on our own data set and 91.2% professional video detection rate is achieved on a public professional video set. Our experiments also show that the features that are particularly for videos are shown most effective for this task.
Yuzhen Niu, Feng Liu 0015
IEEE Trans. Circuits Syst. Video Technol.1
2012 Aesthetics-Based Stereoscopic Photo Cropping for Heterogeneous Displays
abstract
Stereoscopic displays are becoming ubiquitous, ranging from large 3-D TVs to small mobile phones. Stereoscopic photos need to be carefully adapted to be effectively viewed on the displays other than originally intended. In this paper, we present a method that can automatically crop and scale an existing stereoscopic photo to a variety of displays while preserving its aesthetic value. We formulate stereoscopic photo adaptation as an optimization problem that aims to preserve the aesthetic value of the input photo. We define a wide range of energy terms to preserve the stereoscopic photo aesthetics by borrowing rules from stereoscopic photography. Our experiments on a wide variety of stereoscopic photos demonstrate that our method can robustly produce display-dependent stereoscopic photos that deliver pleasant viewing experiences.
Yuzhen Niu, Feng Liu 0015, Wu-chi Feng, Hailin Jin
IEEE Trans. Multim.1
2012 Enabling warping on stereoscopic images
abstract
Warping is one of the basic image processing techniques. Directly applying existing monocular image warping techniques to stereoscopic images is problematic as it often introduces vertical disparities and damages the original disparity distribution. In this paper, we show that these problems can be solved by appropriately warping both the disparity map and the two images of a stereoscopic image. We accordingly develop a technique for extending existing image warping algorithms to stereoscopic images. This technique divides stereoscopic image warping into three steps. Our method first applies the user-specified warping to one of the two images. Our method then computes the target disparity map according to the user specified warping. The target disparity map is optimized to preserve the perceived 3D shape of image content after image warping. Our method finally warps the other image using a spatially-varying warping method guided by the target disparity map. Our experiments show that our technique enables existing warping methods to be effectively applied to stereoscopic images, ranging from parametric global warping to non-parametric spatially-varying warping.
Yuzhen Niu, Wu-chi Feng, Feng Liu 0015
ACM Trans. Graph.1
2011 Rule of Thirds Detection from Photograph
abstract
The rule of thirds is one of the most important composition rules used by photographers to create high-quality photos. The rule of thirds states that placing important objects along the imagery thirds lines or around their intersections often produces highly aesthetic photos. In this paper, we present a method to automatically determine whether a photo respects the rule of thirds. Detecting the rule of thirds from a photo requires semantic content understanding to locate important objects, which is beyond the state of the art. This paper makes use of the recent saliency and generic objectness analysis as an alternative and accordingly designs a range of features. Our experiment with a variety of saliency and generic objectness methods shows that an encouraging performance can be achieved in detecting the rule of thirds from photos.
Long Mai, Hoang Le, Yuzhen Niu, Feng Liu 0015
ISM3
2011 Systems support for stereoscopic video compression
abstract
In this paper, we propose a content-based, threaded stereoscopic video compression algorithm. We believe that it is likely future stereoscopic imaging systems will contain more than two lenses, allowing the display system to optimize the stereoscopic viewing experience. Furthermore, it will allow for the adjustment of framing (scene) composition errors that can arise from stereoscopic capture. Our proposed system uses 10 linearly aligned lenses. Upon capture, the images are run through a feature detection and matching algorithm in order to determine disparity between the stereoscopic images. During compression, the disparity measurements can be used to drive the selection of key frames within the image sets to provide better retrieval of data for display.
Wu-chi Feng, Feng Liu 0015, Yuzhen Niu, Scott Price
NOSSDAV3
2010 Warp propagation for video resizing
abstract
This paper presents a video resizing approach that provides both efficiency and temporal coherence. Prior approaches either sacrifice temporal coherence (resulting in jitter), or require expensive spatio-temporal optimization. By assessing the requirements for video resizing we observe a fundamental tradeoff between temporal coherence in the background and shape preservation for the moving objects. Understanding this tradeoff enables us to devise a novel approach that is efficient, because it warps each frame independently, yet can avoid introducing jitter. Like previous approaches, our method warps frames so that the background are distorted similarly to prior frames while avoiding distortion of the moving objects. However, our approach introduces a motion history map that propagates information about the moving objects between frames, allowing for graceful tradeoffs between temporal coherence in the background and shape preservation for the moving objects. The approach can handle scenes with significant camera and object motion and avoid jitter, yet warp each frame sequentially for efficiency. Experiments with a variety of videos demonstrate that our approach can efficiently produce high-quality video resizing results.
Yuzhen Niu, Feng Liu 0015, Michael Gleicher
CVPR1
2010 Animation rendering with Population Monte Carlo image-plane sampler
Yu-Chi Lai, Stephen Chenney, Feng Liu 0015, Yuzhen Niu, Shaohua Fan
Vis. Comput.4
2009 Using Web Photos for Measuring Video Frame Interestingness
Feng Liu 0015, Yuzhen Niu, Michael Gleicher
IJCAI2