Nana Yu

dblp:185/9890 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An evolutionary algorithm with neighborhood structure incorporating reinforcement learning for dual resource constrained flexible job shop scheduling problem with worker transportation fatigue
Guohui Zhang 0002, Zhixiao Li, Nana Yu
Expert Syst. Appl.4
2026 Enhanced visual prompt meets low-light saliency detection
Nana Yu, Jie Wang 0095, Yahong Han
Pattern Recognit.1
2026 Low-Light Salient Object Detection via Representation Decoupling
abstract
In low-light scenes, images often suffer low contrast, poor distinction between the target and the background, and loss of regional information. This makes it challenging for Salient Object Detection (SOD) algorithms to identify and locate the targets. Most existing methods address low-light SOD by first enhancing low-light images and then performing saliency detection. However, splitting these into two sub-tasks may result in negative transfer effects. Additionally, some methods attempt to integrate these two sub-tasks into an end-to-end framework. However, conflicts may arise during the training process due to the differing features required by low-light enhancement and saliency detection. To address this conflict, We propose a decoupled representation network called DRNet, whose core is the construction of a Visual Center Decoupler (VCD). The VCD decouples the learned representations into enhancement-specific and SOD-specific embeddings. This module provides an implicit way to balance the specific requirements of the two subtasks. On the one hand, the enhancement-specific embeddings use RGB illumination constraints and Local Binary Patterns (LBP) feature aggregation to constrain illumination and maintain texture feature stability. On the other hand, the SOD-specific embeddings utilize dynamic multi-scale convolutions to integrate fine-grained details required at multiple scales. Finally, to validate the performance advantages of DRNet, we conduct extensive comparative experiments between the proposed DRNet and existing single-modal methods. Comparative experiments with some representative bi-modal methods further demonstrate the merits of our proposed method.
Nana Yu, Jie Wang 0095, Yahong Han
IEEE Trans. Circuits Syst. Video Technol.1
2025 Visual Consensus Prompting for Co-Salient Object Detection
abstract
Existing co-salient object detection (CoSOD) methods generally employ a three-stage architecture (i.e., encoding, consensus extraction & dispersion, and prediction) along with a typical full fine-tuning paradigm. Although they yield certain benefits, they exhibit two notable limitations: 1) This architecture relies on encoded features to facilitate consensus extraction, but the meticulously extracted consensus does not provide timely guidance to the encoding stage. 2) This paradigm involves globally updating all parameters of the model, which is parameter-inefficient and hinders the effective representation of knowledge within the foundation model for this task. Therefore, in this paper, we propose an interaction-effective and parameter-efficient concise architecture for the CoSOD task, addressing two key limitations. It introduces, for the first time, a parameter-efficient prompt tuning paradigm and seamlessly embeds consensus into the prompts to formulate task-specific Visual Consensus Prompts (VCP). Our VCP aims to induce the frozen foundation model to perform better on CoSOD tasks by formulating task-specific visual consensus prompts with minimized tunable parameters. Concretely, the primary insight of the purposeful Consensus Prompt Generator (CPG) is to enforce limited tunable parameters to focus on co-salient representations and generate consensus prompts. The formulated Consensus Prompt Disperser (CPD) leverages consensus prompts to form task-specific visual consensus prompts, thereby arousing the powerful potential of pre-trained models in addressing CoSOD tasks. Extensive experiments demonstrate that our concise VCP outperforms 13 cutting-edge full fine-tuning models, achieving the new state of the art (with 6.8% improvement in Fmmetrics on the most challenging CoCA dataset). Source code has been available at https://github.com/WJ-CV/VCP.
Jie Wang 0095, Nana Yu, Yahong Han
CVPR2
2025 Leveraging LLMs to Improve Human Annotation Efficiency with INCEpTION
Luís Filipe Cunha, Nana Yu, Purificação Silvano, Ricardo Campos 0001, Alípio Mário Jorge
ECIR (5)2
2025 Progressive expansion for semi-supervised bi-modal salient object detection
Jie Wang 0095, Nana Yu, Yahong Han
Pattern Recognit.3
2025 Explicitly Disentangling and Exclusively Fusing for Semi-Supervised Bi-Modal Salient Object Detection
abstract
Bi-modal (RGB-T and RGB-D) salient object detection (SOD) aims to enhance detection performance by leveraging the complementary information between modalities. While significant progress has been made, two major limitations persist. Firstly, mainstream fully supervised methods come with a substantial burden of manual annotation, while weakly supervised or unsupervised methods struggle to achieve satisfactory performance. Secondly, the indiscriminate modeling of local detailed information (object edge) and global contextual information (object body) often results in predicted objects with incomplete edges or inconsistent internal representations. In this work, we propose a novel paradigm to effectively alleviate the above limitations. Specifically, we first enhance the consistency regularization strategy to build a basic semi-supervised architecture for the bi-modal SOD task, which ensures that the model can benefit from massive unlabeled samples while effectively alleviating the annotation burden. Secondly, to ensure detection performance (i.e., complete edges and consistent bodies), we disentangle the SOD task into two parallel sub-tasks: edge integrity fusion prediction and body consistency fusion prediction. Achieving these tasks involves two key steps: 1) the explicitly disentangling scheme decouples salient object features into edge and body features, and 2) the exclusively fusing scheme performs exclusive integrity or consistency fusion for each of them. Eventually, our approach demonstrates significant competitiveness compared to 26 fully supervised methods, while effectively alleviating 90% of the annotation burden. Furthermore, it holds a substantial advantage over 15 non-fully supervised methods.
Jie Wang 0095, Xiangji Kong, Nana Yu, Yahong Han
IEEE Trans. Circuits Syst. Video Technol.3
2025 Single-Group Generalized RGB and RGB-D Co-Salient Object Detection
abstract
Co-salient object detection (CoSOD) aims to segment the co-occurring salient objects in a given group of relevant images. Existing methods typically rely on extensive group training data to enhance the model’s CoSOD capabilities. However, fitting prior knowledge of the extensive group results in a significant performance gap between the seen and out-of-sample image groups. Relaxing such a fitting with fewer prior groups may improve the generalization ability of CoSOD while alleviating the annotation burdens. Hence, it is essential to explore the use of fewer groups during the training phase, such as using only single group, to pursue a highly generalized CoSOD model. We term this new setting as Sg-CoSOD, which aims to train a model using only a single group and effectively apply it to any unseen RGB and RGB-D CoSOD test groups. Towards Sg-CoSOD, it is important to ensure detection performance with limited data and release class dependency with only a single-group. Thus, we present a method, i.e., cross-excitation between saliency and ‘Co’, which decouples the CoSOD task into two parallel branches: ‘Co’ To Saliency (CTS) and Saliency To ‘Co’ (STC). The CTS branch focuses on mining group consensus to guide image co-saliency predictions, while the STC branch is dedicated to using saliency priors to motivate group consensus mining. Furthermore, we propose a Class-Agnostic Triplet (CAT) loss to constrain intra-group consensus while suppressing the model from acquiring class prior knowledge. Extensive experiments on RGB and RGB-D CoSOD tasks with multiple unknown groups show that our model has higher generalization capabilities (e.g., for large-scale datasets CoSOD3k and CoSal1k with multiple generalized groups, we obtain a gain of over 15% in$F_{m}$). Further experimental analyses also reveal that the proposed Sg-CoSOD paradigm has significant potential and promising prospects.
Jie Wang 0095, Nana Yu, Yahong Han
IEEE Trans. Circuits Syst. Video Technol.2
2025 Semantic Prompt Enhancement for Semi-Supervised Low-Light Salient Object Detection
abstract
Most existing salient object detection (SOD) models are designed based on data collected in well-lit scenes, which is entirely inadequate for low-light conditions. Although recent models are designed for low-light conditions, they still have limitations. First, they simply integrate features without considering the impact of low-light scenes and fail to enhance the contextual information around salient objects. Second, in extremely dark scenes, it is difficult for the human eye to distinguish between the foreground and background, posing significant challenges for data labeling. To address these issues, we design a brightness Retinex enhancer (BRE) tailored for low-light SOD tasks and, for the first time, explore performing low-light SOD within a semi-supervised framework. By using sparse labeled semantic prompts to augment a large amount of unlabeled data, we mitigate the annotation burden while avoiding ineffective labeling in low-light conditions. More specifically, we first use Retinex decomposition to filter out the influence of illumination, while the semantic features extracted by a large model serve as semantic prompts to assist in enhancement. In addition, we introduce a context-guided encoder (CGE) to improve the model's understanding of salient objects. Finally, both labeled and unlabeled data undergo joint consistency training between the shared decoder (SD) and the perturbation decoder. The semi-supervised model enhances low-light SOD performance while also alleviating the burden of data annotation. Extensive experiments demonstrate that, compared with state-of-the-art fully supervised SOD models, the proposed semi-supervised model achieves highly competitive results across multiple test datasets.
Nana Yu, Jie Wang 0095, Yahong Han, Weiping Ding 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Degradation-removed multiscale fusion for low-light salient object detection
Nana Yu, Jie Wang 0095, Yahong Han
Pattern Recognit.1
2024 Depth-Assisted Semi-Supervised RGB-D Rail Surface Defect Inspection
abstract
Visual-based methods for rail surface defect inspection (RSDI) effectively improve the limitations of manual inspection, as they can intuitively display the locations and segmented areas of sensitive defects. The RGB-D RSDI task, which leverages the complementarity between RGB and depth (D) image information to enhance detection performance, has attracted widespread attention and achieved significant development. However, existing methods primarily depend on fully supervised training strategies that necessitate a substantial number of manually annotated pixel-level labels to supervise model training. Undoubtedly, extensive manual annotation is exceedingly time-consuming and labor-intensive, particularly considering the irregular shapes and textures of surface defects on rails, further compounding the burden of manual labeling. Therefore, in this paper, we aim to introduce the semi-supervised learning paradigm into this task. Towards the semi-supervised RGB-D RSDI task, a specific semi-supervised network for this task and an effective cross-modal fusion module are crucial to ensuring detection performance under the constraints of limited labeled samples. Thus, we propose a Depth-assisted Semi-Supervised RGB-D RSDI network (DSSNet) to simultaneously alleviate the annotation burden and achieve satisfactory detection performance. Specifically, adhering to the consistency training paradigm, we construct a semi-supervised RGB-D RSDI architecture for this task by optimizing structures, perturbation mechanisms, loss settings, etc. Furthermore, we propose a Depth-assisted Multi-scale Cross-modal Fusion Module (DMCFM) that conducts multi-scale exploration and cross-modal complementary fusion with the assistance of depth. Comprehensive experiments demonstrate that, compared to the latest 14 state-of-the-art fully supervised methods, the proposed DSSNet achieves highly competitive results while effectively alleviating an 80$\%$annotation burden.
Jie Wang 0095, Guanwen Qiu, Jinwen Xi, Nana Yu
IEEE Trans. Intell. Transp. Syst.6
2024 Joint Correcting and Refinement for Balanced Low-Light Image Enhancement
abstract
Low-light image enhancement tasks demand an appropriate balance among brightness, color, and illumination. While existing methods often focus on one aspect of the image without considering how to pay attention to this balance, which will cause problems of color distortion and overexposure etc. This seriously affects both human visual perception and the performance of high-level visual models. In this work, a novel synergistic structure is proposed which can balance brightness, color, and illumination more effectively. Specifically, the proposed method, so-called Joint Correcting and Refinement Network (JCRNet), which mainly consists of three stages to balance brightness, color, and illumination of enhancement. Stage 1: we utilize a basic encoder-decoder and local supervision mechanism to extract local information and more comprehensive details for enhancement. Stage 2: cross-stage feature transmission and spatial feature transformation further facilitate color correction and feature refinement. Stage 3: we employ a dynamic illumination adjustment approach to embed residuals between predicted and ground truth images into the model, adaptively adjusting illumination balance. Extensive experiments demonstrate that the proposed method exhibits comprehensive performance advantages over 21 state-of-the-art methods on 9 benchmark datasets. Furthermore, a more persuasive experiment has been conducted to validate our approach the effectiveness in downstream visual tasks (e.g., saliency detection). Compared to several enhancement models, the proposed method effectively improves the segmentation results and quantitative metrics of saliency detection.
Nana Yu, Yahong Han
IEEE Trans. Multim.1
2023 FLA-Net: multi-stage modular network for low-light image enhancement
Nana Yu, Jinjiang Li 0001, Zhen Hua
Vis. Comput.1
2022 LBP-based progressive feature aggregation network for low-light image enhancement
abstract
Abstract At night or in other low‐illumination environments, optical imaging devices cannot capture details and color information in images accurately because of the reduced number of photons captured and the low signal‐to‐noise ratio. Consequently, the image is very noisy with low contrast and inaccurate color information, which affects human visual perception and creates significant challenges in computer vision tasks. Low‐light image enhancement has great research value because it aims to reduce image noise and improve image quality. In this study, we propose an LBP‐based progressive feature aggregation network (P‐FANet) for low‐light image enhancement. The LBP feature has insensitivity to illumination, and it contains rich texture information. In the network, we input the LBP feature into each iteration of the network in an accompanying manner, which helps to restore some detailed information of the low‐light image. First, we input the low‐light image into the dual attention mechanism model to extract global features. Second, the extracted different features enter the feature aggregation module (FAM) for feature fusion. Third, we use the recurrent layer to share the features extracted at different stages, and use the residual layer to further extract deeper features. Finally, the enhanced image is output. The rationality of the method in this study has been verified through ablation experiments. Many experimental results show that the method in this study has greater advantages in subjective and objective evaluations compared with many other advanced methods.
Nana Yu, Jinjiang Li 0001, Zhen Hua
IET Image Process.1
2022 Detail enhancement decolorization algorithm based on rolling guided filtering
Nana Yu, Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.1
2022 Attention based dual path fusion networks for multi-focus image
Nana Yu, Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.1
2022 Decolorization algorithm based on contrast pyramid transform fusion
Nana Yu, Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.1
2019 Using Dictionary Pair Learning for Seizure Detection
abstract
Automatic seizure detection is extremely important in the monitoring and diagnosis of epilepsy. The paper presents a novel method based on dictionary pair learning (DPL) for seizure detection in the long-term intracranial electroencephalogram (EEG) recordings. First, for the EEG data, wavelet filtering and differential filtering are applied, and the kernel function is performed to make the signal linearly separable. In DPL, the synthesis dictionary and analysis dictionary are learned jointly from original training samples with alternating minimization method, and sparse coefficients are obtained by using of linear projection instead of costly [Formula: see text]-norm or [Formula: see text]-norm optimization. At last, the reconstructed residuals associated with seizure and nonseizure sub-dictionary pairs are calculated as the decision values, and the postprocessing is performed for improving the recognition rate and reducing the false detection rate of the system. A total of 530[Formula: see text]h from 20 patients with 81 seizures were used to evaluate the system. Our proposed method has achieved an average segment-based sensitivity of 93.39%, specificity of 98.51%, and event-based sensitivity of 96.36% with false detection rate of 0.236/h.
Nana Yu
Int. J. Neural Syst.2