EDBT 2026 Demo / reviewers in the wild / expert
Keren Fu
dblp:126/7553
· DBLP profile ↗
93ranked-venue papers
22as first author
43since 2021 · last 2026
0000-0002-3195-2077ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 14 first-author · 28 since 2021Artificial intelligence and machine learning · 45 · 12 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DiAPR: Dimensionally-Allocated Prototype Refinement for Non-Exemplar Class Incremental LearningabstractNon-Exemplar Class Incremental Learning (NECIL) strives to preserve classification performance in an evolving data stream without revisiting old-class exemplars. Current methods mitigate catastrophic forgetting by replaying and augmenting historical prototypes as surrogates for old classes. However, they treat prototypes as holistic representations for global-level augmentations, which overlook dimensional semantic disparity and old-new class relationships, failing to maintain old-class discriminability and adaptability to the evolving feature space. To address this challenge, we propose Dimensionally-Allocated Prototype Refinement (DiAPR), a granular framework that progressively refines prototypes to exhibit class separability in the new feature space through three modules. Specifically, Distribution-aware Pairing (DAP) captures old-new class semantic consistency to guide Granular Semantic Allocation (GSA) in dimension-wise conflation, while Cross-Dimensional Transition (CDT) enhances cross-dimensional dependencies. The resulting prototypes sharpen classifier decision boundaries. Moreover, CDT inherently enables softened feature alignment, thereby yielding a more compatible feature space. Extensive experiments demonstrate DiAPR’s superiority, with improvements over SOTA by 2.35%, 0.70%, 0.96% on three CIFAR-100 settings, 1.03%, 0.54%, 0.40% on Tiny-ImageNet, and 0.60% on ImageNet-Subset. Ruixuan Gao, Qijun Zhao, Keren Fu |
AAAI | 3 |
| 2026 | Unleashing the power of motion and depth: A selective fusion strategy for RGB-D video salient object detection
Daerji Suolang, Keren Fu, Qijun Zhao |
Knowl. Based Syst. | 3 |
| 2025 | Samba: A Unified Mamba-based Framework for General Salient Object DetectionabstractExisting salient object detection (SOD) models primarily resort to convolutional neural networks (CNNs) and Transformers. However, the limited receptive fields of CNNs and quadratic computational complexity of transformers both constrain the performance of current models on discovering attention-grabbing objects. The emerging state space model, namely Mamba, has demonstrated its potential to balance global receptive fields and computational complexity. Therefore, we propose a novel unified framework based on the pure Mamba architecture, dubbed saliency Mamba (Samba), to flexibly handle general SOD tasks, including RGB/RGB-D/RGB-T SOD, video SOD (VSOD), and RGB-D VSOD. Specifically, we rethink Mamba’s scanning strategy from the perspective of SOD, and identify the importance of maintaining spatial continuity of salient patches within scanning sequences. Based on this, we propose a saliency-guided Mamba block (SGMB), incorporating a spatial neighboring scanning (SNS) algorithm to preserve spatial continuity of salient patches. Additionally, we propose a context-aware upsampling (CAU) method to promote hierarchical feature alignment and aggregations by modeling contextual dependencies Experimental results show that our Samba outperforms existing methods across five SOD tasks on 21 datasets with lower computational cost, confirming the superiority of introducing Mamba to the SOD areas. Our code is available at https://github.com/Jia-Hao999/Samba. Keren Fu, Xiaohong Liu 0001, Qijun Zhao |
CVPR | 2 |
| 2025 | MoEdit: On Learning Quantity Perception for Multi-object Image EditingabstractMulti-object images are prevalent in various real-world scenarios, including augmented reality, advertisement design, and medical imaging. Efficient and precise editing of these images is critical for these applications. With the advent of Stable Diffusion (SD), high-quality image generation and editing have entered a new era. However, existing methods often struggle to consider each object both individually and part of the whole image editing, both of which are crucial for ensuring consistent quantity perception, resulting in suboptimal perceptual performance. To address these challenges, we propose MoEdit, an auxiliaryfree multi-object image editing framework. MoEdit facilitates high-quality multi-object image editing in terms of style transfer, object reinvention, and background regeneration, while ensuring consistent quantity perception between inputs and outputs, even with a large number of objects. To achieve this, we introduce the Feature Compensation (FeCom) module, which ensures the distinction and separability of each object attribute by minimizing the in-between interlacing. Additionally, we present the Quantity Attention (QTTN) module, which perceives and preserves quantity consistency by effective control in editing, without relying on auxiliary tools. By leveraging the SD model, MoEdit enables customized preservation and modification of specific concepts in inputs with high quality. Experimental results demonstrate that our MoEdit achieves State-Of-The-Art (SOTA) performance in multi-object image editing. Data and codes are available at https://github.com/Tear-kitty/MoEdit. Ka-Hou Chan, Yue Sun 0001, Chan-Tong Lam, Tong Tong 0001, Zitong Yu, Keren Fu, Xiaohong Liu 0001, Tao Tan 0002 |
CVPR | 7 |
| 2025 | Camouflaged Object Detection via Neural Architecture SearchabstractThe core challenge in camouflaged object detection (COD) is identifying objects that blend seamlessly with their surroundings. Existing methods emulate the strategies biological organisms break camouflage by manually constructing modules with expert knowledge from existing segmentation tasks, making it difficult to accurately understand complex and unique camouflage semantics. We are the first to apply neural architecture search (NAS) to COD, introducing an automatic localization and refinement network called ALRNet. It explores a large search space to discover more effective camouflage-specific modules. Specifically, we propose a search-based automatic receptive field block (ARFB) to adaptively excavate hierarchical discriminative cues and decouple features in a multi-branch architecture. Moreover, we introduce an edge-assisted explicit and implicit refinement (EEIR) module, combining explicit priors with implicit search to create a dual-task structure for edge and segmentation knowledge interaction. Experiments on four benchmarks demonstrate that ALRNet outperforms 15 state-of-the-art methods. Codes are available at https://github.com/BoydeLi/ALRNet. Keren Fu, Qijun Zhao |
ICASSP | 2 |
| 2025 | Lightweight Multi-Frequency Enhancement Network for RGB-D Video Salient Object DetectionabstractRGB-D Video Salient Object Detection has gained increasing interest, but existing models often struggle to balance efficiency and accuracy, hindering their applications on resource-constrained devices. A key challenge in designing lightweight models is maintaining accuracy while reducing parameters. To address this issue and bridge the gap in lightweight RGB-D VSOD research, we propose a lightweight network architecture using MobileNetV2 as the backbone. We introduce an Improved Cross-Shift Module (ICSM) to extract the fused depth and flow features with minimal overhead and a Multi-Frequency Enhancement Module (MFEM) to separate high-and low-frequency information and enhance the resulting feature maps using different techniques for final saliency prediction. Experimental results demonstrate that our method achieves competitive accuracy compared to non-efficient models, running at 80 FPS on a GPU with only 4.75M parameters, making it suitable for real-time applications. Code will be available at https://github.com/Tibetsonam/MFENet. Daerji Suolang, Wangchuk Tsering, Keren Fu, Qijun Zhao |
ICASSP | 4 |
| 2025 | Promoting Segment Anything Model towards Highly Accurate Dichotomous Image SegmentationabstractThe Segment Anything Model (SAM) represents a significant breakthrough into foundation models for computer vision, providing a large-scale image segmentation model. However, despite SAM’s zero-shot performance, its segmentation masks lack fine-grained details, particularly in accurately delineating object boundaries. Therefore, it is both interesting and valuable to explore whether SAM can be improved towards highly accurate object segmentation, which is known as the dichotomous image segmentation (DIS) task. To address this issue, we propose DIS-SAM, which advances SAM towards DIS with extremely accurate details. DIS-SAM is a framework specifically tailored for highly accurate segmentation, maintaining SAM’s promptable design. DIS-SAM employs a two-stage approach, integrating SAM with a modified advanced network that was previously designed to handle the prompt-free DIS task. To better train DIS-SAM, we employ a ground truth enrichment strategy by modifying original mask annotations. Despite its simplicity, DIS-SAM significantly advances the SAM, HQ-SAM, and Pi-SAM by ~8.5%, ~6.9%, and ~3.7% maximum F-measure. Our code at https://github.com/Tennine2077/DIS-SAM. Xianjie Liu, Keren Fu, Yao Jiang 0002, Qijun Zhao |
ICME | 2 |
| 2025 | Dynamic patch-aware enrichment transformer for occluded person re-identification
Xin Zhang 0125, Keren Fu, Qijun Zhao |
Knowl. Based Syst. | 2 |
| 2025 | Explicit Motion Handling and Interactive Prompting for Video Camouflaged Object DetectionabstractCamouflage poses notable challenges in distinguishing a static target, as it usually blends seamlessly with the background. However, any movement by the target can disrupt this disguise, making it detectable. Existing video camouflaged object detection (VCOD) approaches take noisy motion estimation as input or model motion implicitly, restricting detection performance in complex dynamic scenes. In this paper, we propose a novel Explicit Motion handling and Interactive Prompting framework for VCOD, dubbed EMIP, which handles motion cues explicitly using a frozen pre-trained optical flow fundamental model. EMIP is characterized by a two-stream architecture for simultaneously conducting camouflaged segmentation and optical flow estimation. Interactions across the dual streams are realized in an interactive prompting way that is inspired by emerging visual prompt learning. Two learnable modules, i.e. the camouflaged feeder and motion collector, are designed to incorporate segmentation-to-motion and motion-to-segmentation prompts, respectively, and enhance outputs of the both streams. The prompt fed to the motion stream is learned by supervising optical flow in a self-supervised manner. Furthermore, we show that long-term historical information can also be incorporated as a prompt into EMIP and achieve more robust results with temporal consistency. By leveraging promoting techniques based on EMIP, the proposed long-term model EMIP ${}^{\dagger }$ incurs lower training cost with only 8.5M trainable parameters (less than 8% of the total model parameters). Experimental results demonstrate that both EMIP and EMIP ${}^{\dagger }$ set new state-of-the-art records on popular VCOD benchmarks. Additionally, comparative evaluations against other video segmentation models on a wider range of video segmentation tasks demonstrate the robustness and superior generalization capabilities of EMIP. Our code is made publicly available at https://github.com/zhangxin06/EMIP. Xin Zhang 0125, Ge-Peng Ji, Keren Fu, Qijun Zhao |
IEEE Trans. Image Process. | 5 |
| 2024 | Uncertainty-Aware Sign Language Video Retrieval with Probability Distribution Modeling
Yuanjiang Luo, Xuxin Cheng, Xianwei Zhuang, Keren Fu |
ECCV (44) | 7 |
| 2024 | Depth-Aware Dual-Stream Interactive Transformer Network for Facial Expression Recognition
Yiben Jiang, Xiao Yang 0029, Keren Fu, Hongyu Yang 0002 |
PRCV (11) | 3 |
| 2024 | Identity-Preserving Animal Image Generation for Animal Individual Identification
Zongming Peng, Yangqianqian Chen, Keren Fu, Qijun Zhao |
PRCV (15) | 5 |
| 2024 | Key Object Detection: Unifying Salient and Camouflaged Object Detection Into One Task
Pengyu Yin, Keren Fu, Qijun Zhao |
PRCV (12) | 2 |
| 2024 | Species-Aware Guidance for Animal Action Recognition with Vision-Language Knowledge
Zhen Zhai, Qijun Zhao, Keren Fu |
PRCV (7) | 4 |
| 2024 | Prior Mask-Guided Highly Accurate Dichotomous Image Segmentation
Shanfeng Zhou, Keren Fu, Qijun Zhao |
PRICAI (4) | 3 |
| 2024 | LSTPNet: Long short-term perception network for dynamic facial expression recognition in the wild
Chengcheng Lu, Yiben Jiang, Keren Fu, Qijun Zhao, Hongyu Yang 0002 |
Image Vis. Comput. | 3 |
| 2024 | Parallax-Aware Network for Light Field Salient Object DetectionabstractMulti-view images capture scene details from different views, making them advantageous for light field salient object detection (LF SOD). However, most existing LF SOD methods neglect effective modeling and utilization of parallax information inherent in multi-view images. To address this limitation, we propose to explicitly model parallax information and conduct SOD in a parallax-aware manner, resulting in a novel network called PANet. Our model initiates by generating horizontal and vertical visual parallax maps from four border views using optical flow estimation. We then introduce a parallax-aware network, incorporating a parallax processing module (PPM) that handles both parallax quality assessment and parallax correction. In the parallax correction phase, we design a channel-based correction unit (CCU) and a graph-based correction unit (GCU) to rectify deviations of parallax features in a direction-specific manner. Additionally, a parallax supplement module (PSM) seamlessly fuses the parallax information from different directions and embeds it into the center view, thereby improving SOD accuracy. Experiments on three benchmark datasets demonstrate the superiority of our PANet model over 15 state-of-the-art models. Our code for the model will be publicly available soon. Yao Jiang 0002, Keren Fu, Qijun Zhao |
IEEE Signal Process. Lett. | 3 |
| 2024 | Transformer-Based Light Field Salient Object Detection and Its Application to AutofocusabstractExisting light field salient object detection (LFSOD) models predominantly rely on convolutional neural networks or local attention to process light field data, consequently encountering difficulties in modeling intra-slice and cross-slice long-range dependencies within focal stacks. In this paper, we ponder the feasibility of relying solely on the pure Transformer architecture to address this dilemma and propose a novel quasi-pure Transformer-based framework for LFSOD, termed TLFNet. TLFNet incorporates innovative Transformer-based fusion modules (PGFormer) along with an edge enhancement module. The PGFormer employs a perpendicular self-attention (PSA) mechanism to capture long-range dependencies along both cross-slice and intra-slice axes within the focal stack, and integrates multi-modal features using a guided feature fusion (GFF) module. To address the issue of blurry edges arising from the Transformer-based encoder-decoder architecture, the edge enhancement module combines detailed texture and body information and employs focal loss to improve the edge precision of salient objects. TLFNet is a nearly pure Transformer-based approach (with approximately 99.01% of its parameters belonging to the Transformer), while the edge enhancement module significantly boosts accuracy with only around 0.99% of parameters. Comprehensive benchmarks demonstrate that TLFNet outperforms 14 light field models and achieves new state-of-the-art performance. Last but not least, we show in this paper a new application scheme of TLFNet, by cooperating with the deep autofocus technique proposed in [1], leading to light field salient object autofocus (LFSOA). LFSOA aims to identify and output the focal slice with a salient object in focus while keeping other irrelevant background blurred (out-of-focus), yielding an autonomous bokeh effect in photography. The code for the model and application will be publicly available soon. Yao Jiang 0002, Keren Fu, Qijun Zhao |
IEEE Trans. Image Process. | 3 |
| 2024 | Salient Object Detection in RGB-D VideosabstractGiven the widespread adoption of depth-sensing acquisition devices, RGB-D videos and related data/media have gained considerable traction in various aspects of daily life. Consequently, conducting salient object detection (SOD) in RGB-D videos presents a highly promising and evolving avenue. Despite the potential of this area, SOD in RGB-D videos remains somewhat under-explored, with RGB-D SOD and video SOD (VSOD) traditionally studied in isolation. To explore this emerging field, this paper makes two primary contributions: the dataset and the model. On one front, we construct the RDVS dataset, a new RGB-D VSOD dataset with realistic depth and characterized by its diversity of scenes and rigorous frame-by-frame annotations. We validate the dataset through comprehensive attribute and object-oriented analyses, and provide training and testing splits. Moreover, we introduce DCTNet+, a three-stream network tailored for RGB-D VSOD, with an emphasis on RGB modality and treats depth and optical flow as auxiliary modalities. In pursuit of effective feature enhancement, refinement, and fusion for precise final prediction, we propose two modules: the multi-modal attention module (MAM) and the refinement fusion module (RFM). To enhance interaction and fusion within RFM, we design a universal interaction module (UIM) and then integrate holistic multi-modal attentive paths (HMAPs) for refining multi-modal low-level features before reaching RFMs. Comprehensive experiments, conducted on pseudo RGB-D video datasets alongside our proposed RDVS, highlight the superiority of DCTNet+ over 19 VSOD models and 14 RGB-D SOD models. Additionally, insightful ablation experiments were performed on both pseudo and realistic RGB-D video datasets to demonstrate the advantages of individual modules as well as the necessity of introducing realistic depth into VSOD. Our code together with RDVS dataset will be available at https://github.com/kerenfu/RDVS/. Ao Mou, Yukang Lu, Dingyao Min, Keren Fu, Qijun Zhao |
IEEE Trans. Image Process. | 5 |
| 2024 | Fusion-Embedding Siamese Network for Light Field Salient Object DetectionabstractLight field salient object detection (SOD) has shown remarkable success and gained considerable attention from the computer vision community. Existing methods usually employ a single-/two-stream network to detect saliency. However, these methods can only handle up to two different modalities at a time, preventing them from being able to fully explore the rich information in multi-modal light field derived data. To address this, we propose the first joint multi-modal learning framework, called FES-Net, for light field SOD, which can take rich inputs not limited to two modalities. Specifically, we propose an attention-aware adaptation module to first transform the multi-modal inputs for use in our joint learning framework. The transformed inputs are then fed to a Siamese network along with multiple embedded feature fusion modules to extract informative multi-modal features. Finally, we predict saliency maps from the high-level extracted features using a saliency decoder module. Our joint multi-modal learning framework effectively resolves the limitations of existing methods, providing efficient and effective multi-modal learning that can fully explore the valuable information in light field data for accurate saliency detection. Furthermore, we improve the performance by introducing the Transformer as our backbone network. To the best of our knowledge, the improved version of our model, called FES-Trans, is the first attempt to address the challenging light field SOD with the powerful Transformer technique. Extensive experiments on benchmark datasets demonstrate that our models are superior light field SOD approaches and outperform cutting-edge models remarkably. Geng Chen 0001, Huazhu Fu, Tao Zhou 0002, Guobao Xiao, Keren Fu, Yong Xia 0001, Yanning Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | 3-D Convolutional Neural Networks for RGB-D Salient Object Detection and BeyondabstractRGB-depth (RGB-D) salient object detection (SOD) recently has attracted increasing research interest, and many deep learning methods based on encoder-decoder architectures have emerged. However, most existing RGB-D SOD models conduct explicit and controllable cross-modal feature fusion either in the single encoder or decoder stage, which hardly guarantees sufficient cross-modal fusion ability. To this end, we make the first attempt in addressing RGB-D SOD through 3-D convolutional neural networks. The proposed model, named RD3D, aims at prefusion in the encoder stage and in-depth fusion in the decoder stage to effectively promote the full integration of RGB and depth streams. Specifically, RD3D first conducts prefusion across RGB and depth modalities through a 3-D encoder obtained by inflating 2-D ResNet and later provides in-depth feature fusion by designing a 3-D decoder equipped with rich back-projection paths (RBPPs) for leveraging the extensive aggregation ability of 3-D convolutions. Toward an improved model RD3D+, we propose to disentangle the conventional 3-D convolution into successive spatial and temporal convolutions and, meanwhile, discard unnecessary zero padding. This eventually results in a 2-D convolutional equivalence that facilitates optimization and reduces parameters and computation costs. Thanks to such a progressive-fusion strategy involving both the encoder and the decoder, effective and thorough interactions between the two modalities can be exploited and boost detection accuracy. As an additional boost, we also introduce channel-modality attention and its variant after each path of RBPP to attend to important features. Extensive experiments on seven widely used benchmark datasets demonstrate that RD3D and RD3D+ perform favorably against 14 state-of-the-art RGB-D SOD approaches in terms of five key evaluation metrics. Our code will be made publicly available at https://github.com/PPOLYpubki/RD3D. Yanye Lu, Keren Fu, Qijun Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Guided Focal Stack Refinement Network for Light Field Salient Object DetectionabstractLight field salient object detection (SOD) is an emerging research direction attributed to the richness of light field data. However, most existing methods lack effective handling of focal stacks, therefore making the latter involved in a lot of interfering information and degrade the performance of SOD. To address this limitation, we propose to utilize multi-modal features to refine focal stacks in a guided manner, resulting in a novel guided focal stack refinement network called GFRNet. To this end, we propose a guided refinement and fusion module (GRFM) to refine focal stacks and aggregate multi-modal features. In GRFM, all-in-focus (AiF) and depth modalities are utilized to refine focal stacks separately, leading to two novel sub-modules for different modalities, namely AiF-based refinement module (ARM) and depth-based refinement module (DRM). Such refinement modules enhance structural and positional information of salient objects in focal stacks, and are able to improve SOD accuracy. Experimental results on four benchmark datasets demonstrate the superiority of our GFRNet model against 12 state-of-the-art models. Yao Jiang 0002, Keren Fu, Qijun Zhao |
ICME | 3 |
| 2023 | Long Short-Term Perception Network for Dynamic Facial Expression Recognition
Chengcheng Lu, Yiben Jiang, Keren Fu, Qijun Zhao, Hongyu Yang 0002 |
PRCV (5) | 3 |
| 2023 | Full-duplex strategy for video object segmentationabstractPrevious video object segmentation approaches mainly focus on simplex solutions linking appearance and motion, limiting effective feature collaboration between these two cues. In this work, we study a novel and efficient full-duplex strategy network (FSNet) to address this issue, by considering a better mutual restraint scheme linking motion and appearance allowing exploitation of cross-modal features from the fusion and decoding stage. Specifically, we introduce a relational cross-attention module (RCAM) to achieve bidirectional message propagation across embedding sub-spaces. To improve the model’s robustness and update inconsistent features from the spatiotemporal embeddings, we adopt a bidirectional purification module after the RCAM. Extensive experiments on five popular benchmarks show that our FSNet is robust to various challenging scenarios (e.g., motion blur and occlusion), and compares well to leading methods both for video object segmentation and video salient object detection. The project is publicly available at https://github.com/GewelsJI/FSNet . Ge-Peng Ji, Deng-Ping Fan, Keren Fu, Jianbing Shen, Ling Shao 0001 |
Comput. Vis. Media | 3 |
| 2022 | Depth-Cooperated Trimodal Network for Video Salient Object DetectionabstractDepth can provide useful geographical cues for salient object detection (SOD), and has been proven helpful in recent RGB-D SOD methods. However, existing video salient object detection (VSOD) methods only utilize spatiotemporal information and seldom exploit depth information for detection. In this paper, we propose a depth-cooperated trimodal network, called DCTNet for VSOD, which is a pioneering work to incorporate depth information to assist VSOD. To this end, we first generate depth from RGB frames, and then propose an approach to treat the three modalities unequally. Specifically, a multi-modal attention module (MAM) is designed to model multi-modal long-range dependencies between the main modality (RGB) and the two auxiliary modalities (depth, optical flow). We also introduce a refinement fusion module (RFM) to suppress noises in each modality and select useful information dynamically for further feature refinement. Lastly, a progressive fusion strategy is adopted after the refined features to achieve final cross-modal fusion. Experiments on five benchmark datasets demonstrate the superiority of our depth-cooperated model against 12 state-of-the-art methods, and the necessity of depth is also validated. Yukang Lu, Dingyao Min, Keren Fu, Qijun Zhao |
ICIP | 3 |
| 2022 | Local-Global Interaction and Progressive Aggregation for Video Salient Object Detection
Dingyao Min, Chao Zhang 0072, Yukang Lu, Keren Fu, Qijun Zhao |
ICONIP (6) | 4 |
| 2022 | Light field salient object detection: A review and benchmarkabstractSalient object detection (SOD) is a long-standing research topic in computer vision with increasing interest in the past decade. Since light fields record comprehensive information of natural scenes that benefit SOD in a number of ways, using light field inputs to improve saliency detection over conventional RGB inputs is an emerging trend. This paper provides the first comprehensive review and a benchmark for light field SOD, which has long been lacking in the saliency community. Firstly, we introduce light fields, including theory and data forms, and then review existing studies on light field SOD, covering ten traditional models, seven deep learning-based models, a comparative study, and a brief review. Existing datasets for light field SOD are also summarized. Secondly, we benchmark nine representative light field SOD models together with several cutting-edge RGB-D SOD models on four widely used light field datasets, providing insightful discussions and analyses, including a comparison between light field SOD and RGB-D SOD models. Due to the inconsistency of current datasets, we further generate complete data and supplement focal stacks, depth maps, and multi-view images for them, making them consistent and uniform. Our supplemental data make a universal benchmark possible. Lastly, light field SOD is a specialised problem, because of its diverse data representations and high dependency on acquisition hardware, so it differs greatly from other saliency detection tasks. We provide nine observations on challenges and future directions, and outline several open issues. All the materials including models, datasets, benchmarking results, and supplemented light field datasets are publicly available at https://github.com/kerenfu/LFSOD-Survey . Keren Fu, Yao Jiang 0002, Ge-Peng Ji, Tao Zhou 0002, Qijun Zhao, Deng-Ping Fan |
Comput. Vis. Media | 1 |
| 2022 | Few-shot learning-based RGB-D salient object detection: A case study
Keren Fu, Xiao Yang 0029 |
Neurocomputing | 1 |
| 2022 | MEANet: Multi-modal edge-aware network for light field salient object detection
Yao Jiang 0002, Wenbo Zhang 0009, Keren Fu, Qijun Zhao |
Neurocomputing | 3 |
| 2022 | Siamese Network for RGB-D Salient Object Detection and BeyondabstractExisting RGB-D salient object detection (SOD) models usually treat RGB and depth as independent information and design separate networks for feature extraction from each. Such schemes can easily be constrained by a limited amount of training data or over-reliance on an elaborately designed training process. Inspired by the observation that RGB and depth modalities actually present certain commonality in distinguishing salient objects, a novel joint learning and densely cooperative fusion (JL-DCF) architecture is designed to learn from both RGB and depth inputs through a shared network backbone, known as the Siamese architecture. In this paper, we propose two effective components: joint learning (JL), and densely cooperative fusion (DCF). The JL module provides robust saliency feature learning by exploiting cross-modal commonality via a Siamese network, while the DCF module is introduced for complementary feature discovery. Comprehensive experiments using 5 popular metrics show that the designed framework yields a robust RGB-D saliency detector with good generalization. As a result, JL-DCF significantly advances the SOTAs by an average of ~2.0% (F-measure) across 7 challenging datasets. In addition, we show that JL-DCF is readily applicable to other related multi-modal detection tasks, including RGB-T SOD and video SOD, achieving comparable or better performance. Keren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun Zhao, Jianbing Shen, Ce Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Fast Camouflaged Object Detection via Edge-based Reversible Re-calibration Network
Ge-Peng Ji, Lei Zhu 0012, Mingchen Zhuge, Keren Fu |
Pattern Recognit. | 4 |
| 2022 | Mutual-Guidance Transformer-Embedding Network for Video Salient Object DetectionabstractVideo salient object detection (VSOD) aims at locating the most attractive objects presented in video sequences by exploiting spatial and temporal cues. Previous methods mainly utilize convolutional neural networks (CNNs) to fuse or complement across RGB and optical flow cues via simple strategies. To take full advantage of CNNs and recently emerged Transformers, this letter proposes a novel mutual-guidance Transformer-embedding network, called MGT-Net, where a mutual-guidance multi-head attention mechanism (MGMA) explores more sophisticated long-range cross-modal interactions. Such a mechanism is designed into a new mutual-guidance Transformer (MGTrans) module that can propagate long-range contextual dependencies based on information of the other modality. To the best of our knowledge, MGT-Net is the first VSOD model that embeds Transformers as modules into CNNs for improved performance. Prior to MGTrans, we also propose and deploy a feature purification module (FPM) to purify noisy backbone features. Experimental results on five benchmark datasets demonstrate the state-of-the-art performance of MGT-Net. Dingyao Min, Chao Zhang 0072, Yukang Lu, Keren Fu, Qijun Zhao |
IEEE Signal Process. Lett. | 4 |
| 2022 | Toward Identity Preserving Face Synthesis Between Sketches and Photos Using Deep Feature InjectionabstractIdentity preservation is crucial for the practical application of face sketch and photo synthesis, such as law enforcement and entertainment. However, existing methods mainly focus on generating faces with good visual quality, leading the evaluation of identity preservation is partial and insufficient. Besides, it is difficult to simultaneously synthesize photos and sketches that have good visual quality and identity preservation. In this article, we propose to utilize auxiliary deep features extracted by an off-the-shelf face classifier to inject into the synthesis processes, so that the synthesized faces have more identity information. Furthermore, we propose a light interpolated convolutional neural network that has a shared encoder and two face-specialized decoders to simultaneously complete the transformations between sketches and photos. Evaluations on identity preservation and visual quality show our method is superior to existing methods in synthesizing both sketches and photos, and is qualified in practical application. Keren Fu, Shenggui Ling, Jiang Wang 0005, Peng Cheng 0006 |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | RGB-D Salient Object Detection via 3D Convolutional Neural NetworksabstractRGB-D salient object detection (SOD) recently has attracted increasing research interest and many deep learning methods based on encoder-decoder architectures have emerged. However, most existing RGB-D SOD models conduct feature fusion either in the single encoder or the decoder stage, which hardly guarantees sufficient cross-modal fusion ability. In this paper, we make the first attempt in addressing RGB-D SOD through 3D convolutional neural networks. The proposed model, named RD3D, aims at pre-fusion in the encoder stage and in-depth fusion in the decoder stage to effectively promote the full integration of RGB and depth streams. Specifically, RD3D first conducts pre-fusion across RGB and depth modalities through an inflated 3D encoder, and later provides in-depth feature fusion by designing a 3D decoder equipped with rich back-projection paths (RBPP) for leveraging the extensive aggregation ability of 3D convolutions. With such a progressive fusion strategy involving both the encoder and decoder, effective and thorough interaction between the two modalities can be exploited and boost the detection accuracy. Extensive experiments on six widely used benchmark datasets demonstrate that RD3D performs favorably against 14 state-of-the-art RGB-D SOD approaches in terms of four key evaluation metrics. Our code will be made publicly available: https://github.com/PPOLYpubki/RD3D. Yi Zhang 0076, Keren Fu, Qijun Zhao, Hongwei Du 0004 |
AAAI | 4 |
| 2021 | Full-Duplex Strategy for Video Object SegmentationabstractAppearance and motion are two important sources of information in video object segmentation (VOS). Previous methods mainly focus on using simplex solutions, lowering the upper bound of feature collaboration among and across these two cues. In this paper, we study a novel framework, termed the FSNet (Full-duplex Strategy Network), which designs a relational cross-attention module (RCAM) to achieve the bidirectional message propagation across embedding subspaces. Furthermore, the bidirectional purification module (BPM) is introduced to update the inconsistent features between the spatial-temporal embeddings, effectively improving the model robustness. By considering the mutual restraint within the full-duplex strategy, our FSNet performs the cross-modal feature-passing (i.e., transmission and receiving) simultaneously before the fusion and decoding stage, making it robust to various challenging scenarios (e.g., motion blur, occlusion) in VOS. Extensive experiments on five popular benchmarks (i.e., DAVIS16, FBMS, MCL, SegTrack-V2, and DAVSOD19) show that our FSNet outperforms other state-of-the-arts for both the VOS and video salient object detection tasks. Ge-Peng Ji, Keren Fu, Deng-Ping Fan, Jianbing Shen, Ling Shao 0001 |
ICCV | 2 |
| 2021 | BTS-Net: Bi-Directional Transfer-And-Selection Network for RGB-D Salient Object DetectionabstractDepth information has been proved beneficial in RGB-D salient object detection (SOD). However, depth maps obtained often suffer from low quality and inaccuracy. Most existing RGB-D SOD models have no cross-modal interactions or only have unidirectional interactions from depth to RGB in their encoder stages, which may lead to inaccurate encoder features when facing low quality depth. To address this limitation, we propose to conduct progressive bidirectional interactions as early in the encoder stage, yielding a novel bi-directional transfer-and-selection network named BTS-Net, which adopts a set of bi-directional transfer-and-selection (BTS) modules to purify features during encoding. Based on the resulting robust encoder features, we also design an effective light-weight group decoder to achieve accurate final saliency prediction. Comprehensive experiments on six widely used datasets demonstrate that BTS-Net surpasses 16 latest state-of-the-art approaches in terms of four key metrics. Wenbo Zhang 0009, Yao Jiang 0002, Keren Fu, Qijun Zhao |
ICME | 3 |
| 2021 | Depth Quality-Inspired Feature Manipulation for Efficient RGB-D Salient Object DetectionabstractRGB-D salient object detection (SOD) recently has attracted increasing research interest by benefiting conventional RGB SOD with extra depth information. However, existing RGB-D SOD models often fail to perform well in terms of both efficiency and accuracy, which hinders their potential applications on mobile devices and real-world problems. An underlying challenge is that the model accuracy usually degrades when the model is simplified to have few parameters. To tackle this dilemma and also inspired by the fact that depth quality is a key factor influencing the accuracy, we propose a novel depth quality-inspired feature manipulation (DQFM) process, which is efficient itself and can serve as a gating mechanism for filtering depth features to greatly boost the accuracy. DQFM resorts to the alignment of low-level RGB and depth features, as well as holistic attention of the depth stream to explicitly control and enhance cross-modal fusion. We embed DQFM to obtain an efficient light-weight model called DFM-Net, where we also design a tailored depth backbone and a two-stage decoder for further efficiency consideration. Extensive experimental results demonstrate that our DFM-Net achieves state-of-the-art accuracy when comparing to existing non-efficient models, and meanwhile runs at 140ms on CPU (2.2x faster than the prior fastest efficient model) with only ~8.5Mb model size (14.9% of the prior lightest). Our code will be available at https://github.com/zwbx/DFM-Net. Wenbo Zhang 0009, Ge-Peng Ji, Zhuo Wang 0004, Keren Fu, Qijun Zhao |
ACM Multimedia | 4 |
| 2021 | Unsupervised many-to-many image-to-image translation across multiple domainsabstractAbstract Unsupervised multi‐domain image‐to‐image translation aims to synthesize images among multiple domains without labelled data, which is more general and complicated than one‐to‐one image mapping. However, existing methods mainly focus on reducing the large costs of modelling and do not pay enough attention to the quality of generated images. In some target domains, their translation results may not be expected or even cause the model collapse. To improve the image quality, an effective many‐to‐many mapping framework for unsupervised multi‐domain image‐to‐image translation is proposed. There are two key aspects to the proposed method. The first is a many‐to‐many architecture with only one domain‐shared encoder and several domain‐specialized decoders to effectively and simultaneously translate images across multiple domains. The second is two proposed constraints extended from one‐to‐one mappings to further help improve the generation. All the evaluations demonstrate that the proposed framework is superior to existing methods and provides an effective solution for multi‐domain image‐to‐image translation. Keren Fu, Shenggui Ling, Peng Cheng 0006 |
IET Image Process. | 2 |
| 2021 | BCNet: Bidirectional collaboration network for edge-guided salient object detection
Bo Dong 0001, Chuanfei Hu, Keren Fu, Geng Chen 0001 |
Neurocomputing | 4 |
| 2021 | Identity-and-pose-guided generative adversarial network for face rotation
Yi Zhang 0018, Keren Fu, Peng Cheng 0006 |
Neurocomputing | 2 |
| 2021 | PGM-face: Pose-guided margin loss for cross-pose face recognition
Yi Zhang 0018, Keren Fu, Peng Cheng 0006, Shanmin Yang, Xiao Yang 0029 |
Neurocomputing | 2 |
| 2021 | EF-Net: A novel enhancement and fusion network for RGB-D saliency detection
Keren Fu, Geng Chen 0001, Hongwei Du 0004, Bensheng Qiu, Ling Shao 0001 |
Pattern Recognit. | 2 |
| 2021 | SAL: Selection and Attention Losses for Weakly Supervised Semantic SegmentationabstractTraining a fully supervised semantic segmentation network requires a large amount of expensive pixel-level annotations in manual labor. In this work, we focus on studying the semantic segmentation problem using only image-level supervision. An effective scheme for weakly supervised segmentation is employed to produce the proxy annotations via image tags firstly. Then the segmentation network is retrained on the generated noisy proxy annotations. However, learning from noisy annotations is risky, as proxy annotations of poor quality may deteriorate the performance of the baseline segmentation and classification networks. In order to train the segmentation network using noisy annotations more effectively, two novel loss functions are proposed in this paper, namely, the selection loss and attention loss. Firstly, a selection loss is designed by weighting the proxy annotations based on a coarse-to-fine strategy for evaluating the quality of segmentation masks. Secondly, an attention loss taking the clean image tags as supervision is utilized to correct the classification errors caused by ambiguous pixel-level labels. Finally, we propose an end-to-end semantic segmentation network SAL-Net guided by the above two losses. From the extensive experiments conducted on PASCAL VOC 2012 dataset, SAL-Net reaches state-of-the-art performance with mean IoU (mIoU) as 62.5% and 66.6% on the test set by taking VGG16 network and ResNet101 network as the baselines respectively, which demonstrates the superiority of the proposed algorithm over eight representative weakly supervised segmentation methods. The code and models are available at https://github.com/zmbhou/SALTMM. Lei Zhou 0003, Chen Gong 0002, Zhi Liu 0003, Keren Fu |
IEEE Trans. Multim. | 4 |
| 2020 | JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object DetectionabstractThis paper proposes a novel joint learning and densely-cooperative fusion (JL-DCF) architecture for RGB-D salient object detection. Existing models usually treat RGB and depth as independent information and design separate networks for feature extraction from each. Such schemes can easily be constrained by a limited amount of training data or over-reliance on an elaborately-designed training process. In contrast, our JL-DCF learns from both RGB and depth inputs through a Siamese network. To this end, we propose two effective components: joint learning (JL), and densely-cooperative fusion (DCF). The JL module provides robust saliency feature learning, while the latter is introduced for complementary feature discovery. Comprehensive experiments on four popular metrics show that the designed framework yields a robust RGB-D saliency detector with good generalization. As a result, JL-DCF significantly advances the top-1 D3Net model by an average of ~1.9% (S-measure) across six challenging datasets, showing that the proposed framework offers a potential solution for real-world applications and could provide more insight into the cross-modality complementarity task. The code will be available at https://github.com/kerenfu/JLDCF/. Keren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun Zhao |
CVPR | 1 |
| 2020 | MDT: Unsupervised Multi-Domain Image-To-Image Translator Based On Generative Adversarial NetworksabstractIn recent years, many methods have been proposed to tackle unsupervised image-to-image translation. However, they mainly focus on two-domain scenarios. To address issues like large cost of training time and resources in translation across multiple (more than two) number of domains, we propose an unsupervised method based on generative adversarial networks. Our method has only one encoder for the consideration of efficiency, together with several domain-specified decoders to transform an image into multiple domains without needing an input domain label. In addition, we propose to employ two constraints namely reconstruction loss and identity loss to further improve the generation. We conduct experiments on two datasets. The results demonstrate the effectiveness and efficiency of our proposed method against state-of-the-art methods. Keren Fu, Shenggui Ling, Peng Cheng 0006 |
ICIP | 2 |
| 2020 | Learning from discrete Gaussian label distribution and spatial channel-aware residual attention for head pose estimation
Yi Zhang 0018, Keren Fu, Jiang Wang 0005, Peng Cheng 0006 |
Neurocomputing | 2 |
| 2020 | Deep 3D Facial Landmark Localization on position maps
Jingchen Zhang, Kangkang Gao, Keren Fu, Peng Cheng 0006 |
Neurocomputing | 3 |
| 2020 | Joint patch and instance discrimination learning for unsupervised person re-identification
Yu Zhao 0034, Qiaoyuan Shu, Keren Fu, Pengcheng Wei, Jian Zhan |
Image Vis. Comput. | 3 |
| 2020 | An Identity-Preserved Model for Face Sketch-Photo SynthesisabstractFace sketch-photo synthesis can be regarded as an image-to-image translation problem. Although many generative models achieve good translations from sketches to photos, they still have limitations in preserving face identity due to the huge modality gap of the two domains. To this end, we propose an identity-preserved adversarial model (IPAM), which includes an extended U-Net to increase the weight of the original sketch in translation, two discriminators focusing on the real or fake image concatenation of two domains to learn more styles of the target domain, and an identity constraint to request the fakes and the real targets to have zero cosine distance in feature space. We evaluate our method on two face sketch databases with face recognition. The results demonstrate our translation method is superior to the existing methods in maintaining face identity information. Shenggui Ling, Keren Fu, Peng Cheng 0006 |
IEEE Signal Process. Lett. | 3 |
| 2019 | Learning to Focus and Discriminate for Fine-Grained ClassificationabstractExisting state-of-the-art fine-grained classification methods usually use separated networks for discriminative region localization and feature learning/classification, and are thus complicated to implement and optimize. In this paper, we aim to provide a compact solution by deepening the collaboration between the region localization, feature learning and classification modules during the learning process of fine-grained classification. We thus propose a method that can learn to simultaneously localize discriminative regions and extract discriminative features by exploring the localization ability of classification convolutional neural networks and joint optimization of different modules. Our method, while being built upon a single backbone network and trained with only softmax losses, achieves state-of-the-art performance on three benchmark fine-grained datasets, which proves that our method is simple but effective for fine-grained classification. Zhicong Feng, Keren Fu, Qijun Zhao |
ICIP | 2 |
| 2019 | Deepside: A general deep framework for salient object detection
Keren Fu, Qijun Zhao, Irene Y. H. Gu, Jie Yang 0002 |
Neurocomputing | 1 |
| 2019 | Superpixel based continuous conditional random field neural network for semantic segmentation
Lei Zhou 0003, Keren Fu, Zhi Liu 0003, Zhimin Yin, Jianli Zheng |
Neurocomputing | 2 |
| 2019 | Fast spatial-temporal stereo matching for 3D face reconstruction under speckle pattern projection
Keren Fu, Yijiang Xie, Hailong Jing, Jiangping Zhu |
Image Vis. Comput. | 1 |
| 2019 | Refinet: A Deep Segmentation Assisted Refinement Network for Salient Object DetectionabstractCompared to conventional saliency detection by handcrafted features, deep convolutional neural networks (CNNs) recently have been successfully applied to saliency detection field with superior performance on locating salient objects. However, due to repeated sub-sampling operations inside CNNs such as pooling and convolution, many CNN-based saliency models fail to maintain fine-grained spatial details and boundary structures of objects. To remedy this issue, this paper proposes a novel end-to-end deep learning-based refinement model named Refinet, which is based on fully convolutional network augmented with segmentation hypotheses. Intermediate saliency maps that are edge-aware are computed from segmentation-based pooling and then feed to a two-tier fully convolutional network for effective fusion and refinement, leading to more precise object details and boundaries. In addition, the resolution of feature maps in the proposed Refinet is carefully designed to guarantee sufficient boundary clarity of the refined saliency output. Compared to widely employed dense conditional random field, Refinet is able to enhance coarse saliency maps generated by existing models with more accurate spatial details, and its effectiveness is demonstrated by experimental results on seven benchmark datasets. Keren Fu, Qijun Zhao, Irene Y. H. Gu |
IEEE Trans. Multim. | 1 |
| 2018 | Spectral salient object detection
Keren Fu, Irene Y. H. Gu, Jie Yang 0002 |
Neurocomputing | 1 |
| 2018 | Inverse Nonnegative Local Coordinate Factorization for Visual TrackingabstractRecently, nonnegative matrix factorization (NMF) with part-based representation has been widely used for appearance modeling in visual tracking. Unfortunately, not all the targets can be successfully decomposed as “parts” unless some rigorous conditions are satisfied. To avoid this problem, this paper introduces NMF's variants into the visual tracking framework in the view of data clustering for appearance modeling. First, an initial target appearance model based on NMF is proposed to describe the target's appearance with the incorporated local coordinate factorization constraint, orthogonality of the bases, and L1,1norm regularized sparse residual error constraint. Second, an inverse NMF model is proposed in which each learned base vector is regarded as a clustering center in a low-dimensional subspace. Potential target samples (from the foreground) will be clustered around base vectors, while the candidate samples (from the background) are very likely to spread irregularly over the entire clustering space. Such differences can be fully exploited by the inverse NMF model to produce more discriminative encoding vectors than the conventional NMF method. Furthermore, incremental updating model is introduced into the tracking framework for online updating the initial appearance model. Experiments on object tracking benchmark suggest that our tracker is able to achieve promising performance when compared with some state-of-the-art methods in deformation, occlusion, and other challenging situations. Fanghui Liu 0001, Tao Zhou 0002, Chen Gong 0002, Keren Fu, Li Bai 0001, Jie Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Kernelized temporal locality learning for real-time visual tracking
Fanghui Liu 0001, Tao Zhou 0002, Keren Fu, Jie Yang 0002 |
Pattern Recognit. Lett. | 3 |
| 2017 | Saliency Detection by Fully Learning a Continuous Conditional Random FieldabstractSalient object detection is aimed at detecting and segmenting objects that human eyes are most focused on when viewing a scene. Recently, conditional random field (CRF) is drawn renewed interest, and is exploited in this field. However, when utilizing a CRF with unary and pairwise potentials having essential parameters, most existing methods only employ manually designed parameters, or learn parameters partly for the unary potentials. Observing that the saliency estimation is a continuous labeling issue, this paper proposes a novel data-driven scheme based on a special CRF framework, the so-called continuous CRF (C-CRF), where parameters for both unary and pairwise potentials are jointly learned. The proposed C-CRF learning provides an optimal way to integrate various unary saliency features with pairwise cues to discover salient objects. To the best of our knowledge, the proposed scheme is the first to completely learn a C-CRF for saliency detection. In addition, we propose a novel formulation of pairwise potentials that enables learning weights for different spatial ranges on a superpixel graph. The proposed C-CRF learning-based saliency model is tested on 6 benchmark datasets and compared with 11 existing methods. Our results and comparisons have provided further support to the proposed method in terms of precision-recall and F-measure. Furthermore, incorporating existing saliency models with pairwise cues through the C-CRF are shown to provide marked boosting performance over individual models. Keren Fu, Irene Y. H. Gu, Jie Yang 0002 |
IEEE Trans. Multim. | 1 |
| 2017 | Visual Tracking via Nonnegative Multiple CodingabstractIt has been extensively observed that an accurate appearance model is critical to achieving satisfactory performance for robust object tracking. Most existing top-ranked methods rely on linear representation over a single dictionary, which brings about improper understanding on the target appearance. To address this problem, in this paper, we propose a novel appearance model named as “nonnegative multiple coding” (NMC) to accurately represent a target. First, a series of local dictionaries are created with different predefined numbers of nearest neighbors, and then the contributions of these dictionaries are automatically learned. As a result, this ensemble of dictionaries can comprehensively exploit the appearance information carried by all the constituted dictionaries. Second, the existing methods explicitly impose the nonnegative constraint to coefficient vectors, but in the proposed model, we directly deploy an efficient 12 norm regularization to achieve the similar nonnegative purpose with theoretical guarantees. Moreover, an efficient occlusion detection scheme is designed to alleviate tracking drifts, which investigates whether negative templates are selected to represent the severely occluded target. Experimental results on two benchmarks demonstrate that our NMC tracker are able to achieve superior performance to state-of-the-art methods. Fanghui Liu 0001, Chen Gong 0002, Tao Zhou 0002, Keren Fu, Xiangjian He, Jie Yang 0002 |
IEEE Trans. Multim. | 4 |
| 2016 | Learning full-range affinity for diffusion-based saliency detectionabstractIn this paper we address the issue of enhancing salient object detection through diffusion-based techniques. For reliably diffusing the energy from labeled seeds, we propose a novel graph-based diffusion scheme called affinity learning-based diffusion (ALD), which is based on learning full-range affinity between two arbitrary graph nodes. The method differs from the previous existing work where implicit diffusion was formulated as a ranking problem on a graph. In the proposed method, the affinity learning is achieved in a unified graph-based semi-supervised manner, whose outcome is leveraged for global propagation. By properly selecting an affinity learning model, the proposed ALD outperforms the ranking-based diffusion in terms of accurately detecting salient objects and enhancing the correct salient objects under a range of background scenarios. By utilizing the ALD, we propose an enhanced saliency detector that outperforms 7 recent state-of-the-art saliency models on 3 benchmark datasets. Keren Fu, Irene Y. H. Gu, Jie Yang 0002 |
ICASSP | 1 |
| 2016 | Robust visual tracking via inverse nonnegative matrix factorizationabstractThe establishment of robust target appearance model over time is an overriding concern in visual tracking. In this paper, we propose an inverse nonnegative matrix factorization (NMF) method for robust appearance modeling. Rather than using a linear combination of nonnegative basis vectors for each target image patch in conventional NMF, the proposed method is a reverse thought to conventional NMF tracker. It utilizes both the foreground and background information, and imposes a local coordinate constraint, where the basis matrix is sparse matrix from the linear combination of candidates with corresponding nonnegative coefficient vectors. Inverse NMF is used as a feature encoder, where the resulting coefficient vectors are fed into a SVM classifier for separating the target from the background. The proposed method is tested on several videos and compared with seven state-of-the-art methods. Our results have provided further support to the effectiveness and robustness of the proposed method. Fanghui Liu 0001, Tao Zhou 0002, Keren Fu, Irene Y. H. Gu, Jie Yang 0002 |
ICASSP | 3 |
| 2016 | Abnormal event detection using spatio-temporal feature and nonnegative locality-constrained linear codingabstractIn this paper, an approach using the spatio-temporal feature and nonnegative locality-constrained linear coding (NLLC) is proposed to detect abnormal events in videos. This approach utilizes position-based spatio-temporal descriptors as the low-level representations of a video clip. Each descriptor consists of the position information of a space-time interest point and an appearance feature vector. To obtain the high-level video representations, the nonnegative locality-constrained linear coding is adopted to encode each spatio-temporal descriptor. Then, the max pooling integrates all NLLC codes of a video clip to produce a feature vector. Finally, the support vector machine (SVM) is employed to classify the feature vector as abnormal or normal. Experimental results on two datasets have demonstrated the promising performance of the proposed approach in the detection of both global and local abnormal events. Yu Zhao 0034, Lei Zhou 0003, Keren Fu, Jie Yang 0002 |
ICIP | 3 |
| 2016 | Geodesic distance transform-based salient region segmentation for automatic traffic sign recognitionabstractVisual-based traffic sign recognition (TSR) requires first detecting and then classifying signs from captured images. In such a cascade system, classification accuracy is often affected by the detection results. This paper proposes a method for extracting a salient region of traffic sign within a detection window for more accurate sign representation and feature extraction, hence enhancing the performance of classification. In the proposed method, a superpixel-based distance map is firstly generated by applying a signed geodesic distance transform from a set of selected foreground and background seeds. An effective method for obtaining a final segmentation from the distance map is then proposed by incorporating the shape constraints of signs. Using these two steps, our method is able to automatically extract salient sign regions of different shapes. The proposed method is tested and validated in a complete TSR system. Test results show that the proposed method has led to a high classification accuracy (97.11%) on a large dataset containing street images. Comparing to the same TSR system without using saliency-segmented regions, the proposed method has yielded a marked performance improvement (about 12.84%). Future work will be on extending to more traffic sign categories and comparing with other benchmark methods. Keren Fu, Irene Y. H. Gu, Anders C. E. Ödblom, Feng Liu 0013 |
Intelligent Vehicles Symposium | 1 |
| 2016 | Robust manifold-preserving diffusion-based saliency detection by adaptive weight construction
Keren Fu, Irene Y. H. Gu, Chen Gong 0002, Jie Yang 0002 |
Neurocomputing | 1 |
| 2016 | Robust visual tracking via constrained correlation filter coding
Fanghui Liu 0001, Tao Zhou 0002, Keren Fu, Jie Yang 0002 |
Pattern Recognit. Lett. | 3 |
| 2016 | Co-saliency detection via inter and intra saliency propagation
Chenjie Ge, Keren Fu, Fanghui Liu 0001, Li Bai 0001, Jie Yang 0002 |
Signal Process. Image Commun. | 2 |
| 2015 | Saliency propagation from simple to difficultabstractSaliency propagation has been widely adopted for identifying the most attractive object in an image. The propagation sequence generated by existing saliency detection methods is governed by the spatial relationships of image regions, i.e., the saliency value is transmitted between two adjacent regions. However, for the inhomogeneous difficult adjacent regions, such a sequence may incur wrong propagations. In this paper, we attempt to manipulate the propagation sequence for optimizing the propagation quality. Intuitively, we postpone the propagations to difficult regions and meanwhile advance the propagations to less ambiguous simple regions. Inspired by the theoretical results in educational psychology, a novel propagation algorithm employing the teaching-to-learn and learning-to-teach strategies is proposed to explicitly improve the propagation quality. In the teaching-to-learn step, a teacher is designed to arrange the regions from simple to difficult and then assign the simplest regions to the learner. In the learning-to-teach step, the learner delivers its learning confidence to the teacher to assist the teacher to choose the subsequent simple regions. Due to the interactions between the teacher and learner, the uncertainty of original difficult regions is gradually reduced, yielding manifest salient objects with optimized background suppression. Extensive experimental results on benchmark saliency datasets demonstrate the superiority of the proposed algorithm over twelve representative saliency detectors. Chen Gong 0002, Dacheng Tao, Wei Liu 0005, Stephen J. Maybank, Keren Fu, Jie Yang 0002 |
CVPR | 6 |
| 2015 | Small target detection using an optimization-based filterabstractSmall target detection is a critical problem in the Infrared Search And Track (IRST) system. Although it has been studied for years, there are some challenges remained, e.g. cloud edges and horizontal lines are likely to cause false alarms. This paper proposes a novel method using an optimization-based filter to detect infrared small target in heavy clutter. First, we design a certain pixel area as active area. Second, a weighted quadratic cost function is performed in the active area. Finally, a filter based on statistics of active area is derived from the cost function. Our method could preserve heterogeneous area, meanwhile, remove target region. Experimental results show our method achieves satisfied performance in heavy clutter. Keren Fu, Tao Zhou 0002, Jie Yang 0002, Qiang Wu 0001, Xiangjian He |
ICASSP | 2 |
| 2015 | Salient object detection using normalized cut and geodesicsabstractRecently the Normalized cut (Ncut) has been introduced to salient object detection [1, 2]. In this paper we validate that instead of proposing new detection models that leverage the Ncut, the previous geodesic saliency detection model which computes shortest paths on a graph can be adapted to eigenvectors of the Ncut to produce superior performance. Since the Ncut partitions a graph in a normalized energy minimization fashion, resulting eigenvectors contain decent cluster information that can group visual contents. Combining it with the existing geodesic saliency detection is conducive to highlighting salient objects uniformly, yielding to improved detection accuracy. Experiments by comparing with 12 existing methods on four benchmark datasets show the proposed method significantly outperforms the original geodesic saliency model and achieves comparable performance to state-of-the-art methods. Keren Fu, Chen Gong 0002, Irene Y. H. Gu, Jie Yang 0002 |
ICIP | 1 |
| 2015 | Co-saliency detection via similarity-based saliency propagationabstractIn this paper, we present a method for discovering the common salient objects from a set of images. We treat co-saliency detection as a pairwise saliency propagation problem, which utilizes the similarity between each pair of images to measure the common property with the guidance of a single saliency map image. Given the pairwise co-salient foreground maps, pairwise saliency is optimized by combining the initial background cues. Pairwise co-salient maps are then fused according to a novel fusion strategy based on the focus of human attention. Finally we adopt an integrated multi-scale scheme to obtain the pixel-level saliency map. Our proposed model makes the existing single saliency model perform well in co-saliency detection and is not overly sensitive to the initial saliency model selected. Extensive experiments on two benchmark databases show the superiority of our co-saliency model against the state-of-the-art methods both subjectively and objectively. Chenjie Ge, Keren Fu, Yijun Li 0003, Jie Yang 0002, Li Bai 0001 |
ICIP | 2 |
| 2015 | Traffic sign recognition using salient region features: A novel learning-based coarse-to-fine schemeabstractTraffic sign recognition, including sign detection and classification, is essential for advanced driver assistance systems and autonomous vehicles. This paper introduces a novel machine learning-based sign recognition scheme. In the proposed scheme, detection and classification are realized through learning in a coarse-to-fine manner. Based on the observation that signs in the same category share some common attributes in appearance, the proposed scheme first distinguishes each individual sign category from the background in the coarse learning stage (i.e. sign detection) followed by distinguishing different sign classes within each category in the fine learning stage (i.e. sign classification). Both stages are realized through machine learning techniques. A complete recognition scheme is developed that is effective for simultaneously recognizing multiple categories of traffic signs. In addition, a novel saliency-based feature extraction method is proposed for sign classification. The method segments salient sign regions by leveraging the geodesic energy propagation. Compared with the conventional feature extraction, our method provides more reliable feature extraction from salient sign regions. The proposed scheme is tested and validated on two categories of Chinese traffic signs from Tencent street view. Evaluations on the test dataset show reasonably good performance, with an average of 97.5% true positive and 0.3% false positive on two categories of traffic signs. Keren Fu, Irene Y. H. Gu, Anders C. E. Ödblom |
Intelligent Vehicles Symposium | 1 |
| 2015 | Scalable Semi-Supervised Classification via Neumann Series
Chen Gong 0002, Keren Fu, Lei Zhou 0003, Jie Yang 0002, Xiangjian He |
Neural Process. Lett. | 2 |
| 2015 | Robust visual tracking via efficient manifold ranking with low-dimensional compressive features
Tao Zhou 0002, Xiangjian He, Keren Fu, Jie Yang 0002 |
Pattern Recognit. | 4 |
| 2015 | Efficient Saliency-Model-Guided Visual Co-Saliency DetectionabstractThis letter proposes a novel framework to detect common salient objects in a group of images automatically and efficiently. Different from most existing co-saliency models which directly redesign algorithms for multiple images, the saliency model for a single image is fully exploited under the proposed framework to guide the co-saliency detection. Given single image saliency maps, a two-stage guided detection pipeline led by queries is proposed to obtain the guided saliency maps of the image set through a ranking scheme. Then the guided saliency maps generated by different queries are fused in a way that takes advantages of both averaging and multiplication. The proposed model makes existing saliency models work well in co-saliency scenarios. Experimental results on two benchmark databases demonstrate that the proposed framework outperforms the state-of-the-art models in terms of both accuracy and efficiency. Yijun Li 0003, Keren Fu, Zhi Liu 0003, Jie Yang 0002 |
IEEE Signal Process. Lett. | 2 |
| 2015 | Normalized Cut-Based Saliency Detection by Adaptive Multi-Level Region MergingabstractExisting salient object detection models favor over-segmented regions upon which saliency is computed. Such local regions are less effective on representing object holistically and degrade emphasis of entire salient objects. As a result, the existing methods often fail to highlight an entire object in complex background. Toward better grouping of objects and background, in this paper, we consider graph cut, more specifically, the normalized graph cut (Ncut) for saliency detection. Since the Ncut partitions a graph in a normalized energy minimization fashion, resulting eigenvectors of the Ncut contain good cluster information that may group visual contents. Motivated by this, we directly induce saliency maps via eigenvectors of the Ncut, contributing to accurate saliency estimation of visual clusters. We implement the Ncut on a graph derived from a moderate number of superpixels. This graph captures both intrinsic color and edge information of image data. Starting from the superpixels, an adaptive multi-level region merging scheme is employed to seek such cluster information from Ncut eigenvectors. With developed saliency measures for each merged region, encouraging performance is obtained after across-level integration. Experiments by comparing with 13 existing methods on four benchmark datasets, including MSRA-1000, SOD, SED, and CSSD show the proposed method, Ncut saliency, results in uniform object enhancement and achieves comparable/better performance to the state-of-the-art methods. Keren Fu, Chen Gong 0002, Irene Y. H. Gu, Jie Yang 0002 |
IEEE Trans. Image Process. | 1 |
| 2015 | Deformed Graph Laplacian for Semisupervised LearningabstractGraph Laplacian has been widely exploited in traditional graph-based semisupervised learning (SSL) algorithms to regulate the labels of examples that vary smoothly on the graph. Although it achieves a promising performance in both transductive and inductive learning, it is not effective for handling ambiguous examples (shown in Fig. 1). This paper introduces deformed graph Laplacian (DGL) and presents label prediction via DGL (LPDGL) for SSL. The local smoothness term used in LPDGL, which regularizes examples and their neighbors locally, is able to improve classification accuracy by properly dealing with ambiguous examples. Theoretical studies reveal that LPDGL obtains the globally optimal decision function, and the free parameters are easy to tune. The generalization bound is derived based on the robustness analysis. Experiments on a variety of real-world data sets demonstrate that LPDGL achieves top-level performance on both transductive and inductive settings by comparing it with popular SSL algorithms, such as harmonic functions, AnchorGraph regularization, linear neighborhood propagation, Laplacian regularized least square, and Laplacian support vector machine. Chen Gong 0002, Tongliang Liu, Dacheng Tao, Keren Fu, Enmei Tu, Jie Yang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2015 | Fick's Law Assisted Propagation for Semisupervised LearningabstractHow to propagate the label information from labeled examples to unlabeled examples is a critical problem for graph-based semisupervised learning. Many label propagation algorithms have been developed in recent years and have obtained promising performance on various applications. However, the eigenvalues of iteration matrices in these algorithms are usually distributed irregularly, which slow down the convergence rate and impair the learning performance. This paper proposes a novel label propagation method called Fick's law assisted propagation (FLAP). Unlike the existing algorithms that are directly derived from statistical learning, FLAP is deduced on the basis of the theory of Fick's First Law of Diffusion, which is widely known as the fundamental theory in fluid-spreading. We prove that FLAP will converge with linear rate and show that FLAP makes eigenvalues of the iteration matrix distributed regularly. Comprehensive experimental evaluations on synthetic and practical datasets reveal that FLAP obtains encouraging results in terms of both accuracy and efficiency. Chen Gong 0002, Dacheng Tao, Keren Fu, Jie Yang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | ReLISH: Reliable Label Inference via Smoothness HypothesisabstractThe smoothness hypothesis is critical for graph-based semi-supervised learning. This paper defines local smoothness, based on which a new algorithm, Reliable Label Inference via Smoothness Hypothesis (ReLISH), is proposed. ReLISH has produced smoother labels than some existing methods for both labeled and unlabeled examples. Theoretical analyses demonstrate good stability and generalizability of ReLISH. Using real-world datasets, our empirical analyses reveal that ReLISH is promising for both transductive and inductive tasks, when compared with representative algorithms, including Harmonic Functions, Local and Global Consistency, Constraint Metric Learning, Linear Neighborhood Propagation, and Manifold Regularization. Chen Gong 0002, Dacheng Tao, Keren Fu, Jie Yang 0002 |
AAAI | 3 |
| 2014 | Signed Laplacian Embedding for Supervised Dimension ReductionabstractManifold learning is a powerful tool for solving nonlinear dimension reduction problems. By assuming that the high-dimensional data usually lie on a low-dimensional manifold, many algorithms have been proposed. However, most algorithms simply adopt the traditional graph Laplacian to encode the data locality, so the discriminative ability is limited and the embedding results are not always suitable for the subsequent classification. Instead, this paper deploys the signed graph Laplacian and proposes Signed Laplacian Embedding (SLE) for supervised dimension reduction. By exploring the label information, SLE comprehensively transfers the discrimination carried by the original data to the embedded low-dimensional space. Without perturbing the discrimination structure, SLE also retains the locality.Theoretically, we prove the immersion property by computing the rank of projection, and relate SLE to existing algorithms in the frame of patch alignment. Thorough empirical studies on synthetic and real datasets demonstrate the effectiveness of SLE. Chen Gong 0002, Dacheng Tao, Jie Yang 0002, Keren Fu |
AAAI | 4 |
| 2014 | Adaptive Multi-Level Region Merging for Salient Object Detection
Keren Fu, Chen Gong 0002, Yixiao Yun, Yijun Li 0003, Irene Y. H. Gu, Jie Yang 0002, Jingyi Yu 0001 |
BMVC | 1 |
| 2014 | Effective small DIM target detection by local connectedness constraintabstractThe main drawback of conventional filtering based methods for small dim target (SDT) detection is they could not guarantee sufficient suppression ability towards trivial high frequency component which belongs to background, such as strong corners and edges. To overcome this bottleneck, this paper proposes an effective SDT detection algorithm by using local connectedness constraint. Our method provides direct control for target size, ensure high accuracy and could be easily embedded into the classical sliding-window based framework. The effectiveness of the proposed method is validated using images with cluttered background. Keren Fu, Chen Gong 0002, Irene Y. H. Gu, Jie Yang 0002 |
ICASSP | 1 |
| 2014 | Saliency detection based on extended boundary prior with foci of attentionabstractIn this paper, we propose a novel bottom-up paradigm for detecting visual saliency. Regarding the boundary as potential background (boundary prior), we firstly transfer the input color image into a graph with additional four virtual nodes. With a new type of edge called feature edge defined considering both color information and spatial distribution, geodesic saliency measure is used to obtain four saliency maps. Then a combination strategy of four maps is proposed, rendering a uniform saliency map to better suppress background and avoid over-suppression of salient object. Finally, we introduce a way of determining foci of attention based on maximal deviation from norm (MDN) to enhance the quality of saliency map. Experimental results on a benchmark dataset demonstrate the better performance of our proposed approach compared with several state-of-art methods. Yijun Li 0003, Keren Fu, Lei Zhou 0003, Yu Qiao 0001, Jie Yang 0002 |
ICASSP | 2 |
| 2014 | Saliency detection via foreground rendering and background exclusionabstractIn this paper, a novel approach for image visual saliency detection is proposed from both the salient object (foreground) and the background perspective. To better highlight the salient object, we start from what is a salient object and adopt priors including contrast prior and center prior to measure the dissimilarity between different image elements. To better suppress the background, we focus on what is the background and measure the pixel-wise saliency by the minimum seam cost where the seam is an optimal 8-connected path from the pixel to some boundary pixel. The final saliency map is obtained by the combination of two measure systems which leads to the goal of both highlighting the salient object and suppressing the background. Both qualitative and quantitative experiments conducted on a benchmark dataset show that our approach outperforms seven state-of-the-art methods. Yijun Li 0003, Keren Fu, Lei Zhou 0003, Yu Qiao 0001, Jie Yang 0002 |
ICIP | 2 |
| 2014 | Visual object tracking with online learning on Riemannian manifolds by one-class support vector machinesabstractThis paper addresses issues in video object tracking. We propose a novel method where tracking is regarded as a one-class classification problem of domain-shift objects. The proposed tracker is inspired by the fact that the positive samples can be bounded by a closed hypersphere generated by one-class support vector machines (SVM), leading to a solution for robust learning of target model online. The main novelties of the paper include: (a) represent the target model by a set of positive samples as a cluster of points on Riemannian manifolds; (b) perform online learning of target model as a dynamic cluster of points flowing on the manifold, in an alternate manner with tracking; (c) formulate geodesic-based kernel function for one-class SVM on Riemannian manifolds under the log-Euclidean metric. Experiments are conducted on several videos, results have provided support to the proposed method. Yixiao Yun, Keren Fu, Irene Y. H. Gu, Jie Yang 0002 |
ICIP | 2 |
| 2014 | Spectral salient object detectionabstractMany existing methods for salient object detection are performed by over-segmenting images into non-overlapping regions, which facilitate local/global color statistics for saliency computation. In this paper, we propose a new approach: spectral salient object detection, which is benefited from selected attributes of normalized cut, enabling better retaining of holistic salient objects as comparing to conventionally employed pre-segmentation techniques. The proposed saliency detection method recursively bi-partitions regions that render the lowest cut cost in each iteration, resulting in binary spanning tree structure. Each segmented region is then evaluated under criterion that fit Gestalt laws and statistical prior. Final result is obtained by integrating multiple intermediate saliency maps. Experimental results on three benchmark datasets demonstrate the effectiveness of the proposed method against 13 state-of-the-art approaches to salient object detection. Keren Fu, Chen Gong 0002, Irene Y. H. Gu, Jie Yang 0002, Xiangjian He |
ICME | 1 |
| 2014 | Visual tracking via graph-based efficient manifold ranking with low-dimensional compressive featuresabstractIn this paper, a novel and robust tracking method based on efficient manifold ranking is proposed. For tracking, tracked results are taken as labeled nodes while candidate samples are taken as unlabeled nodes, and the goal of tracking is to search the unlabeled sample that is the most relevant with existing labeled nodes by manifold ranking algorithm. Meanwhile, we adopt non-adaptive random projections to preserve the structure of original image space, and a very sparse measurement matrix is used to efficiently extract low-dimensional compressive features for object representation. Furthermore, spatial context is used to improve the robustness to appearance variations. Experimental results on some challenging video sequences show the proposed algorithm outperforms six state-of-the-art methods in terms of accuracy and robustness. Tao Zhou 0002, Xiangjian He, Keren Fu, Jie Yang 0002 |
ICME | 4 |
| 2014 | Graph Construction for Salient Object Detection in VideosabstractRecently many graph-based salient region/object detection methods have been developed. They are rather effective for still images. However, little attention has been paid to salient region detection in videos. This paper addresses salient region detection in videos. A unified approach towards graph construction for salient object detection in videos is proposed. The proposed method combines static appearance and motion cues to construct graph, enabling a direct extension of original graph-based salient region detection to video processing. To maintain coherence in both intra- and inter-frames, a spatial-temporal smoothing operation is proposed on a structured graph derived from consecutive frames. The effectiveness of the proposed method is tested and validated using seven videos from two video datasets. Keren Fu, Irene Y. H. Gu, Yixiao Yun, Chen Gong 0002, Jie Yang 0002 |
ICPR | 1 |
| 2014 | Semi-supervised classification with pairwise constraints
Chen Gong 0002, Keren Fu, Qiang Wu 0001, Enmei Tu, Jie Yang 0002 |
Neurocomputing | 2 |
| 2014 | Bayesian salient object detection based on saliency driven clustering
Lei Zhou 0003, Keren Fu, Yijun Li 0003, Yu Qiao 0001, Xiangjian He, Jie Yang 0002 |
Signal Process. Image Commun. | 2 |
| 2014 | PageRank Tracker: From Ranking to TrackingabstractVideo object tracking is widely used in many real-world applications, and it has been extensively studied for over two decades. However, tracking robustness is still an issue in most existing methods, due to the difficulties with adaptation to environmental or target changes. In order to improve adaptability, this paper formulates the tracking process as a ranking problem, and the PageRank algorithm, which is a well-known webpage ranking algorithm used by Google, is applied. Labeled and unlabeled samples in tracking application are analogous to query webpages and the webpages to be ranked, respectively. Therefore, determining the target is equivalent to finding the unlabeled sample that is the most associated with existing labeled set. We modify the conventional PageRank algorithm in three aspects for tracking application, including graph construction, PageRank vector acquisition and target filtering. Our simulations with the use of various challenging public-domain video sequences reveal that the proposed PageRank tracker outperforms mean-shift tracker, co-tracker, semiboosting and beyond semiboosting trackers in terms of accuracy, robustness and stability. Chen Gong 0002, Keren Fu, Artur Loza, Qiang Wu 0001, Jie Yang 0002 |
IEEE Trans. Cybern. | 2 |
| 2013 | Geodesic saliency propagation for image salient region detectionabstractThis paper proposes a novel geodesic saliency propagation method where detected salient objects may be isolated from both the background and other clutter by adding global considerations in the detection process. The method transmits saliency energy from a coarse saliency map to all image parts rather than from image boundaries in conventional cases. The coarse saliency map is computed using the combination of global contrast and Harris convex hull. Superpixels from pre-segmented image are used as pre-processing to further enhance the efficiency. The proposed propagation is geodesic distance assisted and retains the local connectivity of objects. It is capable of rendering a uniform saliency map while suppressing the background, leading to salient objects being popped out. Experiments were conducted on a benchmark dataset, visual comparisons and performance evaluations with 9 existing methods have shown that the proposed method is robust and achieves the state-of-the-art performance. Keren Fu, Chen Gong 0002, Irene Y. H. Gu, Jie Yang 0002 |
ICIP | 1 |
| 2013 | Superpixel based color contrast and color distribution driven salient object detection
Keren Fu, Chen Gong 0002, Jie Yang 0002, Yue Zhou 0005, Irene Y. H. Gu |
Signal Process. Image Commun. | 1 |
| 2012 | Salient Object Detection via Color Contrast and Color Distribution
Keren Fu, Chen Gong 0002, Jie Yang 0002, Yue Zhou 0005 |
ACCV (1) | 1 |