Kaifu Yang

dblp:19/9488 · also Kai-Fu Yang · DBLP profile ↗
← Back
39ranked-venue papers
9as first author
25since 2021 · last 2027
0000-0002-3696-5889ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021
YearPublicationVenuePosition
2027 A novel multi-instance learning framework for screening retinopathy of prematurity
Gui-Na Liu, Yubo Tan, Kaifu Yang, Yongjie Li 0001
Expert Syst. Appl.3
2026 Degraded image enhancement using learnable double opponency
abstract
Real-world environmental conditions such as low-light, haze, and underwater typically result in degraded images characterized by low contrast, color distortion, inconsistent lighting, and high noise levels. This paper introduces a novel approach, the Double Opponency-based Transformer network (DOT), to address common limitations like local details loss and spatial misalignment in existing image enhancement methods. Moreover, unlike the traditional encoder–decoder architectures based models, through decomposing input images into different color components and employing learnable double-opponent processing and non-bottleneck Transformer network, DOT captures intricate color relationships and enhances visual quality while maintaining reasonable global brightness and color accuracy. In the experiments, we demonstrate the effectiveness of the DOT model through extensive quantitative and qualitative evaluations against state-of-the-art methods on various datasets, including low-light, underwater, and hazy image enhancement.
Houwang Zhang, Kaifu Yang, Chong Wu 0007, Leanne Lai Chan
Eng. Appl. Artif. Intell.3
2026 A biological vision inspired framework for machine perception of illusory contours
Kaifu Yang, Hong-Zhi You
Pattern Recognit.2
2026 Retina adaptation network for low-light image enhancement
Houwang Zhang, Kaifu Yang, Yongjie Li 0001, Leanne Lai Chan
Pattern Recognit.3
2025 Nighttime Object Detection with Contextual Auxiliary Learning
Xiangrui Hu, Yongjie Li 0001, Kaifu Yang
ICIG (2)4
2025 Weakly Supervised Micro- and Macro-Expression Spotting Based on Multi-Level Consistency
abstract
Most micro- and macro-expression spotting methods in untrimmed videos suffer from the burden of video-wise collection and frame-wise annotation. Weakly supervised expression spotting (WES) based on video-level labels can potentially mitigate the complexity of frame-level annotation while achieving fine-grained frame-level spotting. However, we argue that existing weakly supervised methods are based on multiple instance learning (MIL) involving inter-modality, inter-sample, and inter-task gaps. The inter-sample gap is primarily from the sample distribution and duration. Therefore, we propose a novel and simple WES framework, MC-WES, using multi-consistency collaborative mechanisms that include modal-level saliency, video-level distribution, label-level duration and segment-level feature consistency strategies to implement fine frame-level spotting with only video-level labels to alleviate the above gaps and merge prior knowledge. The modal-level saliency consistency strategy focuses on capturing key correlations between raw images and optical flow. The video-level distribution consistency strategy utilizes the difference of sparsity in temporal distribution. The label-level duration consistency strategy exploits the difference in the duration of facial muscles. The segment-level feature consistency strategy emphasizes that features under the same labels maintain similarity. Experimental results on three challenging datasets-CAS(ME)$^{2}$2, CAS(ME)$^{3}$3, and SAMM-LV-demonstrate that MC-WES is comparable to state-of-the-art fully supervised methods.
Wang-Wang Yu, Kaifu Yang, Yongjie Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Self-supervised network for low-light traffic image enhancement based on deep noise and artifacts removal
abstract
In the intelligent transportation system (ITS), detecting vehicles and pedestrians in low-light conditions is challenging due to the low contrast between objects and the background. Recently, many works have enhanced low-light images using deep learning-based methods, but these methods require paired images during training, which are impractical to obtain in real-world traffic scenarios. Therefore, we propose a self-supervised network (SSN) for low-light traffic image enhancement that can be trained without paired images. To avoid amplifying noise and artifacts in the processed image during enhancement, we first proposed a denoising net to reduce the noise and artifacts in the input image. Then the processed image can be enhanced by the enhancement net. Considering the compression of the traffic image, we designed an artifacts removal net to improve the quality of the enhanced image. We proposed several effective and differential losses to make SSN trainable with low-light images only. To better integrate the extracted features from different levels in the network, we also proposed an attention module named the multi-head non-local block. In experiments, we evaluated SSN and other low-light image enhancement methods on two low-light traffic image sets: the Berkeley Deep Drive (BDD) dataset and the Hong Kong night-time multi-class vehicle (HK) dataset. The results indicated that SSN significantly improves upon other methods in visual comparison and some blind image quality metrics. We also conducted comparisons on classical ITS tasks like vehicle detection on the images enhanced by SSN and other methods, which further verified its effectiveness.
Houwang Zhang, Kaifu Yang, Yongjie Li 0001, Leanne Lai Chan
Comput. Vis. Image Underst.2
2024 Deep matched filtering for retinal vessel segmentation
Yubo Tan, Kaifu Yang, Shixuan Zhao 0001, Jianglan Wang, Longqian Liu, Yongjie Li 0001
Knowl. Based Syst.2
2024 LGSNet: A Two-Stream Network for Micro- and Macro-Expression Spotting With Background Modeling
abstract
Micro- and macro-expression spotting in an untrimmed video is a challenging task, due to the mass generation of false positive samples. Most existing methods localize higher response areas by extracting hand-crafted features or cropping specific regions from all or some key raw images. However, these methods either neglect the continuous temporal information or model the inherent human motion paradigms (background) as foreground. Consequently, we propose a novel two-stream network, named Local suppression and Global enhancement Spotting Network (LGSNet), which takes segment-level features from optical flow and videos as input. LGSNet adopts anchors to encode expression intervals and selects the encoded deviations as the object of optimization. Furthermore, we introduce a Temporal Multi-Receptive Field Feature Fusion Module (TMRF$^{3}$M) and a Local Suppression and Global Enhancement Module (LSGEM), which help spot short intervals more precisely and suppress background information. To further highlight the differences between positive and negative samples, we set up a large number of random pseudo ground truth intervals (background clips) on some discarded sliding windows to accomplish background clips modeling to counteract the effect of non-expressive face and head movements. Experimental results show that our proposed network achieves state-of-the-art performance on the CAS(ME)$^{2}$, CAS(ME)$^{3}$and SAMM-LV datasets.
Wang-Wang Yu, Kaifu Yang, Yongjie Li 0001
IEEE Trans. Affect. Comput.3
2024 Nighttime Thermal Infrared Image Colorization With Feedback-Based Object Appearance Learning
abstract
Stable imaging in adverse environments (e.g., total darkness) makes thermal infrared (TIR) cameras a prevalent option for night scene perception. However, the low contrast and lack of chromaticity of TIR images are detrimental to human interpretation and subsequent deployment of RGB-based vision algorithms. Therefore, it makes sense to colorize the nighttime TIR images by translating them into the corresponding daytime color images (NTIR2DC). Despite the impressive progress made in the NTIR2DC task, how to improve the translation performance of small object classes is under-explored. To address this problem, we propose a generative adversarial network incorporating feedback-based object appearance learning (FoalGAN). Specifically, an occlusion-aware mixup module and corresponding appearance consistency loss are proposed to reduce the context dependence of object translation. As a representative example of small objects in nighttime street scenes, we illustrate how to enhance the realism of traffic light by designing a traffic light appearance loss. To further improve the appearance learning of small objects, we devise a dual feedback learning strategy to selectively adjust the learning frequency of different samples. In addition, we provide pixel-level annotation for a subset of the Brno dataset, which can facilitate the research of NTIR image understanding under multiple weather conditions. Extensive experiments illustrate that the proposed FoalGAN is not only effective for appearance learning of small objects, but also outperforms other image translation methods in terms of semantic preservation and edge consistency for the NTIR2DC task. Compared with the state-of-the-art NTIR2DC approach, FoalGAN achieves at least 5.4% improvement in semantic consistency and at least 2% lead in edge consistency.
Fuya Luo, Yijun Cao, Kaifu Yang, Chang-Yong Xie, Yongjie Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Weakly Supervised Fixated Object Detection in Traffic Videos Based on Driver's Selective Attention Mechanism
abstract
Traffic scene perception has a significant impact on driving safety. Inexperienced or distracted drivers usually do not allocate enough attention to the objects closely related to the driving task, which causes potential road hazards. In contrast, experienced drivers pay close attention to the objects highly relevant to the driving task under the guidance of visual selective attention, thus achieving driving safety. However, apart from traffic saliency prediction, few existing works have integrated human driver’s perception with computer models to detect the objects attracting the attention of experienced drivers in traffic videos. In this work, we aim to detect these objects, specifically referred to as traffic fixated objects. To achieve this goal, a new eye-tracking-based video fixated object detection dataset (ET-VFOD) is firstly built, which can be as a benchmark for researchers interested in attention-inspired fixated object detection. Then, we propose a traffic video fixated object detection network named VFOD-Net. VFOD-Net decodes the information closely related to the driving task from the reference frames. The information is used as a top-down prior to modulate the model’s encoding process of the current frame, thus improving the detection performance. Considering the high cost of manual annotation, a weakly supervised traffic video fixated object detection pipeline is developed. Experimental results on the ET-VFOD dataset show that our proposed weakly supervised method achieves detection performance close to that of the fully supervised model, which verifies the effectiveness of the proposed method. Our work combines bottom-up and top-down attention to detect the vital objects in traffic videos from the perspective of human drivers, showing potential applications in intelligent driving, such as driver monitoring and warning systems. The dataset and code are available inhttps://github.com/YiShi701/VFOD_Net.
Yi Shi 0012, Long Qin 0002, Shixuan Zhao 0001, Kaifu Yang, Yuyong Cui
IEEE Trans. Circuits Syst. Video Technol.4
2024 Memory-Guided Collaborative Attention for Nighttime Thermal Infrared Image Colorization of Traffic Scenes
abstract
Robust imaging under challenging conditions, such as starlit nights, has broadened the adoption of thermal infrared (TIR) cameras for nighttime driving scenes. Given that TIR images are monochromatic, which makes them difficult to interpret by humans and limits the applicability of RGB-based algorithms, it is reasonable to perform colorization of nighttime TIR (NTIR) images by converting them into corresponding daytime color images (NTIR2DC). Despite the impressive results achieved by previous NTIR2DC methods, how to improve the colorization performance of small-sample categories without semantic annotation is under-explored. To address this issue, we propose a novel learning framework called Memory-guided cOllaboRative atteNtion Generative Adversarial Network (MornGAN), which is inspired by the analogical reasoning mechanisms of humans. Specifically, we first propose an online semantic distillation module to mine and refine the semantic cues of NTIR images. Then, a memory-guided sample selection strategy and adaptive collaborative attention loss are devised to enhance the semantic preservation of small-sample categories. Further, a new conditional gradient repair loss is introduced for reducing edge distortion during translation. Extensive experiments on the NTIR2DC task show that the proposed MornGAN significantly outperforms other image-to-image translation methods in terms of semantic preservation and edge consistency, which helps improve the object detection accuracy remarkably.
Fuya Luo, Yijun Cao, Kaifu Yang, Gang Wang 0031, Yongjie Li 0001
IEEE Trans. Intell. Transp. Syst.3
2024 Night-Time Vehicle Detection Based on Hierarchical Contextual Information
abstract
Night-time vehicle detection, which forms a basic component of the intelligent transportation system, is a topic of intense research interest with multifarious challenges. Due to the presence of low-light conditions, vehicles are typically indistinguishable from the background, and interference from light sources often arises in this complex environment. Existing widely used deep learning-based object detection models are designed for daytime scenarios and they have seldom considered these problems. Based on an investigation of current detection techniques and an analysis of the specific challenges of night-time vehicle detection, we propose a hierarchical contextual information (HCI) framework that can be used as a plug-and-play component to improve existing deep learning-based detection models under night-time conditions. Our HCI consists of three parts, an estimation branch, a segmentation branch and a detection branch, and can be applied to excavate hierarchical contextual clues and fuse them for the detection of vehicles in night-time environments. In each module, the context and predictions are extracted at the image-level, the pixel-level, and the object-level, respectively, and the results from each are complementary and beneficial to each other. Comprehensive experiments on two scenes from the Berkeley Deep Drive (BDD) dataset are presented to demonstrate the flexibility and generalization ability of our HCI. The significant improvements offered by HCI over main-stream detectors such as YOLOX, Faster RCNN, SSD, and EfficientDet also highlight the effectiveness of our approach for night-time vehicle detection.
Houwang Zhang, Kaifu Yang, Yongjie Li 0001, Leanne Lai Chan
IEEE Trans. Intell. Transp. Syst.2
2024 Retinal Layer Segmentation in OCT Images With Boundary Regression and Feature Polarization
abstract
The geometry of retinal layers is an important imaging feature for the diagnosis of some ophthalmic diseases. In recent years, retinal layer segmentation methods for optical coherence tomography (OCT) images have emerged one after another, and huge progress has been achieved. However, challenges due to interference factors such as noise, blurring, fundus effusion, and tissue artifacts remain in existing methods, primarily manifesting as intra-layer false positives and inter-layer boundary deviation. To solve these problems, we propose a method called Tightly combined Cross-Convolution and Transformer with Boundary regression and feature Polarization (TCCT-BP). This method uses a hybrid architecture of CNN and lightweight Transformer to improve the perception of retinal layers. In addition, a feature grouping and sampling method and the corresponding polarization loss function are designed to maximize the differentiation of the feature vectors of different retinal layers, and a boundary regression loss function is devised to constrain the retinal boundary distribution for a better fit to the ground truth. Extensive experiments on four benchmark datasets demonstrate that the proposed method achieves state-of-the-art performance in dealing with problems of false positives and boundary distortion. The proposed method ranked first in the OCT Layer Segmentation task of GOALS challenge held by MICCAI 2022. The source code is available at https://www.github.com/tyb311/TCCT.
Yubo Tan, Wen-Da Shen, Ming-Yuan Wu, Gui-Na Liu, Shixuan Zhao 0001, Yang Chen 0060, Kaifu Yang, Yongjie Li 0001
IEEE Trans. Medical Imaging7
2023 Orientation and Context Entangled Network for Retinal Vessel Segmentation
Xinxu Wei, Kaifu Yang, Danilo Bzdok, Yongjie Li 0001
Expert Syst. Appl.2
2023 Learning to Adapt to Light
Kaifu Yang, Shixuan Zhao 0001, Yongjie Li 0001
Int. J. Comput. Vis.1
2023 Correction to: Learning to Adapt to Light
Kaifu Yang, Shixuan Zhao 0001, Yongjie Li 0001
Int. J. Comput. Vis.1
2023 Learning generalized visual odometry using position-aware optical flow and geometric bundle adjustment
abstract
Recent visual odometry (VO) methods incorporating geometric algorithm into deep-learning architecture have shown outstanding performance on the challenging monocular VO task. Despite encouraging results are shown, previous methods ignore the requirement of generalization capability under noisy environment and various scenes. To address this challenging issue, this work first proposes a novel optical flow network (PANet). Compared with previous methods that predict optical flow as a direct regression task , our PANet computes optical flow by predicting it into the discrete position space with optical flow probability volume, and then converting it to optical flow. Next, we improve the bundle adjustment module to fit the self-supervised training pipeline by introducing multiple sampling, ego-motion initialization, dynamic damping factor adjustment, and Jacobi matrix weighting. In addition, a novel normalized photometric loss function is advanced to improve the depth estimation accuracy. The experiments show that the proposed system not only achieves comparable performance with other state-of-the-art self-supervised learning-based methods on the KITTI dataset, but also significantly improves the generalization capability compared with geometry-based, learning-based and hybrid VO systems on the noisy KITTI and the challenging outdoor (KAIST) scenes.
Yijun Cao, Fuya Luo, Chuan Lin 0003, Kaifu Yang, Yongjie Li 0001
Pattern Recognit.6
2022 A new representation of scene layout improves saliency detection in traffic scenes
De-Huai He, Kaifu Yang, Xue-Mei Wan, Fen Xiao, Yongjie Li 0001
Expert Syst. Appl.2
2022 Global-prior-guided fusion network for salient object detection
Kaifu Yang, Yongjie Li 0001
Expert Syst. Appl.2
2022 Contour-guided saliency detection with long-range interactions
Kaifu Yang, Si-Qin Liang, Yongjie Li 0001
Neurocomputing2
2022 A Fish Retina-Inspired Single Image Dehazing Method
abstract
Outdoor images are significantly degraded by bad weather conditions, such as fog and dust, which seriously limits efficient information extraction by computer vision applications. Most existing weather-degraded hazy image enhancement methods ignore the wavelength dependence of the scattering coefficient and therefore cannot well handle the colorized haze in which the medium transmission varies in different color channels. In this work, we propose a fish retina-inspired method to solve this problem. Fish have special retinal mechanisms for extracting useful information from underwater environments where photons are scattered and absorbed spatially introducing wavelength-dependent degradation of visibility. Due to the similarity between suspensions in water and aerosols in the air, these mechanisms should also be suitable for processing hazy scenes. The proposed method imitates bipolar cells with scene-dependent center-surround receptive fields to eliminate redundant information, ganglion cells with nonlinear processing to boost contrast, and the centrifugal pathway from amacrine cells to the horizontal cells to adaptively protect visual information from excessive lateral suppression. Compared with state-of-the-art methods on various images degraded by fog, haze, or dust, the proposed method shows quite competitive performance both qualitatively and quantitatively.
Yong-Bo Yu, Kaifu Yang, Yongjie Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Retinal Vessel Segmentation With Skeletal Prior and Contrastive Loss
abstract
The morphology of retinal vessels is closely associated with many kinds of ophthalmic diseases. Although huge progress in retinal vessel segmentation has been achieved with the advancement of deep learning, some challenging issues remain. For example, vessels can be disturbed or covered by other components presented in the retina (such as optic disc or lesions). Moreover, some thin vessels are also easily missed by current methods. In addition, existing fundus image datasets are generally tiny, due to the difficulty of vessel labeling. In this work, a new network called SkelCon is proposed to deal with these problems by introducing skeletal prior and contrastive loss. A skeleton fitting module is developed to preserve the morphology of the vessels and improve the completeness and continuity of thin vessels. A contrastive loss is employed to enhance the discrimination between vessels and background. In addition, a new data augmentation method is proposed to enrich the training samples and improve the robustness of the proposed model. Extensive validations were performed on several popular datasets (DRIVE, STARE, CHASE, and HRF), recently developed datasets (UoA-DR, IOSTAR, and RC-SLO), and some challenging clinical images (from RFMiD and JSIEC39 datasets). In addition, some specially designed metrics for vessel segmentation, including connectivity, overlapping area, consistency of vessel length, revised sensitivity, specificity, and accuracy were used for quantitative evaluation. The experimental results show that, the proposed model achieves state-of-the-art performance and significantly outperforms compared methods when extracting thin vessels in the regions of lesions or optic disc. Source code is available at https://www.github.com/tyb311/SkelCon.
Yubo Tan, Kaifu Yang, Shixuan Zhao 0001, Yongjie Li 0001
IEEE Trans. Medical Imaging2
2022 A Local and Global Feature Disentangled Network: Toward Classification of Benign-Malignant Thyroid Nodules From Ultrasound Image
abstract
Thyroid nodules are one of the most common nodular lesions. The incidence of thyroid cancer has increased rapidly in the past three decades and is one of the cancers with the highest incidence. As a non-invasive imaging modality, ultrasonography can identify benign and malignant thyroid nodules, and it can be used for large-scale screening. In this study, inspired by the domain knowledge of sonographers when diagnosing ultrasound images, a local and global feature disentangled network (LoGo-Net) is proposed to classify benign and malignant thyroid nodules. This model imitates the dual-pathway structure of human vision and establishes a new feature extraction method to improve the recognition performance of nodules. We use the tissue-anatomy disentangled (TAD) block to connect the dual pathways, which decouples the cues of local and global features based on the self-attention mechanism. To verify the effectiveness of the model, we constructed a large-scale dataset and conducted extensive experiments. The results show that our method achieves an accuracy of 89.33%, which has the potential to be used in the clinical practice of doctors, including early cancer screening procedures in remote or resource-poor areas.
Shixuan Zhao 0001, Yang Chen 0060, Kaifu Yang, Bu-Yun Ma, Yongjie Li 0001
IEEE Trans. Medical Imaging3
2021 Saliency Detection Inspired by Topological Perception Theory
Kaifu Yang, Fuya Luo, Yongjie Li 0001
Int. J. Comput. Vis.2
2020 A Biological Vision Inspired Framework for Image Enhancement in Poor Visibility Conditions
abstract
Image enhancement is an important pre-processing step for many computer vision applications especially regarding the scenes in poor visibility conditions. In this work, we develop a unified two-pathway model inspired by the biological vision, especially the early visual mechanisms, which contributes to image enhancement tasks including low dynamic range (LDR) image enhancement and high dynamic range (HDR) image tone mapping. Firstly, the input image is separated and sent into two visual pathways: structure-pathway and detail-pathway, corresponding to the M-and P-pathway in the early visual system, which code the low-and high-frequency visual information, respectively. In the structure-pathway, an extended biological normalization model is used to integrate the global and local luminance adaptation, which can handle the visual scenes with varying illuminations. On the other hand, the detail enhancement and local noise suppression are achieved in the detail-pathway based on local energy weighting. Finally, the outputs of structure-and detail-pathway are integrated to achieve the low-light image enhancement. In addition, the proposed model can also be used for tone mapping of HDR images with some fine-tuning steps. Extensive experiments on three datasets (two LDR image datasets and one HDR scene dataset) show that the proposed model can handle the visual enhancement tasks mentioned above efficiently and outperform the related state-of-the-art methods.
Kaifu Yang, Yongjie Li 0001
IEEE Trans. Image Process.1
2019 An Adaptive Method for Image Dynamic Range Adjustment
abstract
In this paper, we relate the operation of image dynamic range adjustment to the following two tasks: 1) for a high dynamic range (HDR) image, its dynamic range will be mapped to the available dynamic range of display devices and 2) for a low dynamic range (LDR) image, its distribution of intensity will be extended to adequately utilize the full dynamic range of display devices. The common goal of both tasks is to preserve or even enhance the details and improve the visibility of scenes when being matched to the available dynamic range of a display device. In this paper, we propose an efficient method for image dynamic range adjustment with three adaptive steps. First, according to the histogram of the luminance map separated from the given RGB image, two suitable Gamma functions are adaptively selected to separately adjust the luminance of the dark and bright components. Second, an adaptive fusion strategy is proposed to combine the two adjusted luminance maps in order to balance the enhancement of the details in different regions. Third, an adaptive luminance-dependent color restoration method is designed to combine the fused luminance map with the original color components to obtain more consistent color saturation between the images before and after dynamic range adjustment. Extensive experiments show that the proposed method can efficiently compress the dynamic range of HDR scenes with good contrast, clear details, and high structural fidelity of the original image appearance. In addition, the proposed method can also obtain promising performance when being used to enhance LDR nighttime images and greatly facilitates the object (car) detection in nighttime traffic scenes.
Kaifu Yang, Hulin Kuang, Chao-Yi Li, Yongjie Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Criteria to evaluate the fidelity of image enhancement by MSRCR
abstract
Image fidelity refers to the ability of a process to render an image accurately. As image enhancement algorithms have been developed in recent years, how to assess the performances of different image enhancement algorithms has become an important question. Some objective image quality assessment (IQA) methods have been proposed, but there is little research on the image fidelity evaluation when comparing the performances of enhancement algorithms. Therefore, the authors proposed a new image fidelity assessment framework consisting of three components: the information entropy fidelity, constituent fidelity and colour fidelity. To verify the rationality of the fidelity criteria, they used the popular IQA database (LIVE), and the results indicated that the method matched better with the subjective assessment. Then, they verified the effectiveness of their method with the famous technique of image enhancement: multi‐scale Retinex with colour restoration (MSRCR). The experimental results demonstrate MSRCR can improve the image quality, but it gives rise to obvious distortions. It is necessary to keep a moderate balance between image fidelity and image quality when they assess the enhanced images. Their results showed that the proposed objective fidelity index could provide an additional objective basis for the quality evaluation of image enhancement algorithms.
Shaobing Gao, Kaifu Yang
IET Image Process.4
2018 Bayes Saliency-Based Object Proposal Generator for Nighttime Traffic Images
abstract
Object proposal is one of the most key pre-processing steps for nighttime vehicle detection systems in intelligent transportation systems. However, most current object proposal methods are developed on daytime data sets, and these methods demonstrate unsatisfactory results when they are used on nighttime images. Therefore, this paper presents a novel Bayes saliency-based object proposal generator for nighttime RGB traffic images to generate a modest and accurate set of proposals, which are more likely to be vehicles for preceding vehicle detection. First, we propose a new Bayes saliency detection approach in which prior estimation, feature extraction, weight estimation, and Bayes rule are used to compute saliency maps. Then, we propose a simple but effective object proposal generator based on the Bayes saliency map. Multi-scale sliding window, proposal rejecting, scoring, and non-maximum suppression are combined to generate a modest and effective set of proposals. Experimental results demonstrate that our proposed approach generates a modest set of proposals and outperforms some state-of-the-art methods on nighttime images in terms of various evaluation metrics. Furthermore, our proposed object proposal approach can improve the detection performance and the speed of several state-of-the-art vehicle detection approaches.
Hulin Kuang, Kaifu Yang, Long Chen 0005, Yongjie Li 0001, Leanne Lai Chan, Hong Yan 0001
IEEE Trans. Intell. Transp. Syst.2
2016 A Unified Framework for Salient Structure Detection by Contour-Guided Visual Search
abstract
We define the task of salient structure (SS) detection to unify the saliency-related tasks, such as fixation prediction, salient object detection, and detection of other structures of interest in cluttered environments. To solve such SS detection tasks, a unified framework inspired by the two-pathway-based search strategy of biological vision is proposed in this paper. First, a contour-based spatial prior (CBSP) is extracted based on the layout of edges in the given scene along a fast non-selective pathway, which provides a rough, task-irrelevant, and robust estimation of the locations where the potential SSs are present. Second, another flow of local feature extraction is executed in parallel along the selective pathway. Finally, Bayesian inference is used to auto-weight and integrate the local cues guided by CBSP and to predict the exact locations of SSs. This model is invariant to the size and features of objects. The experimental results on six large datasets (three fixation prediction datasets and three salient object datasets) demonstrate that our system achieves competitive performance for SS detection (i.e., both the tasks of fixation prediction and salient object detection) compared with the state-of-the-art methods. In addition, our system also performs well for salient object construction from saliency maps and can be easily extended for salient edge detection.
Kaifu Yang, Chao-Yi Li, Yongjie Li 0001
IEEE Trans. Image Process.1
2016 Where Does the Driver Look? Top-Down-Based Saliency Detection in a Traffic Driving Environment
abstract
A traffic driving environment is a complex and dynamically changing scene. When driving, drivers always allocate their attention to the most important and salient areas or targets. Traffic saliency detection, which computes the salient and prior areas or targets in a specific driving environment, is an indispensable part of intelligent transportation systems and could be useful in supporting autonomous driving, traffic sign detection, driving training, car collision warning, and other tasks. Recently, advances in visual attention models have provided substantial progress in describing eye movements over simple stimuli and tasks such as free viewing or visual search. However, to date, there exists no computational framework that can accurately mimic a driver's gaze behavior and saliency detection in a complex traffic driving environment. In this paper, we analyzed the eye-tracking data of 40 subjects consisted of nondrivers and experienced drivers when viewing 100 traffic images. We found that a driver's attention was mostly concentrated on the end of the road in front of the vehicle. We proposed that the vanishing point of the road can be regarded as valuable top-down guidance in a traffic saliency detection model. Subsequently, we build a framework of a classic bottom-up and top-down combined traffic saliency detection model. The results show that our proposed vanishing-point-based top-down model can effectively simulate a driver's attention areas in a driving environment.
Tao Deng 0002, Kaifu Yang, Yongjie Li 0001
IEEE Trans. Intell. Transp. Syst.2
2015 Efficient illuminant estimation for color constancy using grey pixels
abstract
Illuminant estimation is a key step for computational color constancy. Instead of using the grey world or grey edge assumptions, we propose in this paper a novel method for illuminant estimation by using the information of grey pixels detected in a given color-biased image. The underlying hypothesis is that most of the natural images include some detectable pixels that are at least approximately grey, which can be reliably utilized for illuminant estimation. We first validate our assumption through comprehensive statistical evaluation on diverse collection of datasets and then put forward a novel grey pixel detection method based on the illuminant-invariant measure (IIM) in three logarithmic color channels. Then the light source color of a scene can be easily estimated from the detected grey pixels. Experimental results on four benchmark datasets (three recorded under single illuminant and one under multiple illuminants) show that the proposed method outperforms most of the state-of-the-art color constancy approaches with the inherent merit of low computational cost.
Kaifu Yang, Shao-Bing Gao, Yongjie Li 0001
CVPR1
2015 Color Constancy Using Double-Opponency
abstract
The double-opponent (DO) color-sensitive cells in the primary visual cortex (V1) of the human visual system (HVS) have long been recognized as the physiological basis of color constancy. In this work we propose a new color constancy model by imitating the functional properties of the HVS from the single-opponent (SO) cells in the retina to the DO cells in V1 and the possible neurons in the higher visual cortexes. The idea behind the proposed double-opponency based color constancy (DOCC) model originates from the substantial observation that the color distribution of the responses of DO cells to the color-biased images coincides well with the vector denoting the light source color. Then the illuminant color is easily estimated by pooling the responses of DO cells in separate channels in LMS space with the pooling mechanism of sum or max. Extensive evaluations on three commonly used datasets, including the test with the dataset dependent optimal parameters, as well as the intra- and inter-dataset cross validation, show that our physiologically inspired DOCC model can produce quite competitive results in comparison to the state-of-the-art approaches, but with a relative simple implementation and without requiring fine-tuning of the method for each different dataset.
Shao-Bing Gao, Kaifu Yang, Chao-Yi Li, Yongjie Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Boundary Detection Using Double-Opponency and Spatial Sparseness Constraint
abstract
Brightness and color are two basic visual features integrated by the human visual system (HVS) to gain a better understanding of color natural scenes. Aiming to combine these two cues to maximize the reliability of boundary detection in natural scenes, we propose a new framework based on the color-opponent mechanisms of a certain type of color-sensitive double-opponent (DO) cells in the primary visual cortex (V1) of HVS. This type of DO cells has oriented receptive field with both chromatically and spatially opponent structure. The proposed framework is a feedforward hierarchical model, which has direct counterpart to the color-opponent mechanisms involved in from the retina to V1. In addition, we employ the spatial sparseness constraint (SSC) of neural responses to further suppress the unwanted edges of texture elements. Experimental results show that the DO cells we modeled can flexibly capture both the structured chromatic and achromatic boundaries of salient objects in complex scenes when the cone inputs to DO cells are unbalanced. Meanwhile, the SSC operator further improves the performance by suppressing redundant texture edges. With competitive contour detection accuracy, the proposed model has the additional advantage of quite simple implementation with low computational cost.
Kaifu Yang, Shao-Bing Gao, Ce-Feng Guo, Chao-Yi Li, Yongjie Li 0001
IEEE Trans. Image Process.1
2014 Efficient Color Constancy with Local Surface Reflectance Statistics
Shaobing Gao, Wangwang Han, Kaifu Yang, Chaoyi Li, Yongjie Li 0001
ECCV (2)3
2014 Multifeature-Based Surround Inhibition Improves Contour Detection in Natural Images
abstract
To effectively perform visual tasks like detecting contours, the visual system normally needs to integrate multiple visual features. Sufficient physiological studies have revealed that for a large number of neurons in the primary visual cortex (V1) of monkeys and cats, neuronal responses elicited by the stimuli placed within the classical receptive field (CRF) are substantially modulated, normally inhibited, when difference exists between the CRF and its surround, namely, non-CRF, for various local features. The exquisite sensitivity of V1 neurons to the center-surround stimulus configuration is thought to serve important perceptual functions, including contour detection. In this paper, we propose a biologically motivated model to improve the performance of perceptually salient contour detection. The main contribution is the multifeature-based center-surround framework, in which the surround inhibition weights of individual features, including orientation, luminance, and luminance contrast, are combined according to a scale-guided strategy, and the combined weights are then used to modulate the final surround inhibition of the neurons. The performance was compared with that of single-cue-based models and other existing methods (especially other biologically motivated ones). The results show that combining multiple cues can substantially improve the performance of contour detection compared with the models using single cue. In general, luminance and luminance contrast contribute much more than orientation to the specific task of contour extraction, at least in gray-scale natural images.
Kaifu Yang, Chao-Yi Li, Yongjie Li 0001
IEEE Trans. Image Process.1
2013 Efficient Color Boundary Detection with Color-Opponent Mechanisms
abstract
Color information plays an important role in better understanding of natural scenes by at least facilitating discriminating boundaries of objects or areas. In this study, we propose a new framework for boundary detection in complex natural scenes based on the color-opponent mechanisms of the visual system. The red-green and blue-yellow color opponent channels in the human visual system are regarded as the building blocks for various color perception tasks such as boundary detection. The proposed framework is a feed forward hierarchical model, which has direct counterpart to the color-opponent mechanisms involved in from the retina to the primary visual cortex (V1). Results show that our simple framework has excellent ability to flexibly capture both the structured chromatic and achromatic boundaries in complex scenes.
Kaifu Yang, Shaobing Gao, Chaoyi Li, Yongjie Li 0001
CVPR1
2013 A Color Constancy Model with Double-Opponency Mechanisms
abstract
The double-opponent color-sensitive cells in the primary visual cortex (V1) of the human visual system (HVS) have long been recognized as the physiological basis of color constancy. We introduce a new color constancy model by imitating the functional properties of the HVS from the retina to the double-opponent cells in V1. The idea behind the model originates from the observation that the color distribution of the responses of double-opponent cells to the input color-biased images coincides well with the light source direction. Then the true illuminant color of a scene is easily estimated by searching for the maxima of the separate RGB channels of the responses of double-opponent cells in the RGB space. Our systematical experimental evaluations on two commonly used image datasets show that the proposed model can produce competitive results in comparison to the complex state-of-the-art approaches, but with a simple implementation and without the need for training.
Shaobing Gao, Kaifu Yang, Chaoyi Li, Yongjie Li 0001
ICCV2
2011 Contour detection based on a non-classical receptive field model with butterfly-shaped inhibition subregions
Chi Zeng, Yongjie Li 0001, Kaifu Yang, Chaoyi Li
Neurocomputing3