Yongjie Li 0001

dblp:37/1386-1 · also Yong-Jie Li 0001 · DBLP profile ↗
← Back
63ranked-venue papers
1as first author
42since 2021 · last 2027
0000-0002-7395-3131ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 1 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2027 A novel multi-instance learning framework for screening retinopathy of prematurity
Gui-Na Liu, Yubo Tan, Kaifu Yang, Yongjie Li 0001
Expert Syst. Appl.5
2026 Retina adaptation network for low-light image enhancement
Houwang Zhang, Kaifu Yang, Yongjie Li 0001, Leanne Lai Chan
Pattern Recognit.4
2026 Infrared and Visible Image Fusion Using Bimodal Neuron and Dynamic Receptive Field Mechanisms
abstract
Infrared and visible image fusion (IVIF) significantly enhances scene interpretation by integrating broad-spectrum information. Drawing inspiration from specific snakes that possess an evolutionarily optimized bimodal sensory system capable of parallel processing infrared and visible radiation, we propose a novel IVIF framework incorporating two key elements: nonlinear cross-modal interactions across six distinct classes of snake bimodal neurons and dynamic center-surround receptive field organization. These biological principles are mathematically formalized and integrated within a deep neural network (DNN), optimized through an object detection region-guided loss and a frequency-dependent fusion loss that enable data-driven fusion strategy learning. Experimental results demonstrate that the optimized model effectively emulates the infrared-visible information integration observed in snake bimodal neurons. Critically, the nonlinear bimodal neurons capture a significantly greater amount of edge information and finer mid-to-high-frequency details, which are essential for the subsequent reconstruction of the fused image. Furthermore, a comprehensive evaluation of visual quality, encompassing both qualitative and quantitative assessments on six datasets, along with extensive object detection and semantic segmentation experiments using the fused images in both daytime and nighttime scenarios, demonstrates that our model outperforms traditional biologically-inspired IVIF algorithms, achieving performance comparable to SOTA DNN-based methods. The code and weights are available at https://github.com/rwerwer2024/SBNF.
Shaobing Gao, Minjie Tan, Shun Lv, Yiguang Liu, Yongjie Li 0001
IEEE Trans. Image Process.5
2025 Nighttime Object Detection with Contextual Auxiliary Learning
Xiangrui Hu, Yongjie Li 0001, Kaifu Yang
ICIG (2)3
2025 Contrastive Hierarchical Graph Based Multiple Instance Learning for Fundus Screening
Yubo Tan, Shiye Wang, Wen-Da Shen, Yongjie Li 0001
ICIG (1)4
2025 A good teacher learns while teaching: Heterogeneous architectural knowledge distillation for fast MRI reconstruction
Chenghao Qiu, Yongjie Li 0001
Knowl. Based Syst.3
2025 Weakly Supervised Micro- and Macro-Expression Spotting Based on Multi-Level Consistency
abstract
Most micro- and macro-expression spotting methods in untrimmed videos suffer from the burden of video-wise collection and frame-wise annotation. Weakly supervised expression spotting (WES) based on video-level labels can potentially mitigate the complexity of frame-level annotation while achieving fine-grained frame-level spotting. However, we argue that existing weakly supervised methods are based on multiple instance learning (MIL) involving inter-modality, inter-sample, and inter-task gaps. The inter-sample gap is primarily from the sample distribution and duration. Therefore, we propose a novel and simple WES framework, MC-WES, using multi-consistency collaborative mechanisms that include modal-level saliency, video-level distribution, label-level duration and segment-level feature consistency strategies to implement fine frame-level spotting with only video-level labels to alleviate the above gaps and merge prior knowledge. The modal-level saliency consistency strategy focuses on capturing key correlations between raw images and optical flow. The video-level distribution consistency strategy utilizes the difference of sparsity in temporal distribution. The label-level duration consistency strategy exploits the difference in the duration of facial muscles. The segment-level feature consistency strategy emphasizes that features under the same labels maintain similarity. Experimental results on three challenging datasets-CAS(ME)$^{2}$2, CAS(ME)$^{3}$3, and SAMM-LV-demonstrate that MC-WES is comparable to state-of-the-art fully supervised methods.
Wang-Wang Yu, Kaifu Yang, Yongjie Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 JOANet: An Integrated Joint Optimization Architecture Making Medical Image Segmentation Really Helped by Super-Resolution Pre-Processing
abstract
Conventional computer vision pipelines typically treat low-level enhancement and high-level semantic tasks as isolated processes, focusing on optimizing enhancement for perceptual quality rather than computational utility, neglecting semantic task requirements. To bridge this gap, this paper proposes an integrated joint optimization architecture that aligns the objectives of enhancement tasks with the practical needs of semantic tasks. Specifically, the architecture ensures that medical image segmentation (the semantic task) benefits directly from super-resolution pre-processing (the enhancement task). This integrated architecture fundamentally differs from conventional sequential frameworks by enabling joint training of super-resolution and segmentation networks. Guided by its own content reconstruction loss and semantic loss transferred from segmentation, the super-resolution network prioritizes semantically significant regions for segmentation-driven reconstruction. Comprehensive comparative and ablation studies demonstrate that the network, trained jointly, markedly enhances segmentation performance in low-resolution images, even outperforming those directly from referenced high-resolution images. The code is available at https://github.com/kldys/JOANet.
Chenghao Qiu, Yongjie Li 0001
IEEE Trans. Image Process.3
2024 Self-supervised network for low-light traffic image enhancement based on deep noise and artifacts removal
abstract
In the intelligent transportation system (ITS), detecting vehicles and pedestrians in low-light conditions is challenging due to the low contrast between objects and the background. Recently, many works have enhanced low-light images using deep learning-based methods, but these methods require paired images during training, which are impractical to obtain in real-world traffic scenarios. Therefore, we propose a self-supervised network (SSN) for low-light traffic image enhancement that can be trained without paired images. To avoid amplifying noise and artifacts in the processed image during enhancement, we first proposed a denoising net to reduce the noise and artifacts in the input image. Then the processed image can be enhanced by the enhancement net. Considering the compression of the traffic image, we designed an artifacts removal net to improve the quality of the enhanced image. We proposed several effective and differential losses to make SSN trainable with low-light images only. To better integrate the extracted features from different levels in the network, we also proposed an attention module named the multi-head non-local block. In experiments, we evaluated SSN and other low-light image enhancement methods on two low-light traffic image sets: the Berkeley Deep Drive (BDD) dataset and the Hong Kong night-time multi-class vehicle (HK) dataset. The results indicated that SSN significantly improves upon other methods in visual comparison and some blind image quality metrics. We also conducted comparisons on classical ITS tasks like vehicle detection on the images enhanced by SSN and other methods, which further verified its effectiveness.
Houwang Zhang, Kaifu Yang, Yongjie Li 0001, Leanne Lai Chan
Comput. Vis. Image Underst.3
2024 Biologically inspired image invariance guided illuminant estimation using shallow and deep models
abstract
Estimating the illuminant from a color-biased image is an ill-posed problem without prior information or invariance about the surfaces of a scene. Based on the classical image formation model, we have developed a heuristic approach to obtain the illuminant color by computing the ratio of the average of all pixels over the color-biased scene to that over the roughly recovered scene obtained by local normalization in each channel. The computed ratio represents an estimated invariance across color channels (IACC), ranging between the average reflectance of surfaces and the maximum reflectance of surfaces of a scene, modulated by the illuminant color. This work builds a mathematical foundation for IACC and explains why it is suitable for illuminant estimation. The core discovery is that the magnitude relationship of the average reflectances of surfaces between any two color channels is opposite to that of the maximum reflectances of surfaces for most natural scenes. As a result, we have designed two approaches for explicitly learning IACC of an image, resulting in very accurate illuminant estimation. The first approach involves a novel shallow model based on diagonal or non-diagonal matrices, together with the learned model parameters, to improve IACC performance. The second approach applies IACC as a constraint to optimize a novel deep learning approach, which has achieved state-of-the-art performance on two benchmarks. An interesting finding is that the output of the learned network, constrained only by IACC loss, provides a coarse estimation of intrinsic images such as albedo from the input color-biased image
Shaobing Gao, Liangtian He, Yongjie Li 0001
Expert Syst. Appl.3
2024 LVP-net: A deep network of learning visual pathway for edge detection
Chuan Lin 0003, Fuzhang Li, Yijun Cao, Yongjie Li 0001
Image Vis. Comput.5
2024 Deep matched filtering for retinal vessel segmentation
Yubo Tan, Kaifu Yang, Shixuan Zhao 0001, Jianglan Wang, Longqian Liu, Yongjie Li 0001
Knowl. Based Syst.6
2024 LGSNet: A Two-Stream Network for Micro- and Macro-Expression Spotting With Background Modeling
abstract
Micro- and macro-expression spotting in an untrimmed video is a challenging task, due to the mass generation of false positive samples. Most existing methods localize higher response areas by extracting hand-crafted features or cropping specific regions from all or some key raw images. However, these methods either neglect the continuous temporal information or model the inherent human motion paradigms (background) as foreground. Consequently, we propose a novel two-stream network, named Local suppression and Global enhancement Spotting Network (LGSNet), which takes segment-level features from optical flow and videos as input. LGSNet adopts anchors to encode expression intervals and selects the encoded deviations as the object of optimization. Furthermore, we introduce a Temporal Multi-Receptive Field Feature Fusion Module (TMRF$^{3}$M) and a Local Suppression and Global Enhancement Module (LSGEM), which help spot short intervals more precisely and suppress background information. To further highlight the differences between positive and negative samples, we set up a large number of random pseudo ground truth intervals (background clips) on some discarded sliding windows to accomplish background clips modeling to counteract the effect of non-expressive face and head movements. Experimental results show that our proposed network achieves state-of-the-art performance on the CAS(ME)$^{2}$, CAS(ME)$^{3}$and SAMM-LV datasets.
Wang-Wang Yu, Kaifu Yang, Yongjie Li 0001
IEEE Trans. Affect. Comput.5
2024 Nighttime Thermal Infrared Image Colorization With Feedback-Based Object Appearance Learning
abstract
Stable imaging in adverse environments (e.g., total darkness) makes thermal infrared (TIR) cameras a prevalent option for night scene perception. However, the low contrast and lack of chromaticity of TIR images are detrimental to human interpretation and subsequent deployment of RGB-based vision algorithms. Therefore, it makes sense to colorize the nighttime TIR images by translating them into the corresponding daytime color images (NTIR2DC). Despite the impressive progress made in the NTIR2DC task, how to improve the translation performance of small object classes is under-explored. To address this problem, we propose a generative adversarial network incorporating feedback-based object appearance learning (FoalGAN). Specifically, an occlusion-aware mixup module and corresponding appearance consistency loss are proposed to reduce the context dependence of object translation. As a representative example of small objects in nighttime street scenes, we illustrate how to enhance the realism of traffic light by designing a traffic light appearance loss. To further improve the appearance learning of small objects, we devise a dual feedback learning strategy to selectively adjust the learning frequency of different samples. In addition, we provide pixel-level annotation for a subset of the Brno dataset, which can facilitate the research of NTIR image understanding under multiple weather conditions. Extensive experiments illustrate that the proposed FoalGAN is not only effective for appearance learning of small objects, but also outperforms other image translation methods in terms of semantic preservation and edge consistency for the NTIR2DC task. Compared with the state-of-the-art NTIR2DC approach, FoalGAN achieves at least 5.4% improvement in semantic consistency and at least 2% lead in edge consistency.
Fuya Luo, Yijun Cao, Kaifu Yang, Chang-Yong Xie, Yongjie Li 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Memory-Guided Collaborative Attention for Nighttime Thermal Infrared Image Colorization of Traffic Scenes
abstract
Robust imaging under challenging conditions, such as starlit nights, has broadened the adoption of thermal infrared (TIR) cameras for nighttime driving scenes. Given that TIR images are monochromatic, which makes them difficult to interpret by humans and limits the applicability of RGB-based algorithms, it is reasonable to perform colorization of nighttime TIR (NTIR) images by converting them into corresponding daytime color images (NTIR2DC). Despite the impressive results achieved by previous NTIR2DC methods, how to improve the colorization performance of small-sample categories without semantic annotation is under-explored. To address this issue, we propose a novel learning framework called Memory-guided cOllaboRative atteNtion Generative Adversarial Network (MornGAN), which is inspired by the analogical reasoning mechanisms of humans. Specifically, we first propose an online semantic distillation module to mine and refine the semantic cues of NTIR images. Then, a memory-guided sample selection strategy and adaptive collaborative attention loss are devised to enhance the semantic preservation of small-sample categories. Further, a new conditional gradient repair loss is introduced for reducing edge distortion during translation. Extensive experiments on the NTIR2DC task show that the proposed MornGAN significantly outperforms other image-to-image translation methods in terms of semantic preservation and edge consistency, which helps improve the object detection accuracy remarkably.
Fuya Luo, Yijun Cao, Kaifu Yang, Gang Wang 0031, Yongjie Li 0001
IEEE Trans. Intell. Transp. Syst.5
2024 Night-Time Vehicle Detection Based on Hierarchical Contextual Information
abstract
Night-time vehicle detection, which forms a basic component of the intelligent transportation system, is a topic of intense research interest with multifarious challenges. Due to the presence of low-light conditions, vehicles are typically indistinguishable from the background, and interference from light sources often arises in this complex environment. Existing widely used deep learning-based object detection models are designed for daytime scenarios and they have seldom considered these problems. Based on an investigation of current detection techniques and an analysis of the specific challenges of night-time vehicle detection, we propose a hierarchical contextual information (HCI) framework that can be used as a plug-and-play component to improve existing deep learning-based detection models under night-time conditions. Our HCI consists of three parts, an estimation branch, a segmentation branch and a detection branch, and can be applied to excavate hierarchical contextual clues and fuse them for the detection of vehicles in night-time environments. In each module, the context and predictions are extracted at the image-level, the pixel-level, and the object-level, respectively, and the results from each are complementary and beneficial to each other. Comprehensive experiments on two scenes from the Berkeley Deep Drive (BDD) dataset are presented to demonstrate the flexibility and generalization ability of our HCI. The significant improvements offered by HCI over main-stream detectors such as YOLOX, Faster RCNN, SSD, and EfficientDet also highlight the effectiveness of our approach for night-time vehicle detection.
Houwang Zhang, Kaifu Yang, Yongjie Li 0001, Leanne Lai Chan
IEEE Trans. Intell. Transp. Syst.3
2024 Retinal Layer Segmentation in OCT Images With Boundary Regression and Feature Polarization
abstract
The geometry of retinal layers is an important imaging feature for the diagnosis of some ophthalmic diseases. In recent years, retinal layer segmentation methods for optical coherence tomography (OCT) images have emerged one after another, and huge progress has been achieved. However, challenges due to interference factors such as noise, blurring, fundus effusion, and tissue artifacts remain in existing methods, primarily manifesting as intra-layer false positives and inter-layer boundary deviation. To solve these problems, we propose a method called Tightly combined Cross-Convolution and Transformer with Boundary regression and feature Polarization (TCCT-BP). This method uses a hybrid architecture of CNN and lightweight Transformer to improve the perception of retinal layers. In addition, a feature grouping and sampling method and the corresponding polarization loss function are designed to maximize the differentiation of the feature vectors of different retinal layers, and a boundary regression loss function is devised to constrain the retinal boundary distribution for a better fit to the ground truth. Extensive experiments on four benchmark datasets demonstrate that the proposed method achieves state-of-the-art performance in dealing with problems of false positives and boundary distortion. The proposed method ranked first in the OCT Layer Segmentation task of GOALS challenge held by MICCAI 2022. The source code is available at https://www.github.com/tyb311/TCCT.
Yubo Tan, Wen-Da Shen, Ming-Yuan Wu, Gui-Na Liu, Shixuan Zhao 0001, Yang Chen 0060, Kaifu Yang, Yongjie Li 0001
IEEE Trans. Medical Imaging8
2023 Toward Better SSIM Loss for Unsupervised Monocular Depth Estimation
Yijun Cao, Fuya Luo, Yongjie Li 0001
ICIG (1)3
2023 Orientation and Context Entangled Network for Retinal Vessel Segmentation
Xinxu Wei, Kaifu Yang, Danilo Bzdok, Yongjie Li 0001
Expert Syst. Appl.4
2023 Learning to Adapt to Light
Kaifu Yang, Shixuan Zhao 0001, Yongjie Li 0001
Int. J. Comput. Vis.6
2023 Correction to: Learning to Adapt to Light
Kaifu Yang, Shixuan Zhao 0001, Yongjie Li 0001
Int. J. Comput. Vis.6
2023 Learning generalized visual odometry using position-aware optical flow and geometric bundle adjustment
abstract
Recent visual odometry (VO) methods incorporating geometric algorithm into deep-learning architecture have shown outstanding performance on the challenging monocular VO task. Despite encouraging results are shown, previous methods ignore the requirement of generalization capability under noisy environment and various scenes. To address this challenging issue, this work first proposes a novel optical flow network (PANet). Compared with previous methods that predict optical flow as a direct regression task , our PANet computes optical flow by predicting it into the discrete position space with optical flow probability volume, and then converting it to optical flow. Next, we improve the bundle adjustment module to fit the self-supervised training pipeline by introducing multiple sampling, ego-motion initialization, dynamic damping factor adjustment, and Jacobi matrix weighting. In addition, a novel normalized photometric loss function is advanced to improve the depth estimation accuracy. The experiments show that the proposed system not only achieves comparable performance with other state-of-the-art self-supervised learning-based methods on the KITTI dataset, but also significantly improves the generalization capability compared with geometry-based, learning-based and hybrid VO systems on the noisy KITTI and the challenging outdoor (KAIST) scenes.
Yijun Cao, Fuya Luo, Chuan Lin 0003, Kaifu Yang, Yongjie Li 0001
Pattern Recognit.7
2023 Hierarchical nearest neighbor descent, in-tree, and clustering
abstract
Recently, we have proposed a physically-inspired graph-theoretical method, called the Nearest Descent (ND), which is capable of organizing a dataset into an in-tree graph structure. Due to some beautiful and effective features, the constructed in-tree proves well-suited for data clustering. Although there exist some undesired edges (i.e., the inter-cluster edges) in this in-tree, those edges are usually very distinguishable, in sharp contrast to the cases in the famous Minimal Spanning Tree (MST). Here, we propose another graph-theoretical method, called the Hierarchical Nearest Neighbor Descent (HNND). Like ND, HNND also organizes a dataset into an in-tree, but in a more efficient way. Consequently, HNND-based clustering (HNND-C) is more efficient than ND-based clustering (ND-C) as well. This is well proved by the experimental results on five high-dimensional and large-size mass cytometry datasets. The experimental results also show that HNND-C achieves overall better performance than some state-of-the-art clustering methods.
Teng Qiu, Yongjie Li 0001
Pattern Recognit.2
2023 Unsupervised Visual Odometry and Action Integration for PointGoal Navigation in Indoor Environment
abstract
PointGoal navigation in indoor environment is a fundamental task for personal robots to navigate to a specified point. Recent studies solved this PointGoal navigation task with near-perfect success rate in photo-realistically simulated environments, under the assumptions with noiseless actuation and most importantly, perfect localization with GPS and compass sensors. However, accurate GPS signalis difficult to be obtained in real indoor environment. To improve the PointGoal navigation accuracy without GPS signal, we use visual odometry (VO) and propose a novel action integration module (AIM) trained in unsupervised manner. Sepecifically, unsupervised VO computes the relative pose of the agent from the re-projection error of two adjacent frames, and then replaces the accurate GPS signal with the path integration. The pseudo position estimated by VO is used to train action integration which assists agent to update their internal perception of location and helps improve the success rate of navigation. The training and inference process only use RGB, depth, collision as well as self-action information. The experiments show that the proposed system achieves satisfactory results and outperforms the partially supervised learning algorithms on the popular Gibson dataset.
Yijun Cao, Fuya Luo, Chuan Lin 0003, Yongjie Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Fast LDP-MST: An Efficient Density-Peak-Based Clustering Method for Large-Size Datasets
abstract
Recently, a new density-peak-based clustering method, called clustering with local density peaks-based minimum spanning tree (LDP-MST), was proposed, which has several attractive merits, e.g., being able to detect arbitrarily shaped clusters and not very sensitive to noise and parameters. Nevertheless, we also found the limitation of LDP-MST in efficiency. Specifically, LDP-MST has$O(N\log N+M^{2})$time, where$N$denotes the dataset size and$M$is an intermediate variable denoting the number of local density peaks. As our experimental results reveal, when processing large-size datasets, the value of$M$could be very large and consequently those steps of LDP-MST involving$O(M^{2})$time term would be time-consuming. And in the worst case, the value of$M$could be very close to that of$N$, which means that the time complexity of LDP-MST could be$O(N^{2})$in the worst case of$M$. In this study, we use more efficient algorithms to implement those steps of LDP-MST that involve the$O(M^{2})$time term such that the proposed method, Fast LDP-MST, has$O(N\log N)$time complexity even if$M\approx N$. Our experiments demonstrate that Fast LDP-MST is overall more efficient than LDP-MST on large-size datasets, without sacrificing the merits of LDP-MST in effectiveness, robustness, and user-friendliness.
Teng Qiu, Yongjie Li 0001
IEEE Trans. Knowl. Data Eng.2
2022 Deep Pneumonia: Attention-Based Contrastive Learning for Class-Imbalanced Pneumonia Lesion Recognition in Chest X-rays
abstract
Computer-aided X-ray pneumonia lesion recognition is important for accurate diagnosis of pneumonia. With the emergence of deep learning, the identification accuracy of pneumonia has been greatly improved, but there are still some challenges due to the fuzzy appearance of chest X-rays. In this paper, we propose a deep learning framework named Attention-Based Contrastive Learning for Class-Imbalanced X-Ray Pneumonia Lesion Recognition (denoted as Deep Pneumonia). We adopt self-supervised contrastive learning strategy to pre-train the model without using extra pneumonia data for fully mining the limited available dataset. In order to leverage the location information of the lesion area that the doctor has painstakingly marked, we propose mask-guided hard attention strategy and feature learning with contrastive regularization strategy which are applied on the attention map and the extracted features respectively to guide the model to focus more attention on the lesion area where contains more discriminative features for improving the recognition performance. In addition, we adopt Class-Balanced Loss instead of traditional Cross-Entropy as the loss function of classification to tackle the problem of serious class imbalance between different classes of pneumonia in the dataset. The experimental results show that our proposed framework can be used as a reliable computer-aided pneumonia diagnosis system to assist doctors to better diagnose pneumonia cases accurately.
Xinxu Wei, Xiangke Niu, Yongjie Li 0001
IEEE Big Data4
2022 TSN-CA: A Two-Stage Network with Channel Attention for Low-Light Image Enhancement
Xinxu Wei, Yongjie Li 0001
ICANN (3)3
2022 A new representation of scene layout improves saliency detection in traffic scenes
De-Huai He, Kaifu Yang, Xue-Mei Wan, Fen Xiao, Yongjie Li 0001
Expert Syst. Appl.6
2022 Global-prior-guided fusion network for salient object detection
Kaifu Yang, Yongjie Li 0001
Expert Syst. Appl.3
2022 Contour-guided saliency detection with long-range interactions
Kaifu Yang, Si-Qin Liang, Yongjie Li 0001
Neurocomputing4
2022 A Fish Retina-Inspired Single Image Dehazing Method
abstract
Outdoor images are significantly degraded by bad weather conditions, such as fog and dust, which seriously limits efficient information extraction by computer vision applications. Most existing weather-degraded hazy image enhancement methods ignore the wavelength dependence of the scattering coefficient and therefore cannot well handle the colorized haze in which the medium transmission varies in different color channels. In this work, we propose a fish retina-inspired method to solve this problem. Fish have special retinal mechanisms for extracting useful information from underwater environments where photons are scattered and absorbed spatially introducing wavelength-dependent degradation of visibility. Due to the similarity between suspensions in water and aerosols in the air, these mechanisms should also be suitable for processing hazy scenes. The proposed method imitates bipolar cells with scene-dependent center-surround receptive fields to eliminate redundant information, ganglion cells with nonlinear processing to boost contrast, and the centrifugal pathway from amacrine cells to the horizontal cells to adaptively protect visual information from excessive lateral suppression. Compared with state-of-the-art methods on various images degraded by fog, haze, or dust, the proposed method shows quite competitive performance both qualitatively and quantitatively.
Yong-Bo Yu, Kaifu Yang, Yongjie Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Thermal Infrared Image Colorization for Nighttime Driving Scenes With Top-Down Guided Attention
abstract
Benefitting from insensitivity to light and high penetration of foggy environments, infrared cameras are widely used for sensing in nighttime traffic scenes. However, the low contrast and lack of chromaticity of thermal infrared (TIR) images hinder the human interpretation and portability of high-level computer vision algorithms. Colorization to translate a nighttime TIR image into a daytime color (NTIR2DC) image may be a promising way to facilitate nighttime scene perception. Despite recent impressive advances in image translation, semantic encoding entanglement and geometric distortion in the NTIR2DC task remain under-addressed. Hence, we propose a toP-down attEntion And gRadient aLignment based generative adversarial network, referred to as PearlGAN. A top-down guided attention module and an elaborate attentional loss are first designed to reduce the semantic encoding ambiguity during translation. Then, a structured gradient alignment loss is introduced to encourage edge consistency between the translated and input images. In addition, pixel-level annotation is carried out on a subset of FLIR and KAIST datasets to evaluate the semantic preservation performance of multiple translation methods. Furthermore, a new metric is devised to evaluate the geometric consistency in the translation process. Extensive experiments demonstrate the superiority of the proposed PearlGAN over other image translation methods for the NTIR2DC task. The source code and labeled segmentation masks will be available athttps://github.com/FuyaLuo/PearlGAN/.
Fuya Luo, Yunhan Li, Gang Wang 0031, Yongjie Li 0001
IEEE Trans. Intell. Transp. Syst.6
2022 ID-YOLO: Real-Time Salient Object Detection Based on the Driver's Fixation Region
abstract
Object detection is an important task for self-driving vehicles or advanced driver assistant systems (ADASs). Additionally, visual selective attention is a crucial neural mechanism in a driver’s vision system that can rapidly filter out unnecessary visual information in a driving scene. Some existing models detect all objects in driving scenes from the aspect of computer vision. However, in a rapidly changing driving environment, detecting salient or critical objects appearing in drivers’ interested or safety-relevant areas is more useful for ADASs. In this paper, we managed to detect salient and critical objects based on drivers’ fixation regions. To this end, we built an augmented eye tracking object detection (ETOD) dataset based on driving videos with multiple drivers’ eye movement collected by Denget al.Furthermore, we proposed a real-time salient object detection network named increase-decrease YOLO (ID-YOLO) to discriminate the critical objects within the drivers’ fixation region. The proposed ID-YOLO shows excellent detection of major objects that drivers are concerned about during driving. Compared with the present object detection models in autonomous and assisted driving systems, our object detection framework simulates the selective attention mechanism of drivers. Thus, it does not detect all of the objects appearing in the driving scenes but only detects the most relevant ones for driving safety. It can largely reduce the interference of irrelevant scene information, showing potential practical applications in intelligent or assisted driving systems.
Long Qin 0002, Yi Shi 0012, Yahui He, Junrui Zhang 0010, Yongjie Li 0001, Tao Deng 0002
IEEE Trans. Intell. Transp. Syst.6
2022 Retinal Vessel Segmentation With Skeletal Prior and Contrastive Loss
abstract
The morphology of retinal vessels is closely associated with many kinds of ophthalmic diseases. Although huge progress in retinal vessel segmentation has been achieved with the advancement of deep learning, some challenging issues remain. For example, vessels can be disturbed or covered by other components presented in the retina (such as optic disc or lesions). Moreover, some thin vessels are also easily missed by current methods. In addition, existing fundus image datasets are generally tiny, due to the difficulty of vessel labeling. In this work, a new network called SkelCon is proposed to deal with these problems by introducing skeletal prior and contrastive loss. A skeleton fitting module is developed to preserve the morphology of the vessels and improve the completeness and continuity of thin vessels. A contrastive loss is employed to enhance the discrimination between vessels and background. In addition, a new data augmentation method is proposed to enrich the training samples and improve the robustness of the proposed model. Extensive validations were performed on several popular datasets (DRIVE, STARE, CHASE, and HRF), recently developed datasets (UoA-DR, IOSTAR, and RC-SLO), and some challenging clinical images (from RFMiD and JSIEC39 datasets). In addition, some specially designed metrics for vessel segmentation, including connectivity, overlapping area, consistency of vessel length, revised sensitivity, specificity, and accuracy were used for quantitative evaluation. The experimental results show that, the proposed model achieves state-of-the-art performance and significantly outperforms compared methods when extracting thin vessels in the regions of lesions or optic disc. Source code is available at https://www.github.com/tyb311/SkelCon.
Yubo Tan, Kaifu Yang, Shixuan Zhao 0001, Yongjie Li 0001
IEEE Trans. Medical Imaging4
2022 A Local and Global Feature Disentangled Network: Toward Classification of Benign-Malignant Thyroid Nodules From Ultrasound Image
abstract
Thyroid nodules are one of the most common nodular lesions. The incidence of thyroid cancer has increased rapidly in the past three decades and is one of the cancers with the highest incidence. As a non-invasive imaging modality, ultrasonography can identify benign and malignant thyroid nodules, and it can be used for large-scale screening. In this study, inspired by the domain knowledge of sonographers when diagnosing ultrasound images, a local and global feature disentangled network (LoGo-Net) is proposed to classify benign and malignant thyroid nodules. This model imitates the dual-pathway structure of human vision and establishes a new feature extraction method to improve the recognition performance of nodules. We use the tissue-anatomy disentangled (TAD) block to connect the dual pathways, which decouples the cues of local and global features based on the self-attention mechanism. To verify the effectiveness of the model, we constructed a large-scale dataset and conducted extensive experiments. The results show that our method achieves an accuracy of 89.33%, which has the potential to be used in the clinical practice of doctors, including early cancer screening procedures in remote or resource-poor areas.
Shixuan Zhao 0001, Yang Chen 0060, Kaifu Yang, Bu-Yun Ma, Yongjie Li 0001
IEEE Trans. Medical Imaging6
2021 Nighttime Thermal Infrared Image Colorization with Dynamic Label Mining
Fuya Luo, Yijun Cao, Yongjie Li 0001
ICIG (3)3
2021 LSSNet: A Two-stream Convolutional Neural Network for Spotting Macro- and Micro-expression in Long Videos
abstract
Macro- and micro-expression spotting is a very challenging task to locate their occurrence intervals in long face videos. In this paper, we propose an efficient two-stream network named location suppression based spotting network (LSSNet), which includes three parts. First, the optical flow is extracted using the traditional TV-L1 algorithm which captures subtle facial movements while adding temporal information to alleviate the problem of insufficient samples. Then, fixed length features are extracted from the sampled optical flow and raw images by an I3D model, which is used to set sliding windows. Finally, location suppression modules (LSMs) are added to the pyramidal convolutional neural network (CNN) to reduce the proposals with too long and too short intervals. In addition, we use two different methods, named top_k and top_threshold, for validation. We adopt leave-one-subject-out (LOSO) to train our model on CAS(ME)2 and SAMM-LV. Experimental results show that our LSSNet achieves the state-of-the-art result with top_threshold, especially on the CAS(ME)2 dataset. The code is available at https://github.com/williamlee91/mer_spot.
Wang-Wang Yu, Yongjie Li 0001
ACM Multimedia3
2021 A deep-learning-based framework for severity assessment of COVID-19 with CT images
Shixuan Zhao 0001, Yang Chen 0060, Fuya Luo, Zhiqing Kang, Shengping Cai, Wei Zhao 0040, Jun Liu 0075, Yongjie Li 0001
Expert Syst. Appl.10
2021 Saliency Detection Inspired by Topological Perception Theory
Kaifu Yang, Fuya Luo, Yongjie Li 0001
Int. J. Comput. Vis.4
2021 Enhancing in-tree-based clustering via distance ensemble and kernelization
Teng Qiu, Yongjie Li 0001
Pattern Recognit.2
2021 SCOAT-Net: A novel network for segmenting COVID-19 lung opacification from CT images
Shixuan Zhao 0001, Yang Chen 0060, Wei Zhao 0040, Xingzhi Xie, Jun Liu 0075, Yongjie Li 0001
Pattern Recognit.8
2021 Learning Crisp Boundaries Using Deep Refinement Network and Adaptive Weighting Loss
abstract
Significant progress has been made in boundary detection with the help of convolutional neural networks. Recent boundary detection models not only focus on real object boundary detection but also “crisp” boundaries (precisely localized along the object's contour). There are two methods to evaluate crisp boundary performance. One uses more strict tolerance to measure the distance between the ground truth and the detected contour. The other focuses on evaluating the contour map without any postprocessing. In this study, we analyze both methods and conclude that both methods are two aspects of crisp contour evaluation. Accordingly, we propose a novel network named deep refinement network (DRNet) that stacks multiple refinement modules to achieve richer feature representation and a novel loss function, which combines cross-entropy and dice loss through effective adaptive fusion. Experimental results demonstrated that we achieve state-of-the-art performance for several available datasets.
Yijun Cao, Chuan Lin 0003, Yongjie Li 0001
IEEE Trans. Multim.3
2020 A Biological Vision Inspired Framework for Image Enhancement in Poor Visibility Conditions
abstract
Image enhancement is an important pre-processing step for many computer vision applications especially regarding the scenes in poor visibility conditions. In this work, we develop a unified two-pathway model inspired by the biological vision, especially the early visual mechanisms, which contributes to image enhancement tasks including low dynamic range (LDR) image enhancement and high dynamic range (HDR) image tone mapping. Firstly, the input image is separated and sent into two visual pathways: structure-pathway and detail-pathway, corresponding to the M-and P-pathway in the early visual system, which code the low-and high-frequency visual information, respectively. In the structure-pathway, an extended biological normalization model is used to integrate the global and local luminance adaptation, which can handle the visual scenes with varying illuminations. On the other hand, the detail enhancement and local noise suppression are achieved in the detail-pathway based on local energy weighting. Finally, the outputs of structure-and detail-pathway are integrated to achieve the low-light image enhancement. In addition, the proposed model can also be used for tone mapping of HDR images with some fine-tuning steps. Extensive experiments on three datasets (two LDR image datasets and one HDR scene dataset) show that the proposed model can handle the visual enhancement tasks mentioned above efficiently and outperform the related state-of-the-art methods.
Kaifu Yang, Yongjie Li 0001
IEEE Trans. Image Process.3
2019 An Adaptive Method for Image Dynamic Range Adjustment
abstract
In this paper, we relate the operation of image dynamic range adjustment to the following two tasks: 1) for a high dynamic range (HDR) image, its dynamic range will be mapped to the available dynamic range of display devices and 2) for a low dynamic range (LDR) image, its distribution of intensity will be extended to adequately utilize the full dynamic range of display devices. The common goal of both tasks is to preserve or even enhance the details and improve the visibility of scenes when being matched to the available dynamic range of a display device. In this paper, we propose an efficient method for image dynamic range adjustment with three adaptive steps. First, according to the histogram of the luminance map separated from the given RGB image, two suitable Gamma functions are adaptively selected to separately adjust the luminance of the dark and bright components. Second, an adaptive fusion strategy is proposed to combine the two adjusted luminance maps in order to balance the enhancement of the details in different regions. Third, an adaptive luminance-dependent color restoration method is designed to combine the fused luminance map with the original color components to obtain more consistent color saturation between the images before and after dynamic range adjustment. Extensive experiments show that the proposed method can efficiently compress the dynamic range of HDR scenes with good contrast, clear details, and high structural fidelity of the original image appearance. In addition, the proposed method can also obtain promising performance when being used to enhance LDR nighttime images and greatly facilitates the object (car) detection in nighttime traffic scenes.
Kaifu Yang, Hulin Kuang, Chao-Yi Li, Yongjie Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2019 Combining Bottom-Up and Top-Down Visual Mechanisms for Color Constancy Under Varying Illumination
abstract
Multi-illuminant-based color constancy (MCC) is quite a challenging task. In this paper, we proposed a novel model motivated by the bottom-up and top-down mechanisms of human visual system (HVS) to estimate the spatially varying illumination in a scene. The motivation for bottom-up based estimation is from our finding that the bright and dark parts in a scene play different roles in encoding illuminants. However, handling the color shift of large colorful objects is difficult using pure bottom-up processing. Thus, we further introduce a top-down constraint inspired by the findings in visual psychophysics, in which high-level information (e.g., the prior of light source colors) plays a key role in visual color constancy. In order to implement the top-down hypothesis, we simply learn a color mapping between the illuminant distribution estimated by bottom-up processing and the ground truth maps provided by the dataset. We evaluated our model on four datasets and the results show that our method obtains very competitive performance compared with the state-of-the-art MCC algorithms. Moreover, the robustness of our model is more tangible considering that our results were obtained using the same parameters for all the datasets or the parameters of our model were learned from the inputs, that is, mimicking how HVS operates. We also show the color correction results on some real-world images taken from the web.
Shao-Bing Gao, Yanze Ren, Ming Zhang 0016, Yongjie Li 0001
IEEE Trans. Image Process.4
2019 Underwater Image Enhancement Using Adaptive Retinal Mechanisms
abstract
We propose an underwater image enhancement model inspired by the morphology and function of the teleost fish retina. We aim to solve the problems of underwater image degradation raised by the blurring and nonuniform color biasing. In particular, the feedback from color-sensitive horizontal cells to cones and a red channel compensation are used to correct the nonuniform color bias. The center-surround opponent mechanism of the bipolar cells and the feedback from amacrine cells to interplexiform cells then to horizontal cells serve to enhance the edges and contrasts of the output image. The ganglion cells with color-opponent mechanism are used for color enhancement and color correction. Finally, we adopt a luminance-based fusion strategy to reconstruct the enhanced image from the outputs of ON and OFF pathways of fish retina. Our model utilizes the global statistics (i.e., image contrast) to automatically guide the design of each low-level filter, which realizes the self-adaption of the main parameters. Extensive qualitative and quantitative evaluations on various underwater scenes validate the competitive performance of our technique. Our model also significantly improves the accuracy of transmission map estimation and local feature point matching using the underwater image. Our method is a single image approach that does not require the specialized prior about the underwater condition or scene structure.
Shao-Bing Gao, Ming Zhang 0016, Qian Zhao 0019, Yongjie Li 0001
IEEE Trans. Image Process.5
2018 D-NND: A Hierarchical Density Clustering Method via Nearest Neighbor Descent
abstract
Most density-based clustering methods largely rely on how well the underlying density is estimated. However, like clustering, density estimation is also a challenging unsupervised learning problem, especially the determination of the kernel bandwidth. In this paper, we propose a density-based multilayer hierarchical clustering method, called the Deep Nearest Neighbor Descent (D-NND), which can largely alleviate the impact of the density estimation. Unlike previous density-based methods, D-NND learns the underlying density distribution layer by layer and at the same time makes the dataset sparsely and effectively organized into a directed Tree. The experiments on three real-world datasets and several challenging synthetic datasets demonstrate that the proposed method has strong ability to discover the underlying cluster structures and is not very sensitive to the density estimation method, the parameters and the clusters of multiple scales.
Teng Qiu, Chaoyi Li, Yongjie Li 0001
ICPR3
2018 Learning to Boost Bottom-Up Fixation Prediction in Driving Environments via Random Forest
abstract
Saliency detection, an important step in many computer vision applications, can, for example, predict where drivers look in a vehicular traffic environment. While many bottom-up and top-down saliency detection models have been proposed for fixation prediction in outdoor scenes, no specific attempt has been made for traffic images. Here, we propose a learning saliency detection model based on a random forest (RF) to predict drivers' fixation positions in a driving environment. First, we extract low-level (color, intensity, orientation, etc.) and high-level (e.g., the vanishing point and center bias) features and then predict the fixation points via RF-based learning. Finally, we evaluate the performance of our saliency prediction model qualitatively and quantitatively. We use quantitative evaluation metrics that include the revised receiver operating characteristic (ROC), the area under the ROC curve value, and the normalized scan-path saliency score. The experimental results on real traffic images indicate that our model can more accurately predict a driver's fixation area, while driving than the state-of-the-art bottom-up saliency models.
Tao Deng 0002, Yongjie Li 0001
IEEE Trans. Intell. Transp. Syst.3
2018 Bayes Saliency-Based Object Proposal Generator for Nighttime Traffic Images
abstract
Object proposal is one of the most key pre-processing steps for nighttime vehicle detection systems in intelligent transportation systems. However, most current object proposal methods are developed on daytime data sets, and these methods demonstrate unsatisfactory results when they are used on nighttime images. Therefore, this paper presents a novel Bayes saliency-based object proposal generator for nighttime RGB traffic images to generate a modest and accurate set of proposals, which are more likely to be vehicles for preceding vehicle detection. First, we propose a new Bayes saliency detection approach in which prior estimation, feature extraction, weight estimation, and Bayes rule are used to compute saliency maps. Then, we propose a simple but effective object proposal generator based on the Bayes saliency map. Multi-scale sliding window, proposal rejecting, scoring, and non-maximum suppression are combined to generate a modest and effective set of proposals. Experimental results demonstrate that our proposed approach generates a modest set of proposals and outperforms some state-of-the-art methods on nighttime images in terms of various evaluation metrics. Furthermore, our proposed object proposal approach can improve the detection performance and the speed of several state-of-the-art vehicle detection approaches.
Hulin Kuang, Kaifu Yang, Long Chen 0005, Yongjie Li 0001, Leanne Lai Chan, Hong Yan 0001
IEEE Trans. Intell. Transp. Syst.4
2017 Nighttime Vehicle Detection Based on Bio-Inspired Image Enhancement and Weighted Score-Level Feature Fusion
abstract
This paper presents an effective nighttime vehicle detection system that combines a novel bioinspired image enhancement approach with a weighted feature fusion technique. Inspired by the retinal mechanism in natural visual processing, we develop a nighttime image enhancement method by modeling the adaptive feedback from horizontal cells and the center-surround antagonistic receptive fields of bipolar cells. Furthermore, we extract features based on the convolutional neural network, histogram of oriented gradient, and local binary pattern to train the classifiers with support vector machine. These features are fused by combining the score vectors of each feature with the learnt weights. During detection, we generate accurate regions of interest by combining vehicle taillight detection with object proposals. Experimental results demonstrate that the proposed bioinspired image enhancement method contributes well to vehicle detection. Our vehicle detection method demonstrates a 95.95% detection rate at 0.0575 false positives per image and outperforms some state-of-the-art techniques. Our proposed method can deal with various scenes including vehicles of different types and sizes and those with occlusions and in blurred zones. It can also detect vehicles at various locations and multiple vehicles.
Hulin Kuang, Yongjie Li 0001, Leanne Lai Chan, Hong Yan 0001
IEEE Trans. Intell. Transp. Syst.3
2016 A Unified Framework for Salient Structure Detection by Contour-Guided Visual Search
abstract
We define the task of salient structure (SS) detection to unify the saliency-related tasks, such as fixation prediction, salient object detection, and detection of other structures of interest in cluttered environments. To solve such SS detection tasks, a unified framework inspired by the two-pathway-based search strategy of biological vision is proposed in this paper. First, a contour-based spatial prior (CBSP) is extracted based on the layout of edges in the given scene along a fast non-selective pathway, which provides a rough, task-irrelevant, and robust estimation of the locations where the potential SSs are present. Second, another flow of local feature extraction is executed in parallel along the selective pathway. Finally, Bayesian inference is used to auto-weight and integrate the local cues guided by CBSP and to predict the exact locations of SSs. This model is invariant to the size and features of objects. The experimental results on six large datasets (three fixation prediction datasets and three salient object datasets) demonstrate that our system achieves competitive performance for SS detection (i.e., both the tasks of fixation prediction and salient object detection) compared with the state-of-the-art methods. In addition, our system also performs well for salient object construction from saliency maps and can be easily extended for salient edge detection.
Kaifu Yang, Chao-Yi Li, Yongjie Li 0001
IEEE Trans. Image Process.4
2016 A Retinal Mechanism Inspired Color Constancy Model
abstract
In this paper, we propose a novel model for the computational color constancy, inspired by the amazing ability of the human vision system (HVS) to perceive the color of objects largely constant as the light source color changes. The proposed model imitates the color processing mechanisms in the specific level of the retina, the first stage of the HVS, from the adaptation emerging in the layers of cone photoreceptors and horizontal cells (HCs) to the color-opponent mechanism and disinhibition effect of the non-classical receptive field in the layer of retinal ganglion cells (RGCs). In particular, HC modulation provides a global color correction with cone-specific lateral gain control, and the following RGCs refine the processing with iterative adaptation until all the three opponent channels reach their stable states (i.e., obtain stable outputs). Instead of explicitly estimating the scene illuminant(s), such as most existing algorithms, our model directly removes the effect of scene illuminant. Evaluations on four commonly used color constancy data sets show that the proposed model produces competitive results in comparison with the state-of-the-art methods for the scenes under either single or multiple illuminants. The results indicate that single opponency, especially the disinhibitory effect emerging in the receptive field's subunit-structured surround of RGCs, plays an important role in removing scene illuminant(s) by inherently distinguishing the spatial structures of surfaces from extensive illuminant(s).
Shao-Bing Gao, Ruo-Xuan Li, Chao-Yi Li, Yongjie Li 0001
IEEE Trans. Image Process.6
2016 Where Does the Driver Look? Top-Down-Based Saliency Detection in a Traffic Driving Environment
abstract
A traffic driving environment is a complex and dynamically changing scene. When driving, drivers always allocate their attention to the most important and salient areas or targets. Traffic saliency detection, which computes the salient and prior areas or targets in a specific driving environment, is an indispensable part of intelligent transportation systems and could be useful in supporting autonomous driving, traffic sign detection, driving training, car collision warning, and other tasks. Recently, advances in visual attention models have provided substantial progress in describing eye movements over simple stimuli and tasks such as free viewing or visual search. However, to date, there exists no computational framework that can accurately mimic a driver's gaze behavior and saliency detection in a complex traffic driving environment. In this paper, we analyzed the eye-tracking data of 40 subjects consisted of nondrivers and experienced drivers when viewing 100 traffic images. We found that a driver's attention was mostly concentrated on the end of the road in front of the vehicle. We proposed that the vanishing point of the road can be regarded as valuable top-down guidance in a traffic saliency detection model. Subsequently, we build a framework of a classic bottom-up and top-down combined traffic saliency detection model. The results show that our proposed vanishing-point-based top-down model can effectively simulate a driver's attention areas in a driving environment.
Tao Deng 0002, Kaifu Yang, Yongjie Li 0001
IEEE Trans. Intell. Transp. Syst.3
2015 Efficient illuminant estimation for color constancy using grey pixels
abstract
Illuminant estimation is a key step for computational color constancy. Instead of using the grey world or grey edge assumptions, we propose in this paper a novel method for illuminant estimation by using the information of grey pixels detected in a given color-biased image. The underlying hypothesis is that most of the natural images include some detectable pixels that are at least approximately grey, which can be reliably utilized for illuminant estimation. We first validate our assumption through comprehensive statistical evaluation on diverse collection of datasets and then put forward a novel grey pixel detection method based on the illuminant-invariant measure (IIM) in three logarithmic color channels. Then the light source color of a scene can be easily estimated from the detected grey pixels. Experimental results on four benchmark datasets (three recorded under single illuminant and one under multiple illuminants) show that the proposed method outperforms most of the state-of-the-art color constancy approaches with the inherent merit of low computational cost.
Kaifu Yang, Shao-Bing Gao, Yongjie Li 0001
CVPR3
2015 Color Constancy Using Double-Opponency
abstract
The double-opponent (DO) color-sensitive cells in the primary visual cortex (V1) of the human visual system (HVS) have long been recognized as the physiological basis of color constancy. In this work we propose a new color constancy model by imitating the functional properties of the HVS from the single-opponent (SO) cells in the retina to the DO cells in V1 and the possible neurons in the higher visual cortexes. The idea behind the proposed double-opponency based color constancy (DOCC) model originates from the substantial observation that the color distribution of the responses of DO cells to the color-biased images coincides well with the vector denoting the light source color. Then the illuminant color is easily estimated by pooling the responses of DO cells in separate channels in LMS space with the pooling mechanism of sum or max. Extensive evaluations on three commonly used datasets, including the test with the dataset dependent optimal parameters, as well as the intra- and inter-dataset cross validation, show that our physiologically inspired DOCC model can produce quite competitive results in comparison to the state-of-the-art approaches, but with a relative simple implementation and without requiring fine-tuning of the method for each different dataset.
Shao-Bing Gao, Kaifu Yang, Chao-Yi Li, Yongjie Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2015 Boundary Detection Using Double-Opponency and Spatial Sparseness Constraint
abstract
Brightness and color are two basic visual features integrated by the human visual system (HVS) to gain a better understanding of color natural scenes. Aiming to combine these two cues to maximize the reliability of boundary detection in natural scenes, we propose a new framework based on the color-opponent mechanisms of a certain type of color-sensitive double-opponent (DO) cells in the primary visual cortex (V1) of HVS. This type of DO cells has oriented receptive field with both chromatically and spatially opponent structure. The proposed framework is a feedforward hierarchical model, which has direct counterpart to the color-opponent mechanisms involved in from the retina to V1. In addition, we employ the spatial sparseness constraint (SSC) of neural responses to further suppress the unwanted edges of texture elements. Experimental results show that the DO cells we modeled can flexibly capture both the structured chromatic and achromatic boundaries of salient objects in complex scenes when the cone inputs to DO cells are unbalanced. Meanwhile, the SSC operator further improves the performance by suppressing redundant texture edges. With competitive contour detection accuracy, the proposed model has the additional advantage of quite simple implementation with low computational cost.
Kaifu Yang, Shao-Bing Gao, Ce-Feng Guo, Chao-Yi Li, Yongjie Li 0001
IEEE Trans. Image Process.5
2014 Efficient Color Constancy with Local Surface Reflectance Statistics
Shaobing Gao, Wangwang Han, Kaifu Yang, Chaoyi Li, Yongjie Li 0001
ECCV (2)5
2014 Multifeature-Based Surround Inhibition Improves Contour Detection in Natural Images
abstract
To effectively perform visual tasks like detecting contours, the visual system normally needs to integrate multiple visual features. Sufficient physiological studies have revealed that for a large number of neurons in the primary visual cortex (V1) of monkeys and cats, neuronal responses elicited by the stimuli placed within the classical receptive field (CRF) are substantially modulated, normally inhibited, when difference exists between the CRF and its surround, namely, non-CRF, for various local features. The exquisite sensitivity of V1 neurons to the center-surround stimulus configuration is thought to serve important perceptual functions, including contour detection. In this paper, we propose a biologically motivated model to improve the performance of perceptually salient contour detection. The main contribution is the multifeature-based center-surround framework, in which the surround inhibition weights of individual features, including orientation, luminance, and luminance contrast, are combined according to a scale-guided strategy, and the combined weights are then used to modulate the final surround inhibition of the neurons. The performance was compared with that of single-cue-based models and other existing methods (especially other biologically motivated ones). The results show that combining multiple cues can substantially improve the performance of contour detection compared with the models using single cue. In general, luminance and luminance contrast contribute much more than orientation to the specific task of contour extraction, at least in gray-scale natural images.
Kaifu Yang, Chao-Yi Li, Yongjie Li 0001
IEEE Trans. Image Process.3
2013 Efficient Color Boundary Detection with Color-Opponent Mechanisms
abstract
Color information plays an important role in better understanding of natural scenes by at least facilitating discriminating boundaries of objects or areas. In this study, we propose a new framework for boundary detection in complex natural scenes based on the color-opponent mechanisms of the visual system. The red-green and blue-yellow color opponent channels in the human visual system are regarded as the building blocks for various color perception tasks such as boundary detection. The proposed framework is a feed forward hierarchical model, which has direct counterpart to the color-opponent mechanisms involved in from the retina to the primary visual cortex (V1). Results show that our simple framework has excellent ability to flexibly capture both the structured chromatic and achromatic boundaries in complex scenes.
Kaifu Yang, Shaobing Gao, Chaoyi Li, Yongjie Li 0001
CVPR4
2013 A Color Constancy Model with Double-Opponency Mechanisms
abstract
The double-opponent color-sensitive cells in the primary visual cortex (V1) of the human visual system (HVS) have long been recognized as the physiological basis of color constancy. We introduce a new color constancy model by imitating the functional properties of the HVS from the retina to the double-opponent cells in V1. The idea behind the model originates from the observation that the color distribution of the responses of double-opponent cells to the input color-biased images coincides well with the light source direction. Then the true illuminant color of a scene is easily estimated by searching for the maxima of the separate RGB channels of the responses of double-opponent cells in the RGB space. Our systematical experimental evaluations on two commonly used image datasets show that the proposed model can produce competitive results in comparison to the complex state-of-the-art approaches, but with a simple implementation and without the need for training.
Shaobing Gao, Kaifu Yang, Chaoyi Li, Yongjie Li 0001
ICCV4
2011 Contour detection based on a non-classical receptive field model with butterfly-shaped inhibition subregions
Chi Zeng, Yongjie Li 0001, Kaifu Yang, Chaoyi Li
Neurocomputing2
2005 Ant colony system for the beam angle optimization problem in radiotherapy planning: a preliminary study
abstract
Intensity-modulated radiotherapy (IMRT) is being increasingly used for treatment of malignant cancer. Beam angle optimization (BAO) is an important problem in IMRT. In this paper, an emerging population-based meta-heuristic algorithm named ant colony optimization (ACO) is introduced to solve the BAO problem. In the proposed algorithm, a multi-layered graph is designed to map the BAO problem to ACO, and a heuristic function based on the beam's-eye-view dosimetrics (BEVD) score is introduced. In order to verify the feasibility of the presented algorithm, a clinical prostate tumor case is employed, and the preliminary results demonstrate that ACO appears more effcient than genetic algorithm (GA) and can find the optimal beam angles within a clinically acceptable computation time.
Yongjie Li 0001, Dezhong Yao 0001, Wufan Chen, Jiancheng Zheng, Jonathan Yao
Congress on Evolutionary Computation1
2005 A feasibility study of EEG dipole source localization using particle swarm optimization
abstract
Interpretation of the clinical electroencephalographs (EEGs) almost always involves speculation as to the possible locations of the sources inside the brain that are responsible for the observed activity on the scalp. Dipoles are widely used to approximate the sources of electrical activity inside the brain. In this paper, we introduce a novel particle swarm optimization (PSO) algorithm to the EEG dipole source localization problem. A three-concentric-shell model is chosen as our head model, and the dipole number is restricted to 2. The 2 dipoles, each of which has 3 position elements, are combined and represented as a 6-element particle. Initialized by randomly setting the positions and velocities, the particle swarm evolves iteratively. Reported here are simulated cases to demonstrate the feasibility of the proposed PSO-based algorithm. Four groups of dipoles with different physiological meanings are chosen as the tested source models. Simulated cases with 10% noise level are also tested. The results show that PSO is feasible and efficient for the source localization in EEG. Furthermore, compared with the generally accepted genetic algorithm (GA), the PSO algorithm appears to be more accurate and needs less computation time
Lijun Qiu, Yongjie Li 0001, Dezhong Yao 0001
Congress on Evolutionary Computation2