Bo Li 0090

dblp:50/3402-90 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Wideband and Low Insertion Loss Microwave Bandpass Filter Inverse Design Method Based on Convolutional Neural Network Fitting Model and Genetic Algorithm
abstract
To overcome the limitations of traditional bandpass filter design methods, which heavily rely on designer expertise and predefined topology, this paper proposes an inverse design methodology for microwave bandpass filter based on deep learning and genetic algorithms (GA). The filter structures are encoded as two-dimensional pixelated matrices, and a database is constructed through electromagnetic (EM) simulations. To reduce the computational cost of EM simulation, the dataset is further expanded using symmetry, mirroring, and rotation techniques. A convolutional neural network (CNN) is trained to model the mapping between S-parameters and arbitrary pixelated structural matrices, thereby broadening the design space. Subsequently, an improved GA is employed to synthesize filter structure that meets specific frequency response requirement. A filter operating within 0.5-2.08 GHz is successfully synthesized, fabricated, and measured. The experimental results demonstrate that the proposed filter achieves wide bandwidth, low insertion losses and good spurious rejection (20 dB @5.03f0), benefiting from the removal of constraints imposed by predefined topology.
Xing Quan, Shancheng Luan, Yuqin Huang, Bo Li 0090, Jinsong Zhan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 T-Person-GAN: Text-to-Person image generation with identity-consistency and manifold mix-up
Deyin Liu, Lin Wu 0001, Bo Li 0090, Ye Zhao 0001, ZongYuan Ge
Expert Syst. Appl.3
2024 Improving Depth Completion via Depth Feature Upsampling
abstract
The encoder-decoder network (ED-Net) is a commonly employed choice for existing depth completion methods, but its working mechanism is ambiguous. In this paper, we vi-sualize the internal feature maps to analyze how the net-work densifies the input sparse depth. We find that the en-coder feature of ED-Net focus on the areas with input depth points around. To obtain a dense feature and thus esti-mate complete depth, the decoder feature tends to comple-ment and enhance the encoder feature by skip-connection to make the fused encoder-decoder feature dense, resulting in the decoder feature also exhibits sparse. However, ED-Net obtains the sparse decoder feature from the dense fused feature at the previous stage, where the “dense-i-sparse‘’ process destroys the completeness of features and loses in-formation. To address this issue, we present a depth feature upsampling network (DFU) that explicitly utilizes these dense features to guide the upsampling of a low-resolution (LR) depth feature to a high-resolution (HR) one. The completeness of features is maintained throughout the up-sampling process, thus avoiding information loss. Fur-thermore, we propose a confidence-aware guidance module (CGM), which is confidence-aware and performs guidance with adaptive receptive fields (GARF), to fully exploit the potential of these dense features as guidance. Experimental results show that our DFU, a plug-and-play module, can significantly improve the performance of existing ED-Net based methods with limited computational overheads, and new SOTA results are achieved. Besides, the generalization capability on sparser depth is also enhanced. Project page: https://npucvr.github.iolDFU.
Ge Zhang 0006, Shaoqian Wang, Bo Li 0090, Qi Liu 0054, Le Hui, Yuchao Dai
CVPR4
2024 3D Focusing-and-Matching Network for Multi-Instance Point Cloud Registration
abstract
Multi-instance point cloud registration aims to estimate the pose of all instances of a model point cloud in the whole scene. Existing methods all adopt the strategy of first obtaining the global correspondence and then clustering to obtain the pose of each instance. However, due to the cluttered and occluded objects in the scene, it is difficult to obtain an accurate correspondence between the model point cloud and all instances in the scene. To this end, we propose a simple yet powerful 3D focusing-and-matching network for multi-instance point cloud registration by learning the multiple pair-wise point cloud registration. Specifically, we first present a 3D multi-object focusing module to locate the center of each object and generate object proposals. By using self-attention and cross-attention to associate the model point cloud with structurally similar objects, we can locate potential matching instances by regressing object centers. Then, we propose a 3D dual-masking instance matching module to estimate the pose between the model point cloud and each object proposal. It performs instance mask and overlap mask masks to accurately predict the pair-wise correspondence. Extensive experiments on two public benchmarks, Scan2CAD and ROBI, show that our method achieves a new state-of-the-art performance on the multi-instance point cloud registration task.
Le Hui, Qi Liu 0054, Bo Li 0090, Yuchao Dai
NeurIPS4
2024 Jacobian norm with Selective Input Gradient Regularization for interpretable adversarial defense
abstract
Deep neural networks (DNNs) can be easily deceived by imperceptible alterations known as adversarial examples. These examples can lead to misclassification , posing a significant threat to the reliability of deep learning systems in real-world applications. Adversarial training (AT) is a popular technique used to enhance robustness by training models on a combination of corrupted and clean data. However, existing AT-based methods often struggle to handle transferred adversarial examples that can fool multiple defense models, thereby falling short of meeting the generalization requirements for real-world scenarios. Furthermore, AT typically fails to provide interpretable predictions, which are crucial for domain experts seeking to understand the behavior of DNNs. To overcome these challenges, we present a novel approach called Jacobian norm and Selective Input Gradient Regularization (J-SIGR). Our method leverages Jacobian normalization to improve robustness and introduces regularization of perturbation-based saliency maps, enabling interpretable predictions. By adopting J-SIGR, we achieve enhanced defense capabilities and promote high interpretability of DNNs. We evaluate the effectiveness of J-SIGR across various architectures by subjecting it to powerful adversarial attacks. Our experimental evaluations provide compelling evidence of the efficacy of J-SIGR against transferred adversarial attacks, while preserving interpretability. The project code can be found at https://github.com/Lywu-github/jJ-SIGR.git .
Deyin Liu, Lin Wu 0001, Bo Li 0090, Farid Boussaïd, Mohammed Bennamoun, Xianghua Xie, Chengwu Liang
Pattern Recognit.3
2024 Efficient Multi-View Stereo by Dynamic Cost Volume and Cross-Scale Propagation
abstract
Currently, learning-based multi-view stereo (MVS) has been dominated by the pipeline of 3D cost volume and regularization network over thestatic cost volumefor depth regression. However, this methodology is plagued by heavy time and memory consumption, which greatly hinders the applications of these methods for real-world high-resolution images. To address these challenges, we present Effi-MVS+, an efficient multi-scaledynamic cost volumebased MVS method. Firstly, instead of constructing a static cost volume and predicting a probability distribution map for depth regression, we update the depth map by iteratively predicting depth residuals. In each iteration, we construct a lightweight dynamic cost volume by encoding local matching and regularization information. The dynamic cost volume is subsequently processed using a 2D convolution-based GRU, which owns significant advantages in computational complexity and efficiency. Secondly, we propose a cross-scale propagation mechanism to enhance the multi-scale dynamic cost volume. This mechanism facilitates the progressive aggregation of multi-scale information, thereby providing enhanced matching and regularization information. Thirdly, to further improve the efficiency, we provide a reliable initial depth map to launch the framework and guarantee fast convergence. Extensive experiments on the DTU and Tanks & Temples benchmarks demonstrate the superiority of our method, which outperforms other state-of-the-art methods by a large margin in terms ofreconstruction quality, speed, and memory usage. Code will be released at https://github.com/npucvr/Effi-MVS-plus.
Shaoqian Wang, Bo Li 0090, Yuchao Dai
IEEE Trans. Circuits Syst. Video Technol.2
2024 Target Before Shooting: Accurate Anomaly Detection and Localization Under One Millisecond via Cascade Patch Retrieval
abstract
In this work, by re-examining the "matching" nature of Anomaly Detection (AD), we propose a novel AD framework that simultaneously enjoys new records of AD accuracy and dramatically high running speed. In this framework, the anomaly detection problem is solved via a cascade patch retrieval procedure that retrieves the nearest neighbors for each test image patch in a coarse-to-fine fashion. Given a test sample, the top-K most similar training images are first selected based on a robust histogram matching process. Secondly, the nearest neighbor of each test patch is retrieved over the similar geometrical locations on those "most similar images", by using a carefully trained local metric. Finally, the anomaly score of each test image patch is calculated based on the distance to its "nearest neighbor" and the "non-background" probability. The proposed method is termed "Cascade Patch Retrieval" (CPR) in this work. Different from the previous patch-matching-based AD algorithms, CPR selects proper "targets" (reference images and patches) before "shooting" (patch-matching). On the well-acknowledged MVTec AD, BTAD and MVTec-3D AD datasets, the proposed algorithm consistently outperforms all the comparing SOTA methods by remarkable margins, measured by various AD metrics. Furthermore, CPR is extremely efficient. It runs at the speed of 113 FPS with the standard setting while its simplified version only requires less than 1 ms to process an image at the cost of a trivial accuracy drop. The code of CPR is available at https://github.com/flyinghu123/CPR.
Jianfei Hu, Bo Li 0090, Hao Chen 0041, Yongbin Zheng, Chunhua Shen
IEEE Trans. Image Process.3
2023 LRRU: Long-short Range Recurrent Updating Networks for Depth Completion
abstract
Existing deep learning-based depth completion methods generally employ massive stacked layers to predict the dense depth map from sparse input data. Although such approaches greatly advance this task, their accompanied huge computational complexity hinders their practical applications. To accomplish depth completion more efficiently, we propose a novel lightweight deep network framework, the Long-short Range Recurrent Updating (LRRU) network. Without learning complex feature representations, LRRU first roughly fills the sparse input to obtain an initial dense depth map, and then iteratively updates it through learned spatially-variant kernels. Our iterative update process is content-adaptive and highly flexible, where the kernel weights are learned by jointly considering the guidance RGB images and the depth map to be updated, and large-to-small kernel scopes are dynamically adjusted to capture long-to-short range dependencies. Our initial depth map has coarse but complete scene depth information, which helps relieve the burden of directly regressing the dense depth from sparse ones, while our proposed method can effectively refine it to an accurate depth map with less learnable parameters and inference time. Experimental results demonstrate that our proposed LRRU variants achieve state-of-the-art performance across different parameter regimes. In particular, the LRRU-Base model outperforms competing approaches on the NYUv2 dataset, and ranks 1st on the KITTI depth completion benchmark at the time of submission. Project page: https://npucvr.github.io/LRRU/.
Bo Li 0090, Ge Zhang 0006, Qi Liu 0054, Tao Gao 0001, Yuchao Dai
ICCV2
2023 Image Template Matching via Dense and Consistent Contrastive Learning
abstract
Image template matching refers to localizing a small query image as opposed to a large reference image map. The query image a.k.a template has to be screened across every equal-sized region in the reference map to perform inner-product at pixel-level and the resulting similarity indicates the template location. Due to the domain heterogeneity between template and reference images, the matching performance degrades under dramatic appearance changes. More severely, the asymmetric matching easily leads to over-fitting by suggesting excessively false positive regions. To these ends, we propose an effective template matching method based on contrastive learning to perform a dense and consistent InfoNCEloss during matching. This can increase the matching at finer details, and thus effectively regularizes network training to prevent over-fitting. Extensive experiments on the synthetic aperture radar (SAR) and optical datasets, i.e., SEN1-2 and OS datasets demonstrate that our proposed method outperforms state-of-the-art methods by a large margin.
Bo Li 0090, Lin Wu 0001, Deyin Liu, Hongyang Chen 0001, Yuanxin Ye, Xianghua Xie
ICME1
2023 Continuous Parametric Optical Flow
abstract
In this paper, we present continuous parametric optical flow, a parametric representation of dense and continuous motion over arbitrary time interval. In contrast to existing discrete-time representations (i.e., flow in between consecutive frames), this new representation transforms the frame-to-frame pixel correspondences to dense continuous flow. In particular, we present a temporal-parametric model that employs B-splines to fit point trajectories using a limited number of frames. To further improve the stability and robustness of the trajectories, we also add an encoder with a neural ordinary differential equation (NODE) to represent features associated with specific times. We also contribute a synthetic dataset and introduce two evaluation perspectives to measure the accuracy and robustness of continuous flow estimation. Benefiting from the combination of explicit parametric modeling and implicit feature optimization, our model focuses on motion continuity and outperforms the flow-based and point-tracking approaches for fitting long-term and variable sequences.
Jianqin Luo, Zhexiong Wan, Yuxin Mao, Bo Li 0090, Yuchao Dai
NeurIPS4
2023 DSP-Based Traffic Target Detection for Intelligent Transportation
abstract
Internet of Things (IoT)-based intelligent transportation is attracting more and more attention. As a key component of intelligent transportation, traffic video monitoring is very important, in which vehicle and pedestrian detection on the road is a crucial task. Although vehicle and pedestrian detection through deep learning (DL) may achieve high accuracy, it tends to require high computing resources, which hinders its use on IoT devices. As an important class of IoT devices, digital signal processor (DSP) has the characteristics of low energy consumption, small size, and strong performance, which has been widely used in intelligent transportation. In order to use DL on DSP for accurate vehicle and pedestrian detection, we first propose a series of general tactics to optimize the object detection convolutional neural network (CNN) model, including convolution layer optimization, cache optimization, compiler optimization, intrinsics optimization and direct memory access (DMA) acceleration, and then a parallel scheme to extend the model to run on multicore, and further quantize the implementation of the model. We evaluate it on UA-DETRAC and KITTI datasets. Experimental results show that our method achieves a faster speed than running the same CNN model on a mainstream desktop CPU, with only 0.06% accuracy loss.
Jianhua Zhang 0002, Rucen Wang, Ruyu Liu, Dongyan Guo, Bo Li 0090, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.5
2022 Efficient Multi-view Stereo by Iterative Dynamic Cost Volume
abstract
In this paper, we propose a novel iterative dynamic cost volume for multi-view stereo. Compared with other works, our cost volume is much lighter, thus could be processed with 2D convolution based GRU. Notably, the every-step output of the GRU could be further used to generate new cost volume. In this way, an iterative GRU-based optimizer is constructed. Furthermore, we present a cascade and hierarchical refinement architecture to utilize the multiscale information and speed up the convergence. Specifically, a lightweight 3D CNN is utilized to generate the coarsest initial depth map which is essential to launch the GRU and guarantee a fast convergence. Then the depth map is refined by multi-stage GRUs which work on the pyramid feature maps. Extensive experiments on the DTU and Tanks & Temples benchmarks demonstrate that our method could achieve state-of-the-art results in terms of accuracy, speed and memory usage. Code will be released at https://github.com/bdwsq1996/Effi-MVS.
Shaoqian Wang, Bo Li 0090, Yuchao Dai
CVPR2
2022 A Multiscale Framework With Unsupervised Learning for Remote Sensing Image Registration
abstract
Registration for multisensor or multimodal image pairs with a large degree of distortions is a fundamental task for many remote sensing applications. To achieve accurate and low-cost remote sensing image registration, we propose a multiscale framework with unsupervised learning, named MU-Net. Without costly ground truth labels, MU-Net directly learns the end-to-end mapping from the image pairs to their transformation parameters. MU-Net stacks several deep neural network (DNN) models on multiple scales to generate a coarse-to-fine registration pipeline, which prevents the backpropagation from falling into a local extremum and resists significant image distortions. We design a novel loss function paradigm based on structural similarity, which makes MU-Net suitable for various types of multimodal images. MU-Net is compared with traditional feature-based and area-based methods, as well as supervised and other unsupervised learning methods on the optical-optical, optical-infrared, optical-synthetic aperture radar (SAR), and optical-map datasets. Experimental results show that MU-Net achieves more comprehensive and accurate registration performance between these image pairs with geometric and radiometric distortions. We share the code implemented by Pytorch athttps://github.com/yeyuanxin110/MU-Net.
Yuanxin Ye, Tengfeng Tang, Bai Zhu, Chao Yang 0028, Bo Li 0090, Siyuan Hao
IEEE Trans. Geosci. Remote. Sens.5
2021 Hybrid 2-D-3-D Deep Residual Attentional Network With Structure Tensor Constraints for Spectral Super-Resolution of RGB Images
abstract
RGB image spectral super-resolution (SSR) is a challenging task due to its serious ill-posedness, which aims at recovering a hyperspectral image (HSI) from a corresponding RGB image. In this article, we propose a novel hybrid 2-D-3-D deep residual attentional network (HDRAN) with structure tensor constraints, which can take fully advantage of the spatial-spectral context information in the reconstruction progress. Previous works improve the SSR performance only through stacking more layers to catch local spatial correlation neglecting the differences and interdependences among features, especially band features; different from them, our novel method focuses on the context information utilization. First, the proposed HDRAN consists of a 2D-RAN following by a 3D-RAN, where the 2D-RAN mainly focuses on extracting abundant spatial features, whereas the 3D-RAN mainly simulates the interband correlations. Then, we introduce 2-D channel attention and 3-D band attention mechanisms into the 2D-RAN and 3D-RAN, respectively, to adaptively recalibrate channelwise and bandwise feature responses for enhancing context features. Besides, since structure tensor represents structure and spatial information, we apply structure tensor constraint to further reconstruct more accurate high-frequency details during the training process. Experimental results demonstrate that our proposed method achieves the state-of-the-art performance in terms of mean relative absolute error (MRAE) and root mean square error (RMSE) on both the “clean” and “real world” tracks in the NTIRE 2018 Spectral Reconstruction Challenge. As for competitive ranking metric MRAE, our method separately achieves a 16.06% and 2.90% relative reduction on two tracks over the first place. Furthermore, we investigate HDRAN on the other two HSI benchmarks noted as the CAVE and Harvard data sets, also demonstrating better results than state-of-the-art methods.
Jiaojiao Li 0001, Chaoxiong Wu, Rui Song 0003, Weiying Xie, Chiru Ge, Bo Li 0090, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.6
2020 Novel View Synthesis from only a 6-DoF Camera Pose by Two-stage Networks
Bo Li 0090, Yuchao Dai, Tongxin Zhang
ICPR2
2020 Hyperspectral Image Super-Resolution by Band Attention Through Adversarial Learning
abstract
Hyperspectral image (HSI) super-resolution (SR) is a challenging task due to the problems of texture blur and spectral distortion when the upscaling factor is large. To meet these two challenges, band attention through the adversarial learning method is proposed in this article. First, we put the SR process in a generative adversarial network (GAN) framework, so that the resulted high-resolution HSI can keep more texture details. Second, different from the other band-by-band SR method, the input of our method is of full bands. In order to explore the correlation of spectral bands and avoid the spectral distortion, a band attention mechanism is proposed in our generative network. A series of spatial-spectral constraints or loss functions is imposed to guide the training of our generative network so as to further alleviate spectral distortion and texture blur. The experiments on the Pavia and Cave data sets demonstrate that the proposed GAN-based SR method can yield very high-quality results, even under large upscaling factor (e.g., $8\times $ ). More importantly, it can outperform the other state-of-the-art methods by a margin which demonstrates its superiority and effectiveness.
Jiaojiao Li 0001, Ruxing Cui, Bo Li 0090, Rui Song 0003, Yunsong Li 0001, Yuchao Dai, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 MVS2: Deep Unsupervised Multi-View Stereo with Multi-View Symmetry
abstract
The success of existing deep-learning based multi-view stereo (MVS) approaches greatly depends on the availability of large-scale supervision in the form of dense depth maps. Such supervision, while not always possible, tends to hinder the generalization ability of the learned models in never-seen-before scenarios. In this paper, we propose the first unsupervised learning based MVS network, which learns the multi-view depth maps from the input multi-view images and does not need ground-truth 3D training data. Our network is symmetric in predicting depth maps for all views simultaneously, where we enforce cross-view consistency of multi-view depth maps during both training and testing stages. Thus, the learned multi-view depth maps naturally comply with the underlying 3D scene geometry. Besides, our network also learns the multi-view occlusion maps, which further improves the robustness of our network in handling real-world occlusions. Experimental results on multiple benchmarking datasets demonstrate the effectiveness of our network and the excellent generalization ability.
Yuchao Dai, Zhidong Zhu, Zhibo Rao, Bo Li 0090
3DV4
2019 Dual 1D-2D Spatial-Spectral CNN for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image (HSI) spatial super-resolution(SR) is a challenging task. Compared with a RGB images, the mapping between the low-high HSI pairs is more difficult since much more spectral bands are involved. In this paper, a novel dual 1D-2D spatial-spectral convolutional neural network (CNN) architecture is proposed for spatial SR of HSIs. Specifically, by differential treatment over redundancy in spectral and spatial domains of an HSI, the spectral and spatial context are first separately explored by 1D and 2D convolution. These two kinds of feature information are then fused using a novel hierarchical side connection, which impose the spectral information to the spatial path gradually. Experimental results over benchmark Pavia data set demonstrate that the proposed architecture clearly outperform state-of-the-art 3D CNN based works in terms of both visual quality and quantitative assessment.
Jiaojiao Li 0001, Ruxing Cui, Bo Li 0090, Yunsong Li 0001, Shaohui Mei, Qian Du 0001
IGARSS3
2018 3D skeleton based action recognition by video-domain translation-scale invariant mapping and multi-scale dilated CNN
Bo Li 0090, Mingyi He, Yuchao Dai, Xuelian Cheng
Multim. Tools Appl.1
2018 Monocular depth estimation with hierarchical fusion of dilated CNNs and soft-weighted-sum inference
Bo Li 0090, Yuchao Dai, Mingyi He
Pattern Recognit.1
2018 Label Distribution-Based Facial Attractiveness Computation by Deep Residual Learning
abstract
Two key challenges lie in the facial attractiveness computation research: the lack of discriminative face representations, and the scarcity of sufficient and complete training data. Motivated by recent promising work in face recognition using deep neural networks to learn effective features, the first challenge is expected to be addressed from a deep learning point of view. A very deep residual network is utilized to enable automatic learning of hierarchical aesthetics representation. The inspiration to deal with the second challenge comes from the natural representation of the training data, where each training face can be associated with a label (score) distribution given by human raters rather than a single label (average score). This paper, therefore, recasts facial attractiveness computation as a label distribution learning problem. Integrating these two ideas, an end-to-end attractiveness learning framework is established. We also perform feature-level fusion by incorporating the low-level geometric features to further improve the computational performance. Extensive experiments are conducted on a standard benchmark, the SCUT-FBP dataset, where our approach shows significant advantages over the other state-of-the-art work.
Yangyu Fan, Shu Liu 0002, Bo Li 0090, Ashok Samal, Jun Wan 0001, Stan Z. Li
IEEE Trans. Multim.3
2017 Multi-scale 3D deep convolutional neural network for hyperspectral image classification
abstract
Research in deep neural network (DNN) and deep learning has great progress for 1D (speech), 2D (image) and 3D (3D-object) recognition/classification problems. As HSI that with 2D spatial and 1D spectral information is quite different from 3D object image, the existing DNN cannot be directly extended to hyperspectral image (HSI) classification. A Multiscale 3D deep convolutional neural network (M3D-DCNN) is proposed for HSI classification, which could jointly learn both 2D Multi-scale spatial feature and 1D spectral feature from HSI data in an end-to-end approach, promising to achieve better results with large-scale dataset. Although without any hand-craft features or pre/post-processing like PCA, sparse coding etc, we achieve the state-of-the-art results on the standard datasets, which shows the technical validity and advancement of our method.
Mingyi He, Bo Li 0090, Huahui Chen 0002
ICIP2
2017 Integrated deep and shallow networks for salient object detection
abstract
Deep convolutional neural network (CNN) based salient object detection methods have achieved state-of-the-art performance and outperform those unsupervised methods with a wide margin. In this paper, we propose to integrate deep and unsupervised saliency for salient object detection under a unified framework. Specifically, our method takes results of unsupervised saliency (Robust Background Detection, RBD) and normalized color images as inputs, and directly learns an end-to-end mapping between inputs and the corresponding saliency maps. The color images are fed into a Fully Convolutional Neural Networks (FCNN) adapted from semantic segmentation to exploit high-level semantic cues for salient object detection. Then the results from deep FCNN and RBD are concatenated to feed into a shallow network to map the concatenated feature maps to saliency maps. Finally, to obtain a spatially consistent saliency map with sharp object boundaries, we fuse superpixel level saliency map at multi-scale. Extensive experimental results on 8 benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches with a margin.
Jing Zhang 0052, Bo Li 0090, Yuchao Dai, Fatih Porikli, Mingyi He
ICIP2
2017 Facial attractiveness computation by label distribution learning with deep CNN and geometric features
abstract
Facial attractiveness computation is a challenging task because of the lack of labeled data and discriminative features. In this paper, an end-to-end label distribution learning (LDL) framework with deep convolutional neural network (CNN) and geometric features is proposed to meet these two challenges. Different from the previous work, we recast this task as an LDL problem. Compared with the single label regression, the LDL could improve the generalization ability of our model significantly. In addition, we propose some kinds of geometric features as well as an incremental feature selection method, which could select hundred-dimensional discriminative geometric features from an exhaustive pool of raw features. More importantly, we find these selected geometric features are complementary to CNN features. Extensive experiments are carried out on the SCUT-FBP dataset, where our approach achieves superior performance in comparison to the state-of-the-arts.
Shu Liu 0002, Bo Li 0090, Yangyu Fan, Ashok Samal
ICME2
2017 A Classified Slot Re-allocation Algorithm for Synchronous Directional Ad Hoc Networks
Zhicheng Bai, Bo Li 0090, Zhongjiang Yan, Mao Yang 0001, Xiaofei Jiang, Hang Zhang 0006
QSHINE2
2015 Depth and surface normal estimation from monocular images using regression on deep features and hierarchical CRFs
abstract
Predicting the depth (or surface normal) of a scene from single monocular color images is a challenging task. This paper tackles this challenging and essentially underdetermined problem by regression on deep convolutional neural network (DCNN) features, combined with a post-processing refining step using conditional random fields (CRF). Our framework works at two levels, super-pixel level and pixel level. First, we design a DCNN model to learn the mapping from multi-scale image patches to depth or surface normal values at the super-pixel level. Second, the estimated super-pixel depth or surface normal is refined to the pixel level by exploiting various potentials on the depth or surface normal map, which includes a data term, a smoothness term among super-pixels and an auto-regression term characterizing the local structure of the estimation map. The inference problem can be efficiently solved because it admits a closed-form solution. Experiments on the Make3D and NYU Depth V2 datasets show competitive results compared with recent state-of-the-art methods.
Bo Li 0090, Chunhua Shen, Yuchao Dai, Anton van den Hengel, Mingyi He
CVPR1