Yiguang Liu

dblp:35/5030 · DBLP profile ↗
← Back
100ranked-venue papers
21as first author
34since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 16 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Systems, architecture and hardware · 5 · 3 since 2021Computer networks · 2Theory of computation · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 A decoupled framework for low-light image enhancement
Shuangli Du, Yichun Wen, Minghua Zhao, Zhenghao Shi, Yiguang Liu
Expert Syst. Appl.6
2026 Channel-specific spatial importance map generation and lost-detail recovery via asymmetric gradient projection for infrared small target detection
Ling Lei 0003, Tianhang Tang, Shaobing Gao, Yiguang Liu
Neurocomputing5
2026 Unsupervised light field images depth estimation with multi-encoder and graph convolutional networks
Jie Li 0091, Xinjia Li, Shuangli Du, Yiguang Liu
Neurocomputing7
2026 Efficient guided diffusion toward adverse weather image restoration
Tianhang Tang, Shun Lv, Ling Lei 0003, Yiguang Liu
Neurocomputing5
2026 Single-Photon Imaging in Complex Scenarios via Physics-Informed Deep Neural Networks
abstract
Single-photon imaging uses single-photon-sensitive picosecond-resolution sensors to capture 3D structure and supports diverse applications, but success remains mostly limited to simple scenes. In complex scenarios, traditional methods degrade and deep learning methods lack flexibility and generalization. Here, we propose a physics-informed deep neural network (PIDNN) framework that effectively addresses both aspects, adapting to complex and variable sensing environments by embedding imaging physics into the deep neural network for unsupervised learning. Within this framework, by tailoring the number of U-Net skip connections, we impose multi-scale spatiotemporal priors that improve photon-utilization efficiency, laying the foundation for addressing the inherent low-signal-to-background ratio (SBR) problem in subsequent complex scenarios. Additionally, we introduce volume rendering into the PIDNN framework and design a dual-branch structure, further extending its applicability to multiple-depth and fog occlusion. We validated the performance of this method in various complex environments through numerical simulations and real-world experiments. The results of photon-efficient imaging with multiple returns show robust performance under low SBR and large fields of view. The method attains lower root mean-squared error than traditional methods and exhibits stronger generalization than supervised approaches. Further multiple depths and fog interference experiments confirm that its reconstruction quality surpasses existing techniques, demonstrating its flexibility and scalability. Both simulation and experimental results validate its exceptional reconstruction performance and flexibility.
Siao Cai, Shaobing Gao, Yiguang Liu
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 A new framework for realizing fraction-order filters with robust performance
Yiguang Liu
Signal Process. Image Commun.1
2026 Multi-Scale K-Cosine Curvature Fusion With Nonlinear Measure for Robust Corner Detection
Tianhang Tang, Shun Lv, Yiguang Liu
IEEE Signal Process. Lett.4
2026 E2SRLF: End-to-End Light Field Image Super-Resolution Depth Estimation
abstract
Light field camera usually sacrifice spatial resolution for increasing angular resolution. Although it can capture rich spatial and angular information, leads to the spatial resolution reduced greatly. However, existing super-resolution methods usually focus on spatial and angular super-resolution, but ignore the end-to-end disparity map super-resolution estimation. Meanwhile, recent depth estimation methods can not extract a higher resolution disparity map from the low-resolution light field images directly. Motivated by this issue, we propose an end-to-end light field image super-resolution depth estimation network, named E2SRLF. First, a multi-dimensional channel attention mechanism is introduced to reduce the influence of occlusion and weak textures, and enhance the learning of local and global feature. Then, a spatial super-resolution fusion upsampling method is proposed for constructing the super-resolution dimension and acquiring precise high-resolution information. Additionally, we introduce a high-low resolution collaborative constraint based loss function to enforce network training efficiency. Experimental results demonstrate that E2SRLF can generate a high accuracy high-resolution depth map from the low-resolution light field images directly. Comparing to most of the state-of-the-art light field image depth estimation methods, E2SRLF directly achieved more accuracy high resolution disparity map with the lower resolution light field images input. The code of our method are available at: https://github.com/sansi-zhang/E2SRLF.
Jie Li 0091, Chuanlun Zhang, Xiaoyan Wang 0005, Xinjia Li, Yuxin Zeng, Yiguang Liu
IEEE Trans. Circuits Syst. Video Technol.7
2026 RE-LFDE: A Resource-Efficient Hardware Accelerator for Low-Bit Light Field Image Depth Estimation
abstract
Light field image depth estimation methods involve a large number of parameters and floating-point operations. This makes FPGA-based acceleration design a huge challenge, especially when pursuing high-precision acceleration design on resource-limited FPGAs. Motivated by this issue, a resource-efficient hardware accelerator based on a high-precision, low-bit lightweight light field image depth estimation network scheme is proposed, named RE-LFDE. First, we present a parameter-sharing, low-bit and lightweight network. It is able to improve accuracy, simplify the network structure and reduce network parameters. Secondly, we design a time-division multiplexing hardware-software co-design dataflow structure and build a resource-efficient acceleration engine, which can be deployed on the resource-limited FPGAs efficiently. Experimental results show that the average MSE on the 4D LF benchmark of RE-LFDE and its full-precision networks can be reduced to 3.265 and 3.363, respectively, while the weight parameters can be as low as 0.30MB and 0.04MB. Furthermore, on the ZCU104 platform, the consumption of BRAM and LUTRAM can be reduced to 22.44% and 13.64%, respectively. The code and model of the proposed method are available at https://github.com/sansi-zhang/RE-LFDE .
Jie Li 0091, Chuanlun Zhang, Shuangli Du, Wenxuan Yang, Xiaoyan Wang 0005, Yiguang Liu
ACM Trans. Embed. Comput. Syst.7
2026 Infrared and Visible Image Fusion Using Bimodal Neuron and Dynamic Receptive Field Mechanisms
abstract
Infrared and visible image fusion (IVIF) significantly enhances scene interpretation by integrating broad-spectrum information. Drawing inspiration from specific snakes that possess an evolutionarily optimized bimodal sensory system capable of parallel processing infrared and visible radiation, we propose a novel IVIF framework incorporating two key elements: nonlinear cross-modal interactions across six distinct classes of snake bimodal neurons and dynamic center-surround receptive field organization. These biological principles are mathematically formalized and integrated within a deep neural network (DNN), optimized through an object detection region-guided loss and a frequency-dependent fusion loss that enable data-driven fusion strategy learning. Experimental results demonstrate that the optimized model effectively emulates the infrared-visible information integration observed in snake bimodal neurons. Critically, the nonlinear bimodal neurons capture a significantly greater amount of edge information and finer mid-to-high-frequency details, which are essential for the subsequent reconstruction of the fused image. Furthermore, a comprehensive evaluation of visual quality, encompassing both qualitative and quantitative assessments on six datasets, along with extensive object detection and semantic segmentation experiments using the fused images in both daytime and nighttime scenarios, demonstrates that our model outperforms traditional biologically-inspired IVIF algorithms, achieving performance comparable to SOTA DNN-based methods. The code and weights are available at https://github.com/rwerwer2024/SBNF.
Shaobing Gao, Minjie Tan, Shun Lv, Yiguang Liu, Yongjie Li 0001
IEEE Trans. Image Process.4
2025 Low-light stereo image enhancement and de-noising in the low-frequency information enhanced image space
Minghua Zhao, Xiangdong Qin, Shuangli Du, Jiahao Lyu 0001, Yiguang Liu
Expert Syst. Appl.6
2025 Biological Vision Inspired Context-Awareness Network for Various Non-Generic Object Detection
abstract
Object detection approaches are expanding by leaps and bounds with recent progress in deep learning. However, there is a considerable amount of environments hampering and challenging generic detectors in open-world scenarios, which received quite limited attention. In this paper, we focus on three specific challenging conditions: 1) targets presented with low lightness, 2) camouflaged objects merged in backgrounds, 3) complex acquisition scenarios, and present a novel end-to-end detector accordingly, termed Context-awareness Network (CANet). Specifically, we propose Global Context Encoder and Context Feature Fusion module to model the context-awareness (CA) mechanism that plays a crucial role in the human visual system (HVS) in an explicit way, which integrates both latent global and local context information to make each region of interest (RoI) more informative, and thus more discriminative. To our knowledge, such high-level mechanisms are under-explored for object detection in the literature. In addition, Global Semantic Awareness module is designed to regress positions and classify better in the process of extracting the feature. Experiments demonstrate that CANet achieves very competitive performance on the ExDark, DARK FACE, COD10K, and CURE-TSD, suggesting the effectiveness and efficiency of CANet in various challenging conditions as well as common scenarios.
Shaobing Gao, Liangtian He, Yiguang Liu
IEEE Trans. Circuits Syst. Video Technol.4
2025 FPGA-Based Low-Bit and Lightweight Fast Light Field Depth Estimation
abstract
The 3-D vision computing is a key application in unmanned systems, satellites, and planetary rovers. Learning-based light field (LF) depth estimation is one of the major research directions in 3-D vision computing. However, conventional learning-based depth estimation methods involve a large number of parameters and floating-point operations, making it challenging to achieve low-power, fast, and high-precision LF depth estimation on a field-programmable gate array (FPGA). Motivated by this issue, an FPGA-based low-bit, lightweight LF depth estimation network (L$^{3}\text {FNet}$) is proposed. First, a hardware-friendly network is designed, which has small weight parameters, low computational load, and a simple network architecture with minor accuracy loss. Second, we apply efficient hardware unit design and software-hardware collaborative dataflow architecture to construct an FPGA-based fast, low-bit acceleration engine. Experimental results show that compared with the state-of-the-art works with lower mean-square error (mse), L$^{3}\text {FNet}$can reduce the computational load by more than 109 times and weight parameters by approximately 78 times. Moreover, on the ZCU104 platform, it requires 95.65% lookup tables (LUTs), 80.67% digital signal processors (DSPs), 80.93% BlockRAM (BRAM), 58.52% LUTRAM, and 9.493-W power consumption to achieve an efficient acceleration engine with a latency as low as 272 ns. The code and model of the proposed method are available athttps://github.com/sansi-zhang/L3FNet.
Chuanlun Zhang, Wenxuan Yang, Chuanjun Zhao, Shuangli Du, Yiguang Liu
IEEE Trans. Very Large Scale Integr. Syst.8
2024 Improved YOLOv5 Algorithm for Small Object Detection in Drone Images
Yitong Lin, Yiguang Liu
CVM (2)2
2024 AST: An Attention-Guided Segment Transformer for Drone-Based Cross-View Geo-Localization
Zichuan Zhao, Tianhang Tang, Xuelei Shi, Yiguang Liu
CVM (2)5
2024 Multi-Modality Speech Recognition Driven by Background Visual Scenes
abstract
Visual information is often used as a complementary cue for automatic speech recognition in noisy environments. Most previous studies utilize visual information of target speakers (e.g., lip movements) to improve the recognition performance of audio-visual speech recognition (AVSR) models. However, it remains unclear whether visual information of background sound can benefit automatic speech recognition. Our study proceeds in this regard by constructing a new audiovisual dataset and devising an AVSR model. The new dataset, Audio-Visual Natural Scenes (abbreviated as AVNS) dataset, consists of 11 types of natural scenes (around 31.3 hours) and was recorded through professional recording devices. The AVNS dataset provides audio and visual signals of common background noises in natural acoustic scenes. The AVSR model was designed based on a representation learning framework called AV-HuBERT, which could fuse representations of audio and visual modalities for automatic speech recognition. In this work, we combined the AVNS dataset (providing background sound) with the largest benchmark LRS3 dataset (providing target speech) to create adverse noise conditions for the AVSR model. The results showed that incorporating visual information synchronized with background noises greatly improved model performance (reducing WER by up to 4.9%) in noisy environments. These findings demonstrate that noise-related visual information can contribute to model performance in automatic speech recognition.
Yiguang Liu, Wenhui Sun, Zhoujian Sun
ICASSP2
2024 A Point-Line Features Fusion Method for Fast and Robust Monocular Visual-Inertial Initialization
abstract
Fast and robust initialization is essential for highly accurate monocular visual-inertial odometer (VIO), but at present majority of initialization methods rely only on point features, unstable in low texture and blurring situations. Therefore, we propose a novel point-line features fusion method for monocular visual-inertial initialization, as line features are more stable and provide richer geometric information than point features: 1) a closed-form line features initialization method is presented, and combined with point features to obtain a more integrated and robust linear system; 2) a monocular depth network is adopted to provide learned affine-invariant depth map, requiring only one prior depth map for the first frame, which can improve performance under low-parallax scenarios; 3) we can easily use RANSAC to reject outliers in solving linear system based on our formulation. Moreover, line feature re-projection residual is added to visual-inertial bundle adjustment (VI-BA) to obtain more accurate initial parameters. The proposed method is more accurate and robust than state-of-the-art methods due to the line features, especially under extreme low-parallax scenarios, and extensive experiments on popular datasets have confirmed, 0.5s initialization window on EuRoC MAV, 0.3s initialization window on TUM-VI, while the standard method normally waits for a window of 2s.
Guoqiang Xie, Tianhang Tang, Ling Lei 0003, Yiguang Liu
IROS6
2024 Disentanglement then reconstruction: Unsupervised domain adaptation by twice distribution alignments
Lihua Zhou, Mao Ye 0001, Xinpeng Li 0005, Ce Zhu, Yiguang Liu, Xue Li 0001
Expert Syst. Appl.5
2024 Precise Depth Estimation by Calculating Affine Transformation Parameters
abstract
When estimating depth from Unmanned Aerial Vehicle (UAV) consecutive frames, usually only narrow baseline stereo is satisfied. So the estimation needs a higher image matching accuracy often not achieved in real applications. To tackle this problem, we propose a precise and robust depth estimation scheme: first, by multi-view geometry theory, we prove that two corresponding local patches in two consecutive frames satisfy an affine transformation no matter how high the UAV flies and whether it is parallel to the ground or not; second, normal curvatures, local Fourier moment constraints and spatial smoothness constraints are combined to calculate the affine transformation matrix, and the Stepwise Least-Squares Fitting-based PC (SLSF-PC) method is used to calculate translation vectors because it can suppress the illumination effect; third, the normal and the position of the 3D patch corresponding to two image patches are calculated by the obtained affine transformation matrix and translation vector. The method is very precise and robust due to the combination of 2D image features and 3D differential geometry characters such as curvatures and normals. The developed method is tested using real UAV images and images of scale model scenes, confirming the superior performance of the proposed method against the state-of-the-art.
Xuelei Shi, Tianhang Tang, Shun Lv, Yiguang Liu
IEEE Trans. Geosci. Remote. Sens.5
2023 HGIM: Influence maximization in diffusion cascades from the perspective of heterogeneous graph
Yunan Zheng, Yiguang Liu
Appl. Intell.3
2023 Feature similarity rank-based information distillation network for lightweight image superresolution
Haoran Yang 0008, Gwanggil Jeon, Kai Liu 0012, Yiguang Liu, Xiaomin Yang
Knowl. Based Syst.4
2023 A new image decomposition approach using pixel-wise analysis sparsity model
Shuangli Du, Yiguang Liu, Minghua Zhao, Zhenzhen You
Pattern Recognit.2
2023 Motion-Driven Spatial and Temporal Adaptive High-Resolution Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
Graph convolutional networks (GCN) have attracted increasing interest in action recognition in recent years. GCN models human skeleton sequences as spatio-temporal graphs. Also, attention mechanisms are often jointly used with GCNs to highlight important frames or body joints in a sequence. However, attention modules learn parameters offline and are fixed, so may not adapt well to unseen samples. In this paper, we propose a simple but effective motion-driven spatial and temporal adaptation strategy to dynamically strengthen the features of important frames and joints for skeleton-based action recognition. The rationale is that the joints and frames with dramatic motions are generally more informative and discriminative. We combine the spatial and temporal refinements by using a two-branch structure, in which the joint and frame-wise feature refinements perform in parallel. Such a structure can lead to learn more complementary feature representations. Moreover, we propose to use the fully connected graph convolution to learn the long-range spatial dependencies. Besides, we investigate two high-resolution skeleton graphs by creating virtual joints, aiming to improve the representation of skeleton features. By combining the above proposals, we develop a novel motion-driven spatial and temporal adaptive high-resolution GCN. Experimental results demonstrate that the proposed model achieves state-of-the-art (SOTA) results on the challenging large-scale Kinetics-Skeleton and UAV-Human datasets, and it is on par with the SOTA methods on the two NTU-RGB+D 60&120 datasets. Additionally, our motion-driven adaptation method shows encouraging performance when compared with the attention mechanisms.
Zengxi Huang, Yusong Qin, Xiaobing Lin, Tianlin Liu, Zhenhua Feng 0001, Yiguang Liu
IEEE Trans. Circuits Syst. Video Technol.6
2022 TeCM-CLIP: Text-Based Controllable Multi-attribute Face Image Manipulation
Xudong Lou, Yiguang Liu, Xuwei Li
ACCV (7)2
2022 Low-Resource Similar Case Matching in Legal Domain
Jingxin Fang, Xuwei Li, Yiguang Liu
ICANN (2)3
2022 Cyclical Fusion: Accurate 3D Reconstruction via Cyclical Monotonicity
abstract
The dense correspondence estimation is crucial to RGB-D reconstruction systems. However, the projective correspondences are highly unreliable due to sensor depth and pose uncertainties. To tackle this challenge, we introduce a geometry-driven fusion framework, Cyclical Fusion. It pushes the correspondence finding forward to the 3D space instead of searching for candidates on the 2.5D projective map. Moreover, it establishes precise correspondence in two phases, coarse to fine. 1) First, the local surface (represented by a voxel) is characterized by Gaussian distribution. The Karcher-Frechet barycenter is adapted to conduct the robust approximation of covariance. Then, the metric between distributions is calculated via the L2-Wasserstein distance, and the correspondence voxel can be discovered through the nearest distribution-to-distribution model. 2) Our method utilizes an effective correspondence verification scheme derived from cyclical monotonicity related to Rockafellar's theorem. The concept of cyclical monotonicity reveals the geometrical nature of correspondences. A substantial constraint prevents the correspondences from twisting during the fusion process. Accordingly, precise point-to-point correspondence can be discovered. 3) The advection between correspondences is used to form a smooth manifold under regularization terms. Finally, Cyclical Fusion is integrated into a prototype reconstruction system (utilize multiple streams: depth, pose, RGB, and infrared). Experimental results on different benchmarks and real-world scanning verify the superior performance of the proposed method. Cyclical Fusion accomplishes the most authentic reconstruction for which the original projective correspondence-based scheme failed (See Fig.1). Our new techniques make the reconstruction applicable for multimedia content creation and many others.
Duo Chen 0001, Yiguang Liu
ACM Multimedia3
2022 Class Discriminative Adversarial Learning for Unsupervised Domain Adaptation
abstract
As a state-of-the-art family of Unsupervised Domain Adaptation (UDA), bi-classifier adversarial learning methods are formulated in an adversarial (minimax) learning framework with a single feature extractor and two classifiers. Model training alternates between two steps: (I) constraining the learning of the two classifiers to maximize the prediction discrepancy of unlabeled target domain data, and (II) constraining the learning of the feature extractor to minimize this discrepancy. Despite being an elegant formulation, this approach has a fundamental limitation: Maximizing and minimizing the classifier discrepancy is not class discriminative for the target domain, finally leading to a suboptimal adapted model. To solve this problem, we propose a novel Class Discriminative Adversarial Learning (CDAL) method characterized by discovering class discrimination knowledge and leveraging this knowledge to discriminatively regulate the classifier discrepancy constraints on-the-fly. This is realized by introducing an evaluation criterion for judging each classifier's capability and each target domain sample's feature reorientation via objective loss reformulation. Extensive experiments on three standard benchmarks show that our CDAL method yields new state-of-the-art performance. Our code is made available at https://github.com/buerzlh/CDAL.
Lihua Zhou, Mao Ye 0001, Xiatian Zhu, Shuaifeng Li, Yiguang Liu
ACM Multimedia5
2022 A comprehensive survey: Image deraining and stereo-matching task-driven performance analysis
abstract
Abstract Deraining has been attracting a lot of attention from researchers, and various methods have been proposed, especially deep‐networks are widely adopted in recent years. Their structures and learning become more and more complicated and diverse, making it difficult to analyze the contributions and improvements. In this paper, a comprehensive review for current rain removal methods is first provided to show their contributions. Specifically, they are reviewed in terms of handing rain streaks and rain mist. Second, besides evaluating their rain removal ability, they are also evaluated in terms of their impact on subsequent stereo‐matching task. To this end, a new deraining dataset is first prepared, called Rain‐Kitti2012 and Rain‐Kitti2015. They are created by adding rain part to clean image‐pairs in Kitti2012 and Kitti2015. By then, nine state‐of‐the‐art deraining methods are evaluated with full‐reference and no‐reference image quality assessment metrics. Furthermore, the blurriness and distortion types introduced during deraining are measured. Finally, three learning‐based stereo matching methods are compared, and they take the outputs of deraining methods as inputs. It is further discussed how derained images influence the accuracy of stereo matching, which can provide some insight for jointly handling rain removal and stereo matching. 1: A comprehensive review for the current rain removal methods is provided. They are categorized into rain‐streak‐oriented and rain‐mist‐oriented approaches in terms of degradation type, and are categorized into model‐driven and data‐driven approaches in terms of methodology. 2: A new image deraining dataset is introduced, which is the first dataset that can be used to perform stereo‐matching‐driven evaluation for deraining methods. The dataset is created by adding rain part to clean images in KITTI2012 and KITTI2015. 3: We evaluate 9 deep learning based deraining methods with full‐reference and no‐ reference metrics. In addition, the types of distortions produced by these methods are discussed and measured quantitatively. And, the impact of 9 deraining methods on the subsequent stereo matching task is evaluated, which can provide some insight on how to design stereo matching task‐driven deraining methods.
Shuangli Du, Yiguang Liu, Minghua Zhao, Zhenghao Shi, Zhenzhen You
IET Image Process.2
2022 Relay knowledge distillation for efficiently boosting the performance of shallow networks
Shipeng Fu, Zhibing Lai, Yulun Zhang 0001, Yiguang Liu, Xiaomin Yang
Neurocomputing4
2022 Cross-Domain Few-Shot Classification based on Lightweight Res2Net and Flexible GNN
Yu Chen 0119, Yunan Zheng, Tianhang Tang, Yiguang Liu
Knowl. Based Syst.7
2022 Real-time 3D reconstruction using point-dependent pose graph optimization framework
Tianhang Tang, Yiguang Liu
Mach. Vis. Appl.4
2022 Source data-free domain adaptation for a faster R-CNN
Mao Ye 0001, Yan Gan, Yiguang Liu
Pattern Recognit.5
2021 Gaussian Fusion: Accurate 3D Reconstruction via Geometry-Guided Displacement Interpolation
abstract
Reconstructing delicate geometric details with consumer RGB-D sensors is challenging due to sensor depth and poses uncertainties. To tackle this problem, we propose a unique geometry-guided fusion framework: 1) First, we characterize fusion correspondences with the geodesic curves derived from the mass transport problem, also known as the Monge-Kantorovich problem. Compared with the depth map back-projection methods, the geodesic curves reveal the geometric structures of the local surface. 2) Moving the points along the geodesic curves is the core of our fusion approach, guided by local geometric properties, i.e., Gaussian curvature and mean curvature. Compared with the state-of-the-art methods, our novel geometry-guided displacement interpolation fully utilizes the meaningful geometric features of the local surface. It makes the reconstruction accuracy and completeness improved. Finally, a significant number of experimental results on real object data verify the superior performance of the proposed method. Our technique achieves the most delicate geometric details on thin objects for which the original depth map back-projection fusion scheme suffers from severe artifacts (See Fig. 1).
Duo Chen 0001, Yunan Zheng, Yiguang Liu
ICCV5
2021 Rain streaks removal from single image based on texture constraint of background scene
Shuangli Du, Yiguang Liu, Mao Ye 0001, Minghua Zhao
Neurocomputing2
2020 MARMVS: Matching Ambiguity Reduced Multiple View Stereo for Efficient Large Scale Scene Reconstruction
abstract
The ambiguity in image matching is one of main factors decreasing the quality of the 3D model reconstructed by PatchMatch based multiple view stereo. In this paper, we present a novel method, matching ambiguity reduced multiple view stereo (MARMVS) to address this issue. The MARMVS handles the ambiguity in image matching process with three newly proposed strategies: 1) The matching ambiguity is measured by the differential geometry property of image surface with epipolar constraint, which is used as a critical criterion for optimal scale selection of every single pixel with corresponding neighbouring images. 2) The depth of every pixel is initialized to be more close to the true depth by utilizing the depths of its surrounding sparse feature points, which yields faster convergency speed in the following PatchMatch stereo and alleviates the ambiguity introduced by self similar structures of the image. 3) In the last propagation of the PatchMatch stereo, higher priorities are given to those planes with the related 2D image patch possesses less ambiguity, this strategy further propagates a correctly reconstructed surface to raw texture regions. In addition, the proposed method is very efficient even running on consumer grade CPUs, due to proper parameterization and discretization in the depth map computation step. The MARMVS is validated on public benchmarks, and experimental results demonstrate competing performance against the state of the art.
Yiguang Liu, Xuelei Shi, Yunan Zheng
CVPR2
2020 Generic wavelet-based image decomposition and reconstruction framework for multi-modal data analysis in smart camera applications
abstract
Effective acquisition, analysis and reconstruction of multi‐modal data such as colour and multi‐/hyper‐spectral imagery is crucial in smart camera applications, where wavelet‐based coding and compression of images are highly demanded. Many existing discrete wavelet filtering banks have fixed coefficients hence their performance is highly dependent on the signal/image being processed. To tackle this problem, a unified framework is proposed in this study, which can produce a series of discrete wavelet filtering banks, where many existing discrete wavelet filtering banks become special cases of the framework. For each generated filtering bank, it consists of two decomposition filters and two reconstruction filters through an optimisation process. The efficacy of the filtering banks produced by the framework has been validated in two case studies, including colour image decomposition and reconstruction, and hyperspectral image classification. Comprehensive experiments have demonstrated the superior performance of the proposed framework, which will benefit the efficacy of smart camera and camera network applications.
Yijun Yan, Yiguang Liu, Huimin Zhao 0001, Yanmei Chai, Jinchang Ren
IET Comput. Vis.2
2019 A Graph-Structured Representation with BRNN for Static-based Facial Expression Recognition
abstract
Facial expression is controlled by facial muscle and can be considered as appearance and geometric variation of key parts. One key challenging issue of static-based facial expression recognition is to capture effective information from a single facial image. In this paper, we propose a graph representation with Bidirectional RNN (BRNN) for static-based facial expression recognition. Each node on the graph represents appearance information around the facial landmarks. Edges represent the geometric information encoded by the distance between two nodes. A bidirectional recurrent neural network utilized to process the graph extracts the appearance and geometric representation. The final representation from BRNN is fed into a fully connected layer and a Softmax layer to infer expressions. Experimental results show that this method achieves significant improvements over the state-of-art methods on three widely used facial databases (Oulu-CASIA, CK+, and MMI), and our method reduces the error rates of the previous best methods by 42.2%, 35.9% and 18.7%, respectively.
Changmin Bai, Jianfeng Li 0003, Tong Chen 0008, Shigang Li 0001, Yiguang Liu
FG6
2019 A 3D Surface Reconstruction Method Based on Delaunay Triangulation
Wenjuan Miao, Yiguang Liu, Xuelei Shi, Jingming Feng, Kai Xue
ICIG (2)2
2019 Camera Pose Free Depth Sensing Based on Focus Stacking
Kai Xue, Yiguang Liu, Weijie Hong, Wenjuan Miao
ICIG (3)2
2019 Adaptive Deep Convolutional Neural Networks for Scene-Specific Object Detection
abstract
A deep convolutional neural network (CNN) becomes a widely used tool for object detection. Many previous works have achieved excellent performance on object detection benchmarks. However, these works present generic detectors whose performance will drop rapidly when they are applied to a surveillance scene. In this paper, we propose an efficient method to construct a scene-specific regression model based on a generic CNN-based classifier. Our regression model is an adaptive deep CNN (ADCNN), which can predict object locations in the surveillance scene. First, we transfer the generic CNN-based classifier to the surveillance scene by selecting useful kernels. Second, we learn the context information of the surveillance scene in our regression model for accurate location prediction. Our main contributions are: 1) a transfer learning method that selects useful kernels in the generic CNN-based classifier; 2) a special architecture that can effectively learn the local and global context information in the surveillance scene; and 3) a new objective function to effectively train parameters in ADCNN. Compared with some state-of-the-art models, ADCNN achieves the best performance on three surveillance data sets for pedestrian detection and one surveillance data set for vehicle detection.
Xudong Li 0001, Mao Ye 0001, Yiguang Liu, Ce Zhu
IEEE Trans. Circuits Syst. Video Technol.3
2018 Logistic Regression of Point Matches for Accurate Transformation Estimation
abstract
Feature extraction and matching (FEM) has been widely used for the registration of partially overlapping 3D shapes. Due to various factors such as imaging noise, simple geometry, or clutter, it usually introduces false positive ones. To reliably estimate the underlying transformation that brings one partial shape into the best possible alignment with another, it is critical to estimate the extent to which the established point matches are correct. To this end, we propose to use the logit function for the regression of the errors of these point matches. The novel method includes three steps: (i) normalization of the errors of the point matches, (ii) logistic regression of the point matches for the estimation of their reliabilities/weights, and (iii) estimation of the underlying transformation in the weighted least squares sense. These steps are repeated until either the maximum number of iterations has been reached or the weighted average of the errors of the point matches has been below the scanning resolution. A comparative study using real data captured by different range sensors shows that the proposed method outperforms two state-of-the-art ones for more accurate estimation of the underlying transformation.
Yonghuai Liu, Yitian Zhao, Yanquan Zhou, Jiwan Han, Wanneng Yang, Yiguang Liu
3DV8
2018 A New Monocular 3D Object Detection with Neural Network
Weijie Hong, Yiguang Liu, Yunan Zheng, Xuelei Shi
PRCV (4)2
2018 Improve the Spoofing Resistance of Multimodal Verification with Representation-Based Measures
Zengxi Huang, Zhenhua Feng 0001, Josef Kittler, Yiguang Liu
PRCV (3)4
2018 A sparse representation based pansharpening method
Xiaomin Yang, Lihua Jian, Binyu Yan, Kai Liu 0012, Lei Zhang 0005, Yiguang Liu
Future Gener. Comput. Syst.6
2018 Single image deraining via decorrelating the rain streaks and background scene in gradient domain
Shuangli Du, Yiguang Liu, Mao Ye 0001, Jian Guo Liu 0005
Pattern Recognit.2
2018 Retinex-based image enhancement framework by using region covariance filter
Fuyu Tao, Xiaomin Yang, Wei Wu 0002, Kai Liu 0012, Yiguang Liu
Soft Comput.6
2018 Multiple rotation symmetry group detection via saliency-based visual attention and Frieze expansion pattern
Ronggang Huang, Yiguang Liu, Pengfei Wu 0002, Yongtao Shi
Signal Process. Image Commun.2
2018 A Fractional-Order Variational Framework for Retinex: Fractional-Order Partial Differential Equation-Based Formulation for Multi-Scale Nonlocal Contrast Enhancement with Texture Preserving
abstract
This paper discusses a novel conceptual formulation of the fractional-order variational framework for retinex, which is a fractional-order partial differential equation (FPDE) formulation of retinex for the multi-scale nonlocal contrast enhancement with texture preserving. The well-known shortcomings of traditional integer-order computation-based contrast-enhancement algorithms, such as ringing artefacts and staircase effects, are still in great need of special research attention. Fractional calculus has potentially received prominence in applications in the domain of signal processing and image processing mainly because of its strengths like long-term memory, nonlocality, and weak singularity, and because of the ability of a fractional differential to enhance the complex textural details of an image in a nonlinear manner. Therefore, in an attempt to address the aforementioned problems associated with traditional integer-order computation-based contrast-enhancement algorithms, we have studied here, as an interesting theoretical problem, whether it will be possible to hybridize the capabilities of preserving the edges and the textural details of fractional calculus with texture image multi-scale nonlocal contrast enhancement. Motivated by this need, in this paper, we introduce a novel conceptual formulation of the fractional-order variational framework for retinex. First, we implement the FPDE by means of the fractional-order steepest descent method. Second, we discuss the implementation of the restrictive fractional-order optimization algorithm and the fractional-order Courant-Friedrichs-Lewy condition. Third, we perform experiments to analyze the capability of the FPDE to preserve edges and textural details, while enhancing the contrast. The capability of the FPDE to preserve edges and textural details is a fundamental important advantage, which makes our proposed algorithm superior to the traditional integer-order computation-based contrast enhancement algorithms, especially for images rich in textural details.
Yi-Fei Pu, Patrick Siarry, Amitava Chatterjee, Zhengning Wang, Zhang Yi 0001, Yiguang Liu, Jiliu Zhou, Yan Wang 0015
IEEE Trans. Image Process.6
2018 Maximizing Sensor Lifetime with the Minimal Service Cost of a Mobile Charger in Wireless Sensor Networks
abstract
Wireless energy transfer technology based on magnetic resonant coupling has emerged as a promising technology for wireless sensor networks, by providing controllable yet continual energy to sensors. In this paper, we study the use of a mobile charger to wirelessly charge sensors in a rechargeable sensor network so that the sum of sensor lifetimes is maximized while the travel distance of the mobile charger is minimized. Unlike existing studies that assumed a mobile charger must charge a sensor to its full energy capacity before moving to charge the next sensor, we here assume that each sensor can be partially charged so that more sensors can be charged before their energy depletions. Under this new energy charging model, we first formulate two novel optimization problems of scheduling a mobile charger to charge a set of sensors, with the objectives to maximize the sum of sensor lifetimes and to minimize the travel distance of the mobile charger while achieving the maximum sum of sensor lifetimes, respectively. We then propose efficient algorithms for the problems. We finally evaluate the performance of the proposed algorithms through experimental simulations. Simulation results demonstrate that the proposed algorithms are very promising. Especially, the average energy expiration duration per sensor by the proposed algorithm for maximizing the sum of sensor lifetimes is only 9 percent of that by the state-of-the-art algorithm while the travel distance of the mobile charger by the second proposed algorithm is only about from 1 to 15 percent longer than that by the state-of-the-art benchmark.
Wenzheng Xu, Weifa Liang, Xiaohua Jia, Zichuan Xu, Yiguang Liu
IEEE Trans. Mob. Comput.6
2017 Disparity Refinement Using Merged Super-Pixels for Stereo Matching
Jianyu Heng, Yunan Zheng, Yiguang Liu
ICIG (1)4
2017 Memory-based pedestrian detection through sequence learning
abstract
Human recognize an object through eyes scanning in a certain order. We think that the proper order is helpful for capturing useful characteristics, which makes our recognition process rapidly and accurately. Therefore, we propose a memory-based sequence learning model to simulate the human recognition process. Firstly, we divide the image without overlapping to generate the sequence. Then, a convolutional neural network is used for feature extraction. Next, the sequence is re-sorted by order of importance. Finally, a long short-term memory successively receives the sequence to memorize the sequential patterns and predict the sequence label. In addition, we propose a joint learning method to make our model efficiently learn both of the sequence order and the sequence patterns. Our model is applied in the region-based detection framework for pedestrian detection. Compared with the state-of-the-art methods on two pedestrian datasets, our method achieves the comparable performance in term of accuracy and speed.
Xudong Li 0001, Mao Ye 0001, Yiguang Liu, Ce Zhu
ICME3
2017 The Euclidean embedding learning based on convolutional neural network for stereo matching
Menglong Yang, Yiguang Liu, Zhisheng You
Neurocomputing2
2017 Accurate object detection using memory-based models in surveillance scenes
Xudong Li 0001, Mao Ye 0001, Yiguang Liu, Feng Zhang 0052, Song Tang 0001
Pattern Recognit.3
2017 A dense flow-based framework for real-time object registration under compound motion
Songfan Yang, Yinjie Lei, Mingyang Li 0001, Ninad Thakoor, Bir Bhanu, Yiguang Liu
Pattern Recognit.7
2017 The Mapping-Adaptive Convolution: A Fundamental Theory for Homography or Perspective Invariant Matching Methods
abstract
If the local area of a three-dimensional object surface can be considered as a plane, the deformation between its two images captured from different camera placements is modelled by a homography. By tuning the parameters in a homographic mapping, all possible deformations caused by the change of camera placement can be simulated for the local feature matching method. Since aliasing may happen when resampling the original image to the geometry of the simulated image, an antialiasing convolution must be applied before resampling. However, the antialiasing convolution itself must also be homography-adaptive. In the scale invariant feature transform (SIFT) or affine-SIFT (ASIFT) method, the similitude or affine rectification scheme of the convolution is applied to solve this problem under similitude or affine mapping. However, these schemes will not work under homographic or perspective mapping. Although the perspective invariant matching method (perspective-SIFT or PSIFT) has been proposed in some references, the antialiasing scheme with perspective-adaption has not been proposed. This paper will show that the standard convolution is not adaptive to the change of planar mapping, and the simulated images under the same simulated camera placement will not be identical if they are resampled from different original images captured from different camera placements. To solve this issue, a natural extension of the standard convolution, the mapping-adaptive convolution (MA-convolution), is proposed, and its mapping-adaption is proved mathematically in this paper. Based on this novel convolution, the homography invariant simulation scheme can be modelled. We have applied the MA-convolution to the antialiasing scheme in the PSIFT method, and the effectiveness of the MA-convolution has been verified experimentally.
Yiguang Liu, Jipeng Li, Wenzheng Xu
SIAM J. Imaging Sci.2
2017 Geometry Guided Multi-Scale Depth Map Fusion via Graph Optimization
abstract
In depth discontinuous and untextured regions, depth maps created by multiple view stereopsis are with heavy noises, but existing depth map fusion methods cannot handle it explicitly. To tackle the problem, two novel strategies are proposed: 1) a more discriminative fusion method, which is based on geometry consistency, measuring the consistency, and stability of surface geometry computed on both partial and global surfaces, different from traditional methods only using visibility consistency; 2) a graph optimization method which fuses pyramids of depth maps as mutual complementary information is available in different scales, and differs from existing multi-scale fusion methods. The method considers both sampling scale of a point and relations among points, and is proven to be solvable by graph cuts. Experimental results verify the superior performance of the proposed method to the traditional visibility consistency-based methods, and the proposed method is also compared favorably with a number of state-of-the-art methods. Moreover, the proposed method achieves the highest completeness among all the methods compared.
Pengfei Wu 0002, Yiguang Liu, Mao Ye 0001, Yunan Zheng
IEEE Trans. Image Process.2
2017 Maximizing Charging Satisfaction of Smartphone Users via Wireless Energy Transfer
abstract
Smartphones now become an indispensable part of our daily life. However, maintaining a smartphone's continuing operation consumes lots of battery energy. For example, a fully-charged smartphone usually cannot support its continuing operation for a whole day. A fundamental issue on a smartphone is its energy issue. That is, how to prolong the lifetime of a smartphone so that it can run as long as possible to meet its user needs. Wireless energy transfer has been demonstrated as a promising technique to address this issue. In this paper, we study a novel smartphone charging problem, through wireless chargers deployed on public commuters, e.g., subway trains, to charge energy-critical smartphones when their users take subway trains to work or go home. Since the amounts of residual energy of different smartphones are significantly different, the charging satisfactions of different users are essentially different. In this paper, we formulate this charging satisfaction problem as a novel optimization problem that schedules the limited number of wireless chargers on subway trains to charge energy-critical smartphones such that the overall charging satisfaction of smartphone users is maximized, for a given monitoring period (e.g., one day). Forthis problem, we first devise a 1/3-approximation algorithm if the travel trajectory of each smartphone user is given. We then propose an online algorithm to deal with dynamic energy-critical smartphone charging requests. We also propose a nontrivial distributed scheduling algorithm for a variant of the problem where the global knowledge of user energy information is unknown. We finally evaluate the performance of the proposed algorithms through experimental simulations, using a real dataset of subway-taking in San Francisco. The experimental results show that the proposed algorithms are very promising, and over 90 percent of energy-critical user smartphones can be satisfactorily charged in a one-day monitoring period.
Wenzheng Xu, Weifa Liang, Jian Peng 0002, Yiguang Liu, Yan Wang 0015
IEEE Trans. Mob. Comput.4
2017 Fast and Adaptive 3D Reconstruction With Extensively High Completeness
abstract
The seed-and-expand scheme is appropriate for multiple view stereo, since it can build dense point clouds adaptively by avoiding unnecessary computation. However, due to the irregularity of the algorithm, it is not suitable for parallel computing on general public utilities (GPU). This paper is the first attempt to implement the irregular seed-and-expand method on GPU for multiple view stereo problems. Meanwhile, a hierarchical parallel computing architecture is also proposed to maximize the usage of both CPU and GPU. The adaptivity of the seed-and-expand scheme is pushed further by processing a pixel several rounds while, in order to maintain regularity for GPU implementation, every seed has exactly the same behavior in a single round of optimization. The high adaptivity also improves the robustness of the proposed method, thus aggressive matching score and a view selection method can be used to improve the reconstruction completeness extensively, without smearing out local details and lowering the accuracy. Compared with the state of the art, the proposed method achieves higher accuracy and completeness on standard datasets. The proposed method is also very fast. It is maximally five times faster than other methods running on a CPU and is on par with the regular depth map-based methods on GPU, which are naturally suitable for GPU acceleration.
Pengfei Wu 0002, Yiguang Liu, Mao Ye 0001, Shuangli Du
IEEE Trans. Multim.2
2016 Stereo matching based on classification of materials
Menglong Yang, Yiguang Liu, Ying Cai 0002, Zhisheng You
Neurocomputing2
2016 A new framework for remote sensing image super-resolution: Sparse representation-based method by processing dictionaries with multi-type features
Wei Wu 0002, Xiaomin Yang, Kai Liu 0012, Yiguang Liu, Binyu Yan, Hua Hua
J. Syst. Archit.4
2016 DFOB: Detecting and describing features by octagon filter bank for fast image matching
Yiguang Liu, Shuangli Du, Pengfei Wu 0002
Signal Process. Image Commun.2
2016 Hierarchical and Adaptive Phase Correlation for Precise Disparity Estimation of UAV Images
abstract
When using fixed-window phase correlation (PC) to estimate the disparity of stereo images, the precision is usually rather poor due to large depth differences of scenes and noise, and this problem is specially severe when using unmanned aerial vehicle (UAV) image pairs to extract the digital elevation model of mountain land. To tackle this problem, this paper proposes a hierarchical and adaptive PC, which includes three steps: First, PC with the initialized window is performed to coarsely estimate a disparity value, along with the peak of the Dirichlet function for each pixel; then, an additional round of PC is performed for each pixel using the window of smaller size and with being guided by the coarsely estimated disparity; finally, the previous two steps are iteratively performed until convergence. In particular, using the peak of the Dirichlet function of each pixel in step two, we can drop out the influence of dramatically changing areas such as river; moreover, the scheme can minimize the influence of boundary overreach. The novel scheme has been tested on a large number of UAV images captured at mountainous regions in southwest China, showing that the proposed method is superior to the state-of-the-art methods, especially in handling UAV images of the high mountains and rivers.
Yiguang Liu, Shuangli Du, Pengfei Wu 0002
IEEE Trans. Geosci. Remote. Sens.2
2016 L1-Norm Low-Rank Matrix Decomposition by Neural Networks and Mollifiers
abstract
The L1-norm cost function of the low-rank approximation of the matrix with missing entries is not smooth, and also cannot be transformed into a standard linear or quadratic programming problem, and thus, the optimization of this cost function is still not well solved. To tackle this problem, first, a mollifier is used to smooth the cost function. High closeness of the smoothed function to the original one can be obtained by tuning the parameters contained in the mollifier. Next, a recurrent neural network is proposed to optimize the mollified function, which will converge to a local minimum. In addition, to boost the speed of the system, the mollifying process is implemented by a filtering procedure. The influence of two mollifier parameters is theoretically analyzed and experimentally confirmed, showing that one of the parameters is critical to computational efficiency and accuracy, while the other not. A large number of experiments on synthetic data show that the proposed method is competitive to the state-of-the-art methods. In particular, the experiments on large matrices and a real application in the structure from motion indicate that the memory requirement of the proposed algorithm is mild, making it suitable for real applications that often involve large-scale matrix decomposition.
Yiguang Liu, Songfan Yang, Pengfei Wu 0002, Chunguang Li 0001, Menglong Yang
IEEE Trans. Neural Networks Learn. Syst.1
2015 Improving PART algorithm with K-L divergence for imbalanced classification
abstract
Rule-learning extracts the knowledge from a dataset and represent it in a form that is easy for people to understand. RIPPER (Repeated Incremental Pruning to Produce Error Reduction) and PART (Partial Decision Trees) are two well-known schemes for rule-learning. However, due to overpruning of RIPPE R and skew-sensitivity of PART, it is difficult to use two methods to learn from imbalanced datasets. To bypass these difficulties, we propose a K-L divergence-based PART (KLPART) that use K-L divergence as a splitting criterion to build partial decision trees. An experimental framework is carried out with a wide range of imbalanced datasets over RIPPER, PART, KLPART and the combination of these methods for classification with SMOTE processing. The results obtained, which contrasted through nonparametric statistical tests, show that KLPART is robust in the presence of class imbalance, especially when combined with SMOTE. We thereby recommend the use of KLPART with SMOTE when learning from imbalanced datasets.
Chong Su, Shenggen Ju, Yiguang Liu, Zhonghua Yu
Intell. Data Anal.3
2015 Improving Random Forest and Rotation Forest for highly imbalanced datasets
abstract
Decision tree is a simple and effective method and it can be supplemented with ensemble methods to improve its performance. Random Forest and Rotation Forest are two approaches which are perceived as ``classic'' at present. They can build more accurate and diverse classifiers than Bagging and Boost ing by introducing the diversities namely randomly chosen a subset of features or rotated feature space. However, the splitting criteria used for constructing each tree in Random Forest and Rotation Forest are Gini index and information gain ratio respectively, which are skew-sensitive. When learning from highly imbalanced datasets, class imbalance impedes their ability to learn the minority class concept. Hellinger distance decision tree (HDDT) was proposed by Chawla, which is skew-insensitive. Especially, bagged unpruned HDDT has proven to be an effective way to deal with highly imbalanced problem. Nevertheless, the bootstrap sampling used in Bagging can lead to ensembles of low diversity compared to Random Forest and Rotation Forest. In order to combine the skew-insensitivity of HDDT and the diversities of Random Forest and Rotation Forest, we use Hellinger distance as the splitting criterion for building each tree in Random Forest and Rotation Forest respectively. An experimental framework is performed across a wide range of highly imbalanced datasets to investigate the effectiveness of Hellinger distance, information gain ratio and Gini index which are used as the splitting criteria in ensembles of decision trees including Bagging, Boosting, Random Forest and Rotation Forest. In addition, Balanced Random Forest is also included in the experiment since it is designed to tackle class imbalance problem. The experimental results, which contrasted through nonparametric statistical tests, demonstrate that using Hellinger distance as the splitting criterion to build individual decision tree in forest can improve the performances of Random Forest and Rotation Forest for highly imbalanced classification.
Chong Su, Shenggen Ju, Yiguang Liu, Zhonghua Yu
Intell. Data Anal.3
2015 An adaptive bimodal recognition framework using sparse coding for face and ear
Zengxi Huang, Yiguang Liu, Xuwei Li
Pattern Recognit. Lett.2
2015 Coarse-to-fine outlier correction with applications in structure from motion
Shuangli Du, Yiguang Liu, Zengxi Huang, Pengfei Wu 0002
Signal Process. Image Commun.2
2015 A Random Algorithm for Low-Rank Decomposition of Large-Scale Matrices With Missing Entries
abstract
A random submatrix method (RSM) is proposed to calculate the low-rank decomposition U(m×r)V(n×r)(T) (r < m, n) of the matrix Y∈R(m×n) (assuming m > n generally) with known entry percentage 0 < ρ ≤ 1. RSM is very fast as only O(mr(2)ρ(r)) or O(n(3)ρ(3r)) floating-point operations (flops) are required, compared favorably with O(mnr+r(2)(m+n)) flops required by the state-of-the-art algorithms. Meanwhile, RSM has the advantage of a small memory requirement as only max(n(2),mr+nr) real values need to be saved. With the assumption that known entries are uniformly distributed in Y, submatrices formed by known entries are randomly selected from Y with statistical size k×nρ(k) or mρ(l)×l , where k or l takes r+1 usually. We propose and prove a theorem, under random noises the probability that the subspace associated with a smaller singular value will turn into the space associated to anyone of the r largest singular values is smaller. Based on the theorem, the nρ(k)-k null vectors or the l-r right singular vectors associated with the minor singular values are calculated for each submatrix. The vectors ought to be the null vectors of the submatrix formed by the chosen nρ(k) or l columns of the ground truth of V(T). If enough submatrices are randomly chosen, V and U can be estimated accordingly. The experimental results on random synthetic matrices with sizes such as 13 1072 ×10(24) and on real data sets such as dinosaur indicate that RSM is 4.30 ∼ 197.95 times faster than the state-of-the-art algorithms. It, meanwhile, has considerable high precision achieving or approximating to the best.
Yiguang Liu, Yinjie Lei, Chunguang Li 0001, Wenzheng Xu, Yi-Fei Pu
IEEE Trans. Image Process.1
2015 Robust Prostate Segmentation Using Intrinsic Properties of TRUS Images
abstract
Accurate segmentation is usually crucial in transrectal ultrasound (TRUS) image based prostate diagnosis; however, it is always hampered by heavy speckles. Contrary to the traditional view that speckles are adverse to segmentation, we exploit intrinsic properties induced by speckles to facilitate the task, based on the observations that sizes and orientations of speckles provide salient cues to determine the prostate boundary. Since the speckle orientation changes in accordance with a statistical prior rule, rotation-invariant texture feature is extracted along the orientations revealed by the rule. To address the problem of feature changes due to different speckle sizes, TRUS images are split into several arc-like strips. In each strip, every individual feature vector is sparsely represented, and representation residuals are obtained. The residuals, along with the spatial coherence inherited from biological tissues, are combined to segment the prostate preliminarily via graph cuts. After that, the segmentation is fine-tuned by a novel level sets model, which integrates (1) the prostate shape prior, (2) dark-to-light intensity transition near the prostate boundary, and (3) the texture feature just obtained. The proposed method is validated on two 2-D image datasets obtained from two different sonographic imaging systems, with the mean absolute distance on the mid gland images only 1.06±0.53 mm and 1.25±0.77 mm, respectively. The method is also extended to segment apex and base images, producing competitive results over the state of the art.
Pengfei Wu 0002, Yiguang Liu, Yongzhong Li
IEEE Trans. Medical Imaging2
2014 A Probabilistic Framework for Multitarget Tracking with Mutual Occlusions
abstract
Mutual occlusions among targets can cause track loss or target position deviation, because the observation likelihood of an occluded target may vanish even when we have the estimated location of the target. This paper presents a novel probability framework for multitarget tracking with mutual occlusions. The primary contribution of this work is the introduction of a vectorial occlusion variable as part of the solution. The occlusion variable describes occlusion states of the targets. This forms the basis of the proposed probability framework, with the following further contributions: 1) Likelihood: A new observation likelihood model is presented, in which the likelihood of an occluded target is computed by referring to both of the occluded and oc-cluding targets. 2) Priori: Markov random field (MRF) is used to model the occlusion priori such that less likely "circular" or "cascading" types of occlusions have lower priori probabilities. Both the occlusion priori and the motion priori take into consideration the state of occlusion. 3) Optimization: A realtime RJMCMC-based algorithm with a newmove type called "occlusion state update" is presented. Experimental results show that the proposed framework can handle occlusions well, even including long-duration full occlusions, which may cause tracking failures in the traditional methods.
Menglong Yang, Yiguang Liu, Longyin Wen, Zhisheng You, Stan Z. Li
CVPR2
2014 Fractional partial differential equation denoising models for texture image
Yi-Fei Pu, Patrick Siarry, Jiliu Zhou, Yiguang Liu, Guo Huang
Sci. China Inf. Sci.4
2014 Supervised methods for symptom name recognition in free-text clinical records of traditional Chinese medicine: An empirical study
Yaqiang Wang, Zhonghua Yu, Yunhui Chen, Yiguang Liu, Xiaoguang Hu, Yongguang Jiang
J. Biomed. Informatics5
2014 An improved SOM algorithm and its application to color feature extraction
abstract
Reducing the redundancy of dominant color features in an image and meanwhile preserving the diversity and quality of extracted colors is of importance in many applications such as image analysis and compression. This paper presents an improved self-organization map (SOM) algorithm namely MFD-SOM and its application to color feature extraction from images. Different from the winner-take-all competitive principle held by conventional SOM algorithms, MFD-SOM prevents, to a certain degree, features of non-principal components in the training data from being weakened or lost in the learning process, which is conductive to preserving the diversity of extracted features. Besides, MFD-SOM adopts a new way to update weight vectors of neurons, which helps to reduce the redundancy in features extracted from the principal components. In addition, we apply a linear neighborhood function in the proposed algorithm aiming to improve its performance on color feature extraction. Experimental results of feature extraction on artificial datasets and benchmark image datasets demonstrate the characteristics of the MFD-SOM algorithm.
Yiguang Liu, Zengxi Huang, Yongtao Shi
Neural Comput. Appl.2
2014 A homography transform based higher-order MRF model for stereo matching
Menglong Yang, Yiguang Liu, Zhisheng You, Yi Zhang 0018
Pattern Recognit. Lett.2
2013 Dynamics of a mean-shift-like algorithm and its applications on clustering
Yiguang Liu, Stan Z. Li, Wei Wu 0002, Ronggang Huang
Inf. Process. Lett.1
2013 Classification using distances from samples to linear manifolds
Yiguang Liu, Xiaochun Cao, Jian Guo Liu 0005
Pattern Anal. Appl.1
2013 Classification by nearness in complementary subspaces
Menglong Yang, Yiguang Liu, Baojiang Zhong
Pattern Anal. Appl.2
2013 A robust face and ear based multimodal biometric system using sparse representation
Zengxi Huang, Yiguang Liu, Chunguang Li 0001, Menglong Yang
Pattern Recognit.2
2013 Complex-Valued Filtering Based on the Minimization of Complex-Error Entropy
abstract
In this paper, we consider the training of complex-valued filter based on the information theoretic method. We first generalize the error entropy criterion to complex domain to present the complex error entropy criterion (CEEC). Due to the difficulty in estimating the entropy of complex-valued error directly, the entropy bound minimization (EBM) method is used to compute the upper bounds of the entropy of the complex-valued error, and the tightest bound selected by the EBM algorithm is used as the estimator of the complex-error entropy. Then, based on the minimization of complex-error entropy (MCEE) and the complex gradient descent approach, complex-valued learning algorithms for both the (linear) transverse filter and the (nonlinear) neural network are derived. The algorithms are applied to complex-valued linear filtering and complex-valued nonlinear channel equalization to demonstrate their effectiveness and advantages.
Songyan Huang, Chunguang Li 0001, Yiguang Liu
IEEE Trans. Neural Networks Learn. Syst.3
2013 Recovering shape and motion by a dynamic system for low-rank matrix approximation in L 1 norm
Yiguang Liu, Liping Cao, Yi-Fei Pu, Hong Cheng 0002
Vis. Comput.1
2012 Quasi Monte Carlo localization for mobile robots
abstract
In this paper, the authors will present a novel approach called Quasi Monte Carlo localization (QMCL) to address the problem of the inefficient uniform random sequence generated by the Monte Carlo method and hence the unnecessary large sample set used for initialization of Monte Carlo localization. With an additional new motion model, the performance of QMCL is even improved. The efficiency of the QMCL is evaluated by both simulation and experimental tests.
Yiguang Liu
ICARCV4
2012 A Pyramid Nearest Neighbor Search Kernel for object categorization
Hong Cheng 0002, Rongchao Yu, Zicheng Liu 0001, Yiguang Liu
ICPR4
2012 Low-rank matrix decomposition in L1-norm by dynamic systems
Yiguang Liu, Yi-Fei Pu, Hong Cheng 0002
Image Vis. Comput.1
2012 A framework and its empirical study of automatic diagnosis of traditional Chinese medicine utilizing raw free-text clinical records
Yaqiang Wang, Zhonghua Yu, Yongguang Jiang, Yiguang Liu
J. Biomed. Informatics6
2012 Almost periodic solution of impulsive Hopfield neural networks with finite distributed delays
Yiguang Liu, Zengxi Huang
Neural Comput. Appl.1
2012 Video synchronization based on events alignment
Yiguang Liu, Menglong Yang, Zhisheng You
Pattern Recognit. Lett.1
2011 The almost periodic solution of Lotka-Volterra recurrent neural networks with delays
Yiguang Liu, Sai-Ho Ling
Neurocomputing1
2011 Estimating the fundamental matrix based on least absolute deviation
Menglong Yang, Yiguang Liu, Zhisheng You
Neurocomputing2
2011 k-NS: A Classifier by the Distance to the Nearest Subspace
abstract
To improve the classification performance of k-NN, this paper presents a classifier, called k -NS, based on the Euclidian distances from a query sample to the nearest subspaces. Each nearest subspace is spanned by k nearest samples of a same class. A simple discriminant is derived to calculate the distances due to the geometric meaning of the Grammian, and the calculation stability of the discriminant is guaranteed by embedding Tikhonov regularization. The proposed classifier, k-NS, categorizes a query sample into the class whose corresponding subspace is proximal. Because the Grammian only involves inner products, the classifier is naturally extended into the high-dimensional feature space induced by kernel functions. The experimental results on 13 publicly available benchmark datasets show that k-NS is quite promising compared to several other classifiers founded on nearest neighbors in terms of training and test accuracy and efficiency.
Yiguang Liu, Shuzhi Sam Ge, Chunguang Li 0001, Zhisheng You
IEEE Trans. Neural Networks1
2010 The Reliability of Travel Time Forecasting
abstract
Travel time is a fundamental measure in transportation, and accurate travel time forecasting is crucial in intelligent transportation systems (ITSs). Currently, many techniques have been applied to travel time forecasting; however, the reliability of the prediction has not been studied in these approaches. In this paper, we propose an approach using the generalized autoregressive conditional heteroscedasticity (GARCH) model to study the volatility of travel time and supply the information about reliability for travel time forecasting. Three examples on real urban vehicular traffic data show the whole modeling process. In the experiments, we utilize the conditional predicted standard deviation (PSD) to express the reliability of travel time forecasting and screen out the sample points that are thought to be reliable forecasting. The results show that the root-mean-square error (RMSE), mean absolute error (MAE), and mean absolute percent error (MAPE) are all decreasing with an increase in the demand of the reliability. It proves that the model well depicts the reliability of travel time forecasting and that the proposed approach is feasible.
Menglong Yang, Yiguang Liu, Zhisheng You
IEEE Trans. Intell. Transp. Syst.2
2008 Stability analysis for the generalized Hopfield neural networks with multi-level activation functions
Yiguang Liu, Zhisheng You
Neurocomputing1
2007 On the Almost Periodic Solution of Cellular Neural Networks With Distributed Delays
abstract
By exponential dichotomy about differential equations, a formal almost periodic solution (APS) of a class of cellular neural networks (CNNs) with distributed delays is obtained. Then, within different normed spaces, several sufficient conditions guaranteeing the existence and uniqueness of an APS are proposed using two fixed-point theorems. Based on the continuity property and some inequality techniques, two theorems insuring the global stability of the unique APS are given. Comparing with known literatures, all conclusions are drawn with slacker restrictions, e.g., do not require the integral of the kernel function determining the distributed delays from zero to positive infinity to be one, and the activation functions to be bounded, etc.; besides, all criteria are obtained by different ways. Finally, two illustrative examples show the validity and that all criteria are easy to check and apply.
Yiguang Liu, Zhisheng You, Liping Cao
IEEE Trans. Neural Networks1
2006 A Concise Functional Neural Network for Computing the Extremum Eigenpairs of Real Symmetric Matrices
Yiguang Liu, Zhisheng You
ISNN (1)1
2006 On stability of disturbed Hopfield neural networks with time delays
Yiguang Liu, Zhisheng You, Liping Cao
Neurocomputing1
2006 On the almost periodic solution of generalized Hopfield neural networks with time-varying delays
Yiguang Liu, Zhisheng You, Liping Cao
Neurocomputing1
2006 A novel and quick SVM-based multi-class classifier
Yiguang Liu, Zhisheng You, Liping Cao
Pattern Recognit.1
2006 A concise functional neural network computing the largest modulus eigenvalues and their corresponding eigenvectors of a real skew matrix
Yiguang Liu, Zhisheng You, Liping Cao
Theor. Comput. Sci.1
2005 A simple functional neural network for computing the largest and smallest eigenvalues and corresponding eigenvectors of a real symmetric matrix
Yiguang Liu, Zhisheng You, Liping Cao
Neurocomputing1
2005 A functional neural network for computing the largest modulus eigenvalues and their corresponding eigenvectors of an anti-symmetric matrix
Yiguang Liu, Zhisheng You, Liping Cao
Neurocomputing1
2005 A functional neural network computing some eigenvalues and eigenvectors of a special real matrix
Yiguang Liu, Zhisheng You, Liping Cao
Neural Networks1