Cuihua Li

dblp:21/5657 · DBLP profile ↗
← Back
46ranked-venue papers
0as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Cross-Cloud Consistency for Weakly Supervised Point Cloud Semantic Segmentation
abstract
Weakly supervised point cloud semantic segmentation is an increasingly active topic, because fully supervised learning acquires well-labeled point clouds and entails high costs. The existing weakly supervised methods either need meticulously designed data augmentation for self-supervised learning or ignore the negative effects of learning on pseudolabel noises. In this article, by designing different granularity of cross-cloud structures, we propose a cross-cloud consistency method for weakly supervised point cloud semantic segmentation which forms the expectation-maximum (EM) framework. Benefiting from the cross-cloud constraints, our method allows effective learning alternatively between refining pseudolabels and updating network parameters. Specifically, in E-step, we propose a pseudolabel selecting (PLS) strategy based on cross subcloud consistency, improving the credibility of selected pseudolabels explicitly. In M-step, a cross-scene contrastive regularization enforces cross-scene prototypes with the same label in different scenes to be more similar, while keeping prototypes with different labels to be a clear margin, reducing the noise fitting. Finally, we give some insight into the optimization of our method in the EM theoretical way. The proposed method is evaluated on three challenging datasets, where experimental results demonstrate that our method significantly outperforms state-of-the-art weakly supervised competitors. Our code is available online: https://github.com/Yachao-Zhang/Cross-Cloud-Consistency.
Yachao Zhang 0001, Yuxiang Lan, Yuan Xie 0006, Cuihua Li, Yanyun Qu
IEEE Trans. Neural Networks Learn. Syst.4
2024 Adversarial Learning for Coordinate Regression Through $k$k-Layer Penetrating Representation
abstract
Adversarial attack is a crucial step when evaluating the reliability and robustness of deep neural networks (DNNs) models. Most existing attack approaches apply an end-to-end gradient update strategy to generate adversarial examples for a classification or regression problem. However, few of them consider the non-differentiable DNN models (e.g., coordinate regression model) that prevent end-to-end backpropagation resulting in the failure of gradient calculation. In this article, we present a new adversarial example generation approach for both untargeted and targeted attacks on coordinate regression models with non-differentiable operations. The novelty of our approach lies in a$k$-layer penetrating representation, on which we perturb the hidden feature distribution of the$k$th layer through relational guidance to influence the final output, in which end-to-end backpropagation is not required. Rather than modifying a large portion of the pixels in an image, the proposed approach only modifies a very small set of the input pixels. These pixels are carefully and precisely selected by three correlations between the input pixels and hidden features of the$k$th layer of a DNN, thus significantly reducing the adversarial perturbation on a clean image. We successfully apply the proposed approach to two different tasks (i.e., 2D and 3D human pose estimation) which are typical applications of the coordinate regression learning. The comprehensive experiments demonstrate that our approach achieves better performance while using much less adversarial perturbation on clean images.
Mengxi Jiang, Yulei Sui, Xiaofei Xie, Cuihua Li, Yang Liu 0003, Ivor W. Tsang
IEEE Trans. Dependable Secur. Comput.5
2024 Learning All-In Collaborative Multiview Binary Representation for Clustering
abstract
Multiview clustering via binary representation has attracted intensive attention due to its effectiveness in handling large-scale multiple view data. However, these kind of clustering approaches usually ignore a very important potential high-order correlation in discrete representation learning. In this article, we propose a novel all-in collaborative multiview binary representation for clustering (AC-MVBC) framework, where multiview collaborative binary representation and clustering structure are learned in a joint manner. Specifically, using a new type of tensor low-rank constraint, the high-order collaborations, i.e., cross-view and inner view collaborations, can be effectively captured in our model. Moreover, by incorporating the Bregman discrepancy, the projective consistency among different views can be guaranteed to achieve a more powerful binary representation. An efficient optimization algorithm is also proposed to solve the objective function with fast convergence empirically. Experimental results on several challenge datasets demonstrate that the proposed method has achieved highly competent performance compared with the state-of-the-art multiview clustering (MVC) methods while maintaining low computational and memory requirements.
Yachao Zhang 0001, Yuan Xie 0006, Cuihua Li, Zongze Wu 0001, Yanyun Qu
IEEE Trans. Neural Networks Learn. Syst.3
2023 Lattice Network for Lightweight Image Restoration
abstract
Deep learning has made unprecedented progress in image restoration (IR), where residual block (RB) is popularly used and has a significant effect on promising performance. However, the massive stacked RBs bring about burdensome memory and computation cost. To tackle this issue, we aim to design an economical structure for adaptively connecting pair-wise RBs, thereby enhancing the model representation. Inspired by the topological structure of lattice filter in signal processing theory, we elaborately propose the lattice block (LB), where couple butterfly-style topological structures are utilized to bridge pair-wise RBs. Specifically, each candidate structure of LB relies on the combination coefficients learned through adaptive channel reweighting. As a basic mapping block, LB can be plugged into various IR models, such as image super-resolution, image denoising, image deraining, etc. It can avail the construction of lightweight IR models accompanying half parameter amount reduced, while keeping the considerable reconstruction accuracy compared with RBs. Moreover, a novel contrastive loss is exploited as a regularization constraint, which can further enhance the model representation without increasing the inference expenses. Experiments on several IR tasks illustrate that our method can achieve more favorable performance than other state-of-the-art models with lower storage and computation.
Xiaotong Luo, Yanyun Qu, Yuan Xie 0006, Yulun Zhang 0001, Cuihua Li, Yun Fu 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Uncertainty-Driven Dehazing Network
abstract
Deep learning has made remarkable achievements for single image haze removal. However, existing deep dehazing models only give deterministic results without discussing the uncertainty of them. There exist two types of uncertainty in the dehazing models: aleatoric uncertainty that comes from noise inherent in the observations and epistemic uncertainty that accounts for uncertainty in the model. In this paper, we propose a novel uncertainty-driven dehazing network (UDN) that improves the dehazing results by exploiting the relationship between the uncertain and confident representations. We first introduce an Uncertainty Estimation Block (UEB) to predict the aleatoric and epistemic uncertainty together. Then, we propose an Uncertainty-aware Feature Modulation (UFM) block to adaptively enhance the learned features. UFM predicts a convolution kernel and channel-wise modulation cofficients conitioned on the uncertainty weighted representation. Moreover, we develop an uncertainty-driven self-distillation loss to improve the uncertain representation by transferring the knowledge from the confident one. Extensive experimental results on synthetic datasets and real-world images show that UDN achieves significant quantitative and qualitative improvements, outperforming the state-of-the-arts.
Ming Hong, Jianzhuang Liu, Cuihua Li, Yanyun Qu
AAAI3
2022 Cross-Domain and Cross-Modal Knowledge Distillation in Domain Adaptation for 3D Semantic Segmentation
abstract
With the emergence of multi-modal datasets where LiDAR and camera are synchronized and calibrated, cross-modal Unsupervised Domain Adaptation (UDA) has attracted increasing attention because it reduces the laborious annotation of target domain samples. To alleviate the distribution gap between source and target domains, existing methods conduct feature alignment by using adversarial learning. However, it is well-known to be highly sensitive to hyperparameters and difficult to train. In this paper, we propose a novel model (Dual-Cross) that integrates Cross-Domain Knowledge Distillation (CDKD) and Cross-Modal Knowledge Distillation (CMKD) to mitigate domain shift. Specifically, we design the multi-modal style transfer to convert source image and point cloud to target style. With these synthetic samples as input, we introduce a target-aware teacher network to learn knowledge of the target domain. Then we present dual-cross knowledge distillation when the student is learning on source domain. CDKD constrains teacher and student predictions under same modality to be consistent. It can transfer target-aware knowledge from the teacher to the student, making the student more adaptive to the target domain. CMKD generates hybrid-modal prediction from the teacher predictions and constrains it to be consistent with both 2D and 3D student predictions. It promotes the information interaction between two modalities to make them complement each other. From the evaluation results on various domain adaptation settings, Dual-Cross significantly outperforms both uni-modal and cross-modal state-of-the-art methods.
Miaoyu Li, Yachao Zhang 0001, Yuan Xie 0006, Zuodong Gao, Cuihua Li, Zhizhong Zhang 0001, Yanyun Qu
ACM Multimedia5
2022 Self-supervised Exclusive Learning for 3D Segmentation with Cross-Modal Unsupervised Domain Adaptation
abstract
2D-3D unsupervised domain adaptation (UDA) tackles the lack of annotations in a new domain by capitalizing the relationship between 2D and 3D data. Existing methods achieve considerable improvements by performing cross-modality alignment in a modality-agnostic way, failing to exploit modality-specific characteristic for modeling complementarity. In this paper, we present self-supervised exclusive learning for cross-modal semantic segmentation under the UDA scenario, which avoids the prohibitive annotation. Specifically, two self-supervised tasks are designed, named "plane-to-spatial'' and "discrete-to-textured''. The former helps the 2D network branch improve the perception of spatial metrics, and the latter supplements structured texture information for the 3D network branch. In this way, modality-specific exclusive information can be effectively learned, and the complementarity of multi-modality is strengthened, resulting in a robust network to different domains. With the help of the self-supervised tasks supervision, we introduce a mixed domain to enhance the perception of the target domain by mixing the patches of the source and target domain samples. Besides, we propose a domain-category adversarial learning with category-wise discriminators by constructing the category prototypes for learning domain-invariant features. We evaluate our method on various multi-modality domain adaptation settings, where our results significantly outperform both uni-modality and multi-modality state-of-the-art competitors.
Yachao Zhang 0001, Miaoyu Li, Yuan Xie 0006, Cuihua Li, Cong Wang 0039, Zhizhong Zhang 0001, Yanyun Qu
ACM Multimedia4
2022 JSL3d: Joint subspace learning with implicit structure supervision for 3D pose estimation
Mengxi Jiang, Shihao Zhou 0003, Cuihua Li
Pattern Recognit.3
2021 Weakly Supervised Semantic Segmentation for Large-Scale Point Cloud
abstract
Existing methods for large-scale point cloud semantic segmentation require expensive, tedious and error-prone manual point-wise annotation. Intuitively, weakly supervised training is a direct solution to reduce the labeling costs. However, for weakly supervised large-scale point cloud semantic segmentation, too few annotations will inevitably lead to ineffective learning of network. We propose an effective weakly supervised method containing two components to solve the above problem. Firstly, we construct a pretext task, \textit{i.e.,} point cloud colorization, with a self-supervised training manner to transfer the learned prior knowledge from a large amount of unlabeled point cloud to a weakly supervised network. In this way, the representation capability of the weakly supervised network can be improved by knowledge from a heterogeneous task. Besides, to generative pseudo label for unlabeled data, a sparse label propagation mechanism is proposed with the help of generated class prototypes, which is used to measure the classification confidence of unlabeled point. Our method is evaluated on large-scale point cloud datasets with different scenarios including indoor and outdoor. The experimental results show the large gain against existing weakly supervised methods and comparable results to fully supervised methods.
Yachao Zhang 0001, Yuan Xie 0006, Yanyun Qu, Cuihua Li, Tao Mei 0001
AAAI5
2021 Perturbed Self-Distillation: Weakly Supervised Large-Scale Point Cloud Semantic Segmentation
abstract
Large-scale point cloud semantic segmentation has wide applications. Current popular researches mainly focus on fully supervised learning which demands expensive and tedious manual point-wise annotation. Weakly supervised learning is an alternative way to avoid this exhausting an-notation. However, for large-scale point clouds with few labeled points, the network is difficult to extract discriminative features for unlabeled points, as well as the regularization of topology between labeled and unlabeled points is usually ignored, resulting in incorrect segmentation results.To address this problem, we propose a perturbed self-distillation (PSD) framework. Specifically, inspired by self-supervised learning, we construct the perturbed branch and enforce the predictive consistency among the perturbed branch and original branch. In this way, the graph topology of the whole point cloud can be effectively established by the introduced auxiliary supervision, such that the in-formation propagation between the labeled and unlabeled points will be realized. Besides point-level supervision, we present a well-integrated context-aware module to explicitly regularize the affinity correlation of labeled points. Therefore, the graph topology of the point cloud can be further refined. The experimental results evaluated on three large-scale datasets show the large gain (3.0% on average) against recent weakly supervised methods and comparable results to some fully supervised methods.
Yachao Zhang 0001, Yanyun Qu, Yuan Xie 0006, Zonghao Li, Shanshan Zheng, Cuihua Li
ICCV6
2021 KeypointNet: Ranking Point Cloud for Convolution Neural Network
Zuodong Gao, Yanyun Qu, Cuihua Li
ICIG (3)3
2021 SDM3d: shape decomposition of multiple geometric priors for 3D pose estimation
Mengxi Jiang, Zhu Liang Yu, Cuihua Li
Neural Comput. Appl.3
2021 Joint Deep Multi-View Learning for Image Clustering
abstract
In this paper, a novelDeepMulti-viewJointClustering (DMJC) framework is proposed, where multiple deep embedded features, multi-view fusion mechanism, and clustering assignments can be learned simultaneously. Through the joint learning strategy, the clustering-friendly multi-view features and useful multi-view complementary information can be exploited effectively to improve the clustering performance. Under the proposed joint learning framework, we design two ingenious variants of deep multi-view joint clustering models, whose multi-view fusion is implemented by two kinds of simple yet effective schemes. The first model, called DMJC-S, performs multi-view fusion in an implicit way via a novel multi-view soft assignment distribution. The second model, termed DMJC-T, defines a novel multi-view auxiliary target distribution to conduct the multi-view fusion explicitly. Both DMJC-S and DMJC-T are optimized under a KL divergence objective. Experiments on eight challenging image datasets demonstrate the superiority of both DMJC-S and DMJC-T over single/multi-view baselines and the state-of-the-art multi-view clustering methods, which proves the effectiveness of the proposed DMJC framework. To the best of our knowledge, this is the first work to model the multi-view clustering in a deep joint framework, which will provide a meaningful thinking in unsupervised multi-view learning.
Yuan Xie 0006, Bingqian Lin, Yanyun Qu, Cuihua Li, Wensheng Zhang 0002, Lizhuang Ma, Yonggang Wen 0001, Dacheng Tao
IEEE Trans. Knowl. Data Eng.4
2020 Patch Proposal Network for Fast Semantic Segmentation of High-Resolution Images
abstract
Despite recent progress on the segmentation of high-resolution images, there exist an unsolved problem, i.e., the trade-off among the segmentation accuracy, memory resources and inference speed. So far, GLNet is introduced for high or ultra-resolution image segmentation, which has reduced the computational memory of the segmentation network. However, it ignores the importances of different cropped patches, and treats tiled patches equally for fusion with the whole image, resulting in high computational cost. To solve this problem, we introduce a patch proposal network (PPN) in this paper, which adaptively distinguishes the critical patches from the trivial ones to fuse with the whole image for refining segmentation. PPN is a classification network which alleviates network training burden and improves segmentation accuracy. We further embed PPN in a global-local segmentation network, instructing global branch and refinement branch to work collaboratively. We implement our method on four image datasets:DeepGlobe, ISIC, CRAG and Cityscapes, the first two are ultra-resolution image datasets and the last two are high-resolution image datasets. The experimental results show that our method achieves almost the best segmentation performance compared with the state-of-the-art segmentation methods and the inference speed is 12.9 fps on DeepGlobe and 10 fps on ISIC. Moreover, we embed PPN with the general semantic segmentation network and the experimental results on Cityscapes which contains more object classes demonstrate the generalization ability on general semantic segmentation.
Zhenzhen Lei, Bingqian Lin, Cuihua Li, Yanyun Qu, Yuan Xie 0006
AAAI4
2020 Distilling Image Dehazing With Heterogeneous Task Imitation
abstract
State-of-the-art deep dehazing models are often difficult in training. Knowledge distillation paves a way to train a student network assisted by a teacher network. However, most knowledge distill methods are used for image classification and segmentation as well as object detection, and few investigate distilling image restoration and use different task for knowledge transfer. In this paper, we propose a knowledge-distill dehazing network which distills image dehazing with the heterogeneous task imitation. In our network, the teacher is an off-the-shelf auto-encoder network and is used for image reconstruction. The dehazing network is trained assisted by the teacher network with the process-oriented learning mechanism. The student network imitates the task of image reconstruction in the teacher network. Moreover, we design a spatial-weighted channel-attention residual block for the student image dehazing network to adaptively learn the content-aware channel level attention and pay more attention to the features for dense hazy regions reconstruction. To evaluate the effectiveness of the proposed method, we compare our method with several state-of-the-art methods on two synthetic and real-world datasets, as well as real hazy images.
Ming Hong, Yuan Xie 0006, Cuihua Li, Yanyun Qu
CVPR3
2020 LatticeNet: Towards Lightweight Image Super-Resolution with Lattice Block
Xiaotong Luo, Yuan Xie 0006, Yulun Zhang 0001, Yanyun Qu, Cuihua Li, Yun Fu 0001
ECCV (22)5
2020 Single-image super-resolution via joint statistic models-guided deep auto-encoder network
Yanyun Qu, Cuihua Li, Yuan Xie 0006, Ce Li 0001
Neural Comput. Appl.3
2020 Scale robust deep oriented-text detection network
Yuqiang Zheng, Yuan Xie 0006, Yanyun Qu, Cuihua Li, Yan Zhang 0059
Pattern Recognit.5
2019 Joint-attention Discriminator for Accurate Super-resolution via Adversarial Training
abstract
Tremendous progress has been witnessed on single image super-resolution (SR), where existing deep SR models achieve impressive performance in objective criteria, e.g., PSNR and SSIM. However, most of the SR methods are limited in visual perception, for example, they look too smooth. Generative adversarial network (GAN) favors SR visual effects over most of the deep SR models but is poor in objective criteria. In order to trade off the objective and subjective SR performance, we design a joint-attention discriminator with which GAN improves the SR performance in PSNR and SSIM, as well as maintaining the visual effect compared with non-attention GAN based SR models. The joint-attention discriminator contains dense channel-wise attention and cross-layer attention blocks. The former is applied in the shallow layers of the discriminator for channel-wise weighting combination of feature maps. The latter is employed to select feature maps in some middle and deep layers for effective discrimination. Extensive experiments are conducted on six benchmark datasets and the experimental results show that our proposed discriminator combining with different generators can achieve more realistic visual performances.
Yuan Xie 0006, Xiaotong Luo, Yanyun Qu, Cuihua Li
ACM Multimedia5
2019 Reweighted sparse representation with residual compensation for 3D human pose estimation from a single RGB image
Mengxi Jiang, Zhu Liang Yu, Yan Zhang 0059, Qicong Wang, Cuihua Li
Neurocomputing5
2019 Ontology-driven hierarchical sparse coding for large-scale image classification
Yan Zhang 0059, Yanyun Qu, Cuihua Li, Jianping Fan 0001
Neurocomputing3
2019 Robust ℓ2-Hypergraph and its applications
Taisong Jin, Zhengtao Yu 0001, Yue Gao 0002, Shengxiang Gao, Xiaoshuai Sun, Cuihua Li
Inf. Sci.6
2018 Deeptongue: Tongue Segmentation Via Resnet
abstract
Accurate tongue image segmentation is helpful to acquire correct automatic tongue diagnosis result. However, traditional methods cannot bring satisfying results in most cases. This paper proposes an end-to-end trainable tongue image segmentation method using deep convolutional neural network based on ResNet. The proposed method, named DeepTongue, segments tongue by using a forward network without preprocessing. The proposed method has no restrictions of the illumination and size of tongue images. Experimental results show that the proposed DeepTongue improves the segmentation accuracy by a noticeable margin. In addition, DeepTongue is much faster than the existing tongue image segmentation methods.
Bingqian Lin, Junwei Xle, Cuihua Li, Yanyun Qu
ICASSP3
2018 Traffic-Sign Spotting in the Wild via Deep Features
abstract
This paper focuses on traffic sign spotting (TSS which automatically recognizes not only the conventional traffic signs but also information, facility and service signs, and traffic lights. TSS is divided into two sequential tasks: detecting traffic sign candidate regions in an image and recognizing the traffic signs in the regions. It is a very challenging task. We make the following contributions: 1) we create a traffic sign collection from the driverless car. The traffic signs are shot under the natural environment which covers large variation in illuminance and weather conditions. It not only contains the common traffic signs but also contains the information, facility and service signs which are called signposts, as well as traffic lights. 2) we proposed a systematic solution. We construct an Inception convolutional neural network. We use Faster-RCNN for traffic sign detection and make it suitable to detect small targets. 3) We adopt three schemes for the common traffic signs, the signposts and the traffic lights, respectively. The experimental results demonstrate the effectiveness and efficiency of our methods. Our methods won the first place in the traffic sign recognition task of Intelligent Vehicle Future Challenge 2017, China.
Jinkang Guo, Jianyun Lu, Yanyun Qu, Cuihua Li
Intelligent Vehicles Symposium4
2018 A Fast Forgery Detection Algorithm Based on Exponential-Fourier Moments for Video Region Duplication
abstract
Region duplication is one of the most common methods of video forgery. Existing forgery detection algorithms generally suffer from inefficiency and are not effective for the forged regions with mirroring. To address these problems, we present a fast forgery detection algorithm based on Exponential-Fourier moments (EFMs) for detecting region duplication in videos. The algorithm first extracts EFMs features from each block in the current frame, and performs a fast match to find potential matching pairs. Then, a postverification scheme is designed to eliminate falsely matched pairs and locate the altered regions in the current frame. Finally, an adaptive parameter-based fast compression tracking algorithm is used to track the tampered regions in the subsequent frames. The experimental results show that our proposed algorithm has higher detection accuracy and computational efficiency than those of previous algorithms.
Lichao Su, Cuihua Li, Yuecong Lai, Jianmei Yang
IEEE Trans. Multim.2
2017 Coupled Deep Autoencoder for Single Image Super-Resolution
abstract
Sparse coding has been widely applied to learning-based single image super-resolution (SR) and has obtained promising performance by jointly learning effective representations for low-resolution (LR) and high-resolution (HR) image patch pairs. However, the resulting HR images often suffer from ringing, jaggy, and blurring artifacts due to the strong yet ad hoc assumptions that the LR image patch representation is equal to, is linear with, lies on a manifold similar to, or has the same support set as the corresponding HR image patch representation. Motivated by the success of deep learning, we develop a data-driven model coupled deep autoencoder (CDA) for single image SR. CDA is based on a new deep architecture and has high representational capability. CDA simultaneously learns the intrinsic representations of LR and HR image patches and a big-data-driven function that precisely maps these LR representations to their corresponding HR representations. Extensive experimentation demonstrates the superior effectiveness and efficiency of CDA for single image SR compared to other state-of-the-art methods on Set5 and Set14 datasets.
Jun Yu 0002, Ruxin Wang 0002, Cuihua Li, Dacheng Tao
IEEE Trans. Cybern.4
2016 Layered multiple description video coding using dual-tree discrete wavelet transform and H.264/AVC
Jing Chen 0001, Canhui Cai, Cuihua Li
Multim. Tools Appl.4
2016 Image super-resolution base on multi-kernel regression
Yanyun Qu, Cuihua Li, Yuan Xie 0006
Multim. Tools Appl.3
2015 Multiple graph regularized sparse coding and multiple hypergraph regularized sparse coding for image representation
Taisong Jin, Zhengtao Yu 0001, Cuihua Li
Neurocomputing4
2015 Learning local Gaussian process regression for image super-resolution
Yanyun Qu, Cuihua Li, Yuan Xie 0006, Yang Wu 0001, Jianping Fan 0001
Neurocomputing3
2015 Low-rank matrix factorization with multiple Hypergraph regularizer
Taisong Jin, Jun Yu 0002, Jane You, Cuihua Li, Zhengtao Yu 0001
Pattern Recognit.5
2015 Novel Graph Cuts Method for Multi-Frame Super-Resolution
abstract
In this letter, we propose a new graph cuts multi-frame super resolution method. The method is carried out in 3 steps. First, we project each high-resolution pixel p onto the low-resolution images and select low-resolution pixels which fall within the zone of influence of p. Second, we weigh the contribution of the low-resolution pixels via a soft switching function and add them to construct a virtual low resolution pixel. The high resolution image is then recovered after minimizing a Maximum a posteriori Markov Random Field (MAP-MRF) energy function. This is done by approximating our energy function to make it graph representable and minimize it with a graph cuts α-expansion algorithm. Experimental results show that our approach outperforms state-of-the-art methods.
Dongxiao Zhang, Pierre-Marc Jodoin, Cuihua Li, Yun-Dong Wu, Guo-Rong Cai
IEEE Signal Process. Lett.3
2014 Online co-training ranking SVM for visual tracking
abstract
Online learned tracking is widely used to handle the appearance changes of object because of its adaptive ability. Learning to rank technique has attracted much attention recently in visual tracking. But the tracking method with online learning to rank suffers from the error accumulation problem during the self-training process. To solve this problem, we propose an online learning to rank algorithm in the co-training framework for robust visual tracking. A co-training algorithm combined with ranking SVM collects features and unlabeled data for training. Two ranking SVMs are built with different types of features accordingly and dynamically fused into a semi-supervised learning process. This semi-supervised learning approach is updated online to resist the occlusion and adapt to the changes of object's appearance. Many experiments on challenging sequences have shown that the proposed algorithm is more effective than the state-of-the-art methods.
Pingyang Dai, Yi Xie 0004, Cuihua Li
ICASSP4
2014 Image clustering by hyper-graph regularized non-negative matrix factorization
Jun Yu 0002, Cuihua Li, Jane You, Taisong Jin
Neurocomputing3
2014 Discriminative Object Tracking via Sparse Representation and Online Dictionary Learning
abstract
We propose a robust tracking algorithm based on local sparse coding with discriminative dictionary learning and new keypoint matching schema. This algorithm consists of two parts: the local sparse coding with online updated discriminative dictionary for tracking (SOD part), and the keypoint matching refinement for enhancing the tracking performance (KP part). In the SOD part, the local image patches of the target object and background are represented by their sparse codes using an over-complete discriminative dictionary. Such discriminative dictionary, which encodes the information of both the foreground and the background, may provide more discriminative power. Furthermore, in order to adapt the dictionary to the variation of the foreground and background during the tracking, an online learning method is employed to update the dictionary. The KP part utilizes refined keypoint matching schema to improve the performance of the SOD. With the help of sparse representation and online updated discriminative dictionary, the KP part are more robust than the traditional method to reject the incorrect matches and eliminate the outliers. The proposed method is embedded into a Bayesian inference framework for visual tracking. Experimental results on several challenging video sequences demonstrate the effectiveness and robustness of our approach.
Yuan Xie 0006, Wensheng Zhang 0002, Cuihua Li, Shuyang Lin, Yanyun Qu
IEEE Trans. Cybern.3
2013 Robust visual tracking via part-based sparsity model
abstract
The sparse representation has been widely used in many areas including visual tracking. The part-based representation performs outstandingly by using non-holistic templates to against occlusion. This paper combined them and proposed a robust object tracking method using part-based sparsity model for tracking an object in a video sequence. In the proposed model, one object is represented by image patches. The candidates of these patches are sparsely represented in the space which is spanned by the patch templates and trivial templates. The part-based method takes the spatial information of each patch into consideration, where the vote maps of multiple patches are used. Furthermore, the update scheme keeps the representative templates of each part dynamically. Therefore, trackers can effectively deal with the changes of appearances and heavy occlusion. On various public benchmark videos, the abundant results of experiments demonstrate that the proposed tracking method outperforms many existing state-of-the-arts algorithms.
Pingyang Dai, Yanlong Luo, Weisheng Liu, Cuihua Li, Yi Xie 0004
ICASSP4
2013 Visual Tracking Based on Compressive Sensing MCMC Sampling
abstract
Real time visual tracking is a challenge problem in computer vision. In this paper, we propose a real-time tracking method based on compressive sensing Markov Chain Monte Carlo (MCMC) sampling. To extract the features of objects, non-adaptive random projections are employed in the object appearance model which adopts a very sparse random measurement matrix using compress sensing. These projection preserve the structure of objects in the image feature space. A Bayesian classifier is learnt from the object appearance model and the scores of this classifier are integrated into Markov Chain Monte Carlo acceptance mechanism. Furthermore, a two-stage tracking scheme is used to alleviate the drift problem. The experimental results demonstrate that the proposed method is real time and outperforms some start-of-the-art algorithms on public benchmark sequences in terms of accuracy and robustness.
Pingyang Dai, Yanlong Luo, Cuihua Li, Yi Xie 0004
SMC4
2013 An effective unconstrained correlation filter and its kernelization for face recognition
Yan Yan 0001, Hanzi Wang, Cuihua Li, Chenhui Yang, Bineng Zhong 0001
Neurocomputing3
2012 Image super-resolution based on multikernel regression
Yanyun Qu, Tian-Zhu Fang, Cuihua Li, Hanzi Wang
ICPR4
2012 Online multiple instance gradient feature selection for robust visual tracking
Yuan Xie 0006, Yanyun Qu, Cuihua Li, Wensheng Zhang 0002
Pattern Recognit. Lett.3
2012 Image classification by multimodal subspace learning
Jun Yu 0002, Feng Lin 0002, Seah Hock Soon, Cuihua Li, Ziyu Lin
Pattern Recognit. Lett.4
2011 Research on License Plate Detection Based on Salient Feature under Complex Background
abstract
For plate detection, we found that when the plate is disturbed by complex upright borderlines, location can be inaccurate or even missed. On the basis of this, a new vehicle license plate location algorithm based on salient feature is introduced. It makes use of the texture feature, geometric characteristics and color information. By raw location and precise location, it can quickly locate the license plate position accurately and distinguishes the color type. Our experiment results show that this algorithm has a fast, efficient performance of locating vehicle license plate under complex background.
Qun-Wei Yang, Chenhui Yang, Cuihua Li, Hanzi Wang
ICIG3
2010 A fast electronic components orientation and identify method via radon transform
abstract
This paper presents a method which combined radon transform with machine learning technology for electronic component orientation and identification in product line scenes. This method can fetch electronic components' positions and yawing angles and enables the full automation of electronic product line. Firstly, it take images contain a single electronic component as training samples and retrieve its features. Secondly, it uses thresholds to segment objects in overlapping status. Finally, it use radon transform to detect the axis of object and then according to the component features acquired from training sample and classifier, the algorithm can identify electronic component pin's orientation. To increase the method's detection accuracy and speed in factory product line environment, this paper also proposed a strategy for the combination of the method with mechanism. Experiments show that this method has a perfect performance and completely fulfills the requirements of factory product line environment. This method achieves a recall rate of 81.7% and precision rate of 95.1%, after combined the algorithm with mechanism, the precision rate enhance to 98.5% and detection speed lifting strikingly.
Shuyang Lin, Shengrui Li, Cuihua Li
SMC3
2009 Symmetry Detection for Multi-object Using Local Polar Coordinate
Yuanhao Gong, Qicong Wang, Chenhui Yang, Yahui Gao, Cuihua Li
CAIP5
2009 Object Tracking Based on the Combination of Learning and Cascade Particle Filter
abstract
The problem of object tracking in dense clutter is a challenge in computer vision. This paper proposes a method for tracking object robustly by combining the online selection of discriminative color features and the offline selection of discriminative Haar features. Furthermore, the cascade particle filter which has four stages of importance sampling is used to fuse two kinds of features efficiently. When the illumination changes dramatically, the Haar features selected offline play a major role. When the object is occluded, or its rotation angle is very large, the color features selected online play a major role. The experimental results show that the proposed method performs well under the conditions of illumination change, occlusion, object scale change and abrupt motion of object or camera.
Hanjie Gong, Cuihua Li, Pingyang Dai, Yi Xie 0004
SMC2
2004 Sequential updating algorithm for extracting the basis of karhunen loeve transformation
Yanyun Qu, Nanning Zheng 0001, Cuihua Li, Zejian Yuan
ICIP3