EDBT 2026 Demo / reviewers in the wild / expert
Chunna Tian
dblp:23/3700
· DBLP profile ↗
45ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0002-3217-0368ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CRPE-Net: Infrared small target detection transformer with cross-layer relative-position embedding
Shiguo Chen, Chunna Tian, Yuan Yue |
Pattern Recognit. | 3 |
| 2026 | Exploring Self-Image and Cross-Image Consistency Learning for Remote Sensing Burned Area SegmentationabstractThe increasing frequency of global wildfires has led to the destruction of vast forests and wetlands. Non-contact remote sensing technologies provide an effective means for accurate burned area segmentation (BAS). However, existing BAS methods often treat each image independently, focusing primarily on local pixel contexts while neglecting the broader semantic consistency of burned regions across different scenes. The lack of global context modeling limits their robustness, as burned areas typically exhibit distinctive and consistent visual characteristics such as color and texture across diverse environments. To address this limitation, we propose a Self-image and Cross-image Consistency Learning (SCCL) framework, which captures both local pixel-level relationships within a single image and global semantic dependencies across multiple images. By enforcing consistent and compact representations of burned regions within and across images, SCCL enhances segmentation robustness under varying weather and terrain conditions. Additionally, to refine boundary delineation between burned and unburned areas, we introduce a Burned Edge Injector (BEI) and an Edge-Injected Decoder (EID). We further construct two large-scale BAS benchmark datasets, BAS-AUS and BAS-EUR, for comprehensive evaluation. Experiments on these benchmarks demonstrate that our method achieves state-of-the-art performance, significantly outperforming previous approaches, with MAE reduced to 0.017 and 0.016, respectively. The new BAS benchmarks and code are available at https://github.com/VisionVerse/SCCL. Heng Zhou 0006, Chengyang Li 0001, Chunna Tian, Yongqiang Xie, Zhongbo Li, Xiaojun Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Frequency-space enhanced and temporal adaptative RGBT object tracking
Shiguo Chen, Linzhi Xu, Chunna Tian |
Neurocomputing | 4 |
| 2025 | RGBT Fusion-Based Airborne Synthetic Aperture Occlusion Removal Target ImagingabstractDrones are gaining increasing popularity in bush search and rescue operations, where enhancing the accuracy of occluded target detection is crucial to finding the targets. However, airborne LiDAR and multispectral cameras are limited by cost, efficiency, and imaging quality, while RGB- or thermal-based detection methods rely heavily on model performance and prior knowledge. Recovering the target in imaging is a new frontier for occluded target perception. Here, we propose a novel Synthetic Aperture Fusion Imaging (SAFI) method for drones, combining an RGBT fusion network with synthetic aperture imaging to recover occluded targets. Firstly, we propose a Color-contrast RGBT fusion Network (C-RGBTNet) based on transformer. C-RGBTNet captures target thermal features from thermal images and assigns RGB color features to occlusions, generating high-quality fusion images. Then, we apply HSV-based color segmentation to isolate the target with strong thermal intensity (high-brightness) from the cluttered background. Finally, we perform Synthetic Aperture Imaging (SAI), synthesizing the different segmented target parts to yield an occlusion-free target image. Experimental results in our home-brew Downward-view Multimodal Bush (DMB) dataset demonstrate that SAFI achieves high-quality target imaging with effective occlusion removal, significantly outperforming traditional methods. Yuan Yue, Chunna Tian, Shiguo Chen, Dongliang Hou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Deformation-Resilient Multigranularity Learning for Unaligned RGB-T Semantic SegmentationabstractRGB-Thermal semantic segmentation (SS) aims to combine visual light and thermal images to determine the semantic category for each pixel and create an object mask. While existing methods typically rely on well-aligned RGB-T image pairs, real-world RGB-T pairs are often unaligned, and pixel-by-pixel alignment is both challenging and time-consuming. To address this critical issue, we introduce a new unaligned RGB-T SS benchmark and propose the deformation-resilient multigranularity learning (DML) method. DML explores the spatial consistency and modal complementarity of RGB-T and mitigates the interference of warped modalities by aligning multimodal features in a coarse-to-fine multigranularity strategy. Specifically, DML constructs a deformation-aware complementary feature enhancer (DCFE), which consists of deformation-aware feature alignment (DFA) and complementary feature aggregation (CFA) modules. DFA enhances the spatial alignment of RGB-T by estimating the deformation field of warped features. Then, CFA aggregates complementary contexts of modal differences across multiple scales to produce deformation-resilient and robust RGB-T feature representations. Finally, we design the multigranularity mask refinement engine (MMFE), which combines class-agnostic saliency prediction (CSP) and class-aware edge generation (CEG) auxiliary tasks to provide useful boundary and positional cues for SS decoders. The MMFE enhances semantic alignment and interclass separability, yielding object masks with sharp boundaries. Quantitative and qualitative experiments on aligned and unaligned datasets validate the effectiveness of our proposed DML, consistently outperforming existing methods designed for aligned RGB-T data. The new unaligned RGB-T SS benchmark and code are available at https://github.com/VisionVerse/Unaligned-RGBT-Semantic-Segmentation. Heng Zhou 0006, Chengyang Li 0001, Chunna Tian, Yongqiang Xie, Zhongbo Li, Xiaojun Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Accurate SAR Aircraft Detection Algorithm Based on Feature EnhancementabstractAircraft detection in synthetic aperture radar (SAR) images is of much significance because of its all-weather, all-day, and strong penetrating characteristics. However, existing algorithms exhibit inadequate capacity for feature extraction due to imaging discontinuities and background interference of SAR images. To overcome these task-specific issues, we proposed a feature enhancement-based SAR aircraft detection algorithm. In detail, we employed Adaptive Contrast Enhancement (ACE) in the preprocessing stage to reduce noises, and then we embed Scatter Point Focused Module (SPFM) into network to enhance the feature extraction of aircraft scattering points. Furthermore, we devised Background Interference Suppression Module (BISM) to accentuate salient points and suppress non-essential pixels. Experimental results on the GaoFen-3 SAR aircraft dataset demonstrate the effectiveness of the proposed feature enhancement-based method. Yizun Wang, Lei Zhang 0054, Chen Ding 0002, Chunna Tian, Wei Wei 0008 |
IGARSS | 6 |
| 2024 | FDE-Net: A memory-efficiency densely connected network inspired from fractional-order differential equations for single image super-resolution
Xiao Zhang 0058, Lei Zhang 0054, Wei Wei 0008, Chunna Tian, Yanning Zhang 0001 |
Neurocomputing | 5 |
| 2024 | Frequency-aware feature aggregation network with dual-task consistency for RGB-T salient object detection
Heng Zhou 0006, Chunna Tian, Chengyang Li 0001, Yongqiang Xie, Zhongbo Li |
Pattern Recognit. | 2 |
| 2024 | An Implicit-Explicit Prototypical Alignment Framework for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning methods have been explored to mitigate the scarcity of pixel-level annotation in medical image segmentation tasks. Consistency learning, serving as a mainstream method in semi-supervised training, suffers from low efficiency and poor stability due to inaccurate supervision and insufficient feature representation. Prototypical learning is one potential and plausible way to handle this problem due to the nature of feature aggregation in prototype calculation. However, the previous works have not fully studied how to enhance the supervision quality and feature representation using prototypical learning under the semi-supervised condition. To address this issue, we propose an implicit-explicit alignment (IEPAlign) framework to foster semi-supervised consistency training. In specific, we develop an implicit prototype alignment method based on dynamic multiple prototypes on-the-fly. And then, we design a multiple prediction voting strategy for reliable unlabeled mask generation and prototype calculation to improve the supervision quality. Afterward, to boost the intra-class consistency and inter-class separability of pixel-wise features in semi-supervised segmentation, we construct a region-aware hierarchical prototype alignment, which transmits information from labeled to unlabeled and from certain regions to uncertain regions. We evaluate IEPAlign on three medical image segmentation tasks. The extensive experimental results demonstrate that the proposed method outperforms other popular semi-supervised segmentation methods and achieves comparable performance with fully-supervised training methods. Chunna Tian, Xinbo Gao 0001, Heng Zhou 0006, Zhicheng Jiao |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | M2FNet: Mask-Guided Multi-Level Fusion for RGB-T Pedestrian DetectionabstractRGB-Thermal pedestrian detection has shown many notable advantages in various lighting and weather conditions by combining the information from RGB-T images. Due to distinct imaging principles, RGB-T modalities consist of modality-specific and modality-consistent information. However, most existing RGB-T pedestrian detection methods indiscriminately integrate these two types of information, which leads to the pollution of modality information. To address this issue, we propose a novel mask-guided multi-level fusion network (M2FNet) for RGB-T pedestrian detection. M2FNet independently explores consistent and specific features in RGB-T modalities at three different levels, utilizing pixel-level positional information in masks to exclusively focus on pedestrian-related features. Specifically, at the feature extraction level, we selectively embed cross-modality differential compensation (CDC) modules and design the bidirectional multiscale fusion (BMF) module to fully utilize the complementary modality-specific information and enhance the precision of predicted pedestrian masks. At the feature fusion level, the mask-guided global consistency mining (MGCM) module is introduced to capture intra-modal and inter-modal consistent information of pedestrians, which generates highly discriminative RGB-T features. Finally, to further reduce inter-modal differences, we propose a mask-guided pixel-level decision fusion (MPDF) strategy to dynamically weight the RGB-T predictions. Extensive experiments and comparisons demonstrate that our proposed M2FNet, with different backbones, outperforms the state-of-the-art detectors on both publicly available KAIST and CVC-14 RGB-T pedestrian detection datasets. Shiguo Chen, Chunna Tian, Heng Zhou 0006 |
IEEE Trans. Multim. | 3 |
| 2024 | An efficient and universal polygon prediction method based on derivable analytic geometry for arbitrary-shaped text detection
Xiangnan Zhang, Chunna Tian, Xinbo Gao 0001 |
Vis. Comput. | 2 |
| 2023 | Self-aware and Cross-Sample Prototypical Learning for Semi-supervised Medical Image Segmentation
Chunna Tian, Heng Zhou 0006, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao |
MICCAI (2) | 3 |
| 2023 | The CLIP Model is Secretly an Image-to-Prompt ConverterabstractThe Stable Diffusion model is a prominent text-to-image generation model that relies on a text prompt as its input, which is encoded using the Contrastive Language-Image Pre-Training (CLIP). However, text prompts have limitations when it comes to incorporating implicit information from reference images. Existing methods have attempted to address this limitation by employing expensive training procedures involving millions of training samples for image-to-image generation. In contrast, this paper demonstrates that the CLIP model, as utilized in Stable Diffusion, inherently possesses the ability to instantaneously convert images into text prompts. Such an image-to-prompt conversion can be achieved by utilizing a linear projection matrix that is calculated in a closed form. Moreover, the paper showcases that this capability can be further enhanced by either utilizing a small amount of similar-domain training data (approximately 100 images) or incorporating several online training steps (around 30 iterations) on the reference images. By leveraging these approaches, the proposed method offers a simple and flexible solution to bridge the gap between images and text prompts. This methodology can be applied to various tasks such as image variation and image editing, facilitating more effective and seamless interaction between images and textual prompts. Chunna Tian, Haoxuan Ding, Lingqiao Liu |
NeurIPS | 2 |
| 2023 | Balanced image captioning with task-aware decoupled learning and fusion
Lingqiao Liu, Chunna Tian, Xiangnan Zhang, Xilan Tian |
Neurocomputing | 3 |
| 2023 | Model-driven self-aware self-training framework for label noise-tolerant medical image segmentation
Chunna Tian, Xinbo Gao 0001, Yanyu Ye, Heng Zhou 0006, Zhuo Tong |
Signal Process. | 2 |
| 2023 | Position-Aware Relation Learning for RGB-Thermal Salient Object DetectionabstractSalient object detection (SOD) is an important task in computer vision that aims to identify visually conspicuous regions in images. RGB-Thermal SOD combines two spectra to achieve better segmentation results. However, most existing methods for RGB-T SOD use boundary maps to learn sharp boundaries, which lead to sub-optimal performance as they ignore the interactions between isolated boundary pixels and other confident pixels. To address this issue, we propose a novel position-aware relation learning network (PRLNet) for RGB-T SOD. PRLNet explores the distance and direction relationships between pixels by designing an auxiliary task and optimizing the feature structure to strengthen intra-class compactness and inter-class separation. Our method consists of two main components: A signed distance map auxiliary module (SDMAM), and a feature refinement approach with direction field (FRDF). SDMAM improves the encoder feature representation by considering the distance relationship between foreground-background pixels and boundaries, which increases the inter-class separation between foreground and background features. FRDF rectifies the features of boundary neighborhoods by exploiting the features inside salient objects. It utilizes the direction relationship of object pixels to enhance the intra-class compactness of salient features. In addition, we constitute a transformer-based decoder to decode multispectral feature representation. Experimental results on three public RGB-T SOD datasets demonstrate that our proposed method not only outperforms the state-of-the-art methods, but also can be integrated with different backbone networks in a plug-and-play manner. Ablation study and visualizations further prove the validity and interpretability of our method. Heng Zhou 0006, Chunna Tian, Chengyang Li 0001, Yongqiang Xie, Zhongbo Li |
IEEE Trans. Image Process. | 2 |
| 2022 | Dynamic prototypical feature representation learning framework for semi-supervised skin lesion segmentation
Chunna Tian, Xinbo Gao 0001, Xue Feng 0001, Harrison X. Bai, Zhicheng Jiao |
Neurocomputing | 2 |
| 2022 | Multispectral Fusion Transformer Network for RGB-Thermal Urban Scene Semantic SegmentationabstractSemantic segmentation plays a vital role in autonomous vehicles. Fusing the rich details of RGB image and the illumination robustness of thermal image has great potential to improve the performance of RGB-T semantic segmentation. In multispectral feature fusion, the current main methods are less effective in the characterization of correlations and complementarities of RGB-T. In order to generate robust cross-spectral fusion features, we propose a multispectral fusion transformer network (MFTNet). Specifically, we first design an MFT module to handle the intraspectra correlation and the interspectra complementarity of RGB-T in the multispectral fusion encoder. MFT effectively enhances the RGB-T feature representation under various challenges. Then, an optimization strategy with progressive deep supervision (PDS) loss is proposed to directly supervise the upper and lower layers of the decoder. This strategy can guide the decoder to achieve precise segmentation in a coarse-to-fine manner. Finally, plenty of experimental results prove the effectiveness of our method. On the MFNet dataset, MFNet achieved 74.7 mAcc and 57.3 mIoU, outperforming the state-of-the-art methods. Heng Zhou 0006, Chunna Tian, Qizheng Huo, Yongqiang Xie, Zhongbo Li |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Discriminative error prediction network for semi-supervised colon gland segmentation
Chunna Tian, Harrison X. Bai, Zhicheng Jiao, Xilan Tian |
Medical Image Anal. | 2 |
| 2022 | Collaborative boundary-aware context encoding networks for error map prediction
Chunna Tian, Xinbo Gao 0001, Jie Li 0001, Zhicheng Jiao, Zhusi Zhong |
Pattern Recognit. | 2 |
| 2021 | Quality-driven deep active learning method for 3D brain MRI segmentation
Jie Li 0001, Chunna Tian, Zhusi Zhong, Zhicheng Jiao, Xinbo Gao 0001 |
Neurocomputing | 3 |
| 2019 | Deep Spectral Super-Resolution with Noisy InputabstractLearning based methods, e.g., sparse coding or deep convolutional neural networks (DCNNs) have underpinned much of recent progress in increasing the spectral resolution of an RGB image for hyperspectral image (HSI) super-resolution. However, these methods suffer severe performance loss, when the test RGB image distributed differently from the training set, e.g., being corrupted with random noise. To mitigate this problem, we propose an unsupervised deep spectral super-resolution method, which employs a DCNN to generate the latent HSI from an input RGB and encourages it to fit the input RGB image through down-sampling in spectral domain as well as a sparse gradient prior in spatial domain. Due to the powerful capacity of DCNN in capturing the low-level image statistics, the proposed method is able to automatically accommodate the noise corruption in the input RGB image. Experimental results shows the superior performance of the proposed method. Zhiqiang Lang, Lei Zhang 0054, Wei Wei 0008, Jiangtao Nie, Chunna Tian, Yanning Zhang 0001 |
IGARSS | 5 |
| 2019 | Jointing Cross-Modality Retrieval to Reweight Attributes for Image Caption Generation
Mengmeng Jiang, Donghu Deng, Wei Wei 0008, Chunna Tian |
PRCV (3) | 7 |
| 2019 | SS-GANs: Text-to-Image via Stage by Stage Generative Adversarial Networks
Ming Tian, Yuting Xue, Chunna Tian, Donghu Deng, Wei Wei 0008 |
PRCV (2) | 3 |
| 2019 | How much do cross-modal related semantics benefit image captioning by weighting attributes and re-ranking sentences?
Chunna Tian, Ming Tian, Mengmeng Jiang, Donghu Deng |
Pattern Recognit. Lett. | 1 |
| 2019 | Intracluster Structured Low-Rank Matrix Analysis Method for Hyperspectral DenoisingabstractHyperspectral images (HSIs) denoising aims at eliminating the noise generated during the acquisition and transmission of HSIs. Since denoising is an ill-posed problem, utilizing proper knowledge of HSIs as regularization is essential for a good denoiser. Many HSI denoising methods have been proposed to leverage various prior knowledge, e.g., total variation, sparsity, and so on. Among those knowledge, a low-rank property has been shown to be effective for HSI denoising since it has the ability to deal with the missing values. However, most existing low-rank methods seldom consider mining the useful structures inside the low-rank matrix for a better denoising result. In addition, the rank number needs to be assigned manually. To address these problems, we propose an intracluster structured low-rank matrix analysis method for HSI denoising. First, we divide the original HSI into some clusters by taking advantages of both local similarity and nonlocal similarity structures, with which the resulted clusters are simpler and show more obvious low-rank property. Second, with singular value decomposition on the low-rank matrix in each cluster, the structured sparsity is modeled among the singular values to capture the structure of the low-rank matrix. Finally, an efficient optimization method is proposed to learn the structured sparsity adaptively from the data, as well as to inversely estimate the latent clean HSI from the noisy counterpart. The proposed method can not only obtain better denoising results compared with the-state-of-the-art methods but also automatically determine the rank number. Extensive experimental results demonstrate the effectiveness of the proposed method. Wei Wei 0008, Lei Zhang 0054, Yining Jiao, Chunna Tian, Cong Wang 0013, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | A Novel Semantic Attribute-Based Feature for Image Caption GenerationabstractImage captioning is challenging because it connects computer vision and natural language processing. It requires not only sensing objects but also the interrelations and context in an image to generate natural language descriptions. In this paper, we propose to extract a novel visual feature weighted by salient semantic attributes, which is fed to the encoder of Long Short Term Memory (LSTM). Semantic attributes are important to exploit more semantic-related information in images and describe the salient scenes to enhance the accuracy of generating image captions. Based on the Multiple Instance Learning (MIL) architecture on VGG-16 network, we design transferring rules that map high probability attributes to the feature vector in fc7 layer. It results in more semantic-related visual features. Our model can recognize richer details of images effectively and achieve the state-of-the-art performance on MSCOCO 2014 dataset under standard metrics. Chunna Tian |
ICASSP | 3 |
| 2018 | Photo-realistic 2D expression transfer based on FFT and modified Poisson image editing
Chunna Tian, Xinbo Gao 0001 |
Neurocomputing | 1 |
| 2018 | Text detection in natural scene images based on color prior guided MSER
Xiangnan Zhang, Xinbo Gao 0001, Chunna Tian |
Neurocomputing | 3 |
| 2018 | Color pornographic image detection based on color-saliency preserved mixture deformable part model
Chunna Tian, Xiangnan Zhang, Wei Wei 0008, Xinbo Gao 0001 |
Multim. Tools Appl. | 1 |
| 2017 | Hyperspectral image super-resolution extending: An effective fusion based method without knowing the spatial transformation matrixabstractHyperspectral image (HSI) super-resolution, a technique to obtain higher (often spatial) resolution image from the original image, has been extensively studied and applied to lots of fields such as computer vision, remote sensing, etc. Though fusion based method has achieved state-of-the-art result, it always assume the spatial transformation matrix is given in advance, whereas such a matrix is actually unknown in reality. An unsuitable given matrix will deteriorate the superresolution result greatly. To address this issue, we propose a novel fusion based HSI super-resolution method without knowing the spatial transformation matrix. Specifically, we incorporate super-resolution and spatial transformation matrix estimation into a unified framework. We alternately estimate the matrix and the higher spatial resolution HSI. We find that without given the spatial transformation matrix, the proposed method can obtain more accurate reconstruction result compared with other competing methods. Experimental results demonstrate the effectiveness of the proposed method. Lei Zhang 0054, Chunna Tian, Chen Ding 0002, Yanning Zhang 0001, Wei Wei 0008 |
ICME | 3 |
| 2017 | Natural scene text detection with MC-MR candidate extraction and coarse-to-fine filtering
Chunna Tian, Xiangnan Zhang, Xinbo Gao 0001 |
Neurocomputing | 1 |
| 2017 | Structured Sparse Coding-Based Hyperspectral Imagery Denoising With Intracluster FilteringabstractSparse coding can exploit the intrinsic sparsity of hyperspectral images (HSIs) by representing it as a group of sparse codes. This strategy has been shown to be effective for HSI denoising. However, how to effectively exploit the structural information within the sparse codes (structured sparsity) has not been widely studied. In this paper, we propose a new method for HSI denoising, which uses structured sparse coding and intracluster filtering. First, due to the high spectral correlation, the HSI is represented as a group of sparse codes by projecting each spectral signature onto a given dictionary. Then, we cast the structured sparse coding into a covariance matrix estimation problem. A latent variable-based Bayesian framework is adopted to learn the covariance matrix, the sparse codes, and the noise level simultaneously from noisy observations. Although the considered strategy is able to perform denoising through accurately reconstructing spectral signatures, an inconsistent recovery of sparse codes may corrupt the spectral similarity in each spatial homogeneous cluster within the scene. To address this issue, an intracluster filtering scheme is further employed to restore the spectral similarity in each spatial cluster, which results in better denoising results. Our experimental results, conducted using both simulated and real HSIs, demonstrate that the proposed method outperforms several state-of-the-art denoising methods. Wei Wei 0008, Lei Zhang 0054, Chunna Tian, Antonio Plaza, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Facial expression transfer method based on frequency analysis
Wei Wei 0008, Chunna Tian, Stephen J. Maybank, Yanning Zhang 0001 |
Pattern Recognit. | 2 |
| 2016 | Exploring Structured Sparsity by a Reweighted Laplace Prior for Hyperspectral Compressive SensingabstractHyperspectral compressive sensing (HCS) can greatly reduce the enormous cost of hyperspectral images (HSIs) on imaging, storage, and transmission by only collecting a few compressive measurements in the image acquisition. One of the most challenging problems for HCS is how to reconstruct the HSI accurately from such a few measurements. It has been proved that introducing structure information into sparsity prior can improve the reconstruction performance of standard compressive sensing models. However, the structured sparsity of HSIs is unknown in reality and easily affected by random noise, which makes it difficult to explore the structured sparsity in HCS. To address this problem, we propose a novel reweighted Laplace prior-based HCS method in this paper. First, a hierarchical reweighted Laplace prior is proposed to model the distribution of sparsity in an HSI, which relieves the undemocratic penalization of traditional Laplace prior on nonzero coefficients of a sparse signal. Then, a latent variable-based Bayesian model is employed to learn the optimal configuration of the reweighted Laplace prior from the measurements. This model unifies signal recovery, sparsity prior learning, and noise estimation into a variational framework, where these three tasks are alternatively optimized till convergence. The finally learned sparsity prior can well represent the underlying structure in the sparse signal and is adaptive to the unknown noise. These advantages together improve the reconstruction accuracy of HCS obviously. Moreover, the proposed method is extended to learn a matrix normal distribution-based prior with a full covariance matrix, which depicts the underlying structure in the sparse signal better. As a result, the reconstruction accuracy is further improved. Extensive experimental results on three hyperspectral data sets demonstrate that the proposed method outperforms several state-of-the-art HCS methods in terms of the reconstruction accuracy. Lei Zhang 0054, Wei Wei 0008, Chunna Tian, Fei Li 0011, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Reweighted laplace prior based hyperspectral compressive sensing for unknown sparsityabstractCompressive sensing(CS) has been exploited for hype-spectral image(HSI) compression in recent years. Though it can greatly reduce the costs of computation and storage, the reconstruction of HSI from a few linear measurements is challenging. The underlying sparsity of HSI is crucial to improve the reconstruction accuracy. However, the sparsity of HSI is unknown in reality and varied with different noise, which makes the sparsity estimation difficult. To address this problem, a novel reweighted Laplace prior based hyperspectral compressive sensing method is proposed in this study. First, the reweighted Laplace prior is proposed to model the distribution of sparsity in HSI. Second, the latent variable Bayes model is employed to learn the optimal configuration of the reweighted Laplace prior from the measurements. The model unifies signal recovery, prior learning and noise estimation into a variational framework to infer the parameters automatically. The learned sparsity prior can represent the underlying structure of the sparse signal very well and is adaptive to the unknown noise, which improves the reconstruction accuracy of HSI. The experimental results on three hyperspectral datasets demonstrate the proposed method outperforms several state-of-the-art hyperspectral CS methods on the reconstruction accuracy. Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Chunna Tian, Fei Li 0011 |
CVPR | 4 |
| 2015 | Visual Tracking Based on the Adaptive Color Attention Tuned Sparse Generative Object ModelabstractThis paper presents a new visual tracking framework based on an adaptive color attention tuned local sparse model. The histograms of sparse coefficients of all patches in an object are pooled together according to their spatial distribution. A particle filter methodology is used as the location model to predict candidates for object verification during tracking. Since color is an important visual clue to distinguish objects from background, we calculate the color similarity between objects in the previous frames and the candidates in current frame, which is adopted as color attention to tune the local sparse representation-based appearance similarity measurement between the object template and candidates. The color similarity can be calculated efficiently with hash coded color names, which helps the tracker find more reliable objects during tracking. We use a flexible local sparse coding of the object to evaluate the degeneration degree of the appearance model, based on which we build a model updating mechanism to alleviate drifting caused by temporal varying multi-factors. Experiments on 76 challenging benchmark color sequences and the evaluation under the object tracking benchmark protocol demonstrate the superiority of the proposed tracker over the state-of-the-art methods in accuracy. Chunna Tian, Xinbo Gao 0001, Wei Wei 0008 |
IEEE Trans. Image Process. | 1 |
| 2014 | Robust face pose classification method based on geometry-preserving visual phraseabstractConstructing the discriminative feature for face pose classification is challenging. Since the key facial points are co-occurring with different spatial layout in different poses, we propose a pose classification framework based on the local geometry-preserving visual phrase (GVP). The weighting strategy on GVP enhances the discriminability of the single word and the spatial layout in the high order phrase simultaneously. Thus, the co-occurring words and local geometric structure of phrase are discriminative to distinguish pose, flexible and robust to the multi-factor variations. The experimental results on Oriental Face database and PIE database show the superiority of our method compared with the PCA, LDA based methods and the tf-idf weighted BoW method. Wei Wei 0008, Chunna Tian, Yanning Zhang 0001 |
ICIP | 2 |
| 2012 | Multiview Face Recognition: From TensorFace to V-TensorFace and K-TensorFaceabstractFace images under uncontrolled environments suffer from the changes of multiple factors such as camera view, illumination, expression, etc. Tensor analysis provides a way of analyzing the influence of different factors on facial variation. However, the TensorFace model creates a difficulty in representing the nonlinearity of view subspace. In this paper, to break this limitation, we present a view-manifold-based TensorFace (V-TensorFace), in which the latent view manifold preserves the local distances in the multiview face space. Moreover, a kernelized TensorFace (K-TensorFace) for multiview face recognition is proposed to preserve the structure of the latent manifold in the image space. Both methods provide a generative model that involves a continuous view manifold for unseen view representation. Most importantly, we propose a unified framework to generalize TensorFace, V-TensorFace, and K-TensorFace. Finally, an expectation-maximization like algorithm is developed to estimate the identity and view parameters iteratively for a face image of an unknown/unseen view. The experiment on the PIE database shows the effectiveness of the manifold construction method. Extensive comparison experiments on Weizmann and Oriental Face databases for multiview face recognition demonstrate the superiority of the proposed V- and K-TensorFace methods over the view-based principal component analysis and other state-of-the-art approaches for such purpose. Chunna Tian, Xinbo Gao 0001, Qi Tian 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2010 | Chinese text detection and location for images in multimedia messaging serviceabstractText detection and recognition for images in multimedia messaging service is a very important task. Since Chinese characters are composed of four kinds of strokes, i.e., horizontal line, top-down vertical line, left-downward slope line and short pausing stroke, we present Gabor filters with scale and direction varied to describe the strokes of Chinese characters for candidate text area extraction. By establishing four sub-neural networks to learn the texture of text area, the learnt classifiers are used to detect candidate text areas. Experimental results show that the proposed approach can improve the accuracy of text detection and the recognition rate of images in multimedia messaging service. Jianqiang Yan, Dacheng Tao, Chunna Tian, Xinbo Gao 0001, Xuelong Li 0001 |
SMC | 3 |
| 2009 | Multi-view face recognition based on tensor subspace analysis and view manifold modeling
Xinbo Gao 0001, Chunna Tian |
Neurocomputing | 2 |
| 2008 | Multi-view face recognition by nonlinear tensor decompositionabstractWe discuss a new multi-view face recognition method that extends a recently proposed nonlinear tensor decomposition technique. We use this technique to provide a generative face model that can deal with both the linearity and nonlinearity in multi-view face images. Particularly, we study the effectiveness of three kinds of view manifold for multi-view face representation, i.e., the concept-driven, data-driven and hybrid data-concept-driven view manifolds. An EM-like algorithm is developed to estimate the identity and view factors iteratively. The new face generative model can successfully recognize face images captured under unseen views, and the experimental results provide the new method is superior to the traditional TensorFace-based algorithm and the view-based PCA method. Chunna Tian, Xinbo Gao 0001 |
ICPR | 1 |
| 2008 | Face Sketch Synthesis Algorithm Based on E-HMM and Selective EnsembleabstractSketch synthesis plays an important role in face sketch-photo recognition system. In this manuscript, an automatic sketch synthesis algorithm is proposed based on embedded hidden Markov model (E-HMM) and selective ensemble strategy. First, the E-HMM is adopted to model the nonlinear relationship between a sketch and its corresponding photo. Then based on several learned models, a series of pseudo-sketches are generated for a given photo. Finally, these pseudo-sketches are fused together with selective ensemble strategy to synthesize a finer face pseudo-sketch. Experimental results illustrate that the proposed algorithm achieves satisfactory effect of sketch synthesis with a small set of face training samples. Xinbo Gao 0001, Juanjuan Zhong, Jie Li 0001, Chunna Tian |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2007 | Face Sketch Synthesis using E-HMM and Selective EnsembleabstractIn this manuscript, we propose an automatic sketch synthesis algorithm based on embedded hidden Markov model (E-HMM) and selective ensemble strategy. The E-HMM is used to model the nonlinear relationship between a photo-sketch pair firstly, and then a series of pseudo-sketches, which are generated based on several learned models for a given photo, are integrated together with selective ensemble strategy to synthesize a finer face pseudo-sketch. The experimental results illustrate that the proposed algorithm achieves satisfactory effect of sketch synthesis. Juanjuan Zhong, Xinbo Gao 0001, Chunna Tian |
ICASSP (1) | 3 |
| 2006 | A Valid Multi-View Face Detection Tree Based on Floatboost LearningabstractA novel face detection tree based on floatboost learning is proposed to accommodate the in-class variability of multi-view faces. The tree splitting procedure is realized through dividing face training examples into the optimal sub-clusters using the fuzzy c-means (FCM) algorithm together with a new cluster validity function based on the modified partition fuzzy degree. Then each sub-cluster of face examples is conquered with the floatboost learning to construct branches in the node of the detection tree. During training, the proposed algorithm is much faster than the original detection tree. The experimental results on the CMU and our home-brew test database illustrate that the proposed detection tree is more efficient than the original one while keeping its detection speed. Chunna Tian, Xinbo Gao 0001, Jie Li 0001 |
ICIP | 1 |