Yihua Tan

dblp:20/781 · DBLP profile ↗
← Back
47ranked-venue papers
10as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 8 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Deep multi-view clustering based on sample-level adaptive fusion and clustering structure aggregation
Zengbiao Yang, Yihua Tan
Neurocomputing2
2026 Tailoring knowledge for empowered cooperative actions in multi-agent reinforcement learning
Yihua Tan, Pengyi Li 0001
Neural Networks2
2026 Cross-modal transformer fusion via local sampling for drone RGB-infrared object detection
Herong Qi, Xuanyu Xiang, Hui Qin, Yuan Tai, Yihua Tan
Neural Networks5
2026 Deep multi-view clustering based on instance-level adaptive structural contrastive learning
Zengbiao Yang, Yihua Tan
Neural Networks2
2026 Deep Multi-View Clustering via Dynamic Anchors-Driven Global Structure Exploration
Zengbiao Yang, Yihua Tan
IEEE Trans. Knowl. Data Eng.2
2026 Improving Wi-Fi Cooperative Broadcast With Fine-Grained Channel Estimation
abstract
Cooperative broadcast is an efficient approach to improve Wi-Fi broadcast performance in crowded scenarios with densely deployed access points (APs). However, existing concurrent transmission MAC protocols cannot perfectly synchronize APs for the receiving user, and the superimposed channels at users vary over time due to multi-path effects with different carrier frequency offsets (CFOs) from the APs. Traditional channel estimation methods, which treat the superimposed channels as a whole and use a portion of the superimposed channels to derive the rest, are unsuitable. To solve the problem, we propose a fine-grained channel estimation approach that first estimates channel taps and CFOs of each AP, and then reconstructs the superimposed channels. We first study a benchmark channel estimation algorithm that utilizes a widely adopted compressed sensing (CS) technique. However, through analysis and simulations, we show that the CS-based algorithm suffers from high correlation problems in the constructed sensing matrix and the non-sparse channel problem in practice, leading to an estimation error floor at high SNRs. To solve these problems, we present a two-stage channel estimation algorithm. It first estimates the CFOs by identifying the most likely CFO combination matching the received signals, and then estimates the time-domain channel taps. Simulation and experimental results show that the two-stage channel estimation algorithm achieves much lower bit error rate (BER) and packet error rate (PER) than the traditional IEEE 802.11 approach, and the two-stage algorithm outperforms the CS-based algorithm, especially at high SNRs. The network-layer simulation results further demonstrate that, empowered by the proposed two-stage channel estimation algorithm, the cooperative broadcast scheme improves throughput by at least 1.4× (up to 46.8×) compared with the unicast-based broadcast schemes, and by approximately 0.6× to 0.8× compared with the simple uncooperative broadcast scheme.
Lizhao You, Shuoling Liu, Yihua Tan, Zhaorui Wang 0001, Soung Chang Liew
IEEE Trans. Mob. Comput.4
2025 High-confidence pseudo-label graph guided multi-view clustering structure discovery
Zengbiao Yang, Yihua Tan
Knowl. Based Syst.2
2025 A Global Spatial-Temporal Detection Framework for Infrared Small Targets in Complex Ground Scenes
Xuanyu Xiang, Yihua Tan
IEEE Trans. Geosci. Remote. Sens.3
2024 Dynamic Cues-Assisted Transformer for Robust Point Cloud Registration
abstract
Point Cloud Registration is a critical and challenging task in computer vision. Recent advancements have pre-dominantly embraced a coarse-to-fine matching mechanism, with the key to matching the superpoints located in patches with interframe consistent structures. How-ever, previous methods still face challenges with ambiguous matching, because the interference information aggregated from irrelevant regions may disturb the capture of interframe consistency relations, leading to wrong matches. To address this issue, we propose Dynamic Cues-Assisted Transformer (DCATr). Firstly, the interference from irrelevant regions is greatly reduced by constraining attention to certain cues, i.e., regions with highly correlated structures of potential corresponding superpoints. Secondly, cues-assisted attention is designed to mine the interframe consistency relations, while more attention is assigned to pairs with high consistent confidence in feature aggregation. Finally, a dynamic updating fashion is proposed to facilitate mining richer consistency information, further improving aggregated features' distinctiveness and relieving matching ambiguity. Extensive evaluations on indoor and outdoor standard benchmarks demonstrate that DCATr outperforms all state-of-the-art methods.
Hong Chen 0019, Pei Yan, Sihe Xiang, Yihua Tan
CVPR4
2024 MonoCD: Monocular 3D Object Detection with Complementary Depths
abstract
Monocular 3D object detection has attracted widespread attention due to its potential to accurately obtain object 3D localization from a single image at a low cost. Depth estimation is an essential but challenging subtask of monocular 3D object detection due to the ill-posedness of 2D to 3D mapping. Many methods explore multiple local depth clues such as object heights and keypoints and then formulate the object depth estimation as an ensemble of multiple depth predictions to mitigate the insufficiency of single-depth information. However, the errors of existing multiple depths tend to have the same sign, which hinders them from neutralizing each other and limits the overall accuracy of combined depth. To alleviate this problem, we propose to increase the complementarity of depths with two novel designs. First, we add a new depth prediction branch named complementary depth that utilizes global and efficient depth clues from the entire image rather than the local clues to reduce the similarity of depth predictions. Second, we propose to fully exploit the geometric relations between multiple depth clues to achieve complementarity in form. Benefiting from these designs, our method achieves higher complementarity. Experiments on the KITTI bench-mark demonstrate that our method achieves state-of-the-art performance without introducing extra data. In addition, complementary depth can also be a lightweight and plug-and-play module to boost multiple existing monocular 3d object detectors. Code is available at https://github.com/elvintanhust/MonoCD.
Longfei Yan 0003, Pei Yan, Shengzhou Xiong, Xuanyu Xiang, Yihua Tan
CVPR5
2024 Improving Cooperative Wi-Fi Broadcast with Fine-Grained Channel Estimation
abstract
Cooperative broadcast is an efficient approach to improve Wi-Fi broadcast performance in a crowded scenario with densely deployed access points (APs). However, the current concurrent transmission MAC protocols cannot synchronize multi-APs’ signals perfectly for all users. As a result, the superimposed signal from APs is time-varying at the users due to the multiple time-domain channels and carrier frequency offsets (CFOs) from multiple APs. The traditional channel estimation approach that estimates the superimposed channel as a whole is ill-suited for the superimposed signal. In this paper, we propose a fine-grained channel estimation approach to first estimate these channel parameters for each AP, and then reconstruct the superimposed channel. Specifically, we present a two-stage channel estimation algorithm that first estimates the CFOs by discretizing the CFO range and matching the most possible CFOs, and then computes the time-domain channels. Experiment and simulation results show the new channel estimation approach achieves much lower bit error rate (BER) and packet error rate (PER) than the traditional IEEE 802.11 approach. In addition, we propose a distributed mechanism to choose the master AP that initializes multi-APs’ simultaneous transmission, which the current concurrent transmission MAC protocols lack. Network-layer simulation results show that the proposed cooperative broadcast scheme improves the throughput by 64% to 82% compared with the traditional uncooperative broadcast scheme.
Lizhao You, Shuoling Liu, Wenjun Xie, Zhaorui Wang 0001, Yihua Tan, Soung Chang Liew
IWQoS5
2024 Global Structural Consistency Set Transformer
Zengbiao Yang, Yihua Tan
PRCV (2)2
2024 Where to model the epistemic uncertainty of Bayesian convolutional neural networks for classification
Yuan Tai, Yihua Tan, Erbo Zou
Neurocomputing2
2024 Learning feature relationships in CNN model via relational embedding convolution layer
Shengzhou Xiong, Yihua Tan, Guoyou Wang, Pei Yan, Xuanyu Xiang
Neural Networks2
2024 Learning to Holistically Detect Bridges From Large-Size VHR Remote Sensing Imagery
abstract
Bridge detection in remote sensing images (RSIs) plays a crucial role in various applications, but it poses unique challenges compared to the detection of other objects. In RSIs, bridges exhibit considerable variations in terms of their spatial scales and aspect ratios. Therefore, to ensure the visibility and integrity of bridges, it is essential to perform holistic bridge detection in large-size very-high-resolution (VHR) RSIs. However, the lack of datasets with large-size VHR RSIs limits the deep learning algorithms' performance on bridge detection. Due to the limitation of GPU memory in tackling large-size images, deep learning-based object detection methods commonly adopt the cropping strategy, which inevitably results in label fragmentation and discontinuous prediction. To ameliorate the scarcity of datasets, this paper proposes a large-scale dataset named GLH-Bridge comprising 6,000 VHR RSIs sampled from diverse geographic locations across the globe. These images encompass a wide range of sizes, varying from 2,048 × 2,048 to 16,384 × 16,384 pixels, and collectively feature 59,737 bridges. These bridges span diverse backgrounds, and each of them has been manually annotated, using both an oriented bounding box (OBB) and a horizontal bounding box (HBB). Furthermore, we present an efficient network for holistic bridge detection (HBD-Net) in large-size RSIs. The HBD-Net presents a separate detector-based feature fusion (SDFF) architecture and is optimized via a shape-sensitive sample re-weighting (SSRW) strategy. The SDFF architecture performs inter-layer feature fusion (IFF) to incorporate multi-scale context in the dynamic image pyramid (DIP) of the large-size image, and the SSRW strategy is employed to ensure an equitable balance in the regression weight of bridges with various aspect ratios. Based on the proposed GLH-Bridge dataset, we establish a bridge detection benchmark including the OBB and HBB tasks, and validate the effectiveness of the proposed HBD-Net. Additionally, cross-dataset generalization experiments on two publicly available datasets illustrate the strong generalization capability of the GLH-Bridge dataset.
Yansheng Li 0001, Yongjun Zhang 0002, Yihua Tan, Jin-Gang Yu, Song Bai 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Large-scale multi-view clustering via matrix factorization of consensus graph
Zengbiao Yang, Yihua Tan
Pattern Recognit.2
2024 Cloud-Guided Fusion With SAR-to-Optical Translation for Thick Cloud Removal
abstract
Deep learning has been widely used in thick cloud removal (TCR) for optical satellite images. Since thick clouds completely block the surface, synthetic aperture radar (SAR) images have recently been used to assist in the recovery of occluded information. However, this approach faces several challenges: 1) the significant domain gap between SAR and optical features can cause interference in recovering occluded optical information from SAR and 2) the TCR methods need to distinguish between cloudy and non-cloudy regions; otherwise, inconsistencies may arise between the recovered regions and the remaining non-cloudy regions. To this end, we propose a new SAR-assisted TCR method based on a two-step fusion framework, which consists of the feature alignment translation (FAT) network and the cloud-guided fusion (CGF) network. First, the FAT leverages the common features between SAR and optical images to translate SAR images into corresponding optical images, thus recovering the occluded information. Considering the gap between the translated images and the real cloud-free images, the CGF utilizes cloudy images to further refine the translated images, resulting in the cloud-removed images. In the CGF, cloud distribution is predicted to distinguish between cloudy and non-cloudy regions. Then, the cloud distribution is used to guide the refinement of recovered regions using non-cloudy regions. Extensive experiments on both simulated and real datasets show that the proposed algorithm achieves better performance compared with the state-of-the-art methods.
Xuanyu Xiang, Yihua Tan, Longfei Yan 0003
IEEE Trans. Geosci. Remote. Sens.2
2023 Mine-Distill-Prototypes for Complete Few-Shot Class-Incremental Learning in Image Classification
abstract
Recently, few-shot learning (FSL) has received increasing attention because of difficulties in sample collection in some application scenarios, such as maritime surveillance using synthetic aperture radar (SAR) or infrared images. In real situations of such scenarios, it is a common requirement that the model can recognize novel classes incrementally, namely class-incremental learning (CIL). Considering the above requirement, a novel problem that recognizes novel classes incrementally when both the base and novel class samples are scarce is proposed in this article. It is called complete few-shot CIL (C-FSCIL) for distinguishing from the FSCIL that assumes sufficient samples of base classes. Specifically, the following challenges of C-FSCIL are focused on: 1) distance measurement is used for recognizing novel classes incrementally, but the encoder is difficult to be learned well when base class samples are scarce, making some features unsuitable for calculating the distance, decreasing the performance and 2) the catastrophic forgetting problem becomes more difficult to be alleviated than that in FSCIL because of the scarcity of base class samples. To tackle both challenges, mine-distill-prototypes (MDP) algorithm is proposed, which consists of two parts: 1) prototypes-distillation (PD) network is proposed to learn to distill the features and prototypes into a lower dimensional in which ineffective features are eliminated and 2) the prototypes-weight (PW) network and the prototypes-selection (PS) training strategy are proposed for the catastrophic forgetting problem, which aims to capture the relationship between the base and novel prototypes. The superior performance of the proposed algorithm is demonstrated by the experiments on three datasets.
Yuan Tai, Yihua Tan, Shengzhou Xiong, Jinwen Tian
IEEE Trans. Geosci. Remote. Sens.2
2022 Learning Soft Estimator of Keypoint Scale and Orientation with Probabilistic Covariant Loss
abstract
Estimating keypoint scale and orientation is crucial to extracting invariant features under significant geometric changes. Recently, the estimators based on self-supervised learning have been designed to adapt to complex imaging conditions. Such learning-based estimators generally predict a single scalar for the keypoint scale or orientation, called hard estimators. However, hard estimators are difficult to handle the local patches containing structures of different objects or multiple edges. In this paper, a Soft Self-Supervised Estimator (S3Esti) is proposed to overcome this problem by learning to predict multiple scales and orientations. S3Esti involves three core factors. First, the estimator is constructed to predict the discrete distributions of scales and orientations. The elements with high confidence will be kept as the final scales and orientations. Second, a probabilistic covariant loss is proposed to improve the consistency of the scale and orientation distributions under different transformations. Third, an optimization algorithm is designed to minimize the loss function, whose convergence is proved in theory. When combined with different keypoint extraction models, S3Esti generally improves over 50% accuracy in image matching tasks under significant viewpoint changes. In the 3D reconstruction task, S3Esti decreases more than 10% reprojection error and improves the number of registered images. [code release]
Pei Yan, Yihua Tan, Shengzhou Xiong, Yuan Tai, Yansheng Li 0001
CVPR2
2022 Repeatable adaptive keypoint detection via self-supervised learning
Pei Yan, Yihua Tan, Yuan Tai
Sci. China Inf. Sci.2
2021 Explore Visual Concept Formation for Image Classification
abstract
Human beings acquire the ability of image classification through visual concept learning, in which the process of concept formation involves intertwined searches of common properties and concept descriptions. However, in most image classification algorithms using deep convolutional neural network (ConvNet), the representation space is constructed under the premise that concept descriptions are fixed as one-hot codes, which limits the mining of properties and the ability of identifying unseen samples. Inspired by this, we propose a learning strategy of visual concept formation (LSOVCF) based on the ConvNet, in which the two intertwined parts of concept formation, i.e. feature extraction and concept description, are learned together. First, LSOVCF takes sample response in the last layer of ConvNet to induct concept description being assumed as Gaussian distribution, which is part of the training process. Second, the exploration and experience loss is designed for optimization, which adopts experience cache pool to speed up convergence. Experiments show that LSOVCF improves the ability of identifying unseen samples on cifar10, STL10, flower17 and ImageNet based on several backbones, from the classic VGG to the SOTA Ghostnet. The code is available at \url{https://github.com/elvintanhust/LSOVCF}.
Shengzhou Xiong, Yihua Tan, Guoyou Wang
ICML2
2021 Subspace reconstruction based correlation filter for object tracking
Yuan Tai, Yihua Tan, Shengzhou Xiong, Jinwen Tian
Comput. Vis. Image Underst.2
2021 Unsupervised learning framework for interest point detection and description via properties optimization
Pei Yan, Yihua Tan, Yuan Tai, Dongrui Wu, Hanbin Luo, Xiaolong Hao
Pattern Recognit.2
2020 Multi-branch convolutional neural network for built-up area extraction from remote sensing image
Yihua Tan, Shengzhou Xiong, Pei Yan
Neurocomputing1
2020 Optimize TSK Fuzzy Systems for Regression Problems: Minibatch Gradient Descent With Regularization, DropRule, and AdaBound (MBGD-RDA)
abstract
Takagi–Sugeno–Kang (TSK) fuzzy systems are very useful machine learning models for regression problems. However, to our knowledge, there has not existed an efficient and effective training algorithm that ensures their generalization performance and also enables them to deal with big data. Inspired by the connections between TSK fuzzy systems and neural networks, we extend three powerful neural network optimization techniques, i.e., minibatch gradient descent (MBGD), regularization, and AdaBound, to TSK fuzzy systems, and also propose three novel techniques (DropRule, DropMF, and DropMembership) specifically for training TSK fuzzy systems. Our final algorithm, MBGD with regularization, DropRule, and AdaBound, can achieve fast convergence in training TSK fuzzy systems, and also superior generalization performance in testing. It can be used for training TSK fuzzy systems on datasets of any size; however, it is particularly useful for big datasets, on which currently no other efficient training algorithms exist.
Dongrui Wu, Ye Yuan 0002, Jian Huang 0001, Yihua Tan
IEEE Trans. Fuzzy Syst.4
2019 Visual Saliency Based Ship Extraction Using Improved Bing
abstract
In this paper, a novel ship extraction algorithm is proposed to acquire the precise segmentation. First, Binarized Normed Gradients (BING) is improved to locate the potential ship regions according to the remote sensing application. Second, the visual saliency computation by combining Hypercomplex Frequency Domain Transform (HFT) and Phase Quaternion Fourier Transform (PQFT) presented to analyze the located regions. Third, segmentation on the computed saliency map is conducted to extract precise blobs of the ship candidates. Finally, the prior knowledge about ship is utilized to remove the false alarms in the discrimination stage. The experimental results of typical remote sensing images show that the algorithm is very effective.
Yihua Tan, Zengrong Guan, Airong Sun
IGARSS1
2019 Infrared Small Target Detection Algorithm Based on Robust Tensor Decomposition Model within Bayesian Framework
abstract
Small targets detection in infrared video can be further improved by considering that the background has high correlation and low rank characteristics while foreground objects maintain sparsity. In this paper, a new infrared small target detection algorithm within Bayesian framework is proposed. A three-dimensional tensor structure of the video sequence is supposed to be decomposed into low rank background, sparse foreground and noise. The corresponding probabilistic models for the three parts form a Bayesian network which is solved by using variational Bayesian inference. Finally, the isolated sparse component is utilized for further target detecition. Experimental results show that the proposed method is suitable for the detection of small infrared target with good detection accuracy and robustness.
Yihua Tan
IGARSS1
2019 Generalized Compute-Compress-and-Forward
abstract
Compute-and-forward (CF) harnesses interference in wireless communications by exploiting structured coding. The key idea of CF is to compute integer combinations of code words from multiple source nodes, rather than to decode individual code words by treating others as noise. Compute-compress-and-forward (CCF) can further enhance the network performance by introducing compression operations at receivers. In this paper, we develop a more general compression framework, termed generalized CCF (GCCF), where the compression function involves the selection of message segments over finite fields. We show that GCCF achieves a broader compression rate region than CCF. We also compare our compression rate region with the fundamental Slepian-Wolf (SW) region. We show that GCCF is optimal in the sense of achieving the minimum total compression rate. We also establish the criteria under which GCCF achieves the SW region. In addition, we consider a two-hop relay network employing the GCCF scheme. We formulate a sum-rate maximization problem and develop an approximate algorithm to solve the problem. Numerical results are presented to demonstrate the performance superiority of GCCF over CCF and other schemes.
Hai Cheng, Xiaojun Yuan 0002, Yihua Tan
IEEE Trans. Inf. Theory3
2018 Precise Extraction of Built-Up Area Using Deep Features
abstract
Built-up area is one of the most important objects in remote sensing image analysis, therefore extracting built-up area automatically has attracted wide attention. Deep convolution neural network (CNN) was proposed to improve poor generalization ability of artificial features which had been adopted by traditional automatic extraction methods. In this paper, a more efficient CNN model is proposed to extract the deep features of remote sensing images, and then a graph model based on deep features is constructed to the full image for built-up area extraction. The experiments demonstrate that it has very good performance on the satellite remote sensing image data set.
Yihua Tan, Shengzhou Xiong, Yaming Li
IGARSS1
2018 Mobile Lattice-Coded Physical-Layer Network Coding with Practical Channel Alignment
abstract
Physical-layer network coding (PNC) is a communications paradigm that exploits overlapped transmissions to boost the throughput of wireless relay networks. A high point of PNC research was a theoretical proof that PNC that makes use of nested lattice codes could approach the information-theoretic capacity of a two-way relay network (TWRN), where two end nodes communicate via a relay node. The capacity cannot be achieved by conventional methods of time-division or straightforward network coding. Many practical challenges, however, remain to be addressed before the full potential of lattice-coded PNC can be realized. Two major challenges are: (1) for good performance in lattice-coded PNC, channels of simultaneously transmitting nodes must be aligned; (2) for lattice-coded PNC to be practical, the complexity of lattice encoding at the transmitters and lattice decoding at the receiver must be reduced. We address these challenges and implement a first lattice-coded PNC system on a software-defined radio (SDR) platform. Specifically, we design and implement a low-overhead channel precoding system that accurately aligns the channels of distributed nodes. In our implementation, the nodes use low-cost temperature-compensated oscillators (TCXO) only-a consequent challenge is that the channel alignment must be done more frequently and more accurately compared with the use of expensive oscillators. The low overhead and accurate channel alignment are achieved by (1) a channel precoding system implemented over FPGA to realize fast feedback of channel state information; (2) a highly-accurate carrier frequency offset (CFO) estimation method; and (3) a partial-feedback channel estimation method that significantly reduces the amount of feedback information from the receiver to the transmitters for channel precoding at the transmitters. To reduce lattice encoding and decoding complexities, we adapt the low-density lattice code (LDLC) for use in PNC systems. Experiments show that our implemented lattice-coded PNC achieves better bit error rate performance compared with timedivision and straightforward network coding systems. It also has good throughput performance in mobile non-LoS scenarios.
Yihua Tan, Soung Chang Liew, Tao Huang 0007
IEEE Trans. Mob. Comput.1
2017 Compute-Compress-and-Forward: New Results
abstract
In this paper, we consider the design of compression functions for compute-and-forward (CF) based relaying. We develop a general compressing framework, termed generalized compute-compress-and-forward (GCCF), where the compression function involves multiple quantization- and-modulo lattice operations. We show that GCCF achieves a broader compression rate region than CCF. We also compare our compression rate region with the fundamental Slepian-Wolf (SW) region. We show that GCCF is optimal in the sense of achieving the minimum total compression rate. We also establish the criteria under which GCCF achieves the SW region. In addition, we consider a two-hop relay network employing the GCCF scheme. Numerical results are presented to demonstrate the performance superiority of GCCF over other schemes.
Hai Cheng, Xiaojun Yuan 0002, Yihua Tan
GLOBECOM3
2017 Automatic extraction of built-up area based on deep convolution neural network
abstract
Built-up area has been one of the most important objects to be extracted in remote sensing images. Several factors such as complex structure, diverse texture and varied background, bring the challenges for the task of built-up area extraction. In this paper, a multiple input structure of deep convolution neural network (CNN) is proposed to extract built-up area automatically, which can fuse the information of panchromatic and multispectral remote sensing image. The image patch based classification results are further refined by postprocessing of segmentation techniques. The experiments demonstrate that the proposed method has better generalization ability compared to the state-of-the-art method, and the overall classification accuracy is above 98%.
Yihua Tan, Feifei Ren, Shengzhou Xiong
IGARSS1
2016 Real-time cloud detection in high resolution images using Maximum Response Filter and Principle Component Analysis
abstract
In order to maximally make use of the limited capacities of storage and transmission, it's necessary for onboard computer to adaptively compress the cloud regions with lower quality compared with the regions without cloud cover. Therefore, the information processing unit needs to recognize the cloud regions before compress the image. To meet this requirement of satellite imaging payload, a novel approach for real time cloud detection is proposed. First, the visual dictionary is learnt from the training features extracted using Maximum Response (MR) Filter. Second, Principle Component Analysis (PCA) is utilized to reduce the dimensions of the visual words for the quick word search. Third, the MR feature of an image patch is converted into the histogram of visual word, in which the MR feature of a pixel is replaced by the index of the most similar visual word. Finally, the histogram is fed into the trained SVM classifier to detect cloud patch. The experimental results verify that the proposed approach can highly precisely detect the cloud region in patch unit.
Yihua Tan, Feifei Ren
IGARSS1
2016 Semi-automatic building extraction from very high resolution remote sensing imagery via energy minimization model
abstract
Extraction of objects such as buildings from very high resolution(VHR) remote sensing imagery is an important task nowadays. In practical application, the extraction precision is normally satisfied through interactive manual input. Therefore, we propose a semi-automatic building extraction framework with an energy minimization model, which includes two stages: the first stage generates the coarse information of foreground and background, and the second stage extracts the building object by applying energy minimization model. More specifically, in the first stage, the VHR imagery is grouped into small superpixels which are roughly merged into background or foreground according to a line drawn manually. Based on the coarse information of the first stage, we apply the energy minimization model to obtain the precise building object using graph cut optimization. Experimental results indicate our building extraction framework can precisely extract the buildings with different shapes.
Yihua Tan, Yujie Yu, Shengzhou Xiong, Jinwen Tian
IGARSS1
2016 A novel spatio-temporal saliency approach for robust dim moving target detection from airborne infrared image sequences
Yansheng Li 0001, Yongjun Zhang 0002, Jin-Gang Yu, Yihua Tan, Jinwen Tian, Jiayi Ma 0001
Inf. Sci.4
2016 Unsupervised Multilayer Feature Learning for Satellite Image Scene Classification
abstract
This letter proposes a simple but effective approach to automatically learn a multilayer image feature for satellite image scene classification. Different from the hand-crafted features which are empirically designed but lack high generalization ability, the proposed approach can autonomously extract the data-dependent feature. The presented feature extraction algorithm is composed of two layers, and the bases of these two layers are uniformly learned by a plain $K$-means clustering algorithm. Coincidentally, the feature extraction performance of the aforementioned two layers is consistent with visual processing of human visual cortex. More specifically, the first layer can generate edgelike bases, which are analogous to the neuron responses of primary visual cortex (V1), and the second layer can produce cornerlike bases, which resemble the neuron responses of visual extrastriate cortical area two (V2). The proposed feature extraction approach can automatically extract not only simple structure features (e.g., edges) but also complex structure features (e.g., corners and junctions). The learned feature is further discriminated by the linear support vector machine classifier for scene classification. In order to fairly demonstrate the validity of the proposed feature extraction approach, its satellite image scene classification performance is evaluated on the public UCM-21 data set. Experimental results show that the proposed approach can outperform several recent state-of-the-art approaches.
Yansheng Li 0001, Chao Tao 0001, Yihua Tan, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.3
2015 Compute-compress-and-forward
abstract
Compute-and-forward (CF) can harness interference in a multi-hop relay network by allowing a relay to decode and forward a combination of source messages. However, the total forwarding rate of the relays belonging to one hop may far exceed the total information rate of the sources, which implies information redundancy and spectral inefficiency. To tackle this problem, we propose a novel relaying strategy termed compute-compress-and-forward (CCF). Compared with CF, our proposed CCF scheme includes an extra compressing stage in which the computed combinations of the relays are compressed to reduce forwarding rates. We design the compressing function and develop a successive recovering algorithm to recover source messages at a destination. Numerical results are presented to demonstrate the performance advantage of CCF over CF.
Yihua Tan, Xiaojun Yuan 0002
ISIT1
2015 Kernel regression in mixed feature spaces for spatio-temporal saliency detection
Yansheng Li 0001, Yihua Tan, Jin-Gang Yu, Shengxiang Qi, Jinwen Tian
Comput. Vis. Image Underst.2
2015 Built-Up Area Detection From Satellite Images Using Multikernel Learning, Multifield Integrating, and Multihypothesis Voting
abstract
This letter proposes a novel supervised approach for accurate built-up area detection from high-resolution remote sensing images. In existing supervised built-up area detection approaches based on block-based image interpretation, the determination of the block size and the pursuit of the pixel-level result are not well addressed. Concerning these issues, this letter proposes a complete and systematic approach. It first utilizes multikernel learning to incorporate multiple features to implement the block-level image interpretation. Then, multifield integrating (i.e., the image interpretation results using different block sizes are fused) is proposed to obtain the block-level result. On the basis of the achieved result of the second step, multihypothesis voting is finally presented for working toward the pixel-level built-up area detection result through multihypothesis superpixel representation and graph smoothing. The proposed approach has been validated in the ZY-3 and GF-1 satellite images, and experimental results show that the proposed approach can outperform the state-of-the-art approaches.
Yansheng Li 0001, Yihua Tan, Shengxiang Qi, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.2
2014 Urban building extraction via visual graphical topic model
abstract
This paper addresses the automatic building extraction problem from high-resolution remote sensing images. The buildings in remote sensing images generally represent different shapes (i.e., simple rectangular or complex hybrid shape), it is intractable to extract all the buildings with different shapes once. Therefore, we adopt a hierarchical extraction style with a visual graphical topic model embedded, which includes two stages: the first stage detects the simple rectangular buildings and the second stage extracts the complex hybrid buildings. More specifically, the first stage is mainly responsible for regular buildings detection, and unsupervised visual graphical topic model (i.e., replicated softmax restricted boltzmann machine) and supervised discriminative model learning, and the second stage is mainly in charge of complex buildings extraction using the learned semantic feature mapping and discriminative model. Experimental results show that the second stage can obviously improve the building detection rate with slightly increasing the false alarm rate.
Yansheng Li 0001, Yihua Tan, Jinwen Tian
IGARSS2
2014 Adaptive and fast target detection in high-resolution SAR image
abstract
In this paper, a new adaptive and fast Constant false alarm rate (CFAR) target detection algorithm based on two level CFAR (TL-CFAR) detectors in high-resolution synthetic aperture radar (SAR) images is proposed. In the first level, the initial mask of targets is obtained by Cell Averaging CFAR (CA-CFAR) detector. In the second level, the precise parameters estimation of CFAR in the local window is implemented by removing those pixels that may belong to the neighboring targets which is identified from the first detector. The problem of high computational complexity of two levels CFAR detector is mitigated by introducing the integral image. Real SAR image data is used to verify the effectiveness of the proposed algorithm, and the results indicate that this algorithm can detect targets fast and precisely.
Yihua Tan, Airong Sun, Qingyun Li
IGARSS1
2014 Maximal Entropy Random Walk for Region-Based Visual Saliency
abstract
Visual saliency is attracting more and more research attention since it is beneficial to many computer vision applications. In this paper, we propose a novel bottom-up saliency model for detecting salient objects in natural images. First, inspired by the recent advance in the realm of statistical thermodynamics, we adopt a novel mathematical model, namely, the maximal entropy random walk (MERW) to measure saliency. We analyze the rationality and superiority of MERW for modeling visual saliency. Then, based on the MERW model, we establish a generic framework for saliency detection. Different from the vast majority of existing saliency models, our method is built on a purely region-based strategy, which is able to yield high-resolution saliency maps with well preserved object shapes and uniformly highlighted salient regions. In the proposed framework, the input image is first over-segmented into superpixels, which are taken as the primary units for subsequent procedures, and regional features are extracted. Then, saliency is measured according to two principles, i.e., uniqueness and visual organization, both implemented in a unified approach, i.e., the MERW model based on graph representation. Intensive experimental results on publicly available datasets demonstrate that our method outperforms the state-of-the-art saliency models.
Jin-Gang Yu, Ji Zhao 0001, Jinwen Tian, Yihua Tan
IEEE Trans. Cybern.4
2013 Unsupervised Detection of Built-Up Areas From Multiple High-Resolution Remote Sensing Images
abstract
Given a set of high-resolution remote sensing images covering different scenes, we propose an unsupervised approach to simultaneously detect possible built-up areas from them. The motivation behind is that the frequently recurring appearance patterns or repeated textures corresponding to common objects of interest (e.g., built-up areas) in the input image data set can help us discriminate built-up areas from others. With this inspiration, our method consists of two steps. First, we extract a large set of corners from each input image by an improved Harris corner detector. Afterward, we incorporate the extracted corners into a likelihood function to locate candidate regions in each input image. Given a set of candidate build-up regions, in the second stage, we formulate the problem of build-up area detection as an unsupervised grouping problem. The candidate regions are modeled through texture histogram, and the grouping problem is solved by spectrum clustering and graph cuts. Experimental results show that the proposed approach outperforms the existing algorithms in terms of detection accuracy.
Chao Tao 0001, Yihua Tan, Zhengrong Zou, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.2
2012 Urban area detection using multiple Kernel Learning and graph cut
abstract
This paper presents a new method for urban detection from high-spatial-resolution satellite images. Unlike traditional approaches using only texture information for urban detection, we integrate several complementary image features through multiple Kernel Learning framework, and demonstrate that fusing multiple features can help improving urban detection accuracy rate. Furthermore, since that most of supervised urban classification approaches are mainly based on block-based image interpretation, the resulting urban boundary is very coarse. To handle this, we formulate the urban boundary refinement as a binary labeling problem, and propose a graph cut based approach to solve it. Experimental results show that the proposed approach outperforms the existing algorithm in terms of detection accuracy.
Chao Tao 0001, Yihua Tan, Jin-Gang Yu, Jin-Wen Tian
IGARSS2
2011 Airport Detection From Large IKONOS Images Using Clustered SIFT Keypoints and Region Information
abstract
This letter presents a new method for airport detection from large high-spatial-resolution IKONOS images. To this end, we describe airport by a set of scale-invariant feature transform (SIFT) keypoints and detect it using an improved SIFT matching strategy. After obtaining SIFT matched keypoints, to both discard the redundant matched points and locate the possible regions of candidates that contain the target, a novel region-location algorithm is proposed, which exploits the clustering information from matched SIFT keypoints, as well as the region information extracted through the image segmentation. Finally, airport recognition is achieved by applying the prior knowledge to the candidate regions. Experimental results show that the proposed approach outperforms the existing algorithms in terms of detection accuracy.
Chao Tao 0001, Yihua Tan, Huajie Cai, Jin-Wen Tian
IEEE Geosci. Remote. Sens. Lett.2
2011 Efficient Multi-Input/Multi-Output VLSI Architecture for Two-Dimensional Lifting-Based Discrete Wavelet Transform
abstract
This brief paper proposes an efficient multi-input/multi-output VLSI architecture (MIMOA) for two-dimensional lifting-based discrete wavelet transform (DWT). The novelty is the simplicity and generality to construct the MIMOA, which is a high-speed architecture with computing time as low as N2/M for an N × N image with controlled increase of hardware cost. M is the throughput rate.
Xin Tian 0006, Yihua Tan, Jin-Wen Tian
IEEE Trans. Computers3
2007 Accurate Dynamic Scene Model for Moving Object Detection
abstract
Adaptive pixel-wise Gaussian mixture model (GMM) is a popular method to model dynamic scenes viewed by a fixed camera. However, it is not a trivial problem for GMM to capture the accurate mean and variance of a complex pixel. This paper presents a two-layer Gaussian mixture model (TLGMM) of dynamic scenes for moving object detection. The first layer, namely real model, deals with gradually changing pixels specially; the second layer, called on-ready model, focuses on those pixels changing significantly and irregularly. TLGMM can represent dynamic scenes more accurately and effectively. Additionally, a long term and a short term variance are taken into account to alleviate the transparent problems faced by pixel-based methods.
Yihua Tan, Jin-Wen Tian, Jian Liu 0011
ICIP (6)2