VLDB 2026 Research / reviewers in the wild / expert
Yuan Yuan 0001
dblp:64/5845-1
· DBLP profile ↗
365ranked-venue papers
63as first author
149since 2021 · last 2026
0000-0002-0404-5498ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 159 · 31 first-author · 48 since 2021Graphics, computer vision, multimedia, augmented reality and games · 123 · 16 first-author · 47 since 2021Applied, interdisciplinary, general and emerging computing · 96 · 19 first-author · 61 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow DerivativesabstractReconstructing controllable Gaussian splats for articulated objects from monocular video is especially challenging due to its inherently insufficient constraints. Existing methods address this by relying on dense masks and manually defined control signals, limiting their real-world applications. In this paper, we propose an annotation-free method, FreeGaussian, which mathematically disentangles camera egomotion and articulated movements via flow derivatives. By establishing a connection between 2D flows and 3D Gaussian dynamic flow, our method enables optimization and continuity of dynamic Gaussian motions from flow priors without any control signals. Furthermore, we introduce a 3D spherical vector controlling scheme, which represents the state as a 3D Gaussian trajectory, thereby eliminating the need for complex 1D control signal calculations and simplifying controllable Gaussian modeling. Extensive experiments on articulated objects demonstrate the state-of-the-art visual performance and precise, part-aware controllability of our method. Delin Qu, Junli Liu, Haoming Song, Dong Wang 0028, Yuan Yuan 0001, Bin Zhao 0001 |
AAAI | 7 |
| 2026 | Multimodal Graph Conditioned Diffusion Model for Video CaptioningabstractVideo captioning aims to describe the content of a given video with condensed natural language sentences. Such a captioning task is full of challenges since the high requirements for visual-textual relevance and multimodal fusion understanding. Previous works primarily focus on visual content modeling, often overlooking the rich semantic correlations between visual and textual modalities, which results in incomplete understanding of the multimodal context and suboptimal caption accuracy. In this paper, we propose a multimodal graph conditioned diffusion model for video captioning, named MGCDVc. The idea behind our model is to incorporate graph-based relational reasoning with diffusion-based generative modeling to jointly model cross-modal relationships and capture latent semantic structure. Specifically, we learn a set of latent concept anchors to bridge the visual and textual modality nodes, enabling the construction of a weighted multimodal graph. Then we introduce the graph conditioned diffusion strategy which generates the textual semantic nodes and associated edges under the graph structure awareness condition. Furthermore, a soft pruning mechanism is designed to filter out low-quality nodes, thus further refining the generated multimodal graph to provide more accurate semantic structural guidance for caption generation. Experimental results on several popular datasets demonstrate that our model achieves better performance in video captioning task. Benhui Zhang 0001, Junyu Gao 0001, Yuan Yuan 0001 |
WWW | 3 |
| 2026 | Spectral consistency learning for cross-domain hyperspectral image classification
Zhiyu Jiang, Dandan Ma, Yuan Yuan 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Quantum-inspired interpretable deep learning architecture for text sentiment analysis
Bingyu Li 0002, Da Zhang 0010, Zhiyuan Zhao 0005, Yuan Yuan 0001, Junyu Gao 0001, Xuelong Li 0001 |
Neural Networks | 4 |
| 2026 | Distantly supervised reinforcement localization for real-world object distribution estimation
Haojie Guo, Junyu Gao 0001, Yuan Yuan 0001 |
Pattern Recognit. | 3 |
| 2026 | Hybrid texture-structural learning for hyperspectral image classification
Mingxin Jin, Cong Wang 0033, Yuan Yuan 0001 |
Pattern Recognit. | 3 |
| 2026 | GLGF-CR: A Gated Local-Global Fusion approach for cloud removal in real-world remote sensing
Ganchao Liu, Jiawei Qiu, Yuan Yuan 0001 |
Pattern Recognit. | 4 |
| 2026 | Enhancing visual inertial odometry with efficient dynamic PerceptionNet and consistency improvement fusion
Ganchao Liu, Haozhe Tian, Yuan Yuan 0001 |
Pattern Recognit. | 4 |
| 2026 | FSO-VO: Visual odometry based on dense optical flow prediction and sequence optimization
Ganchao Liu, Sihang Zhang, Yuan Yuan 0001 |
Pattern Recognit. | 4 |
| 2026 | Fully PolSAR image reconstruction for enhanced land cover mapping
Junyu Gao 0001, Yuan Yuan 0001 |
Pattern Recognit. | 3 |
| 2026 | Efficient greedy optimization method for k-means
Yuan Yuan 0001, Lin Zhao 0003, Shenfei Pei, Feiping Nie 0001 |
Pattern Recognit. | 1 |
| 2026 | Hierarchical textual-visual guidance for referring remote sensing segmentation
Qi Wang 0009, Yuan Yuan 0001, Junyu Gao 0001 |
Pattern Recognit. | 4 |
| 2026 | Efficient hierarchical multi-resolution k-means clustering
Lin Zhao 0003, Yuan Yuan 0001, Feiping Nie 0001 |
Pattern Recognit. | 2 |
| 2026 | BEMN: Balanced Bias Enhanced Multi-Branch Network for Cross-View Geo-LocalizationabstractCross-view geo-localization (CVGL) offers a promising alternative for positioning in GNSS-constrained environments through visual matching techniques. Extreme viewpoint variations and the complexity of real-world scenes present significant challenges to this task. However, current methods primarily focus on learning single-scale features, which may be inadequate for practical applications. Although some approaches attempt to incorporate multi-scale representations, they may suffer from unimodal bias arising from structural discrepancies among model branches, limiting effective multi-scale feature extraction. To address these issues, we propose a fully multi-branch network architecture, named BEMN, which is designed to learn multi-scale robust feature representations. Specifically, we construct a multi-branch backbone network based on pretrained visual models and design a two-stage training strategy. In the first stage, a separate training scheme is employed to thoroughly optimize each branch of the network, and a joint feature alignment (JFA) module is introduced to align cross-view features. The entire network is fine-tuned in the second stage, where a frequency domain adjustment (FDA) module is designed to improve performance. To further assess the generalization ability of CVGL methods, we establish Xian-37, a highly challenging CVGL test dataset featuring complex real scenes captured from diverse platforms and viewpoints. Experimental results across multiple public benchmarks validate the superiority of our approach, achieving state-of-the-art performance and demonstrating outstanding generalization capabilities. Our code and model are available at https://github.com/VERYBC/BEMN. Bo Sun 0017, Yuan Yuan 0001, Ganchao Liu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | PRFCM: Poisson-Specific Residual-Driven Fuzzy $C$-Means Clustering for Image SegmentationabstractA Fuzzy$C$-Means (FCM) algorithm has been widely applied to image segmentation due to its simplicity and effectiveness. However, conventional FCM and its variants often struggle to maintain robustness and accuracy when dealing with complex noise environments, particularly Poisson and mixed Poisson-Gaussian noise. To address this shortcoming, this work proposes a novel Poisson-specific Residual-driven FCM (PRFCM) algorithm for robust image segmentation, which is the first work to develop a dedicated residual regularization mechanism that effectively realizes the robust estimation of Poisson noise (regarded as residual between noisy and noise-free images). It incorporates a weighted$\ell _{2}$-norm regularization term with respect to Poisson noise distribution into FCM. An iterative residual approximation method is introduced to solve the minimization problem about residual, thus simplifying PRFCM's optimization procedure and enhancing its computational efficiency. PRFCM is also extended to cope with mixed Poisson-Gaussian noise scenarios without compromising performance. Experimental results on both simulated and real-scene images demonstrate that the proposed approach outperforms other FCM-related methods in terms of segmentation accuracy, noise resilience, and structural preservation, especially in challenging noise conditions. Cong Wang 0033, Shengnan Jiang, Yuan Yuan 0001, Junfeng Jing, MengChu Zhou, Witold Pedrycz |
IEEE Trans. Fuzzy Syst. | 3 |
| 2026 | Hyperspectral Image Super-Resolution via Boundary Perception and Topology Inference
Cong Wang 0033, Yuan Yuan 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | Text-Pass Filter: An Efficient Scene Text DetectorabstractTo pursue an efficient text assembling process, existing methods detect texts via the shrink-mask expansion strategy. However, the shrinking operation loses the visual features of text margins and confuses the foreground and background difference, which brings intrinsic limitations to recognize text features. We follow this issue and design Text-Pass Filter (TPF) for arbitrary-shaped text detection. It segments the whole text directly, which avoids the intrinsic limitations. It is noteworthy that different from previous whole text region-based methods, TPF can separate adhesive texts naturally without complex decoding or post-processing processes, which makes it possible for real-time text detection. Concretely, we find that the band-pass filter allows through components in a specified band of frequencies, called its passband but blocks components with frequencies above or below this band. It provides a natural idea for extracting whole texts separately. By simulating the band-pass filter, TPF constructs a unique feature-filter pair for each text. In the inference stage, every filter extracts the corresponding matched text by passing its pass-feature and blocking other features. Meanwhile, considering the large aspect ratio problem of ribbon-like texts makes it hard to recognize texts wholly, a Reinforcement Ensemble Unit (REU) is designed to enhance the feature consistency of the same text and to enlarge the filter's recognition field to help recognize whole texts. Furthermore, a Foreground Prior Unit (FPU) is introduced to encourage TPF to discriminate the difference between the foreground and background, which improves the feature-filter pair quality. Experiments demonstrate the effectiveness of REU and FPU while showing the TPF's superiority. Chuang Yang 0003, Haozhao Ma, Xu Han 0019, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 4 |
| 2025 | Enhancing Low-Rank Adaptation with Recoverability-Based Reinforcement Pruning for Object CountingabstractObject counting is crucial for understanding the distribution of objects in different scenarios. Recently, many object counting networks have been designed to be more complex to achieve marginal improvements, leading to excessive time spent on model design. With the development of large models (LMs), various visual tasks can be accomplished by transferring pre-trained weights from LMs and fine-tuning them. However, tens of millions of training data make the pre-training parameters of LMs not entirely necessary. Moreover, if unnecessary parameters in the large model are not removed, it may lead to decreased performance on the tasks to be transferred. Motivated by this, this paper proposes an Enhancing low-Rank adaptation with Recoverability-based Reinforcement Pruning (E3RP) method to balance the complexity of large model and the accuracy of counting tasks. Firstly, we design a new reward mechanism based on the feature similarity of large model before and after globally unstructured pruning of specific parameters. Additionally, we propose a Patch Query Flip Attention (PQFA) mechanism to align multi-scale features through bidirectional interaction of features. Finally, the parameters of large model are pruned utilizing the pruning rate autonomously determined by the reinforcement learning network, and the large model is fine-tuned to counting tasks by a simple decoding head. Extensive experiments on four cross-scenario datasets demonstrate that the proposed method can remove redundant network parameters while ensuring network performance, with a maximum reduction of up to 63%. Haojie Guo, Junyu Gao 0001, Yuan Yuan 0001 |
AAAI | 3 |
| 2025 | Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial AttacksabstractRecent Customized Portrait Generation (CPG) methods, taking a facial image and a textual prompt as inputs, have attracted substantial attention. Although these methods generate high-fidelity portraits, they fail to prevent the generated portraits from being tracked and misused by malicious face recognition systems. To address this, this paper proposes a Customized Portrait Generation framework with facial Adversarial attacks (Adv-Cpg). Specifically, to achieve facial privacy protection, we devise a lightweight local ID encryptor and an encryption enhancer. They implement progressive double-layer encryption protection by directly injecting the target identity and adding additional identity guidance, respectively. Furthermore, to accomplish fine-grained and personalized portrait generation, we develop a multi-modal image customizer capable of generating controlled fine-grained facial features. To the best of our knowledge, Adv-Cpg is the first study that introduces facial adversarial attacks into CPG. Extensive experiments demonstrate the superiority of Adv-Cpg, e.g., the average attack success rate of the proposed Adv-Cpg is 28.1% and 2.86% higher compared to the SOTA noise-based attack methods and unconstrained attack methods, respectively. Hongyuan Zhang 0001, Yuan Yuan 0001 |
CVPR | 3 |
| 2025 | Where Does It Exist from the Low-Altitude: Spatial Aerial Video GroundingabstractThe task of localizing an object's spatial tube based on language instructions and video, known as spatial video grounding (SVG), has attracted widespread interest. Existing SVG tasks have focused on ego-centric fixed front perspective and simple scenes, which only involved a very limited view and environment. However, UAV-based SVG remains underexplored, which neglects the inherent disparities in drone movement and the complexity of aerial object localization. To facilitate research in this field, we introduce the novel spatial aerial video grounding (SAVG) task. Specifically, we meticulously construct a large-scale benchmark, UAV-SVG, which contains over 2 million frames and offers 216 highly diverse target categories. To address the disparities and challenges posed by complex aerial environments, we propose a new end-to-end transformer architecture, coined SAVG-DETR. The innovations are three-fold. 1) To overcome the computational explosion of self-attention when introducing multi-scale features, our encoder efficiently decouples the multi-modality and multi-scale spatio-temporal modeling into intra-scale multi-modality interaction and cross-scale visual-only fusion. 2) To enhance small object grounding ability, we propose the language modulation module to integrate multi-scale information into language features and the multi-level progressive spatial decoder to decode from high to low level. The decoding stage for the lower-level vision-language features is gradually increased. 3) To improve the prediction consistency across frames, we design the decoding paradigm based on offset generation. At each decoding stage, we utilize reference anchors to constrict the grounding region, use context-rich object queries to predict offsets, and update reference anchors for the next stage. From coarse to fine, our SAVG-DETR gradually bridges the modality gap and iteratively refines reference anchors of the referred object, eventually grounding the spatial tube. Extensive experiments demonstrate that our SAVG-DETR significantly outperforms existing state-of-the-art methods. The dataset and code will be available at here. Yang Zhan 0007, Yuan Yuan 0001 |
NeurIPS | 2 |
| 2025 | Building extraction from remote sensing images with deep learning: A survey on vision techniques
Yuan Yuan 0001, Junyu Gao 0001 |
Comput. Vis. Image Underst. | 1 |
| 2025 | H3T: Hierarchical Transferable Transformer with TokenMix for Unsupervised Domain Adaptation
Yihua Ren, Junyu Gao 0001, Yuan Yuan 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Memory-enhanced hierarchical transformer for video paragraph captioning
Benhui Zhang 0001, Junyu Gao 0001, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2025 | Distance-aware network for physical-world object distribution estimation and counting
Yuan Yuan 0001, Haojie Guo, Junyu Gao 0001 |
Pattern Recognit. | 1 |
| 2025 | Dual Heterogeneous Network for Hyperspectral Image ClassificationabstractModeling discriminative spectral-spatial features is a key to improving hyperspectral image classification performance. However, existing methods cannot fully characterize the spatial specificity of hyperspectral images, thus making them unable to fully explore the useful information within the image and further improve the discriminative power of features. To address this issue, this work proposes a dual heterogeneous network (DHNet) for hyperspectral image classification. Specifically, the network consists of spatial-specific and spectral-specific branches and captures spectral-spatial features with complementarity by combining convolution and spectral-spatial involution. To better characterize spatial specificity, the spectral-spatial involution modifies the weight parameters based on the center spectral information and neighborhood spatial information of various spatial locations. Besides, two feature calibration modules are proposed. Spatial-specific and spectral-specific weights are generated from the respective branches to calibrate the features captured by the other branches to improve the information interaction between the two branches. The center spectral mapping integrates the spectral features of the target pixel into the feature to suppress the influence of the neighboring disturbing pixels. Experimental results on four datasets indicate that DHNet achieves an accuracy improvement of 1.23%, 2.03%, 2.52%, and 1.77% over the state-of-the-art peers, respectively. Mingxin Jin, Cong Wang 0033, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Dual-Stage Prior-Driven Diffusion Model for Remote Sensing Spectral Super-ResolutionabstractSpectral super-resolution (SSR) is a key technology for generating high spatial resolution hyperspectral images (HSIs). However, deep learning approaches for SSR, especially generative models like diffusion, often rely heavily on large training datasets. Furthermore, their stochastic generation process can compromise the precise spectral fidelity required in remote sensing. To address this limitation, we propose a dual-stage prior-driven diffusion model (DPDM), for SSR tasks. DPDM comprises two modules: the prior-driven diffusion module (PDM) and the spectral refinement module (SRM). PDM replaces the conventional pure noise input with a structured prior, which we term the prior-informed noise (PIN). This PIN is deterministically generated by projecting the input multispectral image (MSI) onto a spectral basis, which is extracted from derived from the spectral response function. By initializing the reverse process with this information-rich starting point, our model significantly reduces its dependence on large training datasets and inherently enforces spectral consistency. The SRM is subsequently introduced to specifically target and correct residual artifacts and coarse features from the initial stage. Employing a hierarchical multi-scale architecture and a single-sample optimization framework, the SRM meticulously restores fine-grained details while suppressing noise in the PDM’s output. By integrating these two modules, DPDM progressively enhances both spatial and spectral fidelity. Extensive experimental results demonstrate that DPDM achieves competitive performance in SSR tasks. Zengyi Li, Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Real-Time Text Detection With Similar Mask in Traffic, Industrial, and Natural ScenesabstractTexts on the intelligent transportation scene include mass information. Fully harnessing this information is one of the critical drivers for advancing intelligent transportation. Unlike the general scene, detecting text in transportation has extra demand, such as a fast inference speed, except for high accuracy. Most existing real-time text detection methods are based on the shrink mask, which loses some geometry semantic information and needs complex post-processing. In addition, the previous method usually focuses on correct output, which ignores feature correction and lacks guidance during the intermediate process. To this end, we propose an efficient multi-scene text detector that contains an effective text representation similar mask (SM) and a feature correction module (FCM). Unlike previous methods, the former aims to preserve the geometric information of the instances as much as possible. Its post-progressing saves 50% of the time, accurately and efficiently reconstructing text contours. The latter encourages false positive features to move away from the positive feature center, optimizing the predictions from the feature level. Some ablation studies demonstrate the efficiency of the SM and the effectiveness of the FCM. Moreover, the deficiency of existing traffic datasets (such as the low-quality annotation or closed source data unavailability) motivated us to collect and annotate a traffic text dataset, which introduces motion blur. In addition, to validate the scene robustness of the SM-Net, we conduct experiments on traffic, industrial, and natural scene datasets. Extensive experiments verify it achieves (SOTA) performance on several benchmarks. The code and dataset are available at:https://github.com/fengmulin/SMNet. Xu Han 0019, Junyu Gao 0001, Chuang Yang 0003, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Focus Entirety and Perceive Environment for Arbitrary-Shaped Text DetectionabstractDue to the diversity of scene text in aspects such as font, color, shape, and size, accurately and efficiently detecting text is still a formidable challenge. Among the various detection approaches, segmentation-based approaches have emerged as prominent contenders owing to their flexible pixel-level predictions. However, these methods typically model text instances in a bottom-up manner, which is highly susceptible to noise. In addition, the prediction of pixels is isolated without introducing pixel-feature interaction, which also influences the detection performance. To alleviate these problems, we propose a multi-information level arbitrary-shaped text detector consisting of a focus entirety module (FEM) and a perceive environment module (PEM). The former extracts instance-level features and adopts a top-down scheme to model texts to reduce the influence of noises. Specifically, it assigns consistent entirety information to pixels within the same instance to improve their cohesion. In addition, it emphasizes the scale information, enabling the model to distinguish varying scale texts effectively. The latter extracts region-level information and encourages the model to focus on the distribution of positive samples in the vicinity of a pixel, which perceives environment information. It treats the kernel pixels as positive samples and helps the model differentiate text and kernel features. Extensive experiments demonstrate the FEM's ability to efficiently support the model in handling different scale texts and confirm the PEM can assist in perceiving pixels more accurately by focusing on pixel vicinities. Comparisons show the proposed model outperforms existing state-of-the-art approaches on four public datasets. Xu Han 0019, Junyu Gao 0001, Chuang Yang 0003, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 4 |
| 2025 | Spotlight Text Detector: Spotlight on Candidate Regions Like a CameraabstractThe irregular contour representation is one of the tough challenges in scene text detection. Although segmentation-based methods have achieved significant progress with the help of flexible pixel prediction, the overlap of geographically close texts hinders detecting them separately. To alleviate this problem, some shrink-based methods predict text kernels and expand them to restructure texts. However, the text kernel is an artificial object with incomplete semantic features that are prone to incorrect or missing detection. In addition, different from the general objects, the geometry features (aspect ratio, scale, and shape) of scene texts vary significantly, which makes it difficult to detect them accurately. To consider the above problems, we propose an effective spotlight text detector (STD), which consists of a spotlight calibration module (SCM) and a multivariate information extraction module (MIEM). The former concentrates efforts on the candidate kernel, like a camera focus on the target. It obtains candidate features through a mapping filter and calibrates them precisely to eliminate some false positive samples. The latter designs different shape schemes to explore multiple geometric features for scene texts. It helps extract various spatial relationships to improve the model's ability to recognize kernel regions. Ablation studies prove the effectiveness of the designed SCM and MIEM. Extensive experiments verify that our STD is superior to existing state-of-the-art methods on various datasets, including ICDAR2015, CTW1500, MSRA-TD500, and Total-Text. Xu Han 0019, Junyu Gao 0001, Chuang Yang 0003, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 4 |
| 2025 | Confident Multi-View StereoabstractSolving the Multi-View Stereo (MVS) problem is a cornerstone in computer vision, with depth map estimation and fusion being one of the most critical approaches. The depth confidence map is pivotal in ensuring the precision and completeness of the reconstruction outcomes. These algorithms frequently encounter a trade-off between completeness and accuracy in the confidence map, which can significantly impair the final reconstruction results. This paper analyzes the causes and phenomena of these issues, namely Confidence Jitter, Confidence Gap, and Confidence Disappearance. From these insights, a multi-view stereo network named CF-MVSNet is introduced, comprising three essential components. Firstly, the method mitigates the Confidence Jitter problem through two confidence fusion strategies. Secondly, it narrows the depth sampling space to near sub-pixel levels, addressing the Confidence Gap through neighborhood-average pooling. Lastly, the algorithm tackles the Confidence Disappearance problem resulting from multi-scale classification and regression with a loss function named CL. Our proposed method demonstrates superior performance across two critical metrics: the completeness of the depth map and the accuracy of the reconstructed point cloud, outperforming current state-of-the-art MVS methods. Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 3 |
| 2025 | Hierarchical Context Measurement Network for Single Hyperspectral Image Super-ResolutionabstractSingle hyperspectral image super-resolution aims to enhance the spatial resolution of a hyperspectral image without relying on any auxiliary information. Despite the abundant spectral information, the inherent high-dimensionality in hyperspectral images still remains a challenge for memory efficiency. Recently, recursion-based methods have been proposed to reduce memory requirements. However, these methods utilize the reconstruction features as feedback embedding to explore context information, leading to sub-optimal performance as they ignore the complementarity of different hierarchical levels of information in the context. Additionally, existing methods equivalently compensate the previous feedback information to the current band, resulting in an indistinct and untargeted introduction of the context. In this paper, we propose a hierarchical context measurement network to construct corresponding measurement strategies for different hierarchical information, capturing comprehensive and powerful complementary knowledge from the context. Specifically, a feature-wise similarity measurement module is designed to calculate global cross-layer relationships between the middle features of the current band and those of the context, so as to explore the embedded middle features discriminatively through generated global dependencies. Furthermore, considering the pixel-wise correspondence between the reconstruction features and the super-resolved results, we propose a pixel-wise similarity measurement module for the complementary reconstruction features embedding, exploring detailed complementary information within the embedded reconstruction features by dynamically generating a spatially adaptive filter for each pixel. Experimental results reported on three benchmark hyperspectral datasets reveal that the proposed method outperforms other state-of-the-art peers in both visual and metric evaluations. Cong Wang 0033, Yuan Yuan 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Graph Convolutional Network With Self-Augmented Weights for Semi-Supervised Multi-View LearningabstractRecently, owing to the effectiveness in exploiting inherent connections between data in different views, graph-based deep learning approaches have gained widespread popularity in semi-supervised multi-view tasks. Generally, the existing approaches fuse the information from different views via the linear or nonlinear weight strategies, which distinguish the importance of different views by attributing their weights between $[{0, 1}]$ , i.e., some less important views are discarded since assigned with 0 and the pivotal views are not enhanced. However, these view-weighting strategies ignore the complementary information from the less important views. To address this issue, a superior-performing graph convolutional network (GCN) with self-augmented weights is proposed. The proposed self-augmented weight strategy is based on exponential series integration, which preserves the less important views and simultaneously strengthens the key views for multi-view fusion. Specifically, the designed weight strategy can adaptively preserve the complementary information from the less important views by assigning nonzero weights and strengthen the pivotal views by assigning higher weights based on exponential series integration. Besides, to further improve the model performance, an orthogonal constraint layer with a forced orthogonal weight is introduced, which is capable of making the representation more discriminative. Extensive experiments demonstrate the superiority of the proposed method. Hongyuan Zhang 0001, Yuan Yuan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Mono3DVG: 3D Visual Grounding in Monocular ImagesabstractWe introduce a novel task of 3D visual grounding in monocular RGB images using language descriptions with both appearance and geometry information. Specifically, we build a large-scale dataset, Mono3DRefer, which contains 3D object targets with their corresponding geometric text descriptions, generated by ChatGPT and refined manually. To foster this task, we propose Mono3DVG-TR, an end-to-end transformer-based network, which takes advantage of both the appearance and geometry information in text embeddings for multi-modal learning and 3D object localization. Depth predictor is designed to explicitly learn geometry features. The dual text-guided adapter is proposed to refine multiscale visual and geometry features of the referred object. Based on depth-text-visual stacking attention, the decoder fuses object-level geometric cues and visual appearance into a learnable query. Comprehensive benchmarks and some insightful analyses are provided for Mono3DVG. Extensive comparisons and ablation studies show that our method significantly outperforms all baselines. The dataset and code will be released. Yang Zhan 0007, Yuan Yuan 0001, Zhitong Xiong |
AAAI | 2 |
| 2024 | Cyclic Learning for Binaural Audio Generation and LocalizationabstractBinaural audio is obtained by simulating the biological structure of human ears, which plays an important role in artificial immersive spaces. A promising approach is to utilize mono audio and corresponding vision to synthesize binaural audio, thereby avoiding expensive binaural audio recording. However, most existing methods di-rectly use the entire scene as a guide, ignoring the corre-spondence between sounds and sounding objects. In this paper, we advocate generating binaural audio using fine-grained raw waveform and object-level visual information as guidance. Specifically, we propose a Cyclic Locating-and-Ul'mixing (CLUP) framework that jointly learns vi-sual sounding object localization and binaural audio generation. Visual sounding object localization establishes the correspondence between specific visual objects and sound modalities, which provides object-aware guidance to improve binaural generation performance. Meanwhile, the spatial information contained in the generated binaural au-dio can further improve the performance of sounding object localization. In this case, visual sounding object localization and binaural audio generation can achieve cyclic learning and benefit from each other. Experimental re-sults demonstrate that on the FAIR-Play benchmark dataset, our method is significantly ahead of the existing baselines in multiple evaluation metrics (STFTJ↓: 0.787 vs. 0.851, ENVJ↑: 0.128 vs. 0.134, WAVJ↓: 5.244 vs. 5.684, SNR↑: 7.546 vs. 7.044). Zhaojian Li 0002, Bin Zhao 0001, Yuan Yuan 0001 |
CVPR | 3 |
| 2024 | TAS: Personalized Text-guided Audio Spatialization
Zhaojian Li 0002, Bin Zhao 0001, Yuan Yuan 0001 |
ACM Multimedia | 3 |
| 2024 | A Descriptive Basketball Highlight Dataset for Automatic Commentary GenerationabstractThe emergence of video captioning makes it possible to automatically generate natural language description for a given video. However, generating detailed video descriptions that incorporate domain-specific information remains an unsolved challenge, holding significant research and application value, particularly in domains such as sports commentary generation. Moreover, sports event commentary goes beyond being a mere game report, it involves entertaining, metaphorical, and emotional descriptions. To promote the field of sports commentary automatic generation, in this paper, we introduce a novel dataset, the Basketball Highlight Commentary (BH-Commentary), comprising approximately 4K basketball highlight videos with groundtruth commentaries from professional commentators. In addition, we propose an end-to-end framework as a benchmark for basketball highlight commentary generation task, in which a lightweight and effective prompt strategy is designed to enhance alignment fusion among visual and textual features. Experimental results on the BH-Commentary dataset demonstrate the validity of the dataset and the effectiveness of the proposed benchmark for sports highlight commentary generation. Benhui Zhang 0001, Junyu Gao 0001, Yuan Yuan 0001 |
ACM Multimedia | 3 |
| 2024 | RRTrN: A lightweight and effective backbone for scene text recognition
Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
Expert Syst. Appl. | 3 |
| 2024 | Iterative Edge Enhancing Framework for Building Change DetectionabstractThe building change detection (BCD) task serves urban planning by monitoring land use. However, due to the complexity of remote-sensing images and high foreground–background similarity, it leads to inaccurate detection of building edge regions. Existing methods deal with this problem by fusing features of different layers. But the fusing operation cannot separate details information from the overall information of buildings, resulting in inaccurate detection of building edge area. To address the above challenges, we propose an iterative edge-enhancing framework (IEEF). The IEEF alleviates the building edge detection difficulty by densely implementing a detail semantic enhancement module (DSEM) in the decoding part. This module takes differential features between adjacent scales to explicitly represent the building edge information. Simultaneously, to deal with the class imbalance problem, a Density-Guided Sampling method dedicated to change detection is proposed to increase the proportion of positive samples during training. Our proposed method achieves state-of-the-art performance on the LEarning, VIsion and Remote sensing laboratory building Change Detection (LEVIR-CD) dataset and the Wuhan University (WHU) dataset and obtains accurate changed building edges. Shuai Song, Yuanlin Zhang 0003, Yuan Yuan 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Multi-view scene matching with relation aware feature perception
Bo Sun 0017, Ganchao Liu, Yuan Yuan 0001 |
Neural Networks | 3 |
| 2024 | Center-enhanced video captioning model with multimodal semantic alignment
Benhui Zhang 0001, Junyu Gao 0001, Yuan Yuan 0001 |
Neural Networks | 3 |
| 2024 | Detail-Preserving and Diverse Image Translation for Adverse Visual Object DetectionabstractThe effectiveness of object detection is significantly hampered in challenging nighttime or rainy scenarios. This is due to the severe domain shifts between daytime and adverse-visual images. Previous methods have demonstrated that using image-to-image translation methods for data augmentation can effectively address domain shifts, but they may still fail in preserving image objects when faced with extreme adverse images like rainy nights. In addition, achieving diversity in the generated results remains challenging. To this end, we propose a Progressive Adverse Image Translation (PAIT) framework that tackles domain shifts by generating diverse and detail-preserving images. The main contributions of this paper are as follows. 1) We propose a novel PAIT framework, which incorporates an iterative mapping module and a slicing layer. This framework enables the progressive generation of increasingly challenging images in a fine-to-coarse manner. 2) To preserve the details of the images, we innovatively introduce an iterative mapping module to generate smooth style transform curves. 3) To enhance the diversity of synthesized images, a simple but efficient end-to-end optimization method is proposed. 4) We found a strong correlation between the style diversity of augmented images and the performance of the detection model through a quantitative analysis, highlighting the crucial role of style diversity in enhancing the model’s generalizability. Our framework achieves state-of-the-art performance on multiple challenging visual datasets, surpassing the current state-of-the-art methods by 27%(+8.0AP). Moreover, our approach and modules can be easily extended to different detectors and other domain adaptation methods, making it a versatile solution for object detection in adverse visual environments. Our code will be available athttps://github.com/ssunguotu/Diverse-Aug. Guolong Sun, Zhitong Xiong, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Balanced Density Regression Network for Remote Sensing Object CountingabstractCounting objects in remote sensing is crucial for analyzing their distribution in images. Compared to surveillance perspectives, counting dense objects in remote sensing images is more challenging due to the smaller sizes of these targets. Recently, many methods utilize Gaussian convolution regression to estimate the count of dense objects in remote sensing images. However, most methods ignore the issue of regression imbalance inherent in Gaussian distribution, which is caused by the numerical differences in the center and edge regions. To tackle this challenge, we propose a Balanced Density Regression Network (BDRNet) to mitigate regression inaccuracies in Gaussian distributions due to numerical variances. Different from other methods, we divide the regression problem into two steps: first focusing on the regions of interest, then achieving precise regression. BDRNet consists of an Adaptive Kernel Weighting Attention (AKWA) mechanism and a Pixel-wise Occupancy Prediction (PwOE) module. Firstly, AKWA is designed to acquire accurate semantic feature information, which is obtained by learning the weights of dilated convolutions with different sizes of receptive fields. Secondly, the PwOE module applies Gaussian position embeddings to point labels to constrain the network to focus on the object region without increasing annotation cost. Finally, the integration of pixel-wise occupancy prediction features and kernel weighting features forms multi-layer cross-attention mechanisms, facilitating channel-level feature interaction and improving density regression predictions. Thus, the center and edge regions of the Gaussian kernel are treated equally, and the regression is balanced. Additionally, Extensive experiments on diverse datasets validate the effectiveness of the method, resulting in preferable performance. The code is available at: https://github.com/HotChieh/BDRNet. Haojie Guo, Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Contrastive Tokens and Label Activation for Remote Sensing Weakly Supervised Semantic SegmentationabstractIn recent years, there has been remarkable progress in Weakly Supervised Semantic Segmentation (WSSS), with Vision Transformer (ViT) architectures emerging as a natural fit for such tasks due to their inherent ability to leverage global attention for comprehensive object information perception. However, directly applying ViT to WSSS tasks can introduce challenges. The characteristics of ViT can lead to an over-smoothing problem, particularly in dense scenes of remote sensing images, significantly compromising the effectiveness of Class Activation Maps (CAM) and posing challenges for segmentation. Moreover, existing methods often adopt multi-stage strategies, adding complexity and reducing training efficiency. To overcome these challenges, a comprehensive framework CTFA (Contrastive Token and Foreground Activation) based on the ViT architecture for WSSS of remote sensing images is presented. Our proposed method includes a Contrastive Token Learning Module (CTLM), incorporating both patch-wise and class-wise token learning to enhance model performance. In patch-wise learning, we leverage the semantic diversity preserved in intermediate layers of ViT and derive a relation matrix from these layers and employ it to supervise the final output tokens, thereby improving the quality of CAM. In class-wise learning, we ensure the consistency of representation between global and local tokens, revealing more entire object regions. Additionally, by activating foreground features in the generated pseudo label using a dual-branch decoder, we further promote the improvement of CAM generation. Our approach demonstrates outstanding results across three well-established datasets, providing a more efficient and streamlined solution for WSSS. Code will be available at: https://github.com/ZaiyiHu/CTFA. Zaiyi Hu, Junyu Gao 0001, Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Statistical Texture Awareness Network for Hyperspectral Image ClassificationabstractThe distribution of ground objects in hyperspectral images predominantly reveals spatial indications of both order and disorder, encapsulating a wealth of texture information. This texture information encompasses not only local structural details but also global statistical priors of an image. Nevertheless, convolutional-neural-network-based methods for hyperspectral image classification (HIC) primarily use skip connections to incorporate shallow features abundant in texture information into deeper layers. They face challenges in effectively capturing the statistical properties of texture information, and the traditional method of modeling statistical attributes struggles to seamlessly integrate into parameter learning of convolutional neural networks (CNNs). To do so, this work proposes a statistical texture awareness network (STANet) for HIC. It achieves the exploration of learnable texture features. Through multilevel quantization and quantization encoding, a statistical texture learning module (STLM) is constructed to represent texture information from low-level features in a statistical manner. As a result, it augments the discriminatory power of such features. In addition, a complete feature fusion module (CFFM) is designed to intelligently combine multiscale contextual semantic and statistical texture features, thereby bolstering the discrimination of spectral-spatial ones. Experimental results reported for three public datasets demonstrate the superior performance of the proposed network over other peers. Mingxin Jin, Cong Wang 0033, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | VL-MFL: UAV Visual Localization Based on Multisource Image Feature LearningabstractObtaining the earth-fixed coordinates is a fundamental requirement for long-distance unmanned aerial vehicle (UAV) flight. Global navigation satellite systems are the most common location model, but their signals are susceptible to interference from obstacles and complex electromagnetic environments. To solve this issue, a visual localization framework based on multi-source image feature learning (VL-MFL) is proposed. In the proposed framework, the UAV is located by mapping airborne images to the satellite images with absolute coordinate positions. Firstly, for the heterogeneity issues caused by the different imaging environments of drone and satellite images, a lightweight Siamese network based on 3-D attention mechanism is proposed to extract the consistent features from the multi-source images. Secondly, to overcome the problem of inaccurate localization caused by the large receptive field of traditional convolutional neural networks, the cell-divided strategy is imported to strengthen the position mapping relationship of multi-source images features. Finally, based on similarity measurement, a confidence evaluation mechanism is established and a search region prediction method is proposed, which is effectively improved the accuracy and efficiency in matching localization. To evaluate the location performance of the proposed framework, several related methods are compared and analysed in details. The results on the real-world datasets indicate that the proposed method has achieved outstanding location accuracy and real-time performance. Ganchao Liu, Sihang Zhang, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Enhancing Unimodal Features Matters: A Multimodal Framework for Building ExtractionabstractIn recent years, deep learning and multi-modal data have substantially propelled the development of building extraction models. However, prevailing multi-modal methods are difficult to cope with two challenges: 1) modal laziness: the training error is minimized before the model has learned extensive uni-modal patterns; 2) modal imbalance: the backpropagation process is easily dominated by a certain modality. As a result, the uni-modal features learning is insufficient, leading to limited performance of the model when dealing with the intricate foreground and background contexts surrounding the buildings. In this paper, we deal with this problem from the perspective of algorithm and model evaluation. At the algorithmic level, we propose a Uni-modal Feature Enhancement (UFE) framework. Specifically, UFE is model-agnostic, comprising two distinct components: Adaptive Gradient Enhancement (AGE) for modal laziness and Consistency Constraint Loss (CCL) for modal imbalance. AGE dynamically modulates the original gradient by monitoring the representation effects of uni-modal features and multi-modal fusion features. CCL imposes mutual constraints on diverse modal branches at the semantic level to reconcile the optimization process. At the model evaluation level, a new metric, named Uni-modal Utilization Ratio (UUR), is presented to assess models through the learning efficacy of uni-modal features. The experimental results including the variants of UUR on two building extraction datasets demonstrate a substantial performance improvement by UFE. Moreover, UFE also exhibits its adaptability when integrated with various model components and its generalization on other multi-modal image-related tasks. Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Alignment and Fusion Using Distinct Sensor Data for Multimodal Aerial Scene ClassificationabstractMultimodal sensors offer a wealth of rich and diverse data, which is helpful for classify similar and complex aerial scenes. However, the heterogeneity of the data collected from different sensors brings a great challenge for alignment and fusion. For this, we present a multimodal aerial scene classification approach for extracting distinct modal information representations, realizing alignment and fusion of semantic information at both the data and feature levels. Firstly, an Adaptive Zero-Crossing Rate (AZCR) module is proposed to convert the sequential data into images, achieving alignment at the data level. This module is proficient at extracting temporal and frequency domain features from sequential data through adaptive parameter adjustments. Secondly, we propose a Multi-Modal Alignment and Fusion (MMAF) module to facilitate the alignment and fusion of distinct data, thereby achieving comprehensive modality integration at the feature level. Finally, the multimodal alignment loss function is designed to assess the alignment outcomes and constrain the training process. Our approach has been proven effective in accurately classifying aerial scenes, as demonstrated by the results of our experiments on two public datasets. The proposed method achieves 81.32% and 59.80% F1 score on the ADVANCE and URFC datasets. Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Dimensionally Unified Metric Model for Multisource and Multiview Scene MatchingabstractThe core challenge of multi-view scene matching is to effectively extract features from multi-view images, which is a key factor to achieve robust scene matching. This paper delves into methodologies for enhancing the multi-view robustness of scene matching, with the following key contributions: 1) This paper propose the metric feature consistency principle which emphasize the necessity of implementing accurate correspondence between drone and satellite images. To verify the principle, consistent feature enhancement for channel, spatial, and hybrid dimensions is explored. As a counterexample, the shuffle operation is used to break the consistency of dimension semantic information. Experiments show that the mAP is reduced by about 20% after breaking the consistency relationship of metric features. 2) Built upon the principle, this paper introduces a scene matching framework named DUMM (Dimensionally Unified Metric Model). Utilizing the multi-dimensional feature enhancement model, it effectively enhances the correspondence of the dual-branch features within the siamese network. 3) This paper is the first to introduce the feature dimension, frequency domain feature. By the upper and lower bounds of the cosine transform function, the negative effects of cross-view variations can be effectively mitigated, thus enhancing the overall robustness. The introduction of frequency features improves the mAP by 2.54% compared with the method of improving the dimensional consistency ofH×W×C, verifying the positive impact of the fusion of frequency domain features on enhancing the robustness of scene matching. Nevertheless, it remains imperative to adhere to the principle of feature dimension consistency across all additional frequency dimensions. Bo Sun 0017, Ganchao Liu, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Recreating Brightness From Remote Sensing Shadow AppearanceabstractShadow removal from remote sensing images is still an open issue. Recently, deep network training on unpaired data is preferable since corresponding ground truths of shadow images are not available in practice. Nevertheless, unsupervised shadow removal research for remote sensing imagery is limited by the scarcity of publicly available benchmarks. This paper proposes an unsupervised progressive network (UP-ShadowGAN) to jointly learn decoupled features for shadow removal and color transfer. UP-ShadowGAN explores the mapping between shadow and shadow-free domains through adversarial learning and cycle consistency constraint. In particular, we employ progressive learning to decompose the overall mapping process into more manageable shadow removal and color transfer steps. Specifically, the realistic illumination is restored by propagating spatial context between shadow and shadow-free nodes. Coupled with a multi-color space aggregation strategy, diverse color space representations alleviate color deviation caused by spatial inconsistency. More importantly, we contribute the first unpaired remote sensing shadow removal dataset (URSSR), which encourages future exploration. Extensive experiments demonstrate that UP-ShadowGAN competes favorably with state-of-the-art methods. The dataset and code are available at https://github.com/chi-kaichen/UP-ShadowGAN. Qi Wang 0009, Kaichen Chi, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Cross-Difference Semantic Consistency Network for Semantic Change DetectionabstractThe objective of Semantic Change Detection (SCD) is to discern intricate changes in land cover while simultaneously identifying their semantic categories. Prior research has shown that using multiple independent branches for the distinct tasks of change localization and semantic recognition is a reliable approach to solving the SCD problem. Nevertheless, conventional SCD architectures rely heavily on a high degree of consistency within the bi-temporal feature space when modeling difference features, inevitably resulting in false positives or missed alerts within change areas. In this paper, we introduce a SCD framework called the Cross-Differential Semantic Consistency (CdSC) network. CdSC is designed to mine deep discrepancies in bi-temporal instance features while preserving their semantic consistency. Specifically, the 3D-Cross-Difference module, incorporating 3D convolutions, explores the interaction of cross-temporal features, revealing inherent differences among various land features. Simultaneously, deep semantic representations are further utilized to enhance the local correlation of difference information, thereby improving the model’s discriminative capabilities within change regions. Incorporating principles from contrastive learning, a Semantic Co-Alignment loss is introduced to increase intra-class consistency and inter-class distinctiveness of dual-temporal semantic features, thereby addressing the challenges posed by semantic disparities. Extensive experiments on two SCD datasets demonstrate that CdSC outperforms other state-of-the-art SCD methods significantly in both qualitative and quantitative evaluations. The code and dataset are available at https://github.com/weiAI1996/CdSC. Qi Wang 0009, Kaichen Chi, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Neighbor Spectra Maintenance and Context Affinity Enhancement for Single Hyperspectral Image Super-ResolutionabstractSingle hyperspectral image super-resolution aims to improve the spatial resolution of a hyperspectral image without relying on auxiliary information. By taking advantage of the high similarity among neighbor bands, some recent methods have employed a recursive structure to super-resolve a hyperspectral image band-by-band. They are usually memory-efficient and perform well. However, they tend to introduce feedback information without distinction so as to weaken the utilization of complementary information in the context. Additionally, the spectral structure is inevitably destroyed when spatial information is extracted from neighbor bands, which hampers the effective exploration of spectral information in the subsequent process. To this end, we propose a two-stage network based on neighbor spectra maintenance and context affinity enhancement, which is composed of two sub-networks: neighbor network and context network. The former utilizes several neighbor bands to generate the neighbor spatial-spectral feature, incorporating a parallel processing scheme designed to reduce spectral distortion. Then we construct a relationship representation between the neighbor feature and feedback context information in the context network. By referring to the representation, the contents with higher complementarity will be highlighted in this stage. Experimental results on five public hyperspectral image datasets demonstrate that the proposed network not only outperforms state-of-the-art methods in terms of spatial reconstruction accuracy and spectral fidelity, but also requires less memory usage. Cong Wang 0033, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | HCNet: Hierarchical Feature Aggregation and Cross-Modal Feature Alignment for Remote Sensing Image CaptioningabstractRemote sensing image captioning aims to describe the crucial objects from remote sensing images in the form of natural language. The inefficient utilization of object texture and semantic features in images, along with the ineffective cross-modal alignment between image and text features, are the primary factors that impact the model to generate high-quality captions. To alleviate this trouble, this paper presents a network for remote sensing image captioning, namely HCNet, including hierarchical feature aggregation and cross-modal feature alignment. Specifically, a hierarchical feature aggregation module is proposed to obtain a comprehensive representation of vision features, which is beneficial for producing accurate descriptions. Considering the disparities between different modal features, we design a cross-modal feature interaction module in the decoder to facilitate feature alignment. It can fully utilize cross-modal features to localize critical objects. Besides, a cross-modal feature align loss is introduced to realize the alignment between image and text features. Extensive experiments show our HCNet can achieve satisfactory performance. Especially, we demonstrate significant performance improvements of +14.15% CIDEr score on NWPU datasets compared to existing approaches. The source code is publicly available at https://github.com/CVer-Yang/HCNet. Zhigang Yang 0002, Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Boosting Binary Object Change Detection via Unpaired Image Prototypes ContrastabstractBinary object change detection aims to monitor the evolution of the object of interest in a fixed region. Constructing a relevant dataset for deep learning models is strenuous. In the existing datasets, there is usually an imbalance between changed and unchanged samples, as well as a restricted diversity within the changed samples. Aiming at that, some methods utilize unpaired images used for object segmentation to generate pseudo-bitemporal images for change detection. However, due to the existence of the domain gap between different data sources, the model obtained by these methods can not well generalize to the real bitemporal images. Inspired by them but to avoid the domain difference, we explore how to directly use the unpaired images within a real change detection dataset to complement changed samples. In detail, a concise metric-based framework is designed, which consists of two branches, a projector and a predictor. The framework obtains the change map by computing the distance between the bitemporal embedding outputted by the projector. Meanwhile, instructed by an indirect semantic supervision module (ISSM) specially designed, the predictor can generate the semantic confidence map distinguishing the pixels in an image into two categories. Based on the output of the framework, an unpaired image prototype contrast module (UIPCM) is proposed. It enriches the diversity of the change samples for training by combining the prototypes in unpaired images at the feature level, leading to alleviating the imbalance between changed and unchanged samples. Besides, a dual margin contrastive loss (DMCL) is adopted during training. It can reduce the constraint on the consistency of bitemporal embedding in unchanged regions. The benefits and the superiority of the proposed method are demonstrated on two well-recognized datasets. The code is available at https://github.com/ptdoge/UIPC. Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Visual Consistency Enhancement for Multiview Stereo Reconstruction in Remote SensingabstractLearnable multiview stereo (MVS) aerial image depth estimation has obtained great success in 3-D digital urban reconstruction. Currently, most depth estimation methods in the large-scale sense heavily involve adapting the general MVS framework. However, these methods often overlook the cross-view interval and limited viewpoint inherent in aerial images data. In this article, we introduce an learning-based MVS method for aerial image depth estimation, which enhances visual consistency to address the insufficient accuracy caused by the characteristics of aerial image data, namely, AggrMVS. First, an optical flow-guided feature extraction module is introduced to map the dynamic relationship between reference and source images. It explicitly captures edge information of different depth components to guide the cost volume regularization. Second, a cross-view volume fusion module is proposed to enhance the interaction among reference volumes, further improving the aggregation ability of the source volume. Furthermore, AggrMVS achieves refined aerial image depth estimation results with a lightweight cascade architecture. Since low-altitude oblique aerial datasets currently lack, we reconstruct a multicategory synthetic aerial scene benchmark from general MVS datasets. The benchmark dataset is available athttps://github.com/ToscW/BlendedUAV. Experiments on public and proposed datasets confirm that AggrMVS outperforms other MVS depth estimation methods in terms of qualitative and quantitative aspects. Wei Zhang 0250, Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Single-Stream Extractor Network With Contrastive Pre-Training for Remote-Sensing Change CaptioningabstractRemote sensing (RS) image change captioning is a visual semantic understanding task that has received increasing attention. The change captioning methods are required to understand the visual information of the images and capture the most significant difference between them, then describe it in natural language. Most existing methods mainly focus on improving the difference feature encoder or language decoder, while ignoring the visual feature extractor. The current feature extractors suffer from several issues, including 1) domain gap between pre-training on single temporal natural images and downstream bi-temporal RS task, 2) limited difference feature modeling in the implicit single-stream network, and 3) high computational costs caused by extracting features for each temporal phase image under the dual-stream extractor. To address these issues, we propose a Single-stream Extractor Network (SEN). It consists of a single-stream extractor pre-trained on bi-temporal RS images using contrastive learning to mitigate the domain gap and high computational cost. Additionally, to improve feature modeling for difference information, we propose a shallow feature embedding (SFE) module and a cross attention guided difference (CAGD) module, which enhance the representation of temporal features and extract the difference features explicitly. Extensive experiments and visualizations demonstrate the effectiveness and advanced performance of SEN. The code and model weights are available at https://github.com/mrazhou/SEN. Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | FF-LPD: A Real-Time Frame-by-Frame License Plate Detector With Knowledge Distillation and Feature PropagationabstractWith the increasing availability of cameras in vehicles, obtaining license plate (LP) information via on-board cameras has become feasible in traffic scenarios. LPs play a pivotal role in vehicle identification, making automatic LP detection (ALPD) a crucial area within traffic analysis. Recent advancements in deep learning have spurred a surge of studies in ALPD. However, the computational limitations of on-board devices hinder the performance of real-time ALPD systems for moving vehicles. Therefore, we propose a real-time frame-by-frame LP detector focusing on real-time accurate LP detection. Specifically, video frames are categorized into keyframes and non-keyframes. Keyframes are processed by a deeper network (high-level stream), while non-keyframes are handled by a lightweight network (low-level stream), significantly enhancing efficiency. To achieve accurate detection, we design a knowledge distillation strategy to boost the performance of low-level stream and a feature propagation method to introduce the temporal clues in video LP detection. Our contributions are: (1) A real-time frame-by-frame LP detector for video LP detection is proposed, achieving a competitive performance with popular one-stage LP detectors. (2) A simple feature-based knowledge distillation strategy is introduced to improve the low-level stream performance. (3) A spatial-temporal attention feature propagation method is designed to refine the features from non-keyframes guided by the memory features from keyframes, leveraging the inherent temporal correlation in videos. The ablation studies show the effectiveness of knowledge distillation strategy and feature propagation method. Haoxuan Ding, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Image Process. | 3 |
| 2024 | An End-to-End Contrastive License Plate DetectorabstractAs a unique identity of vehicle, License Plate (LP) facilitates the intelligent transportation in many fields, such as traffic enforcement, intelligent transportation dispatching, etc. Recently, the LP detectors are trained by supervised learning which is directly guided by manual annotations and lacks the use of visual knowledge in image content, limiting the further development of detection performance. Inspired by the contrast and comparison in perception of human beings, a contrastive learning method is introduced into license plate detection task and we propose an end-to-end Contrastive License Plate Detector (CLPD). In CLPD, a special contrastive triad for contrastive learning is designed which aims to decouple the foregrounds and backgrounds. Based on this triad, a contrastive learning branch is introduced into the license plate detection pipeline to prompt the feature expression ability of backbone and extracting more discriminative features for detection. This contrastive learning branch is jointly trained with supervised learning branch for detection and it is only used in training, keeping the efficiency in inference. The experiment results show that the proposed CLPD improves the detection accuracy compared to baselines and other license plate detectors significantly on three datasets. The ablation studies further explore the potential of CLPD. In addition, the proposed CLPD has generalization to improve the performance on different baselines. And the visualization results in latent space verify our proposed CLPD aggregates features tightly and extracts discriminative features effectively. Haoxuan Ding, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Multi-Domain Adaptation for Motion DeblurringabstractMotion deblurring is an important topic in the field of image enhancement, which has widespread applications including video surveillance, object detection, etc. Many algorithms are designed for motion deblurring and achieve remarkable performance. However, mainstream motion blur datasets are collected under normal weather and illuminance conditions, i.e., normal domain, ignoring their variations. As a result, current methods perform poorly in dynamic real-world scenes. To address these issues, we study the work in two aspects. First, we collect the real-world motion blur dataset with a well-designed collection device from various angles, focal lengths, and street scenes. Considering its domain is single, it is augmented via a Domain Transfer Strategy (DTS) to construct a Multi-Domain dataset (MD dataset), expanding the domains of the collected dataset. Second, we propose a Multi-Domain Adaptive Deblur Network (MDADNet) with two modules. The one is the Domain Adaptation (DA) module that exploits domain invariant features to stabilize the performance of the MDADNet in multiple domains. The other is the Meta Deblurring (MDB) module that employs the auxiliary branch to enhance the deblurring ability. It also enables the MDADNet to update parameters during the testing stage, improving the generalizations of the MDADNet. Extensive experimental results demonstrate that the MD-trained methods significantly strengthen the motion deblurring ability in multiple domains. Particularly, the proposed MDADNet achieves state-of-the-art performance on the MD dataset and public motion blur datasets. Kai Zhuang, Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 3 |
| 2024 | Zoom Text DetectorabstractTo pursue comprehensive performance, recent text detectors improve detection speed at the expense of accuracy. They adopt shrink-mask-based text representation strategies, which leads to a high dependence of detection accuracy on shrink-masks. Unfortunately, three disadvantages cause unreliable shrink-masks. Specifically, these methods try to strengthen the discrimination of shrink-masks from the background by semantic information. However, the feature defocusing phenomenon that coarse layers are optimized by fine-grained objectives limits the extraction of semantic features. Meanwhile, since both shrink-masks and the margins belong to texts, the detail loss phenomenon that the margins are ignored hinders the distinguishment of shrink-masks from the margins, which causes ambiguous shrink-mask edges. Moreover, false-positive samples enjoy similar visual features with shrink-masks. They aggravate the decline of shrink-masks recognition. To avoid the above problems, we propose a zoom text detector (ZTD) inspired by the zoom process of the camera. Specifically, zoomed-out view module (ZOM) is introduced to provide coarse-grained optimization objectives for coarse layers to avoid feature defocusing. Meanwhile, zoomed-in view module (ZIM) is presented to enhance the margins recognition to prevent detail loss. Furthermore, sequential-visual discriminator (SVD) is designed to suppress false-positive samples by sequential and visual features. Experiments verify the superior comprehensive performance of ZTD. Chuang Yang 0003, Mulin Chen, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Text kernel calculation for arbitrary shape text detection
Xu Han 0019, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
Vis. Comput. | 3 |
| 2023 | Weakly-Supervised Scene-Specific Crowd Counting Using Real-Synthetic Hybrid DataabstractDue to the domain gap between the public large-scale datasets and actual scenes, the crowd counting models trained on the common datasets have a significant performance degradation when applying in practical applications. To address the above issue, one of the solution is to label additional data from the novel scenes, which is time-consuming and impractical for multiple scenes. Another solution is to utilize domain adaptation approaches to adapt a well-trained model to novel scenes. However, most of these approaches focus on appearance adaptation while the background and the crowd distribution is not adapted. In this paper, we propose a weakly-supervised method with real-synthetic hybrid data which only requires a small portion of unlabelled real images and auto-generated synthetic labelled images for training. First, the hybrid data is generated based on background from the real scene and random distributed synthetic persons. Second, an initialized counter is trained based on the hybrid data and the crowd distribution is predicted based on the predictions on real images. Then, a better crowd counter is trained based on new hybrid data generated from updated crowd distribution. The process is iterated until convergence. Extensive experiments demonstrate the effectiveness of the proposed method. Yaowu Fan, Jia Wan 0001, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 3 |
| 2023 | Optimal Kernel for Real-Time Arbitrary-Shaped Text DetectionabstractRecently, segmentation-based text detection methods develop rapidly, which achieve competitive accuracy and detection speed. However, these methods are hard to fit text instances accurately, which leads to the decrease of model performance. Meanwhile, the poor perception of the text center by the boundary pixels further affects the detection accuracy. We follow the issues and design an efficient framework for arbitrary-shaped text detection, which is constructed based on Optimal Kernel Representation (OKR) and Pixel Enhancement Module (PEM). Specifically, OKR is proposed to fit texts with optimal kernels. It erodes texts according to the corresponding geometric characteristics, which is simpler and more accurate compared with previous methods. PEM is used to enhance the perception of boundary pixels to the virtual character centers of text, thus improving the cohesion of the whole instance. Particularly, PEM only participates in the training process, which brings no extra computation costs to inference. Ablation experiments show the effectiveness of OKR and PEM. Comparisons on serveral benchmarks verify that our efficient detector is superior to the existing state-of-the-art (SOTA) methods. Haozhao Ma, Chuang Yang 0003, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 3 |
| 2023 | Difference Guided VHR Remote Sensing Image Change DetectionabstractVHR remote sensing images have abundant ground features and details, but it is a great challenge for machine understanding. The "same object with different spectral" problem caused by environment changes, such as seasonal alternation, bad weather and shadow, is the biggest challenge in multitemporal image change detection, which is more prominent in VHR images. For this problem, a novel difference guided VHR image change detection (DGCD) method is proposed in this paper. In the feature learning stage of DGCD model, difference features are used to guide the feature extraction to suit with the change detection task. In order to make the model focus on the change features, both of the spatial and channel attention mechanism are introduced. Finally, for the edge region of VHR image which is hard to be discriminated, a new edge enhanced loss function based on BCL loss is designed. Experiments on public datasets show the superiority of proposed DGCD method. It has good generalization ability in different classical challenging scenarios. Compared with the representative methods in recent years, the proposed DGCD method performs better on VHR image change detection. Jiukai Sun, Ganchao Liu, Xuelong Li 0001, Yuan Yuan 0001 |
ICASSP | 4 |
| 2023 | Multi-level Graph Contrastive Prototypical ClusteringabstractRecently, graph neural networks (GNNs) have drawn a surge of investigations in deep graph clustering. Nevertheless, existing approaches predominantly are inclined to semantic-agnostic since GNNs exhibit inherent limitations in capturing global underlying semantic structures. Meanwhile, multiple objectives are imposed within one latent space, whereas representations from different granularities may presumably conflict with each other, yielding severe performance degradation for clustering. To this end, we propose a novel Multi-Level Graph Contrastive Prototypical Clustering (MLG-CPC) framework for end-to-end clustering. Specifically, a Prototype Discrimination (ProDisc) objective function is proposed to explicitly capture semantic information via cluster assignments. Moreover, to alleviate the issue of objectives conflict, we introduce to perceive representations of different granularities within individual feature-, prototypical-, and cluster-level spaces by the feature decorrelation, prototype contrast, and cluster space consistency respectively. Extensive experiments on four benchmarks demonstrate the superiority of the proposed MLG-CPC against the state-of-the-art graph clustering approaches. Yuan Yuan 0001, Qi Wang 0009 |
IJCAI | 2 |
| 2023 | Bio-Inspired Audiovisual Multi-Representation Integration via Self-Supervised LearningabstractAudiovisual self-supervised representation learning has made significant strides in various audiovisual tasks. Existing methods mostly focus on single representation modeling between audio and visual modalities, ignoring the complex correspondence between them, resulting in the inability to execute cross-modal understanding in a more natural audiovisual scene. Several biological studies have shown that human learning is influenced by multi-layered synchronization of perception. To this end, inspired by biology, we argue to exploit the naturally existing relationships in audio and visual modalities to learn audiovisual representations under multilayer perceptual integration. Firstly, we introduce an audiovisual multi-representation pretext task that integrates semantic consistency, temporal alignment, and spatial correspondence. Secondly, we propose a self-supervised audiovisual multi-representation learning approach, which simultaneously learns the perceptual relationship between visual and audio modalities at semantic, temporal, and spatial levels. To establish fine-grained correspondence between visual objects and sounds, an audiovisual object detection module is proposed, which detects potential sounding objects by combining unsupervised knowledge at multiple levels. In addition, we propose a modality-wise loss and a task-wise loss to learn a subspace-orthogonal representation space that makes representation relations more discriminative. Finally, experimental results demonstrate that collectively understanding the semantic, temporal, and spatial correspondence between audiovisual modalities enables the model to perform better on downstream tasks such as sound separation, sound spatialization, and audiovisual segmentation. Zhaojian Li 0002, Bin Zhao 0001, Yuan Yuan 0001 |
ACM Multimedia | 3 |
| 2023 | Projection concept factorization with self-representation for data clustering
Chenyu Shao, Mulin Chen, Yuan Yuan 0001, Qi Wang 0009 |
Neurocomputing | 3 |
| 2023 | Uncertainty-Aware Graph Reasoning With Global Collaborative Learning for Remote Sensing Salient Object DetectionabstractRecently, fully convolutional networks (FCNs) have contributed significantly to salient object detection in optical remote sensing images (RSIs). However, owing to the limited receptive fields of FCNs, accurate and integral detection of salient objects in RSIs with complex edges and irregular topology is still challenging. Moreover, suffering from the low contrast and complicated background of RSIs, existing models often occur ambiguous or uncertain recognition. To remedy the above problems, we propose a novel hybrid modeling approach, i.e., uncertainty-aware graph reasoning with global collaborative learning (UG2L) framework. Specifically, we propose a graph reasoning pipeline to model the intricate relations among RSI patches instead of pixels, and introduce an efficient graph reasoning block (GRB) to build graph representations. On top of it, a global context block (GCB) with a linear attention mechanism is proposed to explore the multiscale and global context collaboratively. Finally, we design a simple yet effective uncertainty-aware loss (UAL) to enhance the model’s reliability for better prediction of saliency or non-saliency. Experimental and visual results on three datasets show the superiority of the proposed UG2L. Code is available at https://github.com/lyf0801/UG2L. Yanfeng Liu, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Edge Neighborhood Contrastive Learning for Building Change DetectionabstractBuilding change detection aims to identify the change in buildings in the same geographic area. Recently, many methods based on deep learning (DL) have achieved encouraging performance. However, some challenges remain in effectively exploiting the temporal–spatial correlation and achieving good discrimination in the neighborhood of the edge. To relieve these issues, we develop a selective attention module (SAM) to model the relationship between the semantic and the state (i.e., unchanged or changed) of the pixel, which is integrated into an existing metric learning-based architecture. Moreover, inspired by recent advances in contrastive learning, we present a novel edge neighborhood contrastive learning method to force the network to learn discriminative and compact features, leading to improving the accuracy of building change detection. Experimental results demonstrate that our method achieves competitive performance in terms of objective metrics and visual comparisons. Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Hierarchical Information Enhancing Detector for Remotely Sensed Object DetectionabstractFor the remote sensing object detection task, two-stage networks are widely used due to their high accuracy. These networks roughly predict the proposal regions containing potential objects. It is assumed in these methods that the sizes of these regions are close to that of the corresponding real object. However, this assumption is not always true. Consequently, the detector is affected by the size-unfitting proposal regions. In this letter, a hierarchical information enhancing detector (HIE-Det) is advocated to deal with this issue. First, the important semantic reinjection (ISR) module is proposed to mitigate the lack of object semantics caused by the size-unfitting problem. Compared with the normal detectors, the ISR module increases the proportion of information on objects and improves the effectiveness of the detection model. Second, the object boundary enhancing (OBE) module is proposed to improve the robustness of the regression. The OBE module introduces the convolutional branch stacking multigranularity grids for the same proposal region. Multiple granularity levels improve the robustness of the model to the different degrees of proposal size unfitting. Finally, to evaluate the effectiveness of the HIE-Det on multiscale datasets in a balanced and effective manner, we propose the scale-modulating scores (S-scores), i.e., scale-modulating average precision (sAP) and scale-modulating average recall (sAR). Compared with the other comprehensive scores, the S-scores are rid of the sample amounts and give priority to weaker indices. Implementing the proposed HIE-Det, S-scores {sAP, sAR} are, respectively, improved from {17.3%, 29.7%} to {34.8%, 43.2%}, reaching the state-of-the-art performance on the HRRSD dataset. These experiments verify the effectiveness of the proposed HIE-Det. Yuanlin Zhang 0003, Yuan Yuan 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | LSV-LP: Large-Scale Video-Based License Plate Detection and RecognitionabstractIn the past few decades, license plate detection and recognition (LPDR) systems have made great strides relying on Convolutional Neural Networks (CNN). However, these methods are evaluated on small and non-representative datasets that perform poorly in complex natural scenes. Besides, most of existing license plate datasets are based on a single image, while the information source in the actual application of license plates is frequently based on video. The mainstream algorithms also ignore the dynamic clue between consecutive frames in the video, which makes the LPDR system have a lot of room for improvement. In order to solve these problems, this paper constructs a large-scale video-based license plate dataset named LSV-LP, which consists of 1,402 videos, 401,347 frames and 364,607 annotated license plates. Compared with other data sets, LSV-LP has stronger diversity, and at the same time, it has multiple sources due to different collection methods. There may be multiple license plates in a frame, which is more in line with complex natural scenes. Based on the proposed dataset, we further design a new framework that explores the information between adjacent frames, called MFLPR-Net. In addition to these, we release the annotation tools for license plates or vehicles in videos. By evaluating the performance of MFLPR-Net and some mainstream methods, it is proved that the proposed model is superior to other LPDR systems.In order to be more intuitive, we put some samples on https://drive.google.com/file/d/1udqRddpJZMpTdHHQdwZRll6vaYALUiql/view?usp=sharingGoogle Drive. The whole dataset is available at https://github.com/Forest-art/LSV-LP. Qi Wang 0009, Xiaocheng Lu, Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Boosting One-Stage License Plate Detector via Self-Constrained Contrastive AggregationabstractScene Text Detection (STD) has applied in many fields successfully. One of the important applications of STD is License Plate Detection (LPD). As a unique identity of vehicle, License Plate (LP) facilitates the intelligent transportation in many fields, such as traffic enforcement, intelligent transportation dispatching, etc. However, there are many scene texts similar to LPs causing misjudgment of LP detector. To alleviate these disturbances, more discriminative features are necessary. In latent feature space, discriminative features should aggregate into a tight cluster to widen decision boundary. We assume three perspectives about how to aggregate features and boost feature expression. From these assumptions, a special contrastive triad is designed. Then, we propose a Self-Constrained Contrastive Aggregation (SCCA) method to lead the feature aggregation in latent space and boost the feature expression of backbone. The proposed SCCA is jointly trained with supervised learning for detection to improve the detection performance. The experiments show that our proposed SCCA prompts the baseline significantly and exceeds recent LP detectors, reaching 99.7 on both F1-score and AP on UFPR-ALPR dataset. Meanwhile, we compare the self-constrained contrastive learning with vanilla contrastive learning in experiments and visualize their LP features. The results show that our proposed SCCA reaches better performance and verifies our assumptions are reasonable. Haoxuan Ding, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Asymmetric Dual-Direction Quasi-Recursive Network for Single Hyperspectral Image Super-ResolutionabstractSingle hyperspectral image super-resolution aims to reconstruct a high-resolution hyperspectral image from a low-resolution one, which does not use any auxiliary images. For now, existing super-resolution methods often ignore the difference between the features of neighbor and non-neighbor spectral bands, leaving the feature exploration untargeted. As a result, the complementary information of such bands has not been effectively exploited. To do so, we propose an asymmetric dual-direction quasi-recursive network for single hyperspectral image super-resolution, which separately explores the features among neighbor and non-neighbor bands via forward and backward units. By considering the high similarity among neighbor bands, each forward unit thoroughly exploits spatial-spectral features among such bands through two kinds of correspondence aggregation modules. It also preserves a spectral structure by a spectral band grouping strategy and a spatial-spectral consistency module. Owning to the inconsecutive spectra among non-neighbor bands, backward units focus on extracting spatial features in such bands. With the aid of a global feature context fusion module, the information of global non-neighbor context and neighbor bands are adaptively fused, thus improving information completeness and complementarity. Experimental results reported for natural and remote sensing hyperspectral image datasets demonstrate the proposed network not only outperforms the state-of-the-art methods in terms of reconstruction quality and noise suppression, but also requires a smaller memory footprint. Cong Wang 0033, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Trinity-Net: Gradient-Guided Swin Transformer-Based Remote Sensing Image Dehazing and BeyondabstractHaze superimposes a veil over remote sensing images, which severely limits the extraction of valuable military information. To this end, we present a novel trinity model to restore realistic surface information by integrating the merits of both prior-based and deep learning-based strategies. Concretely, the critical insight of our Trinity-Net is to investigate how to incorporate prior information into CNNs and Swin Transformer for reasonable estimation of haze parameters. Then, haze-free images are obtained by reconstructing the remote sensing image formation model. Although Swin Transformer has shown tremendous potential in the dehazing task, which typically results in ambiguous details. We devise a gradient guidance module that naturally inherits structure priors of gradient maps, guiding the deep model to generate visually pleasing details. In light of the generality of image formation parameters, we successfully promote Trinity-Net to natural image dehazing and underwater image enhancement tasks. Notably, the acquisition of large-scale remote sensing hazy images and natural hazy images in military scenes is not feasible in practice. To bridge this gap, we construct aRemote Sensing Image Dehazing Benchmark(RSID) and aNatural Image Dehazing Benchmark(NID), including 1000 real-world hazy images with corresponding ground truth images, respectively. To our knowledge, this is the first exploration to develop dehazing benchmarks in the military field, alleviating the dilemma of data scarcity. Extensive experiments on three vision tasks illustrate the superiority of our Trinity-Net against multiple state-of-the-art methods. The datasets and code are available at https://github.com/chi-kaichen/Trinity-Net. Kaichen Chi, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | LGNet: Location-Guided Network for Road Extraction From Satellite ImagesabstractRoad connectivity is vital in road extraction for accurate vehicle navigation. However, the segmentation-based methods fail to model the connectivity resulting in broken road segments. Therefore, we propose a Location-Guided Network (LGNet) for promoting connectivity performance in a very effective and efficient way. Specifically, an auxiliary Road Location Prediction (RLP) task is designed to obtain global road connectivity information, which improves the performance of road segmentation. The RLP can predict the location coordinates of the whole roads with row anchors and column anchors. By aggregating the global location context to the segmentation branch with a location-guided decoder (LG-Decoder), the features can finally capture the connectivity of each road segment. Overall, LGNet has the following advantages: 1) The proposed RLP and LCG can plug into any encoder-decoder network and achieve an impressive performance. 2) High computational efficiency. In comparison with the multi-branch method, our proposed LGNet requires about 6× fewer GFLOPs. 3) The superior road connectivity performance. A series of experiments are conducted on two road extraction data sets (SpaceNet and DeepGlobe), confirming the effectiveness of the LGNet. Jingtao Hu, Junyu Gao 0001, Yuan Yuan 0001, Jocelyn Chanussot, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Exploring Hard Samples in Multiview for Few-Shot Remote Sensing Scene ClassificationabstractFew-shot remote sensing scene classification is of high practical value in real situations where data are scarce and annotated costly. The few-shot learner needs to identify new categories with limited examples, and the core issue of this assignment is how to prompt the model to learn transferable knowledge from a large-scale base dataset. Although current approaches based on transfer learning or meta-learning have achieved significant performance on this task, there are still two problems to be addressed: (i) as an essential characteristic of remote sensing images, spatial rotation insensitivity surprisingly remains largely unexplored; (ii) the high distribution uncertainty of hard samples reduces the discriminative power of the model decision boundary. Stimulated by these, we propose a corresponding end-to-end framework termed a Hard Sample Learning (HSL) and Multi-view Integration (MI) Network (HSL-MINet). First, the MI module contains a pretext task introduced to guide the knowledge transfer, and a multiview-attention mechanism used to extract correlational information across different rotation views of images. Second, aiming at increasing the discrimination of the model decision boundary, the HSL module is designed to evaluate and select hard samples via a class-wise adaptive threshold strategy, and then decrease the uncertainty of their feature distributions by a devised triplet loss. Extensive evaluations on NWPU-RESISC45, WHU-RS19, and UCM datasets show that the effectiveness of our HSL-MINet surpasses the former state-of-the-art approaches. Yuyu Jia, Junyu Gao 0001, Wei Huang 0068, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Holistic Mutual Representation Enhancement for Few-Shot Remote Sensing SegmentationabstractFew-shot segmentation endeavors to utilize a minimal amount of annotated samples (support) to guide the segmentation of unseen objects (query). Previous techniques primarily employ asupport-to-queryparadigm, neglecting to sufficiently leverage the mutual representation between query and support images, which leaves models suffering from intra-class variations and background interference in remote sensing images. This paper proposes a Holistic Mutual Representation Enhancement (HMRE) method to bridge these gaps. First, a Dual Activation (DA) module is devised to establish information symmetry between the two branches and forms the foundation for mutual representation enhancement. Subsequently, the holistic mutual enhancement is jointly constructed by the Global Semantic (GS) and Spatial Dense (SD) mutual enhancement modules. In the prediction stage for segmentation, we integrate the enhanced mutual representation into the Mutual-Fusion Decoder to activate the homologous object regions bidirectionally. To expedite the replication of investigation in this task, we further create a corresponding benchmark Flood-3i. The whole dataset is attainable at https://drive.google.com/drive/folders/1FMAKf2sszoFKjq0UrUmSLnJDbwQSpfxR. Extensive experiments on two benchmarks iSAID-5i and Flood-3i demonstrate the superiority of our proposed method, which also sets a new state-of-the-art. Yuyu Jia, Junyu Gao 0001, Wei Huang 0068, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Dual-Field-of-View Context Aggregation and Boundary Perception for Airport Runway Extraction
Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | RGB-Induced Feature Modulation Network for Hyperspectral Image Super-ResolutionabstractSuper-resolution (SR) is one of the powerful techniques to improve image quality for low-resolution (LR) hyperspectral image (HSI) with insufficient detail and noise. Traditional methods typically perform simple cascade or addition during the fusion of the auxiliary high-resolution RGB and LR HSI. As a result, the abundant HR RGB details are not utilized as a priori information to enhance the HSI feature representation, leaving room for further improvements. To address this issue, we propose an RGB-induced feature modulation network for HSI SR (IFMSR). Considering that similar patterns are common in images, a multi-corresponding patch aggregation is designed to globally assemble this contextual information, which is beneficial for feature learning. Besides, to adequately exploit plentiful HR RGB details, an RGB-induced detail enhancement (RDE) module and a deep cross-modality feature modulation (CFM) module are proposed to transfer the supplementary materials from RGB to HSI. These modules can provide a more direct and instructive representation, leading to further edge recovery. Experiments on several datasets demonstrate that our approach achieves comparable performance under more realistic degradation condition. Our code is publicly available at https://github.com/qianngli/IFMSR. Qiang Li 0042, Maoguo Gong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Tensor-Based Hyperspectral Anomaly Detection Method Under Prior Physical ConstraintsabstractEfficient and precise modeling of the background to accurately identify anomalies is the cornerstone to hyperspectral anomaly detection. Hyperspectral image (HSI) can be regarded as 3-D cube data, which contains both spatial and spectral information. The data are converted into 2-D matrices for processing in most existing method, which loses a large amount of structural information. In addition, it is difficult to construct a model with strong representation ability without enough prior knowledge constraints, and the modeling of the background is easily polluted by anomalies. This article introduces a tensor-based hyperspectral anomaly detection method that takes into account prior physical constraints as a solution to the aforementioned problems. The proposed method uses a tensor representation of the image, which preserves its geometrical properties and adheres to fundamental physical principles. After separating them from the image tensor, the background and anomaly tensors are treated separately. For the background tensor, we introduce segmented smoothness constraints in both spatial and spectral dimensions by applying linear total variation (TV) norm regularization. This can improve resistance to complicated backgrounds and lessen the introduction of extra noise when the background is restored. A low-rank constraint based on image eigenvalues is intended to generate a more realistic background model in its spatial dimension, making the method more sensitive to tiny anomalies. For the anomaly tensor, there is sparsity in its spectral dimension, and the background is effectively separated from the anomaly by$l_{1} $-norm constraints. Eventually, the anomaly detection map is decided by the anomaly tensor computed iteratively. Comprehensive experiments on several genuine and simulated datasets show that the proposed method performs significantly better at anomaly detection than the state-of-the-art methods. Xin Li 0188, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Multiscale Factor Joint Learning for Hyperspectral Image Super-ResolutionabstractHyperspectral image super-resolution (SR) using auxiliary RGB image has obtained great success. Currently, most methods respectively train single model to handle different scale factors, which may lead to the inconsistency of spatial and spectral contents when converted to the same size. In fact, the manner ignores the exploration of potential interdependence among different scale factors in single model. To this end, we propose a multi-scale factor joint learning for hyperspectral image super-resolution (MulSR). Specifically, to take advantage of the inherent priors of spatial and spectral information, a deep architecture using single scale factor is designed by terms of symmetrical guided encoder (SGE) to explore the hyperspectral image and RGB image. Considering that there are obvious differences in texture details at various scale factors, another architecture is proposed which is basically the same as above, except that its scale factor is larger. On this basis, a multi-scale information interaction (MII) unit is modeled between two architectures by a direction-aware spatial context aggregation (DSCA) module. Besides, the contents generated by the model with multi-scale factor are combined to build a learnable feedback compensation correction (LFCC). The difference is fed back to the architecture with large scale factor, forming an interactive feedback joint optimization pattern. This calibrates the representation of spatial and spectral contents in the reconstruction process. Experiments on synthetic and real datasets demonstrate that our MulSR shows superior performance in terms of qualitative and quantitative aspects. Our code is publicly available at https://github.com/qianngli/MulSR. Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Distilling Knowledge From Super-Resolution for Efficient Remote Sensing Salient Object DetectionabstractCurrent state-of-the-art remote sensing salient object detectors always require high-resolution spatial context to ensure excellent performance, which incurs enormous computation costs and hinders real-time efficiency. In this work, we propose a universal super-resolution assisted learning (SRAL) framework to boost performance and accelerate the inference efficiency of existing approaches. To this end, we propose to reduce the spatial resolution of the input remote sensing images (RSIs), which is model-agnostic, and can be applied to existing algorithms without extra computation cost. Specifically, a transposed saliency detection decoder (TSDD) is designed to upsample interim features progressively. On top of it, an auxiliary super-resolution decoder (ASRD) is proposed to build a multitask learning (MTL) framework to investigate an efficient complementary paradigm of saliency detection and super-resolution. Furthermore, a novel task-fusion guidance module (TFGM) is proposed to effectively distill domain knowledge from the super-resolution auxiliary task to the salient object detection task in optical RSIs. The presented ASRD and TFGM can be omitted in the inference phase without any extra computational budget. Extensive experiments on three datasets show that the presented SRAL with 224×224 input is superior to more than 20 algorithms. Moreover, it can be successfully generalized to existing typical networks with significant accuracy improvements in a parameter-free manner. Codes and models are available at https://github.com/lyf0801/SRAL. Yanfeng Liu, Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Transcending Pixels: Boosting Saliency Detection via Scene Understanding From Aerial ImageryabstractExisting remote sensing image salient object detection (RSI-SOD) methods widely perform object-level semantic understanding with pixel-level supervision, but ignore the image-level scene information. As a fundamental attribute of RSIs, the scene has a complex intrinsic correlation with salient objects, which may bring hints to improve saliency detection performance. However, existing RSI-SOD datasets lack both pixel- and image-level labels, and it is non-trivial to effectively transfer the scene domain knowledge for more accurate saliency localization. To address these challenges, we first annotate the image-level scene labels of three RSI-SOD datasets inspired by remote sensing scene classification. On top of it, we present a novel scene-guided dual-stream network (SDNet), which can perform cross-task knowledge distillation from the scene classification to facilitate accurate saliency detection. Specifically, a scene knowledge transfer module (SKTM) and a conditional dynamic guidance module (CDGM) are designed for extracting saliency key area as spatial attention from the scene subnet and guiding the saliency subnet to generate scene-enhanced saliency features, respectively. Finally, an object contour awareness module (OCAM) is introduced to enable the model to focus more on irregular spatial details of salient objects from the complicated background. Extensive experiments reveal that our SDNet outperforms over 20 state-of-the-art algorithms on three datasets. Moreover, we prove that the proposed framework is model-agnostic, and its extension to six baselines can bring significant performance benefits. Code will be available at https://github.com/lyf0801/SDNet. Yanfeng Liu, Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Enhancing Prospective Consistency for Semisupervised Object Detection in Remote-Sensing ImagesabstractDeep learning-based object detection has recently played a vital role in both computer vision and Earth observation communities. However, the performance of modern object detectors is highly limited by the quantity and quality of manually labeled training samples. Furthermore, compared to object detection in natural scenes, Remote Sensing Object Detection (RSOD) faces two specific critical challenges. 1) Densely arranged instances: geospatial objects tend to be densely packed in remote sensing scenarios. 2) Large variations in object scale: the wide field of the bird’s eye view leads to dramatic variations in object scale across various categories. The above issues bring significant difficulties to attaining manual annotations for deep learning-based RSOD. To this end, in this paper, we turn our attention from fully-supervised RSOD to semi-supervised RSOD, and propose a novel framework based on the teacher-student paradigm, namely Prospective Consistent Teacher (PCT), which includes three crucial components,i.e., Weighted Dense-Proposal Learning (WDPL), Mean-Consistency-based Proposal Pruning (MCPP), and EM-based Fitting Policy (EFP). Specifically, WDPL re-weights the dense proposals with box confidences, while MCPP ranks the student proposals with consistency analysis to select discriminative and consistent boxes. EFP can automatically set thresholds for pseudo labels and improve the consistent information of the teacher network. Extensive experimental results on two challenging public datasets,i.e., DOTA and DIOR, have demonstrated the reduced reliance of our proposed method on large amounts of labeled data for the task of RSOD. Jinhao Shen, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | F3-Net: Multiview Scene Matching for Drone-Based Geo-LocalizationabstractScene matching involves establishing a mapping relationship between heterogeneous images, which is crucial for drone visual geo-localization. However, it poses a significant challenge for multi-view images such as those captured by drones and satellites. To address this issue, this paper proposes an end-to-end geo-localization framework named F3-Net for calculating the similarity of multi-source and multi-view images. The key contributions of F3-Net are as follows: 1) The Split and Fusion (SF) module is designed to fully exploit the features through the global self-attention mechanism. 2) To improve the multi-view semantic features, a Target Feature Enhancement (TFE) module is introduced, based on the principle of invariance target semantic consistency. 3) After multi-view feature learning, a Feature Alignment and Unity (FAU) module with Earth Mover distance is used to calculate the similarity of non-aligned features. F3-Net fully exploits the multi-source image feature correspondence and multi-view image semantic consistency. Different from the traditional siamese network, the features of multi-view images are regarded as probability distribution, so F3-Net can quantify and eliminate the feature differences of multi-view images in the learning process. Experiments show that F3-Net can effectively overcome multi-view changes and achieve high accuracy on University-1652 dataset. Bo Sun 0017, Ganchao Liu, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | NACAD: A Noise-Adaptive Context-Aware Detector for Remote Sensing Small ObjectsabstractSmall object detection in remote sensing faces significant challenges such as their offset-sensitivity caused by the small area coverage, the dim targets in images, and their vulnerability to complex backgrounds, which often result in missed detections and false alarms. In this work, we propose aNoise-Adaptive Context-Aware Detector(NACAD) to alleviate the above problems, which mainly consists of a region proposal network withNoise Adaptive Module(NAM), aContext Aware Module(CAM) and aPosition Refined Module(PRM). The main contributions are threefold: 1) We leverage the information around small objects as positive-incentive noise (also known as π-noise), through enlarging the range of small objects by NAM, more anchors of them are preserved as positive samples, thus stimulating the model to detect small objects. 2) The CAM is designed to provide multiple observation perspectives and abundant contextual representations for the enhancement of object features. 3) To reduce the interference of pure noise in the complicated backgrounds around small objects, the spatial calibration along two coordinate axes is devised by PRM to optimally use information beyond object regions. The effectiveness of our proposed detector, particularly on small objects, has been validated by the experiments on two public datasets, ITCVD and HRRSD. In particular, the NAM improves the recall of small objects, CAM enhances small object features, and PRM helps address the pure noise in complicated backgrounds around small objects. Yuan Yuan 0001, Yiru Zhao, Dandan Ma |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Parameter-Efficient Transfer Learning for Remote Sensing Image-Text RetrievalabstractVision-and-language pre-training (VLP) models have experienced a surge in popularity recently. By fine-tuning them on specific datasets, significant performance improvements have been observed in various tasks. However, full fine-tuning of VLP models not only consumes a significant amount of computational resources but also has a significant environmental impact. Moreover, as remote sensing (RS) data is constantly being updated, full fine-tuning may not be practical for real-world applications. To address this issue, in this work, we investigate the parameter-efficient transfer learning (PETL) method to effectively and efficiently transfer visual-language knowledge from the natural domain to the RS domain on the image-text retrieval task. To this end, we make the following contributions. 1) We construct a novel and sophisticated PETL framework for the RS image-text retrieval (RSITR) task, which includes the pretrained CLIP model, a multimodal remote sensing adapter, and a hybrid multi-modal contrastive (HMMC) learning objective; 2) To deal with the problem of high intra-modal similarity in RS data, we design a simple yet effective HMMC loss; 3) We provide comprehensive empirical studies for PETL-based RS image-text retrieval. Our results demonstrate that the proposed method is promising and of great potential for practical applications. 4) We benchmark extensive state-of-the-art PETL methods on the RSITR task. Our proposed model only contains 0.16M training parameters, which can achieve a parameter reduction of 98.9% compared to full fine-tuning, resulting in substantial savings in training costs. Our retrieval performance exceeds traditional methods by 7-13% and achieves comparable or better performance than full fine-tuning. This work can provide new ideas and useful insights for RS vision-language tasks. Yuan Yuan 0001, Yang Zhan 0007, Zhitong Xiong |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing DataabstractIn this article, we introduce the task of visual grounding for remote sensing data (RSVG). RSVG aims to localize the referred objects in remote sensing (RS) images with the guidance of natural language. To retrieve rich information from RS imagery using natural language, many research tasks, such as RS image visual question answering, RS image captioning, and RS image–text retrieval, have been investigated a lot. However, the object-level visual grounding on RS images is still underexplored. Thus, in this work, we propose to construct the dataset and explore deep learning models for the RSVG task. Specifically, our contributions can be summarized as follows. First, we build the new large-scale benchmark of RSVG based on detection in optical remote sensing (DIOR) dataset, termed DIOR-RSVG, to fully advance the research of RSVG. This new dataset includes image/expression/box triplets for training and evaluating visual grounding models. Second, we benchmark extensive state-of-the-art (SOTA) natural image visual grounding methods on the constructed DIOR-RSVG dataset, and some insightful analyses are provided based on the results. Third, a novel transformer-based multigranularity visual language fusion (MGVLF) module is proposed. Remotely sensed images are usually with large-scale variations and cluttered backgrounds. To deal with the scale-variation problem, the MGVLF module takes advantage of multiscale visual features and multigranularity textual embeddings to learn more discriminative representations. To cope with the cluttered background problem, MGVLF adaptively filters irrelevant noise and enhances salient features. In this way, our proposed model can incorporate more effective multilevel and multimodal features to boost performance. This work can provide useful insights for developing better RSVG models. Yang Zhan 0007, Zhitong Xiong, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Difference-Guided Aggregation Network With Multiimage Pixel Contrast for Change DetectionabstractChange detection is a critical task in remote sensing to monitor the state of the surface on Earth. This field has been dominated by deep learning-based methods recently. Many models that model the temporal-spatial correlation in bitemporal images through the non-local interaction between bitemporal features achieve impressive performance. However, under complex scenes including multiple change types or weakly discriminate objects, they suffer from achieving discriminative fusion of information due to the weak semantic discrimination of the bitemporal representations. Aiming at this problem, a difference-guided aggregation network (DGANet) is proposed, where two key modules are injected, i.e., a difference-guided aggregation module (DGAM) and a weighted metric module (WMM). The bitemporal features in DGAM are aggregated with the guidance of their differences, which focuses on their change relevance and relaxes their semantic distinction. Therefore, the fused features are change-relevant and discriminative. WMM aims to achieve adaptive distance computation between the bitemporal features by dynamic feature attention in different dimensions. It is helpful to suppress the pseudo-changes. Besides, a change magnitude contrastive loss (CMCL) is introduced to employ the dependency of bitemporal pixels in different bitemporal images, which further enhances the representation quality of the model. Meanwhile, it is further extended in this work. The effectiveness of the three improvements is demonstrated by extensive ablation studies. The results on three datasets widely used illustrate that our method achieves satisfactory performance. Qiang Li 0042, Yanling Miao, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Crowd Localization From Gaussian Mixture Scoped Knowledge and Scoped TeacherabstractCrowd localization is to predict each instance head position in crowd scenarios. Since the distance of pedestrians being to the camera are variant, there exists tremendous gaps among scales of instances within an image, which is called the intrinsic scale shift. The core reason of intrinsic scale shift being one of the most essential issues in crowd localization is that it is ubiquitous in crowd scenes and makes scale distribution chaotic. To this end, the paper concentrates on access to tackle the chaos of the scale distribution incurred by intrinsic scale shift. We propose Gaussian Mixture Scope (GMS) to regularize the chaotic scale distribution. Concretely, the GMS utilizes a Gaussian mixture distribution to adapt to scale distribution and decouples the mixture model into sub-normal distributions to regularize the chaos within the sub-distributions. Then, an alignment is introduced to regularize the chaos among sub-distributions. However, despite that GMS is effective in regularizing the data distribution, it amounts to dislodging the hard samples in training set, which incurs overfitting. We assert that it is blamed on the block of transferring the latent knowledge exploited by GMS from data to model. Therefore, a Scoped Teacher playing a role of bridge in knowledge transform is proposed. What' s more, the consistency regularization is also introduced to implement knowledge transform. To that effect, the further constraints are deployed on Scoped Teacher to derive feature consistence between teacher and student end. With proposed GMS and Scoped Teacher implemented on four mainstream datasets of crowd localization, the extensive experiments demonstrate the superiority of our work. Moreover, comparing with existing crowd locators, our work achieves state-of-the-art via F1-measure comprehensively on four datasets. Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Image Process. | 3 |
| 2023 | Reinforcement Shrink-Mask for Text DetectionabstractExisting real-time text detectors reconstruct text contours by shrink-masks only. Though they simplify the framework and can make the model run fast, the strong dependence on shrink-masks leads to unreliable detection results (e.g., miss detection and overdetection). Moreover, these methods ignore the information from surrounding pixels, which causes sensitive shrink-masks and accelerates the reliability decline of detection results. Considering the above problems, we construct an effective and efficient text detection network, termed as Reinforcement Shrink-Mask for Text Detection (RSMTD), which strengthens the model's ability to recognize texts while enjoying a high detection speed. Specifically, an effective text representation strategy (Reinforcement Shrink-Mask, RSM) is designed to decouple texts and shrink-masks. RSM builds texts through shrink-masks and reinforcement offsets to ensure stable detection results encountering shrink-masks that deviate from the ground-truth. It is worth noting that reinforcement offsets can force our method to focus on the foreground shapes to bring precise shrink-mask edges. For the robustness improvement of shrink-masks, Super-pixel Window (SPW) is proposed to encourage RSMTD to utilize the surroundings of each pixel to predict shrink-masks. Particularly, SPW treats the interval regions between texts and shrink-masks as background, which helps to suppress interval regions and to avoid text adhesion. Moreover, a lightweight feature merging branch is constructed to further accelerate the inference process. As demonstrated in the experiments, our method is superior to existing state-of-the-art (SOTA) methods in both detection accuracy and speed on multiple benchmarks. Chuang Yang 0003, Mulin Chen, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 3 |
| 2023 | Text Growing on LeafabstractIrregular-shaped texts bring challenges to Scene Text Detection (STD). Although existing regression-based approaches achieve comparable performances, they fail to cover some highly curved ribbon-like text lines. Inspired by morphology, we found that the leaf vein can easily cover various geometries. Specifically, lateral and thin veins are emitted to margin along main vein gradually with the leaf growth. This process can decompose a concave object into consecutive convex regions, which are easier to fit. Hence, the leaf vein is suitable for representing highly curved texts. Considering the aforementioned advantage, we design a leaf vein-based text representation method (LVT), where text contour is treated as leaf margin and represented through main, lateral, and thin veins. We further construct a detection framework based on LVT, namely LeafText. In the text reconstruction stage, LeafText simulates the leaf growth process to rebuild text contours. It grows main veins in Cartesian coordinates to locate texts roughly at first. Then, lateral and thin veins are generated along the main vein growth direction in polar coordinates. They are responsible for generating the coarse contour and refining it, respectively. Meanwhile, Multi-Oriented Smoother (MOS) is designed to smooth the main vein for ensuring reliable growth directions of lateral and thin veins. Additionally, a global incentive loss is proposed to enhance the predictions of lateral and thin veins. Ablation experiments demonstrate LVT can fit irregular-shaped texts precisely and verify the effectiveness of MOS and global incentive loss. Comparisons show that LeafText is superior to existing state-of-the-art (SOTA) methods on MSRA-TD500, CTW1500, Total-Text, and ICDAR2015 datasets. Chuang Yang 0003, Mulin Chen, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 3 |
| 2023 | Domain-Adaptive Crowd Counting via High-Quality Image Translation and Density ReconstructionabstractRecently, crowd counting using supervised learning achieves a remarkable improvement. Nevertheless, most counters rely on a large amount of manually labeled data. With the release of synthetic crowd data, a potential alternative is transferring knowledge from them to real data without any manual label. However, there is no method to effectively suppress domain gaps and output elaborate density maps during the transferring. To remedy the above problems, this article proposes a domain-adaptive crowd counting (DACC) framework, which consists of a high-quality image translation and density map reconstruction. To be specific, the former focuses on translating synthetic data to realistic images, which prompts the translation quality by segregating domain-shared/independent features and designing content-aware consistency loss. The latter aims at generating pseudo labels on real scenes to improve the prediction quality. Next, we retrain a final counter using these pseudo labels. Adaptation experiments on six real-world datasets demonstrate that the proposed method outperforms the state-of-the-art methods. Junyu Gao 0001, Tao Han 0002, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Single-Shot Balanced Detector for Geospatial Object DetectionabstractGeospatial object detection is an essential task in remote sensing community. One-stage methods based on deep learning have faster running speed but cannot reach higher detection accuracy than two-stage methods. In this paper, to achieve excellent speed/accuracy trade-off for geospatial object detection, a single-shot balanced detector is presented. First, a balanced feature pyramid network (BFPN) is designed, which can balance semantic information and spatial information between high-level and shallow-level features adaptively. Second, we propose a task-interactive head (TIH). It can reduce the task misalignment between classification and regression. Extensive experiments show that the improved detector obtains significant detection accuracy with considerable speed on two benchmark datasets. Yanfeng Liu, Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 3 |
| 2022 | BiP-Net: Bidirectional Perspective Strategy Based Arbitrary-Shaped Text Detection NetworkabstractDetecting irregular-shaped text instances is the main challenge for text detection. Existing approaches can be roughly divided into top-down and bottom-up perspective methods. The former encodes text contours into unified units, which always fails to fit highly curved text contours. The latter represents text instances by a number of local units, where the complicated network and post-processing lead to slow detection speed. In this paper, to detect arbitrary-shaped text instances with high detection accuracy and speed simultaneously, we propose a Bidirectional Perspective strategy based Network (BiP-Net). Specifically, a new text representation strategy is proposed to represent text contours from a topdown perspective, which can fit highly curved text contours effectively. Moreover, a contour connecting (CC) algorithm is proposed to avoid the information loss of text contours by rebuilding interval contours from a bottom-up perspective. The experimental results on MSRA-TD500, CTW1500, and ICDAR2015 datasets demonstrate the superiority of BiP-Net against several state-of-the-art methods. Chuang Yang 0003, Mulin Chen, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 3 |
| 2022 | ACP: Adaptive Channel Pruning for Efficient Neural NetworksabstractIn recent years, deep convolutional neural networks have achieved amazing results on multiple tasks. However, these complex network models often require significant computation resources and energy costs, so that they are difficult to deploy to power-constrained devices, such as IoT systems, mobile phones, embedded devices, etc. Aforementioned challenges can be overcome through model compression like network pruning. In this paper, we propose an adaptive channel pruning module (ACPM) to automatically adjust the pruning rate with respect to each channel, which is more efficient to prune redundant channel parameters, as well as more robust to datasets and backbones. With one-shot pruning strategy design, the model compression time can be saved significantly. Extensive experiments demonstrate that ACPM makes tremendous improvement on both pruning rate and accuracy, and also achieves the state-of-the-art results on a series of different networks and benchmarks. Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2022 | A Multi-Source Image Matching Network for UAV Visual LocationabstractVisual localization is an important but challenging task for unmanned aerial vehicles (UAV). Matching real-time UAV orthophotos to pre-existing georeferenced satellite images is the key problem for this task. However, UAV and satellite images are inconsistent in image styles, perspectives, and times. In this paper, a new fully convolutional siamese network is proposed to extract similar features for multi-source images. The Squeeze-and-Excitation structure is integrated into the densely connected network to adapt to multi-scale features and the texture differences of different regions. Besides, a loss function with a progressive sampling strategy is utilized to mine the similarity of matching multi-source images and improve the description compactness among dimensions. Extensive experimental results with in-depth analysis are provided, which indicate that the proposed framework can significantly improve the matching performance of the learned descriptor. Ganchao Liu, Yuan Yuan 0001 |
ICIP | 3 |
| 2022 | Efficient Deblurring Via High-Frequency and Low-Frequency Information FusionabstractThe distortion of high-frequency information is the most fundamental problem of dynamic scene blur, which leads to the degradation of image quality. However, most deep-based methods fail to show satisfactory results because of ignoring the importance of image structural information (high-frequency and low-frequency preception) in deblurring. In this paper, we propose a high-frequency and low-frequency information fusion deblurring network (HLFNet) that uses edge perception as a guide. The proposed HLFNet consists of the high-frequency information network (HFNet) and the low-frequency information network (LFNet). Besides, we adopt the proposed multi-scale atrous convolution (MSA) block into LFNet, which can effectively reduce the number of model parameters while expanding the receptive fields. Extensive experiments show that the proposed model can achieve state-of-the-art results with smaller parameters and shorter inference time on the public datasets. Ruilong Lu, Yuan Yuan 0001, Qi Wang 0009 |
ICIP | 2 |
| 2022 | Adaptive Detail Injection-Based Feature Pyramid Network for Pan-SharpeningabstractMany remarkable works have been proposed to deal with distortions problems in image fusion to date. However, the spectral distortion and the spatial distortion cannot always be well addressed at the same time. To deal with this, we propose an Adaptive Feature Pyramid Network (AFPN) to efficiently embed an Adaptive Detail Injection (ADI) module at different scales. Feature-domain injection gains are proposed in the ADI module to adaptively modulate spatial information and guide a refined detail injection. Furthermore, we propose a texture loss function to further guide our model to learn detail perception in each band. Experiments on QuickBird and GaoFen-1 datasets show that our method achieves superior performance and produces visually pleasing fusion images. Our code is available at https://github.com/yisun98/AFPN. Yi Sun 0009, Yuanlin Zhang 0003, Yuan Yuan 0001 |
ICIP | 3 |
| 2022 | Prototype Queue Learning for Multi-Class Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation aims to undertake the segmentation task of novel classes with only a few annotated images. However, most existing methods tend to segment the foreground and background in the image, which limits practical application. In this paper, we present a Prototype Queue Network, which performs few-shot segmentation on multiclass in the images by aggregating binary classes into multiple classes. A prototype queue learning module is proposed to achieve multi-class segmentation by mining the relationship among features of different classes with queue and pseudo labels. In addition, a background latent class distribution refinement module is proposed to prevent the latent novel class in the background from being incorrectly predicted, which refines the boundary among different classes. Furthermore, we propose a two-steps segmentation module to optimize the process of extracting feature representation by adding progressive constraints, which can further improve the accuracy of segmentation. Experiments on the UDD and Vaihingen datasets demonstrate that our method achieves state-of-the-art performance. Zichao Wang 0005, Zhiyu Jiang, Yuan Yuan 0001 |
ICIP | 3 |
| 2022 | Learning From Synthetic Data for Crowd Instance Segmentation in the WildabstractCrowd understanding has widespread applications, including video surveillance, crowd monitoring. Unlike existing coarse-grained crowd understanding methods(e.g., counting people in images), crowd instance segmentation can provide more precise results (pixel-wise segmentation for each person in images). However, crowd instance segmentation demands a considerable amount of pixel-wise labeled data, which is very time-consuming and challenging to annotate accurate human instance masks in the crowd scene. In this paper, we propose a data generator and labeler to automatically generate synthetic crowd instance segmentation data. Then based on it, we build a large-scale synthetic crowd instance segmentation dataset called "GCIS Dataset". Besides, we demonstrate two approaches that utilize the synthetic GCIS dataset to advance the performance of crowd instance segmentation: 1)supervised crowd instance segmentation: pretrain crowd instance segmentation models on GCIS dataset, then finetune on other real data. It can remarkably boost the model’s real-world performance; 2) crowd instance segmentation via domain adaption: transfer the synthetic GCIS dataset to photo-realistic images, then train the model together with transformed data and real data, which shows better performance when tested on real-world data. Extensive experiments show the validity of the synthetic GCIS dataset for crowd instance segmentation. The dataset and source code will be released online. Yuan Yuan 0001, Qi Wang 0009 |
ICIP | 2 |
| 2022 | DTransGAN: Deblurring Transformer Based on Generative Adversarial NetworkabstractMotion deblurring is challenging due to the fast movements of the object or the camera itself. Existing methods usually try to liberate it by training CNN model or Generative Adversarial Networks(GAN). However, their methods can’t restore the details very well. In this paper, a Deblurring Transformer based on Generative Adversarial Network(DTransGAN) is proposed to improve the deblurring performance of the vehicles under the surveillance camera scene. The proposed DTransGAN combines the low-level information and the high-level information through skip connection, which saves the original information of the image as much as possible to restore the details. Besides, we replace the convolution layer in the generator with the swin transformer block, which could pay more attention to the reconstruction of details. Finally, we create the vehicle motion blur dataset. It contains two parts, namely the clear image and the corresponding blurry image. Experiments on public datasets and the collected dataset report that DTransGAN achieves the state-of-the-art for motion deblurring task. Kai Zhuang, Yuan Yuan 0001, Qi Wang 0009 |
ICIP | 2 |
| 2022 | EMGC²F: Efficient Multi-view Graph Clustering with Comprehensive FusionabstractThis paper proposes an Efficient Multi-view Graph Clustering with Comprehensive Fusion (EMGC²F) model and a corresponding efficient optimization algorithm to address multi-view graph clustering tasks effectively and efficiently. Compared to existing works, our proposals have the following highlights: 1) EMGC²F directly finds a consistent cluster indicator matrix with a Super Nodes Similarity Minimization module from multiple views, which avoids time-consuming spectral decomposition in previous works. 2) EMGC²F comprehensively mines information from multiple views. More formally, it captures the consistency of multiple views via a Cross-view Nearest Neighbors Voting (CN²V) mechanism, meanwhile capturing the importance of multiple views via an adaptive weighted-learning mechanism. 3) EMGC²F is a parameter-free model and the time complexity of the proposed algorithm is far less than existing works, demonstrating the practicability. Empirical results on several benchmark datasets demonstrate that our proposals outperform SOTA competitors both in effectiveness and efficiency. Danyang Wu, Jitao Lu, Feiping Nie 0001, Rong Wang 0001, Yuan Yuan 0001 |
IJCAI | 5 |
| 2022 | Hyperspectral image super-resolution via multi-domain feature learning
Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
Neurocomputing | 2 |
| 2022 | Multi-type spectral spatial feature for hyperspectral image classification
Yuan Yuan 0001, Mingxin Jin |
Neurocomputing | 1 |
| 2022 | Dual attention and dual fusion: An accurate way of image-based geo-localization
Yuan Yuan 0001, Bo Sun 0017, Ganchao Liu |
Neurocomputing | 1 |
| 2022 | Bipartite graph based spectral rotation with fuzzy anchors
Yuan Yuan 0001, Chengze Wang |
Neurocomputing | 1 |
| 2022 | Locate Where You Are by Block Joint Learning NetworkabstractUnmanned aerial vehicles (UAVs) are widely applied in various fields, which is located by the global position system (GPS) in most cases. However, in the GPS-denied cases, visual localization becomes very important. In complicated environments, such as weak illumination and ground objects changed, visual localization is unstable. This letter presents a UAV visual localization model with a block joint learning network (BJN). Different from the traditional feature extraction-comparison paradigm, the proposed BJN extracts the joint features of the input images at the same time, so as to fully mine the coupling relationship between the multi-source images. Besides this, to overcome the challenges of inconsistency style changes in image matching, the saliency feature based on the attention mechanism and the traditional edge feature operator are introduced in joint feature learning. To evaluate the performance of the proposed model, the experiments on the simulated dataset and the real dataset are given. Both the results on simulated and real datasets indicate that the proposed model is effective on multi-source image matching and UAV visual localization. Ganchao Liu, Yuan Yuan 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | A New Multiscale Residual Learning Network for HSI Inconsistent Noise RemovalabstractHyperspectral image (HSI) often suffers from various noise disturbances which makes the interpretation difficult. To solve this problem, a lot of HSI denoising algorithms have been proposed and widely used. Although many convolution neural network (CNN) based denoising methods have achieved successful performance in independently and identically distributed (i.i.d) Gaussian denoising, they are still limited to remove inconsistent noise, and even worse than traditional representative algorithms like total variation regularized low-rank matrix factorization (LRTV) and low-rank matrix recovery (LRMR). In this letter, a new multiscale residual learning network (MSRHSID) is proposed for HSIs denoising. In this network, a noise estimation network is used in our proposed method to obtain the image noise prior to achieve the reduction of inconsistent noise. At the same time, an efficient multiscale residual module (MRM) is employed to further improve the denoising effect. Both of the denoising experiments on synthetic and real-data HSIs show that this proposed MSRHSID outperforms the state-of-the-art methods. Yuan Yuan 0001, Hanwen Ma, Ganchao Liu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | OLCN: An Optimized Low Coupling Network for Small Objects DetectionabstractIn remotely sensed images, it is quite common to run into small objects, such as cars and small storage tanks. However, these small objects are quite easy to get ignored because of the positioning difficulty. Thus, small objects detection is very challenging for the remote sensing object detection task. In order to deal with this challenge, theoptimized low coupling network(OLCN) is proposed. First, alow coupling robust regression(LCRR) module improves the positioning accuracy to avoid small objects getting missed. Second, areceptive field optimizing layer(RFOL) is proposed to train better classifiers by providing more accurateregions of interest(RoIs). Experimental results on the public dataset HRRSD verify the effectiveness of the proposed OLCN. Small objects detection metric is improved from 5.70% of the baseline to 22.90% of the OLCN. Moreover, the proposed method has reached state-of-the-art performance on the HRRSD dataset. Yuan Yuan 0001, Yuanlin Zhang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | SDUNet: Road extraction via spatial enhanced and densely connected UNet
Mengxing Yang, Yuan Yuan 0001, Ganchao Liu |
Pattern Recognit. | 2 |
| 2022 | Adaptive open domain recognition by coarse-to-fine prototype-based network
Yuan Yuan 0001, Xinxing He, Zhiyu Jiang |
Pattern Recognit. | 1 |
| 2022 | A Novel NMF Guided for Hyperspectral Unmixing From Incomplete and Noisy DataabstractThe nonnegative matrix factorization (NMF)-combined spatial–spectral information has been widely applied in the unmixing of hyperspectral images (HSIs). However, how to select the appropriate similarity pixels and explore the spatial information and how to adapt the unmixing algorithm to complex data are both great challenges. In this article, we propose a novel unmixing method named spatial–spectral neighborhood preserving NMF (SSNPNMF) for incomplete and noisy HSI data. First, a spatial–spectral kernel regularizer is introduced to preprocess the HSI, which can reduce noise and complete missing elements. Second, a distance metric SSD based on spatial–spectral information is designed to select similar pixels in the image. Subsequently, the spatial–spectral relationship of the selected first$k$similar pixels is used to reconstruct the image and obtain the reconstruction matrix. Finally, the reconstruction matrix is used to constrain the abundances and improve the unmixing performance. Experimental results on synthetic data and Cuprite data indicate that SSNPNMF has a more effective unmixing performance compared with the state-of-the-art methods. Xiaoqiang Lu, Ganchao Liu, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Symmetrical Feature Propagation Network for Hyperspectral Image Super-ResolutionabstractSingle hyperspectral image (HSI) super-resolution (SR) methods using a auxiliary high-resolution (HR) RGB image have achieved great progress recently. However, most existing methods aggregate the information of RGB image and HSI early during input or shallow feature extraction, whose difference between two images has not been treated and discussed. Although a few methods combine both the image features in the middle layer of the network, they fail to make full use of the two inherent properties, i.e., rich spectra of HSI and HR content of RGB image, to guide model representation learning. To address these issues, in this article, we propose a dual-stage learning approach for HSI SR to learn a general spatial–spectral prior and image-specific details, respectively. In the coarse stage, we fully take advantage of two adjacent bands and RGB image to build the model. During coarse SR, a symmetrical feature propagation approach is developed to learn the inherent content of each image over a relatively long range. The symmetrical structure encourages the two streams to better retain their particularity. Meanwhile, it can realize the information interaction by the adaptive local block aggregation (ALBA) module. To learn image-specific details, a back-projection refinement network is embedded in the structure, which further improves the performance in fine stage. The experiments on four benchmark datasets demonstrate that the proposed approach presents excellent performance over the existing methods. Our code is publicly available athttps://github.com/qianngli/SFPN. Qiang Li 0042, Maoguo Gong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Adaptive Relationship Preserving Sparse NMF for Hyperspectral UnmixingabstractHyperspectral unmixing is an essential research topic for spectral data analysis due to the existence of mixed pixels. Recently, many methods based on sparse nonnegative matrix factorization (NMF) have been widely used for unmixing by incorporating similarity preserving. However, most of them conduct the similarity learning and unmixing by using two separate steps, which may lead to the case that the learned similarity matrix is not the optimal one for unmixing. Thus, the performance of unmixing would be influenced and become undesirable. To alleviate this problem, we propose an adaptive relationship preserving-based sparse NMF (ARP-NMF) for hyperspectral unmixing. Typically, we regard similarity learning and unmixing as an alternative optimization process. During this process, the learned similarity weights and unmixing results can be mutually improved. Thus, our proposed method can learn the optimal similarity weights for unmixing and obtain better generalization ability for different hyperspectral images than traditional methods. Moreover, by using the spectral and spatial local structure, our ARP-NMF method effectively preserves structure consistency between pixels and abundances. Experimental results both on the synthetic data and the real data reveal that our proposed method outperforms several representative sparse NMF-based methods. Xuelong Li 0001, Yuan Yuan 0001, Yongsheng Dong 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | ABNet: Adaptive Balanced Network for Multiscale Object Detection in Remote Sensing ImageryabstractBenefiting from the development of convolutional neural networks (CNNs), many excellent algorithms for object detection have been presented. Remote sensing object detection (RSOD) is a challenging task mainly due to: 1) complicated background of remote sensing images (RSIs) and 2) extremely imbalanced scale and sparsity distribution of remote sensing objects. Existing methods cannot effectively solve these problems with excellent detection accuracy and rapid speed. To address these issues, we propose an adaptive balanced network (ABNet) in this article. First, we design an enhanced effective channel attention (EECA) mechanism to improve the feature representation ability of the backbone, which can alleviate the obstacles of complex background on foreground objects. Then, to combine multiscale features adaptively in different channels and spatial positions, an adaptive feature pyramid network (AFPN) is designed to capture more discriminative features. Furthermore, considering that the original FPN ignores rich deep-level features, a context enhancement module (CEM) is proposed to exploit abundant semantic information for multiscale object detection. Experimental results on three public datasets demonstrate that our approach exhibits superior performance over baseline by only introducing less than 1.5M extra parameters. Yanfeng Liu, Qiang Li 0042, Yuan Yuan 0001, Qian Du 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Style Transformation-Based Spatial-Spectral Feature Learning for Unsupervised Change DetectionabstractDue to the inconsistent imaging environment, the styles of multitemporal multispectral images (MSIs) are quite different, such as image brightness and transparency. For multitemporal MSIs with different styles, the “same object with different spectra” problem is one of the biggest challenges in change detection. To overcome the challenge, a novel unsupervised spatial–spectral feature learning (FL) framework based on style transformation (ST) (called STFL-CD) is proposed for MSI change detection in this article. For dual-temporal MSIs, the proposed STFl-CD algorithm consists of two phases: ST and spatial–spectral FL. Since the image styles are inconsistent under different imaging environments, the first innovation is to transform the image styles through unmixing and reconstruction. Through ST, the challenge of the “same object with different spectra” problem will be reduced fundamentally. By introducing the attention mechanism, the other innovation is to extract the joint spectral–spatial change features based on a 3-D convolutional neural network with spatial and channel attention. In addition, for multitemporal MSIs, a multitemporal version STFL-CD (MT-STFL-CD) framework is designed based on a recurrent neural network to learn the correlation features between multitemporal remote sensing images. Both of the visual and quantitative results on the real MSI datasets indicate that the proposed unsupervised STFL-CD frameworks have significant advantages on multitemporal MSI change detection. In particular, the performance of the proposed unsupervised STFL-CD algorithm is even comparable to that of the state-of-the-art supervised or semisupervised methods. Ganchao Liu, Yuan Yuan 0001, Yuelin Zhang, Yongsheng Dong 0002, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Semantics-Consistent Representation Learning for Remote Sensing Image-Voice RetrievalabstractWith the development of earth observation technology, massive amounts of remote sensing (RS) images are acquired. To find useful information from these images, cross-modal RS image–voice retrieval provides a new insight. This article aims to study the task of RS image–voice retrieval so as to search effective information from massive amounts of RS data. Existing methods for RS image–voice retrieval rely primarily on the pairwise relationship to narrow the heterogeneous semantic gap between images and voices. However, apart from the pairwise relationship included in the data sets, the intramodality and nonpaired intermodality relationships should also be considered simultaneously since the semantic consistency among nonpaired representations plays an important role in the RS image–voice retrieval task. Inspired by this, a semantics-consistent representation learning (SCRL) method is proposed for RS image–voice retrieval. The main novelty is that the proposed method takes the pairwise, intramodality, and nonpaired intermodality relationships into account simultaneously, thereby improving the semantic consistency of the learned representations for the RS image–voice retrieval. The proposed SCRL method consists of two main steps: 1) semantics encoding and 2) SCRL. First, an image encoding network is adopted to extract high-level image features with a transfer learning strategy, and a voice encoding network with dilated convolution is devised to obtain high-level voice features. Second, a consistent representation space is conducted by modeling the three kinds of relationships to narrow the heterogeneous semantic gap and learn semantics-consistent representations across two modalities. Extensive experimental results on three challenging RS image–voice data sets, including Sydney, UCM, and RSICD image–voice data sets, show the effectiveness of the proposed method. Hailong Ning, Bin Zhao 0001, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Hybrid Feature Aligned Network for Salient Object Detection in Optical Remote Sensing ImageryabstractRecently, salient object detection in optical remote sensing images (RSI-SOD) has attracted great attention. Benefiting from the success of deep learning and the inspiration of natural SOD task, RSI-SOD has achieved fast progress over the past two years. However, existing methods usually suffer from the intrinsic problems of optical RSIs, 1) cluttered background; 2) scale variation of salient objects; 3) complicated edges and irregular topology. To remedy these problems, we propose a hybrid feature aligned network (HFANet) jointly modeling boundary learning to detect salient objects effectively. Specifically, we design a hybrid encoder by unifying two components to capture global context for mitigating the disturbance of complex background. Then, to detect multiscale salient objects effectively, we propose a Gated Fold-ASPP (GF-ASPP) to extract abundant context in the deep semantic features. Furthermore, an adjacent feature aligned module (AFAM) is presented for integrating adjacent features with unparameterized alignment strategy. Finally, we propose a novel interactive guidance loss (IGLoss) to combine saliency and edge detection, which can adaptively perform mutual supervision of the two sub-tasks to facilitate detection of salient objects with blurred edges and irregular topology. Adequate experimental results on three optical RSI-SOD datasets reveal that the presented approach exceeds 18 state-of-the-art ones. All codes and detection results are available athttps://github.com/lyf0801/HFANet. Qi Wang 0009, Yanfeng Liu, Zhitong Xiong, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Spatial-Spectral Clustering With Anchor Graph for Hyperspectral ImageabstractHyperspectral image (HSI) clustering, which aims at dividing hyperspectral pixels into clusters without labeled training data, has drawn significant attention in practical applications. Recently, many graph-based clustering methods, which construct an adjacent graph to model the data relationship, have shown dominant performance. However, the high dimensionality of HSI data makes it hard to construct the pairwise adjacent graph. Besides, abundant spatial structures are often overlooked during the clustering procedure. In order to better handle the high dimensionality problem and preserve the spatial structures, this paper proposes a novel unsupervised approach called spatial-spectral clustering with anchor graph (SSCAG) for HSI data clustering. The SSCAG has the following contributions: 1) the multiscale filtering module is utilized to smooth the homogeneous regions, so that it can increase the similarity and consistency of neighboring pixels and capture the multiple views of a local region with different scales; 2) a new similarity metric is proposed to embed the spatial-spectral features into the combined adjacent graph, which can mine the intrinsic property structure of HSI data; 3) the AG-based strategy is adopted to construct the adjacent graph by a neighbor assignment scheme without hyperparameters, and its optimization employs SVD to replace eigenvalue decomposition to reduce the computational complexity. Extensive experiments on three public HSI datasets show that the proposed SSCAG is competitive against the state-of-the-art approaches. Qi Wang 0009, Yanling Miao, Mulin Chen, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Hyperspectral Unmixing Using Nonlocal Similarity-Regularized Low-Rank Tensor FactorizationabstractRecently, methods based on nonnegative tensor factorization (NTF), which benefits from the tensor representation of hyperspectral imagery (HSI) without any information loss, have attracted increasing attention. However, most existing methods fail to explore the internal spatial structure of data, resulting in low unmixing performance. Moreover, when the algorithm is optimized, the solution is unstable. In this article, a regularizer based on nonlocal tensor similarity is proposed, which can not only fully preserve the global information of HSI but also mine the internal information of data in the spatial domain. HSI is regarded as a 3-D tensor and is directly subjected to endmember extraction and abundance estimation. To fully explore the structural characteristics of data, we simultaneously use the local smoothing and low tensor rank prior of the data to constrain the unmixing model. First, several 4-D tensor groups can be obtained after the nonlocal similarity structure of HSI is learned. Subsequently, a low tensor rank prior is applied to each 4-D tensor, which can fully simulate the nonlocal similarity in the image. In addition, total variation (TV) is also used to explore the local spatial relationship of data, which can generate a smooth abundance map through edge preservation. The optimization is solved by the ADMM algorithm. Experiments on synthetic and real data illustrate the superiority of the proposed method. Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Feature-Aligned Single-Stage Rotation Object Detection With Continuous BoundaryabstractRecently, rotation detection has gained much attention and shown its potential for accurate localization in remote sensing scenes. However, the objects in remote sensing images have a variety of directions, sizes, and aspect ratios, which makes it difficult to locate and classify objects. Therefore, the object detection task is still facing great challenges in the field of remote sensing. In this paper, we propose a novel single-stage detector, which includes feature alignment block (FAB), double regression branches (DRBs), and a circumcircle rotation box (CRB). FAB utilizes the deformable convolution to flexibly obtain the features of different aspect ratio objects, and aligns the regression features with the corresponding classification features through fusion. Consequently, it can make the extracted feature information have stronger discrimination, which is conducive to improving the object position and classification accuracy. DRBs consist of a main regression branch and an auxiliary regression branch. The main regression branch is used to fine-tune the result of the auxiliary regression branch to obtain a more accurate regression result. Moreover, to eliminate the boundary discontinuity problem faced by regression-based detectors, we construct CRB through designing an angle of rotation and the radius of the circumcircle. Extensive experiments and visual analysis are conducted on three public benchmarks, i.e., DOTA, HRSC2016, and DIOR-R. The results show that our proposed method has excellent localization and classification performance for oriented objects on all the representative datasets. Yuan Yuan 0001, Dandan Ma |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Partial-DNet: A Novel Blind Denoising Model With Noise Intensity Estimation for HSIabstractBecause of the inevitable noise interference in hyperspectral images (HSIs), the understanding and application of HSIs are seriously restricted. To solve this problem, the research on the data-driven neural network-based denoising method has become a hotspot in recent years. However, for HSIs with inconsistent and mixed noises, the traditional data-driven denoising algorithms have obvious limitations on generalization ability. In order to overcome these drawbacks, a novel partial densenet (Partial-DNet) model with noise intensity estimation is proposed for HSIs blind denoising in this article. In the proposed Partial-DNet model, the noise intensity of each band is estimated in the first and then fused with the observed images to generate the feature maps by introducing the channel attention mechanism. Finally, a novel multiscale neural network is explored to extract the spatial–spectral joint features. The contributions of this article can be summarized as follows: 1) the noise intensity of each band is estimated as a prior to guide the blind denoising framework suit with different data sets adaptively; 2) a Partial-DNet model is proposed to extract multiscale spatial–spectral features more efficient to maintain the details better while denoising; and 3) the experiments on both simulated and real HSI data sets indicate that the proposed denoising framework can be used to remove the inconsistent mixed noises in HSIs adaptively. Compared with the other state-of-the-art denoising methods, HSIs denoised by the proposed Partial-DNet algorithm not only have a higher PSNR index but also have higher classification accuracy under the same situation. Yuan Yuan 0001, Hanwen Ma, Ganchao Liu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Proxy-Based Deep Learning Framework for Spectral-Spatial Hyperspectral Image Classification: Efficient and RobustabstractDeep convolutional networks have been extensively deployed in hyperspectral image (HSI) classification. Reaching for high accuracy, the existing deep-learning-based methods commonly deepen or widen their networks for better performance, which brings higher computational complexity and the risk of overfitting. Although the introduction of the residual module and batch-normalization reduces the generalization degradation in complex networks, the mainstream methods still suffer from low robustness to the noise. To tackle these issues, a compact proxy-based deep learning framework is proposed to perform highly accurate HSI classification with superb efficiency and robustness. In this article: 1) novel deep proxies are integrated to replace the dense classifier layers in conventional networks, which represents specific classes in deep embedding space and enables fast and reliable convergence; 2) the proxy-based feature embedding is studied in distance metric and similarity metric, and compatible dual-metric loss functions are designed for further optimized embedding distribution, which leads to more robust generalization; and 3) state-of-the-art performance and robustness are demonstrated by the proposed framework on mainstream HSI data sets with the minimal network scale and time complexity. Yuan Yuan 0001, Chengze Wang, Zhiyu Jiang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Reweighted Low-Rank and Joint-Sparse Unmixing With Library PruningabstractSparse unmixing (SU) is a semisupervised learning problem, which performs abundance estimation when a spectral library is given. In this way, the essence of SU is to select the most suitable subset from the spectral library for representing all mixed pixels. Many SU methods adopt joint-sparse and low-rank constraints to guide the abundance estimation. However, the spatial correlation learning in these algorithms is not accurate enough, which seriously affects the unmixing performance. Besides, most pruning-based unmixing methods suffer from complicated pruning strategies and ignore the relationship between the spectral library and mixed pixels. This article proposes a reweighted low-rank and joint-sparse unmixing approach, which combines an effective pruning strategy (RLSU-LP). The RLSU-LP approach consists ofrough unmixing stage, library pruning, andfine-tuning unmixing stage. First, the proposed method utilizes image segmentation to obtain different homogeneous regions, i.e., superpixels. A confidence index is introduced to describe the superpixel homogeneity, which is conducive to learning the meticulous spatial correlation. The RLSU-LP method reasonably relaxes or tightens the sparse and low-rank constraints of the abundance matrix by using the confidence index. Furthermore, a supervised library pruning strategy is proposed, which aims to eliminate the inactive endmembers by considering the contribution of representing mixed pixels. Experiments on the synthesized dataset and authentic hyperspectral images verify the effectiveness of our proposed algorithm. Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Sparse Unmixing Based on Adaptive Loss MinimizationabstractSparse unmixing (SU) algorithms use the existing spectral library as prior knowledge to analyze the endmembers and estimate abundance maps. The majority of SU algorithms utilize loss functions based onL2,1-norm orF-norm to minimize reconstruction error. They have different advantages and shortcomings. In short,F-norm has a differentiable characteristic, and it is easy to minimize as a loss function. However, it is very sensitive to heavy noise and outliers. While theL2,1-norm emphasizes the reconstruction error on each band and is robust to noise with different intensities in different bands. But theL2,1-norm is non-differentiable at zero-point. This paper introduces an adaptive loss function based on σ-norm for SU, which combines the advantages ofL2,1-norm andF-norm. The adaptive loss function is related to a non-negative parameter σ. By adjusting the parameter σ, the adaptive loss function can approachF-norm orL2,1-norm. To the best of our knowledge, it is the first time to apply an adaptive loss function to SU. Moreover, the adaptive loss function is globally differentiable, and we propose an optimization algorithm for the adaptive loss function and verify its convergence. Experiments on the real-world and synthetic HSIs show that the adaptive loss function effectively enhances the performance of the sparse unmixing algorithms. Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Dual-Stage Approach Toward Hyperspectral Image Super-ResolutionabstractHyperspectral image produces high spectral resolution at the sacrifice of spatial resolution. Without reducing the spectral resolution, improving the resolution in the spatial domain is a very challenging problem. Motivated by the discovery that hyperspectral image exhibits high similarity between adjacent bands in a large spectral range, in this paper, we explore a new structure for hyperspectral image super-resolution (DualSR), leading to a dual-stage design, i.e., coarse stage and fine stage. In coarse stage, five bands with high similarity in a certain spectral range are divided into three groups, and the current band is guided to study the potential knowledge. Under the action of alternative spectral fusion mechanism, the coarse SR image is super-resolved in band-by-band. In order to build model from a global perspective, an enhanced back-projection method via spectral angle constraint is developed in fine stage to learn the content of spatial-spectral consistency, dramatically improving the performance gain. Extensive experiments demonstrate the effectiveness of the proposed coarse stage and fine stage. Besides, our network produces state-of-the-art results against existing works in terms of spatial reconstruction and spectral fidelity. Our code is publicly available at https://github.com/qianngli/DualSR. Qiang Li 0042, Yuan Yuan 0001, Xiuping Jia, Qi Wang 0009 |
IEEE Trans. Image Process. | 2 |
| 2022 | CM-Net: Concentric Mask Based Arbitrary-Shaped Text DetectionabstractRecently fast arbitrary-shaped text detection has become an attractive research topic. However, most existing methods are non-real-time, which may fall short in intelligent systems. Although a few real-time text methods are proposed, the detection accuracy is far behind non-real-time methods. To improve the detection accuracy and speed simultaneously, we propose a novel fast and accurate text detection framework, namely CM-Net, which is constructed based on a new text representation method and a multi-perspective feature (MPF) module. The former can fit arbitrary-shaped text contours by concentric mask (CM) in an efficient and robust way. The latter encourages the network to learn more CM-related discriminative features from multiple perspectives and brings no extra computational cost. Benefiting the advantages of CM and MPF, the proposed CM-Net only needs to predict one CM of the text instance to rebuild the text contour and achieves the best balance between detection accuracy and speed compared with previous works. Moreover, to ensure that multi-perspective features are effectively learned, the multi-factor constraints loss is proposed. Extensive experiments demonstrate the proposed CM is efficient and robust to fit arbitrary-shaped text instances, and also validate the effectiveness of MPF and constraints loss for discriminative text features recognition. Furthermore, experimental results show that the proposed CM-Net is superior to existing state-of-the-art (SOTA) real-time text detection methods in both detection speed and accuracy on MSRA-TD500, CTW1500, Total-Text, and ICDAR2015 datasets. Chuang Yang 0003, Mulin Chen, Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Image Process. | 4 |
| 2022 | Disentangled Representation Learning for Cross-Modal Biometric MatchingabstractCross-modal biometric matching (CMBM) aims to determine the corresponding voice from a face, or identify the corresponding face from a voice. Recently, many CMBM methods have been proposed by forcing the distance between two modal features to be narrowed. However, these methods ignore the alignability between the two modal features. Because the feature is extracted under the supervision of identity information from single modal data, it can only reflect the identity information of single modal data. In order to address this problem, a disentangled representation learning method is proposed to disentangle the alignable latent identity factors and nonalignable the modality-dependent factors for CMBM. The proposed method consists of two main steps: 1) feature extraction and 2) disentangled representation learning. Firstly, an image feature extraction network is adopted to obtain face features, and a voice feature extraction network is applied to learn voice features. Secondly, a disentangled latent variable is explored to disentangle the latent identity factors that are shared across the modalities from the modality-dependent factors. The modality-dependent factors are filtered out, while the latent identity factors from the two modalities are enforced to be narrowed to align the same identity information. Then, the disentangled latent identity factors are considered as pure identity information to bridge the two modalities for cross-modal verification, 1:$N$matching, and retrieval. Note that the proposed method learns the identity information from the input face images and voice segments with only identity label as supervised information. Extensive experiments on the challenging VoxCeleb dataset demonstrate the proposed method outperforms the state-of-the-art methods. Hailong Ning, Xiangtao Zheng, Xiaoqiang Lu, Yuan Yuan 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Neuron Linear Transformation: Modeling the Domain Shift for Crowd CountingabstractCross-domain crowd counting (CDCC) is a hot topic due to its importance in public safety. The purpose of CDCC is to alleviate the domain shift between the source and target domain. Recently, typical methods attempt to extract domain-invariant features via image translation and adversarial learning. When it comes to specific tasks, we find that the domain shifts are reflected in model parameters' differences. To describe the domain gap directly at the parameter level, we propose a neuron linear transformation (NLT) method, exploiting domain factor and bias weights to learn the domain shift. Specifically, for a specific neuron of a source model, NLT exploits few labeled target data to learn domain shift parameters. Finally, the target neuron is generated via a linear transformation. Extensive experiments and analysis on six real-world data sets validate that NLT achieves top performance compared with other domain adaptation methods. An ablation study also shows that the NLT is robust and more effective than supervised and fine-tune training. Code is available at https://github.com/taohan10200/NLT. Qi Wang 0009, Tao Han 0002, Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Task-Related Self-Supervised Learning For Remote Sensing Image Change DetectionabstractChange detection for remote sensing images is widely applied for urban change detection, disaster assessment and other fields. However, most of the existing CNN-based change detection methods still suffer from the problem of inadequate pseudo-changes suppression and insufficient feature representation. In this work, an unsupervised change detection method based on Task-related Self-supervised Learning Change Detection network with smooth mechanism(TSLCD) is proposed to eliminate it. The main contributions include: (1) the task-related self-supervised learning module is introduced to extract spatial features more effectively. (2) a hard-sample-mining loss function is applied to pay more attention to the hard-to-classify samples. (3) a smooth mechanism is utilized to remove some of pseudo-changes and noise. Experiments on four remote sensing change detection datasets reveal that the proposed TSLCD method achieves the state-of-the-art for change detection task. Zhinan Cai, Zhiyu Jiang, Yuan Yuan 0001 |
ICASSP | 3 |
| 2021 | Selection Based on Statistical Characteristics for Object DetectionabstractIn the domain of object detection, automatically selecting positive and negative samples methods have become a hot research topic in recent years. However, most of them focus on improving the sampling process but ignore the relationship between object size and feature map, in which the shallow and deep feature layers can capture small and large size objects well respectively. In this paper, we propose a multi-scale sample selection based on statistical characteristics for object detection. To improve the robustness of the Intersection over Union (IoU) threshold, we design a multi-scale sample selection module (MSSM), which takes full advantage of different feature layers. Besides, we introduce a multi-scale attention module (MSAM) by embedding in the feature pyramid networks (FPN) to improve the efficiency of feature fusion. Experiments on MS COCO dataset demonstrate that our method achieves significant improvement over the state-of-the-art methods. Yuan Yuan 0001, Dandan Ma |
ICASSP | 2 |
| 2021 | Deep Feature Selection-And-Fusion for RGB-D Semantic SegmentationabstractScene depth information can help visual information for more accurate semantic segmentation. However, how to effectively integrate multi-modality information into representative features is still an open problem. Most of the existing work uses DCNNs to implicitly fuse multi-modality information. But as the network deepens, some critical distinguishing features may be lost, which reduces the segmentation performance. This work proposes a unified and efficient feature selection-and-fusion network (FSFNet), which contains a symmetric cross-modality residual fusion module used for explicit fusion of multi-modality information. Besides, the network includes a detailed feature propagation module, which is used to maintain low-level detailed information during the forward process of the network. Compared with the state-of-the-art methods, experimental evaluations demonstrate that the proposed model achieves competitive performance on two public datasets. Yuejiao Su, Yuan Yuan 0001, Zhiyu Jiang |
ICME | 2 |
| 2021 | Weighted Sparsity Constraint Tensor Factorization for Hyperspectral UnmixingabstractRecently, the unmixing methods based on non-negative tensor factorization (NTF) have received a lot of attention. Many NTF-based methods combine total variation (TV) regularization, aiming at maintaining the smoothness of the abundance maps to improve the performance of unmixing. However, the existing TV regularization ignores the sparsity sharing on the spatial difference images among different bands. To tackle this issue, a weighted total variation regularizer on the spatial difference maps of abundances is proposed in this paper, which uses the$L_{2,1}$norm to explore the sparse structure in abundances along the spectral dimension. In addition, the$L_{1/2}$norm is used to enhance the spatial sparsity of abundances. The proposed method can not only enhance the sparsity in abundances, but also keep the spatial similarity characteristics of data. Compared with the existing popular methods, the proposed method has superior performance on both synthetic data and real data. Yuan Yuan 0001 |
IGARSS | 1 |
| 2021 | One-Stage Detector from Coarse to Fine for Rotating Object of Remote SensingabstractRotation detection has become a popular topic in the field of remote sensing in recent years. Although quite a few progress has been made, some challenges still exist in feature alignment and regression accuracy due to large aspect ratio and arbitrary orientations of remote sensing objects, especially for the one-stage detectors. To address these problems, we propose a novel one-stage detector from coarse to fine for rotating objects. To alleviate misalignment problem between regression features and classification features, we construct the Feature Alignment Block (FAB). It can flexibly extract the features of objects with different aspect ratios by the deformable convolution and align the regression features with the corresponding classification features. Moreover, to obtain a more accurate regression estimate, we design the refined regression head (RRH) that can effectively fine-tune the coarse regression position. Experiments on the public DOTA and HRSC2016 datasets demonstrate that our proposed method shows excellent detection performance for rotating objects. Yuan Yuan 0001, Dandan Ma |
IGARSS | 2 |
| 2021 | Hyperspectral Anomaly Detection Based on Adaptive Weighted Sparse Dictionary LearningabstractThe background estimation and modeling are the core of hyperspectral anomaly detection. However, the complex hyperspectral image does not conform to the assumption of multivariate normal distribution in most methods. At the same time, the existence of unknown abnormal targets in the background will also affect the modeling of the background. To solve the above problems, a hyperspectral anomaly detection method based on adaptive weighted sparse dictionary learning (AWSDLD) is proposed in this paper. Firstly, the dictionary learning framework based on adaptive weights is used to learn more representative background dictionaries without considering the background distribution. Secondly, due to the capped norm property, the proposed method can effectively suppress the influence of abnormal targets on background modeling. Finally, the abnormal targets are more significant and easier to be detected in the residual image between the reconstructed image and the original image. The experimental results on three real datasets show the effectiveness of the proposed method. Yuan Yuan 0001 |
IGARSS | 2 |
| 2021 | Adaptive Spectral and Spatial Feature Extraction Framework for Hyperspectral ClassificationabstractHyperspectral image (HSI) classification is an important research topic in the field of remote sensing. In addition to discriminative spectral information, spatial information also plays an important part in HSI data. So jointly extracting spectral-spatial features is popular to achieve better classification in most recent research. However, simply directly introducing the spatial information without analyzing its necessity will result in some problems. In some cases, spectra have enough material discrimination ability and spatial feature is indeed unneceseary which will brings additional computational burden and even adversely affect the classification results. In order to address these problems, we propose an adaptive spectral spatial feature extraction framework with early prediction strategy for HSI classification. Our method can not only perform high efficiency but also reduce the potential interference of spatial information to improve classification accuracy. Specifically, it mainly consists of two classification branches and a small gate network which is utilized to adaptively determine the necessity of spatial features. Experimental results on the public HSI datasets demonstrate that our approach obtains better performance in both accuracy and efficiency than the comparative state-of-the-art level methods. Yuan Yuan 0001, Dandan Ma |
IGARSS | 2 |
| 2021 | Self-Supervised Spectral Matching Network for Hyperspectral Target DetectionabstractHyperspectral target detection is a pixel-level recognition problem. Given a few target samples, it aims to identify the specific target pixels such as airplane, vehicle, ship, from the entire hyperspectral image. In general, the background pixels take the majority of the image and complexly distributed. As a result, the datasets are weak annotated and extremely imbalanced. To address these problems, a spectral mixing based self-supervised paradigm is designed for hyperspectral data to obtain an effective feature representation. The model adopts a spectral similarity based matching network framework. In order to learn more discriminative features, a pair-based loss is adopted to minimize the distance between target pixels while maximizing the distances between target and background. Furthermore, through a background separated step, the complex unlabeled spectra are downsampled into different sub-categories. The experimental results on three real hyperspectral datasets demonstrate that the proposed framework achieves better results compared with the existing detectors. Can Yao, Yuan Yuan 0001, Zhiyu Jiang |
IGARSS | 2 |
| 2021 | AWFA-LPD: Adaptive Weight Feature Aggregation for Multi-frame License Plate DetectionabstractFor license plate detection (LPD), most of the existing work is based on images as input. If these algorithms can be applied to multiple frames or videos, they can be adapted to more complex unconstrained scenes. In this paper, we propose a LPD framework for detecting license plates in multiple frames or videos, called AWFA-LPD, which effectively integrates the features of nearby frames. Compared with image based detection models, our network integrates optical flow extraction module, which can propagate the features of local frames and fuse with the reference frame. Moreover, we concatenate a non-link suppression module after the detection results to post-process the bounding boxes. Extensive experiments demonstrate the effectiveness and efficiency of our framework. Xiaocheng Lu, Yuan Yuan 0001, Qi Wang 0009 |
ICMR | 2 |
| 2021 | Pixel-Wise Crowd Understanding via Synthetic Data
Qi Wang 0009, Junyu Gao 0001, Wei Lin 0018, Yuan Yuan 0001 |
Int. J. Comput. Vis. | 4 |
| 2021 | A semi-supervised learning algorithm via adaptive Laplacian graph
Yuan Yuan 0001, Qi Wang 0009, Feiping Nie 0001 |
Neurocomputing | 1 |
| 2021 | Distribution equalization learning mechanism for road crack detection
Jie Fang 0001, Bo Qu, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2021 | Audio description from image by modal translation network
Hailong Ning, Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
Neurocomputing | 3 |
| 2021 | Local and correlation attention learning for subtle facial expression recognition
Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu |
Neurocomputing | 2 |
| 2021 | A novel hyperspectral unmixing model based on multilayer NMF with Hoyer's projection
Yuan Yuan 0001, Ganchao Liu |
Neurocomputing | 1 |
| 2021 | Person Reidentification via Unsupervised Cross-View Metric LearningabstractPerson reidentification (Re-ID) aims to match observations of individuals across multiple nonoverlapping camera views. Recently, metric learning-based methods have played important roles in addressing this task. However, metrics are mostly learned in supervised manners, of which the performance relies heavily on the quantity and quality of manual annotations. Meanwhile, metric learning-based algorithms generally project person features into a common subspace, in which the extracted features are shared by all views. However, it may result in information loss since these algorithms neglect the view-specific features. Besides, they assume person samples of different views are taken from the same distribution. Conversely, these samples are more likely to obey different distributions due to view condition changes. To this end, this paper proposes an unsupervised cross-view metric learning method based on the properties of data distributions. Specifically, person samples in each view are taken from a mixture of two distributions: one models common prosperities among camera views and the other focuses on view-specific properties. Based on this, we introduce a shared mapping to explore the shared features. Meanwhile, we construct view-specific mappings to extract and project view-related features into a common subspace. As a result, samples in the transformed subspace follow the same distribution and are equipped with comprehensive representations. In this paper, these mappings are learned in an unsupervised manner by clustering samples in the projected space. Experimental results on five cross-view datasets validate the effectiveness of the proposed method. Yachuang Feng, Yuan Yuan 0001, Xiaoqiang Lu |
IEEE Trans. Cybern. | 2 |
| 2021 | Feature-Aware Adaptation and Density Alignment for Crowd Counting in Video SurveillanceabstractWith the development of deep neural networks, the performance of crowd counting and pixel-wise density estimation is continually being refreshed. Despite this, there are still two challenging problems in this field: 1) current supervised learning needs a large amount of training data, but collecting and annotating them is difficult and 2) existing methods cannot generalize well to the unseen domain. A recently released synthetic crowd dataset alleviates these two problems. However, the domain gap between the real-world data and synthetic images decreases the models' performance. To reduce the gap, in this article, we propose a domain-adaptation-style crowd counting method, which can effectively adapt the model from synthetic data to the specific real-world scenes. It consists of multilevel feature-aware adaptation (MFA) and structured density map alignment (SDA). To be specific, MFA boosts the model to extract domain-invariant features from multiple layers. SDA guarantees the network outputs fine density maps with a reasonable distribution on the real domain. Finally, we evaluate the proposed method on four mainstream surveillance crowd datasets, Shanghai Tech Part B, WorldExpo'10, Mall, and UCSD. Extensive experiments are evidence that our approach outperforms the state-of-the-art methods for the same cross-domain counting problem. Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Cybern. | 2 |
| 2021 | Bio-Inspired Representation Learning for Visual Attention PredictionabstractVisual attention prediction (VAP) is a significant and imperative issue in the field of computer vision. Most of the existing VAP methods are based on deep learning. However, they do not fully take advantage of the low-level contrast features while generating the visual attention map. In this article, a novel VAP method is proposed to generate the visual attention map via bio-inspired representation learning. The bio-inspired representation learning combines both low-level contrast and high-level semantic features simultaneously, which are developed by the fact that the human eye is sensitive to the patches with high contrast and objects with high semantics. The proposed method is composed of three main steps: 1) feature extraction; 2) bio-inspired representation learning; and 3) visual attention map generation. First, the high-level semantic feature is extracted from the refined VGG16, while the low-level contrast feature is extracted by the proposed contrast feature extraction block in a deep network. Second, during bio-inspired representation learning, both the extracted low-level contrast and high-level semantic features are combined by the designed densely connected block, which is proposed to concatenate various features scale by scale. Finally, the weighted-fusion layer is exploited to generate the ultimate visual attention map based on the obtained representations after bio-inspired representation learning. Extensive experiments are performed to demonstrate the effectiveness of the proposed method. Yuan Yuan 0001, Hailong Ning, Xiaoqiang Lu |
IEEE Trans. Cybern. | 1 |
| 2021 | Spectral-Spatial Joint Sparse NMF for Hyperspectral UnmixingabstractThe nonnegative matrix factorization (NMF) combining with spatial-spectral contextual information is an important technique for extracting endmembers and abundances of hyperspectral image (HSI). Most methods constrain unmixing by the local spatial position relationship of pixels or search spectral correlation globally by treating pixels as an independent point in HSI. Unfortunately, they ignore the complex distribution of substance and rich contextual information, which makes them effective in limited cases. In this article, we propose a novel unmixing method via two types of self-similarity to constrain sparse NMF. First, we explore the spatial similarity patch structure of data on the whole image to construct the spatial global self-similarity group between pixels. And according to the regional continuity of the feature distribution, the spectral local self-similarity group of pixels is created inside the superpixel. Then based on the sparse expression of the pixel in the subspace, we sparsely encode the pixels in the same spatial group and spectral group respectively. Finally, the abundance of pixels within each group is forced to be similar to constrain the NMF unmixing framework. Experiments on synthetic and real data fully demonstrate the superiority of our method over other existing methods. Yuan Yuan 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | ASK: Adaptively Selecting Key Local Features for RGB-D Scene RecognitionabstractIndoor scene images usually contain scattered objects and various scene layouts, which make RGB-D scene classification a challenging task. Existing methods still have limitations for classifying scene images with great spatial variability. Thus, how to extract local patch-level features effectively using only image label is still an open problem for RGB-D scene recognition. In this article, we propose an efficient framework for RGB-D scene recognition, which adaptively selects important local features to capture the great spatial variability of scene images. Specifically, we design a differentiable local feature selection (DLFS) module, which can extract the appropriate number of key local scene-related features. Discriminative local theme-level and object-level representations can be selected with DLFS module from the spatially-correlated multi-modal RGB-D features. We take advantage of the correlation between RGB and depth modalities to provide more cues for selecting local features. To ensure that discriminative local features are selected, the variational mutual information maximization loss is proposed. Additionally, the DLFS module can be easily extended to select local features of different scales. By concatenating the local-orderless and global-structured multi-modal features, the proposed framework can achieve state-of-the-art performance on public RGB-D scene recognition datasets. Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Image Process. | 2 |
| 2020 | Efficient Dynamic Scene Deblurring Using Spatially Variant Deconvolution Network With Optical Flow Guided TrainingabstractIn order to remove the non-uniform blur of images captured from dynamic scenes, many deep learning based methods design deep networks for large receptive fields and strong fitting capabilities, or use multi-scale strategy to deblur image on different scales gradually. Restricted by the fixed structures and parameters, these methods are always huge in model size to handle complex blurs. In this paper, we start from the deblurring deconvolution operation, then design an effective and real-time deblurring network. The main contributions are three folded, 1) we construct a spatially variant deconvolution network using modulated deformable convolutions, which can adjust receptive fields adaptively according to the blur features. 2) our analysis shows the sampling points of deformable convolution can be used to approximate the blur kernel, which can be simplified to bi-directional optical flows. So the position learning of sampling points can be supervised by bi-directional optical flows. 3) we build a light-weighted backbone for image restoration problem, which can balance the calculations and effectiveness well. Experimental results show that the proposed method achieves state-of-the-art deblurring performance, but with less parameters and shorter running time. Yuan Yuan 0001, Dandan Ma |
CVPR | 1 |
| 2020 | Variational Context-Deformable ConvNets for Indoor Scene ParsingabstractContext information is critical for image semantic segmentation. Especially in indoor scenes, the large variation of object scales makes spatial-context an important factor for improving the segmentation performance. Thus, in this paper, we propose a novel variational context-deformable (VCD) module to learn adaptive receptive-field in a structured fashion. Different from standard ConvNets, which share fixed-size spatial context for all pixels, the VCD module learns a deformable spatial-context with the guidance of depth information: depth information provides clues for identifying real local neighborhoods. Specifically, adaptive Gaussian kernels are learned with the guidance of multimodal information. By multiplying the learned Gaussian kernel with standard convolution filters, the VCD module can aggregate flexible spatial context for each pixel during convolution. The main contributions of this work are as follows: 1) a novel VCD module is proposed, which exploits learnable Gaussian kernels to enable feature learning with structured adaptive-context; 2) variational Bayesian probabilistic modeling is introduced for the training of VCD module, which can make it continuous and more stable; 3) a perspective-aware guidance module is designed to take advantage of multi-modal information for RGB-D segmentation. We evaluate the proposed approach on three widely-used datasets, and the performance improvement has shown the effectiveness of the proposed method. Zhitong Xiong, Yuan Yuan 0001, Nianhui Guo, Qi Wang 0009 |
CVPR | 2 |
| 2020 | Focus on Semantic Consistency for Cross-Domain Crowd UnderstandingabstractFor pixel-level crowd understanding, it is time-consuming and laborious in data collection and annotation. Some domain adaptation algorithms try to liberate it by training models with synthetic data, and the results in some recent works have proved the feasibility. However, we found that a mass of estimation errors in the background areas impede the performance of the existing methods. In this paper, we propose a domain adaptation method to eliminate it. According to the semantic consistency, a similar distribution in deep layer's features of the synthetic and real-world crowd area, we first introduce a semantic extractor to effectively distinguish crowd and background in high-level semantic information. Besides, to further enhance the adapted model, we adopt adversarial learning to align features in the semantic space. Experiments on three representative real datasets show that the proposed domain adaptation scheme achieves the state-of-the-art for cross-domain counting problems. Tao Han 0002, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 3 |
| 2020 | Video Frame Interpolation Via Residue RefinementabstractVideo frame interpolation achieves temporal super-resolution by generating smooth transitions between frames. Although great success has been achieved by deep neural networks, the synthesized images stills suffer from poor visual appearance and unsatisfactory artifacts. In this paper, we propose a novel network structure that leverages residue refinement and adaptive weight to synthesize in-between frames. The residue refinement technique is used for optical flow and image generation for higher accuracy and better visual appearance, while the adaptive weight map combines the forward and backward warped frames to reduce the artifacts. Moreover, all submodules in our method are implemented by U-Net with less depths, so the efficiency is guaranteed. Experiments on public datasets demonstrate the effectiveness and superiority of our method over the state-of-the-art approaches. Haopeng Li 0001, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2020 | Enhanced Non-Local Cascading Network with Attention Mechanism for Hyperspectral Image DenoisingabstractBecause of the complexity of imaging environment, hyper-spectral remote sensing images (HSIs) often suffer from different kinds of noise. Despite the success in natural image denoising, most of the existing CNN-based HSIs denoising methods still suffer from the problem of inadequate noise suppression and insufficient feature extraction. In this paper, a novel HSIs denoising algorithm based on an enhanced non-local cascading network with attention mechanism (ENCAM) is proposed, which can extract the joint spatial-spectral feature more effectively. The main contributions include: (1) the non-local structure is introduced to enlarge the receptive field to extract the spatial features more effectively; (2) multi-scale convolutions and channel attention module are applied to enhance extracted multi-scale features; (3) a cascading residual dense structure is used to extract different frequency features. Both of the theoretical analysis and the experiments indicate that the proposed method is superior to the other state-of-the-art methods on HSIs denoising. Hanwen Ma, Ganchao Liu, Yuan Yuan 0001 |
ICASSP | 3 |
| 2020 | Deep Image Deblurring Using Local Correlation BlockabstractDynamic scene deblurring is a challenging problem due to the various blurry source. Many deep learning based approaches try to train end-to-end deblurring networks, and achieve successful performance. However, the architectures and parameters of these methods are unchanged after training, so they need deeper network architectures and more parameters to adapt different blurry images, which increase the computational complexity. In this paper, we propose a local correlation block (LCBlock), which can adjust the weights of features adaptively according to the blurry inputs. And we use it to construct a dynamic scene deblurring network named LCNet. Experimental results show that the proposed LC-Net produces compariable performance with shorter running time and smaller network size, compared to state-of-the-art learning-based methods. Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2020 | MSPNET: Multi-Supervised Parallel Network for Crowd CountingabstractCrowd counting has a wide range of applications such as video surveillance and public safety. Many existing methods only focus on improving the accuracy of counting but ignore the importance of density maps. It's no doubt that a high-quality density map contains more information such as localization and movement of the crowd. In this paper, we propose a multi-supervised parallel network (MSPNet) to achieve high accuracy of crowd counting and generate high-quality density maps. We conduct multiple supervisions in the training process, which can supplement the details lost in pooling and up-sampling operations to improve the quality of density maps. In addition, to reduce the impact of background noise, the attention mechanism is employed to help the network focus on the crowd. Extensive experiments on two mainstream benchmarks show that MSPNet achieves significantly improvement over the state-of-the-art in terms of counting accuracy and the quality of density maps. Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2020 | A Novel Unsupervised Change Detection Approach Based On Spectral Transformation For Multispectral ImagesabstractChange detection (CD) for multispectral remote sensing images is an important approach to observe the changes of the earth. However, the same object usually has different spectra in multi-temporal images, which is one of the biggest challenges for CD. To overcome this problem, a novel unsupervised CD approach based on spectral transformation and joint spectral-spatial feature learning (STCD) is proposed for multispectral images in this paper. By exploring the relationship between imaging environment and the object spectra, the spectral transformation is used to suppress the phenomenon of “same object with different spectra”. Besides, a detection network with joint spectral-spatial feature learning is designed to extract the spectral-spatial features simultaneously to make the CD algorithm more robust. Both theoretical analyses and experiment results proved that the proposed STCD method is superior to the state-of-the-art unsupervised methods on multispectral images CD. Yuelin Zhang, Ganchao Liu, Yuan Yuan 0001 |
ICIP | 3 |
| 2020 | Open Set Domain Recognition via Attention-Based GCN and Semantic Matching OptimizationabstractOpen set domain recognition has got the attention in recent years. The task aims to specifically classify each sample in the practical unlabeled target domain, which consists of all known classes in the manually labeled source domain and target-specific unknown categories. The absence of annotated training data or auxiliary attribute information for unknown categories makes this task especially difficult. Moreover, exiting domain discrepancy in label space and data distribution further distracts the knowledge transferred from known classes to unknown classes. To address these issues, this work presents an end-to-end model based on attention-based GCN and semantic matching optimization, which first employs the attention mechanism to enable the central node to learn more discriminating representations from its neighbors in the knowledge graph. Moreover, a coarse-to-fine semantic matching optimization approach is proposed to progressively bridge the domain gap. Experimental results validate that the proposed model not only has superiority on recognizing the images of known and unknown classes, but also can adapt to various openness of the target domain. Xinxing He, Yuan Yuan 0001, Zhiyu Jiang |
ICPR | 2 |
| 2020 | Visual Localization Based on Remote Sensing Scene Matching with Siamese Feature Aggregation NetworkabstractThis paper presents a new framework with a siamese feature aggregation network (SFANet) for visual localization based on remote sensing scene matching. Specifically, the presented framework predicts the location of a query image by finding the matching remote sensing images with geographical information. We employ the fully convolutional networks (FCNs) and a siamese network of NetVLAD to aggregate local features and learn the global representations for images from different sources. A new soft margin loss function is established for the network. Geographic coordinates of the query images are obtained by calculating the similarity with satellite images. We also collect a multi-scale dataset that contains 136959 images from 45653 locations. Various experiments are carried out on it. Experimental results show the effectivity of the proposed method. Yuan Yuan 0001, Ganchao Liu |
IGARSS | 2 |
| 2020 | Detect Geographical Location by Multi-View Scene MatchingabstractWith the help of satellite geographical information, we present an innovative framework for localizing precisely while handling the varying views and different sources. We construct a convolutional neural network named Attentive Siamese-like Net (ASN) which can extract the multi-view scene representation. On top of it, a database retrieval system is established to search the realistic geographic coordinates quickly and reliably. For handling a complicated scene, a mass of visual data is collected from google satellite image and Google Earth Software. Experiments are carried out on two datasets of different scales and geographical range, which shows the superiority and effectiveness of our method. Yuan Yuan 0001, Ganchao Liu |
IGARSS | 2 |
| 2020 | Instance-Aware Remote Sensing Image Captioning with Cross-Hierarchy AttentionabstractThe spatial attention is a straightforward approach to enhance the performance for remote sensing image captioning. However, conventional spatial attention approaches consider only the attention distribution on one fixed coarse grid, resulting in the semantics of tiny objects can be easily ignored or disturbed during the visual feature extraction. Worse still, the fixed semantic level of conventional spatial attention limits the image understanding in different levels and perspectives, which is critical for tackling the huge diversity in remote sensing images. To address these issues, we propose a remote sensing image caption generator with instance-awareness and cross-hierarchy attention. 1) The instances awareness is achieved by introducing a multi-level feature architecture that contains the visual information of multi-level instance-possible regions and their surroundings. 2) Moreover, based on this multi-level feature extraction, a cross-hierarchy attention mechanism is proposed to prompt the decoder to dynamically focus on different semantic hierarchies and instances at each time step. The experimental results on public datasets demonstrate the superiority of proposed approach over existing methods. Chengze Wang, Zhiyu Jiang, Yuan Yuan 0001 |
IGARSS | 3 |
| 2020 | Weighted Hierarchical Sparse Representation for Hyperspectral Target DetectionabstractHyperspectral target detection has been widely studied in the field of remote sensing. However, background dictionary building issue and the correlation analysis of target and background dictionary issue have not been well studied. To tackle these issues, a Weighted Hierarchical Sparse Representation for hyperspectral target detection is proposed. The main contributions of this work are listed as follows. 1) Considering the insufficient representation of the traditional background dictionary building by dual concentric window structure, a hierarchical background dictionary is built considering the local and global spectral information simultaneously. 2) To reduce the impureness impact of background dictionary, target scores from target dictionary and background dictionary are weighted considered according to the dictionary quality. Three hyperspectral target detection data sets are utilized to verify the effectiveness of the proposed method. And the experimental results show a better performance when compared with the state-of-the-arts. Chenlu Wei, Zhiyu Jiang, Yuan Yuan 0001 |
IGARSS | 3 |
| 2020 | Unsupervised Semantic Aggregation and Deformable Template Matching for Semi-Supervised LearningabstractUnlabeled data learning has attracted considerable attention recently. However, it is still elusive to extract the expected high-level semantic feature with mere unsupervised learning. In the meantime, semi-supervised learning (SSL) demonstrates a promising future in leveraging few samples. In this paper, we combine both to propose an Unsupervised Semantic Aggregation and Deformable Template Matching (USADTM) framework for SSL, which strives to improve the classification performance with few labeled data and then reduce the cost in data annotating. Specifically, unsupervised semantic aggregation based on Triplet Mutual Information (T-MI) loss is explored to generate semantic labels for unlabeled data. Then the semantic labels are aligned to the actual class by the supervision of labeled data. Furthermore, a feature pool that stores the labeled samples is dynamically updated to assign proxy labels for unlabeled data, which are used as targets for cross-entropy minimization. Extensive experiments and analysis across four standard semi-supervised learning benchmarks validate that USADTM achieves top performance (e.g., 90.46% accuracy on CIFAR-10 with 40 labels and 95.20% accuracy with 250 labels). The code is released at https://github.com/taohan10200/USADTM. Tao Han 0002, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009 |
NeurIPS | 3 |
| 2020 | A dense connection based network for real-time object tracking
Yuwei Lu, Yuan Yuan 0001, Qi Wang 0009 |
Neurocomputing | 2 |
| 2020 | Gated forward refinement network for action segmentation
Dong Wang 0028, Yuan Yuan 0001, Qi Wang 0009 |
Neurocomputing | 2 |
| 2020 | MSN: Modality separation networks for RGB-D scene recognition
Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
Neurocomputing | 2 |
| 2020 | Adaptive forward vehicle collision warning based on driving behavior
Yuan Yuan 0001, Yuwei Lu, Qi Wang 0009 |
Neurocomputing | 1 |
| 2020 | Deep Gabor convolution network for person re-identification
Yuan Yuan 0001, Jian'an Zhang, Qi Wang 0009 |
Neurocomputing | 1 |
| 2020 | Subspace Clustering Constrained Sparse NMF for Hyperspectral UnmixingabstractAs one of the most important information of hyperspectral images (HSI), spatial information is usually simulated with the similarity among pixels to enhance the unmixing performance of nonnegative matrix factorization (NMF). Nevertheless, the similarity is generally calculated based on the Euclidean distance between pairwise pixels, which is sensitive to noise and fails in capturing subspace information of hyperspectral data. In addition, it is independent of the NMF framework. In this article, we propose a novel unmixing method called subspace clustering constrained sparse NMF (SC-NMF) for hyperspectral unmixing to more accurately extract endmembers and correspond abundances. First, the nonnegative subspace clustering is embedded into the NMF framework to learn a similar graph, which takes full advantage of the characteristics of the reconstructed data itself to extract the spatial correlation of pixels for unmixing. It is noteworthy that the similar graph and NMF will be simultaneously updated. Second, to mitigate the influence of noise in HSI, only the k largest values are retained in each self-expression vector. Finally, we use the idea of subspace clustering to extract endmembers by linearly combining of all pixels in spectral subspace, aiming at giving a reasonable physical significance to the endmembers. We evaluate the proposed SC-NMF on both synthetic and real hyperspectral data, and experimental results demonstrate that the proposed method is effective and superior by comparing with the state-of-the-art methods. Xiaoqiang Lu, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Gated and Axis-Concentrated Localization Network for Remote Sensing Object DetectionabstractIn the multicategory object detection task of high-resolution remote sensing images, small objects are always difficult to detect. This happens because the influence of location deviation on small object detection is greater than on large object detection. The reason is that, with the same intersection decrease between a predicted box and a true box, Intersection over Union (IoU) of small objects drops more than those of large objects. In order to address this challenge, we propose a new localization model to improve the location accuracy of small objects. This model is composed of two parts. First, a global feature gating process is proposed to implement a channel attention mechanism on local feature learning. This process takes full advantages of global features’ abundant semantics and local features’ spatial details. In this case, more effective information is selected for small object detection. Second, an axis-concentrated prediction (ACP) process is adopted to project convolutional feature maps into different spatial directions, so as to avoid interference between coordinate axes and improve the location accuracy. Then, coordinate prediction is implemented with a regression layer using the learned object representation. In our experiments, we explore the relationship between the detection accuracy and the object scale, and the results show that the performance improvements of small objects are distinct using our method. Compared with the classical deep learning detection models, the proposed gated axis-concentrated localization network (GACL Net) has the characteristic of focusing on small objects. Xiaoqiang Lu, Yuanlin Zhang 0003, Yuan Yuan 0001, Yachuang Feng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Attribute-Cooperated Convolutional Neural Network for Remote Sensing Image ClassificationabstractRemote sensing image (RSI) classification is one of the most important fields in RSI processing. It is well known that RSIs are very complicated due to its various kinds of contents. Therefore, it is very difficult to distinguish different scene categories with similar visual contents, like desert and bare land. To address hard negative categories, an attribute-cooperated convolutional neural network (ACCNN) is proposed to exploit attributes as additional guiding information. First, the classification branch extracts convolutional neural network feature, which is then utilized to recognize the RSI scene categories. Second, the attribute branch is proposed to make the network distinguish scene categories efficiently. The proposed attribute branch shares feature extraction layers with the classification branch and makes the classification branch aware of extra attribute information. Finally, the relationship branch constraints the relationship between the classification branch and the attribute branch. To exploit the attribute information, three attribute-classification data sets are generated (AC-AID, AC-UCM, and AC-Sydney). Experimental results show that the proposed method is competitive to state-of-the-art methods. The data sets are available at https://github.com/CrazyStoneonRoad/Attribute-Cooperated-Classification-Data sets. Yuanlin Zhang 0003, Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Memory-Augmented Temporal Dynamic Learning for Action RecognitionabstractHuman actions captured in video sequences contain two crucial factors for action recognition, i.e., visual appearance and motion dynamics. To model these two aspects, Convolutional and Recurrent Neural Networks (CNNs and RNNs) are adopted in most existing successful methods for recognizing actions. However, CNN based methods are limited in modeling long-term motion dynamics. RNNs are able to learn temporal motion dynamics but lack effective ways to tackle unsteady dynamics in long-duration motion. In this work, we propose a memory-augmented temporal dynamic learning network, which learns to write the most evident information into an external memory module and ignore irrelevant ones. In particular, we present a differential memory controller to make a discrete decision on whether the external memory module should be updated with current feature. The discrete memory controller takes in the memory history, context embedding and current feature as inputs and controls information flow into the external memory module. Additionally, we train this discrete memory controller using straight-through estimator. We evaluate this end-to-end system on benchmark datasets (UCF101 and HMDB51) of human action recognition. The experimental results show consistent improvements on both datasets over prior works and our baselines. Yuan Yuan 0001, Dong Wang 0028, Qi Wang 0009 |
AAAI | 1 |
| 2019 | ACM: Adaptive Cross-Modal Graph Convolutional Neural Networks for RGB-D Scene RecognitionabstractRGB image classification has achieved significant performance improvement with the resurge of deep convolutional neural networks. However, mono-modal deep models for RGB image still have several limitations when applied to RGB-D scene recognition. 1) Images for scene classification usually contain more than one typical object with flexible spatial distribution, so the object-level local features should also be considered in addition to global scene representation. 2) Multi-modal features in RGB-D scene classification are still under-utilized. Simply combining these modal-specific features suffers from the semantic gaps between different modalities. 3) Most existing methods neglect the complex relationships among multiple modality features. Considering these limitations, this paper proposes an adaptive crossmodal (ACM) feature learning framework based on graph convolutional neural networks for RGB-D scene recognition. In order to make better use of the modal-specific cues, this approach mines the intra-modality relationships among the selected local features from one modality. To leverage the multi-modal knowledge more effectively, the proposed approach models the inter-modality relationships between two modalities through the cross-modal graph (CMG). We evaluate the proposed method on two public RGB-D scene classification datasets: SUN-RGBD and NYUD V2, and the proposed method achieves state-of-the-art performance. Yuan Yuan 0001, Zhitong Xiong, Qi Wang 0009 |
AAAI | 1 |
| 2019 | Learning From Synthetic Data for Crowd Counting in the WildabstractRecently, counting the number of people for crowd scenes is a hot topic because of its widespread applications (e.g. video surveillance, public security). It is a difficult task in the wild: changeable environment, large-range number of people cause the current methods can not work well. In addition, due to the scarce data, many methods suffer from over-fitting to a different extent. To remedy the above two problems, firstly, we develop a data collector and labeler, which can generate the synthetic crowd scenes and simultaneously annotate them without any manpower. Based on it, we build a large-scale, diverse synthetic dataset. Secondly, we propose two schemes that exploit the synthetic data to boost the performance of crowd counting in the wild: 1) pretrain a crowd counter on the synthetic data, then finetune it using the real data, which significantly prompts the model's performance on real data; 2) propose a crowd counting method via domain adaptation, which can free humans from heavy data annotations. Extensive experiments show that the first method achieves the state-of-the-art performance on four real datasets, and the second outperforms our baselines. The dataset and source code are available at https://gjy3035.github.io/GCC-CL/. Qi Wang 0009, Junyu Gao 0001, Wei Lin 0018, Yuan Yuan 0001 |
CVPR | 4 |
| 2019 | Learning by Inertia: Self-supervised Monocular Visual Odometry for Road VehiclesabstractIn this paper, we present iDVO (inertia-embedded deep visual odometry), a self-supervised learning based monocular visual odometry (VO) for road vehicles. When modelling the geometric consistency within adjacent frames, most deep VO methods ignore the temporal continuity of the camera pose, which results in a very severe jagged fluctuation in the velocity curves. With the observation that road vehicles tend to perform smooth dynamic characteristics in most of the time, we design the inertia loss function to describe the abnormal motion variation, which assists the model to learn the consecutiveness from long-term camera ego-motion. Based on the recurrent convolutional neural network (RCNN) architecture, our method implicitly models the dynamics of road vehicles and the temporal consecutiveness by the extended Long Short-Term Memory (LSTM) block. Furthermore, we develop the dynamic hard-edge mask to handle the non-consistency in fast camera motion by blocking the boundary part and which generates more efficiency in the whole non-consistency mask. The proposed method is evaluated on the KITTI dataset, and the results demonstrate state-of-the-art performance with respect to other monocular deep VO and SLAM approaches. Chengze Wang, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2019 | Hyperspectral Unmixing VIA L1/4 Sparsity-Constrained Multilayer NMFabstractHyperspectral unmixing, by extracting the fractional abundances of endmembers from the hyperspectral image (HSI), has raised wide attention in recent years. In last decade, nonnegative matrix factorization (NMF) have been intensively studied for solving spectral unmixing problem. In this paper, we extend the multilayer NMF method by incorporating the L1/4sparsity constraint, named L1/4-MLNMF. The L1/4regularizer induces sparsity effectively. We propose an iterative estimation algorithm for L1/4-MLNMF, which provides sparser and more accurate results than MLNMF. Experiments on a synthetic dataset and a real dataset show that the prposed method outperforms the similar competitors. Qi Wang 0009, Yuan Yuan 0001 |
IGARSS | 3 |
| 2019 | Local and Global Feature Learning for Subtle Facial Expression Recognition from Attention Perspective
Yuan Yuan 0001, Yachuang Feng |
PRCV (2) | 2 |
| 2019 | Metric learning by simultaneously learning linear transformation matrix and weight matrix for person re-identificationabstractMahalanobis metric learning is one of the most popular methods for person re‐identification. Most existing metric learning methods regularly formulate the person re‐identification as an unconstrained optimisation problem and the constraints on the Mahalanobis matrix are seldom imposed. In addition, weights are often used to model the relationships between different variables but they often suffer from boundedness caused by their hand‐designed feature. Taking the above two disadvantages into consideration, the authors propose a new metric learning method for person re‐identification, which formulates the metric learning problem as a constrained optimisation problem by imposing a constraint on the linear transformation matrix. Furthermore, they treat the weights as unknown variables and introduce a weight learning method instead of designing weight intuitively. Finally, they evaluate the proposed method on two challenging person re‐identification databases and show that it performs favourably against the state‐of‐the‐art approaches. Jian'an Zhang, Qi Wang 0009, Yuan Yuan 0001 |
IET Comput. Vis. | 3 |
| 2019 | Muti-stage learning for gender and age prediction
Jie Fang 0001, Yuan Yuan 0001, Xiaoqiang Lu, Yachuang Feng |
Neurocomputing | 2 |
| 2019 | SCAR: Spatial-/channel-wise attention regression networks for crowd counting
Junyu Gao 0001, Qi Wang 0009, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2019 | Robust Space-Frequency Joint Representation for Remote Sensing Image Scene ClassificationabstractRemote sensing image scene classification is a fundamental problem, which aims to label an image with a specific semantic category automatically. Recent progress on remote sensing image scene classification is substantial, benefitting mostly from the powerful feature extraction capability of convolutional neural networks (CNNs). Even though these CNN-based methods have achieved competitive performances, they only construct the representation of the image in location-sensitive space-domain. As a result, their representations are not robust to rotation-variant remote sensing images, which influence the classification accuracy. In this paper, we propose a novel feature representation method by introducing a frequency-domain branch to the traditional only-space-domain architecture. Our framework takes full advantages of discriminative features from space domain and location-robust features from the frequency domain, providing more advanced representations through an additional joint learning module, a property that is critically needed to perform remote sensing image scene classification. Additionally, our method produces satisfactory performances on four public and challenging remote sensing image scene data sets, Sydney, UC-Merced, WHU-RS19, and AID. Jie Fang 0001, Yuan Yuan 0001, Xiaoqiang Lu, Yachuang Feng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Remote Sensing Image Scene Classification Using Rearranged Local FeaturesabstractRemote sensing image scene classification is a fundamental problem, which aims to label an image with a specific semantic category automatically. Recently, deep learning methods have achieved competitive performance for remote sensing image scene classification, especially the methods based on a convolutional neural network (CNN). However, most of the existing CNN methods only use feature vectors of the last fully connected layer. They give more importance to global information and ignore local information of images. It is common that some images belong to different categories, although they own similar global features. The reason is that the category of an image may be highly related to local features, other than the global feature. To address this problem, a method based on rearranged local features is proposed in this paper. First, outputs of the last convolutional layer and the last fully connected layer are employed to depict the local and global information, respectively. After that, the remote sensing images are clustered to several collections using their global features. For each collection, local features of an image are rearranged according to their similarities with local features of the cluster center. In addition, a fusion strategy is proposed to combine global and local features for enhancing the image representation. The proposed method surpasses the state of the arts on four public and challenging data sets: UC-Merced, WHU-RS19, Sydney, and AID. Yuan Yuan 0001, Jie Fang 0001, Xiaoqiang Lu, Yachuang Feng |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Hierarchical and Robust Convolutional Neural Network for Very High-Resolution Remote Sensing Object DetectionabstractObject detection is a basic issue of very high-resolution remote sensing images (RSIs) for automatically labeling objects. At present, deep learning has gradually gained the competitive advantage for remote sensing object detection, especially based on convolutional neural networks (CNNs). Most of the existing methods use the global information in the fully connected feature vector and ignore the local information in the convolutional feature cubes. However, the local information can provide spatial information, which is helpful for accurate localization. In addition, there are variable factors, such as rotation and scaling, which affect the object detection accuracy in RSIs. In order to solve these problems, this paper presents a hierarchical robust CNN. First, multiscale convolutional features are extracted to represent the hierarchical spatial semantic information. Second, multiple fully connected layer features are stacked together so as to improve the rotation and scaling robustness. Experiments on two data sets have shown the effectiveness of our method. In addition, a large-scale high-resolution remote sensing object detection data set is established to make up for the current situation that the existing data set is insufficient or too small. The data set is available athttps://github.com/CrazyStoneonRoad/TGRS-HRRSD-Dataset. Yuanlin Zhang 0003, Yuan Yuan 0001, Yachuang Feng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Hyperspectral Image Denoising by Fusing the Selected Related BandsabstractHyperspectral images (HSIs) convey more useful information than RGB or gray images, which are widely used in many remote sensing tasks. In real scenarios, HSIs are inevitably corrupted by noise because of sensors' imperfectness or atmospheric influence. Recently, many HSI denoising methods have been proposed to utilize the interband information between different spectral bands. However, these methods regard the HSI as a whole and treat the different spectral bands with the same noise level. In fact, the noise levels in different bands are different. Especially, only few certain bands are corrupted by noise, named the target noised bands. Under this circumstance, an HSI denoising method is proposed by considering the band relationship and different noise levels. The target noised bands are adaptively denoised by fusing some selected bands. Specifically, some related but quality superior bands are selected according to the target noised bands. Then, the target noised bands can be denoised by fusing the selected related bands. Experimental results show that the proposed method achieves considerable performances in comparison with several state-of-the-art hyperspectral denoising methods. Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | A Deep Scene Representation for Aerial Scene ClassificationabstractAs a fundamental problem in earth observation, aerial scene classification tries to assign a specific semantic label to an aerial image. In recent years, the deep convolutional neural networks (CNNs) have shown advanced performances in aerial scene classification. The successful pretrained CNNs can be transferable to aerial images. However, global CNN activations may lack geometric invariance and, therefore, limit the improvement of aerial scene classification. To address this problem, this paper proposes a deep scene representation to achieve the invariance of CNN features and further enhance the discriminative power. The proposed method: 1) extracts CNN activations from the last convolutional layer of pretrained CNN; 2) performs multiscale pooling (MSP) on these activations; and 3) builds a holistic representation by the Fisher vector method. MSP is a simple and effective multiscale strategy, which enriches multiscale spatial information in affordable computational time. The proposed representation is particularly suited at aerial scenes and consistently outperforms global CNN activations without requiring feature adaptation. Extensive experiments on five aerial scene data sets indicate that the proposed method, even with a simple linear classifier, can achieve the state-of-the-art performance. Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | VSSA-NET: Vertical Spatial Sequence Attention Network for Traffic Sign DetectionabstractAlthough traffic sign detection has been studied for years and great progress has been made with the rise of deep learning technique, there are still many problems remaining to be addressed. For complicated real-world traffic scenes, there are two main challenges. First, traffic signs are usually small-sized objects, which makes them more difficult to detect than large ones; second, it is hard to distinguish false targets which resemble real traffic signs in complex street scenes without context information. To handle these problems, we propose a novel end-to-end deep learning method for traffic sign detection in complex environments. Our contributions are as follows: 1) we propose a multi-resolution feature fusion network architecture which exploits densely connected deconvolution layers with skip connections, and can learn more effective features for a small-size object and 2) we frame the traffic sign detection as a spatial sequence classification and regression task, and propose a vertical spatial sequence attention module to gain more context information for better detection performance. To comprehensively evaluate the proposed method, we experiment on several traffic sign datasets as well as the general object detection dataset, and the results have shown the effectiveness of our proposed method. Yuan Yuan 0001, Zhitong Xiong, Qi Wang 0009 |
IEEE Trans. Image Process. | 1 |
| 2019 | Spatial Structure Preserving Feature Pyramid Network for Semantic Image SegmentationabstractRecently, progress on semantic image segmentation is substantial, benefiting from the rapid development of Convolutional Neural Networks. Semantic image segmentation approaches proposed lately have been mostly based on Fully convolutional Networks (FCNs). However, these FCN-based methods use large receptive fields and too many pooling layers to depict the discriminative semantic information of the images. Specifically, on one hand, convolutional kernel with large receptive field smooth the detailed edges, since too much contexture information is used to depict the “center pixel.” However, the pooling layer increases the receptive field through zooming out the latest feature maps, which loses many detailed information of the image, especially in the deeper layers of the network. These operations often cause low spatial resolution inside deep layers, which leads to spatially fragmented prediction. To address this problem, we exploit the inherent multi-scale and pyramidal hierarchy of deep convolutional networks to extract the feature maps with different resolutions and take full advantages of these feature maps via a gradually stacked fusing way. Specifically, for two adjacent convolutional layers, we upsample the features from deeper layer with stride of 2 and then stack them on the features from shallower layer. Then, a convolutional layer with kernels of 1× 1 is followed to fuse these stacked features. The fused feature preserves the spatial structure information of the image; meanwhile, it owns strong discriminative capability for pixel classification. Additionally, to further preserve the spatial structure information and regional connectivity of the predicted category label map, we propose a novel loss term for the network. In detail, two graph model-based spatial affinity matrixes are proposed, which are used to depict the pixel-level relationships in the input image and predicted category label map respectively, and then their cosine distance is backward propagated to the network. The proposed architecture, called spatial structure preserving feature pyramid network, significantly improves the spatial resolution of the predicted category label map for semantic image segmentation. The proposed method achieves state-of-the-art results on three public and challenging datasets for semantic image segmentation. Yuan Yuan 0001, Jie Fang 0001, Xiaoqiang Lu, Yachuang Feng |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Self-attention Learning for Person Re-identification
Minyue Jiang, Yuan Yuan 0001, Qi Wang 0009 |
BMVC | 2 |
| 2018 | ROI-wise Reverse Reweighting Network for Road Marking Detection
Yuan Yuan 0001, Qi Wang 0009 |
BMVC | 2 |
| 2018 | Forward Vehicle Collision Warning Based on Quick Camera CalibrationabstractForward Vehicle Collision Warning (FCW) is one of the most important functions for autonomous vehicles. In this procedure, vehicle detection and distance measurement are core components, requiring accurate localization and estimation. In this paper, we propose a simple but efficient forward vehicle collision warning framework by aggregating monocular distance measurement and precise vehicle detection. In order to obtain forward vehicle distance, a quick camera calibration method which only needs three physical points to calibrate related camera parameters is utilized. As for the forward vehicle detection, a multi-scale detection algorithm that regards the result of calibration as distance priori is proposed to improve the precision. Intensive experiments are conducted in our established real scene dataset and the results have demonstrated the effectiveness of the proposed framework. Yuwei Lu, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2018 | Cross-Modal Message Passing for Two-Stream FusionabstractProcessing and fusing information among multi-modal is a very useful technique to achieving high performance in many computer vision problem. In order to tackle multi-modal information more effectively, we introduce a novel framework for multi-modal fusion: Cross-modal Message Passing (CMMP). Specifically, we propose a cross-modal message passing mechanism to fuse two-stream network for action recognition, which composes of an appearance modal network (RGB image) and a motion modal (optical flow image) network. The objectives of individual networks in this framework are two-fold: a standard classification objective and a competing objective. The classification object ensures that each modal network predicts the true action category while the competing objective encourages each modal network to outperform the other one. We quantitatively show that the proposed CMMP fuse the traditional two-stream network more effectively, and outperforms all existing two-stream fusion method on UCF-101 and HMDB-51 datasets. Dong Wang 0028, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2018 | AI-NET: Attention Inception Neural Networks for Hyperspectral Image ClassificationabstractRecently, deep learning methods have dominated many fields thanks to its powerful discriminative feature learning ability. While for hyperspectral images (HSI) analysis, these deep neural networks methods suffer from overfitting as the number of labeled training samples are limited. Thus more efficient neural network architecture should be designed to improve the performance of HSI classification task. In this paper, a novel attention inception module is introduced to extract features dynamically from multi-resolution convolutional filters. The AI-NET constructed by stacking the proposed attention inception module can adaptively learn the network architecture by dynamically routing between the attention inception modules. By exploiting different spatial size convolutional filters and dynamic CNN architecture, more representative feature can be learned with limited training samples. Extensive experimental results have shown that the proposed method can adaptively adjust the network architecture and obtain better classification performance. Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
IGARSS | 2 |
| 2018 | Hyperspectral Band Selection with Convolutional Neural Network
Rui Cai 0002, Yuan Yuan 0001, Xiaoqiang Lu |
PRCV (4) | 2 |
| 2018 | Action recognition using spatial-optical data organization and sequential learning framework
Yuan Yuan 0001, Yang Zhao 0021, Qi Wang 0009 |
Neurocomputing | 1 |
| 2018 | Contour-aware network for semantic segmentation via adaptive depth
Zhiyu Jiang, Yuan Yuan 0001, Qi Wang 0009 |
Neurocomputing | 2 |
| 2018 | Incrementally perceiving hazards in driving
Yuan Yuan 0001, Jianwu Fang, Qi Wang 0009 |
Neurocomputing | 1 |
| 2018 | Spectral clustering based on iterative optimization for large-scale and high-dimensional data
Yang Zhao 0021, Yuan Yuan 0001, Feiping Nie 0001, Qi Wang 0009 |
Neurocomputing | 2 |
| 2018 | Locality constraint distance metric learning for traffic congestion detection
Qi Wang 0009, Jia Wan 0001, Yuan Yuan 0001 |
Pattern Recognit. | 3 |
| 2018 | Structured dictionary learning for abnormal event detection in crowded scenes
Yuan Yuan 0001, Yachuang Feng, Xiaoqiang Lu |
Pattern Recognit. | 1 |
| 2018 | Deep Metric Learning for Crowdedness RegressionabstractCross-scene regression tasks, such as congestion level detection and crowd counting, are useful but challenging. There are two main problems, which limit the performance of existing algorithms. The first one is that no appropriate congestion-related feature can reflect the real density in scenes. Though deep learning has been proved to be capable of extracting high level semantic representations, it is hard to converge on regression tasks, since the label is too weak to guide the learning of parameters in practice. Thus, many approaches utilize additional information, such as a density map, to guide the learning, which increases the effort of labeling. Another problem is that most existing methods are composed of several steps, for example, feature extraction and regression. Since the steps in the pipeline are separated, these methods face the problem of complex optimization. To remedy it, a deep metric learning-based regression method is proposed to extract density related features, and learn better distance measurement simultaneously. The proposed networks trained end-to-end for better optimization can be used for crowdedness regression tasks, including congestion level detection and crowd counting. Extensive experiments confirm the effectiveness of the proposed method. Qi Wang 0009, Jia Wan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | A Joint Convolutional Neural Networks and Context Transfer for Street Scenes LabelingabstractStreet scene understanding is an essential task for autonomous driving. One important step toward this direction is scene labeling, which annotates each pixel in the images with a correct class label. Although many approaches have been developed, there are still some weak points. First, many methods are based on the hand-crafted features whose image representation ability is limited. Second, they cannot label foreground objects accurately due to the data set bias. Third, in the refinement stage, the traditional Markov random filed inference is prone to over smoothness. For improving the above problems, this paper proposes a joint method of priori convolutional neural networks at superpixel level (called as “priori s-CNNs”) and soft restricted context transfer. Our contributions are threefold: 1) a priori s-CNNs model that learns priori location information at superpixel level is proposed to describe various objects discriminatingly; 2) a hierarchical data augmentation method is presented to alleviate data set bias in the priori s-CNNs training stage, which improves foreground objects labeling significantly; and 3) a soft restricted MRF energy function is defined to improve the priori s-CNNs model's labeling performance and reduce the over smoothness at the same time. The proposed approach is verified on CamVid data set (11 classes) and SIFT Flow Street data set (16 classes) and achieves a competitive performance. Qi Wang 0009, Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Embedding Structured Contour and Location Prior in Siamesed Fully Convolutional Networks for Road DetectionabstractRoad detection from the perspective of moving vehicles is a challenging issue in autonomous driving. Recently, many deep learning methods spring up for this task, because they can extract high-level local features to find road regions from raw RGB data, such as convolutional neural networks and fully convolutional networks (FCNs). However, how to detect the boundary of road accurately is still an intractable problem. In this paper, we propose siamesed FCNs (named “s-FCN-loc”), which is able to consider RGB-channel images, semantic contours, and location priors simultaneously to segment the road region elaborately. To be specific, the s-FCN-loc has two streams to process the original RGB images and contour maps, respectively. At the same time, the location prior is directly appended to the siamesed FCN to promote the final detection performance. Our contributions are threefold: 1) An s-FCN-loc is proposed that learns more discriminative features of road boundaries than the original FCN to detect more accurate road regions. 2) Location prior is viewed as a type of feature map and directly appended to the final feature map in s-FCN-loc to promote the detection performance effectively, which is easier than other traditional methods, namely, different priors for different inputs (image patches). 3) The convergent speed of training s-FCN-loc model is 30% faster than the original FCN because of the guidance of highly structured contours. The proposed approach is evaluated on the KITTI road detection benchmark and one-class road detection data set, and achieves a competitive result with the state of the arts. Qi Wang 0009, Junyu Gao 0001, Yuan Yuan 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | Asymmetric cross-view dictionary learning for person re-identificationabstractPerson re-identification is a critical yet challenging task in video surveillance which intends to match people over non-overlapping cameras. Most metric learning algorithms for person re-identification use symmetric matrix to project feature vectors into the same subspace to compute the similarity while ignoring the discrepancy between views. To solve this problem, we proposed an asymmetric cross-view matching algorithm with dictionary learning to alleviate the variations in human appearance across different views. Not only the views' dictionaries but also the persons' dictionary codes are constrained. Moreover, the `between-class' and the `within-class' distance are taken into consideration which makes the forming dictionary codes more robust and discriminative than the original feature vectors. The effectiveness of our approach is validated on the VIPeR and CUHK01 datasets. Experimental results show the proposed algorithm achieves compelling performance and asymmetric model plays an important role in the proposed approach. Minyue Jiang, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2017 | Traffic congestion analysis: A new PerspectiveabstractIn this paper, a new perspective of congestion is presented to promote the development of traffic video analysis. Our main contributions are threefold: a) An unified and quantifiable definition of congestion is proposed to describe the traffic state in video. b) Based on the definition, a congestion dataset which contains multiple traffic scenes is constructed as a platform for the research community. At the same time, a precise labeling method is introduced to get the ground truth of congestion level accurately. c) An algorithm based on Inverse Perspective Mapping (IPM) and pairwise regression is proposed to analyze traffic videos and serves as a baseline. We further compare the proposed method with two deep learning methods. Intensive experiments justify the effectiveness of the proposed method. Jia Wan 0001, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2017 | Largest center-specific margin for dimension reductionabstractDimensionality reduction plays an important role in solving the “curse of the dimensionality” and attracts a number of researchers in the past decades. In this paper, we proposed a new supervised linear dimensionality reduction method named largest center-specific margin (LCM) based on the intuition that after linear transformation, the distances between the points and their corresponding class centers should be small enough, and at the same time the distances between different unknown class centers should be as large as possible. On the basis of this observation, we take the unknown class centers into consideration for the first time and construct an optimization function to formulate this problem. In addition, we creatively transform the optimization objective function into a matrix function and solve the problem analytically. Finally, experiment results on three real datasets show the competitive performance of our algorithm. Jian'an Zhang, Yuan Yuan 0001, Feiping Nie 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2017 | HDPA: Hierarchical deep probability analysis for scene parsingabstractScene parsing is an important task in computer vision and many issues still need to be solved. One problem is about the non-unified framework for predicting things and stuff and the other one refers to the inadequate description of contextual information. In this paper, we address these issues by proposing a Hierarchical Deep Probability Analysis(HDPA) method which particularly exploits the power of probabilistic graphical model and deep convolutional neural network on pixel-level scene parsing. To be specific, an input image is initially segmented and represented through a CNN framework under Gaussian pyramid. Then the graphical models are built under each scale and the labels are ultimately predicted by structural analysis. Three contributions are claimed: unified framework for scene labeling, hierarchical probabilistic graphical modeling and adequate contextual information consideration. Experiments on three benchmarks show that the proposed method outperforms the state-of-the-arts in scene parsing. Yuan Yuan 0001, Zhiyu Jiang, Qi Wang 0009 |
ICME | 1 |
| 2017 | Embedding structured contour and location prior in siamesed fully convolutional networks for road detectionabstractRoad detection from the perspective of moving vehicles is a challenging issue in autonomous driving. Recently, many deep learning methods spring up for this task because they can extract high-level local features to find road regions from raw RGB data, such as Convolutional Neural Networks (CNN) and Fully Convolutional Networks (FCN). However, how to detect the boundary of road accurately is still an intractable problem. In this paper, we propose a siamesed fully convolutional network (named as “s-FCN-loc”) based on VGG-net architecture, which is able to consider RGB-channel, semantic contour and location prior simultaneously to segment road region elaborately. To be specific, the s-FCN-loc has two streams to process original RGB images and contour maps respectively. At the same time, the location prior is directly appended to the last feature map to promote the final detection performance. Experiments demonstrate that the proposed s-FCN-loc can learn more discriminative features of road boundaries and converge 30% faster than the original FCN during the training stage. Finally, the proposed approach is evaluated on KITTI road detection benchmark, and achieves a competitive result. Junyu Gao 0001, Qi Wang 0009, Yuan Yuan 0001 |
ICRA | 3 |
| 2017 | A sparse dictionary learning method for hyperspectral anomaly detection with capped normabstractHyperspectral anomaly detection is playing an important role in remote sensing field. Most conventional detectors based on the Reed-Xiaoli (RX) method assume the background signature obeys a Gaussian distribution. However, it is definitely hard to be satisfied in practice. Moreover, background statistics is susceptible to contamination of anomalies in the processing windows, which may lead to many false alarms and sensitiveness to the size of windows. To solve these problems, a novel sparse dictionary learning hyperspectral anomaly detection method with capped norm constraint is proposed. Contributions are claimed in threefold: 1) requiring no assumptions on the background distribution makes the method more adaptive to different scenes; 2) benefiting from the capped norm our method has a stronger distinctiveness to anomalies; and 3) it also has better adaptability to detect different sizes of anomalies without using the sliding dual window. The extensive experimental results demonstrate the desirable performance of our method. Dandan Ma, Yuan Yuan 0001, Qi Wang 0009 |
IGARSS | 2 |
| 2017 | JM-Net and Cluster-SVM for Aerial Scene ClassificationabstractAerial scene classification, which is a fundamental problem for remote sensing imagery, can automatically label an aerial image with a specific semantic category. Although deep learning has achieved competitive performance for aerial scene classification, training the conventional neural networks with aerial datasets will easily stick in overtting and local minimum. Because the aerial datasets only contain a few hundreds or thousands images, meanwhile the conventional networks usually contain millions of parameters to be trained. To address the problem, a novel convolutional neural network named JM-Net is proposed in this paper, which has different size of convolution kernels in same layer and ignores the fully convolytion layer, so it has fewer parameters and can be trained well on aerial datasets. Additionally, Cluster-SVM, a strategy to improve the accuracy and speed up the classification is used in the specific task. Finally, our method suparssed the state-of-art result on the challenging AID dataset while cost shorter time and used smaller storage space. Xiaoqiang Lu, Yuan Yuan 0001, Jie Fang 0001 |
IJCAI | 2 |
| 2017 | Convolutional 2D LDA for Nonlinear Dimensionality ReductionabstractRepresenting high-volume and high-order data is an essential problem, especially in machine learning field. Although existing two-dimensional (2D) discriminant analysis achieves promising performance, the single and linear projection features make it difficult to analyze more complex data. In this paper, we propose a novel convolutional two-dimensional linear discriminant analysis (2D LDA) method for data representation. In order to deal with nonlinear data, a specially designed Convolutional Neural Networks (CNN) is presented, which can be proved having the equivalent objective function with common 2D LDA. In this way, the discriminant ability can benefit from not only the nonlinearity of Convolutional Neural Networks, but also the powerful learning process. Experiment results on several datasets show that the proposed method performs better than other state-of-the-art methods in terms of classification accuracy. Qi Wang 0009, Zequn Qin, Feiping Nie 0001, Yuan Yuan 0001 |
IJCAI | 4 |
| 2017 | Learning deep event models for crowd anomaly detection
Yachuang Feng, Yuan Yuan 0001, Xiaoqiang Lu |
Neurocomputing | 2 |
| 2017 | Joint Dictionary Learning for Multispectral Change DetectionabstractChange detection is one of the most important applications of remote sensing technology. It is a challenging task due to the obvious variations in the radiometric value of spectral signature and the limited capability of utilizing spectral information. In this paper, an improved sparse coding method for change detection is proposed. The intuition of the proposed method is that unchanged pixels in different images can be well reconstructed by the joint dictionary, which corresponds to knowledge of unchanged pixels, while changed pixels cannot. First, a query image pair is projected onto the joint dictionary to constitute the knowledge of unchanged pixels. Then reconstruction error is obtained to discriminate between the changed and unchanged pixels in the different images. To select the proper thresholds for determining changed regions, an automatic threshold selection strategy is presented by minimizing the reconstruction errors of the changed pixels. Adequate experiments on multispectral data have been tested, and the experimental results compared with the state-of-the-art methods prove the superiority of the proposed method. Contributions of the proposed method can be summarized as follows: 1) joint dictionary learning is proposed to explore the intrinsic information of different images for change detection. In this case, change detection can be transformed as a sparse representation problem. To the authors' knowledge, few publications utilize joint learning dictionary in change detection; 2) an automatic threshold selection strategy is presented, which minimizes the reconstruction errors of the changed pixels without the prior assumption of the spectral signature. As a result, the threshold value provided by the proposed method can adapt to different data due to the characteristic of joint dictionary learning; and 3) the proposed method makes no prior assumption of the modeling and the handling of the spectral signature, which can be adapted to different data. Xiaoqiang Lu, Yuan Yuan 0001, Xiangtao Zheng |
IEEE Trans. Cybern. | 2 |
| 2017 | Statistical Hypothesis Detector for Abnormal Event Detection in Crowded ScenesabstractAbnormal event detection is now a challenging task, especially for crowded scenes. Many existing methods learn a normal event model in the training phase, and events which cannot be well represented are treated as abnormalities. However, they fail to make use of abnormal event patterns, which are elements to comprise abnormal events. Moreover, normal patterns in testing videos may be divergent from training ones, due to the existence of abnormalities. To address these problems, in this paper, an abnormality detector is proposed to detect abnormal events based on a statistical hypothesis test. The proposed detector treats each sample as a combination of a set of event patterns. Due to the unavailability of labeled abnormalities for training, abnormal patterns are adaptively extracted from incoming unlabeled testing samples. Contributions of this paper are listed as follows: 1) we introduce the idea of a statistical hypothesis test into the framework of abnormality detection, and abnormal events are identified as ones containing abnormal event patterns while possessing high abnormality detector scores; 2) due to the complexity of video events, noise seldom follows a simple distribution. For this reason, we approximate the complex noise distribution by employing a mixture of Gaussian. This benefits the modeling of video events and improves abnormality detection accuracies; and 3) because of the existence of abnormalities, there are always some unusually occurring normal events in the testing videos, which differ from the training ones. To represent normal events precisely, an online updating strategy is proposed to cover these cases in the normal event patterns. As a result, false detections are eliminated mostly. Extensive experiments and comparisons with state-of-the-art methods verify the effectiveness of the proposed algorithm. Yuan Yuan 0001, Yachuang Feng, Xiaoqiang Lu |
IEEE Trans. Cybern. | 1 |
| 2017 | Remote Sensing Scene Classification by Unsupervised Representation LearningabstractWith the rapid development of the satellite sensor technology, high spatial resolution remote sensing (HSR) data have attracted extensive attention in military and civilian applications. In order to make full use of these data, remote sensing scene classification becomes an important and necessary precedent task. In this paper, an unsupervised representation learning method is proposed to investigate deconvolution networks for remote sensing scene classification. First, a shallow weighted deconvolution network is utilized to learn a set of feature maps and filters for each image by minimizing the reconstruction error between the input image and the convolution result. The learned feature maps can capture the abundant edge and texture information of high spatial resolution images, which is definitely important for remote sensing images. After that, the spatial pyramid model (SPM) is used to aggregate features at different scales to maintain the spatial layout of HSR image scene. A discriminative representation for HSR image is obtained by combining the proposed weighted deconvolution model and SPM. Finally, the representation vector is input into a support vector machine to finish classification. We apply our method on two challenging HSR image data sets: the UCMerced data set with 21 scene categories and the Sydney data set with seven land-use categories. All the experimental results achieved by the proposed method outperform most state of the arts, which demonstrates the effectiveness of the proposed method. Xiaoqiang Lu, Xiangtao Zheng, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Dimensionality Reduction by Spatial-Spectral Preservation in Selected BandsabstractDimensionality reduction (DR) has attracted extensive attention since it provides discriminative information of hyperspectral images (HSI) and reduces the computational burden. Though DR has gained rapid development in recent years, it is difficult to achieve higher classification accuracy while preserving the relevant original information of the spectral bands. To relieve this limitation, in this paper, a different DR framework is proposed to perform feature extraction on the selected bands. The proposed method uses determinantal point process to select the representative bands and to preserve the relevant original information of the spectral bands. The performance of classification is further improved by performing multiple Laplacian eigenmaps (LEs) on the selected bands. Different from the traditional LEs, multiple Laplacian matrices in this paper are defined by encoding spatial-spectral proximity on each band. A common low-dimensional representation is generated to capture the joint manifold structure from multiple Laplacian matrices. Experimental results on three real-world HSIs demonstrate that the proposed framework can lead to a significant advancement in HSI classification compared with the state-of-the-art methods. Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Discovering Diverse Subset for Unsupervised Hyperspectral Band SelectionabstractBand selection, as a special case of the feature selection problem, tries to remove redundant bands and select a few important bands to represent the whole image cube. This has attracted much attention, since the selected bands provide discriminative information for further applications and reduce the computational burden. Though hyperspectral band selection has gained rapid development in recent years, it is still a challenging task because of the following requirements: 1) an effective model can capture the underlying relations between different high-dimensional spectral bands; 2) a fast and robust measure function can adapt to general hyperspectral tasks; and 3) an efficient search strategy can find the desired selected bands in reasonable computational time. To satisfy these requirements, a multigraph determinantal point process (MDPP) model is proposed to capture the full structure between different bands and efficiently find the optimal band subset in extensive hyperspectral applications. There are three main contributions: 1) graphical model is naturally transferred to address band selection problem by the proposed MDPP; 2) multiple graphs are designed to capture the intrinsic relationships between hyperspectral bands; and 3) mixture DPP is proposed to model the multiple dependencies in the proposed multiple graphs, and offers an efficient search strategy to select the optimal bands. To verify the superiority of the proposed method, experiments have been conducted on three hyperspectral applications, such as hyperspectral classification, anomaly detection, and target detection. The reliability of the proposed method in generic hyperspectral tasks is experimentally proved on four real-world hyperspectral data sets. Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Image Process. | 1 |
| 2017 | Tracking as a Whole: Multi-Target Tracking by Modeling Group Behavior With Sequential DetectionabstractVideo-based vehicle detection and tracking is one of the most important components for intelligent transportation systems. When it comes to road junctions, the problem becomes even more difficult due to the occlusions and complex interactions among vehicles. In order to get a precise detection and tracking result, in this paper we propose a novel tracking-by-detection framework. In the detection stage, we present a sequential detection model to deal with serious occlusions. In the tracking stage, we model group behavior to treat complex interactions with overlaps and ambiguities. The main contributions of this paper are twofold: 1) shape prior is exploited in the sequential detection model to tackle occlusions in crowded scene and 2) traffic force is defined in the traffic scene to model group behavior, and it can assist to handle complex interactions among vehicles. We evaluate the proposed approach on real surveillance videos at road junctions and the performance has demonstrated the effectiveness of our method. Yuan Yuan 0001, Yuwei Lu, Qi Wang 0009 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Anomaly Detection in Traffic Scenes via Spatial-Aware Motion ReconstructionabstractAnomaly detection from a driver's perspective when driving is important to autonomous vehicles. As a part of Advanced Driver Assistance Systems (ADAS), it can remind the driver about dangers in a timely manner. Compared with traditional studied scenes such as a university campus and market surveillance videos, it is difficult to detect an abnormal event from a driver's perspective due to camera waggle, abidingly moving background, drastic change of vehicle velocity, etc. To tackle these specific problems, this paper proposes a spatial localization constrained sparse coding approach for anomaly detection in traffic scenes, which first measures the abnormality of motion orientation and magnitude, respectively, and then fuses these two aspects to obtain a robust detection result. The main contributions are threefold, as follows. 1) This work describes the motion orientation and magnitude of the object, respectively, in a new way, which is demonstrated to be better than the traditional motion descriptors. 2) The spatial localization of an object is taken into account considering the sparse reconstruction framework, which utilizes the scene's structural information and outperforms the conventional sparse coding methods. 3) Results of motion orientation and magnitude are adaptively weighted and fused by a Bayesian model, which makes the proposed method more robust and able to handle more kinds of abnormal events. The efficiency and effectiveness of the proposed method are validated by testing on nine difficult video sequences that we captured ourselves. Observed from the experimental results, the proposed method is more effective and efficient than the popular competitors and yields a higher performance. Yuan Yuan 0001, Dong Wang 0028, Qi Wang 0009 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | An Incremental Framework for Video-Based Traffic Sign Detection, Tracking, and RecognitionabstractVideo-based traffic sign detection, tracking, and recognition is one of the important components for the intelligent transport systems. Extensive research has shown that pretty good performance can be obtained on public data sets by various state-of-the-art approaches, especially the deep learning methods. However, deep learning methods require extensive computing resources. In addition, these approaches mostly concentrate on single image detection and recognition task, which is not applicable in real-world applications. Different from previous research, we introduce a unified incremental computational framework for traffic sign detection, tracking, and recognition task using the mono-camera mounted on a moving vehicle under non-stationary environments. The main contributions of this paper are threefold: (1) to enhance detection performance by utilizing the contextual information, this paper innovatively utilizes the spatial distribution prior of the traffic signs; (2) to improve the tracking performance and localization accuracy under non-stationary environments, a new efficient incremental framework containing off-line detector, online detector, and motion model predictor together is designed for traffic sign detection and tracking simultaneously; and (3) to get a more stable classification output, a scale-based intra-frame fusion method is proposed. We evaluate our method on two public data sets and the performance has shown that the proposed system can obtain results comparable with the deep learning method with less computing resource in a near-real-time manner. Yuan Yuan 0001, Zhitong Xiong, Qi Wang 0009 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | Deep Representation for Abnormal Event Detection in Crowded ScenesabstractAbnormal event detection is extremely important, especially for video surveillance. Nowadays, many detectors have been proposed based on hand-crafted features. However, it remains challenging to effectively distinguish abnormal events from normal ones. This paper proposes a deep representation based algorithm which extracts features in an unsupervised fashion. Specially, appearance, texture, and short-term motion features are automatically learned and fused with stacked denoising autoencoders. Subsequently, long-term temporal clues are modeled with a long short-term memory (LSTM) recurrent network, in order to discover meaningful regularities of video events. The abnormal events are identified as samples which disobey these regularities. Moreover, this paper proposes a spatial anomaly detection strategy via manifold ranking, aiming at excluding false alarms. Experiments and comparisons on real world datasets show that the proposed algorithm outperforms state of the arts for the abnormal event detection problem in crowded scenes. Yachuang Feng, Yuan Yuan 0001, Xiaoqiang Lu |
ACM Multimedia | 2 |
| 2016 | A target detection method for hyperspectral image based on mixture noise model
Xiangtao Zheng, Yuan Yuan 0001, Xiaoqiang Lu |
Neurocomputing | 2 |
| 2016 | Action recognition by joint learning
Yuan Yuan 0001, Lei Qi 0004, Xiaoqiang Lu |
Image Vis. Comput. | 1 |
| 2016 | Congested scene classification via efficient unsupervised feature learning and density estimation
Yuan Yuan 0001, Jia Wan 0001, Qi Wang 0009 |
Pattern Recognit. | 1 |
| 2016 | A discriminative representation for human action recognition
Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu |
Pattern Recognit. | 1 |
| 2016 | Relevance and irrelevance graph based marginal Fisher analysis for image search reranking
Zhong Ji, Yanwei Pang, Yuan Yuan 0001 |
Signal Process. | 3 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 33 |
| 2016 | Hyperspectral Image Classification via Multitask Joint Sparse Representation and Stepwise MRF OptimizationabstractHyperspectral image (HSI) classification is a crucial issue in remote sensing. Accurate classification benefits a large number of applications such as land use analysis and marine resource utilization. But high data correlation brings difficulty to reliable classification, especially for HSI with abundant spectral information. Furthermore, the traditional methods often fail to well consider the spatial coherency of HSI that also limits the classification performance. To address these inherent obstacles, a novel spectral-spatial classification scheme is proposed in this paper. The proposed method mainly focuses on multitask joint sparse representation (MJSR) and a stepwise Markov random filed framework, which are claimed to be two main contributions in this procedure. First, the MJSR not only reduces the spectral redundancy, but also retains necessary correlation in spectral field during classification. Second, the stepwise optimization further explores the spatial correlation that significantly enhances the classification accuracy and robustness. As far as several universal quality evaluation indexes are concerned, the experimental results on Indian Pines and Pavia University demonstrate the superiority of our method compared with the state-of-the-art competitors. Yuan Yuan 0001, Jianzhe Lin, Qi Wang 0009 |
IEEE Trans. Cybern. | 1 |
| 2016 | Hyperspectral Anomaly Detection by Graph Pixel SelectionabstractHyperspectral anomaly detection (AD) is an important problem in remote sensing field. It can make full use of the spectral differences to discover certain potential interesting regions without any target priors. Traditional Mahalanobis-distance-based anomaly detectors assume the background spectrum distribution conforms to a Gaussian distribution. However, this and other similar distributions may not be satisfied for the real hyperspectral images. Moreover, the background statistics are susceptible to contamination of anomaly targets which will lead to a high false-positive rate. To address these intrinsic problems, this paper proposes a novel AD method based on the graph theory. We first construct a vertex- and edge-weighted graph and then utilize a pixel selection process to locate the anomaly targets. Two contributions are claimed in this paper: 1) no background distributions are required which makes the method more adaptive and 2) both the vertex and edge weights are considered which enables a more accurate detection performance and better robustness to noise. Intensive experiments on the simulated and real hyperspectral images demonstrate that the proposed method outperforms other benchmark competitors. In addition, the robustness of the proposed method has been validated by using various window sizes. This experimental result also demonstrates the valuable characteristic of less computational complexity and less parameter tuning for real applications. Yuan Yuan 0001, Dandan Ma, Qi Wang 0009 |
IEEE Trans. Cybern. | 1 |
| 2016 | Unsupervised Band Selection Based on Evolutionary Multiobjective Optimization for Hyperspectral ImagesabstractBand selection is an important preprocessing step for hyperspectral image processing. Many valid criteria have been proposed for band selection, and these criteria model band selection as a single-objective optimization problem. In this paper, a novel multiobjective model is first built for band selection. In this model, two objective functions with a conflicting relationship are designed. One objective function is set as information entropy to represent the information contained in the selected band subsets, and the other one is set as the number of selected bands. Then, based on this model, a new unsupervised band selection method called multiobjective optimization band selection (MOBS) is proposed. In the MOBS method, these two objective functions are optimized simultaneously by a multiobjective evolutionary algorithm to find the best tradeoff solutions. The proposed method shows two unique characters. It can obtain a series of band subsets with different numbers of bands in a single run to offer more options for decision makers. Moreover, these band subsets with different numbers of bands can communicate with each other and have a coevolutionary relationship, which means that they can be optimized in a cooperative way. Since it is unsupervised, the proposed algorithm is compared with some related and recent unsupervised methods for hyperspectral image band selection to evaluate the quality of the obtained band subsets. Experimental results show that the proposed method can generate a set of band subsets with different numbers of bands in a single run and that these band subsets have a stable good performance on classification for different data sets. Maoguo Gong, Mingyang Zhang 0002, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Dual-Clustering-Based Hyperspectral Band Selection by Contextual AnalysisabstractHyperspectral image (HSI) involves vast quantities of information that can help with the image analysis. However, this information has sometimes been proved to be redundant, considering specific applications such as HSI classification and anomaly detection. To address this problem, hyperspectral band selection is viewed as an effective dimensionality reduction method that can remove the redundant components of HSI. Various HSI band selection methods have been proposed recently, and the clustering-based method is a traditional one. This agglomerative method has been considered simple and straightforward, while the performance is generally inferior to the state of the art. To tackle the inherent drawbacks of the clustering-based band selection method, a new framework concerning on dual clustering is proposed in this paper. The main contribution can be concluded as follows: 1) a novel descriptor that reveals the context of HSI efficiently; 2) a dual clustering method that includes the contextual information in the clustering process; 3) a new strategy that selects the cluster representatives jointly considering the mutual effects of each cluster. Experimental results on three real-world HSIs verify the noticeable accuracy of the proposed method, with regard to the HSI classification application. The main comparison has been conducted among several recent clustering-based band selection methods and constraint-based band selection methods, demonstrating the superiority of the technique that we present. Yuan Yuan 0001, Jianzhe Lin, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Ensemble Manifold Rank Preserving for Acceleration-Based Human Activity RecognitionabstractWith the rapid development of mobile devices and pervasive computing technologies, acceleration-based human activity recognition, a difficult yet essential problem in mobile apps, has received intensive attention recently. Different acceleration signals for representing different activities or even a same activity have different attributes, which causes troubles in normalizing the signals. We thus cannot directly compare these signals with each other, because they are embedded in a nonmetric space. Therefore, we present a nonmetric scheme that retains discriminative and robust frequency domain information by developing a novel ensemble manifold rank preserving (EMRP) algorithm. EMRP simultaneously considers three aspects: 1) it encodes the local geometry using the ranking order information of intraclass samples distributed on local patches; 2) it keeps the discriminative information by maximizing the margin between samples of different classes; and 3) it finds the optimal linear combination of the alignment matrices to approximate the intrinsic manifold lied in the data. Experiments are conducted on the South China University of Technology naturalistic 3-D acceleration-based activity dataset and the naturalistic mobile-devices based human activity dataset to demonstrate the robustness and effectiveness of the new nonmetric scheme for acceleration-based human activity recognition. Dapeng Tao, Yuan Yuan 0001, Yang Xue 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Salient Band Selection for Hyperspectral Image Classification via Manifold RankingabstractSaliency detection has been a hot topic in recent years, and many efforts have been devoted in this area. Unfortunately, the results of saliency detection can hardly be utilized in general applications. The primary reason, we think, is unspecific definition of salient objects, which makes that the previously published methods cannot extend to practical applications. To solve this problem, we claim that saliency should be defined in a context and the salient band selection in hyperspectral image (HSI) is introduced as an example. Unfortunately, the traditional salient band selection methods suffer from the problem of inappropriate measurement of band difference. To tackle this problem, we propose to eliminate the drawbacks of traditional salient band selection methods by manifold ranking. It puts the band vectors in the more accurate manifold space and treats the saliency problem from a novel ranking perspective, which is considered to be the main contributions of this paper. To justify the effectiveness of the proposed method, experiments are conducted on three HSIs, and our method is compared with the six existing competitors. Results show that the proposed method is very effective and can achieve the best performance among the competitors. Qi Wang 0009, Jianzhe Lin, Yuan Yuan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Adaptive multi-bit quantization for hashing
Cheng Deng 0002, Huiru Deng, Xianglong Liu 0001, Yuan Yuan 0001 |
Neurocomputing | 4 |
| 2015 | MR image super-resolution via manifold regularized sparse learning
Xiaoqiang Lu, Zihan Huang, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2015 | Adaptive road detection via context-aware label transfer
Qi Wang 0009, Jianwu Fang, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2015 | Image quality assessment: A sparse learning way
Yuan Yuan 0001, Xiaoqiang Lu |
Neurocomputing | 1 |
| 2015 | Video-based road detection via online structural learning
Yuan Yuan 0001, Zhiyu Jiang, Qi Wang 0009 |
Neurocomputing | 1 |
| 2015 | Semi-supervised change detection method for multi-temporal hyperspectral images
Yuan Yuan 0001, Haobo Lv, Xiaoqiang Lu |
Neurocomputing | 1 |
| 2015 | Low-rank representation for 3D hyperspectral images analysis from map perspective
Yuan Yuan 0001, Xiaoqiang Lu |
Signal Process. | 1 |
| 2015 | Multi-spectral pedestrian detection
Yuan Yuan 0001, Xiaoqiang Lu |
Signal Process. | 1 |
| 2015 | Image Pair Analysis With Matrix-Value OperatorabstractImage pair analysis provides significant image pair priori which describes the dependency between training image pairs for various learning-based image processing. For avoiding the information loss caused by vectorizing training images, a novel matrix-value operator learning method is proposed for image pair analysis. Sample-dependent operators, named image pair operators (IPOs) by us, are employed to represent the local image-to-image dependency defined by each of the training image pairs. A linear combination of IPOs is learned via operator regression for representing the global dependency between input and output images defined by all of the training image pairs. The proposed operator learning method enjoys the image-level information of training image pairs because IPOs enable training images to be used without vectorizing during the learning and testing process. By applying the proposed algorithm in learning-based super-resolution, the efficiency and the effectiveness of the proposed algorithm in learning image pair information is verified by experimental results. Yi Tang 0003, Yuan Yuan 0001 |
IEEE Trans. Cybern. | 2 |
| 2015 | Label Image Constrained Multiatlas SelectionabstractMultiatlas based method is commonly used in medical image segmentation. In multiatlas based image segmentation, atlas selection and combination are considered as two key factors affecting the performance. Recently, manifold learning based atlas selection methods have emerged as very promising methods. However, due to the complexity of prostate structures in raw images, it is difficult to get accurate atlas selection results by only measuring the distance between raw images on the manifolds. Although the distance between the regions to be segmented across images can be readily obtained by the label images, it is infeasible to directly compute the distance between the test image (gray) and the label images (binary). This paper tries to address this problem by proposing a label image constrained atlas selection method, which exploits the label images to constrain the manifold projection of raw images. Analyzing the data point distribution of the selected atlases in the manifold subspace, a novel weight computation method for atlas combination is proposed. Compared with other related existing methods, the experimental results on prostate segmentation from T2w MRI showed that the selected atlases are closer to the target structure and more accurate segmentation were obtained by using our proposed method. Pingkun Yan, Yihui Cao, Yuan Yuan 0001, Baris Turkbey, Peter L. Choyke |
IEEE Trans. Cybern. | 3 |
| 2015 | Online Anomaly Detection in Crowd Scenes via Structure AnalysisabstractAbnormal behavior detection in crowd scenes is continuously a challenge in the field of computer vision. For tackling this problem, this paper starts from a novel structure modeling of crowd behavior. We first propose an informative structural context descriptor (SCD) for describing the crowd individual, which originally introduces the potential energy function of particle's interforce in solid-state physics to intuitively conduct vision contextual cueing. For computing the crowd SCD variation effectively, we then design a robust multi-object tracker to associate the targets in different frames, which employs the incremental analytical ability of the 3-D discrete cosine transform (DCT). By online spatial-temporal analyzing the SCD variation of the crowd, the abnormality is finally localized. Our contribution mainly lies on three aspects: 1) the new exploration of abnormal detection from structure modeling where the motion difference between individuals is computed by a novel selective histogram of optical flow that makes the proposed method can deal with more kinds of anomalies; 2) the SCD description that can effectively represent the relationship among the individuals; and 3) the 3-D DCT multi-object tracker that can robustly associate the limited number of (instead of all) targets which makes the tracking analysis in high density crowd situation feasible. Experimental results on several publicly available crowd video datasets verify the effectiveness of the proposed method. Yuan Yuan 0001, Jianwu Fang, Qi Wang 0009 |
IEEE Trans. Cybern. | 1 |
| 2015 | Substance Dependence Constrained Sparse NMF for Hyperspectral UnmixingabstractHyperspectral unmixing is one of the most important problems in analyzing remote sensing images, which aims to decompose a mixed pixel into a collection of constituent materials named endmembers and their corresponding fractional abundances. Recently, various methods have been proposed to incorporate sparse constraints into hyperspectral unmixing and achieve advanced performance. However, most of them ignore the complex distribution of substances in hyperspectral data so that they are only effective in limited cases. In this paper, the concept of substance dependence is introduced to help hyperspectral unmixing. Generally, substance dependence can be considered in a local region by K-nearest neighbors method. However, since substances of hyperspectral images are complicatedly distributed, number K of the most similar substances to each substance is difficult to decide. In this case, substance dependence should be considered in the whole data space, and the number of the K most similar substances to each substance can be adaptively determined by searching from the whole space. Through maintaining the substance dependence during unmixing, the abundances resulted from the proposed method are closer to the real fractions, which lead to better unmixing performance. The following contributions can be summarized. 1) The concept of substance dependence is proposed to describe the complicated relationship between substances in the hyperspectral image. 2) We propose substance dependence constrained sparse nonnegative matrix factorization (SDSNMF) for hyperspectral unmixing. Using SDSNMF, we meet or exceed state-of-the-art unmixing performance. 3) Adequate experiments on both synthetic and real hyperspectral data have been tested. Compared with the state-of-the-art methods, the experimental results prove the superiority of the proposed method. Yuan Yuan 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Fast Hyperspectral Anomaly Detection via High-Order 2-D Crossing FilterabstractAnomaly detection has been an important topic in hyperspectral image analysis. This technique is sometimes more preferable than the supervised target detection because it requires no a priori information for the interested materials. Many efforts have been made in this topic; however, they usually suffer from excessive time cost and a high false-positive rate. There are two major problems that lead to such a predicament. First, the construction of the background model and affinity estimation are often overcomplicated. Second, most of these methods have to impose a stringent assumption on the spectrum distribution of background; however, these assumptions cannot hold for all practical situations. Based on this consideration, this paper proposes a novel method allowing for fast yet accurate pixel-level hyperspectral anomaly detection. We claim the following main contributions: construct a high-order 2-D crossing approach to find the regions of rapid change in the spectrum, which runs without any a priori assumption; and design a low-complexity discrimination framework for fast hyperspectral anomaly detection, which can be implemented by a series of filtering operators with linear time cost. Experiments on three different hyperspectral images containing several pixel-level anomalies demonstrate the superiority of the proposed detector compared with the benchmark methods. Yuan Yuan 0001, Qi Wang 0009, Guokang Zhu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Spectral-Spatial Kernel Regularized for Hyperspectral Image DenoisingabstractNoise contamination is a ubiquitous problem in hyperspectral images (HSIs), which is a challenging and promising theme in many remote sensing applications. A large number of methods have been proposed to remove noise. Unfortunately, most denoising methods fail to take full advantages of the high spectral correlation and to simultaneously consider the specific noise distributions in HSIs. Recently, a spectral-spatial adaptive hyperspectral total variation (SSAHTV) was proposed and obtained promising results. However, the SSAHTV model is insensitive to the image details, which makes the edges blur. To overcome all of these drawbacks, a spectral-spatial kernel method for HSI denoising is proposed in this paper. The proposed method is inspired by the observation that the spectral-spatial information is highly redundant in HSIs, which is sufficient to estimate the clear images. In this paper, a spectral-spatial kernel regularization is proposed to maintain the spectral correlations in spectral dimension and to match the original structure between two spatial dimensions. Moreover, an adaptive mechanism is developed to balance the fidelity term according to different noise distributions in each band. Therefore, it cannot only suppress noise in the high-noise band but also preserve information in the low-noise band. The reliability of the proposed method in removing noise is experimentally proved on both simulated data and real data. Yuan Yuan 0001, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Hyperspectral Band Selection by Multitask Sparsity PursuitabstractHyperspectral images have been proved to be effective for a wide range of applications; however, the large volume and redundant information also bring a lot of inconvenience at the same time. To cope with this problem, hyperspectral band selection is a pertinent technique, which takes advantage of removing redundant components without compromising the original contents from the raw image cubes. Because of its usefulness, hyperspectral band selection has been successfully applied to many practical applications of hyperspectral remote sensing, such as land cover map generation and color visualization. This paper focuses on groupwise band selection and proposes a new framework, including the following contributions: 1) a smart yet intrinsic descriptor for efficient band representation; 2) an evolutionary strategy to handle the high computational burden associated with groupwise-selection-based methods; and 3) a novel MTSP-based criterion to evaluate the performance of each candidate band combination. To verify the superiority of the proposed framework, experiments have been conducted on both hyperspectral classification and color visualization. Experimental results on three real-world hyperspectral images demonstrate that the proposed framework can lead to a significant advancement in these two applications compared with other competitors. Yuan Yuan 0001, Guokang Zhu, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Scene Recognition by Manifold Regularized Deep Learning ArchitectureabstractScene recognition is an important problem in the field of computer vision, because it helps to narrow the gap between the computer and the human beings on scene understanding. Semantic modeling is a popular technique used to fill the semantic gap in scene recognition. However, most of the semantic modeling approaches learn shallow, one-layer representations for scene recognition, while ignoring the structural information related between images, often resulting in poor performance. Modeled after our own human visual system, as it is intended to inherit humanlike judgment, a manifold regularized deep architecture is proposed for scene recognition. The proposed deep architecture exploits the structural information of the data, making for a mapping between visible layer and hidden layer. By the proposed approach, a deep architecture could be designed to learn the high-level features for scene recognition in an unsupervised fashion. Experiments on standard data sets show that our method outperforms the state-of-the-art used for scene recognition. Yuan Yuan 0001, Lichao Mou, Xiaoqiang Lu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | In defense of iterated conditional mode for hyperspectral image classificationabstractHyperspectral image classification is one of the most significant topics in remote sensing. A large number of methods have been proposed to improve the classification accuracy. However, the improvement often comes at the cost of higher complexity. In this work, we mainly focus on the Markov Random Fields related paradigm, which involves a demanding energy minimization procedure. Traditional methods are prone to employ the advanced optimization techniques. On the contrary, this paper is in defense of a simple yet efficient method for hyperspectral image classification, Iterated Conditional Mode, which has been generally considered inferior to other state-of-the-art methods. Our purpose is successfully achieved by tackling two inherent drawbacks of ICM, sensitive label initialization and local minimum. We apply our method to three real-world hyperspectral images, and compare the results with those of state-of-the-art methods. The comparisons show that the proposed method outperforms its competitors. Jianzhe Lin, Qi Wang 0009, Yuan Yuan 0001 |
ICME | 3 |
| 2014 | Statistical quantization for similarity search
Qi Wang 0009, Guokang Zhu, Yuan Yuan 0001 |
Comput. Vis. Image Underst. | 3 |
| 2014 | Tag-Saliency: Combining bottom-up and top-down information for saliency detection
Guokang Zhu, Qi Wang 0009, Yuan Yuan 0001 |
Comput. Vis. Image Underst. | 3 |
| 2014 | Ego motion guided particle filter for vehicle tracking in airborne videos
Xianbin Cao 0001, Changcheng Gao, Jinhe Lan, Yuan Yuan 0001, Pingkun Yan |
Neurocomputing | 4 |
| 2014 | Hybrid structure for robust dimensionality reduction
Xiaoqiang Lu, Yuan Yuan 0001 |
Neurocomputing | 2 |
| 2014 | Multi-cue based tracking
Qi Wang 0009, Jianwu Fang, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2014 | High quality image resizing
Qi Wang 0009, Yuan Yuan 0001 |
Neurocomputing | 2 |
| 2014 | Learning to resize image
Qi Wang 0009, Yuan Yuan 0001 |
Neurocomputing | 2 |
| 2014 | Sparse frontal face image synthesis from an arbitrary profile image
Lin Zhao 0003, Xinbo Gao 0001, Yuan Yuan 0001, Dapeng Tao |
Neurocomputing | 3 |
| 2014 | Part-Based Online Tracking With Geometry Constraint and Attention SelectionabstractVisual tracking in condition of occlusion, appearance or illumination change has been a challenging task over decades. Recently, some online trackers, based on the detection by classification framework, have achieved good performance. However, problems are still embodied in at least one of the three aspects: 1) tracking the target with a single region has poor adaptability for occlusion, appearance or illumination change; 2) lack of sample weight estimation, which may cause overfitting issue; and 3) inadequate motion model to prevent target from drifting. For tackling the above problems, this paper presents the contributions as follows: 1) a novel part-based structure is utilized in the online AdaBoost tracking; 2) attentional sample weighting and selection is tackled by introducing a weight relaxation factor, instead of treating the samples equally as traditional trackers do; and 3) a two-stage motion model, multiple parts constraint, is proposed and incorporated into the part-based structure to ensure a stable tracking. The effectiveness and efficiency of the proposed tracker is validated upon several complex video sequences, compared with seven popular online trackers. The experimental results show that the proposed tracker can achieve increased accuracy with comparable computational cost. Jianwu Fang, Qi Wang 0009, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Robust Superpixel Tracking via Depth FusionabstractAlthough numerous trackers have been designed to adapt to the nonstationary image streams that change over time, it remains a challenging task to facilitate a tracker to accurately distinguish the target from the background in every frame. This paper proposes a robust superpixel-based tracker via depth fusion, which exploits the adequate structural information and great flexibility of mid-level features captured by superpixels, as well as the depth-map's discriminative ability for the target and background separation. By introducing graph-regularized sparse coding into the appearance model, the local geometrical structure of data is considered, and the resulting appearance model has a more powerful discriminative ability. Meanwhile, the similarity of the target superpixels' neighborhoods in two adjacent frames is also incorporated into the refinement of the target estimation, which helps a more accurate localization. Most importantly, the depth cue is fused into the superpixel-based target estimation so as to tackle the cluttered background with similar appearance to the target. To evaluate the effectiveness of the proposed tracker, four video sequences of different challenging situations are contributed by the authors. The comparison results demonstrate that the proposed tracker has more robust and accurate performance than seven ones representing the state-of-the-art. Yuan Yuan 0001, Jianwu Fang, Qi Wang 0009 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Alternatively Constrained Dictionary Learning For Image SuperresolutionabstractDictionaries are crucial in sparse coding-based algorithm for image superresolution. Sparse coding is a typical unsupervised learning method to study the relationship between the patches of high-and low-resolution images. However, most of the sparse coding methods for image superresolution fail to simultaneously consider the geometrical structure of the dictionary and the corresponding coefficients, which may result in noticeable superresolution reconstruction artifacts. In other words, when a low-resolution image and its corresponding high-resolution image are represented in their feature spaces, the two sets of dictionaries and the obtained coefficients have intrinsic links, which has not yet been well studied. Motivated by the development on nonlocal self-similarity and manifold learning, a novel sparse coding method is reported to preserve the geometrical structure of the dictionary and the sparse coefficients of the data. Moreover, the proposed method can preserve the incoherence of dictionary entries and provide the sparse coefficients and learned dictionary from a new perspective, which have both reconstruction and discrimination properties to enhance the learning performance. Furthermore, to utilize the model of the proposed method more effectively for single-image superresolution, this paper also proposes a novel dictionary-pair learning method, which is named as two-stage dictionary training. Extensive experiments are carried out on a large set of images comparing with other popular algorithms for the same purpose, and the results clearly demonstrate the effectiveness of the proposed sparse representation model and the corresponding dictionary learning algorithm. Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan |
IEEE Trans. Cybern. | 2 |
| 2014 | Distributed Object Detection With Linear SVMsabstractIn vision and learning, low computational complexity and high generalization are two important goals for video object detection. Low computational complexity here means not only fast speed but also less energy consumption. The sliding window object detection method with linear support vector machines (SVMs) is a general object detection framework. The computational cost is herein mainly paid in complex feature extraction and innerproduct-based classification. This paper first develops a distributed object detection framework (DOD) by making the best use of spatial-temporal correlation, where the process of feature extraction and classification is distributed in the current frame and several previous frames. In each framework, only subfeature vectors are extracted and the response of partial linear classifier (i.e., subdecision value) is computed. To reduce the dimension of traditional block-based histograms of oriented gradients (BHOG) feature vector, this paper proposes a cell-based HOG (CHOG) algorithm, where the features in one cell are not shared with overlapping blocks. Using CHOG as feature descriptor, we develop CHOG-DOD as an instance of DOD framework. Experimental results on detection of hand, face, and pedestrian in video show the superiority of the proposed method. Yanwei Pang, Yuan Yuan 0001, Kongqiao Wang |
IEEE Trans. Cybern. | 3 |
| 2014 | Learning From Errors in Super-ResolutionabstractA novel framework of learning-based super-resolution is proposed by employing the process of learning from the estimation errors. The estimation errors generated by different learning-based super-resolution algorithms are statistically shown to be sparse and uncertain. The sparsity of the estimation errors means most of estimation errors are small enough. The uncertainty of the estimation errors means the location of the pixel with larger estimation error is random. Noticing the prior information about the estimation errors, a nonlinear boosting process of learning from these estimation errors is introduced into the general framework of the learning-based super-resolution. Within the novel framework of super-resolution, a low-rank decomposition technique is used to share the information of different super-resolution estimations and to remove the sparse estimation errors from different learning algorithms or training samples. The experimental results show the effectiveness and the efficiency of the proposed framework in enhancing the performance of different learning-based algorithms. Yi Tang 0003, Yuan Yuan 0001 |
IEEE Trans. Cybern. | 2 |
| 2014 | NATAS: Neural Activity Trace Aware SaliencyabstractSaliency detection has raised much interest in computer vision recently. Many visual saliency models have been developed for individual images, video clips, and image pairs. However, image sequence, one most general occasion in the real world, is not explored yet. A general image sequence is different from video clips whose temporal continuity is maintained and image pairs where common objects exist. It might contain some similar low-level properties while completely distinct contents. Traditional saliency detection methods will fail on these general sequences. Based on this consideration, this paper investigates the shortcomings of the classical saliency detection methods, which significantly limit their advantages: 1) inability to capture the natural connections among sequential images, 2) over-reliance on motion cues, and 3) restriction to image pairs/videos with common objects. In order to address these problems, we propose a framework that performs the following contributions: 1) construct an image data set as benchmark through a rigorously designed behavioral experiment, 2) propose a neural activity trace aware saliency model to capture the general connections among images, and 3) design a novel measure to handle the low-level clues contained among sequential images. Experimental results demonstrate that the proposed saliency model is associated with a tremendous advancement compared with traditional methods when dealing with the general image sequence. Guokang Zhu, Qi Wang 0009, Yuan Yuan 0001 |
IEEE Trans. Cybern. | 3 |
| 2014 | Double Constrained NMF for Hyperspectral UnmixingabstractGiven only the collected hyperspectral data, unmixing aims at obtaining the latent constituent materials and their corresponding fractional abundances. Recently, manynonnegative matrix factorization(NMF)-based algorithms have been developed to deal with this issue. Considering that the abundances of most materials may be sparse, the sparseness constraint is intuitively introduced into NMF. Although sparse NMF algorithms have achieved advanced performance in unmixing, the result is still susceptible to unstable decomposition and noise corruption. To reduce the aforementioned drawbacks, the structural information of the data is exploited to guide the unmixing. Since similar pixel spectra often imply similar substance constructions, clustering can explicitly characterize this similarity. Through maintaining the structural information during the unmixing, the resulting fractional abundances by the proposed algorithm can well coincide with the real distributions of constituent materials. Moreover, the additional clustering-based regularization term also lessens the interference of noise to some extent. The experimental results on synthetic and real hyperspectral data both illustrate the superiority of the proposed method compared with other state-of-the-art algorithms. Xiaoqiang Lu, Hao Wu 0098, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2014 | Learning Regularized LDA by ClusteringabstractAs a supervised dimensionality reduction technique, linear discriminant analysis has a serious overfitting problem when the number of training samples per class is small. The main reason is that the between- and within-class scatter matrices computed from the limited number of training samples deviate greatly from the underlying ones. To overcome the problem without increasing the number of training samples, we propose making use of the structure of the given training data to regularize the between- and within-class scatter matrices by between- and within-cluster scatter matrices, respectively, and simultaneously. The within- and between-cluster matrices are computed from unsupervised clustered data. The within-cluster scatter matrix contributes to encoding the possible variations in intraclasses and the between-cluster scatter matrix is useful for separating extra classes. The contributions are inversely proportional to the number of training samples per class. The advantages of the proposed method become more remarkable as the number of training samples per class decreases. Experimental results on the AR and Feret face databases demonstrate the effectiveness of the proposed method. Yanwei Pang, Yuan Yuan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2013 | Multi-spectral dataset and its application in saliency detection
Qi Wang 0009, Guokang Zhu, Yuan Yuan 0001 |
Comput. Vis. Image Underst. | 3 |
| 2013 | Pedestrian detection in unseen scenes by dynamically updating visual words
Xianbin Cao 0001, Bo Ning 0003, Yuan Yuan 0001, Pingkun Yan |
Neurocomputing | 4 |
| 2013 | Adaptively post-encoding multiple description video coding
Xuguang Lan, Meng Yang 0002, Yuan Yuan 0001, Songlin Zhao, Nanning Zheng 0001 |
Neurocomputing | 3 |
| 2013 | Face recognition using Weber local descriptors
Shutao Li 0001, Dayi Gong, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2013 | Sparse coding for image denoising using spike and slab prior
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan |
Neurocomputing | 2 |
| 2013 | Robust probabilistic tensor analysis for time-variant collaborative filtering
Zhao Ma, Yanwei Pang, Yuan Yuan 0001 |
Neurocomputing | 4 |
| 2013 | Image registration by normalized mapping
Qi Wang 0009, Cuiming Zou, Yuan Yuan 0001, Hongbing Lu, Pingkun Yan |
Neurocomputing | 3 |
| 2013 | Special issue: Behaviours in video
Huiyu Zhou 0001, Yuan Yuan 0001, Yingzi Du, Pingkun Yan |
Neurocomputing | 2 |
| 2013 | SIFT on manifold: An intrinsic description
Guokang Zhu, Qi Wang 0009, Yuan Yuan 0001, Pingkun Yan |
Neurocomputing | 3 |
| 2013 | Greedy regression in sparse coding space for single-image super-resolution
Yi Tang 0003, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Robust visual tracking with discriminative sparse learning
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan |
Pattern Recognit. | 2 |
| 2013 | Multi-spectral saliency detection
Qi Wang 0009, Pingkun Yan, Yuan Yuan 0001, Xuelong Li 0001 |
Pattern Recognit. Lett. | 3 |
| 2013 | Energy-saving object detection by efficiently rejecting a set of neighboring sub-images
Yanwei Pang, Yuan Yuan 0001, Kongqiao Wang |
Signal Process. | 4 |
| 2013 | Image Super-Resolution Via Double Sparsity Regularized Manifold LearningabstractOver the past few years, high resolutions have been desirable or essential, e.g., in online video systems, and therefore, much has been done to achieve an image of higher resolution from the corresponding low-resolution ones. This procedure of recovering/rebuilding is called single-image super-resolution (SR). Performance of image SR has been significantly improved via methods of sparse coding. That is to say, the image frame patch can be sparse linear combinations of basis elements. However, most of these existing methods fail to consider the local geometrical structure in the space of the training data. To take this crucial issue into account, this paper proposes a method named double sparsity regularized manifold learning (DSRML). DSRML can preserve the properties of the aforementioned local geometrical structure by employing manifold learning, e.g., locally linear embedding. Based on a large amount of experimental results, DSRML is demonstrated to be more robust and more effective than previous efforts in the task of single-image SR. Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Person Re-Identification by Regularized Smoothing KISS Metric LearningabstractWith the rapid development of the intelligent video surveillance (IVS), person re-identification, which is a difficult yet unavoidable problem in video surveillance, has received increasing attention in recent years. That is because computer capacity has shown remarkable progress and the task of person re-identification plays a critical role in video surveillance systems. In short, person re-identification aims to find an individual again that has been observed over different cameras. It has been reported that KISS metric learning has obtained the state of the art performance for person re-identification on the VIPeR dataset. However, given a small size training set, the estimation to the inverse of a covariance matrix is not stable and thus the resulting performance can be poor. In this paper, we present regularized smoothing KISS metric learning (RS-KISS) by seamlessly integrating smoothing and regularization techniques for robustly estimating covariance matrices. RS-KISS is superior to KISS, because RS-KISS can enlarge the underestimated small eigenvalues and can reduce the overestimated large eigenvalues of the estimated covariance matrix in an effective way. By providing additional data, we can obtain a more robust model by RS-KISS. However, retraining RS-KISS on all the available examples in a straightforward way is time consuming, so we introduce incremental learning to RS-KISS. We thoroughly conduct experiments on the VIPeR dataset and verify that 1) RS-KISS completely beats all available results for person re-identification and 2) incremental RS-KISS performs as well as RS-KISS but reduces the computational cost significantly. Dapeng Tao, Yongfei Wang, Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | Visual Saliency by Selective ContrastabstractAutomatic detection of salient objects in visual media (e.g., videos and images) has been attracting much attention. The detected salient objects can be utilized for segmentation, recognition, and retrieval. However, the accuracy of saliency detection remains a challenge. The reason behind this challenge is mainly due to the lack of a well-defined model for interpreting saliency formulation. To tackle this problem, this letter proposes to detect salient objects based on selective contrast. Selective contrast intrinsically explores the most distinguishable component information in color, texture, and location. A large number of experiments are thereafter carried out upon a benchmark dataset, and the results are compared with those of 12 other popular state-of-the-art algorithms. In addition, the advantage of the proposed algorithm is also demonstrated in a retargeting application. Qi Wang 0009, Yuan Yuan 0001, Pingkun Yan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Error Analysis of Stochastic Gradient Descent RankingabstractRanking is always an important task in machine learning and information retrieval, e.g., collaborative filtering, recommender systems, drug discovery, etc. A kernel-based stochastic gradient descent algorithm with the least squares loss is proposed for ranking in this paper. The implementation of this algorithm is simple, and an expression of the solution is derived via a sampling operator and an integral operator. An explicit convergence rate for leaning a ranking function is given in terms of the suitable choices of the step size and the regularization parameter. The analysis technique used here is capacity independent and is novel in error analysis of ranking learning. Experimental results on real-world data have shown the effectiveness of the proposed algorithm in ranking tasks, which verifies the theoretical analysis in ranking error. Hong Chen 0004, Yi Tang 0003, Luoqing Li, Yuan Yuan 0001, Xuelong Li 0001, Yuan Yan Tang |
IEEE Trans. Cybern. | 4 |
| 2013 | Saliency Detection by Multiple-Instance LearningabstractSaliency detection has been a hot topic in recent years. Its popularity is mainly because of its theoretical meaning for explaining human attention and applicable aims in segmentation, recognition, etc. Nevertheless, traditional algorithms are mostly based on unsupervised techniques, which have limited learning ability. The obtained saliency map is also inconsistent with many properties of human behavior. In order to overcome the challenges of inability and inconsistency, this paper presents a framework based on multiple-instance learning. Low-, mid-, and high-level features are incorporated in the detection procedure, and the learning ability enables it robust to noise. Experiments on a data set containing 1000 images demonstrate the effectiveness of the proposed framework. Its applicability is shown in the context of a seam carving application. Qi Wang 0009, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2013 | Learning Saliency by MRF and Differential ThresholdabstractSaliency detection has been an attractive topic in recent years. The reliable detection of saliency can help a lot of useful processing without prior knowledge about the scene, such as content-aware image compression, segmentation, etc. Although many efforts have been spent in this subject, the feature expression and model construction are far from perfect. The obtained saliency maps are therefore not satisfying enough. In order to overcome these challenges, this paper presents a new psychologic visual feature based on differential threshold and applies it in a supervised Markov-random-field framework. Experiments on two public data sets and an image retargeting application demonstrate the effectiveness, robustness, and practicability of the proposed method. Guokang Zhu, Qi Wang 0009, Yuan Yuan 0001, Pingkun Yan |
IEEE Trans. Cybern. | 3 |
| 2013 | Graph-Regularized Low-Rank Representation for Destriping of Hyperspectral ImagesabstractHyperspectral image destriping is a challenging and promising theme in remote sensing. Striping noise is a ubiquitous phenomenon in hyperspectral imagery, which may severely degrade the visual quality. A variety of methods have been proposed to effectively alleviate the effects of the striping noise. However, most of them fail to take full advantage of the high spectral correlation between the observation subimages in distinct bands and consider the local manifold structure of the hyperspectral data space. In order to remedy this drawback, in this paper, a novel graph-regularized low-rank representation (LRR) destriping algorithm is proposed by incorporating the LRR technique. To obtain desired destriping performance, two sides of performing destriping are included: 1) To exploit the high spectral correlation between the observation subimages in distinct bands, the technique of LRR is first utilized for destriping, and 2) to preserve the intrinsic local structure of the original hyperspectral data, the graph regularizer is incorporated in the objective function. The experimental results and quantitative analysis demonstrate that the proposed method can both remove striping noise and achieve cleaner and higher contrast reconstructed results. Xiaoqiang Lu, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2013 | Manifold Regularized Sparse NMF for Hyperspectral UnmixingabstractHyperspectral unmixing is one of the most important techniques in analyzing hyperspectral images, which decomposes a mixed pixel into a collection of constituent materials weighted by their proportions. Recently, many sparse nonnegative matrix factorization (NMF) algorithms have achieved advanced performance for hyperspectral unmixing because they overcome the difficulty of absence of pure pixels and sufficiently utilize the sparse characteristic of the data. However, most existing sparse NMF algorithms for hyperspectral unmixing only consider the Euclidean structure of the hyperspectral data space. In fact, hyperspectral data are more likely to lie on a low-dimensional submanifold embedded in the high-dimensional ambient space. Thus, it is necessary to consider the intrinsic manifold structure for hyperspectral unmixing. In order to exploit the latent manifold structure of the data during the decomposition, manifold regularization is incorporated into sparsity-constrained NMF for unmixing in this paper. Since the additional manifold regularization term can keep the close link between the original image and the material abundance maps, the proposed approach leads to a more desired unmixing performance. The experimental results on synthetic and real hyperspectral data both illustrate the superiority of the proposed method compared with other state-of-the-art approaches. Xiaoqiang Lu, Hao Wu 0098, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2013 | Sparse Coding From a Bayesian PerspectiveabstractSparse coding is a promising theme in computer vision. Most of the existing sparse coding methods are based on either l0 or l1 penalty, which often leads to unstable solution or biased estimation. This is because of the nonconvexity and discontinuity of the l0 penalty and the over-penalization on the true large coefficients of the l1 penalty. In this paper, sparse coding is interpreted from a novel Bayesian perspective, which results in a new objective function through maximum a posteriori estimation. The obtained solution of the objective function can generate more stable results than the l0 penalty and smaller reconstruction errors than the l1 penalty. In addition, the convergence property of the proposed algorithm for sparse coding is also established. The experiments on applications in single image super-resolution and visual tracking demonstrate that the proposed method is more effective than other state-of-the-art methods. Xiaoqiang Lu, Yuan Yuan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2012 | Geometry constrained sparse coding for single image super-resolutionabstractThe choice of the over-complete dictionary that sparsely represents data is of prime importance for sparse coding-based image super-resolution. Sparse coding is a typical unsupervised learning method to generate an over-complete dictionary. However, most of the sparse coding methods for image super-resolution fail to simultaneously consider the geometrical structure of the dictionary and corresponding coefficients, which may result in noticeable super-resolution reconstruction artifacts. In this paper, a novel sparse coding method is proposed to preserve the geometrical structure of the dictionary and the sparse coefficients of the data. Moreover, the proposed method can preserve the incoherence of dictionary entries, which is critical for sparse representation. Inspired by the development on non-local self-similarity and manifold learning, the proposed sparse coding method can provide the sparse coefficients and learned dictionary from a new perspective, which have both reconstruction and discrimination properties to enhance the learning performance. Extensive experimental results on image super-resolution have demonstrated the effectiveness of the proposed method. Xiaoqiang Lu, Pingkun Yan, Yuan Yuan 0001, Xuelong Li 0001 |
CVPR | 4 |
| 2012 | Robust lossless data hiding using clustering and statistical quantity histogram
Lingling An, Xinbo Gao 0001, Yuan Yuan 0001, Dacheng Tao |
Neurocomputing | 3 |
| 2012 | Content-adaptive reliable robust lossless data embedding
Lingling An, Xinbo Gao 0001, Yuan Yuan 0001, Dacheng Tao, Cheng Deng 0002 |
Neurocomputing | 3 |
| 2012 | A method using long digital straight segments for fingerprint recognition
Xiubao Jiang, Xinge You, Yuan Yuan 0001, Mingming Gong |
Neurocomputing | 3 |
| 2012 | Similarity learning for object recognition based on derived kernel
Hong Li 0009, Yantao Wei, Luoqing Li, Yuan Yuan 0001 |
Neurocomputing | 4 |
| 2012 | Incremental threshold learning for classifier selection
Yanwei Pang, Junping Deng, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2012 | Fully affine invariant SURF for image matching
Yanwei Pang, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2012 | Scale invariant image matching using triplewise constraint and weighted voting
Yanwei Pang, Mianyou Shang, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2012 | Learning optimal spatial filters by discriminant analysis for brain-computer-interface
Yanwei Pang, Yuan Yuan 0001, Kongqiao Wang |
Neurocomputing | 2 |
| 2012 | Video super-resolution with 3D adaptive normalized convolution
Kaibing Zhang, Guangwu Mu, Yuan Yuan 0001, Xinbo Gao 0001, Dacheng Tao |
Neurocomputing | 3 |
| 2012 | Efficient image matching using weighted voting
Yuan Yuan 0001, Yanwei Pang, Kongqiao Wang, Mianyou Shang |
Pattern Recognit. Lett. | 1 |
| 2012 | Robust Alternative Minimization for Matrix CompletionabstractRecently, much attention has been drawn to the problem of matrix completion, which arises in a number of fields, including computer vision, pattern recognition, sensor network, and recommendation systems. This paper proposes a novel algorithm, named robust alternative minimization (RAM), which is based on the constraint of low rank to complete an unknown matrix. The proposed RAM algorithm can effectively reduce the relative reconstruction error of the recovered matrix. It is numerically easier to minimize the objective function and more stable for large-scale matrix completion compared with other existing methods. It is robust and efficient for low-rank matrix completion, and the convergence of the RAM algorithm is also established. Numerical results showed that both the recovery accuracy and running time of the RAM algorithm are competitive with other reported methods. Moreover, the applications of the RAM algorithm to low-rank image recovery demonstrated that it achieves satisfactory performance. Xiaoqiang Lu, Tieliang Gong, Pingkun Yan, Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2012 | Robust CoHOG Feature Extraction in Human-Centered Image/Video Management SystemabstractMany human-centered image and video management systems depend on robust human detection. To extract robust features for human detection, this paper investigates the following shortcomings of co-occurrence histograms of oriented gradients (CoHOGs) which significantly limit its advantages: 1) The magnitudes of the gradients are discarded, and only the orientations are used; 2) the gradients are not smoothed, and thus, aliasing effect exists; and 3) the dimensionality of the CoHOG feature vector is very large (e.g., 200,000). To deal with these problems, in this paper, we propose a framework that performs the following: 1) utilizes a novel gradient decomposition and combination strategy to make full use of the information of gradients; (2) adopts a two-stage gradient smoothing scheme to perform efficient gradient interpolation; and (3) employs incremental principal component analysis to reduce the large dimensionality of the CoHOG features. Experimental results on the two different human databases demonstrate the effectiveness of the proposed method. Yanwei Pang, Yuan Yuan 0001, Kongqiao Wang |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2011 | Medical Image Segmentation Using Descriptive Image FeaturesabstractSegmentation of medical images is an important component for diagnosis and treatment of diseases using medical imaging technologies. However, automated accurate medical image segmentation is still a challenge due to the difficulties in finding a robust feature descriptor to describe the object boundaries in medical images. In this paper, a new normal vector feature profile (NVFP) is proposed to describe the local image information of a contour point by concatenating a series of local region descriptors along the normal direction at that point. To avoid trapping by false boundaries caused by nonboundary image features, a modified scale invariant feature transform (SIFT) descriptor is developed. The number and locations of sample points for building NVFP are determined for each contour point, which are constrained by the neighboring anatomical structures and the statistical consistency of the training features. NVFP is incorporated into a model based method for image segmentation. The performance of our proposed method was demonstrated by segmenting prostate MR images. The segmentation results indicated that our method can achieve better performance compared with other existing methods. Meijuan Yang, Yuan Yuan 0001, Xuelong Li 0001, Pingkun Yan |
BMVC | 2 |
| 2011 | Robust Sparse Tensor Decomposition by Probabilistic Latent Semantic AnalysisabstractMovie recommendation system is becoming more and more popular in recent years. As a result, it is becoming increasingly important to develop machine learning algorithm on partially-observed matrix to predict users' preferences on missing data. Motivated by the user ratings prediction problem, we propose a novel robust tensor probabilistic latent semantic analysis (RT-pLSA) algorithm that not only takes time variable into account, but also uses the periodic property of data in time attribute. Different from the previous algorithms of predicting missing values on two-dimensional sparse matrix, we formulize the prediction problem as a probabilistic tensor factorization problem with periodicity constraint on time coordinate. Furthermore, we apply the Tsallis divergence error measure in the context of RT-pLSA tensor decomposition that is able to robustly predict the latent variable in the presence of noise. Our experimental results on two benchmark movie rating dataset: Netflix and Movie lens, show a good predictive accuracy of the model. Yanwei Pang, Zhao Ma, Yuan Yuan 0001 |
ICIG | 4 |
| 2011 | Integrating kAS and SIFT-like Descriptor for Image DescriptionabstractShape-descriptor (e.g. Adjacent Contour Segments, i.e. kAS) and key point-descriptor (e.g. Scale Invariant Feature Transform, i.e. SIFT) are widely used for computer vision. However, few works principally integrate shape-descriptor and key point-descriptor to describe the content of an image. On one hand, in some cases the degree of locality of keying-descriptor is too high to capture semantic characteristics of an object. On the other hand, though the shape has higher semantic level than key point, it contains no texture information because only the information of contour/edge is used. To make full use of the information of both shape and key point for generate robust and distinctive features, in this paper we propose an algorithm to integrate shape and key point descriptor. Specifically, we employ kAS to extract useful shape information. Then key points of a kAS shape are defined at which we propose to extract SIFT-like features. Experimental results on image matching demonstrate the effectiveness of the proposed algorithm. Mianyou Shang, Yanwei Pang, Yuan Yuan 0001 |
ICIG | 4 |
| 2011 | Single-Image Super-Resolution via Sparse Coding RegressionabstractIn this paper, it has been shown that the sparse coding algorithm for single-image super-resolution is equivalent to a linear regression algorithm in the sparse coding space. Following the idea, the sparse coding algorithm are generalized by a novel L2-Boosting-based single-resolution super-resolution algorithm which focuses on the relationship between sparse codings corresponding to the low- and high-resolution image patches. The experimental results demonstrate the effectiveness of the proposed algorithm by comparing with other state-of-the-art algorithms. Yi Tang 0003, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001 |
ICIG | 2 |
| 2011 | Putting images on a manifold for atlas-based image segmentationabstractIn medical image analysis, atlas-based segmentation has become a popular approach. Given a target image, how to select the atlases with the similar shape of anatomical structure to the input image is one of the most critical factors affecting the segmentation accuracy. In this paper, we propose a novel strategy by putting the images on a manifold to analyze the intrinsic similarity between the images. A subset of atlases can be selected and the optimal fusion weights are computed in a low-dimensional manifold space. Finally, it combines the selected atlases by using the corresponding weights for image segmentation. The experimental results demonstrated that our proposed method is robust and accurate especially when a large number of training samples are available. Yihui Cao, Yuan Yuan 0001, Xuelong Li 0001, Pingkun Yan |
ICIP | 2 |
| 2011 | A novel alternative algorithm for limited angle tomographyabstractThis paper studies incomplete data problems of circular cone-beam computed tomography, which occur frequently in medical imaging and industrial imaging. The incomplete data problems in which projection data are only available in an angular range can be attributed to the limited angle tomography. Limited angle tomography is a severely ill-posed inverse problem. In recent years, image reconstruction based on total variation (TV) was employed to reduce the problem and gave better performance on edge-preserving reconstruction. However, the artificial parameter can only be determined through considerable experimentation. In this paper, an alternating minimization method based on TV is proposed to reduce the data insufficiency in tomographic imaging. This novel alternating minimization method provides a robust and effective reconstruction without any artificial parameter in the iterative processes, by using the TV as a multiplicative constraint. The results demonstrate that this new reconstruction method brings satisfactory performance. Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001 |
ICIP | 2 |
| 2011 | Multimodal learning for multi-label image classificationabstractWe tackle the challenge of web image classification using additional tags information. Unlike traditional methods that only use the combination of several low-level features, we try to use semantic concepts to represent images and corresponding tags. At first, we extract the latent topic information by probabilistic latent semantic analysis (pLSA) algorithm, and then use multi-label multiple kernel learning to combine visual and textual features to make a better image classification. In our experiments on PASCAL VOC'07 set and MIR Flickr set, we demonstrate the benefit of using multimodal feature to improve image classification. Specifically, we discover that on the issue of image classification, utilizing latent semantic feature to represent images and associated tags can obtain better classification results than other ways that integrating several low-level features. Yanwei Pang, Zhao Ma, Yuan Yuan 0001, Xuelong Li 0001, Kongqiao Wang |
ICIP | 3 |
| 2011 | Robust color correction in stereo visionabstractThe phenomenon of color discrepancy between image pairs happens frequently in stereo vision systems. This inconsistence in color domain may cause difficulties when identifying point correspondence to reconstruct the scene depth. In this paper, we propose a robust algorithm to correct the color discrepancy between images. The proposed algorithm neither requires a color calibration chart/object which is a tedious procedure, nor explicitly compensates for the image as a whole, which possibly give bad correction results in local areas of an image. Instead, we correct the image region by region. Experiments show that the presented color correction algorithm is effective and efficient. Qi Wang 0009, Pingkun Yan, Yuan Yuan 0001, Xuelong Li 0001 |
ICIP | 3 |
| 2011 | Learning shape statistics for hierarchical 3D medical image segmentationabstractAccurate image segmentation is important for many medical imaging applications, whereas it remains challenging due to the complexity in medical images, such as the complex shapes and varied neighbor structures. This paper proposes a new hierarchical 3D image segmentation method based on patient-specific shape prior and surface patch shape statistics (SURPASS) model. In the segmentation process, a coarse-to-fine, two-stage strategy is designed, which contains global segmentation and local segmentation. In the global segmentation stage, patient-specific shape prior is estimated by using manifold learning techniques to achieve the overall segmentation. In the second stage, SURPASS is computed to solve the problem of poor segmentation at certain surface patches. The effectiveness of the proposed 3D image segmentation method has been demonstrated by the experiments on segmenting the prostate from a series of MR images. Wuxia Zhang, Yuan Yuan 0001, Xuelong Li 0001, Pingkun Yan |
ICIP | 2 |
| 2011 | Segmenting Images by Combining Selected Atlases on Manifold
Yihui Cao, Yuan Yuan 0001, Xuelong Li 0001, Baris Turkbey, Peter L. Choyke, Pingkun Yan |
MICCAI (3) | 2 |
| 2011 | Local learning-based image super-resolutionabstractLocal learning algorithm has been widely used in single-frame super-resolution reconstruction algorithm, such as neighbor embedding algorithm [1] and locality preserving constraints algorithm [2]. Neighbor embedding algorithm is based on manifold assumption, which defines that the embedded neighbor patches are contained in a single manifold. While manifold assumption does not always hold. In this paper, we present a novel local learning-based image single-frame SR reconstruction algorithm with kernel ridge regression (KRR). Firstly, Gabor filter is adopted to extract texture information from low-resolution patches as the feature. Secondly, each input low-resolution feature patch utilizes K nearest neighbor algorithm to generate a local structure. Finally, KRR is employed to learn a map from input low-resolution (LR) feature patches to high-resolution (HR) feature patches in the corresponding local structure. Experimental results show the effectiveness of our method. Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan, Luoqing Li, Xuelong Li 0001 |
MMSP | 3 |
| 2011 | Local semi-supervised regression for single-image super-resolutionabstractIn this paper, we propose a local semi-supervised learning-based algorithm for single-image super-resolution. Different from most of example-based algorithms, the information of test patches is considered during learning local regression functions which map a low-resolution patch to a high-resolution patch. Localization strategy is generally adopted in single-image super-resolution with nearest neighbor-based algorithms. However, the poor generalization of the nearest neighbor estimation decreases the performance of such algorithms. Though the problem can be fixed by local regression algorithms, the sizes of local training sets are always too small to improve the performance of nearest neighbor-based algorithms significantly. To overcome the difficulty, the semi-supervised regression algorithm is used here. Unlike supervised regression, the information about test samples is considered in semi-supervised regression algorithms, which makes the semi-supervised regression more powerful. Noticing that numerous test patches exist, the performance of nearest neighbor-based algorithms can be further improved by employing a semi-supervised regression algorithm. Experiments verify the effectiveness of the proposed algorithm. Yi Tang 0003, Xiaoli Pan, Yuan Yuan 0001, Pingkun Yan, Luoqing Li, Xuelong Li 0001 |
MMSP | 3 |
| 2011 | Summarizing tourist destinations by mining user-generated travelogues and photos
Yanwei Pang, Yuan Yuan 0001, Tanji Hu, Rui Cai 0002, Lei Zhang 0001 |
Comput. Vis. Image Underst. | 3 |
| 2011 | Arbitrary ROI-based wavelet video coding
Xuguang Lan, Nanning Zheng 0001, Yuan Yuan 0001 |
Neurocomputing | 4 |
| 2011 | Image reconstruction by an alternating minimisation
Xiaoqiang Lu, Yi Sun 0009, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2011 | Travelogue Enriching and Scenic Spot Overview Based on Textual and Visual Topic ModelsabstractWe consider the problem of enriching the travelogue associated with a small number (even one) of images with more web images. Images associated with the travelogue always consist of the content and the style of textual information. Relying on this assumption, in this paper, we present a framework of travelogue enriching, exploiting both textual and visual information generated by different users. The framework aims to select the most relevant images from automatically collected candidate image set to enrich the given travelogue, and form a comprehensive overview of the scenic spot. To do these, we propose to build two-layer probabilistic models, i.e. a text-layer model and image-layer models, on offline collected travelogues and images. Each topic (e.g. Sea, Mountain, Historical Sites) in the text-layer model is followed by an image-layer model with sub-topics learnt (e.g. the topic of sea is with the sub-topic like beach, tree, sunrise and sunset). Based on the model, we develop strategies to enrich travelogues in the following steps: (1) remove noisy names of scenic spots from travelogues; (2) generate queries to automatically gather candidate image set; (3) select images to enrich the travelogue; and (4) choose images to portray the visual content of a scenic spot. Experimental results on Chinese travelogues demonstrate the potential of the proposed approach on tasks of travelogue enrichment and the corresponding scenic spot illustration. Yanwei Pang, Yuan Yuan 0001, Xuelong Li 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2011 | Optimization for limited angle tomography in medical image processing
Xiaoqiang Lu, Yi Sun 0009, Yuan Yuan 0001 |
Pattern Recognit. | 3 |
| 2011 | Segmentation of retinal blood vessels using the radial projection and semi-supervised approach
Xinge You, Qinmu Peng, Yuan Yuan 0001, Yiu-Ming Cheung, Jiajia Lei |
Pattern Recognit. | 3 |
| 2011 | Efficient HOG human detection
Yanwei Pang, Yuan Yuan 0001, Xuelong Li 0001 |
Signal Process. | 2 |
| 2011 | Lossless Data Embedding Using Generalized Statistical Quantity HistogramabstractHistogram-based lossless data embedding (LDE) has been recognized as an effective and efficient way for copyright protection of multimedia. Recently, a LDE method using the statistical quantity histogram has achieved good performance, which utilizes the similarity of the arithmetic average of difference histogram (AADH) to reduce the diversity of images and ensure the stable performance of LDE. However, this method is strongly dependent on some assumptions, which limits its applications in practice. In addition, the capacities of the images with the flat AADH, e.g., texture images, are a little bit low. For this purpose, we develop a novel framework for LDE by incorporating the merits from the generalized statistical quantity histogram (GSQH) and the histogram-based embedding. Algorithmically, we design the GSQH driven LDE framework carefully so that it: (1) utilizes the similarity and sparsity of GSQH to construct an efficient embedding carrier, leading to a general and stable framework; (2) is widely adaptable for different kinds of images, due to the usage of the divide-and-conquer strategy; (3) is scalable for different capacity requirements and avoids the capacity problems caused by the flat histogram distribution; (4) is conditionally robust against JPEG compression under a suitable scale factor; and (5) is secure for copyright protection because of the safe storage and transmission of side information. Thorough experiments over three kinds of images demonstrate the effectiveness of the proposed framework. Xinbo Gao 0001, Lingling An, Yuan Yuan 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Identifying camera and processing from cropped JPEG photos via tensor analysisabstractDigital image and video forensics is to detect the authenticity of digital images/videos. So far, existing algorithms are designed by exploring one or several special features, which usually result in limited performance and applicability. With the perspective of entire procedure including image acquisition and image processing, a theory of general blind image forensics is proposed in this paper. Then a new blind image forensic method is designed based on this theory to identify camera source and processing history of cropped compressed digital images. In this method, tensor decomposition analysis is applied to extract features of nonlinear operations, which come from both algorithms embedded within camera and operations done by post-software. Then, the Support Vector Machine is utilized to classify whether the target image is captured by the claimed camera and underwent the declared processing history. Experimental results show that this method has high detection accuracy, which demonstrated that our theory is correct. Weihai Li, Nenghai Yu, Yuan Yuan 0001 |
SMC | 3 |
| 2010 | Rotation invariant iris feature extraction using Gaussian Markov random fields with non-separable wavelet
Jing Huang 0018, Xinge You, Yuan Yuan 0001 |
Neurocomputing | 3 |
| 2010 | No-reference image quality assessment in contourlet domain
Wen Lu 0004, Dacheng Tao, Yuan Yuan 0001, Xinbo Gao 0001 |
Neurocomputing | 4 |
| 2010 | Outlier-resisting graph embedding
Yanwei Pang, Yuan Yuan 0001 |
Neurocomputing | 2 |
| 2010 | Semi-supervised Gaussian process latent variable model with pairwise constraints
Xiumei Wang 0002, Xinbo Gao 0001, Yuan Yuan 0001, Dacheng Tao, Jie Li 0001 |
Neurocomputing | 3 |
| 2010 | Hybrid sampling on mutual information entropy-based clustering ensembles for optimizations
Feng Wang 0048, Zhiyi Lin 0001, Yuanxiang Li 0001, Yuan Yuan 0001 |
Neurocomputing | 5 |
| 2010 | Incremental tensor biased discriminant analysis: A new color-based visual tracking method
Xinbo Gao 0001, Yuan Yuan 0001, Dacheng Tao, Jie Li 0001 |
Neurocomputing | 3 |
| 2010 | Photo-sketch synthesis and recognition based on subspace learning
Bing Xiao 0003, Xinbo Gao 0001, Dacheng Tao, Yuan Yuan 0001, Jie Li 0001 |
Neurocomputing | 4 |
| 2010 | Robust Tensor Analysis With L1-NormabstractTensor analysis plays an important role in modern image and vision computing problems. Most of the existing tensor analysis approaches are based on the Frobenius norm, which makes them sensitive to outliers. In this paper, we propose L1-norm-based tensor analysis (TPCA-L1), which is robust to outliers. Experimental results upon face and other datasets demonstrate the advantages of the proposed approach. Yanwei Pang, Xuelong Li 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Footwear for Gender RecognitionabstractConventionally, biometrics resources, such as face, gait silhouette, footprint, and pressure, have been utilized in gender recognition systems. However, the acquisition and processing time of these biometrics data makes the analysis difficult. This letter demonstrates for the first time how effective the footwear appearance is for gender recognition as a biometrics resource. A footwear database is also established with reprehensive shoes (footwears). Preliminary experimental results suggest that footwear appearance is a promising resource for gender recognition. Moreover, it also has the potential to be used jointly with other developed biometrics resources to boost performance. Yuan Yuan 0001, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | L1-Norm-Based 2DPCAabstractIn this paper, we first present a simple but effective L1-norm-based two-dimensional principal component analysis (2DPCA). Traditional L2-norm-based least squares criterion is sensitive to outliers, while the newly proposed L1-norm 2DPCA is robust. Experimental results demonstrate its advantages. Xuelong Li 0001, Yanwei Pang, Yuan Yuan 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2009 | A fast feature extraction methodabstractA fast subspace analysis and feature extraction algorithm is proposed which is based on fast Haar transform and integral vector. In rapid object detection and conventional binary subspace learning, Haar-like functions have been frequently used but true Haar functions are seldom employed. In this paper we have shown that true Haar functions can be successfully used to accelerate subspace analysis and feature extraction. Both the training and testing speed of the proposed method is higher than conventional algorithms. Experimental results on face database demonstrated its effectiveness. Yanwei Pang, Xuelong Li 0001, Yuan Yuan 0001, Dacheng Tao |
ICASSP | 4 |
| 2009 | Object trajectory clustering via tensor analysisabstractIn this paper we present a new video object trajectory clustering algorithm1, which allows us to model and analyse the patterns of object behaviors based on the extracted features using tensor analysis. The proposed algorithm consists of three steps as follows: extraction of trajectory features by tensor analysis, non-parametric probabilistic mean shift clustering and clustering correction. The performance of the proposed algorithm is evaluated on standard data-sets and compared with classical techniques. Huiyu Zhou 0001, Dacheng Tao, Yuan Yuan 0001, Xuelong Li 0001 |
ICIP | 3 |
| 2009 | Improving Security of an Image Encryption Algorithm based on Chaotic Circular ShiftabstractAn image encryption algorithm based on chaotic circular bit shift is proposed recently. This paper analyses the security of this algorithm and point out that the key space is not as large as they alleged and the algorithm can not resist chosen-plaintext attack or difference attack. The amount of chosen-plaintexts to carry out an attack is very few. This paper also introduces two methods to improve its security by changing chaotic sequences generators, altering orders of permutation and substitution, and applying feedback link mode. The improved algorithm has variable key space and has very good avalanche effect to resist chosen-plaintext attacks, chosen-ciphertext attacks, or difference attacks. The computation cost of improved algorithm is very low. Weihai Li, Yuan Yuan 0001 |
SMC | 2 |
| 2009 | Object tracking using SIFT features and mean shift
Huiyu Zhou 0001, Yuan Yuan 0001, Chunmei Shi |
Comput. Vis. Image Underst. | 2 |
| 2009 | Scene segmentation based on IPCA for visual surveillance
Yuan Yuan 0001, Yanwei Pang, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2009 | Multiscale facial structure representation for face recognition under varying illumination
Taiping Zhang, Bin Fang 0001, Yuan Yuan 0001, Yuan Yan Tang, Zhaowei Shang, Fangnian Lang |
Pattern Recognit. | 3 |
| 2009 | Non-rigid object tracking in complex scenes
Huiyu Zhou 0001, Yuan Yuan 0001, Chunmei Shi |
Pattern Recognit. Lett. | 2 |
| 2009 | Texture image retrieval based on non-tensor product wavelet filter banks
Zhenyu He 0007, Xinge You, Yuan Yuan 0001 |
Signal Process. | 3 |
| 2009 | A novel iris segmentation using radial-suppression edge detection
Jing Huang 0018, Xinge You, Yuan Yan Tang, Yuan Yuan 0001 |
Signal Process. | 5 |
| 2009 | Passive detection of doctored JPEG image via block artifact grid extraction
Weihai Li, Yuan Yuan 0001, Nenghai Yu |
Signal Process. | 2 |
| 2009 | Curvelet based face recognition via dimension reduction
Tanaya Mandal, Q. M. Jonathan Wu, Yuan Yuan 0001 |
Signal Process. | 3 |
| 2009 | Visual information analysis for security
Dacheng Tao, Yuan Yuan 0001, Jialie Shen 0001, Kaiqi Huang, Xuelong Li 0001 |
Signal Process. | 2 |
| 2009 | Binary Sparse Nonnegative Matrix FactorizationabstractThis paper presents a fast part-based subspace selection algorithm, termed the binary sparse nonnegative matrix factorization (B-SNMF). Both the training process and the testing process of B-SNMF are much faster than those of binary principal component analysis (B-PCA). Besides, B-SNMF is more robust to occlusions in images. Experimental results on face images demonstrate the effectiveness and the efficiency of the proposed B-SNMF. Yuan Yuan 0001, Xuelong Li 0001, Yanwei Pang, Dacheng Tao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Fast Haar transform based feature extraction for face representation and recognitionabstractSubspace learning is the process of finding a proper feature subspace and then projecting high-dimensional data onto the learned low-dimensional subspace. The projection operation requires many floating-point multiplications and additions, which makes the projection process computationally expensive. To tackle this problem, this paper proposes twosimple-but-effectivefast subspace learning and image projection methods, fast Haar transform (FHT) based principal component analysis and FHT based spectral regression discriminant analysis. The advantages of these two methods result from employing both the FHT for subspace learning and the integral vector for feature extraction. Experimental results on three face databases demonstrated their effectiveness and efficiency. Yanwei Pang, Xuelong Li 0001, Yuan Yuan 0001, Dacheng Tao |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2009 | Iterative Subspace Analysis Based on Feature Line DistanceabstractNearest feature line-based subspace analysis is first proposed in this paper. Compared with conventional methods, the newly proposed one brings better generalization performance and incremental analysis. The projection point and feature line distance are expressed as a function of a subspace, which is obtained by minimizing the mean square feature line distance. Moreover, by adopting stochastic approximation rule to minimize the objective function in a gradient manner, the new method can be performed in an incremental mode, which makes it working well upon future data. Experimental results on the FERET face database and the UCI satellite image database demonstrate the effectiveness. Yanwei Pang, Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2009 | View-Independent Behavior AnalysisabstractThe motion analysis of the human body is an important topic of research in computer vision devoted to detecting, tracking, and understanding people's physical behavior. This strong interest is driven by a wide spectrum of applications in various areas such as smart video surveillance. Most research in behavior (or gesture) representation focusses on view-dependent representation, and some research on view invariance considers only information from 3-D models, which is effective under considerable changes of viewpoint. This paper introduces a view-independent behavior-analysis framework based on decision fusion in which distance and view angle factors are analyzed. This is a first effort to tackle the problem of behaviors under significant changes in view angle, and a first corresponding video database is built. Kaiqi Huang, Dacheng Tao, Yuan Yuan 0001, Xuelong Li 0001, Tieniu Tan |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2008 | Discriminant adaptive edge weights for graph embeddingabstractMany linear dimensionality reduction (LDR) methods, such as PCA and LDA, can be reformulated in the framework of graph embedding (GE). In this framework, those LDR methods are differentiated by values of edge weights of a graph. This paper first proposes a linear dimensionality reduction method, which assigns edges with discriminant adaptive weights. Specifically, we compute a local decision hyper-plane by using support vector machine (SVM). Then edge weighs corresponding to the local region are expressed as a function of the angle between the direction of the edges and the normal vector of the hyper-plane. Experimental results demonstrate the advantages of this proposed method. Yuan Yuan 0001, Yanwei Pang |
ICASSP | 1 |
| 2008 | Doctored JPEG image detectionabstractNowadays, digital images can be easily modified by using software. In this paper, a new blind approach is proposed to detect copy-paste trail in a doctored JPEG image, i.e., to check whether a copied area came from the same image or not. When a copy-paste procedure is done on an image, especially adding or hiding an object, the block artifact grid contained in the copy-pasted slice is moved together. Since a slice must be placed properly in the target image to avoid obvious vision flaw, the grid in the slice mismatches to the original grid in the target image normally. Our approach utilizes the mismatch information of block artifact grid as a clue of copy-paste forgery. Experiment results demonstrate the efficiency of the proposed approach. Weihai Li, Nenghai Yu, Yuan Yuan 0001 |
ICME | 3 |
| 2008 | Boosting simple projections for multi-class dimensionality reductionabstractThis paper presents a novel method for dimensionality reduction and for multi-class classification tasks. This method iteratively selects a series of simple but effective 1D subspaces, and then combines the corresponding 1D projections by Adaboost.M2. Its major advantages are: 1) it does not impose specific assumptions on data distribution; 2) it minimizes Bayes error estimation in low-dimensional space; 3) it simplifies existing subspace-based methods to eigenvalue decomposition problem; and 4) each of the 1D subspaces (with associated nearest neighbor classifier) has different emphasis - measured by weighted training error. Experiments on both synthetic and real-world data demonstrate the effectiveness of the proposed method. Yuan Yuan 0001, Yanwei Pang |
SMC | 1 |
| 2008 | Visual music and musical vision
Xuelong Li 0001, Dacheng Tao, Stephen J. Maybank, Yuan Yuan 0001 |
Neurocomputing | 4 |
| 2008 | Spatial relationship representation for visual object searching
Lijuan Duan, Laiyun Qing, Wen Gao 0001, Xilin Chen 0001, Yuan Yuan 0001 |
Neurocomputing | 6 |
| 2008 | Level set image segmentation with Bayesian analysis
Huiyu Zhou 0001, Yuan Yuan 0001, Faquan Lin, Tangwei Liu |
Neurocomputing | 2 |
| 2008 | Application of semantic features in face recognition
Huiyu Zhou 0001, Yuan Yuan 0001, Abdul Hamid Sadka |
Pattern Recognit. | 2 |
| 2008 | Gabor-Based Region Covariance Matrices for Face RecognitionabstractThis paper presents a new method for human face recognition by utilizing Gabor-based region covariance matrices as face descriptors. Both pixel locations and Gabor coefficients are employed to form the covariance matrices. Experimental results demonstrate the advantages of this proposed method. Yanwei Pang, Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Binary Two-Dimensional PCAabstractFast training and testing procedures are crucial in biometrics recognition research. Conventional algorithms, e.g., principal component analysis (PCA), fail to efficiently work on large-scale and high-resolution image data sets. By incorporating merits from both two-dimensional PCA (2DPCA)-based image decomposition and fast numerical calculations based on Haarlike bases, this technical correspondence first proposes binary 2DPCA (B-2DPCA). Empirical studies demonstrated the advantages of B-2DPCA compared with 2DPCA and binary PCA. Yanwei Pang, Dacheng Tao, Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2008 | Effective Feature Extraction in High-Dimensional SpaceabstractThis correspondence first kernalizes the region covariance matrix and formalizes the similarity metric using four block matrices. The effectiveness of the proposed methods is proven with experiments on face recognition. Yanwei Pang, Yuan Yuan 0001, Xuelong Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2005 | Improved Matching Pursuits Image CodingabstractThis paper reports improvements in compression of both inter- and intra-frame images by the matching pursuits (MP) algorithm. For both types of image, applying a 2D wavelet decomposition prior to MP coding is beneficial. The MP algorithm is then applied using various separable ID codebooks. MERGE coding with precision limited quantization (PLQ) is used to yield a highly embedded data stream. For inter-frames (residuals) a codebook of only 8 bases with compact footprint is found to give improved fidelity at lower complexity than previous MP methods. Great improvement is also achieved in MP coding of still images (intra-frames). Compared to JPEG 2000, lower distortion is achieved on the residual images tested, and also on intra-frames at low bit rates. Yuan Yuan 0001, Donald M. Monro |
ICASSP (2) | 1 |
| 2005 | Bases for low complexity matching pursuits image codingabstractThe basis size and number of bases in a basis dictionary (codebook) are the dominant parameters in the complexity of coding data by matching pursuits (MP). By basis picking for a range of constrained basis widths on a training image coded by 1D MP, dictionaries are obtained for both still and residual images and applied in both ID and 2D coding. On still (intra) images fidelity measured by PSNR is optimum for widths in the range 9 to 13 and dictionary sizes from 14 to 18. For residual images similar results are obtained but the bias towards narrower, smaller dictionaries is even more marked. The resulting combinations of 1D and 2D dictionaries offer significantly lower computational cost while maintaining high quality. Donald M. Monro, Yuan Yuan 0001 |
ICIP (2) | 2 |
| 2005 | 3D wavelet video coding with replicated matching pursuitsabstractImproved coding of 3D spatio-temporal wavelet video is achieved using matching pursuits (MP) with a separable 2D basis dictionary and temporal replication of atoms. Using two temporal scales and the biorthogonal 5/3 wavelet filter incorporating motion compensation, four sets of four transform planes are combined into a group of temporal planes (GOTP). The planes are wavelet transformed to 4 scales spatially before the 2D MP algorithm is applied over the entire GOTP to select the optimum inner product with a 2D dictionary of 64 bases formed separably from 8 ID basis functions. The selected atom is conditionally replicated in the three corresponding temporal frequency planes within the GOTP. Compared to applying 2D MP to the planes without replication, a gain of up to 2 dB has been achieved. Comparisons with a standard H.263 codec and the classic MP enhanced H.263 MP codec, show significant gains at low bit rates. Yuan Yuan 0001, Donald M. Monro |
ICIP (1) | 1 |
| 2005 | Cast shadow detection in video segmentation
Dong Xu 0001, Xuelong Li 0001, Zhengkai Liu, Yuan Yuan 0001 |
Pattern Recognit. Lett. | 4 |
| 2004 | Low complexity separable matching pursuits [video coding applications]abstractMethods of reducing the complexity of the matching pursuits algorithm with minimal loss of fidelity when coding displaced frame difference (DFD) images in video compression are investigated. A full search using 2D basis functions is used as a benchmark. The use of separable 1D bases greatly reduces the complexity, and significant further reductions are achieved by using only a 1D inner product search to locate the atom position, followed by a further 1D inner product search in the opposite direction to identify the second 1D basis function. To avoid ignoring significant structures orthogonal to the search direction, it is proposed to alternate the initial search direction between horizontal and vertical scanning. This produces a modest increase in distortion compared to the full 2D search, with a complexity reduction in excess of an order of magnitude. Yuan Yuan 0001, Adrian N. Evans, Donald M. Monro |
ICASSP (3) | 1 |
| 2004 | Artistic mosaic generationabstractIn this paper, a new approach to image processing in arts is presented in the domain of generating a series of artistic mosaic pictures. The mosaic pictures consist of the same elements as each other, but might be extremely different from the semantic contents. Some preliminary experimental results are given to show the impact of the proposed special techniques. Xuelong Li 0001, Yuan Yuan 0001 |
ICIG | 2 |
| 2004 | Anchorperson extraction for Picture in Picture news video
Dong Xu 0001, Xuelong Li 0001, Zhengkai Liu, Yuan Yuan 0001 |
Pattern Recognit. Lett. | 4 |
| 2003 | Adaptive color quantization based on perceptive edge protection
Xuelong Li 0001, Tianqiang Yuan, Nenghai Yu, Yuan Yuan 0001 |
Pattern Recognit. Lett. | 4 |