EDBT 2026 Demo / reviewers in the wild / expert
Chunhong Pan
dblp:26/3810
· DBLP profile ↗
267ranked-venue papers
4as first author
53since 2021 · last 2026
0000-0001-7433-4474ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 162 · 3 first-author · 22 since 2021Artificial intelligence and machine learning · 121 · 2 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 39 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph convolutional network with adaptive grouping aggregation strategy
Chunhong Pan |
Neural Networks | 3 |
| 2026 | USVTrack: A Benchmark for Multi-Object Tracking in Complex Water Surface ScenesabstractMulti-object tracking (MOT) in water surface scenes is crucial for the autonomous navigation of Unmanned Surface Vehicles (USVs). However, existing MOT datasets rarely focus on these scenes. Moreover, the few available water surface MOT datasets contain limited data shot onboard and concentrate narrowly on specific marine scenes, creating a significant gap from real-world USV navigation applications. To promote research on USV autonomous navigation, we introduce USVTrack, a fully onboard-shot MOT benchmark that covers diverse and complex water surface scenes, characterized by a high proportion of small objects and varied backgrounds. Then, we propose an innovative end-to-end method specifically designed for MOT in complex water surface scenes, termed as USVMOT. It improves tracking performance through four key contributions: 1) integrating mask information via knowledge distillation to boost feature discriminability; 2) deploying task-specific auxiliary pathways to alleviate the competition between detection and re-identification (ReID) in end-to-end MOT methods; 3) employing an adaptive high-quality mask generation strategy based on the Segment Anything Model (SAM) that obviates extensive manual annotation; and 4) introducing an object-aware association method that dynamically tailors the tracking strategy according to object size and motion speed. Extensive experiments on the USVTrack benchmark demonstrate that USVMOT outperforms existing methods. Our analysis reveals that MOT in complex water surface scenes remains challenging, highlighting the need for further advancements. Yuwei Cheng, Kun Ding 0001, Chunhong Pan, Shiming Xiang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic SegmentationabstractPre-training techniques significantly enhance the performance of semantic segmentation tasks with limited training data. However, the efficacy under a large domain gap between pre-training (e.g. RGB) and fine-tuning (e.g. infrared) remains underexplored. In this study, we first benchmark the infrared semantic segmentation performance of various pre-training methods and reveal several phenomena distinct from the RGB domain. Next, our layerwise analysis of pre-trained attention maps uncovers that: (1) There are three typical attention patterns (local, hybrid, and global); (2) Pre-training tasks notably influence pattern distribution across layers; (3) The hybrid pattern is crucial for semantic segmentation as it attends to both nearby and foreground elements; (4) The texture bias impedes model generalization in infrared tasks. Building on these insights, we propose UNIP, a UNified Infrared Pre-training framework, to enhance the pre-trained model performance. This framework uses the hybrid-attention distillation NMI-HAD as the pre-training target, a large-scale mixed dataset InfMix for pre-training, and a last-layer feature pyramid network LL-FPN for fine-tuning. Experimental results show that UNIP outperforms various pre-training methods by up to 13.5% in average mIoU on three infrared segmentation tasks, evaluated using fine-tuning and linear probing metrics. UNIP-S achieves performance on par with MAE-L while requiring only 1/10 of the computational cost. Furthermore, with fewer parameters, UNIP significantly surpasses state-of-the-art (SOTA) infrared or RGB segmentation methods and demonstrates the broad potential for application in other modalities, such as RGB and depth. Our code is available at https://github.com/casiatao/UNIP. Jinyong Wen, Kun Ding 0001, Shiming Xiang, Chunhong Pan |
ICLR | 6 |
| 2025 | Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference OptimizationabstractPreference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face challenges in handling noisy images of different timesteps and require complex transformations into pixel space. In this work, we show that pre-trained diffusion models are naturally suited for step-level reward modeling in the noisy latent space, as they are explicitly designed to process latent images at various noise levels. Accordingly, we propose the **Latent Reward Model (LRM)**, which repurposes components of the diffusion model to predict preferences of latent images at arbitrary timesteps. Building on LRM, we introduce **Latent Preference Optimization (LPO)**, a step-level preference optimization method conducted directly in the noisy latent space. Experimental results indicate that LPO significantly improves the model's alignment with general, aesthetic, and text-image alignment preferences, while achieving a 2.5-28x training speedup over existing preference optimization methods. Cheng Da, Kun Ding 0001, Huan Yang 0005, Yan Li 0043, Tingting Gao, Di Zhang 0026, Shiming Xiang, Chunhong Pan |
NeurIPS | 10 |
| 2025 | HAN: An efficient hierarchical self-attention network for skeleton-based gesture recognition
Ying Wang 0008, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 4 |
| 2024 | Weak Distribution Detectors Lead to Stronger Generalizability of Vision-Language Prompt TuningabstractWe propose a generalized method for boosting the generalization ability of pre-trained vision-language models (VLMs) while fine-tuning on downstream few-shot tasks. The idea is realized by exploiting out-of-distribution (OOD) detection to predict whether a sample belongs to a base distribution or a novel distribution and then using the score generated by a dedicated competition based scoring function to fuse the zero-shot and few-shot classifier. The fused classifier is dynamic, which will bias towards the zero-shot classifier if a sample is more likely from the distribution pre-trained on, leading to improved base-to-novel generalization ability. Our method is performed only in test stage, which is applicable to boost existing methods without time-consuming re-training. Extensive experiments show that even weak distribution detectors can still improve VLMs' generalization ability. Specifically, with the help of OOD detectors, the harmonic mean of CoOp and ProGrad increase by 2.6 and 1.5 percentage points over 11 recognition datasets in the base-to-novel setting. Kun Ding 0001, Haojian Zhang, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
AAAI | 6 |
| 2024 | An Efficient Graph Autoencoder with Lightweight Desmoothing Decoder and Long-Range ModelingabstractGraph self-supervised learning provides a powerful guarantee for learning high-quality representations in an unsupervised manner. Despite its early birth, the performance of generative graph self-supervised learning has long lagged behind that of up-and-coming contrastive learning, especially on node classification tasks. In this paper, we investigate potential issues in existing graph autoencoders and attribute their poor performance to three main aspects: complex decoder design, lack of desmoothing process in feature remap, and overemphasis on local topological proximity. To tackle these issues, we propose an effective and efficient graph autoencoder framework for unsupervised representation learning, which contains two key components: lightweight smoothness-aware feature reconstructor and global structural dependency catcher. After performing a desmoothing operation on encoded representations via a learnable high-pass filter, the feature decoder reconstructs the original features through a simple linear projection. The lightweight design liberates the decoder from self-supervised pretext tasks and puts the encoder more accountable for achieving optimization objectives, which promotes effective training of the encoder. Global structural dependency catcher utilizes graph diffusion to build a structural regularization to capture long-range topological dependency on a graph. The empirical studies demonstrate the effectiveness of our approach, which can surpass dominant contrastive learning methods. Jinyong Wen, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
ICDM | 5 |
| 2024 | Multi-task prompt tuning with soft context sharing for vision-language models
Kun Ding 0001, Ying Wang 0008, Pengzhang Liu, Haojian Zhang, Shiming Xiang, Chunhong Pan |
Neurocomputing | 7 |
| 2024 | Efficient Remote Sensing Image Super-Resolution via Lightweight Diffusion ModelsabstractWith the emergence of diffusion models, the image generation has experienced a significant advancement. In super-resolution tasks, diffusion models surpass generative adversarial network (GAN)-based methods in generating more realistic samples. However, these models come with significant costs: denoising networks rely on large U-Net, making them computationally intensive for high-resolution (HR) images, and the extensive sampling steps in diffusion models lead to prolonged inference time. This complexity limits their application in remote sensing, due to the high demand for high-resolution images in such scenarios. To address this, we propose a lightweight diffusion model (LWTDM), which simplifies the denoising network and efficiently incorporates conditional information using a cross-attention-based encoder–decoder architecture. Furthermore, LWTDM serves as the pioneering model that incorporates the accelerated sampling technique from denoising diffusion implicit models (DDIMs). This integration involves the meticulous selection of sampling steps, ensuring the quality of the generated images. The experiments confirm that LWTDM strikes a favorable balance between precision and perceptual quality, while its faster inference speed makes it suitable for diverse remote sensing scenarios with specific requirements. The source code is available at:https://github.com/Suanmd/LWTDM. Tai An, Chunlei Huo, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Preformer: Simple and Efficient Design for Precipitation Nowcasting With TransformersabstractThe primary objective of precipitation nowcasting is to predict precipitation patterns several hours in advance. Recent studies have emphasized the potential of deep learning methods for this task. To harness the correlations among various meteorological elements, existing frameworks project multiple meteorological elements into a latent space and then utilize convolutional-recurrent networks for future precipitation prediction. Although effective, the escalating model complexity may impede practical applications. This letter develops the Preformer, a streamlined Transformer framework for precipitation nowcasting that efficiently captures global spatiotemporal dependencies among multiple meteorological elements. The Preformer implements an encoder-translator-decoder architecture, where the encoder integrates spatial features of multiple elements, the translator models spatiotemporal dynamics, and the decoder combines spatiotemporal information to forecast future precipitation. Without introducing complex structures or strategies, the Preformer achieves state-of-the-art performance even with the least parameters. Qizhao Jin, Xinbang Zhang, Xinyu Xiao, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | Reusable Architecture Growth for Continual Stereo MatchingabstractThe remarkable performance of recent stereo depth estimation models benefits from the successful use of convolutional neural networks to regress dense disparity. Akin to most tasks, this needs gathering training data that covers a number of heterogeneous scenes at deployment time. However, training samples are typically acquired continuously in practical applications, making the capability to learn new scenes continually even more crucial. For this purpose, we propose to perform continual stereo matching where a model is tasked to 1) continually learn new scenes, 2) overcome forgetting previously learned scenes, and 3) continuously predict disparities at inference. We achieve this goal by introducing a Reusable Architecture Growth (RAG) framework. RAG leverages task-specific neural unit search and architecture growth to learn new scenes continually in both supervised and self-supervised manners. It can maintain high reusability during growth by reusing previous units while obtaining good performance. Additionally, we present a Scene Router module to adaptively select the scene-specific architecture path at inference. Comprehensive experiments on numerous datasets show that our framework performs impressively in various weather, road, and city circumstances and surpasses the state-of-the-art methods in more challenging cross-dataset settings. Further experiments also demonstrate the adaptability of our method to unseen scenes, which can facilitate end-to-end stereo architecture learning and practical deployment. Chenghao Zhang 0003, Gaofeng Meng, Bin Fan 0001, Zhaoxiang Zhang 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Improving the Homophily of Heterophilic Graphs for Semi-Supervised Node ClassificationabstractGraph Neural Networks (GNNs) have been applied to process the widespread graph data, including social networks and web data, etc. However, lots of GNNs can only perform well on homophilic graphs, while losing their superiority when tackling heterophilic graphs. Recent works try to use spectral theory or attention mechanism to design some more complex learning paradigms for heterophilic graphs. In this paper, we instead utilize some explored properties to construct three new graph structures of high homophily to improve the homophily of heterophilic graphs for better representation learning. Along with the original graph structure, totally four graph structures are injected into a Multi-View Graph Fusion Network (MVGFN) to learn a group of more expressive features for the semi-supervised node classification. Ablation experiments show that all three newly-constructed graph structures obtain higher homophily levels. Comparisons among several baselines indicate the superiority of our method on both homophilic and heterophilic graphs. Yuhu Wang, Shiming Xiang, Chunhong Pan |
ICME | 3 |
| 2023 | Graph Information Interaction on Feature and Structure via Cross-modal Contrastive LearningabstractThe abundant features and structure information on graphs provide a potential guarantee for learning high-quality representations without supervision. Feature attribute represents the inherent properties of nodes, while structure attribute describes their neighborhood relationship. These two types of attributes can be regarded as different modal forms of the same instance and should be consistent in identifying a member. We propose to directly regard feature and structure attributes as two separate views to embed this consistency into contrastive learning method, realizing graph information interaction on feature and structure in a cross-modal contrastive framework. Under this framework, node representations are learned in an unsupervised manner by maximizing the agreement between feature representation and structure representation. In terms of negative samples, instead of randomly sampling points from empirical distribution, a simple yet effective multi-sample mixing strategy is proposed to synthesize true negative samples with greater probability, alleviating the tricky false negative issue. Extensive experiments on multiple types of graphs demonstrate the effectiveness of the proposed method. Jinyong Wen, Yuhu Wang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
ICME | 5 |
| 2023 | Exploring Universal Principles for Graph Contrastive Learning: A Statistical PerspectiveabstractAlthough recent advances have prompted the prosperity in graph contrastive learning, the researches on universal principles for model design and desirable properties of latent representations are still inadequate. From a statistical perspective, this paper proposes two principles for guidance and constructs a general self-supervised framework for negative-free graph contrastive learning. Reformulating data augmentation as a mixture process, the first one, termed consistency principle, lays stress on exploring and mapping cross-view common information to consistent and essence-revealing representations. For the purpose of instantiation, four statistical indicators are employed to estimate and maximize the correlation between representations from various views, whose accordant variation trend during training implies the extraction of common content. With awareness of the insufficiency of a solo consistency principle, suffering from degenerated and coupled solutions, a decorrelation principle is put forward to encourage diverse and informative representations. Accordingly, two specific strategies, performing in representation space and eigen spectral space, respectively, are propounded to decouple various representation channels. Under two principles, various combinations of concrete implementations derive a family of methods. The comparison experiments with current state-of-the-arts demonstrate the effectiveness and sufficiency of two principles for high-quality graph representations. Furthermore, visual studies reveal how certain principles affect learned representations. Jinyong Wen, Shiming Xiang, Chunhong Pan |
ACM Multimedia | 3 |
| 2023 | Graph convolutional network with tree-guided anisotropic message passing
Yuhu Wang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
Neural Networks | 5 |
| 2023 | Patch loss: A generic multi-scale perceptual loss for single image super-resolutionabstractIn single image super-resolution (SISR), although PSNR is a key metric for signal fidelity, images with high PSNR do not necessarily render high visual quality. As a result, current perception-driven SISR methods employ perceptual metrics close to the human eye to measure the quality of the generated images. Unfortunately, the perceptual loss and adversarial loss, widely used by the perception-driven SISR methods, still underperform on these non-differentiable perceptual metrics. To this end, we propose a generic multi-scale perceptual loss, i.e., the patch loss, which can be easily plugged into off-the-shelf SISR methods to improve a broad range of perceptual metrics. Specifically, the proposed patch loss minimizes the multi-scale similarity of image patches and enhances the restoration of regions with complex textures and sharp edges via parameter-free adaptive patch-wise attention. Our proposed patch loss introduces more realistic details compared to the perceptual loss and fewer artifacts compared to the adversarial loss. Tai An, Binjie Mao, Chunlei Huo, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 6 |
| 2023 | AutoMSNet: Multi-Source Spatio-Temporal Network via Automatic Neural Architecture Search for Traffic Flow PredictionabstractRecently the research of traffic flow prediction with deep learning framework has be largely developed, whereas most current methods are still faced with the following shortcomings. For spatial feature extraction, studies have shown that both local and non-local correlations exist on traffic networks. Considering the temporal dependencies, short-term impending and longer periodic components are two most critical patterns of traffic data, which further provide different information for the prediction task. Furthermore, multi-source heterogeneous external data, which naturally holds semantic gap with traffic data, also have impact on traffic flow. To solve the above problems, this paper proposes an AutoMSNet (Multi-Source Spatio-Temporal Network via Automatic neural architecture search). The AutoMSNet is composed of an encoder-decoder structure. The encoder takes neighboring data as inputs, while the decoder captures long-term periodic patterns. Thus, different functions of two temporal features are simultaneously extracted. Moreover, a neural architecture search space is designed for spatial feature extraction. Through architecture search technique, graph convolutions with different receptive fields are automatically selected and combined to form an optimal module structure. Therefore, both local and non-local spatial features can be adaptively captured. Besides, a meta learning feature fusion strategy is proposed to integrate external data, which can alleviate the semantic gap between different data sources. Extensive experiments on three real-world traffic datasets evaluate the superiority of the proposed model. Shen Fang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Learning from the Target: Dual Prototype Network for Few Shot Semantic SegmentationabstractDue to the scarcity of annotated samples, the diversity between support set and query set becomes the main obstacle for few shot semantic segmentation. Most existing prototype-based approaches only exploit the prototype from the support feature and ignore the information from the query sample, failing to remove this obstacle.In this paper, we proposes a dual prototype network (DPNet) to dispose of few shot semantic segmentation from a new perspective. Along with the prototype extracted from the support set, we propose to build the pseudo-prototype based on foreground features in the query image. To achieve this goal, the cycle comparison module is developed to select reliable foreground features and generate the pseudo-prototype with them. Then, a prototype interaction module is utilized to integrate the information of the prototype and the pseudo-prototype based on their underlying correlation. Finally, a multi-scale fusion module is introduced to capture contextual information during the dense comparison between prototype (pseudo-prototype) and query feature. Extensive experiments conducted on two benchmarks demonstrate that our method exceeds previous state-of-the-arts with a sizable margin, verifying the effectiveness of the proposed method. Binjie Mao, Xinbang Zhang, Lingfeng Wang 0002, Qian Zhang 0009, Shiming Xiang, Chunhong Pan |
AAAI | 6 |
| 2022 | Discriminative Graph Representation Learning with Distributed SamplingabstractGraph neural networks (GNNs) have been widely used to accomplish graph classification tasks such as predicting molecular properties and classifying the labels of proteins. Discovering the latent discriminative substructures (e.g., functional groups in molecules) is a vital task to enhance the classification performance. In this paper, this task is addressed as a problem of discriminative graph representation learning. Specifically, a novel node sampling strategy is developed to achieve this goal. To this end, graph-dependent sampling vectors are first learned by a mini-network to exploit various informative substructures on graphs and sample some representative nodes, which could be regarded as performing a distributed sampling on graphs. Then, the sampled nodes are organized together topologically as a subgraph with landing probabilities of random walks. Moreover, a self-adaptive pooling ratio of nodes is obtained via feature smoothness of graphs, eliminating the trouble of manual selection of subgraph size. As a result, these treatments are equivalent to performing the difficult step of down-pooling operation on non-grid graph data. Extensive experiments and ablation studies on multiple benchmark datasets demonstrate the effectiveness and superiority of our proposed approach. Additionally, interpretability studies illustrate the ability of our model to extract discriminative substructures. Jinyong Wen, Yuhu Wang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
BIBM | 5 |
| 2022 | AME: Attention and Memory Enhancement in Hyper-Parameter OptimizationabstractTraining Deep Neural Networks (DNNs) is inherently subject to sensitive hyper-parameters and untimely feedbacks of performance evaluation. To solve these two difficulties, an efficient parallel hyper-parameter optimization model is proposed under the framework of Deep Reinforcement Learning (DRL). Technically, we develop Attention and Memory Enhancement (AME), that includes multi-head attention and memory mechanism to enhance the ability to capture both the short-term and long-term relationships between different hyper-parameter configurations, yielding an attentive sampling mechanism for searching high-performance configurations embedded into a huge search space. During the optimization of transformer-structured configuration searcher, a conceptually intuitive yet powerful strategy is applied to solve the problem of insufficient number of samples due to the untimely feedback. Experiments on three visual tasks, including image classification, object detection, semantic segmentation, demonstrate the effectiveness of AME. Nuo Xu 0006, Jianlong Chang, Xing Nie, Chunlei Huo, Shiming Xiang, Chunhong Pan |
CVPR | 6 |
| 2022 | Continual Stereo Matching of Continuous Driving Scenes with Growing ArchitectureabstractThe deep stereo models have achieved state-of-the-art performance on driving scenes, but they suffer from severe performance degradation when tested on unseen scenes. Although recent work has narrowed this performance gap through continuous online adaptation, this setup requires continuous gradient updates at inference and can hardly deal with rapidly changing scenes. To address these challenges, we propose to perform continual stereo matching where a model is tasked to 1) continually learn new scenes, 2) overcome forgetting previously learned scenes, and 3) continuously predict disparities at deployment. We achieve this goal by introducing a Reusable Architecture Growth (RAG) framework. RAG leverages task-specific neural unit search and architecture growth for continual learning of new scenes. During growth, it can maintain high reusability by reusing previous neural units while achieving good performance. A module named Scene Router is further introduced to adaptively select the scene-specific architecture path at inference. Experimental results demonstrate that our method achieves compelling performance in various types of challenging driving scenes. Chenghao Zhang 0003, Bin Fan 0001, Gaofeng Meng, Zhaoxiang Zhang 0001, Chunhong Pan |
CVPR | 6 |
| 2022 | Stereo Depth Estimation with Echoes
Chenghao Zhang 0003, Bolin Ni, Gaofeng Meng, Bin Fan 0001, Zhaoxiang Zhang 0001, Chunhong Pan |
ECCV (27) | 7 |
| 2022 | Spatiotemporal Contextual Consistency Network for Precipitation NowcastingabstractPrecipitation nowcasting is forecasting rainfall in the short-term conditioned by the known meteorological parameters. Recently, deep neural networks (DNNs) have shown outstanding performance in this task. But, there are several challenges imposed by the multiple meteorological elements, including the multimodal modeling, the considerable variation in scales of precipitation region, as well as the long-tailed distribution of rainfall data. To solve these problems, this paper proposes Spatiotemporal Contextual Consistency Network (SCCN) for learning from the multi meteorological elements. Architecturally, a parameter-shared multimodal fusion CNN encoder, which dynamically exchanges features between different modalities, is used to encode the multimodal meteorological data. To improve the spatial modeling of the multiple meteorological features, we compose the multi-scale filters and deconstruction convolution to modify the gate operators in ConvLSTM to propose a spatial contextual consistency ConvLSTM (SCC-ConvLSTM). Furthermore, considering the temporal consistency in rainfall, a temporal consistency module (TCM) is designed to gear to long-tailed distribution. Under this module, different long-tailed meteorological elements are calculated to encode features and residuals fused with the previous precipitation distribution in sequence. The experimental results of precipitation nowcasting demonstrate the effectiveness of our method on the ERA5 dataset and WeatherBench dataset. Xinyu Xiao, Qizhao Jin, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICDM | 5 |
| 2022 | Components Regulated Generation of Handwritten Chinese Text-lines in Arbitrary LengthabstractGenerating readable images of handwritten Chinese text-lines is very challenging due to complicated topological structures in Chinese. To address this problem, we propose a components regulated model named HCT-GAN to generate the entire lines of Chinese handwriting from text-line labels. Specifically, HCT-GAN is designed as a CGAN-based architecture that additionally integrates a Chinese text encoder (CTE), a sequence recognition module(SRM), and a spatial perception module (SPM). Compared with the one-hot embedding, CTE learns the latent content representation by reusing the structure and component embedding shared among the Chinese characters. SRM provides sequence-level constraints to the generated images. SPM can adaptively constrain the spatial correlation between the generated components, which facilitates the modeling of characters with complicated topological structures. Benefiting from such artful modeling, our model suffices to generate images of handwritten Chinese text-lines in arbitrary length. Extensive experimental results demonstrate that our model achieves state-of-the-art performance in handwritten Chinese lines generation. Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICPR | 5 |
| 2022 | AutoMF: Spatio-temporal Architecture Search for The Meteorological Forecasting TaskabstractDespite years of studies, meteorological forecasting with deep learning still faces several challenges, including the multi-modal correlation, the spatio-temporal dependency, the spatial heterogeneity, and the temporal periodicity. Though with elaborate design, manually designed networks adopted by current methods could be far from optimal in modeling the spatiotemporal dynamics of meteorological data. In this work, we propose AutoMF, i.e. a multi-level architecture search framework for the Meteorological Forecasting task. Working in a data-driven paradigm, AutoMF is capable of generating suitable convolution networks and spatio-temporal networks to capture the multi-modal correlation and the spatio-temporal dependency. Based on this framework, we develop a differential sampling-based architecture search method to optimize the architecture, and introduce the progressive search strategy to facilitate the search process. Furthermore, the spatial heterogeneity and temporal periodicity are explicitly modeled through integrating corresponding status indicators. Extensive experiments exhibit the capability of the proposed method. Xinbang Zhang, Qizhao Jin, Shiming Xiang, Chunhong Pan |
ICPR | 4 |
| 2022 | Task-aware adaptive attention learning for few-shot semantic segmentation
Binjie Mao, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
Neurocomputing | 4 |
| 2022 | TVGCN: Time-variant graph convolutional network for traffic forecasting
Yuhu Wang, Shen Fang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
Neurocomputing | 5 |
| 2022 | AHDet: A dynamic coarse-to-fine gaze strategy for active object detection
Nuo Xu 0006, Chunlei Huo, Xin Zhang 0093, Chunhong Pan |
Neurocomputing | 4 |
| 2022 | PSNet: Perspective-sensitive convolutional network for object detection
Xin Zhang 0093, Chunlei Huo, Nuo Xu 0006, Lingfeng Wang 0002, Chunhong Pan |
Neurocomputing | 6 |
| 2022 | Learning adversarial point-wise domain alignment for stereo matching
Chenghao Zhang 0003, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Neurocomputing | 5 |
| 2022 | Monocular contextual constraint for stereo matching with adaptive weights assignment
Chenghao Zhang 0003, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Image Vis. Comput. | 5 |
| 2022 | HENet: Head-Level Ensemble Network for Very High Resolution Remote Sensing Images Semantic SegmentationabstractSemantic segmentation plays an important role in very high resolution (VHR) image understanding. Despite the potentials of the deep convolutional network in improving performance by end-to-end feature learning, each model has its limitations, and it is hard to discriminate complex features purely by a single model. Ensemble learning is promising for integrating the strengths of different models, however, the ensemble of deep models is challenging due to the huge amount of parameters and computation of the deep model itself as well as the difficulty in capturing complementarity between different models. To tackle these problems, a head-level ensemble network (HENet) is proposed in this letter, which reduces model complexity by sharing feature extraction networks and improves complementarity between models by novel cooperative learning (CL). Experiments on ISPRS 2-D semantic labeling benchmark demonstrate the effectiveness and advantage of the proposed method. Chunlei Huo, Nuo Xu 0005, Xin Zhang 0093, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Train in Dense and Test in Sparse: A Method for Sparse Object Detection in Aerial ImagesabstractApplications of aerial imaging, especially based on unmanned aerial vehicles (UAVs) platform, rapidly explode in recent years. Meanwhile, vision-based sensing, e.g., detection and recognition, for UAVs becomes increasingly important. Objects in aerial images are usually of tiny size, hence occupying a limited area. Terminology speaking, the images are very sparse in spatial. However, existing work in aerial object detection commonly ignores this point. Conversely, we explore the availability of such a property in improving the detection performance of aerial images. Specifically, we propose a general method, train in dense and test in sparse (TDTS), to exploit sparsity in aerial object detection: 1) in the training stage, the possible positions of object are learned by training a fully convolutional network (called prophet head) and 2) in the testing stage, prophet head identifies the possible object locations to reduce redundant computation in classification and box prediction head by sparse convolution. By extensive experiments on the VisDrone2019-Det data set, we find that the sparsity can not only help to speed up inference but also to improve accuracy. Thus, we argue that the sparsity deserves more attention. Kun Ding 0001, Guojin He, Huxiang Gu, Zisha Zhong, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Hyperparameter Configuration Learning for Ship Detection From Synthetic Aperture Radar ImagesabstractDetecting ships from synthetic aperture radar (SAR) images is inherently subject to its imaging mechanism. With the development of deep learning, advanced learning-based techniques have been migrated from optical images to SAR images. However, the default hyperparameters (e.g., learning rate, size of the anchor box) predefined by a heuristic strategy on optical images might be suboptimal for SAR datasets. In addition, the low-quality imaging in SAR images further reduces the portability of hyperparameters. To solve this problem, a new optimization method, named reinforcement learning and hyperband (RLH), is proposed to dynamically learn hyperparameter configurations by deep reinforcement learning (DRL), where a neural network is adopted to capture the relationship between different configurations and predict new configurations to further improve the performance. Hyperparameter configuration is able to be automatically learned to accommodate various SAR image datasets, and experiments on two SAR image datasets demonstrate the effectiveness and advantage of the proposed approach. Nuo Xu 0005, Chunlei Huo, Xin Zhang 0093, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | CMT: Cross Mean Teacher Unsupervised Domain Adaptation for VHR Image Semantic SegmentationabstractSemantic segmentation of remote sensing images has achieved superior results with the supervised deep learning models. However, their performance to unseen data domains could be very bad due to the domain shift between different domains. Recently, a series of unsupervised domain adaptation (UDA) methods has been developed to solve the domain shift problem in semantic segmentation. Most of them use adversarial learning to achieve global cross-domain alignment and use a self-training (ST) strategy to generate pseudo-labels for classwise alignment. However, these methods ignore the pixels that are not assigned pseudo-labels. Those pixels are mostly at the boundaries, which are vital to the final segmentation results. To solve this problem, this letter proposes a cross mean teacher (CMT) UDA method. The whole framework consists of two parts. On the one hand, the global cross-domain distribution alignment is performed, and then, reliable pseudo-labels are assigned to the target data. On the other hand, a cross teacher–student network (CTSN) is developed to effectively use those pixels with and without pseudo-labels. This network contains two student networks ($S_{1}$and$S_{2}$) and two teacher networks ($T_{1}$and$T_{2}$) for cross-consistency constraints that supervises$S_{2}$(or$S_{1}$) by the prediction results of$T_{1}$(or$T_{2}$). The cross supervision by CTSN is helpful to prevent performance bottlenecks caused by the high coupling of teacher–student network in existing methods. Extensive experiments on three different remote sensing adaptation scenes verify the effectiveness and superiority of the proposed method. Bin Fan 0001, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | MFNet: The Spatio-Temporal Network for Meteorological Forecasting With Architecture SearchabstractExploiting deep learning for the meteorological forecasting task is a challenging due to the complex spatio-temporal correlation, non-stationarity and imbalanced data distribution. Though with elaborate design, handcraft hierarchical architectures adopted by current methods could be far from optimal in sufficiently modeling the dynamics of meteorological data. For the Meteorological Forecasting task, this letter presents the MFNet, which is a spatio-temporal network with the Neural Architecture Search (NAS) technique. Working in the data-driven paradigm, our method is capable of automatically generating suitable architecture to model the spatio-temporal correlation. Moreover, the non-stationarity of meteorological data is explicitly modeled through simulating spatio-temporal variations in response to the intrinsic driven force of the meteorological state, and the Error Sensitive Regression (ESR) loss is introduced accounting for the imbalanced data distribution. Extensive experiments exhibit the capability of our method and demonstrate that deep learning is potential for serving as an operational technique for global meteorological forecasting. Xinbang Zhang, Qizhao Jin, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Subgraph-aware graph structure revision for spatial-temporal graph modeling
Yuhu Wang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
Neural Networks | 4 |
| 2022 | Scene captioning with deep fusion of images and point clouds
Chunxia Zhang 0001, Lubin Weng, Shiming Xiang, Chunhong Pan |
Pattern Recognit. Lett. | 5 |
| 2022 | MS-Net: Multi-Source Spatio-Temporal Network for Traffic Flow PredictionabstractPredicting urban traffic flow is a challenging task, due to the complicated spatio-temporal dependencies on traffic networks. Urban traffic flow usually has both short-term neighboring and long-term periodic temporal dependencies. It is also noticed that the spatial correlations over different traffic nodes are both local and non-local. What’s more, the traffic flow is affected by various external factors. To capture the non-local spatial correlations, we propose a Dilated Attentional Graph Convolution (DAGC). The DAGC utilizes a dilated graph convolution kernel to expand the nodes’ receptive field and exploit multi-order neighborhood. Technically, the lower-order neighborhood corresponds to local spatial dependencies, while the higher-order neighborhood corresponds to non-local spatial dependencies between nodes. Based on DAGC, a Multi-Source Spatio-Temporal Network (MS-Net) is designed, which suffices to integrate long-range historical traffic data as well as multi-modal external information. MS-Net consists of four components: a spatial feature extraction module, a temporal feature fusion module, an external factors embedding module, and a multi-source data fusion module. Extensive experiments on three real traffic datasets demonstrates that the proposed model performs well on both the public transportation networks, road networks, and can handle large-scale traffic networks in particular the Beijing bus network which has more than 4,000 traffic nodes. Shen Fang, Véronique Prinet, Jianlong Chang, Michael Werman, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2022 | Decoupled Representation Learning for Character Glyph SynthesisabstractCharacter glyph synthesis is still an open challenging problem, which involves two related aspects,i.e., font style transfer and content consistency. In this paper, we propose a novel model named FontGAN, which integrates the character structure stylization, de-stylization and texture transfer into a unified framework. Specifically, we decouple character images into style representation and content representation, which offers fine-grained control of these two types of variables, thus improving the quality of the generated results. To effectively capture the style information, a style consistency module (SCM) is introduced. Technically, SCM exploits category-guided Kullback-Leibler divergence to explicitly model the style representation into different prior distributions. In this way, our model is capable of implementing transformations between multiple domains in one framework. In addition, we propose content prior module (CPM) to provide content prior for the model to guide the content encoding process and alleviates the problem of stroke deficiency during structure de-stylization. Benefiting from the idea of decoupling and regrouping, our FontGAN suffices to achieve many-to-many translation tasks for glyph structure. Experimental results demonstrate that the proposed FontGAN achieves the state-of-the-art performance in character glyph synthesis. Xiyan Liu, Gaofeng Meng, Jianlong Chang, Ruiguang Hu, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 6 |
| 2021 | Ltaf-Net: Learning Task-Aware Adaptive Features and Refining Mask for Few-Shot Semantic SegmentationabstractFew shot segmentation is a newly-developing and challenging computer vision task which is only provided with few labeled samples of the novel class. Some recent works on this problem focus more on how to design an effective comparison module but ignore how to extract the features passed to compare. In this paper we propose a novel model named LTAF-Net for few-shot segmentation. This model aims to adaptively recalibrate the extracted features which could boost the accuracy of dense comparison between support features and query features. Besides an additional prediction refinement module is designed to refine the initial mask. Meanwhile this method can apply to k-shot setting without developing a new specialized architecture and achieve competitive performance. Experiments on PASCAL-5iand FSS-1000 strongly prove the effectiveness of the proposed model. Our model outperforms the second-best method 1.4% in 1-shot and 0.76% in 5-shot respectively in PASCAL-5i. Binjie Mao, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
ICASSP | 4 |
| 2021 | Reinforcement Stacked Learning with Semantic-Associated Attention for Visual Question AnsweringabstractThe task of visual question answering (VQA) is to generate an answer for a question according to the content of an image being asked. In this process, the critical problems of effectively embedding the question feature and image feature as well as transforming the features to the prediction of answer are still faithfully unresolved. In this paper, depending on these problems, a semantic-associated attention method and a reinforcement stacked learning mechanism are proposed. Firstly, within the associations of high-level semantics, a visual spatial attention model (VSA) and a multi-semantic attention model (MSA) are proposed to extract the low-level image feature and high-level semantic feature, respectively. Furthermore, we develop a reinforcement stacked learning architecture, which splits the transformation process into multiple stages, to gradually approach the answers. At each stage, a new reinforcement learning (RL) method is introduced to directly criticize inappropriate answers to optimize the model. The extensive experiments on the VQA task show that our method can achieve state-of-the-art performance. Xinyu Xiao, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
ICASSP | 4 |
| 2021 | Differentiable Convolution Search for Point Cloud ProcessingabstractExploiting convolutional neural networks for point cloud processing is quite challenging, due to the inherent irregular distribution and discrete shape representation of point clouds. To address these problems, many handcrafted convolution variants have sprung up in recent years. Though with elaborate design, these variants could be far from optimal in sufficiently capturing diverse shapes formed by discrete points. In this paper, we propose PointSeaConv, i.e., a novel differential convolution search paradigm on point clouds. It can work in a purely data-driven manner and thus is capable of auto-creating a group of suitable convolutions for geometric shape modeling. We also propose a joint optimization framework for simultaneous search of internal convolution and external architecture, and introduce epsilon-greedy algorithm to alleviate the effect of discretization error. As a result, PointSeaNet, a deep network that is sufficient to capture geometric shapes at both convolution level and architecture level, can be searched out for point cloud processing. Extensive experiments strongly evidence that our proposed PointSeaNet surpasses current handcrafted deep models on challenging benchmarks across multiple tasks with remarkable margins. Xing Nie, Yongcheng Liu, Shaohong Chen, Jianlong Chang, Chunlei Huo, Gaofeng Meng, Qi Tian 0001, Chunhong Pan |
ICCV | 9 |
| 2021 | Knowledge Mining and Transferring for Domain Adaptive Object DetectionabstractWith the thriving of deep learning, CNN-based object detectors have made great progress in the past decade. However, the domain gap between training and testing data leads to a prominent performance degradation and thus hinders their application in the real world. To alleviate this problem, Knowledge Transfer Network (KTNet) is proposed as a new paradigm for domain adaption. Specifically, KT-Net is constructed on a base detector with intrinsic knowledge mining and relational knowledge constraints. First, we design a foreground/background classifier shared by source domain and target domain to extract the common attribute knowledge of objects in different scenarios. Second, we model the relational knowledge graph and explicitly constrain the consistency of category correlation under source domain, target domain, as well as cross-domain conditions. As a result, the detector is guided to learn object-related and domain-independent representation. Extensive experiments and visualizations confirm that transferring object-specific knowledge can yield notable performance gains. The proposed KTNet achieves state-of-the-art results on three cross-domain detection benchmarks. Chenghao Zhang 0003, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
ICCV | 5 |
| 2021 | AsyNCE: Disentangling False-Positives for Weakly-Supervised Video GroundingabstractWeakly-supervised video grounding has been investigated to ground textual phases in video content with only video-sentence pairs provided during training, for the lack of prohibitively costly bounding box annotations. Existing methods cast this task into a frame-level multiple instance learning (MIL) problem with the ranking loss. While an object might appear sparsely across multiple frames, causing uncertain false-positive frames. Thus, directly computing the average loss of all frames is inadequate in video domain. Moreover, the positive and negative pairs are equally coupling in ranking loss, so that it is impossible to handle false-positive frames individually. Additionally, naive inner production is suboptimal for the similarity measure of cross domains. To solve these issues, we propose a novel AsyNCE loss to flexibly disentangle the positive pairs from negative ones in frame-level MIL, which allows for mitigating the uncertainty of false-positive frames effectively. Besides, a cross-modal transformer block is introduced to purify the text feature by image frame context, generating a visual-guided text feature for better similarity measure. Extensive experiments on YouCook2, RoboWatch and WAB datasets demonstrate the superiority and robustness of our method over state-of-the-art methods. Cheng Da, Yanhao Zhang 0002, Chunhong Pan |
ACM Multimedia | 6 |
| 2021 | Dual Stream Fusion Network for Multi-spectral High Resolution Remote Sensing Image Segmentation
Yiwen Shi, Chunlei Huo, Shiming Xiang, Chunhong Pan |
PRCV (2) | 6 |
| 2021 | 3D-SceneCaptioner: Visual Scene Captioning Network for Three-Dimensional Point Clouds
Xianbing Pan, Shiming Xiang, Chunhong Pan |
PRCV (2) | 4 |
| 2021 | Few-Shot Learning via Feature Hallucination with Variational InferenceabstractDeep learning has achieved huge success in the field of artificial intelligence, but the performance heavily depends on labeled data. Few-shot learning aims to make a model rapidly adapt to unseen classes with few labeled samples after training on a base dataset, and this is useful for tasks lacking labeled data such as medical image processing. Considering that the core problem of few-shot learning is the lack of samples, a straightforward solution to this issue is data augmentation. This paper proposes a generative model (VI-Net) based on a cosine-classifier baseline. Specifically, we construct a framework to learn to define a generating space for each category in the latent space based on few support samples. In this way, new feature vectors can be generated to help make the decision boundary of classifier sharper during the fine-tuning process. To evaluate the effectiveness of our proposed approach, we perform comparative experiments and ablation studies on mini-ImageNet and CUB. Experimental results show that VI-Net does improve performance compared with the baseline and obtains the state-of-the-art result among other augmentation-based methods. Qinxuan Luo, Lingfeng Wang 0002, Jingguo Lv, Shiming Xiang, Chunhong Pan |
WACV | 5 |
| 2021 | Dynamic camera configuration learning for high-confidence active object detection
Nuo Xu 0006, Chunlei Huo, Xin Zhang 0093, Gaofeng Meng, Chunhong Pan |
Neurocomputing | 6 |
| 2021 | DATA: Differentiable ArchiTecture Approximation With Distribution Guided SamplingabstractNeural architecture search (NAS) is inherently subject to the gap of architectures during searching and validating. To bridge this gap effectively, we develop Differentiable ArchiTecture Approximation (DATA) with Ensemble Gumbel-Softmax (EGS) estimator and Architecture Distribution Constraint (ADC) to automatically approximate architectures during searching and validating in a differentiable manner. Technically, the EGS estimator consists of a group of Gumbel-Softmax estimators, which is capable of converting probability vectors to binary codes and passing gradients reversely, reducing the estimation bias in a differentiable way. To narrow the distribution gap between sampled architectures and supernet, further, the ADC is introduced to reduce the variance of sampling during searching. Benefiting from such modeling, architecture probabilities and network weights in the NAS model can be jointly optimized with the standard back-propagation, yielding an end-to-end learning mechanism for searching deep neural architectures in an extended search space. Conclusively, in the validating process, a high-performance architecture that approaches to the learned one during searching is readily built. Extensive experiments on various tasks including image classification, few-shot learning, unsupervised clustering, semantic segmentation and language modeling strongly demonstrate that DATA is capable of discovering high-performance architectures while guaranteeing the required efficiency. Code is available at https://github.com/XinbangZhang/DATA-NAS. Xinbang Zhang, Jianlong Chang, Yiwen Guo, Gaofeng Meng, Shiming Xiang, Zhouchen Lin, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2021 | You Only Search Once: Single Shot Neural Architecture Search via Direct Sparse OptimizationabstractRecently neural architecture search (NAS) has raised great interest in both academia and industry. However, it remains challenging because of its huge and non-continuous search space. Instead of applying evolutionary algorithm or reinforcement learning as previous works, this paper proposes a direct sparse optimization NAS (DSO-NAS) method. The motivation behind DSO-NAS is to address the task in the view of model pruning. To achieve this goal, we start from a completely connected block, and then introduce scaling factors to scale the information flow between operations. Next, sparse regularizations are imposed to prune useless connections in the architecture. Lastly, an efficient and theoretically sound optimization method is derived to solve it. Our method enjoys both advantages of differentiability and efficiency, therefore it can be directly applied to large datasets like ImageNet and tasks beyond classification. Particularly, on the CIFAR-10 dataset, DSO-NAS achieves an average test error 2.74 percent, while on the ImageNet dataset DSO-NAS achieves 25.4 percent test error under 600M FLOPs with 8 GPUs in 18 hours. As for semantic segmentation task, DSO-NAS also achieve competitive result compared with manually designed architectures on the PASCAL VOC dataset. Code is available at https://github.com/XinbangZhang/DSO-NAS. Xinbang Zhang, Zehao Huang, Naiyan Wang, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Meta-MSNet: Meta-Learning Based Multi-Source Data Fusion for Traffic Flow PredictionabstractTraffic flow prediction is a challenging task while most existing works are faced with two main problems in extracting complicated intrinsic and extrinsic features. In terms of intrinsic features, current methods don't fully exploit different functions of short-term neighboring and long-term periodic temporal patterns. As for extrinsic features, recent works mainly employ hand-crafted fusion strategies to integrate external factors but remain generalization issues. To solve these problems, we propose a meta-learning based multi-source spatio-temporal network (Meta-MSNet). The Meta-MSNet is designed with an encoder-decoder structure. The encoder captures neighboring temporal dependencies while the decoder extracts periodic features. Furthermore, two meta-learning based fusion modules are designed to integrate multi-source external data both on temporal and spatial dimensions. Experiments on three real-world traffic datasets have verified the superiority of the proposed model. Shen Fang, Xianbing Pan, Shiming Xiang, Chunhong Pan |
IEEE Signal Process. Lett. | 4 |
| 2021 | Handwritten Text Generation via Disentangled RepresentationsabstractAutomatically generating handwritten text images is a challenging task due to the diverse handwriting styles and the irregular writing in natural scenes. In this paper, we propose an effective generative model called HTG-GAN to synthesize handwritten text images from latent prior. Unlike single-character synthesis, our method is capable of generating images of sequence characters with arbitrary length, which pays more attention to the structural relationship between characters. We model the structural relationship as the style representation to avoid explicitly modeling the stroke layout. Specifically, the text image is disentangled into style representation and content representation, where the style representation is mapped into Gaussian distribution and the content representation is embedded using character index. In this way, our model can generate new handwritten text images with specified contents and various styles to perform data augmentation, thereby boosting handwritten text recognition (HTR). Experimental results show that our method achieves state-of-the-art performance in handwritten text generation. Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
IEEE Signal Process. Lett. | 4 |
| 2020 | Spatio-Temporal Graph Structure Learning for Traffic ForecastingabstractAs an indispensable part in Intelligent Traffic System (ITS), the task of traffic forecasting inherently subjects to the following three challenging aspects. First, traffic data are physically associated with road networks, and thus should be formatted as traffic graphs rather than regular grid-like tensors. Second, traffic data render strong spatial dependence, which implies that the nodes in the traffic graphs usually have complex and dynamic relationships between each other. Third, traffic data demonstrate strong temporal dependence, which is crucial for traffic time series modeling. To address these issues, we propose a novel framework named Structure Learning Convolution (SLC) that enables to extend the traditional convolutional neural network (CNN) to graph domains and learn the graph structure for traffic forecasting. Technically, SLC explicitly models the structure information into the convolutional operation. Under this framework, various non-Euclidean CNN methods can be considered as particular instances of our formulation, yielding a flexible mechanism for learning on the graph. Along this technical line, two SLC modules are proposed to capture the global and local structures respectively and they are integrated to construct an end-to-end network for traffic forecasting. Additionally, in this process, Pseudo three Dimensional convolution (P3D) networks are combined with SLC to capture the temporal dependencies in traffic data. Extensively comparative experiments on six real-world datasets demonstrate our proposed approach significantly outperforms the state-of-the-art ones. Jianlong Chang, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
AAAI | 5 |
| 2020 | AugFPN: Improving Multi-Scale Feature Learning for Object DetectionabstractCurrent state-of-the-art detectors typically exploit feature pyramid to detect objects at different scales. Among them, FPN is one of the representative works that build a feature pyramid by multi-scale features summation. However, the design defects behind prevent the multi-scale features from being fully exploited. In this paper, we begin by first analyzing the design defects of feature pyramid in FPN, and then introduce a new feature pyramid architecture named AugFPN to address these problems. Specifically, AugFPN consists of three components: Consistent Supervision, Residual Feature Augmentation, and Soft RoI Selection. AugFPN narrows the semantic gaps between features of different scales before feature fusion through Consistent Supervision. In feature fusion, ratio-invariant context information is extracted by Residual Feature Augmentation to reduce the information loss of feature map at the highest pyramid level. Finally, Soft RoI Selection is employed to learn a better RoI feature adaptively after feature fusion. By replacing FPN with AugFPN in Faster R-CNN, our models achieve 2.3 and 1.6 points higher Average Precision (AP) when using ResNet50 and MobileNet-v2 as backbone respectively. Furthermore, AugFPN improves RetinaNet by 1.6 points AP and FCOS by 0.9 points AP when using ResNet50 as backbone. Codes are available on https://github.com/Gus-Guo/AugFPN. Chaoxu Guo, Bin Fan 0001, Qian Zhang 0009, Shiming Xiang, Chunhong Pan |
CVPR | 5 |
| 2020 | Decoupled Representation Learning for Skeleton-Based Gesture RecognitionabstractSkeleton-based gesture recognition is very challenging, as the high-level information in gesture is expressed by a sequence of complexly composite motions. Previous works often learn all the motions with a single model. In this paper, we propose to decouple the gesture into hand posture variations and hand movements, which are then modeled separately. For the former, the skeleton sequence is embedded into a 3D hand posture evolution volume (HPEV) to represent fine-grained posture variations. For the latter, the shifts of hand center and fingertips are arranged as a 2D hand movement map (HMM) to capture holistic movements. To learn from the two inhomogeneous representations for gesture recognition, we propose an end-to-end two-stream network. The HPEV stream integrates both spatial layout and temporal evolution information of hand postures by a dedicated 3D CNN, while the HMM stream develops an efficient 2D CNN to extract hand movement features. Eventually, the predictions of the two streams are aggregated with high efficiency. Extensive experiments on SHREC'17 Track, DHG-14/28 and FPHA datasets demonstrate that our method is competitive with the state-of-the-art. Yongcheng Liu, Ying Wang 0008, Véronique Prinet, Shiming Xiang, Chunhong Pan |
CVPR | 6 |
| 2020 | PackDet: Packed Long-Head Object Detector
Kun Ding 0001, Guojin He, Huxiang Gu, Zisha Zhong, Shiming Xiang, Chunhong Pan |
ECCV (13) | 6 |
| 2020 | Learning Where to Focus for Efficient Video Object Detection
Zhengkai Jiang 0001, Yu Liu 0015, Ceyuan Yang, Jihao Liu, Peng Gao 0007, Qian Zhang 0009, Shiming Xiang, Chunhong Pan |
ECCV (16) | 8 |
| 2020 | View-Angle Invariant Object Monitoring Without Image RegistrationabstractObject monitoring can be performed by change detection algorithms. However, for the image pair with a large perspective difference, the change detection performance is usually impacted by inaccurate image registration. To address the above difficulties, a novel object-specific change detection approach is proposed for object monitoring in this paper. In contrast to traditional approaches, the proposed approach is robust to view angle variation and does not require explicit image registration. Experiments demonstrate the effectiveness and advantages of the proposed approach. Xin Zhang 0093, Chunlei Huo, Chunhong Pan |
ICASSP | 3 |
| 2020 | PointSpherical: Deep Shape Context for Point Cloud Learning in Spherical CoordinatesabstractWe propose Spherical Hierarchical modeling of 3D point cloud. Inspired by Shape Context, we design a receptive field on each 3D point by placing a spherical coordinate on it. We sample points using the furthest point method and creating overlapping balls of points. We divide the space into radial, polar angular, and azimuthal angular bins on which we form a Spherical Hierarchy for each ball. We apply 1x1 CNN convolution on points to start the initial feature extraction. Repeated 3D CNN and max-pooling over the Spherical bins propagate contextual information until all the information is condensed in the center bin. Extensive experiments on five datasets strongly evidence that our method outperforms current models on various Point Cloud Learning tasks, including 2D/3D shape classification, 3D part segmentation, and 3D semantic segmentation. Bin Fan 0001, Yongcheng Liu, Yirong Yang, Jianbo Shi, Chunhong Pan, Huiwen Xie |
ICPR | 7 |
| 2020 | Forground-Guided Vehicle Perception FrameworkabstractAs the basis of advanced visual tasks such as vehicle tracking and traffic flow analysis, vehicle detection needs to accurately predict the position and category of vehicle objects. In the past decade, deep learning based methods have made great progress. However, we also notice that some existing cases are not studied thoroughly. First, false positive on the background regions is one of the critical problems. Second, most of the previous approaches only optimize a single vehicle detection model, ignoring the relationship between different visual perception tasks. In response to the above two findings, we introduce a foreground segmentation branch for the first time, which can predict the pixel level of vehicles in advance. Furthermore, two attention modules are designed to guide the work of the detection branch. The proposed method can be easily grafted into the one-stage and two-stage detection framework. We evaluate the effectiveness of our model on LSVH, a dataset with large variations in vehicle scales, and achieve the state-of-the-art detection accuracy. Shiming Xiang, Chunhong Pan |
ICPR | 4 |
| 2020 | Adaptive Remote Sensing Image Attribute Learning for Active Object DetectionabstractIn recent years, deep learning methods bring incredible progress to the field of object detection. However, in the field of remote sensing image processing, existing methods neglect the relationship between imaging configuration and detection performance, and do not take into account the importance of detection performance feedback for improving image quality. Therefore, detection performance is limited by the passive nature of the conventional object detection framework. In order to solve the above limitations, this paper takes adaptive brightness adjustment and scale adjustment as examples, and proposes an active object detection method based on deep reinforcement learning. The goal of adaptive image attribute learning is to maximize the detection performance. With the help of active object detection and image attribute adjustment strategies, low-quality images can be converted into high-quality images, and the overall performance is improved without retraining the detector. Nuo Xu 0006, Chunlei Huo, Jiacheng Guo, Jian Wang 0068, Chunhong Pan |
ICPR | 6 |
| 2020 | Deep Space Probing for Point Cloud Analysisabstract3D points distribute in a continuous 3D space irregularly, thus directly adapting 2D image convolution to 3D points is not an easy job. Previous works often artificially divide the space into regular grids, yet it could be suboptimal to learn geometry. In this paper, we propose SPCNN, namely, Space Probing Convolutional Neural Network, which naturally generalizes image CNN to deal with point clouds. The key idea of SPCNN is learning to probe the 3D space in an adaptive manner. Specifically, we define a pool of learnable convolutional weights, and let each point in the local region learn to choose a suitable convolutional weight from the pool. This is achieved by constructing a geometry guided index-mapping function that implicitly establishes a correspondence between convolutional weights and some local regions in the neighborhood (Fig. 1). In this way, the index-mapping function learns to adaptively partition nearby space for local geometry pattern recognition. With this convolution as a basic operator, SPCNN, a hierarchical architecture can be developed for effective point cloud analysis. Extensive experiments on challenging benchmarks across three tasks demonstrate that SPCNN achieves the state-of-the-art or has competitive performance. Yirong Yang, Bin Fan 0001, Yongcheng Liu, Jiyong Zhang 0001, Xin Liu 0027, Xinyu Cai, Shiming Xiang, Chunhong Pan |
ICPR | 9 |
| 2020 | Detecting Maneuvering Target Accurately Based on a Two-Phase Approach From Remote Sensing ImageryabstractManeuvering target detection in satellite images is difficult due to their small sizes, blurred appearances under various illuminations and shadows, and occlusion by trees and buildings. Recently, a fully convolutional regression network (FCRN) was proposed and achieved by the state-of-the-art performance in the Munich vehicle database. However, such a one-phase approach often makes mistakes at difficult places because of its swift glance and rejecting any second check. In this letter, a new object spatial density building net (SDBN) was designed, and a two-phase detection approach was proposed. It used the first SDBN to generate candidate regions and the second SDBN to proceed with a meticulous check on the object categories. Experiments on four maneuvering target databases, the Munich vehicle database, the Open Vehicle Database of San-Francisco (OVDS), the Overhead Imagery Research Data Set (OIRDS), and the Open Aircraft Database (OAD) show that the proposed method outperforms FCRN by an obvious margin. In addition, the accurate geometrical parameters (positions, orientations, and lengths) of all the objects were computed based on the spatial density maps, and the published experimental result of FCRN in OIRDS was pointed out and the corrected result was given. All source codes and databases are available at http://www.github.com/cxy177/SDBN. Xueyun Chen, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2020 | Deep Self-Evolution ClusteringabstractClustering is a crucial but challenging task in pattern analysis and machine learning. Existing methods often ignore the combination between representation learning and clustering. To tackle this problem, we reconsider the clustering task from its definition to develop Deep Self-Evolution Clustering (DSEC) to jointly learn representations and cluster data. For this purpose, the clustering task is recast as a binary pairwise-classification problem to estimate whether pairwise patterns are similar. Specifically, similarities between pairwise patterns are defined by the dot product between indicator features which are generated by a deep neural network (DNN). To learn informative representations for clustering, clustering constraints are imposed on the indicator features to represent specific concepts with specific representations. Since the ground-truth similarities are unavailable in clustering, an alternating iterative algorithm called Self-Evolution Clustering Training (SECT) is presented to select similar and dissimilar pairwise patterns and to train the DNN alternately. Consequently, the indicator features tend to be one-hot vectors and the patterns can be clustered by locating the largest response of the learned indicator features. Extensive experiments strongly evidence that DSEC outperforms current models on twelve popular image, text and audio datasets consistently. Jianlong Chang, Gaofeng Meng, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Local-Aggregation Graph NetworksabstractConvolutional neural networks (CNNs) provide a dramatically powerful class of models, but are subject to traditional convolution that can merely aggregate permutation-ordered and dimension-equal local inputs. It causes that CNNs are allowed to only manage signals on Euclidean or grid-like domains (e.g., images), not ones on non-Euclidean or graph domains (e.g., traffic networks). To eliminate this limitation, we develop a local-aggregation function, a sharable nonlinear operation, to aggregate permutation-unordered and dimension-unequal local inputs on non-Euclidean domains. In the context of the function approximation theory, the local-aggregation function is parameterized with a group of orthonormal polynomials in an effective and efficient manner. By replacing the traditional convolution in CNNs with the parameterized local-aggregation function, Local-Aggregation Graph Networks (LAGNs) are readily established, which enable to fit nonlinear functions without activation functions and can be expediently trained with the standard back-propagation. Extensive experiments on various datasets strongly demonstrate the effectiveness and efficiency of LAGNs, leading to superior performance on numerous pattern recognition and machine learning tasks, including text categorization, molecular activity detection, taxi flow prediction, and image classification. Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Baselines Extraction from Curved Document Images via Slope Fields RecoveryabstractBaselines estimation is a critical preprocessing step for many tasks of document image processing and analysis. The problem is very challenging due to arbitrarily complicated page layouts and various types of image quality degradations. This paper proposes a method based on slope fields recovery for curved baseline extraction from a distorted document image captured by a hand-held camera. Our method treats the curved baselines as the solution curves of an ordinary differential equation defined on a slope field. By assuming the page shape is a smooth and developable surface, we investigate a type of intrinsic geometric constraints of baselines to estimate the latent slope field. The curved baselines are finally obtained by solving an ordinary differential equation through the Euler method. Unlike the traditional text-lines based methods, our method is free from text-lines detection and segmentation. It can exploit multiple visual cues other than horizontal text-lines available in images for baselines extraction and is quite robust to document scripts, various types of image quality degradation (e.g., image distortion, blur and non-uniform illumination), large areas of non-textual objects and complex page layouts. Extensive experiments on synthetic and real-captured document images are implemented to evaluate the performance of the proposed method. Gaofeng Meng, Chunhong Pan, Shiming Xiang, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Efficient nearest neighbor search in high dimensional hamming space
Bin Fan 0001, Qingqun Kong, Baoqian Zhang, Hongmin Liu 0001, Chunhong Pan, Jiwen Lu |
Pattern Recognit. | 5 |
| 2020 | Geometric rectification of document images using adversarial gated unwarping network
Xiyan Liu, Gaofeng Meng, Bin Fan 0001, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 5 |
| 2020 | A generalized least-squares approach regularized with graph embedding for dimensionality reduction
Xiangjun Shen, Si-Xing Liu, Bing-Kun Bao, Chunhong Pan, Zhengjun Zha, Jianping Fan 0001 |
Pattern Recognit. | 4 |
| 2020 | 3D PostureNet: A unified framework for skeleton-based posture recognition
Ying Wang 0008, Yongcheng Liu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. Lett. | 5 |
| 2020 | Triplet Adversarial Domain Adaptation for Pixel-Level Classification of VHR Remote Sensing ImagesabstractPixel-level classification for very high resolution (VHR) images is a crucial but challenging task in remote sensing. However, since the diverse ways of satellite image acquisition and the distinct structures of various regions, the distributions of the same semantic classes among different data sets are dissimilar. Therefore, the classification model trained on one data set (source domain) may collapse, when it is directly applied to another one (target domain). To solve this problem, many adversarial-based domain adaptation methods have been proposed. However, these methods only consider the source and the target domains independently in the adversarial training, where only the target domain is explicitly contributed to narrow the gap between the distributions of both domains. Unlike previous methods, we propose a triplet adversarial domain adaptation (TriADA) method that jointly considers both domains to learn a domain-invariant classifier by a novel domain similarity discriminator. Specifically, the discriminator takes a triplet of segmentation maps as input, where two segmentation maps from the same domain are to be distinguished from the two maps from the different domains during the adversarial learning. Consequently, it explicitly considers both domains' information to narrow the distribution gap across domains. To enhance the discriminability of the classifier on the target domain, a class-aware self-training strategy, which depends on the output of the discriminator, is proposed to assign pseudo-labels with high adapted confidence on target data to retrain the classifier. Extensive experiments on several VHR pixel-level classification benchmarks demonstrate the effectiveness of our method as well as its superiority to the-state of the art. Bin Fan 0001, Hongmin Liu 0001, Chunlei Huo, Shiming Xiang, Chunhong Pan |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2019 | No-Reference Image Quality Assessment with Reinforcement Recursive List-Wise RankingabstractOpinion-unaware no-reference image quality assessment (NR-IQA) methods have received many interests recently because they do not require images with subjective scores for training. Unfortunately, it is a challenging task, and thus far no opinion-unaware methods have shown consistently better performance than the opinion-aware ones. In this paper, we propose an effective opinion-unaware NR-IQA method based on reinforcement recursive list-wise ranking. We formulate the NR-IQA as a recursive list-wise ranking problem which aims to optimize the whole quality ordering directly. During training, the recursive ranking process can be modeled as a Markov decision process (MDP). The ranking list of images can be constructed by taking a sequence of actions, and each of them refers to selecting an image for a specific position of the ranking list. Reinforcement learning is adopted to train the model parameters, in which no ground-truth quality scores or ranking lists are necessary for learning. Experimental results demonstrate the superior performance of our approach compared with existing opinion-unaware NR-IQA methods. Furthermore, our approach can compete with the most effective opinion-aware methods. It improves the state-of-the-art by over 2% on the CSIQ benchmark and outperforms most compared opinion-aware models on TID2013. Jie Gu 0002, Gaofeng Meng, Cheng Da, Shiming Xiang, Chunhong Pan |
AAAI | 5 |
| 2019 | Video Object Detection with Locally-Weighted Deformable NeighborsabstractDeep convolutional neural networks have achieved great success on various image recognition tasks. However, it is nontrivial to transfer the existing networks to video due to the fact that most of them are developed for static image. Frame-byframe processing is suboptimal because temporal information that is vital for video understanding is totally abandoned. Furthermore, frame-by-frame processing is slow and inefficient, which can hinder the practical usage. In this paper, we propose LWDN (Locally-Weighted Deformable Neighbors) for video object detection without utilizing time-consuming optical flow extraction networks. LWDN can latently align the high-level features between keyframes and keyframes or nonkeyframes. Inspired by (Zhu et al. 2017a) and (Hetang et al. 2017) who propose to aggregate features between keyframes and keyframes, we adopt brain-inspired memory mechanism to propagate and update the memory feature from keyframes to keyframes. We call this process Memory-Guided Propagation. With such a memory mechanism, the discriminative ability of features in keyframes and non-keyframes are both enhanced, which helps to improve the detection accuracy. Extensive experiments on VID dataset demonstrate that our method achieves superior performance in a speed and accuracy trade-off, i.e., 76.3% on the challenging VID dataset while maintaining 20fps in speed on Titan X GPU. Zhengkai Jiang 0001, Peng Gao 0007, Chaoxu Guo, Qian Zhang 0009, Shiming Xiang, Chunhong Pan |
AAAI | 6 |
| 2019 | What and Where the Themes Dominate in ImageabstractThe image captioning is to describe an image with natural language as human, which has benefited from the advances in deep neural network and achieved substantial progress in performance. However, the perspective of human description to scene has not been fully considered in this task recently. Actually, the human description to scene is tightly related to the endogenous knowledge and the exogenous salient objects simultaneously, which implies that the content in the description is confined to the known salient objects. Inspired by this observation, this paper proposes a novel framework, which explicitly applies the known salient objects in image captioning. Under this framework, the known salient objects are served as the themes to guide the description generation. According to the property of the known salient object, a theme is composed of two components: its endogenous concept (what) and the exogenous spatial attention feature (where). Specifically, the prediction of each word is dominated by the concept and spatial attention feature of the corresponding theme in the process of caption prediction. Moreover, we introduce a novel learning method of Distinctive Learning (DL) to get more specificity of generated captions like human descriptions. It formulates two constraints in the theme learning process to encourage distinctiveness between different images. Particularly, reinforcement learning is introduced into the framework to address the exposure bias problem between the training and the testing modes. Extensive experiments on the COCO and Flickr30K datasets achieve superior results when compared with the state-of-the-art methods. Xinyu Xiao, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
AAAI | 4 |
| 2019 | Relation-Shape Convolutional Neural Network for Point Cloud AnalysisabstractPoint cloud analysis is very challenging, as the shape implied in irregular points is difficult to capture. In this paper, we propose RS-CNN, namely, Relation-Shape Convolutional Neural Network, which extends regular grid CNN to irregular configuration for point cloud analysis. The key to RS-CNN is learning from relation, i.e., the geometric topology constraint among points. Specifically, the convolutional weight for local point set is forced to learn a high-level relation expression from predefined geometric priors, between a sampled point from this point set and the others. In this way, an inductive local representation with explicit reasoning about the spatial layout of points can be obtained, which leads to much shape awareness and robustness. With this convolution as a basic operator, RS-CNN, a hierarchical architecture can be developed to achieve contextual shape-aware learning for point cloud analysis. Extensive experiments on challenging benchmarks across three tasks verify RS-CNN achieves the state of the arts. Yongcheng Liu, Bin Fan 0001, Shiming Xiang, Chunhong Pan |
CVPR | 4 |
| 2019 | Guiding the Flowing of Semantics: Interpretable Video Captioning via POS TagabstractXinyu Xiao, Lingfeng Wang, Bin Fan, Shinming Xiang, Chunhong Pan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xinyu Xiao, Lingfeng Wang 0002, Bin Fan 0001, Shiming Xiang, Chunhong Pan |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Adaptive Brightness Learning for Active Object RecognitionabstractState-of-the-art object detection methods based on deep learning achieved promising performances in recent years. However, the performances are limited by the passive nature of the traditional object recognition framework in ignoring the relationship between imaging configuration and recognition performance as well as the importance of recognition performance feedback for improving image quality. To address the above limitations, an active object recognition method based on reinforcement learning is proposed in this paper by taking adaptive brightness adjustment as an example. Progressive brightness adjustment strategy is learned by maximizing recognition performance on reference high-quality training samples. With the help of active object recognition and brightness adjustment strategy, low-quality images can be converted into high-quality images, and overall performances are improved without retraining the detector. Nuo Xu 0006, Chunlei Huo, Chunhong Pan |
ICASSP | 3 |
| 2019 | Progressive Sparse Local Attention for Video Object DetectionabstractTransferring image-based object detectors to the domain of videos remains a challenging problem. Previous efforts mostly exploit optical flow to propagate features across frames, aiming to achieve a good trade-off between accuracy and efficiency. However, introducing an extra model to estimate optical flow can significantly increase the overall model size. The gap between optical flow and high-level features can also hinder it from establishing spatial correspondence accurately. Instead of relying on optical flow, this paper proposes a novel module called Progressive Sparse Local Attention (PSLA), which establishes the spatial correspondence between features across frames in a local region with progressively sparser stride and uses the correspondence to propagate features. Based on PSLA, Recursive Feature Updating (RFU) and Dense Feature Transforming (DenseFT) are proposed to model temporal appearance and enrich feature representation respectively in a novel video object detection framework. Experiments on ImageNet VID show that our method achieves the best accuracy compared to existing methods with smaller model size and acceptable runtime speed. Chaoxu Guo, Bin Fan 0001, Jie Gu 0002, Qian Zhang 0009, Shiming Xiang, Véronique Prinet, Chunhong Pan |
ICCV | 7 |
| 2019 | DensePoint: Learning Densely Contextual Representation for Efficient Point Cloud ProcessingabstractPoint cloud processing is very challenging, as the diverse shapes formed by irregular points are often indistinguishable. A thorough grasp of the elusive shape requires sufficiently contextual semantic information, yet few works devote to this. Here we propose DensePoint, a general architecture to learn densely contextual representation for point cloud processing. Technically, it extends regular grid CNN to irregular point configuration by generalizing a convolution operator, which holds the permutation invariance of points, and achieves efficient inductive learning of local patterns. Architecturally, it finds inspiration from dense connection mode, to repeatedly aggregate multi-level and multi-scale semantics in a deep hierarchy. As a result, densely contextual information along with rich semantics, can be acquired by DensePoint in an organic manner, making it highly effective. Extensive experiments on challenging benchmarks across four tasks, as well as thorough model analysis, verify DensePoint achieves the state of the arts. Yongcheng Liu, Bin Fan 0001, Gaofeng Meng, Jiwen Lu, Shiming Xiang, Chunhong Pan |
ICCV | 6 |
| 2019 | Rotation and Scale-Invariant Object Detector for High Resolution Optical Remote Sensing ImagesabstractObject detection of high-resolution optical remote sensing images is challenging due to two fundamental problems. One is the huge scale variation of objects in images, e.g., small vehicle and cross-sea bridge. The other one is the objects could take on arbitrary orientations because of the high angle shot. In this paper, we propose a Rotation and Scale-invariant Detector (RS-Det) for remote sensing images to solve the above problem in an unified network. Specifically, RS-Det consists of a deformable convolution module to learn spatial transformation (such as rotation, transition, etc) and a feature pyramid architecture for multi-scale feature representation. These two modules enable a better feature learning of convolutional neural network and boost the performance by 3.6% compared with the baseline method. In DOTA, a large-scale dataset for aerial image object detection, our RS-Det achieves the state-of-the-art accuracy, which verifies our method's superiority. Chunlei Huo, Feilong Wei, Chunhong Pan |
IGARSS | 4 |
| 2019 | GSTNet: Global Spatial-Temporal Network for Traffic Flow PredictionabstractPredicting traffic flow on traffic networks is a very challenging task, due to the complicated and dynamic spatial-temporal dependencies between different nodes on the network. The traffic flow renders two types of temporal dependencies, including short-term neighboring and long-term periodic dependencies. What's more, the spatial correlations over different nodes are both local and non-local. To capture the global dynamic spatial-temporal correlations, we propose a Global Spatial-Temporal Network (GSTNet), which consists of several layers of spatial-temporal blocks. Each block contains a multi-resolution temporal module and a global correlated spatial module in sequence, which can simultaneously extract the dynamic temporal dependencies and the global spatial correlations. Extensive experiments on the real world datasets verify the effectiveness and superiority of the proposed method on both the public transportation network and the road network. Shen Fang, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
IJCAI | 5 |
| 2019 | DATA: Differentiable ArchiTecture ApproximationabstractNeural architecture search (NAS) is inherently subject to the gap of architectures during searching and validating. To bridge this gap, we develop Differentiable ArchiTecture Approximation (DATA) with an Ensemble Gumbel-Softmax (EGS) estimator to automatically approximate architectures during searching and validating in a differentiable manner. Technically, the EGS estimator consists of a group of Gumbel-Softmax estimators, which is capable of converting probability vectors to binary codes and passing gradients from binary codes to probability vectors. Benefiting from such modeling, in searching, architecture parameters and network weights in the NAS model can be jointly optimized with the standard back-propagation, yielding an end-to-end learning mechanism for searching deep models in a large enough search space. Conclusively, during validating, a high-performance architecture that approaches to the learned one during searching is readily built. Extensive experiments on a variety of popular datasets strongly evidence that our method is capable of discovering high-performance architectures for image classification, language modeling and semantic segmentation, while guaranteeing the requisite efficiency during searching. Jianlong Chang, Xinbang Zhang, Yiwen Guo, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
NeurIPS | 6 |
| 2019 | Incremental Poisson Surface Reconstruction for Large Scale Three-Dimensional Modeling
Wei Sui, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
PRCV (3) | 5 |
| 2019 | Scene text detection and recognition with advances in deep learning: a survey
Xiyan Liu, Gaofeng Meng, Chunhong Pan |
Int. J. Document Anal. Recognit. | 3 |
| 2019 | Nonlinear Asymmetric Multi-Valued HashingabstractMost existing hashing methods resort to binary codes for large scale similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose Nonlinear Asymmetric Multi-Valued Hashing (NAMVH) supported by two distinct non-binary embeddings. Specifically, a real-valued embedding is used for representing the newly-coming query by an ideally nonlinear transformation. Besides, a multi-integer-embedding is employed for compressing the whole database, which is modeled by Binary Sparse Representation (BSR) with fixed sparsity. With these two non-binary embeddings, NAMVH preserves more precise similarities between data points and enables access to the incremental extension with database samples evolving dynamically. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the pairwise label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by a well-designed alternative optimization method. Extensive experiments on seven large scale datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency. Cheng Da, Gaofeng Meng, Shiming Xiang, Kun Ding 0001, Shibiao Xu, Qing Yang 0002, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2019 | Blind image quality assessment via learnable attention-based pooling
Jie Gu 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 4 |
| 2019 | Visual object tracking via a manifold regularized discriminative dual dictionary model
Lingfeng Wang 0002, Chunhong Pan |
Pattern Recognit. | 2 |
| 2019 | Dense semantic embedding network for image captioning
Xinyu Xiao, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 5 |
| 2019 | Pseudo low rank video representation
Tingzhao Yu, Lingfeng Wang 0002, Chaoxu Guo, Huxiang Gu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 6 |
| 2019 | Learning graph structure via graph convolutional networks
Jianlong Chang, Gaofeng Meng, Shibiao Xu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 6 |
| 2019 | A Performance Evaluation of Local Features for Image-Based 3D ReconstructionabstractThis paper performs a comprehensive and comparative evaluation of the state-of-the-art local features for the task of image-based 3D reconstruction. The evaluated local features cover the recently developed ones by using powerful machine learning techniques and the elaborately designed handcrafted features. To obtain a comprehensive evaluation, we choose to include both float type features and binary ones. Meanwhile, two kinds of datasets have been used in this evaluation. One is a dataset of many different scene types with groundtruth 3D points, containing images of different scenes captured at fixed positions, for quantitative performance evaluation of different local features in the controlled image capturing situation. The other dataset contains Internet scale image sets of several landmarks with a lot of unrelated images, which is used for qualitative performance evaluation of different local features in the free image collection situation. Our experimental results show that binary features are competent to reconstruct scenes from controlled image sequences with only a fraction of processing time compared to using float type features. However, for the case of a large scale image set with many distracting images, float type features show a clear advantage over binary ones. Currently, the most traditional SIFT is very stable with regard to scene types in this specific task and produces very competitive reconstruction results among all the evaluated local features. Meanwhile, although the learned binary features are not as competitive as the handcrafted ones, learning float type features with CNN is promising but still requires much effort in the future. Bin Fan 0001, Qingqun Kong, Xinchao Wang, Zhiheng Wang 0001, Shiming Xiang, Chunhong Pan, Pascal Fua |
IEEE Trans. Image Process. | 6 |
| 2019 | Deep Hierarchical Encoder-Decoder Network for Image CaptioningabstractEncoder-decoder models have been widely used in image captioning, and most of them are designed via single long short term memory (LSTM). The capacity of single-layer network, whose encoder and decoder are integrated together, is limited for such a complex task of image captioning. Moreover, how to effectively increase the “vertical depth” of encoder-decoder remains to be solved. To deal with these problems, a novel deep hierarchical encoder-decoder network is proposed for image captioning, where a deep hierarchical structure is explored to separate the functions of encoder and decoder. This model is capable of efficiently exerting the representation capacity of deep networks to fuse high level semantics of vision and language in generating captions. Specifically, visual representations in top levels of abstraction are simultaneously considered, and each of these levels is associated to one LSTM. The bottom-most LSTM is applied as the encoder of textual inputs. The application of the middle layer in encoder-decoder is to enhance the decoding ability of top-most LSTM. Furthermore, depending on the introduction of semantic enhancement module of image feature and distribution combine module of text feature, variants of architectures of our model are constructed to explore the impacts and mutual interactions among the visual representation, textual representations, and the output of the middle LSTM layer. Particularly, the framework is training under a reinforcement learning method to address the exposure bias problem between the training and the testing by the policy gradient optimization. Qualitative analyses indicate the process that our model “translates” image to sentence and further visualization presents the evolution of the hidden states from different hierarchical LSTMs over time. Extensive experiments demonstrate that our model outperforms current state-of-the-art models on three benchmark datasets: Flickr8K, Flickr30K, and MSCOCO. On both image captioning and retrieval tasks, our method achieves the best results. On MSCOCO captioning Leaderboard, our method also achieves superior performance. Xinyu Xiao, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 5 |
| 2019 | Weakly Semantic Guided Action RecognitionabstractAction recognition plays a fundamental role in computer vision and video analysis. Nevertheless, extracting effective spatial-temporal features remains a challenging task. This paper proposes three simple but effective weakly semantic guided modules (SGMs) for both environment-constrained and cross-domain action recognition. The SGMs are composed of total 3-D convolution and element-wise gated operations; thus, they are efficient and easy to implement. The semantic guidance is obtained in a weakly supervised manner, in which each video clip is labeled with only an action class instead of pixel-level semantics. Benefitting from the semantic guidance, the network [called semantic guided network (SGN)] can focus on the salient parts of the video clips. Consequently, the redundant information can be reduced and the model is more robust to noise. Besides, benefitting from the intrinsic property of SGMs, SGN is totally end-to-end trainable. Quantities of experiments on both environment-constrained (e.g., Penn, HMDB-51, and UCF101) and cross-domain (e.g., ODAR) action recognition datasets demonstrate its effectiveness. Specifically, SGN gets improvements of 3.7%, 2.1%, and 5.2% for Penn, HMDB-51, and UCF-101 than the baseline ResNet3D, respectively, and SGN ranked third place in the ODAR 2017 challenge. Tingzhao Yu, Lingfeng Wang 0002, Cheng Da, Huxiang Gu, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 6 |
| 2018 | Exploiting Vector Fields for Geometric Rectification of Distorted Document Images
Gaofeng Meng, Yuanqi Su, Ying Wu 0001, Shiming Xiang, Chunhong Pan |
ECCV (16) | 5 |
| 2018 | Mgn: Multi-Glimpse Network for Action RecognitionabstractCurrent state-of-the-art action recognition approaches rely on optical flow to extract the local motion information and ignore the importance of global description of the videos. In this paper, we present a novel architecture, named Multi-Glimpse Network (MGN), to boost the performance of action recognition by combining the local and global information of the videos. Specifically, MGN makes predictions through two important modules, Local Glimpse and Global Glimpse. Local Glimpse extracts the local spatiotemporal features of different periods using temporal sampling method. Global Glimpse aggregates the extracted local features to develop global description of the videos. These two modules are complementary and indispensable. Our MGN achieves competitive results on four video action benchmarks of UCF10l, HMDB51, ODAR and Penn. Chaoxu Guo, Tingzhao Yu, Huxiang Gu, Shiming Xiang, Chunhong Pan |
ICASSP | 5 |
| 2018 | Fast Variational Level Set Based Image Segmentation via Two-Scale Filtering ModelabstractOne major difficulty in medical image segmentation is intensity inhomogeneity, which manifests itself with a slow intensity variation over the whole image domain. Recently, a local binary fitting (LBF) model has been proposed to solve this problem within level set segmentation framework. However, the LBF model has two main problems, i.e., high computational cost and sensitivity to initialization. By analyzing the LBF model, we find that the most computational part is the calculation of two cluster images, which need to be updated in each iteration during the evolution of level set function. With this observation in mind, we propose a novel two-scale filtering (TSF) model, in which the two cluster images can be pre-calculated before evolution. Additionally, we implicitly utilize order constraint to restrict the order of two cluster images. As a result, the proposed TSF model is less sensitive to initialization. Extensive experiments on real medical images illustrate the desirable performances, as compared with the state-of-the-art models. Lingfeng Wang 0002, Ying Wang 0008, Chunhong Pan |
ICASSP | 3 |
| 2018 | Two-Stream Designed 2D/3D Residual Networks with Lstms for Action Recognition in VideosabstractConvolutional Neural Networks(CNNs) have achieved great success for object recognition in still images. However, CNNs can't make evident improvement for action recognition in videos, one reason is that many current network architectures are relatively shallow compared with deep models in image domain, and the other reason is that CNNs can't capture effective long-term motion information from videos. Encouraged by the good performance of Residual Network-s(ResNets) for training extremely deep models, and Long-term Recurrent Convolutional Networks(LSTMs) for dealing with tasks involving sequences, we presented an action recognition method based on a two-stream architecture, with 2D ResNets with LSTMs in one stream and designed 3D ResNets with LSTMs in the other stream, which can combine appearance and motion information better. Especially, our proposed method first learns spatiotemporal features of videos through the Residual networks, then models complex temporal dynamics by the Long-term Recurrent Convolutional networks, and with a softmax layer on the top of two streams, the final classification results can be predicted by fusing scores of each stream with weights on score distribution. Furthermore, for better reducing the influence of redundant background information in videos for recognition results, we also applied a center extraction method to generate central regions of videos instead of an entire video into a visual representation. On two video action benchmarks of UCF101 and HMDB51, our method achieved promising performance compared with state-of-the-art. Lifei Song, Liguo Weng, Lingfeng Wang 0002, Min Xia 0002, Chunhong Pan |
ICIP | 5 |
| 2018 | Reconstructed Densenets for Image Super-ResolutionabstractDeep learning has been successfully applied to single image super-resolution problem due to its high data fitting ability. However, the trending of deeper layers and wider receptive field to acquire better performance brings high computation complexity and serious information vanishing. To address this problem, we proposed a new Reconstructed DenseNets model for super-resolution. The basic idea behind Reconstructed DenseNets is to improve the recent DenseNets model by modifying the two core modules, dense blocks and transition blocks, so that the Reconstructed DenseNets can emphasize the quality of data reconstruction. Specifically, on the one hand, the batch normalization layers in dense blocks is ignored to overcome the data shift risk. One the other hand, the pooling layers in transition blocks is also ignored to ensure the ability to reconstruct. Based on the above two improvements, the new DenseNets is named as Reconstructed DenseNets. Extensive experiments evaluate the effectiveness of our model, showing the outperforming of the state-of-the-art approaches. Lingfeng Wang 0002, Linwei Qiu, Wei Sui, Chunhong Pan |
ICIP | 4 |
| 2018 | Adversarial Domain Adaptation with a Domain Similarity Discriminator for Semantic Segmentation of Urban AreasabstractExisting semantic segmentation models of urban areas have shown to perform well in a supervised setting. However, collecting lots of annotated images from each city to train such models is time-consuming or difficult. In addition, when transferring the segmentation model from the trained city (source domain) to an unseen city (target domain), the performance will largely degrade due to the domain shift. For this reason, we propose a domain adaptation method with a domain similarity discriminator to eliminate such domain shift in the framework of adversarial learning. Contrary to the single-input adversarial network, our domain similarity discriminator, which consists of a Siamese network, is able to measure the similarity of the pairwise-input data. In this way, we can use more information about the pairwise-input to measure the similarity between different distributions so as to address the problem of domain shift. Experimental results demonstrate that our approach outperforms the competing methods on three different cities. Bin Fan 0001, Shiming Xiang, Chunhong Pan |
ICIP | 4 |
| 2018 | Semantic Image Synthesis via Conditional Cycle-Generative Adversarial NetworksabstractTraditional approaches for semantic image synthesis mainly focus on text descriptions while ignoring the related structures and attributes in the original images. Therefore, some critical information, e.g., the style, backgrounds, objects shapes and pose, is missed in the generated images. In this paper, we propose a novel framework called Conditional Cycle-Generative Adversarial Network (CCGAN) to address this issue. Our model can generate photo-realistic images conditioned on the given text descriptions, while maintaining the attributes of the original images. The framework mainly consists of two coupled conditional adversarial networks, which are able to learn a desirable image mapping that can keep the structures and attributes in the images. We introduce a conditional cycle consistency loss to prevent the contradiction between two generators. This loss allows the generated images to retain most of the features of the original image, so as to improve the stability of network training. Moreover, benefiting from the mechanism of circular training, the proposed networks can learn the semantic information of the text much accurately. Experiments on Caltech-UCSD Bird dataset and Oxford-102 flower dataset demonstrate that the proposed method significantly outperforms the existing methods in terms of image details reconstruction and semantic information expression. Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICPR | 4 |
| 2018 | Enhancing Pix2Pix for Remote Sensing Image ClassificationabstractRemote sensing image classification is challenging due to low separation between different classes and difficulty in learning discriminative features. GAN (Generative Adversarial Model) is promising for this task due to the generator in reproducing samples and the discriminator for improving the generator. Among GANs variants for image translation and image classification tasks, Pix2Pix performs best. However, Pix2Pix is limited in explicitly capturing the relationship between the source domain and the reconstructed ones from the target domain. To address the above problem, an improved Pix2Pix is proposed in this paper, where a controller is added to Pix2Pix whose role is to improve classification performance and enhance training stability. Experiments demonstrate the effectiveness and advantages of the proposed approach. Hongping Yan, Chunlei Huo, Jiayuan Yu, Chunhong Pan |
ICPR | 5 |
| 2018 | Kernel-Weighted Graph Convolutional Network: A Deep Learning Approach for Traffic ForecastingabstractTraffic forecasting is of great significance and has many applications in Intelligent Traffic System (ITS). In spite of many thoughtful attempts in the past decades, this task still remains far from being solved, due to the diversity, complexity and nonlinearity of traffic situations. Technically, it can be cast on the framework of regressions with spatial-template data. Typically, one may consider to employ the Convolutional Neural Network (CNN) to achieve this goal. Unfortunately, the traditional CNN is developed for grid data. By contrast, here we are facing with non-grid traffic data points that are observed spatially at locations of interest. To this end, this paper proposes a novel Kernel-Weighted Graph Convolutional Network (KW-GCN) for traffic forecasting, which learns simultaneously a group of convolutional kernels and their linear combination weights for each of the nodes in the graph. This yields a mechanism that is able to learn the features locally and exploit the structure information of traffic road-network globally. By introducing additional parameters, our KW-GCN can relax the restriction of weight sharing in classical CNN to better handle the traffic data of non-stationarity. Furthermore, it has been illustrated that the proposed linear weighting of kernels can be viewed as the low-rank decomposition of the well-known locally-connected networks, and thus it avoids over-fitting to some degree. We apply our approach to the real-world GPS data set of about 30,000 taxis in seven months in Beijing. Experiments on both taxi-flow forecasting and road-speed forecasting demonstrate that our method significantly outperforms the state-of-the-art ones. Qizhao Jin, Jianlong Chang, Shiming Xiang, Chunhong Pan |
ICPR | 5 |
| 2018 | Learning Deep Relationship for Image Change DetectionabstractVery high resolution image change detection is difficult due to the low interclass variability and the resulting high overlap between the changed class and the unchanged class. To address the above difficulties, the concept “relationship” is proposed to represent the intraclass similarity and interclass difference, which is established by interclass couples and intraclass couples. By relationship representation and relationship learning, intraclass couples can be compressed into a compact cluster, and the distances between interclass couples are enlarged. To better discover the complex relationship hidden in change features, relationship learning is integrated into a deep learning framework, where the relationship is learned progressively. In consequence, the final change detection performance is improved with the reduced overlap between the changed class and the unchanged class. Experiments demonstrate the effectiveness of the proposed approach. Chunlei Huo, Yushuang Zhang, Jiayuan Yu, Yunpeng Ling, Chunhong Pan |
IGARSS | 5 |
| 2018 | SafeNet: Scale-normalization and Anchor-based Feature Extraction Network for Person Re-identificationabstractPerson Re-identification (ReID) is a challenging retrieval task that requires matching a person's image across non-overlapping camera views. The quality of fulfilling this task is largely determined on the robustness of the features that are used to describe the person. In this paper, we show the advantage of jointly utilizing multi-scale abstract information to learn powerful features over full body and parts. A scale normalization module is proposed to balance different scales through residual-based integration. To exploit the information hidden in non-rigid body parts, we propose an anchor-based method to capture the local contents by stacking convolutions of kernels with various aspect ratios, which focus on different spatial distributions. Finally, a well-defined framework is constructed for simultaneously learning the representations of both full body and parts. Extensive experiments conducted on current challenging large-scale person ReID datasets, including Market1501, CUHK03 and DukeMTMC, demonstrate that our proposed method achieves the state-of-the-art results. Kun Yuan 0003, Qian Zhang 0009, Chang Huang, Shiming Xiang, Chunhong Pan |
IJCAI | 5 |
| 2018 | Multi-Label Image Classification via Knowledge Distillation from Weakly-Supervised DetectionabstractMulti-label image classification is a fundamental but challenging task towards general visual understanding. Existing methods found the region-level cues (e.g., features from RoIs) can facilitate multi-label classification. Nevertheless, such methods usually require laborious object-level annotations (i.e., object labels and bounding boxes) for effective learning of the object-level visual features. In this paper, we propose a novel and efficient deep framework to boost multi-label classification by distilling knowledge from weakly-supervised detection task without bounding box annotations. Specifically, given the image-level annotations, (1) we first develop a weakly-supervised detection (WSD) model, and then (2) construct an end-to-end multi-label image classification framework augmented by a knowledge distillation module that guides the classification model by the WSD model according to the class-level predictions for the whole image and the object-level visual features for object RoIs. The WSD model is the teacher model and the classification model is the student model. After this cross-task knowledge distillation, the performance of the classification model is significantly improved and the efficiency is maintained since the WSD model can be safely discarded in the test phase. Extensive experiments on two large-scale datasets (MS-COCO and NUS-WIDE) show that our framework achieves superior performances over the state-of-the-art methods on both performance and efficiency. Yongcheng Liu, Lu Sheng, Shiming Xiang, Chunhong Pan |
ACM Multimedia | 6 |
| 2018 | Structure-Aware Convolutional Neural NetworksabstractConvolutional neural networks (CNNs) are inherently subject to invariable filters that can only aggregate local inputs with the same topological structures. It causes that CNNs are allowed to manage data with Euclidean or grid-like structures (e.g., images), not ones with non-Euclidean or graph structures (e.g., traffic networks). To broaden the reach of CNNs, we develop structure-aware convolution to eliminate the invariance, yielding a unified mechanism of dealing with both Euclidean and non-Euclidean structured data. Technically, filters in the structure-aware convolution are generalized to univariate functions, which are capable of aggregating local inputs with diverse topological structures. Since infinite parameters are required to determine a univariate function, we parameterize these filters with numbered learnable parameters in the context of the function approximation theory. By replacing the classical convolution in CNNs with the structure-aware convolution, Structure-Aware Convolutional Neural Networks (SACNNs) are readily established. Extensive experiments on eleven datasets strongly evidence that SACNNs outperform current models on various machine learning tasks, including image classification and clustering, text categorization, skeleton-based action recognition, molecular activity detection, and taxi flow prediction. Jianlong Chang, Jie Gu 0002, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
NeurIPS | 6 |
| 2018 | Facade repetition detection in a fronto-parallel view with fiducial lines extraction
Hongfei Xiao, Gaofeng Meng, Lingfeng Wang 0002, Chunhong Pan |
Neurocomputing | 4 |
| 2018 | Deep unsupervised learning with consistent inference of latent representations
Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 5 |
| 2018 | Joint spatial-temporal attention for action recognition
Tingzhao Yu, Chaoxu Guo, Lingfeng Wang 0002, Huxiang Gu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. Lett. | 6 |
| 2018 | Deep generative video prediction
Tingzhao Yu, Lingfeng Wang 0002, Huxiang Gu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. Lett. | 5 |
| 2018 | Self-Paced AutoEncoderabstractAutoencoder, which learns latent representations of samples in an unsupervised manner, has great potential in computer vision and signal processing. However, the diversity of samples makes learning a component autoencoder remaining a challenging task. This letter proposes a novel Self-Paced AutoEncoder (SPAE) for unsupervised feature extraction. The motivation behind this letter is to take samples gradually from simple to complex into consideration during training, which is similar to the mechanism of knowledge acquisition for humans. Under the unsupervised learning framework constructed on the autoencoder infrastructure, our SPAE first learns a weak autoencoder via samples with small losses and, then, elevates itself to a relatively strong autoencoder through samples with large losses. Then, the SPAE is generalized to a temporal domain, resulting to temporal SPAE (TSPAE), where the temporal information is explored and exploited to improve the performance. Typically, a TSPAE is capable of compressing temporal sequences into temporal-independent data. Experiments on the image classification and action recognition demonstrate the effectiveness of SPAE and TSPAE. Tingzhao Yu, Chaoxu Guo, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
IEEE Signal Process. Lett. | 5 |
| 2018 | Blind Image Quality Assessment via Vector Regression and Object Oriented PoolingabstractThis paper presents an effective method based on vector regression and object oriented pooling for blind image quality assessment. Unlike previous models that map the extracted features directly to a quality score, the proposed vector regression framework yields a vector of belief scores for the input image. We explore the uncertainty factors in quality assessment and design the belief scores to measure the confidences of an image to be assigned to the corresponding quality grades. Moreover, we propose an object oriented pooling strategy to further improve the performance by incorporating semantic information of image contents. According to this strategy, regions occupied by objects will be assigned more weights in the pooling phase, leading to a more accurate quality assessment. Extensive experiments on benchmark datasets demonstrate that our approach achieves state-of-the-art performance and shows a great generalization ability. Jie Gu 0002, Gaofeng Meng, Judith Redi, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 5 |
| 2018 | In Defense of Locality-Sensitive HashingabstractHashing-based semantic similarity search is becoming increasingly important for building large-scale content-based retrieval system. The state-of-the-art supervised hashing techniques use flexible two-step strategy to learn hash functions. The first step learns binary codes for training data by solving binary optimization problems with millions of variables, thus usually requiring intensive computations. Despite simplicity and efficiency, locality-sensitive hashing (LSH) has never been recognized as a good way to generate such codes due to its poor performance in traditional approximate neighbor search. We claim in this paper that the true merit of LSH lies in transforming the semantic labels to obtain the binary codes, resulting in an effective and efficient two-step hashing framework. Specifically, we developed the locality-sensitive two-step hashing (LS-TSH) that generates the binary codes through LSH rather than any complex optimization technique. Theoretically, with proper assumption, LS-TSH is actually a useful LSH scheme, so that it preserves the label-based semantic similarity and possesses sublinear query complexity for hash lookup. Experimentally, LS-TSH could obtain comparable retrieval accuracy with state of the arts with two to three orders of magnitudes faster training speed. Kun Ding 0001, Chunlei Huo, Bin Fan 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | Groupwise Retargeted Least-Squares RegressionabstractIn this brief, we propose a new groupwise retargeted least squares regression (GReLSR) model for multicategory classification. The main motivation behind GReLSR is to utilize an additional regularization to restrict the translation values of ReLSR, so that they should be similar within same class. By analyzing the regression targets of ReLSR, we propose a new formulation of ReLSR, where the translation values are expressed explicitly. On the basis of the new formulation, discriminative least-squares regression can be regarded as a special case of ReLSR with zero translation values. Moreover, a groupwise constraint is added to ReLSR to form the new GReLSR model. Extensive experiments on various machine leaning data sets illustrate that our method outperforms the current state-of-the-art approaches. Lingfeng Wang 0002, Chunhong Pan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | AMVH: Asymmetric Multi-Valued hashingabstractMost existing hashing methods resort to binary codes for similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose an asymmetric multi-valued hashing method supported by two different non-binary embeddings. (1) A real-valued embedding is used for representing the newly-coming query. (2) A multi-integer-embedding is employed for compressing the whole database, which is modeled by binary sparse representation with fixed sparsity. With these two non-binary embeddings, the similarities between data points can be preserved precisely. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by alternative optimization. Extensive experiments on three multilabel datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency. Cheng Da, Shibiao Xu, Kun Ding 0001, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
CVPR | 6 |
| 2017 | Learning deep vector regression model for no-reference image quality assessmentabstractThe goal of no-reference image quality assessment (NR-IQA) is to estimate human perceived image quality without access to either reference image or prior knowledge about distortion type. Previous approaches for this problem are typically based on a regression framework that maps the image features directly to a quality score. In contrast, psychological evidence shows that humans prefer to evaluate visual quality with qualitative descriptions, e.g., using a five-grade ordinal scale: “excellent”, “good”, “fair”, “poor” and “bad”. Based on this observation, we propose a vector regression model that predicts five belief scores rather than a single quality score. The belief scores are designed to indicate the confidences of the test image being assigned with these five quality grades. In addition, with the purpose of more extensive applications, a saliency-based pooling strategy is presented to convert the predicted confidences into objective quality scores. Extensive experiments performed on two benchmark datasets demonstrate that our approach achieves state-of-the-art performance and shows great generalization ability. Jie Gu 0002, Gaofeng Meng, Lingfeng Wang 0002, Chunhong Pan |
ICASSP | 4 |
| 2017 | RoDLSR: Robust discriminative least squares regression model for multi-category classificationabstractDiscriminative least squares regression (DLSR) is a simple yet effective method for multi-class classification. One problem of DLSR is that it is lack of robustness to outliers. In order to tackle this difficulty, in this paper, we propose a novel Robust DLSR (RoDLSR) model. The core idea behind RoDLSR is to find and further ignore the outliers among the support vector set. Specifically, we modify the regression targets of outliers by adding an additional item. As a result, the range of regression residuals can be controlled within predefined threshold. Extensive experiments evaluate the effectiveness of RoDLSR, especially on the corrupted databases. Lingfeng Wang 0002, Shuaizheng Liu, Chunhong Pan |
ICASSP | 3 |
| 2017 | Deep Adaptive Image ClusteringabstractImage clustering is a crucial but challenging task in machine learning and computer vision. Existing methods often ignore the combination between feature learning and clustering. To tackle this problem, we propose Deep Adaptive Clustering (DAC) that recasts the clustering problem into a binary pairwise-classification framework to judge whether pairs of images belong to the same clusters. In DAC, the similarities are calculated as the cosine distance between label features of images which are generated by a deep convolutional network (ConvNet). By introducing a constraint into DAC, the learned label features tend to be one-hot vectors that can be utilized for clustering images. The main challenge is that the ground-truth similarities are unknown in image clustering. We handle this issue by presenting an alternating iterative Adaptive Learning algorithm where each iteration alternately selects labeled samples and trains the ConvNet. Conclusively, images are automatically clustered based on the label features. Experimental results show that DAC achieves state-of-the-art performance on five popular datasets, e.g., yielding 97.75% clustering accuracy on MNIST, 52.18% on CIFAR-10 and 46.99% on STL-10. Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICCV | 5 |
| 2017 | Deep Networks for Degraded Document Image Binarization through Pyramid ReconstructionabstractBinarization of document images is an important processing step for document images analysis and recognition. However, this problem is quite challenging in some cases because of the quality degradation of document images, such as varying illumination, complicated backgrounds, image noises due to ink spots, water stains or document creases. In this paper, we propose a framework based on deep convolutional neural-network (DCNN) for adaptive binarization of degraded document images. The basic idea of our method is to decompose a degraded document image into a spatial pyramid structure by using DCNN, with each layer at different scale. Then the foreground image is sequentially reconstructed from these layers in a coarse-to-fine manner by using deconvolutional network. Such kind of decomposition is quite beneficial, since multi-resolution supervision information can be directly introduced into network learning. We also define several loss functions about label consistency and foregrounds smoothing to further regularize the training of the network. Experimental results demonstrate the effectiveness of the proposed method. Gaofeng Meng, Kun Yuan 0003, Ying Wu 0001, Shiming Xiang, Chunhong Pan |
ICDAR | 5 |
| 2017 | Learnable contextual regularization for semantic segmentation of indoor scene imagesabstractSemantic segmentation of indoor scene images has a wide range of applications. However, due to a large number of classes and uneven distribution in indoor scenes, mislabels are often made when facing small objects or boundary regions. Technically, contextual information may benefit for segmentation results, but has not yet been exploited sufficiently. In this paper, we propose a learnable contextual regularization model for enhancing the semantic segmentation results of color indoor scene images. This regularization model is combined with a deep convolutional segmentation network without significantly increasing the number of additional parameters. Our model, derived from the inherent contextual regularization on the indoor scene objects, benefits much from the learnable constraint layers bridging the lower layers and the higher layers in the deep convolutional network. The constraint layers are further integrated with a weighted L1-norm based contextual regularization between the neighboring pixels of RGB values to improve the segmentation results. Experimental results on NYUDv2 indoor scene dataset demonstrate the effectiveness and efficiency of the proposed method. Gaofeng Meng, Lingfeng Wang 0002, Chunhong Pan |
ICIP | 5 |
| 2017 | Efficient similarity learning for asymmetric hashingabstractHashing techniques with asymmetric schemes (e.g., only bi-narizing the database points) have recently attracted wide attention in the circle of image retrieval. In comparison with those methods which binarize simultaneously both of the query and database points, they not only enjoy the storage and search efficiencies, but also provide higher accuracy. Gearing to this line, this paper proposes a metric-embedded asymmetric hashing (MEAH) that learns jointly a bilinear similarity measure and binary codes of database points in an unsupervised manner. Technically, the learned similarity measure is able to bridge the gap between the binary codes and the real-valued codes, which are represented possibly with different dimensions. What is more, this measure is capable of preserving the global structure hidden in the database. Extensive experiments on two public image benchmarks demonstrate the superiority of our approach over the several state-of-the-art unsupervised hashing methods. Cheng Da, Yang Yang 0062, Chunlei Huo, Shiming Xiang, Chunhong Pan |
ICIP | 6 |
| 2017 | Image super-resolution via deep dilated convolutional networksabstractDeep learning techniques have been successfully applied in single image super-resolution (SR). Recently, researches have shown that increasing the depth of network can significantly improve SR performance. Very deep networks for SR achieved a large improvement than former methods. However, simply increasing depths basically introduce more parameters and this lead to cumbersome computational cost. In this paper, we present a general and effective method to accelerate very deep networks for single image SR. Our method is based on dilated convolution operation, which support exponential expansion of the receptive field without increasing filter size. With the help of dilated convolution, shallow networks can achieve large receptive field and exploit contextual information in an efficient way. Based on a very deep network, we propose a 12 layers dilated convolutional network for SR (DCNSR). While accelerating 2x speed, our shallow network achieves better performance than original deep networks and shows state-of-the-art reconstructed results. Zehao Huang, Lingfeng Wang 0002, Gaofeng Meng, Chunhong Pan |
ICIP | 4 |
| 2017 | Context-aware cascade network for semantic labeling in VHR imageabstractSemantic labeling for the very high resolution (VHR) image of urban areas is challenging, because of many complex manmade objects with different materials and fine-structured objects located together. Under the framework of convolutional neural networks (CNNs), this paper proposes a novel end-to-end network for semantic labeling. Specifically, our network not only improves the labeling accuracy of complex manmade objects by aggregating multiple context semantics with a cascaded architecture, but also refines fine-structured objects by utilizing the low-level detail in shallow layers of CNNs with a hierarchical pyramid structure. Throughout the network, a dedicated residual correction scheme is employed to amend the latent fitting residual. As a result of these specific components, the whole model works in a global-to-local and coarse-to-fine manner. Experimental results show that our network outperforms the state-of-the-art methods on the large-scale ISPRS Vaihingen 2D Semantic Labeling Challenge dataset. Yongcheng Liu, Bin Fan 0001, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
ICIP | 6 |
| 2017 | MR images segmentation and bias correction via LIC modelabstractThis paper presents a novel Linear Intrinsic Component (LIC) model for simultaneous estimation of bias field and segmentation of magnetic resonance (MR) images with the intensity inhomogene-ity. The core of LIC model is linear transformation, which is derived from Taylor expansion of non-linear model. Due to the linear transformation, observed image can be decomposed into four components, namely, true image, which characterizes a physical property of the tissues, multiplicative and additive bias fields, which result in intensity inhomogeneity, and Gaussian noises. Based on sub-space constraint on two bias fields, and piecewise smoothness restriction on the true image, we can performing the task of joint bias field estimation and image segmentation. To model the complex noises subject to non-gaussian distribution, we further extend LIC model by introducing the non-gaussian noise term, and propose the Non-Gaussian LIC (NGLIC) model. By adopting L1regularization to our solution, the NGLIC model can be effectively solved by iterative soft-thresholding approach. Both LIC and NGLIC models are evaluated on a lot of MR simulated images downloaded from Brain-Web and real images, showing the superiority to the state-of-the-art approach on both segmentation and bias field correction results. Lingfeng Wang 0002, Bin Lav, Chunhong Pan |
ICIP | 4 |
| 2017 | Cascaded temporal spatial features for video action recognitionabstractExtracting spatial-temporal descriptors is a challenging task for video-based human action recognition. We decouple the 3D volume of video frames directly into a cascaded temporal spatial domain via a new convolutional architecture. The motivation behind this design is to achieve deep nonlinear feature representations with reduced network parameters. First, a 1D temporal network with shared parameters is first constructed to map the video sequences along the time axis into feature maps in temporal domain. These feature maps are then organized into channels like those of RGB image (named as Motion Image here for abbreviation), which is desired to preserve both temporal and spatial information. Second, the Motion Image is regarded as the input of the latter cascaded 2D spatial network. With the combination of the 1D temporal network and the 2D spatial network together, the size of whole network parameters is largely reduced. Benefiting from the Motion Image, our network is an end-to-end system for the task of action recognition, which can be trained with the classical algorithm of back propagation. Quantities of comparative experiments on two benchmark datasets demonstrate the effectiveness of our new architecture. Tingzhao Yu, Huxiang Gu, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
ICIP | 5 |
| 2017 | Efficient cloud detection in remote sensing images using edge-aware segmentation network and easy-to-hard training strategyabstractDetecting cloud regions in remote sensing image (RSI) is very challenging yet of great importance to meteorological forecasting and other RSI-related applications. Technically, this task is typically implemented as a pixel-level segmentation. However, traditional methods based on handcrafted or low-level cloud features often fail to achieve satisfactory performances from images with bright non-cloud and/or semitransparent cloud regions. What is more, the performances could be further degraded due to the ambiguous boundaries caused by complicated textures and non-uniform distribution of intensities. In this paper, we propose a multi-task based deep neural network for cloud detection in RSIs. Architecturally, our network is designed to combine the two tasks of cloud segmentation and cloud edge detection together to encourage a better detection near cloud boundaries, resulting in an end-to-end approach for accurate cloud detection. Accordingly, an efficient sample selection strategy is proposed to train our network in an easy-to-hard manner, in which the number of the selected samples is governed by a weight that is annealed until the entire training samples have been considered. Both visual and quantitative comparisons are conducted on RSIs collected from Google Earth. The experimental results indicate that our method can yield superior performance over the state-of-the-art methods. Kun Yuan 0003, Gaofeng Meng, Dongcai Cheng, Shiming Xiang, Chunhong Pan |
ICIP | 6 |
| 2017 | Structured binary feature extraction for hyperspectral imagery classificationabstractIn this paper, we propose a novel structured binary feature extraction method for hyperspectral image classification. To pursuit high discriminative ability and low memory cost, we resort to applying the learning to hash technique to the traditional spectral-spatial hyperspectral features. We show how the structured information among different kinds of features and different feature groups can be used to learn discriminative binary features for classification. Experiments on two standard benchmark hyperspectral data sets demonstrate the effectiveness of the proposed method. Zisha Zhong, Bin Fan 0001, Shiming Xiang, Chunhong Pan |
ICIP | 5 |
| 2017 | Active Rectification of Curved Document Images Using Structured Beams
Gaofeng Meng, Shiming Xiang, Chunhong Pan, Nanning Zheng 0001 |
Int. J. Comput. Vis. | 3 |
| 2017 | Ordinal pyramid coding for rotation invariant feature extraction
Guoli Wang 0004, Bin Fan 0001, Chunhong Pan |
Neurocomputing | 4 |
| 2017 | Building Regional Covariance Descriptors for Vehicle DetectionabstractWe study the question of building regional covariance descriptors (RCDs) for vehicle detection from high-resolution satellite images. A unified way is proposed to build RCD features by constant convolutional kernels in the forms of 2-D masks. Two novel formulas are designed to construct different RCD types based upon one or two convolutional masks, obtaining ten novel RCD features by four simple constant convolutional masks. Experiments show that such convolutional-mask-based RCDs outperform the previous image-derivative-based RCDs, the popular local binary patterns (LBPs), the histogram of oriented gradients (HOGs), and LBP+HOG. Furthermore, feeding to nonlinear support vector machines (SVMs) of two kernel types [L1kernel and radial basis function (RBF)], these RCDs outperform four known deep convolutional neural networks: AlexNet, GoogLeNet, CaffeNet, and LeNet, as well as their fine-tuned models by their well-trained weights of imageNet classification. Among three popular classic classifiers we have tested in the experiments, nonlinear SVMs outperform BP and Adaboost obviously, and L1kernel exceeds RBF slightly. Xueyun Chen, Ren-Xi Gong, Ling-Ling Xie, Shiming Xiang, Cheng-Lin Liu 0001, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2017 | SeNet: Structured Edge Network for Sea-Land SegmentationabstractSeparating an optical remote sensing image into sea and land areas is very challenging yet of great importance to coastline extraction and subsequent object detection. Traditional methods based on handcrafted feature extraction and image processing often face this dilemma when confronting high-resolution remote sensing images for their complicated texture and intensity distribution. In this letter, we apply the prevalent deep convolutional neural networks to the sea–land segmentation problem and make two innovations on top of the traditional structure. First, we propose a local smooth regularization to achieve better spatially consistent results, which frees us from the complicated morphological operations that are commonly used in traditional methods. Second, we use a multitask loss to simultaneously obtain the segmentation and edge detection results. The attached structured edge detection branch can further refine the segmentation result and dramatically improve edge accuracy. Experiments on a set of natural-colored images from Google Earth demonstrate the effectiveness of our approach in terms of quantitative and visual performances compared with state-of-the-art methods. Dongcai Cheng, Gaofeng Meng, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | Feature Extraction by Rotation-Invariant Matrix Representation for Object Detection in Aerial ImageabstractThis letter proposes a novel rotation-invariant feature for object detection in optical remote sensing images. Different from previous rotation-invariant features, the proposed rotation-invariant matrix (RIM) can incorporate partial angular spatial information in addition to radial spatial information. Moreover, it can be further calculated between different rings for a redundant representation of the spatial layout. Based on the RIM, we further propose an RIM_FV_RPP feature for object detection. For an image region, we first densely extract RIM features from overlapping blocks; then, these RIM features are encoded into Fisher vectors; finally, a pyramid pooling strategy that hierarchically accumulates Fisher vectors in ring subregions is used to encode richer spatial information while maintaining rotation invariance. Both of the RIM and RIM_FV_RPP are rotation invariant. Experiments on airplane and car detection in optical remote sensing images demonstrate the superiority of our feature to the state of the art. Guoli Wang 0004, Xinchao Wang, Bin Fan 0001, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | Edge-directed single image super-resolution via cross-resolution sharpening function learning
Lingfeng Wang 0002, Chunhong Pan |
Multim. Tools Appl. | 4 |
| 2017 | Ensemble based deep networks for image super-resolution
Lingfeng Wang 0002, Zehao Huang, Yongchao Gong, Chunhong Pan |
Pattern Recognit. | 4 |
| 2017 | Automatic Road Detection and Centerline Extraction via Cascaded End-to-End Convolutional Neural NetworkabstractAccurate road detection and centerline extraction from very high resolution (VHR) remote sensing imagery are of central importance in a wide range of applications. Due to the complex backgrounds and occlusions of trees and cars, most road detection methods bring in the heterogeneous segments; besides for the centerline extraction task, most current approaches fail to extract a wonderful centerline network that appears smooth, complete, as well as single-pixel width. To address the above-mentioned complex issues, we propose a novel deep model, i.e., a cascaded end-to-end convolutional neural network (CasNet), to simultaneously cope with the road detection and centerline extraction tasks. Specifically, CasNet consists of two networks. One aims at the road detection task, whose strong representation ability is well able to tackle the complex backgrounds and occlusions of trees and cars. The other is cascaded to the former one, making full use of the feature maps produced formerly, to obtain the good centerline extraction. Finally, a thinning algorithm is proposed to obtain smooth, complete, and single-pixel width road centerline network. Extensive experiments demonstrate that CasNet outperforms the state-of-the-art methods greatly in learning quality and learning speed. That is, CasNet exceeds the comparing methods by a large margin in quantitative performance, and it is nearly 25 times faster than the comparing methods. Moreover, as another contribution, a large and challenging road centerline data set for the VHR remote sensing image will be publicly available for further studies. Ying Wang 0008, Shibiao Xu, Hongzhen Wang, Shiming Xiang, Chunhong Pan |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2017 | Cross-Modal Hashing via Rank-Order PreservingabstractDue to the query effectiveness and efficiency, cross-modal similarity search based on hashing has acquired extensive attention in the multimedia community. Most existing methods do not explicitly employ the ranking information when learning hash functions, which is quite important for building practical retrieval systems. To solve this issue, this paper proposes a rank-order preserving hashing (RoPH) method with a novel regression-based rank-order preserving loss that has provable large margin property and is easy to optimize. Moreover, we jointly learn the binary codes and hash functions instead of using any relaxation trick. To solve the induced optimization problem, the alternating descent technique is adopted and each subproblem can be solved conveniently. Specifically, we show that the involved binary quadratic programming subproblem with respect to an introduced auxiliary binary variable satisfies submodularity, enabling us to use the off-the-shelf graph-cut algorithms to solve it exactly and efficiently. Extensive experiments on three benchmarks demonstrate that RoPH significantly improves the ranking quality over the state of the arts. Kun Ding 0001, Bin Fan 0001, Chunlei Huo, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 5 |
| 2016 | Data-guided random walks for fine-structured object segmentationabstractRandom walks (RW) is a popular technique for object segmentation. Apart from the satisfactory performance in various applications, its most appealing advantage is the computational efficiency. However, RW often fails to produce complete and connected results in fine-structured (FS) object segmentation. To utilize the high efficiency and overcome the drawbacks in tackling FS objects, we develop a novel approach within the RW framework. Specifically, we propose to introduce labeling preference learned from the image data into the RW model to guide the propagation of random walkers. With the help of the guidance, random walkers are more likely to propagate correctly to the FS regions, thus yielding more accurate results. Similar to RW, this approach also bears properties such as computational efficiency, closed-form solution and unique global optimum. Moreover, it has the capacities of handling disconnected objects and transferring segmentation. Comparative experimental results demonstrate that the proposed approach achieves the state-of-the-art performance in FS object segmentation, with a low requirement of runtime. Yongchao Gong, Shiming Xiang, Chunhong Pan |
ICASSP | 3 |
| 2016 | Fine-structured object segmentation via edge-guided graph cut with interaction simplificationabstractFine-structured object segmentation is a challenging problem in object segmentation community. There are mainly two difficulties that can seriously degrade the segmentation quality: 1) insufficient interactions on fine structures due to the high demand of time and manual efforts, and 2) shrinking bias that discourages long object boundaries. To address these two issues, we develop a novel method within the graph cut framework. First, the commonly used operation of scribbling or dragging bounding boxes is replaced by loosely drawing a few rectangles, thus the interaction burden is largely reduced. Second, an edge-guided graph cut model is proposed to mitigate shrinking bias. This model enforces connectivity of fine structures by adjusting the weighting between neighboring pixels. Finally, the segmentation task is formulated as an optimization problem, which can be optimized effectively and efficiently. Comparative experimental results demonstrate the effectiveness of our method. Yongchao Gong, Shiming Xiang, Lingfeng Wang 0002, Chunhong Pan |
ICASSP | 4 |
| 2016 | Learning-based fully 3D face reconstruction from a single imageabstractThis paper presents an algorithm for fully reconstructing a 3D face from a single image. This task is still highly challenging as most current methods only care about the frontal face, ignoring side face, such as the neck, ears etc. In our algorithm, to get the more detailed texture, we deal with the shape reconstruction and texture recovery respectively. For shape, we estimate the deformation of the 3D model by a set of feature points. For texture, due to the similar facial structure, we divide the full texture into patches and show how sparse learning model can be used to fully recover the texture of the 3D face. Extensive experiment results on the CMU-PIE database and images downloaded from the Internet demonstrate that our method outperforms the state-of-the-art methods. Ying Wang 0008, Feiyun Zhu, Chunhong Pan |
ICASSP | 4 |
| 2016 | Enhancement of Low Light Level Images with coupled dictionary learningabstractLow Light Level Images (LLLIs) are captured with exceptionally low brightness and low contrast, and cannot be enhanced satisfactorily with ordinary methods. In this paper, we propose a LLLI enhancement method using coupled dictionary learning. During the training stage, a pair of dictionaries and a linear mapping function are learned simultaneously. The dictionary pair aims to describe the raw LLLIs and their enhanced versions, and the linear mapping function models the correspondence between the representations of the dictionary pair. In the enhancement process, the resulting image is generated through dictionary mapping from patches of the input LLLI. We adopt a clustering strategy to improve the robustness of coupled dictionary learning, and propose an improved algorithm for fast implementation. Experimental results on real images demonstrate the effectiveness of our method. Xinwei Jiang, Chunhong Pan |
ICPR | 3 |
| 2016 | Building extraction from multi-source remote sensing images via deep deconvolution neural networksabstractBuilding extraction from remote sensing images is of great importance in urban planning. Yet it is a longstanding problem for many complicate factors such as various scales and complex backgrounds. This paper proposes a novel supervised building extraction method via deep deconvolution neural networks (DeconvNet). Our method consists of three steps. First, we preprocess the multi-source remote sensing images provided by the IEEE GRSS Data Fusion Contest. A high-quality Vancouver building dataset is created on pansharpened images whose ground-truth are obtained from the OpenStreetMap project. Then, we pretrain a deep deconvolution network on a public large-scale Massachusetts building dataset, which is further fine-tuned by two band combinations (RGB and NRG) of our dataset, respectively. Moreover, the output saliency maps of the fine-tuned models are fused to produce the final building extraction result. Extensive experiments on our Vancouver building dataset demonstrate the effectiveness and efficiency of the proposed method. To the best of our knowledge, it is the first work to use deconvolution networks for building extraction from remote sensing images. Zuming Huang, Hongzhen Wang, Haichang Li, Limin Shi, Chunhong Pan |
IGARSS | 6 |
| 2016 | Efficient sea-land segmentation using seeds learning and edge directed graph cut
Dongcai Cheng, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Neurocomputing | 4 |
| 2016 | Accurate urban road centerline extraction from VHR imagery via multiscale segmentation and tensor voting
Feiyun Zhu, Shiming Xiang, Ying Wang 0008, Chunhong Pan |
Neurocomputing | 5 |
| 2016 | Road Centerline Extraction via Semisupervised Segmentation and Multidirection Nonmaximum SuppressionabstractAccurate road centerline extraction from remotely sensed images plays a significant role in road map generation and updating. In the road extraction problem, the acquisition of labeled data is time consuming and costly; thus, there are only a small amount of labeled samples in reality. In the existing centerline extraction algorithms, the thinning-based algorithms always produce small spurs that reduce the smoothness and accuracy of the road centerline; the regression-based algorithms can extract a smooth road network, but they are time consuming. To solve the aforementioned problems, we propose a novel road centerline extraction method, which is constructed based on semisupervised segmentation and multiscale filtering (MF) and multidirection nonmaximum suppression (M-NMS) (MF&M-NMS). Specifically, a semisupervised method, which explores the intrinsic structures between the labeled samples and the unlabeled ones, is introduced to obtain the segmentation result. Then, a novel MF&M-NMS-based algorithm is proposed to gain a smooth and complete road centerline network. Experimental results on a public data set demonstrate that the proposed method achieves comparable or better performances by comparing with the state-of-the-art methods. In addition, our method is nearly ten times faster than the state-of-the-art methods. Feiyun Zhu, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2016 | Fine-structured object segmentation via neighborhood propagation
Yongchao Gong, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 3 |
| 2016 | Blur-Kernel Bound Estimation From Pyramid StatisticsabstractThis letter presents an approach for automatically estimating the spatial bound of the blur kernel in a motion-blurred image based on the statistics of multilevel image gradients. We observe that blur has a significant impact on the histogram of oriented gradients (HOGs) at higher levels of an image pyramid, but has much less of an impact at coarser levels. Based on this fact, we estimate the spatial bound of the unknown blur kernel using a learning-based approach. We first learn a generic pyramid HOG model from natural sharp images, then given an HOG pyramid of a blurry image, we predict the corresponding model of its latent sharp image. Finally, we learn another model to predict the spatial kernel bound from the difference between the observed and the predicted HOG pyramids. Experimental results show that the proposed method can estimate accurate blur kernel sizes, enabling existing blind deconvolution methods to achieve best possible results. Shaoguo Liu, Jue Wang 0001, Chunhong Pan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Efficient Multiple Feature Fusion With Hashing for Hyperspectral Imagery Classification: A Comparative StudyabstractDue to the complementary properties of different features, multiple feature fusion has a large potential for hyperspectral imagery classification. At the meantime, hashing is promising in representing a high-dimensional float-type feature with extremely low bit binary codes while maintaining the performance. In this paper, we study the possibility of using hashing to fuse multiple features for hyperspectral imagery classification. For this purpose, we propose a multiple feature fusion framework to evaluate the performance of using different hashing methods. For comparison and completeness, we also have an extensive comparison to five subspace-based dimension reduction methods and six fusion-based methods which are popular solutions to deal with multiple features in hyperspectral image classification. Experimental results on four benchmark hyperspectral data sets demonstrate that using hashing to fuse multiple features can achieve comparable or better performance with the traditional subspace-based dimension reduction methods and fusion-based methods. Moreover, the binary features obtained by using hashing need much less storage and are faster to compute distances with the help of machine instructions. Zisha Zhong, Bin Fan 0001, Kun Ding 0001, Haichang Li, Shiming Xiang, Chunhong Pan |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2016 | Cross-Modal Retrieval via Deep and Bidirectional Representation LearningabstractCross-modal retrieval emphasizes understanding inter-modality semantic correlations, which is often achieved by designing a similarity function. Generally, one of the most important things considered by the similarity function is how to make the cross-modal similarity computable. In this paper, a deep and bidirectional representation learning model is proposed to address the issue of image-text cross-modal retrieval. Owing to the solid progress of deep learning in computer vision and natural language processing, it is reliable to extract semantic representations from both raw image and text data by using deep neural networks. Therefore, in the proposed model, two convolution-based networks are adopted to accomplish representation learning for images and texts. By passing the networks, images and texts are mapped to a common space, in which the cross-modal similarity is measured by cosine distance. Subsequently, a bidirectional network architecture is designed to capture the property of the cross-modal retrieval-the bidirectional search. Such architecture is characterized by simultaneously involving the matched and unmatched image-text pairs for training. Accordingly, a learning framework with maximum likelihood criterion is finally developed. The network parameters are optimized via backpropagation and stochastic gradient descent. A great deal of experiments are conducted to sufficiently evaluate the proposed method on three publicly released datasets: IAPRTC-12, Flickr30k, and Flickr8k. The overall results definitely show that the proposed architecture is effective and the learned representations have good semantics to achieve superior cross-modal retrieval performance. Yonghao He, Shiming Xiang, Cuicui Kang, Jian Wang 0068, Chunhong Pan |
IEEE Trans. Multim. | 5 |
| 2016 | MSDLSR: Margin Scalable Discriminative Least Squares Regression for Multicategory ClassificationabstractIn this brief, we propose a new margin scalable discriminative least squares regression (MSDLSR) model for multicategory classification. The main motivation behind the MSDLSR is to explicitly control the margin of DLSR model. We first prove that the DLSR is a relaxation of the traditional L2-support vector machine. Based on this fact, we further provide a theorem on the margin of DLSR. With this theorem, we add an explicit constraint on DLSR to restrict the number of zeros of dragging values, so as to control the margin of DLSR. The new model is called MSDLSR. Theoretically, we analyze the determination of the margin and support vectors of MSDLSR. Extensive experiments illustrate that our method outperforms the current state-of-the-art approaches on various machine leaning and real-world data sets. Lingfeng Wang 0002, Xu-Yao Zhang, Chunhong Pan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Layer-Wise Floorplan Extraction for Automatic Urban Building ReconstructionabstractUrban building reconstruction is an important step for urban digitization and realisticvisualization. In this paper, we propose a novel automatic method to recover urban building geometry from 3D point clouds. The proposed method is suitable for buildings composed of planar polygons and aligned with the gravity direction, which are quite common in the city. Our key observation is that the building shapes are usually piecewise constant along the gravity direction and determined by several dominant shapes. Based on this observation, we formulate building reconstruction as an energy minimization problem under the Markov Random Field (MRF) framework. Specifically, point clouds are first cutinto a sequence of slices along the gravity direction. Then, floorplans are reconstructed by extracting boundaries of these slices, among which dominant floorplans are extracted and propagated to other floors via MRF. To guarantee correct propagation, a new distance measurement for floorplans is designed, which first encodes floorplans into strings and then calculates distances between their corresponding strings. Additionally, an image based editing method is also proposed to recover detailed window structures. Experimental results on both synthetic and real data sets have validated the effectiveness of our method. Wei Sui, Lingfeng Wang 0002, Bin Fan 0001, Hongfei Xiao, Huai-Yu Wu, Chunhong Pan |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2015 | 10, 000+ Times Accelerated Robust Subset SelectionabstractSubset selection from massive data with noised information is increasingly popular for various applications. This problem is still highly challenging as current methods are generally slow in speed and sensitive to outliers. To address the above two issues, we propose an accelerated robust subset selection (ARSS) method. Extensive experiments on ten benchmark datasets verify that our method not only outperforms state of the art methods, but also runs 10,000+ times faster than the most related method. Feiyun Zhu, Bin Fan 0001, Xinliang Zhu, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
AAAI | 6 |
| 2015 | Cross-Modal Similarity Learning: A Low Rank Bilinear FormulationabstractThe cross-media retrieval problem has received much attention in recent years due to the rapid increasing of multimedia data on the Internet. A new approach to the problem has been raised which intends to match features of different modalities directly. In this research, there are two critical issues: how to get rid of the heterogeneity between different modalities and how to match the cross-modal features of different dimensions. Recently metric learning methods show a good capability in learning a distance metric to explore the relationship between data points. However, the traditional metric learning algorithms only focus on single-modal features, which suffer difficulties in addressing the cross-modal features of different dimensions. In this paper, we propose a cross-modal similarity learning algorithm for the cross-modal feature matching. The proposed method takes a bilinear formulation, and with the nuclear-norm penalization, it achieves low-rank representation. Accordingly, the accelerated proximal gradient algorithm is successfully imported to find the optimal solution with a fast convergence rate O(1/t2). Experiments on three well known image-text cross-media retrieval databases show that the proposed method achieves the best performance compared to the state-of-the-art algorithms. Cuicui Kang, Shengcai Liao, Yonghao He, Jian Wang 0068, Wenjia Niu, Shiming Xiang, Chunhong Pan |
CIKM | 7 |
| 2015 | Ordinal pyramid pooling for rotation invariant object recognitionabstractLocal feature descriptor plays a fundamental role in many visual tasks, and its rotation invariance is a key issue for many recognition and detection problems. This paper proposes a novel rotation invariant descriptor by ordinal pyramid pooling of local Fourier transform features based on their radial gradient orientations. Since both the low-level feature and pooling strategy are rotation invariant, the obtained descriptor is rotation invariant by nature. Pooling based on orders of gradient orientations is not only invariant to in-plane rotation, but also encodes gradient orientation information into descriptor as well as spatial information to some extent. Moreover, these information is enhanced by the proposed pyramid pooling structure. Therefore, our method is naturally rotation invariant and has strong discriminative ability. Experimental results on the aerial car dataset demonstrate the effectiveness of our descriptor. Guoli Wang 0004, Bin Fan 0001, Chunhong Pan |
ICASSP | 3 |
| 2015 | Explicit order model for region-based level set segmentationabstractRegion-based level set methods have been widely used for image segmentation. Among them, the method based on local binary fitting (LBF) model is an efficient one. Unfortunately, LBF model is sensitive to initial contour. To overcome this disadvantage, we propose two explicit order models, i.e., the global order preserving and local order smoothness models. The global order preserving model ensures that the binary fitting values have the same order globally, while the local order smoothness model requires that these orders are smooth locally. With these two models, our segmentation results are not sensitive to initializations. Experimental results on synthetic and real images show desirable performances of our method, as compared with the state-of-the-art approaches. Lingfeng Wang 0002, Chunhong Pan |
ICASSP | 2 |
| 2015 | kNN Hashing with Factorized Neighborhood RepresentationabstractHashing is very effective for many tasks in reducing the processing time and in compressing massive databases. Although lots of approaches have been developed to learn data-dependent hash functions in recent years, how to learn hash functions to yield good performance with acceptable computational and memory cost is still a challenging problem. Based on the observation that retrieval precision is highly related to the kNN classification accuracy, this paper proposes a novel kNN-based supervised hashing method, which learns hash functions by directly maximizing the kNN accuracy of the Hamming-embedded training data. To make it scalable well to large problem, we propose a factorized neighborhood representation to parsimoniously model the neighborhood relationships inherent in training data. Considering that real-world data are often linearly inseparable, we further kernelize this basic model to improve its performance. As a result, the proposed method is able to learn accurate hashing functions with tolerable computation and storage cost. Experiments on four benchmarks demonstrate that our method outperforms the state-of-the-arts. Kun Ding 0001, Chunlei Huo, Bin Fan 0001, Chunhong Pan |
ICCV | 4 |
| 2015 | Extraction of Virtual Baselines from Distorted Document Images Using Curvilinear ProjectionabstractThe baselines of a document page are a set of virtual horizontal and parallel lines, to which the printed contents of document, e.g., text lines, tables or inserted photos, are aligned. Accurate baseline extraction is of great importance in the geometric correction of curved document images. In this paper, we propose an efficient method for accurate extraction of these virtual visual cues from a curved document image. Our method comes from two basic observations that the baselines of documents do not intersect with each other and that within a narrow strip, the baselines can be well approximated by linear segments. Based upon these observations, we propose a curvilinear projection based method and model the estimation of curved baselines as a constrained sequential optimization problem. A dynamic programming algorithm is then developed to efficiently solve the problem. The proposed method can extract the complete baselines through each pixel of document images in a high accuracy. It is also scripts insensitive and highly robust to image noises, non-textual objects, image resolutions and image quality degradation like blurring and non-uniform illumination. Extensive experiments on a number of captured document images demonstrate the effectiveness of the proposed method. Gaofeng Meng, Zuming Huang, Yonghong Song, Shiming Xiang, Chunhong Pan |
ICCV | 5 |
| 2015 | Super-resolution reconstruction using graph Laplacian penalizationabstractThis paper proposes to employ graph Laplacian penalization in multi-image super-resolution reconstruction. Most state-of-the-art methods use an optimization model with two items: the data fidelity item and the penalization item. However, these methods often ignore the impact of the penalization item and utilize simple formulations such as high-pass filters to fulfill the super-resolution task. As a result, they can not restore much local structural information of the high resolution image. By using graph Laplacian, the proposed method in this paper can retain more local manifold structures in the high resolution images. Based on this idea, the optimization model is constructed and the solution is presented. Comparative experiments have validated our method. The experiments have also tested our method has much faster convergence speed. Limin Shi, Bangyu Li, Shiming Xiang, Chunhong Pan |
ICIP | 5 |
| 2015 | Road extraction via adaptive graph cuts with multiple featuresabstractAccurate road extraction from complex backgrounds plays a fundamental role in a wide range of remote sensing applications. There are two shortcomings for the existing methods: 1) Most of them ignore the spatially contextual information inherent in images; 2) Few existing methods show robustness to the occlusions of cars or trees. To address these two problems, we propose a novel approach via adaptive graph cuts with multiple features. Specifically, for the former problem, we apply multiple features (spectral feature, spatial feature and gradient feature) to obtain not only the spectral characteristic but also the spatially contextual feature. In this way, the structural information of road network can be effectively captured. For the latter, adaptive graph cuts based algorithm is adopted. These two schemes show better performance than state-of-the-art methods under the conditions of occlusions. Experiments on 25 images indicate the validity and effectiveness of our method by comparing with state-of-the-art approaches. Ying Wang 0008, Feiyun Zhu, Chunhong Pan |
ICIP | 4 |
| 2015 | Fine-structured object segmentation via local and nonlocal neighborhood propagationabstractIn this paper, we present a novel method for the challenging task of fine-structured (FS) object segmentation. This task is formulated as a label propagation problem on an affinity graph. To enhance the completeness and connectivity of the FS objects, we introduce a novel neighborhood system combining both local and nonlocal connections, together with a robust scheme for edge weight calculation. Additionally, region cost is incorporated into the energy function to further maintain the connectivity of fine parts where the propagation is hard to reach. An appealing advantage of the proposed method is that the energy minimization has a closed-form solution and global optimum is guaranteed. Comparative experimental results on three datasets demonstrate the effectiveness of the proposed method. Yongchao Gong, Shiming Xiang, Lingfeng Wang 0002, Chunhong Pan |
ICIP | 4 |
| 2015 | Objectness estimation using edgesabstractGenerating object proposals before object detection has become a common way. In this paper, we present a novel method to measure the objectness of bounding boxes using edges. The contours play an important role in object localization and detection. The number of edges that are close to the boundary of a box has strong relationship with the likelihood of the box covering an object. In our method, we adopt a two-step scheme to generate object proposals. In the first step, we count the number of contours close to the box, where we use the proposed “Tile Algorithm” to wipe off the inner edges of a box. In the second step we re-rank the object proposals with a linear SVM classifier across all aspect-ratios for calibration. Experiments on the VOC2007 dataset show that we achieve 96.47% object detection rate with 1000 proposals. Hongzhen Wang, Zikun Liu 0001, Lingfeng Wang 0002, Lubin Weng, Chunhong Pan |
ICIP | 5 |
| 2015 | Visual tracking via manifold regularized local structured sparse representation modelabstractIn this paper, we propose a new visual tracking method via the manifold regularized local structured sparse representation model under the particle filtering tracking framework. First, in order to tackle the difficulties of partial occlusion and illumination variation, the local structured sparse representation model is incorporated by exploiting both partial and spatial information of the target. Second, the manifold regularization is used to ensure that neighboring particles should share similar representation coefficients, so that these particles can cooperate with each other. Extensive experiments are performed on various video sequences, showing improvement over the state-of-the-art approaches. Lingfeng Wang 0002, Chunhong Pan |
ICIP | 2 |
| 2015 | Adaptive regularization level set evolution for medical image segmentation and bias field correctionabstractIn this paper, we propose a level-set based segmentation method for medical images with intensity inhomogeneity. Maximum a Posteriori estimation is adopted to combine image segmentation and bias field correction into a unified framework. Within this framework, both contour prior and bias field prior can be fully used. In order to restrict bias field, we introduce an adaptive regularization. Based on this new adaptive regularization, the bias field is estimated more smooth and the input medical image with intensity inhomogeneity is recovered more clearly. Especially, the estimated bias field of our method introduces less structure information obtained from input image. Experimental results on both synthetic and real images show the advantages of our method in both segmentation and bias field correction accuracies as compared with the state-of-the-art approaches. Xiaomeng Xin, Lingfeng Wang 0002, Chunhong Pan, Shigang Liu |
ICIP | 3 |
| 2015 | Large Scale Image Annotation via Deep Representation Learning and Tag Embedding LearningabstractIn this paper, we focus on the issue of large scale image annotation, whereas most existing methods are devised for small datasets. A novel model based on deep representation learning and tag embedding learning is proposed. Specifically, the proposed model learns an unified latent space for image visual features and tag embeddings simultaneously. Furthermore, a metric matrix is introduced to estimate the relevance scores between images and tags. Finally, an objective function modeling triplet relationships (irrelevant tag, image, relevant tag) is proposed with maximum margin pursuit. The proposed model is easy to tackle new images and tags via online learning and has a relatively low test computation complexity. Experimental results on NUS-WIDE dataset demonstrate the effectiveness of the proposed model. Yonghao He, Jian Wang 0068, Cuicui Kang, Shiming Xiang, Chunhong Pan |
ICMR | 5 |
| 2015 | Image-Text Cross-Modal Retrieval via Modality-Specific Feature LearningabstractCross-modal retrieval extends the ability of search engines to deal with the massive cross-modal data. The goal of image-text cross-modal retrieval is to search images (texts) by using text (image) queries by computing the similarities of images and texts directly. Many existing methods rely on low-level visual features and textual features for cross-modal retrieval, ignoring the characteristics existing in the raw data of different modalities. In this paper, a novel model based on modality-specific feature learning is proposed. Considering the characteristics of different modalities, the model uses two types of convolutional neural networks to map the raw data to the latent space representations for images and texts, respectively. Particularly, the convolution based network used for texts involves word embedding learning, which has been proved effective to extract meaningful textual features for text classification. In the latent space, the mapped features of images and texts form relevant and irrelevant image-text pairs, which are used by the one-vs-more learning scheme. This learning scheme can achieve ranking functionality by allowing for one relevant and more irrelevant pairs. The standard back-propagation technique is employed to update the parameters of two convolutional networks. Extensive cross-modal retrieval experiments are carried out on three challenging datasets that consist of image-document pairs or image-query click-through data from a search engine, and the results firmly demonstrate that the proposed model is much more effective. Jian Wang 0068, Yonghao He, Cuicui Kang, Shiming Xiang, Chunhong Pan |
ICMR | 5 |
| 2015 | Image Deblurring with Coupled Dictionary Learning
Shiming Xiang, Gaofeng Meng, Ying Wang 0008, Chunhong Pan, Changshui Zhang |
Int. J. Comput. Vis. | 4 |
| 2015 | Image tag-ranking via pairwise supervision based semi-supervised model
Yonghao He, Cuicui Kang, Jian Wang 0068, Shiming Xiang, Chunhong Pan |
Neurocomputing | 5 |
| 2015 | Sparse Hierarchical Clustering for VHR Image Change DetectionabstractThe traditional clustering approaches are limited for the unsupervised change detection of very high resolution images due to the multimodal distribution of change features. To overcome this difficulty, a sparse hierarchical clustering approach is proposed. Discriminative change features are generated by stacking bitemporal multiscale center-symmetric local binary pattern features. In order to explore the multimodal and hierarchical distribution of the change features, a tree-structured dictionary is learned from the pseudotraining set and the unlabeled data. The sparse reconstruction error, a more robust distance compared to the Euclidean distance, is used to determine the label of each change feature. Comparative experiments demonstrate the effectiveness of the proposed method. Kun Ding 0001, Chunlei Huo, Zisha Zhong, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2015 | Multicluster Spatial-Spectral Unsupervised Feature Selection for Hyperspectral Image ClassificationabstractA new unsupervised spatial-spectral feature selection method for hyperspectral images has been proposed in this letter. The key idea is to select the features that better preserve the multicluster structure of the multiple spatial-spectral features. Specifically, the multicluster structure information is obtained through spectral clustering utilizing a weighted combination of the multiple features. Then, such information is preserved in a group-sparsity-based robust linear regression model. The features that contribute more in preserving the multicluster structure information are selected. Comparative experiments on two popular real hyperspectral images validate the effectiveness of the proposed method, showing higher classification accuracy. Haichang Li, Shiming Xiang, Zisha Zhong, Kun Ding 0001, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2015 | Discriminant Tensor Spectral-Spatial Feature Extraction for Hyperspectral Image ClassificationabstractWe propose to integrate spectral-spatial feature extraction and tensor discriminant analysis for hyperspectral image classification. First, we apply remarkable spectral-spatial feature extraction approaches in the hyperspectral cube to extract a feature tensor for each pixel. Then, based on class label information, local tensor discriminant analysis is used to remove redundant information for subsequent classification procedure. The approach not only extracts sufficient spectral-spatial features from original hyperspectral images but also gets better feature representation owing to tensor framework. Comparative results on two benchmarks demonstrate the effectiveness of our method. Zisha Zhong, Bin Fan 0001, Jiangyong Duan, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2015 | BB-Homography: Joint Binary Features and Bipartite Graph Matching for Homography EstimationabstractHomography estimation is a fundamental problem in the field of computer vision. For estimating the homography between two images, one of the key issues is to match keypoints in the reference image to the keypoints in the moving image. To match keypoints in real time, a binary image descriptor, due to its low matching and storage costs, emerges as a more and more popular tool. Upon achieving the low costs, the binary descriptor sacrifices the discriminative power of using floating points. In this paper, we present BB-Homography, a new approach that fuses fast binary descriptor matching and bipartite graph for homography estimation. Starting with binary descriptor matching, BB-Homography uses bipartite graph matching (GM) algorithm to refine the matching results, which are finally passed over to estimate homography. On realizing the correlation between keypoint correspondence and homography estimation, BB-Homography iteratively performs the GM and the homography estimation such that they can refine each other at each iteration. In particular, based on spectral graph, a fast bipartite GM algorithm is developed for lowering the time cost of BB-Homography. BB-Homography is extensively evaluated on both public benchmarks and live-captured video streams that consistently shows that BB-Homography outperforms conventional methods for homography estimation. Shaoguo Liu, Yiyi Wei, Chunhong Pan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Manifold Regularized Local Sparse Representation for Face RecognitionabstractSparse representation-(or sparse coding)-based classification has been successfully applied to face recognition. However, it can become problematic in the presence of illumination variations or occlusions. In this paper, we propose a Manifold Regularized Local Sparse Representation (MRLSR) model to address such difficulties. The key idea behind the MRLSR method is that all coding vectors in sparse representation should be group sparse, which means holding the two properties of both individual sparsity and local similarity. As a consequence, the face recognition rate can be considerably improved. The MRLSR model is optimized by the modified homotopy algorithm, which keeps stable under different choices of the weighting parameter. Extensive experiments are performed on various face databases, which contain illumination variations and occlusions. We show that the proposed method outperforms the state-of-the-art approaches and provides the highest recognition rate. Lingfeng Wang 0002, Huai-Yu Wu, Chunhong Pan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Automatic Spatial-Spectral Feature Selection for Hyperspectral Image via Discriminative Sparse Multimodal LearningabstractSpectral-spatial feature combination for hyperspectral image analysis has become an important research topic in hyperspectral remote sensing applications. A simple and straightforward way to integrate spectral-spatial features is to concatenate heterogeneous features into a long vector. Then, the dimensionality reduction techniques, i.e., feature selection, are applied before subsequent utilizations. However, such representation can introduce redundancy and noise. Moreover, traditional single-feature selection methods treat different features equally and ignore their complementary properties. As a result, the performance of subsequent tasks, i.e., classification, would drop down. In this paper, we propose a novel approach to integrate the spectral-spatial features based on the concatenating strategy, termed discriminative sparse multimodal learning for feature selection (DSML-FS). In the proposed method, joint structured sparsity regularizations are used to exploit the intrinsic data structure and relationships among different features. Discriminative least squares regression is applied to enlarge the distance between classes. Therefore, the weight matrix incorporating the information of feature wise and individual properties is automatically learned for spectral-spatial feature selection. We develop an alternative iterative algorithm to solve the nonlinear optimization problem in DSML-FS with global convergence. We systematically evaluate the proposed algorithm on three available hyperspectral data sets, and the encouraging experimental results demonstrate the effectiveness of DSML-FS. Qian Zhang 0009, Chunhong Pan |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | Robust Hyperspectral Unmixing With Correntropy-Based MetricabstractHyperspectral unmixing is one of the crucial steps for many hyperspectral applications. The problem of hyperspectral unmixing has proved to be a difficult task in unsupervised work settings where the endmembers and abundances are both unknown. In addition, this task becomes more challenging in the case that the spectral bands are degraded by noise. This paper presents a robust model for unsupervised hyperspectral unmixing. Specifically, our model is developed with the correntropy-based metric where the nonnegative constraints on both endmembers and abundances are imposed to keep physical significance. Besides, a sparsity prior is explicitly formulated to constrain the distribution of the abundances of each endmember. To solve our model, a half-quadratic optimization technique is developed to convert the original complex optimization problem into an iteratively reweighted nonnegative matrix factorization with sparsity constraints. As a result, the optimization of our model can adaptively assign small weights to noisy bands and put more emphasis on noise-free bands. In addition, with sparsity constraints, our model can naturally generate sparse abundances. Experiments on synthetic and real data demonstrate the effectiveness of our model in comparison to the related state-of-the-art unmixing models. Ying Wang 0008, Chunhong Pan, Shiming Xiang, Feiyun Zhu |
IEEE Trans. Image Process. | 2 |
| 2015 | Learning Consistent Feature Representation for Cross-Modal Multimedia RetrievalabstractThe cross-modal feature matching has gained much attention in recent years, which has many practical applications, such as the text-to-image retrieval. The most difficult problem of cross-modal matching is how to eliminate the heterogeneity between modalities. The existing methods (e.g., CCA and PLS) try to learn a common latent subspace, where the heterogeneity between two modalities is minimized so that cross-matching is possible. However, most of these methods require fully paired samples and suffer difficulties when dealing with unpaired data. Besides, utilizing the class label information has been found as a good way to reduce the semantic gap between the low-level image features and high-level document descriptions. Considering this, we propose a novel and effective supervised algorithm, which can also deal with the unpaired data. In the proposed formulation, the basis matrices of different modalities are jointly learned based on the training samples. Moreover, a local group-based priori is proposed in the formulation to make a better use of popular block based features (e.g., HOG and GIST). Extensive experiments are conducted on four public databases: Pascal VOC2007, LabelMe, Wikipedia, and NUS-WIDE. We also evaluated the proposed algorithm with unpaired data. By comparing with existing state-of-the-art algorithms, the results show that the proposed algorithm is more robust and achieves the best performance, which outperforms the second best algorithm by about 5% on both the Pascal VOC2007 and NUS-WIDE databases. Cuicui Kang, Shiming Xiang, Shengcai Liao, Changsheng Xu, Chunhong Pan |
IEEE Trans. Multim. | 5 |
| 2015 | Automatic blur-kernel-size estimation for motion deblurring
Shaoguo Liu, Jue Wang 0001, Sunghyun Cho, Chunhong Pan |
Vis. Comput. | 5 |
| 2014 | Active Flattening of Curved Document Images via Two Structured BeamsabstractDocument images captured by a digital camera often suffer from serious geometric distortions. In this paper, we propose an active method to correct geometric distortions in a camera-captured document image. Unlike many passive rectification methods that rely on text-lines or features extracted from images, our method uses two structured beams illuminating upon the document page to recover two spatial curves. A developable surface is then interpolated to the curves by finding the correspondence between them. The developable surface is finally flattened onto a plane by solving a system of ordinary differential equations. Our method is a content independent approach and can restore a corrected document image of high accuracy with undistorted contents. Experimental results on a variety of real-captured document images demonstrate the effectiveness and efficiency of the proposed method. Gaofeng Meng, Ying Wang 0008, Shenquan Qu, Shiming Xiang, Chunhong Pan |
CVPR | 5 |
| 2014 | Urban road extraction via graph cuts based probability propagationabstractIn this paper, we propose a graph cuts (GC) based probability propagation approach to automatically extract road network from complex remote sensing images. First, the support vector machine (SVM) classifier with a sigmoid model is applied to assign each pixel a posterior probability of being labelled as road class, which avoids the weaknesses of hard labels in general SVM. Then a GC based probability propagation algorithm is employed to keep the extracted road results smooth and coherent, which can reduce the connections between roads and road-like objects. Finally, a road-geometrical prior is considered to refine the extraction result, so that the non-road objects in images can be removed. Experimental results on two remote sensing image datasets indicate the validity and effectiveness of our method by comparing with two other approaches. Ying Wang 0008, Yongchao Gong, Feiyun Zhu, Chunhong Pan |
ICIP | 5 |
| 2014 | Image annotation via learning the image-label interrelationsabstractThe goal of image annotation is to automatically assign meaningful and content-related labels to the digital images by using machines. It is beneficial to image search and image sharing in social networks. Various methods for image annotation are proposed in last decade and they have gained much progress. However, most of them are not precise and fast enough for real-world applications. In this paper, we propose a novel and fast image annotation method via learning the image-label interrelation. The main idea of the proposed method is to predict labels by linearly propagating the label information through the image-label interrelation and the image similarities. Thus, we propose a model based on the regression between the label groundtruth and the propagated label information to learn the image-label interrelation. In addition, a label-biased regularization is integrated into our model to learn more effective and meaningful image-label interrelation. Finally, our model can be solved in closed form, therefore it achieves a fast learning process. Experimental results on three benchmark datasets demonstrate that our method shows the comparable performance with state-of-the-art methods and has faster learning time. Yonghao He, Jian Wang 0068, Shiming Xiang, Chunhong Pan |
ICIP | 4 |
| 2014 | Facade repetition extraction using block matrix based modelabstractRepetition extraction plays an important role in facade image analysis. In this paper, this task is handled within the graph cut based image segmentation framework. To model the repetitions, generalized translation symmetry (GTS) is introduced to enable aperiodic repetition layouts. More importantly, GTS is explicitly formulated in terms of matrix multiplication. That is, GTS is viewed as the product of a repetitive pattern and two block matrices. These two block matrices are employed to represent the vertical and horizontal symmetry respectively. On this basis, repetition extraction is formulated as a GTS constrained energy minimization problem. An alternatively optimization algorithm based on graph cut and dynamic programming is finally developed to solve the problem. Experimental results demonstrate the validity of our method. Hongfei Xiao, Gaofeng Meng, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
ICIP | 5 |
| 2014 | 3D object tracking via boundary constrained region-based modelabstractIn this paper, we propose a method for joint 2D segmentation and 2D-3D pose tracking. First, we define a novel energy functional which considers the discrimination between statistical appearance models and the coherence among neighboring pixels simultaneously. And then, a particle filter-like stochastic optimization technique is adopted to solve the energy functional, so that a preferable initial value can be provided for the subsequent damped Newton optimization method. Furthermore, an occlusion-aware updating strategy is utilized for appearance models, which can easily increase the foreground learning rate. As a result, our method is more suitable for the video sequences with occlusion. Experimental results highlight excellent performance on challenging synthetic and real-world sequences as compared with the state-of-the-art approaches. Lingfeng Wang 0002, Wei Sui, Huai-Yu Wu, Chunhong Pan |
ICIP | 5 |
| 2014 | Facade Labeling via Explicit Matrix Factorization
Hongfei Xiao, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICISP | 5 |
| 2014 | Local Label Probability Propagation for Hyperspectral Image ClassificationabstractClassification of hyper spectral images is an important issue in remote sensing image processing systems. Hyper spectral images have advantages in pixel-wise classification owing to the high spectral resolution. However, the pixel-wise classification result often introduces the salt-and-pepper appearance because of the complex noise produced by atmosphere and instrument. An effective way to overcome this phenomenon is to resort to the spatial information. This paper proposes a method to solve the above problem by using spatial similarity information. First, in order to avoid the effect of noisy pixels and mixed pixels, reliable seeds are selected in local windows according to the agreement between the central pixel and its spatial neighbors. Then, the information of the reliable seeds is propagated to their spatial neighbors by a graph Laplacian. Specifically, the graph Laplacian is designed to propagate information among spatial neighbors with close similarity relationship so that some small or long thin objects are identified. Through the seed selection and local reliable information propagation, the problem of noisy labels is solved elegantly. Experiments on three real hyper spectral data sets with different spatial resolution, spectral resolution and land covers demonstrate the effectiveness of our method. Haichang Li, Jiangyong Duan, Shiming Xiang, Lingfeng Wang 0002, Chunhong Pan |
ICPR | 5 |
| 2014 | Cross Modal Deep Model and Gaussian Process Based Model for MSR-Bing ChallengeabstractIn the MSR-Bing Image Retrieval Challenge, the contestants are required to design a system that can score the query-image pairs based on the relevance between queries and images. To address this problem, we propose a regression based cross modal deep learning model and a Gaussian Process scoring model. The regression based cross modal deep learning model takes the image features and query features as inputs respectively and outputs the relevance scores directly. The Gaussian Process scoring model regards the challenge as a ranking problem and utilizes the click (or pseudo click) information from both the training set and the development set to predict the relevance scores. The proposed models are used in different situations: matched and miss-matched queries. Experiments on the development set show the effectiveness of the proposed models. Jian Wang 0068, Cuicui Kang, Yonghao He, Shiming Xiang, Chunhong Pan |
ACM Multimedia | 5 |
| 2014 | Cooperative fusion particle filter tracker
Lingfeng Wang 0002, Hongping Yan, Chunhong Pan |
Sci. China Inf. Sci. | 3 |
| 2014 | Multifocus image fusion via focus segmentation and region reconstruction
Jiangyong Duan, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Neurocomputing | 4 |
| 2014 | Kernel sparse representation with pixel-level and region-level local feature kernels for face recognition
Cuicui Kang, Shengcai Liao, Shiming Xiang, Chunhong Pan |
Neurocomputing | 4 |
| 2014 | Nonrigid medical image registration with locally linear reconstruction
Lingfeng Wang 0002, Chunhong Pan |
Neurocomputing | 2 |
| 2014 | Vehicle Detection in Satellite Images by Hybrid Deep Convolutional Neural NetworksabstractDetecting small objects such as vehicles in satellite images is a difficult problem. Many features (such as histogram of oriented gradient, local binary pattern, scale-invariant feature transform, etc.) have been used to improve the performance of object detection, but mostly in simple environments such as those on roads. Kembhavi et al. proposed that no satisfactory accuracy has been achieved in complex environments such as the City of San Francisco. Deep convolutional neural networks (DNNs) can learn rich features from the training data automatically and has achieved state-of-the-art performance in many image classification databases. Though the DNN has shown robustness to distortion, it only extracts features of the same scale, and hence is insufficient to tolerate large-scale variance of object. In this letter, we present a hybrid DNN (HDNN), by dividing the maps of the last convolutional layer and the max-pooling layer of DNN into multiple blocks of variable receptive field sizes or max-pooling field sizes, to enable the HDNN to extract variable-scale features. Comparative experimental results indicate that our proposed HDNN significantly outperforms the traditional DNN on vehicle detection. Xueyun Chen, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2014 | Change Field: A New Change Measure for VHR ImagesabstractDue to the complexity of very high resolution (VHR) images and the inaccurate correspondence, change feature extraction is the key difficulty of VHR image change detection. In this letter, change field is proposed to represent the complex changes between VHR images. Change field measures the complex changes based on the displacements and the compensated distance. Based on change field, a novel change detection approach is proposed, where the inter-class variability is improved and the changed class and the unchanged class can be separated effectively. Experiments demonstrate the effectiveness of the proposed approach. Leigang Huo, Xiangchu Feng, Chunlei Huo, Zhixin Zhou, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2014 | Robust level set image segmentation via a local correntropy-based K-means clustering
Lingfeng Wang 0002, Chunhong Pan |
Pattern Recognit. | 2 |
| 2014 | Visual Tracking Via Kernel Sparse Representation With Multikernel FusionabstractIt remains a challenging task to track an object robustly due to factors such as pose variation, illumination change, occlusion, and background clutter. In the past decades, a number of researchers have been attracted to tackling these difficulties, and they proposed many effective methods. Among them, sparse representation-based tracking method is a promising. While much success has been demonstrated, there are several issues that still need to be addressed. First, the introduction to trivial occlusion templates brings a high computational cost of this method. Second, the utilization of raw template object representation makes this method difficult to adopt sophisticated object features. To solve these problems, we consider the sparse representation problem in a kernel space and propose a kernel sparse representation (KSR)-based tracking algorithm. Under the kernel representation, it is not necessary to introduce trivial occlusion templates in order to reduce the computational cost. Furthermore, multikernel fusion allows our method to use multiple sophisticated object features, such as spatial color histogram and spatial gradient-orientation histogram, and let these features complement each other during the tracking process. Comparative experiments on challenging scenes demonstrate that our KSR-based tracking algorithm outperforms the state-of-the-art approaches in tracking accuracy. Lingfeng Wang 0002, Hongping Yan, Chunhong Pan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | Receptive Fields Selection for Binary Feature DescriptionabstractFeature description for local image patch is widely used in computer vision. While the conventional way to design local descriptor is based on expert experience and knowledge, learning-based methods for designing local descriptor become more and more popular because of their good performance and data-driven property. This paper proposes a novel data-driven method for designing binary feature descriptor, which we call receptive fields descriptor (RFD). Technically, RFD is constructed by thresholding responses of a set of receptive fields, which are selected from a large number of candidates according to their distinctiveness and correlations in a greedy way. Using two different kinds of receptive fields (namely rectangular pooling area and Gaussian pooling area) for selection, we obtain two binary descriptors RFDR and RFDG .accordingly. Image matching experiments on the well-known patch data set and Oxford data set demonstrate that RFD significantly outperforms the state-of-the-art binary descriptors, and is comparable with the best float-valued descriptors at a fraction of processing time. Finally, experiments on object recognition tasks confirm that both RFDR and RFDG successfully bridge the performance gap between binary descriptors and their floating-point competitors. Bin Fan 0001, Qingqun Kong, Tomasz Trzcinski, Zhiheng Wang 0001, Chunhong Pan, Pascal Fua |
IEEE Trans. Image Process. | 5 |
| 2014 | Fast Image Upsampling via the Displacement FieldabstractIn this paper, we present a fast image upsampling method within a two-scale framework to ensure the sharp construction of upsampled image for both large-scale edges and small-scale structures. In our approach, the low-frequency image is recovered via a novel sharpness preserving interpolation technique based on a well-constructed displacement field, which is estimated by a cross-resolution sharpness preserving model. Within this model, the distances of pixels on edges are preserved, which enables the recovery of sharp edges in the high-resolution result. Likewise, local high-frequency structures are reconstructed via a sharpness preserving reconstruction algorithm. Extensive experiments show that our method outperforms current state-of-the-art approaches, based on quantitative and qualitative evaluations, as well as perceptual evaluation by a user study. Moreover, our approach is very fast so as to be practical for real applications. Lingfeng Wang 0002, Huai-Yu Wu, Chunhong Pan |
IEEE Trans. Image Process. | 3 |
| 2014 | Spectral Unmixing via Data-Guided SparsityabstractHyperspectral unmixing, the process of estimating a common set of spectral bases and their corresponding composite percentages at each pixel, is an important task for hyperspectral analysis, visualization, and understanding. From an unsupervised learning perspective, this problem is very challenging-both the spectral bases and their composite percentages are unknown, making the solution space too large. To reduce the solution space, priors. In practice, these priors would easily lead to some unsuitable solution. This is because they are achieved by applying an identical strength of constraints to all the factors, which does not hold in practice. To overcome this limitation, we propose a novel sparsity-based method by learning a data-guided map (DgMap) to describe the individual mixed level of each pixel. Through this DgMap, the l(p) (0 < p < 1) constraint is applied in an adaptive manner. Such implementation not only meets the practical situation, but also guides the spectral bases toward the pixels under highly sparse constraint. What is more, an elegant optimization scheme as well as its convergence proof have been provided in this paper. Extensive experiments on several datasets also demonstrate that the DgMap is feasible, and high quality unmixing results could be obtained by our method. Feiyun Zhu, Ying Wang 0008, Bin Fan 0001, Shiming Xiang, Gaofeng Meng, Chunhong Pan |
IEEE Trans. Image Process. | 6 |
| 2013 | Indoor frame recovering via line segments refinement and votingabstractFrame structure estimation from line segments is an important yet challenging problem in understanding indoor scenes. In practice, line segment extraction can be affected by occlusions, illumination variations, and weak object boundaries. To address this problem, an approach for frame structure recovery based on line segment refinement and voting is proposed. We refined line segments by the revising, connecting, and adding operations. We then propose an iterative voting mechanism for selecting refined line segments, where a cross ratio constraint is enforced to build crab-like models. Our algorithm outperforms state-of-the-art approaches, especially when considering complex indoor scenes. Luanzheng Guo, Lingfeng Wang 0002, Chunhong Pan, Shiming Xiang |
ICASSP | 4 |
| 2013 | VHR image change detection based on discriminative dictionary learningabstractThe difficulty of Very High Resolution (VHR) image change detection is mainly due to the low separability between the changed and unchanged class. The traditional approaches usually address the problem by solving the feature extraction and classification separately, which cannot ensure that the classification algorithm makes the best use of the features. Considering this, we propose a novel approach that combines the feature extraction and the classification task by utilizing the sparse representation algorithm with discriminative dictionary. Experiments on real data sets show that our method achieves effective results. Kun Ding 0001, Chunlei Huo, Chunhong Pan |
ICASSP | 4 |
| 2013 | Removing out-of-focus blur from similar image pairsabstractThis paper presents a new deblurring method to remove the out-of-focus blur from similar image pairs. The method is motivated by an observation that a blurred structure appearing in one image can often have its corresponding clear one in the similar clear images. Our method first extracts the patch pairs from input images by SIFT matching. Then the constraints on the patch pairs are used to estimate the blur kernel via the RANSAC algorithm. Finally, the non-blind deconvolution is adopted to restore the blurred image. The main advantage is that we can improve the deblurring results with the help of additional similar clear images in many practical applications. Our method is validated on synthetic and real images by comparing with state-of-the-art methods. Jiangyong Duan, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICASSP | 4 |
| 2013 | Learning weighted Hamming distance for binary descriptorsabstractLocal image descriptors are one of the key components in many computer vision applications. Recently, binary descriptors have received increasing interest of the community for its efficiency and low memory cost. The similarity of binary descriptors is measured by Hamming distance which has equal emphasis on all elements of binary descriptors. This paper improves the performance of binary descriptors by learning a weighted Hamming distance for binary descriptors with larger weights assigned to more discriminative elements. What is more, the weighted Hamming distance can be computed as fast as the Hamming distance on the basis of a pre-computed look-up-table. Therefore, the proposed method improves the matching performance of binary descriptors without sacrificing matching speed. Experimental results on two popular binary descriptors (BRIEF [1] and FREAK [2]) validate the effectiveness of the proposed method. Bin Fan 0001, Qingqun Kong, Xiao-Tong Yuan, Zhiheng Wang 0001, Chunhong Pan |
ICASSP | 5 |
| 2013 | Robust VHR image change detection based on local features and multi-scale fusionabstractUrban change detection of Very High Resolution (VHR) remote sensing images is challenging, due to the ill-posed nature of change detection problem, the inherent nature of VHR image, the complex morphology of urban scenes, etc. To address the above difficulties, a robust approach is proposed, which is based on discriminative local features, robust distance metric and novel multi-scale fusion strategy. By integrating these components synergistically, the proposed approach is superior to the traditional approaches in capturing semantic changes and removing the false changes. Comparative experiments demonstrate the effectiveness and advantages of the proposed approach. Chunlei Huo, Shiming Xiang, Chunhong Pan |
ICASSP | 4 |
| 2013 | Active learning based automatic face segmentation for kinect videoabstractThis paper presents a novel segmentation approach for extracting faces from videos. Under an active learning framework, the segmentation is conducted automatically without human interactions. A small portion of pixels are first labeled as face or non-face. Given these labeled samples, a semi-supervised spline regression model is then applied to obtain the face region. Based on the segmentation result, new pixels are selected and labeled. These two steps perform iterately until convergence. The main novelty is that color and depth data are combined to provide the labeling information. Our approach is validated via comparisons with state-of-the-art methods on real videos captured from the commodity Kinect camera. Jixia Zhang, Shaoguo Liu, Franck Davoine, Chunhong Pan, Shiming Xiang |
ICASSP | 5 |
| 2013 | Efficient Image Dehazing with Boundary Constraint and Contextual RegularizationabstractImages captured in foggy weather conditions often suffer from bad visibility. In this paper, we propose an efficient regularization method to remove hazes from a single input image. Our method benefits much from an exploration on the inherent boundary constraint on the transmission function. This constraint, combined with a weighted L_1-norm based contextual regularization, is modeled into an optimization problem to estimate the unknown scene transmission. A quite efficient algorithm based on variable splitting is also presented to solve the problem. The proposed method requires only a few general assumptions and can restore a high-quality haze-free image with faithful colors and fine image details. Experimental results on a variety of haze images demonstrate the effectiveness and efficiency of the proposed method. Gaofeng Meng, Ying Wang 0008, Jiangyong Duan, Shiming Xiang, Chunhong Pan |
ICCV | 5 |
| 2013 | Group sparsity based semi-supervised band selection for hyperspectral imagesabstractIn this paper, we propose a novel group sparsity based semi-supervised band selection method. There are three key features in our method. First, it fulfills the band selection task by employing group sparsity on the regression coefficients in a robust linear regression for classification model, so that the selected bands hold lower classification errors. Second, the spatial smoothness prior is incorporated to preserve the similarity of spatial neighbors in band selection. Third, the objective function is efficiently optimized via an alternative iteration algorithm. Comparative results on two hyper-spectral data sets validate the effectiveness of our method, showing higher classification accuracies. Haichang Li, Ying Wang 0008, Jiangyong Duan, Shiming Xiang, Chunhong Pan |
ICIP | 5 |
| 2013 | Kinect depth restoration via energy minimization with TV21 regularizationabstractDepth maps generated by Kinect cameras often contain a significant amount of missing pixels and strong noise, limiting their usability in many computer vision applications. We present a new energy minimization method to fill the missing regions and remove noise in a depth map, by exploiting the strong correlation between color and depth values in local image neighborhoods. To preserve sharp edges and remove noise from the depth map, we propose to add a TV21regularization term into the energy function. Finally, we show how to effectively minimize the total energy using an alternating optimization approach. Experimental results show that the proposed method outperforms commonly-used depth inpainting approaches. Shaoguo Liu, Ying Wang 0008, Jue Wang 0001, Jixia Zhang, Chunhong Pan |
ICIP | 6 |
| 2013 | Maximum correntropy criterion based 3D head tracking with commodity depth cameraabstract3D head tracking becomes easier with the depth image from Microsoft Kinect. However, the noise from face occlusion and illumination still affects the tracking quality. In this paper, we introduce the robust Maximum Correntropy Criterion (MCC) to the problem of 3D head tracking, to tackle these noises. Fortunately, MCC can handle arbitrarily distributed noises. To solve the MCC based cost function, we develop an effective two-stage optimization scheme with the half-quadric technology. A head tracking system that uses Miscrosoft Kinect is also developed based on the MCC formulation. The system is fully automatic and online, without need of offline training. Experimental results show that the system is very robust against partial occlusion, large motion and sudden illumination variations. Shaoguo Liu, Ying Wang 0008, Jixia Zhang, Chunhong Pan |
ICIP | 5 |
| 2013 | An integrated graph-based face segmentation approach from Kinect videosabstractIn this paper, we present an Integrated Semi-Supervised Graph (IntSSG) approach to automatically segment face from color-depth video. In the first step, IntSSG performs skin color detection and online SIFT matching to initialize some face and non-face pixels. Then, the labels of these pixels are refined by conducting adaptive depth thresholding. Finally, based on a semi-supervised graph framework, IntSSG segments face by propagating the refined labels to other pixels. Experimental results show that IntSSG is able to accurately segment faces in difficult situations such as large pose changes and illumination variations. Jixia Zhang, Shaoguo Liu, Jiangyong Duan, Ying Wang 0008, Chunhong Pan |
ICIP | 6 |
| 2013 | A unified level set framework utilizing parameter priors for medical image segmentation
Lingfeng Wang 0002, Zeyun Yu, Chunhong Pan |
Sci. China Inf. Sci. | 3 |
| 2013 | Registration of Optical and SAR Satellite Images by Exploring the Spatial Relationship of the Improved SIFTabstractAlthough feature-based methods have been successfully developed in the past decades for the registration of optical images, the registration of optical and synthetic aperture radar (SAR) images is still a challenging problem in remote sensing. In this letter, an improved version of the scale-invariant feature transform is first proposed to obtain initial matching features from optical and SAR images. Then, the initial matching features are refined by exploring their spatial relationship. The refined feature matches are finally used for estimating registration parameters. Experimental results have shown the effectiveness of the proposed method. Bin Fan 0001, Chunlei Huo, Chunhong Pan, Qingqun Kong |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2013 | Nonparametric Illumination Correction for Scanned Document Images via Convex HullsabstractA scanned image of an opened book page often suffers from various scanning artifacts known as scanning shading and dark borders noises. These artifacts will degrade the qualities of the scanned images and cause many problems to the subsequent process of document image analysis. In this paper, we propose an effective method to rectify these scanning artifacts. Our method comes from two observations: that the shading surface of most scanned book pages is quasi-concave and that the document contents are usually printed on a sheet of plain and bright paper. Based on these observations, a shading image can be accurately extracted via convex hulls-based image reconstruction. The proposed method proves to be surprisingly effective for image shading correction and dark borders removal. It can restore a desired shading-free image and meanwhile yield an illumination surface of high quality. More importantly, the proposed method is nonparametric and thus does not involve any user interactions or parameter fine-tuning. This would make it very appealing to nonexpert users in applications. Extensive experiments based on synthetic and real-scanned document images demonstrate the efficiency of the proposed method. Gaofeng Meng, Shiming Xiang, Nanning Zheng 0001, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2013 | Level set evolution with locally linear classification for image segmentation
Ying Wang 0008, Shiming Xiang, Chunhong Pan, Lingfeng Wang 0002, Gaofeng Meng |
Pattern Recognit. | 3 |
| 2013 | Region-based image segmentation with local signed difference energy
Lingfeng Wang 0002, Huai-Yu Wu, Chunhong Pan |
Pattern Recognit. Lett. | 3 |
| 2013 | Edge-Directed Single-Image Super-Resolution Via Adaptive Gradient Magnitude Self-InterpolationabstractSuper-resolution from a single image plays an important role in many computer vision systems. However, it is still a challenging task, especially in preserving local edge structures. To construct high-resolution images while preserving the sharp edges, an effective edge-directed super-resolution method is presented in this paper. An adaptive self-interpolation algorithm is first proposed to estimate a sharp high-resolution gradient field directly from the input low-resolution image. The obtained high-resolution gradient is then regarded as a gradient constraint or an edge-preserving constraint to reconstruct the high-resolution image. Extensive results have shown both qualitatively and quantitatively that the proposed method can produce convincing super-resolution images containing complex and sharp features, as compared with the other state-of-the-art super-resolution algorithms. Lingfeng Wang 0002, Shiming Xiang, Gaofeng Meng, Huai-Yu Wu, Chunhong Pan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2013 | A Graph-Based Classification Method for Hyperspectral ImagesabstractThe goal of this paper is to apply graph cut (GC) theory to the classification of hyperspectral remote sensing images. The task is formulated as a labeling problem on Markov random field (MRF) constructed on the image grid, and GC algorithm is employed to solve this task. In general, a large number of user interactive strikes are necessary to obtain satisfactory segmentation results. Due to the spatial variability of spectral signatures, however, hyperspectral remote sensing images often contain many tiny regions. Labeling all these tiny regions usually needs expensive human labor. To overcome this difficulty, a pixelwise fuzzy classification based on support vector machine (SVM) is first applied. As a result, only pixels with high probabilities are preserved as labeled ones. This generates a pseudouser strike map. This map is then employed for GC to evaluate the truthful likelihoods of class labels and propagate them to the MRF. To evaluate the robustness of our method, we have tested our method on both large and small training sets. Additionally, comparisons are made between the results of SVM, SVM with stacking neighboring vectors, SVM with morphological preprocessing, extraction and classification of homogeneous objects, and our method. Comparative experimental results demonstrate the validity of our method. Shiming Xiang, Chunhong Pan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2013 | Forward-Backward Mean-Shift for Visual Tracking With Local-Background-Weighted HistogramabstractObject tracking plays an important role in many intelligent transportation systems. Unfortunately, it remains a challenging task due to factors such as occlusion and target-appearance variation. In this paper, we present a new tracking algorithm to tackle the difficulties caused by these two factors. First, considering the target-appearance variation, we introduce the local-background-weighted histogram (LBWH) to describe the target. In our LBWH, the local background is treated as the context of the target representation. Compared with traditional descriptors, the LBWH is more robust to the variability or the clutter of the potential background. Second, to deal with the occlusion case, a new forward-backward mean-shift (FBMS) algorithm is proposed by incorporating a forward-backward evaluation scheme, in which the tracking result is evaluated by the forward-backward error. Extensive experiments on various scenarios have demonstrated that our tracking algorithm outperforms the state-of-the-art approaches in tracking accuracy. Lingfeng Wang 0002, Hongping Yan, Huai-Yu Wu, Chunhong Pan |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2012 | Shadow-Free TILT for Facade Rectification
Lumei Li, Hongping Yan, Lingfeng Wang 0002, Chunhong Pan |
ACCV (4) | 4 |
| 2012 | Image Guided Tone Mapping with Locally Nonlinear Model
Huxiang Gu, Ying Wang 0008, Shiming Xiang, Gaofeng Meng, Chunhong Pan |
ECCV (4) | 5 |
| 2012 | Classification oriented semi-supervised band selection for hyperspectral images
Shiming Xiang, Chunhong Pan |
ICPR | 3 |
| 2012 | Kernel Homotopy based sparse representation for object classification
Cuicui Kang, Shengcai Liao, Shiming Xiang, Chunhong Pan |
ICPR | 4 |
| 2012 | Skin detection via linear regression tree
Jixia Zhang, Franck Davoine, Chunhong Pan |
ICPR | 4 |
| 2012 | Multilevel SIFT Matching for Large-Size VHR Image RegistrationabstractA fast approach is proposed in this letter for large-size very high resolution image registration, which is accomplished based on coarse-to-fine strategy and blockwise scale-invariant feature transform (SIFT) matching. Coarse registration is implemented at low resolution level, which provides a geometric constraint. The constraint makes the blockwise SIFT matching possible and is helpful for getting more matched keypoints at the latter refined procedure. Refined registration is achieved by blockwise SIFT matching and global optimization on the whole matched keypoints based on iterative reweighted least squares. To improve the efficiency, blockwise SIFT matching is implemented in a parallel manner. Experiments demonstrate the effectiveness of the proposed approach. Chunlei Huo, Chunhong Pan, Leigang Huo, Zhixin Zhou |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2012 | Metric Rectification of Curved Document ImagesabstractIn this paper, we propose a metric rectification method to restore an image from a single camera-captured document image. The core idea is to construct an isometric image mesh by exploiting the geometry of page surface and camera. Our method uses a general cylindrical surface (GCS) to model the curved page shape. Under a few proper assumptions, the printed horizontal text lines are shown to be line convergent symmetric. This property is then used to constrain the estimation of various model parameters under perspective projection. We also introduce a paraperspective projection to approximate the nonlinear perspective projection. A set of close-form formulas is thus derived for the estimate of GCS directrix and document aspect ratio. Our method provides a straightforward framework for image metric rectification. It is insensitive to camera positions, viewing angles, and the shapes of document pages. To evaluate the proposed method, we implemented comprehensive experiments on both synthetic and real-captured images. The results demonstrate the efficiency of our method. We also carried out a comparative experiment on the public CBDAR2007 data set. The experimental results show that our method outperforms the state-of-the-art methods in terms of OCR accuracy and rectification errors. Gaofeng Meng, Chunhong Pan, Shiming Xiang, Jiangyong Duan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Image deblurring with matrix regression and gradient evolution
Shiming Xiang, Gaofeng Meng, Ying Wang 0008, Chunhong Pan, Changshui Zhang |
Pattern Recognit. | 4 |
| 2012 | 3-D Head Tracking via Invariant Keypoint LearningabstractKeypoint matching is a standard tool to solve the correspondence problem in vision applications. However, in 3-D face tracking, this approach is often deficient because the human face complexities, together with its rich viewpoint, nonrigid expression, and lighting variations in typical applications, can cause many variations impossible to handle by existing keypoint detectors and descriptors. In this paper, we propose a new approach to tailor keypoint matching to track the 3-D pose of the user head in a video stream. The core idea is to learn keypoints that are explicitly invariant to these challenging transformations. First, we select keypoints that are stable under randomly drawn small viewpoints, nonrigid deformations, and illumination changes. Then, we treat keypoint descriptor learning at different large angles as an incremental scheme to learn discriminative descriptors. At matching time, to reduce the ratio of outlier correspondences, we use second-order color information to prune keypoints unlikely to lie on the face. Moreover, we integrate optical flow correspondences in an adaptive way to remove motion jitter efficiently. Extensive experiments show that the proposed approach can lead to fast, robust, and accurate 3-D head tracking results even under very challenging scenarios. Franck Davoine, Vincent Lepetit, Christophe Chaillou, Chunhong Pan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2012 | Discriminative Least Squares Regression for Multiclass Classification and Feature SelectionabstractThis paper presents a framework of discriminative least squares regression (LSR) for multiclass classification and feature selection. The core idea is to enlarge the distance between different classes under the conceptual framework of LSR. First, a technique called ε-dragging is introduced to force the regression targets of different classes moving along opposite directions such that the distances between classes can be enlarged. Then, the ε-draggings are integrated into the LSR model for multiclass classification. Our learning framework, referred to as discriminative LSR, has a compact model form, where there is no need to train two-class machines that are independent of each other. With its compact form, this model can be naturally extended for feature selection. This goal is achieved in terms of L2,1 norm of matrix, generating a sparse learning model for feature selection. The model for multiclass classification and its extension for feature selection are finally solved elegantly and efficiently. Experimental evaluation over a range of benchmark datasets indicates the validity of our method. Shiming Xiang, Feiping Nie 0001, Gaofeng Meng, Chunhong Pan, Changshui Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2011 | Effective multi-resolution background subtractionabstractIn this paper, we propose a novel multi-resolution background sub traction method. We adopt coarse to fine strategy, which is the essence the multi-resolution scheme, to obtain the foreground mask. The rough mask is first gained relied on the Single Gaussian Model, which holds minor computation cost. Then, the slightly accuracy mask is calculated by the Saliency-based Extraction Model, which contains high accuracy and stability. Finally, Contour-based Refining Model is used to refine the mask edge. Our algorithm is evaluated against several video sequences, and experimental results show that the proposed method is suitable for various scenes and is appealing with respect to robustness. Lingfeng Wang 0002, Chunhong Pan |
ICASSP | 2 |
| 2011 | Image editing based on Sparse Matrix-Vector multiplicationabstractThis paper presents a unified model for image editing in terms of Sparse Matrix-Vector (SpMV) multiplication. In our framework, we cast image editing as a linear energy minimization problem and address it by solving a sparse linear system, which is able to yield a globally optimal solution. First, three classical image editing operations, including linear filtering, resizing and selecting, are reformulated in the SpMV multiplication form. The SpMV form helps us set up a straightforward mechanism to flexibly and naturally combine various image features (low-level visual features or geometrical features) and constraints together into an integrated energy minimization function under the L2norm. Then, we apply our model to implement the tasks of pan-sharpening, image cloning, image mixed editing and texture transfer, which are now popularly used in the field of digital art. Comparative experiments are reported to validate the effectiveness and efficiency of our model. Ying Wang 0008, Hongping Yan, Chunhong Pan, Shiming Xiang |
ICASSP | 3 |
| 2011 | Kernel sparse representation with local patterns for face recognitionabstractIn this paper we propose a novel kernel sparse representation classification (SRC) framework and utilize the local binary pattern (LBP) descriptor in this framework for robust face recognition. First we develop a kernel coordinate descent (KCD) algorithm for 11 minimization in the kernel space, which is based on the covariance update technique. Then we extract LBP descriptors from each image and apply two types of kernels (χ2distance based and Hamming distance based) with the proposed KCD algorithm under the SRC framework for face recognition. Experiments on both the Extended Yale B and the PIE face databases show that the proposed method is more robust against noise, occlusion, and illumination variations, even with small number of training samples. Cuicui Kang, Shengcai Liao, Shiming Xiang, Chunhong Pan |
ICIP | 4 |
| 2011 | Robust airplane detection in satellite imagesabstractAutomatic target detection in satellite images remains a challenging problem. The main difficulties lie in the cooccurrence of variations of target type, pose, and size in huge satellite image. In this paper, we propose a new airplane detection approach based on visual saliency computation and symmetry detection. The advantages are twofold. First, saliency and symmetry detection perform stably in obtaining target location and orientation information. Second, independent of target type, pose and size, saliency map and symmetry detection are computed only once. This saves a large amount of computational time but does not miss any targets. Experiments show that our method provides a promising way to detect airplanes in complex airport scenes. Shiming Xiang, Chunhong Pan |
ICIP | 4 |
| 2011 | Softferns for homography estimationabstractFerns was recently proposed to address the problem of fast keypoint matching [1]. In estimating homography, it often yields many similar matching scores under large viewpoint change or when the object is highly repetitive. However, we observe that keypoints having similar scores are usually far away from each other. Based on this fact, we propose a Softferns approach, in which Ferns and homography can alternatively refine each other. Extensive experiments prove that Softferns is a very general improvement over Ferns. Shaoguo Liu, Jixia Zhang, Franck Davoine, Chunhong Pan |
ICIP | 5 |
| 2011 | Hierarchical fusion of descriptor matching and L-K optical flowabstractDescriptor matching is recently used in optical flow to handle large motion displacements. While doing so is a success, motion jitter often arises from lack of temporal consistency. However, when accounting for the consistency, it may violate our expectation of handling arbitrary motion shifts. In this paper, we propose a new approach to remove this controversy. At its core is a hierarchical fusion of descriptor matching and optical flow to determine the true local motion displacements. Applications to estimating homography and 3D head pose verify that the approach can well adapt to either large or small motion displacements. Chunhong Pan, Franck Davoine, Shaoguo Liu |
ICIP | 2 |
| 2011 | MEAN-shift tracking algorithm with weight fusion strategyabstractIn this paper, we propose a new Mean-shift algorithm to tackle some tracking difficulties, such as background clutter and partial occlusion. First, we compare all Mean-shift-like tracking algorithms, and indicate that the main difference among them is weight calculation. Then, a new fusion strategy is proposed to unify all weight calculation methods into a framework. Based on this framework, we propose a novel weight calculation method, which takes the candidate model into consideration as well as incorporates the local background. Extensive experiments are conducted to evaluate the proposed approach. Comparative experimental results indicate that the tracking accuracy is improved as compared with the state-of-the-arts. Lingfeng Wang 0002, Chunhong Pan, Shiming Xiang |
ICIP | 2 |
| 2011 | Level set evolution with locally linear classification for image segmentationabstractThis paper presents a novel local region-based level set model for image segmentation. In each local region, we define a locally weighted least squares energy to fit a linear classification function. The local energy is then integrated over the entire image domain to form an energy functional in terms of level set function. The energy minimization is achieved by level set evolution and estimation of parameters of the locally linear function in an iterative process. By introducing the locally linear functions to separate background and foreground in local regions, our model not only ensures the accuracy of the segmentation results, but also be very robust to initialization. Experiments are reported to demonstrate the effectiveness and efficiency of our model. Ying Wang 0008, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
ICIP | 4 |
| 2011 | Interactive blood-coil simulation in real-time during aneurysm embolization
Yiyi Wei, Stephane Cotin, Jérémie Allard, Chunhong Pan, Songde Ma |
Comput. Graph. | 5 |
| 2011 | Interactive Image Segmentation With Multiple Linear Reconstructions in WindowsabstractThis paper proposes an algorithm for interactive image segmentation. The task is formulated as a problem of graph-based transductive classification. Specifically, given an image window, the color of each pixel in it will be reconstructed linearly with those of the remaining pixels in this window. The optimal reconstruction weights will be kept unchanged to linearly reconstruct their class labels. The label reconstruction errors are estimated in each window. These errors are further collected together to develop a learning model. Then, the class information about the user specified foreground and background pixels are integrated into a regularization framework. Under this framework, a globally optimal labeling is finally obtained. The computational complexity is analyzed, and an approach for speeding up the algorithm is presented. Comparative experimental results illustrate the validity of our algorithm. Shiming Xiang, Chunhong Pan, Feiping Nie 0001, Changshui Zhang |
IEEE Trans. Multim. | 2 |
| 2011 | Regression Reformulations of LLE and LTSA With Locally Linear TransformationabstractLocally linear embedding (LLE) and local tangent space alignment (LTSA) are two fundamental algorithms in manifold learning. Both LLE and LTSA employ linear methods to achieve their goals but with different motivations and formulations. LLE is developed by locally linear reconstructions in both high- and low-dimensional spaces, while LTSA is developed with the combinations of tangent space projections and locally linear alignments. This paper gives the regression reformulations of the LLE and LTSA algorithms in terms of locally linear transformations. The reformulations can help us to bridge them together, with which both of them can be addressed into a unified framework. Under this framework, the connections and differences between LLE and LTSA are explained. Illuminated by the connections and differences, an improved LLE algorithm is presented in this paper. Our algorithm learns the manifold in way of LLE but can significantly improve the performance. Experiments are conducted to illustrate this fact. Shiming Xiang, Feiping Nie 0001, Chunhong Pan, Changshui Zhang |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2011 | Partwise Cross-Parameterization via Nonregular Convex Hull DomainsabstractIn this paper, we propose a novel partwise framework for cross-parameterization between 3D mesh models. Unlike most existing methods that use regular parameterization domains, our framework uses nonregular approximation domains to build the cross-parameterization. Once the nonregular approximation domains are constructed for 3D models, different (and complex) input shapes are transformed into similar (and simple) shapes, thus facilitating the cross-parameterization process. Specifically, a novel nonregular domain, the convex hull, is adopted to build shape correspondence. We first construct convex hulls for each part of the segmented model, and then adopt our convex-hull cross-parameterization method to generate compatible meshes. Our method exploits properties of the convex hull, e.g., good approximation ability and linear convex representation for interior vertices. After building an initial cross-parameterization via convex-hull domains, we use compatible remeshing algorithms to achieve an accurate approximation of the target geometry and to ensure a complete surface matching. Experimental results show that the compatible meshes constructed are well suited for shape blending and other geometric applications. Huai-Yu Wu, Chunhong Pan, Hongbin Zha, Qing Yang 0002, Songde Ma |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2010 | Adaptive εLBP for Background Subtraction
Lingfeng Wang 0002, Huai-Yu Wu, Chunhong Pan |
ACCV (3) | 3 |
| 2010 | Medical Image Segmentation Based on Novel Local Order Energy
Lingfeng Wang 0002, Zeyun Yu, Chunhong Pan |
ACCV (2) | 3 |
| 2010 | Real-time object tracking based on the relative hist model within particle filter framework
Lingfeng Wang 0002, Chunhong Pan |
ICASSP | 2 |
| 2010 | Fast and Effective Background Subtraction Based on ELBP
Lingfeng Wang 0002, Chunhong Pan |
ICASSP | 2 |
| 2010 | Model Transduction for Triangle Meshes
Huai-Yu Wu, Chunhong Pan, Hongbin Zha, Songde Ma |
J. Comput. Sci. Technol. | 2 |
| 2010 | Fast Object-Level Change Detection for VHR ImagesabstractA novel approach is presented for change detection of very high resolution images, which is accomplished by fast object-level change feature extraction and progressive change feature classification. Object-level change feature is helpful for improving the discriminability between the changed class and the unchanged class. Progressive change feature classification helps improve the accuracy and the degree of automation, which is implemented by dynamically adjusting the training samples and gradually tuning the separating hyperplane. Experiments demonstrate the effectiveness of the proposed approach. Chunlei Huo, Zhixin Zhou, Hanqing Lu, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2010 | Skew Estimation of Document Images Using BaggingabstractThis paper proposes a general-purpose method for estimating the skew angles of document images. Rather than to derive a skew angle merely from text lines, the proposed method exploits various types of visual cues of image skew available in local image regions. The visual cues are extracted by Radon transform and then outliers of them are iteratively rejected through a floating cascade. A bagging (bootstrap aggregating) estimator is finally employed to combine the estimations on the local image blocks. Our experimental results show significant improvements against the state-of-the-art methods, in terms of execution speed and estimation accuracy, as well as the robustness to short and sparse text lines, multiple different skews and the presence of nontextual objects of various types and quantities. Gaofeng Meng, Chunhong Pan, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2010 | TurboPixel Segmentation Using Eigen-ImagesabstractTurboPixel (TP) is a powerful tool for image over-segmentation. It is fast and can yield a lattice-like structure of superpixel regions with uniform size. This paper presents a method to learn eigen-images from the image to be segmented. Such eigen-images are used to generate the evolution speed in the TP framework. The task is formulated as a problem of pixel clustering. Specifically, for the pixels in each local window, a linear transformation is introduced to map their color vectors to be the cluster indicator vectors. The errors under all such linear transformations are estimated and summed together to obtain an objective function, from which a global optimum is finally obtained. In this process, the eigen-images are constructed. Based on these eigen-images, multidimensional image gradient operator is defined to evaluate the gradient, which is supplied to the TP algorithm to obtain the final superpixel segmentations. The computational issues are discussed, and an image pyramid is introduced to speed up the computation. Comparative experiments illustrate the effectiveness of our method. Shiming Xiang, Chunhong Pan, Feiping Nie 0001, Changshui Zhang |
IEEE Trans. Image Process. | 2 |
| 2009 | Tracking Eye Gaze under Coordinated Head Rotations with an Ordinary Camera
Chunhong Pan, Christophe Chaillou |
ACCV (2) | 2 |
| 2009 | Mean-Shift Object Tracking with a Novel Back-Projection Calculation Method
Lingfeng Wang 0002, Huai-Yu Wu, Chunhong Pan |
ACCV (1) | 3 |
| 2009 | Saliency-based automatic target detection in forward looking infrared imagesabstractA saliency-based target detection method for forward looking infrared (FLIR) image is proposed. Firstly, saliency map is computed using scale-space representation and separated into dark saliency map (DSM) and bright saliency map (BSM). Secondly, dark and bright regions of interest (ROI) are detected by respective type of saliency map using marker-based maximally stable extremal regions (MSER) detection algorithm. Finally, shape matching algorithm is applied after grouping of the two types of ROI for object detection. Experimental results show that this work provides a promising way to solve the problems caused by salient dark parts of target. Chunhong Pan, Li-xiong Liu |
ICIP | 2 |
| 2009 | Toward Real-Time Simulation of Blood-Coil Interaction during Aneurysm Embolization
Yiyi Wei, Stephane Cotin, Jérémie Allard, Chunhong Pan, Songde Ma |
MICCAI (1) | 5 |
| 2008 | Multimodal preserving embedding for face recognitionabstractTraditional dimension reduction approaches always consider the samples in a class are uni-modal. In real world, samples in a class are usually multi-modal, for instance, the manifold of the facial appearance of a person under different illumination, expression, and poses is multi-modal. Recently, dimension reduction approaches based on manifold learning are presented, the main purpose is to preserve the manifold structure on low dimensional space. In this paper, by analyzing the manifold learning methods and the traditional dimension reduction methods, we show that most of these methods can be summarized into a general framework. Based on this framework, we propose a novel dimension reduction approach, called multi-modal preserving embedding (MPE) by utilizing path-based similarity measure. We also describe two useful extensions of our method: KernelMPE and TensorMPE. Comprehensive comparisons and extensive experiments on face recognition are included to demonstrate the effectiveness of our method. Ying Wang 0008, Chunhong Pan |
FG | 2 |
| 2008 | Facial image composition based on active appearance modelabstractIn this paper, based on active appearance model (AAM), we present an easy-to-use framework for facial image composition, which can automatically exchange the source image's face or facial features onto the target image. The manual interaction is simple and the user only needs to input semantic information of ROI (region of interest) to be exchanged, such as 'face' or 'eyes'. Our framework mainly consists of two steps: model fitting and component compositing. Model fitting is designed to interpret each input image and obtain a synthesized model face of the image. Then by using component compositing, visual pleasing result is generated by solving Poisson equation with the boundary condition, produced automatically from model fitting. Furthermore, we propose a solution for eliminating the artifacts when part of the target face is occluded by hair, glasses, etc. The visually satisfactory results demonstrate the effectiveness of our facial image composition system. Chunhong Pan, Haifeng Gong, Huai-Yu Wu |
ICASSP | 2 |
| 2008 | Style preserving Chinese character synthesis based on hierarchical representation of characterabstractwith English, and they are not suitable for Chinese character synthesis. In this paper, we propose an unified approach for modeling and synthesizing Chinese characters. Using a three-level hierarchical representation, each character is decomposed into basic components, which forms the stroke database and radical database. In the synthesis process, we use a wavelet-based approach to select proper strokes and radicals, and some aesthetic constraints are defined based on the relationships between components, then genetic algorithm is employed to search for the optimal results which best match the aesthetic constraints. Experimental results demonstrates the effectiveness of our method. Ying Wang 0008, Chunhong Pan |
ICASSP | 3 |
| 2007 | Consistent Correspondence between Arbitrary Manifold SurfacesabstractWe propose a novel framework for consistent correspondence between arbitrary manifold meshes. Different from most existing methods, our approach directly maps the connectivity of the source mesh onto the target mesh without needing to segment input meshes, thus effectively avoids dealing with unstable extreme conditions (e.g. complex boundaries or high genus). In this paper, firstly, a novel mean-value Laplacian fitting scheme is proposed, which aims at computing a shape-preserving (conformal) correspondence directly in 3D-to-3D space, efficiently avoiding local optimum caused by the nearest-point search, and achieving good results even with only a few marker points. Secondly, we introduce a vertex relocation and projection approach, which refines the initial fitting result in the way of local conformity. Each vertex of the initial result is gradually projected onto the target model's surface to ensure a complete surface match. Furthermore, we provide a fast and effective approach to automatically detect critic points in the context of consistent correspondence. By fitting these critic points that capture the important features of the target mesh, the output compatible mesh matches the target mesh's profiles quite well. Compared with previous approaches, our scheme is robust, fast, and convenient, thus suitable for common applications. Huai-Yu Wu, Chunhong Pan, Qing Yang 0002, Songde Ma |
ICCV | 2 |
| 2007 | Generalized optical flow in the scale space
Haifeng Gong, Chunhong Pan, Qing Yang 0002, Hanqing Lu, Songde Ma |
Comput. Vis. Image Underst. | 2 |
| 2006 | Interactive Contour Extraction Using NURBS-HMM
Debin Lei, Chunhong Pan, Qing Yang 0002, Minyong Shi |
ACCV (1) | 2 |
| 2006 | Neural Network Modeling of Spectral EmbeddingabstractMost of spectral embedding algorithms such as Isomap, LLE and Laplacian Eigenmap only give map on training samples. One main problem of these methods is to find the embedding of new samples, which is known as the outof-sample problem of spectral embedding. In this paper, we propose a neural network based method to solve this problem. Neural network is used to train and perform both the forward map from high dimensional image space to low dimensional embedding space, and the backward map in the reverse direction. Additionally, combining the forward and backward network, this method is able to build auto-association model to retrieve high dimensional data, and cross association model to learn high dimensional correspondences. Experiments are conducted on real images for forward and backward map, auto-association and cross association. Haifeng Gong, Chunhong Pan, Qing Yang 0002, Hanqing Lu, Songde Ma |
BMVC | 2 |
| 2006 | A Region-Based Approach to Building Detection in Densely Build-Up High Resolution Satellite ImageabstractWe propose a novel region-based approach for building detection in high-resolution satellite image with densely build-up buildings. In our method, first the prior building model is constructed with texture and shape features from training building set. After over-segmentation of input image into many small regions, we identify regions which have a similar pattern with prior building model. These regions are called building like regions (BLRs). Then we group BLRs to get candidate building regions (CBRs), which have similar shape with prior building model. Next, lines which have strong relationship with each CBR are extracted. From these lines and CBR boundaries, 2-D rooftop hypotheses are generated. At last, shadows and geometrical rules are used to verify the hypotheses. Experimental results are shown on area with hundreds of buildings. Zongying Song, Chunhong Pan, Qing Yang 0002 |
ICIP | 2 |
| 2005 | A Band-Weighted Landuse Classification Method for Multispectral ImagesabstractLanduse classification is an important problem in the remote sensing field. It can be used in a wide range of applications. In this paper, we propose a hybrid method fusing edges and regions information for the landuse classification of multispectral images. It mainly includes the steps of image pre-processing, initial segmentation and region merging. Especially, a novel spatial mean shift procedure is proposed so that some information can be extracted and used in the successive steps. Aiming at the multispectral images processing, we also design a band weighting strategy that give a proper weight to each band adoptively according to the region to be processed. Experimental results on the Landsat TM and ETM+ images validate the performance of the proposed method. Chunhong Pan, Gang Wu 0017, Véronique Prinet, Qing Yang 0002, Songde Ma |
CVPR (1) | 1 |
| 2005 | A Semi-Supervised Framework for Mapping Data to the Intrinsic ManifoldabstractThis paper presents a novel scheme for manifold learning. Different from the previous work reducing data to Euclidean space which cannot handle the looped manifold well, we map the scattered data to its intrinsic parameter manifold by semisupervised learning. Given a set of partially labeled points, the map to a specified parameter manifold is computed by an iterative neighborhood average method called anchor points diffusion procedure (APD). We explore this idea on the most frequently used close formed manifolds, Stiefel manifolds whose special cases include hyper sphere and orthogonal group. The experiments show that APD can recover the underlying intrinsic parameters of points on scattered data manifold successfully. Haifeng Gong, Chunhong Pan, Qing Yang 0002, Hanqing Lu, Songde Ma |
ICCV | 2 |
| 2005 | A novel registration method for SAR and SPOT imagesabstractIn this paper, we propose a novel mutual information (MI) based method to register SAR and SPOT images. The traditional MI can register SAR and SPOT images well. However, its robustness is weakened by the absence of orientation information. In our approach, we first extract orientation information at four directions by Gabor filters, then MI of each corresponding image pair is calculated and the average value of MI is used as an improved measure for MI. Experiments show that our method is more robust than the traditional MI method. Meanwhile our method maintains comparable accuracy to the traditional MI, which is much better than coarse manual registration. Lixia Shu, Tieniu Tan, Chunhong Pan |
ICIP (2) | 4 |
| 2005 | Comparative research on landuse classification by fusing optical image and DEMabstractThree fusion methods are presented for landuse classification in this paper. They make fully use of the spectral and terrain information obtained from the optical image and DEM, and combine the remarkable features of each source in the final classification. Their performance is demonstrated by experiments and a comparison is made. Keywords-image fusion;landuse classification;DEM Chunhong Pan, Véronique Prinet, Wei Liu 0003 |
IGARSS | 2 |
| 2005 | Parametric reconstruction of generalized cylinders from limb edgesabstractThe three-dimensional (3-D) reconstruction of generalized cylinders (GCs) is an important research field in computer vision. One of the main difficulties is that some contour features in images cannot be reconstructed by traditional stereovision because they do not correspond to reflectance discontinuities of surface in space. In this paper, we present a novel, parametric approach for the 3-D reconstruction of circular generalized cylinders (CGCs) only from the limb edges of CGCs in two images. Instead of exploiting the invariant and quasiinvariant properties of some specific subclasses of GCs in projections, our reconstruction is achieved by some general assumptions on GCs, and can, therefore, be applied to a broader subclass of GCs. In order to improve robustness, we perform the extraction and labeling of the limb edge interactively, and estimate the epipolar geometry between two images by an optimal algorithm. Then, for different types of GCs, three kinds of symmetries (parallel symmetry, skew symmetry, and local smooth symmetry) are employed to compute the symmetry of limb edges. The surface points corresponding to limb edges in images are reconstructed by integrating the recovered epipolar geometry and the properties induced from the assumptions that we make on the GCs. Finally, a homography-based method is exploited to further refine the 3-D description of the GC with a coplanar curved axis. Chunhong Pan, Hongping Yan, Gérard G. Medioni, Songde Ma |
IEEE Trans. Image Process. | 1 |
| 2004 | Generalized optical flow in the scale spaceabstractScale space is a natural way to handle multi-scale problem. Yang and Ma considered the correspondence between scales, and proposed optical flow in the scale space. In this paper, we generalized Yang and Ma's work to generic images. We first generalize the Horn-Schunck algorithm to multidimensional multichannel image sequence. Since the global smoothness constraint for regularization is no longer suitable in general cases, we introduce localized smoothness weight regularization. At last, we apply the proposed method in color image scale-space pullback, together with another localized smoothness trick considering flow density. Haifeng Gong, Qing Yang 0002, Chunhong Pan, Hanqing Lu |
ICIG | 3 |
| 2004 | A land use classification method based on region and edge information fusion
Gang Wu 0017, Chunhong Pan, Véronique Prinet, Songde Ma |
ICIP | 2 |
| 2004 | A novel two-steps strategy for automatic gis-image registration
Zhanwu Yu, Véronique Prinet, Chunhong Pan |
ICIP | 3 |
| 2004 | Parametric Tracking of Legs by Exploiting Intelligent Edge
Chunhong Pan, Hongping Yan, Songde Ma |
J. Comput. Sci. Technol. | 1 |
| 2003 | Fast Construction of Plant Architectural Models Based on Substructure Decomposition
Hongping Yan, Philippe de Reffye, Chunhong Pan, Bao-Gang Hu |
J. Comput. Sci. Technol. | 3 |
| 2000 | 3D Motion Estimation of Human by Genetic AlgorithmabstractIn this paper we show that based on coplanar constraint the motion and structure of the articulated object can be determined within a scale factor from a monocular sequence of images. By genetic algorithm we achieve the unique robust numerical solution to motion estimation of the coplanar links. Then we apply the techniques to human motion analysis, and obtain the 3D motion data of joints, reanimate successfully the data. The experiments with simulated data and real images are included to demonstrate the validity of the theoretic results. Chunhong Pan, Songde Ma |
ICPR | 1 |