VLDB 2026 Research / reviewers in the wild / expert
Shiming Xiang
dblp:81/6575
· DBLP profile ↗
230ranked-venue papers
20as first author
81since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 138 · 10 first-author · 54 since 2021Graphics, computer vision, multimedia, augmented reality and games · 120 · 8 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 9 since 2021Databases, data management, data science and information retrieval · 14 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LookFlow: Training-Free and Efficient High-Resolution Image Synthesis via Dynamic Lookahead Guidance FlowabstractRectification flow Transformers (RFTs) have shown promising performance in diffusion-based image synthesis but are typically confined to lower-resolution scenarios, limiting their ability to generate high-resolution images. Existing resolution extrapolation approaches often suffer from excessive computational overhead, resulting in prolonged inference times. We propose LookFlow, a training-free high-resolution synthesis framework that accelerates inference while preserving visual quality. Building on pretrained text-to-image RFTs, LookFlow employs a dynamic lookahead guidance flow mechanism to refine high-resolution velocity predictions by leveraging multi-timestep lookahead information extracted from a low-resolution flow. Additionally, reusing temporally similar features across consecutive timesteps drastically reduces computation and significantly decreases inference time overhead. Extensive experiments on COCO demonstrate that LookFlow robustly scales resolutions from 4× to 25×, achieving up to a maximum speedup of 2.01× while maintaining competitive visual fidelity. Jianlong Chang, Ying Wang 0008, Kun Ding 0001, Shiming Xiang |
AAAI | 8 |
| 2026 | Beyond Counting: Evaluating Abstract and Emotional Reasoning in Vision-Language ModelsabstractDespite the rapid progress of Vision Language Models (VLMs), existing benchmarks still concentrate on coarse-grained object recognition or simple relational reasoning, leaving the fine-grained and higher-order reasoning abilities of these systems largely unexamined. To bridge this critical evaluation gap, we introduce EmojiGrid, a novel diagnostic benchmark specifically designed to probe these fine-grained and higher-order skills. Leveraging the universal and semantically rich nature of emojis, we synthesize a grid‑based visual dataset paired with 29,000+ QA pairs. Each pair is explicitly anchored in a three-level cognitive taxonomy comprising (i) Perception and Information Extraction, (ii) Relational and Structural Reasoning, and (iii) Abstraction and Advanced Cognition. These dimensions further decompose into nine categories covering a broad range of cognitive skills, including counting, spatial relations, compositional logic, semantic sentiment, and related higher-order reasoning tasks. Our extensive evaluation of 25 state-of-the-art open-source and proprietary VLMs reveals a significant performance gap between foundational perceptual tasks and higher-level cognitive abilities, particularly in abstraction and advanced emotional reasoning. Notably, all models struggle with compositional logic, spatial consistency, and especially emotional and semantic understanding. EmojiGrid provides a quantifiable, fine-grained benchmark to diagnose VLM limitations and guides future progress toward models that can truly perceive, reason about, and interpret complex, symbol-rich visual scenes. Jianlong Chang, Ying Wang 0008, Kun Ding 0001, Shiming Xiang |
AAAI | 8 |
| 2026 | Efficient redundancy reduction for open-vocabulary semantic segmentation
Qi Yang 0015, Kun Ding 0001, Qiyuan Cao, Shiming Xiang |
Neurocomputing | 8 |
| 2026 | From Image to Pixels: Towards Fine-Grained Medical Vision-Language ModelsabstractMultimodal large language models (MLLMs) offer immense potential for biomedical AI, yet current applications remain limited to coarse-grained image understanding and basic textual queries-falling short of the fine-grained reasoning required in clinical contexts. In this work, we present a comprehensive solution spanning data, model, and training innovations to advance pixel-level multimodal intelligence in biomedicine. First, we construct MeCoVQA, a new visual-language benchmark that spans eight medical imaging modalities and four core tasks, supporting both spatially-grounded reasoning and fine-grained diagnostic comprehension. Building on this, we introduce MedPLIB, an end-to-end biomedical MLLM equipped with pixel-level visual understanding. MedPLIB supports diverse multimodal tasks-including VQA, point- and region-based querying, grounding, and segmentation-through unified modeling. To further accommodate the heterogeneous nature of biomedical tasks, we design a task-specialized Mixture-of-Experts (MoE) architecture, where each expert is tailored to a specific task and jointly optimized via unified fine-tuning. This modular design accommodates diverse biomedical tasks while maintaining a unified and efficient architecture. By integrating retrieval-augmented generation (RAG) and in-context learning (ICL), MedPLIB also demonstrates strong generalization on out-of-distribution (OOD) medical image segmentation. Experiments across multiple benchmarks show that MedPLIB sets a new state-of-the-art on biomedical vision-language tasks; notably, it outperforms the best existing small and large models by 19.7 and 15.6 mDice in zero-shot pixel-level grounding, highlighting its clinical utility and generalization strength. Lingdong Shen, Xiaoshuang Huang, Fangxin Shang, Yehui Yang, Bin Fan 0001, Shiming Xiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | AtmosOceanNet: SST forecast method driven by atmospheric-oceanic multimodal data
Qinxuan Wang, Kun Ding 0001, Ying Wang 0008, Yineng Li, Shiming Xiang, Xiaoqing Chu |
Pattern Recognit. | 6 |
| 2026 | USVTrack: A Benchmark for Multi-Object Tracking in Complex Water Surface ScenesabstractMulti-object tracking (MOT) in water surface scenes is crucial for the autonomous navigation of Unmanned Surface Vehicles (USVs). However, existing MOT datasets rarely focus on these scenes. Moreover, the few available water surface MOT datasets contain limited data shot onboard and concentrate narrowly on specific marine scenes, creating a significant gap from real-world USV navigation applications. To promote research on USV autonomous navigation, we introduce USVTrack, a fully onboard-shot MOT benchmark that covers diverse and complex water surface scenes, characterized by a high proportion of small objects and varied backgrounds. Then, we propose an innovative end-to-end method specifically designed for MOT in complex water surface scenes, termed as USVMOT. It improves tracking performance through four key contributions: 1) integrating mask information via knowledge distillation to boost feature discriminability; 2) deploying task-specific auxiliary pathways to alleviate the competition between detection and re-identification (ReID) in end-to-end MOT methods; 3) employing an adaptive high-quality mask generation strategy based on the Segment Anything Model (SAM) that obviates extensive manual annotation; and 4) introducing an object-aware association method that dynamically tailors the tracking strategy according to object size and motion speed. Extensive experiments on the USVTrack benchmark demonstrate that USVMOT outperforms existing methods. Our analysis reveals that MOT in complex water surface scenes remains challenging, highlighting the need for further advancements. Yuwei Cheng, Kun Ding 0001, Chunhong Pan, Shiming Xiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic SegmentationabstractPre-training techniques significantly enhance the performance of semantic segmentation tasks with limited training data. However, the efficacy under a large domain gap between pre-training (e.g. RGB) and fine-tuning (e.g. infrared) remains underexplored. In this study, we first benchmark the infrared semantic segmentation performance of various pre-training methods and reveal several phenomena distinct from the RGB domain. Next, our layerwise analysis of pre-trained attention maps uncovers that: (1) There are three typical attention patterns (local, hybrid, and global); (2) Pre-training tasks notably influence pattern distribution across layers; (3) The hybrid pattern is crucial for semantic segmentation as it attends to both nearby and foreground elements; (4) The texture bias impedes model generalization in infrared tasks. Building on these insights, we propose UNIP, a UNified Infrared Pre-training framework, to enhance the pre-trained model performance. This framework uses the hybrid-attention distillation NMI-HAD as the pre-training target, a large-scale mixed dataset InfMix for pre-training, and a last-layer feature pyramid network LL-FPN for fine-tuning. Experimental results show that UNIP outperforms various pre-training methods by up to 13.5% in average mIoU on three infrared segmentation tasks, evaluated using fine-tuning and linear probing metrics. UNIP-S achieves performance on par with MAE-L while requiring only 1/10 of the computational cost. Furthermore, with fewer parameters, UNIP significantly surpasses state-of-the-art (SOTA) infrared or RGB segmentation methods and demonstrates the broad potential for application in other modalities, such as RGB and depth. Our code is available at https://github.com/casiatao/UNIP. Jinyong Wen, Kun Ding 0001, Shiming Xiang, Chunhong Pan |
ICLR | 5 |
| 2025 | Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models StrongerabstractRecent advancements in Large Vision Language Models (LVLMs) have significantly improved performance in Visual Question Answering (VQA) tasks through multimodal Retrieval-Augmented Generation (RAG). However, existing methods still face challenges, such as the scarcity of knowledge with reasoning examples and erratic responses from retrieved knowledge. To address these issues, in this study, we propose a multimodal RAG framework, termed RCTS, which enhances LVLMs by constructing a Reasoning Context-enriched knowledge base and a Tree Search re-ranking method. Specifically, we introduce a self-consistent evaluation mechanism to enrich the knowledge base with intrinsic reasoning patterns. We further propose a Monte Carlo Tree Search with Heuristic Rewards (MCTS-HR) to prioritize the most relevant examples. This ensures that LVLMs can leverage high-quality contextual reasoning for better and more consistent responses. Extensive experiments demonstrate that our framework achieves state-of-the-art performance on multiple VQA datasets, significantly outperforming In-Context Learning (ICL) and Vanilla-RAG methods. It highlights the effectiveness of our knowledge base and re-ranking method in improving LVLMs. Qi Yang 0015, Chenghao Zhang 0003, Lubin Fan, Kun Ding 0001, Jieping Ye, Shiming Xiang |
ICML | 6 |
| 2025 | EvoVLMA: Evolutionary Vision-Language Model Adaptation
Kun Ding 0001, Ying Wang 0008, Shiming Xiang |
ACM Multimedia | 3 |
| 2025 | Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and FilteringabstractThe task of Knowlegde-Based Visual Question Answering (KB-VQA) requires the model to understand visual features and retrieve external knowledge. Retrieval-Augmented Generation (RAG) have been employed to address this problem through knowledge base querying. However, existing work demonstrate two limitations: insufficient interactivity during knowledge retrieval and ineffective organization of retrieved information for Visual-Language Model (VLM). To address these challenges, we propose a three-stage visual language model with Process, Retrieve and Filter (VLM-PRF) framework. For interactive retrieval, VLM-PRF uses reinforcement learning (RL) to guide the model to strategically process information via tool-driven operations. For knowledge filtering, our method trains the VLM to transform the raw retrieved information into into task-specific knowledge. With a dual reward as supervisory signals, VLM-PRF successfully enable model to optimize retrieval strategies and answer generation capabilities simultaneously. Experiments on two datasets demonstrate the effectiveness of our framework. Yuyang Hong, Qi Yang 0015, Lubin Fan, Ying Wang 0008, Kun Ding 0001, Shiming Xiang, Jieping Ye |
NeurIPS | 8 |
| 2025 | Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference OptimizationabstractPreference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face challenges in handling noisy images of different timesteps and require complex transformations into pixel space. In this work, we show that pre-trained diffusion models are naturally suited for step-level reward modeling in the noisy latent space, as they are explicitly designed to process latent images at various noise levels. Accordingly, we propose the **Latent Reward Model (LRM)**, which repurposes components of the diffusion model to predict preferences of latent images at arbitrary timesteps. Building on LRM, we introduce **Latent Preference Optimization (LPO)**, a step-level preference optimization method conducted directly in the noisy latent space. Experimental results indicate that LPO significantly improves the model's alignment with general, aesthetic, and text-image alignment preferences, while achieving a 2.5-28x training speedup over existing preference optimization methods. Cheng Da, Kun Ding 0001, Huan Yang 0005, Yan Li 0043, Tingting Gao, Di Zhang 0026, Shiming Xiang, Chunhong Pan |
NeurIPS | 9 |
| 2025 | HAN: An efficient hierarchical self-attention network for skeleton-based gesture recognition
Ying Wang 0008, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 3 |
| 2025 | Transformer with token attention and attribute prediction for image captioning
Lifei Song, Ying Wang 0008, Linsu Shi, Jiazhong Yu, Shiming Xiang |
Pattern Recognit. Lett. | 6 |
| 2025 | Leveraging Privileged Information for Partially Observable Reinforcement LearningabstractReinforcement learning has achieved remarkable success across diverse scenarios. However, learning optimal policies within partially observable games remains a formidable challenge. Crucial privileged information in states is often shrouded during gameplay, yet ideally, it should be accessible and exploitable during training. Previous studies have concentrated on formulating policies based wholly on partial observations or oracle states. Nevertheless, these approaches often face hindrances in attaining effective generalization. To surmount this challenge, we propose the actor–cross-critic (ACC) learning framework, integrating both partial observations and oracle states. ACC achieves this by coordinating two critics and invoking a maximization operation mechanism to switch between them dynamically. This approach encourages the selection of the higher values when computing advantages within the actor–critic framework, thereby accelerating learning and mitigating bias under partial observability. Some theoretical analyses show that ACC exhibits better learning ability toward optimal policies than actor–critic learning using the oracle states. We highlight its superior performance through comprehensive evaluations in decision-making tasks, such asQuestBall,Minigrid, andAtari, and the challenging card gameDouDizhu. Jinqiu Li, Enmin Zhao, Junliang Xing, Shiming Xiang |
IEEE Trans. Games | 5 |
| 2024 | Weak Distribution Detectors Lead to Stronger Generalizability of Vision-Language Prompt TuningabstractWe propose a generalized method for boosting the generalization ability of pre-trained vision-language models (VLMs) while fine-tuning on downstream few-shot tasks. The idea is realized by exploiting out-of-distribution (OOD) detection to predict whether a sample belongs to a base distribution or a novel distribution and then using the score generated by a dedicated competition based scoring function to fuse the zero-shot and few-shot classifier. The fused classifier is dynamic, which will bias towards the zero-shot classifier if a sample is more likely from the distribution pre-trained on, leading to improved base-to-novel generalization ability. Our method is performed only in test stage, which is applicable to boost existing methods without time-consuming re-training. Extensive experiments show that even weak distribution detectors can still improve VLMs' generalization ability. Specifically, with the help of OOD detectors, the harmonic mean of CoOp and ProGrad increase by 2.6 and 1.5 percentage points over 11 recognition datasets in the base-to-novel setting. Kun Ding 0001, Haojian Zhang, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
AAAI | 5 |
| 2024 | Defying Imbalanced Forgetting in Class Incremental LearningabstractWe observe a high level of imbalance in the accuracy of different learned classes in the same old task for the first time. This intriguing phenomenon, discovered in replay-based Class Incremental Learning (CIL), highlights the imbalanced forgetting of learned classes, as their accuracy is similar before the occurrence of catastrophic forgetting. This discovery remains previously unidentified due to the reliance on average incremental accuracy as the measurement for CIL, which assumes that the accuracy of classes within the same task is similar. However, this assumption is invalid in the face of catastrophic forgetting. Further empirical studies indicate that this imbalanced forgetting is caused by conflicts in representation between semantically similar old and new classes. These conflicts are rooted in the data imbalance present in replay-based CIL methods. Building on these insights, we propose CLass-Aware Disentanglement (CLAD) as a means to predict the old classes that are more likely to be forgotten and enhance their accuracy. Importantly, CLAD can be seamlessly integrated into existing CIL methods. Extensive experiments demonstrate that CLAD consistently improves current replay-based methods, resulting in performance gains of up to 2.56%. Shixiong Xu, Gaofeng Meng, Xing Nie, Bolin Ni, Bin Fan 0001, Shiming Xiang |
AAAI | 6 |
| 2024 | Enhancing Visual Continual Learning with Language-Guided SupervisionabstractContinual learning (CL) aims to empower models to learn new tasks without forgetting previously acquired knowledge. Most prior works concentrate on the techniques of architectures, replay data, regularization, etc. However, the category name of each class is largely neglected. Existing methods commonly utilize the one-hot labels and randomly initialize the classifier head. We argue that the scarce semantic information conveyed by the one-hot labels hampers the effective knowledge transfer across tasks. In this paper, we revisit the role of the classifier head within the CL paradigm and replace the classifier with semantic knowledge from pretrained language models (PLMs). Specifically, we use PLMs to generate semantic targets for each class, which are frozen and serve as supervision signals during training. Such targets fully consider the semantic correlation between all classes across tasks. Empirical studies show that our approach mitigates forgetting by alleviating representation drifting and facilitating knowledge transfer across tasks. The proposed method is simple to implement and can seamlessly be plugged into existing methods with negligible adjustments. Extensive experiments based on eleven mainstream baselines demonstrate the effectiveness and generalizability of our approach to various protocols. For example, under the class-incremental learning setting on ImageNet-100, our method significantly improves the Top-1 accuracy by 3.2% to 6.1% while reducing the forgetting rate by 2.6% to 13.1%. Bolin Ni, Hongbo Zhao 0006, Chenghao Zhang 0003, Gaofeng Meng, Zhaoxiang Zhang 0001, Shiming Xiang |
CVPR | 7 |
| 2024 | Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual SegmentationabstractRecently, an audio-visual segmentation (AVS) task has been introduced, aiming to group pixels with sounding objects within a given video. This task necessitates a first-ever audio-driven pixel-level understanding of the scene, posing significant challenges. In this paper, we propose an innovative audio-visual transformer framework, termed COMBO, an acronym for COoperation of Multi-order Bi-lateral relatiOns. For the first time, our framework ex-plores three types of bilateral entanglements within AVS: pixel entanglement, modality entanglement, and temporal entanglement. Regarding pixel entanglement, we employ a Siam-Encoder Module (SEM) that leverages prior knowledge to generate more precise visual features from thefoun-dational model. For modality entanglement, we design a Bilateral-Fusion Module (BFM), enabling COMBO to align corresponding visual and auditory signals bi-directionally. As for temporal entanglement, we introduce an innovative adaptive inter-frame consistency loss according to the inherent rules of temporal. Comprehensive experiments and ablation studies on AVSBench-object (84.7 mIoU on S4, 59.2 mIou on MS3) and AVSBench-semantic (42.1 mIoU on AVSS) datasets demonstrate that COMBO surpasses previous state-of-the-art methods. Project page is available at https://yannqi.github.io/AVS-COMBO. Qi Yang 0015, Xing Nie, Ying Guo 0008, Shiming Xiang |
CVPR | 8 |
| 2024 | AddressCLIP: Empowering Vision-Language Models for City-Wide Image Address Localization
Shixiong Xu, Chenghao Zhang 0003, Lubin Fan, Gaofeng Meng, Shiming Xiang, Jieping Ye |
ECCV (28) | 5 |
| 2024 | An Efficient Graph Autoencoder with Lightweight Desmoothing Decoder and Long-Range ModelingabstractGraph self-supervised learning provides a powerful guarantee for learning high-quality representations in an unsupervised manner. Despite its early birth, the performance of generative graph self-supervised learning has long lagged behind that of up-and-coming contrastive learning, especially on node classification tasks. In this paper, we investigate potential issues in existing graph autoencoders and attribute their poor performance to three main aspects: complex decoder design, lack of desmoothing process in feature remap, and overemphasis on local topological proximity. To tackle these issues, we propose an effective and efficient graph autoencoder framework for unsupervised representation learning, which contains two key components: lightweight smoothness-aware feature reconstructor and global structural dependency catcher. After performing a desmoothing operation on encoded representations via a learnable high-pass filter, the feature decoder reconstructs the original features through a simple linear projection. The lightweight design liberates the decoder from self-supervised pretext tasks and puts the encoder more accountable for achieving optimization objectives, which promotes effective training of the encoder. Global structural dependency catcher utilizes graph diffusion to build a structural regularization to capture long-range topological dependency on a graph. The empirical studies demonstrate the effectiveness of our approach, which can surpass dominant contrastive learning methods. Jinyong Wen, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
ICDM | 4 |
| 2024 | Dual Critic Reinforcement Learning under Partial ObservabilityabstractPartial observability in environments poses significant challenges that impede the formation of effective policies in reinforcement learning. Prior research has shown that borrowing the complete state information can enhance sample efficiency. This strategy, however, frequently encounters unstable learning with high variance in practical applications due to the over-reliance on complete information. This paper introduces DCRL, a Dual Critic Reinforcement Learning framework designed to adaptively harness full-state information during training to reduce variance for optimized online performance. In particular, DCRL incorporates two distinct critics: an oracle critic with access to complete state information and a standard critic functioning within the partially observable context. It innovates a synergistic strategy to meld the strengths of the oracle critic for efficiency improvement and the standard critic for variance reduction, featuring a novel mechanism for seamless transition and weighting between them. We theoretically prove that DCRL mitigates the learning variance while maintaining unbiasedness. Extensive experimental analyses across the Box2D and Box3D environments have verified DCRL's superior performance. The source code is available in the supplementary. Jinqiu Li, Enmin Zhao, Junliang Xing, Shiming Xiang |
NeurIPS | 5 |
| 2024 | Compositional Kronecker Context Optimization for vision-language models
Kun Ding 0001, Ying Wang 0008, Haojian Zhang, Shiming Xiang |
Neurocomputing | 6 |
| 2024 | Multi-task prompt tuning with soft context sharing for vision-language models
Kun Ding 0001, Ying Wang 0008, Pengzhang Liu, Haojian Zhang, Shiming Xiang, Chunhong Pan |
Neurocomputing | 6 |
| 2024 | Image captioning: Semantic selection unit with stacked residual attention
Lifei Song, Ying Wang 0008, Yuanhua Wang, Shiming Xiang |
Image Vis. Comput. | 6 |
| 2024 | Efficient Remote Sensing Image Super-Resolution via Lightweight Diffusion ModelsabstractWith the emergence of diffusion models, the image generation has experienced a significant advancement. In super-resolution tasks, diffusion models surpass generative adversarial network (GAN)-based methods in generating more realistic samples. However, these models come with significant costs: denoising networks rely on large U-Net, making them computationally intensive for high-resolution (HR) images, and the extensive sampling steps in diffusion models lead to prolonged inference time. This complexity limits their application in remote sensing, due to the high demand for high-resolution images in such scenarios. To address this, we propose a lightweight diffusion model (LWTDM), which simplifies the denoising network and efficiently incorporates conditional information using a cross-attention-based encoder–decoder architecture. Furthermore, LWTDM serves as the pioneering model that incorporates the accelerated sampling technique from denoising diffusion implicit models (DDIMs). This integration involves the meticulous selection of sampling steps, ensuring the quality of the generated images. The experiments confirm that LWTDM strikes a favorable balance between precision and perceptual quality, while its faster inference speed makes it suitable for diverse remote sensing scenarios with specific requirements. The source code is available at:https://github.com/Suanmd/LWTDM. Tai An, Chunlei Huo, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Preformer: Simple and Efficient Design for Precipitation Nowcasting With TransformersabstractThe primary objective of precipitation nowcasting is to predict precipitation patterns several hours in advance. Recent studies have emphasized the potential of deep learning methods for this task. To harness the correlations among various meteorological elements, existing frameworks project multiple meteorological elements into a latent space and then utilize convolutional-recurrent networks for future precipitation prediction. Although effective, the escalating model complexity may impede practical applications. This letter develops the Preformer, a streamlined Transformer framework for precipitation nowcasting that efficiently captures global spatiotemporal dependencies among multiple meteorological elements. The Preformer implements an encoder-translator-decoder architecture, where the encoder integrates spatial features of multiple elements, the translator models spatiotemporal dynamics, and the decoder combines spatiotemporal information to forecast future precipitation. Without introducing complex structures or strategies, the Preformer achieves state-of-the-art performance even with the least parameters. Qizhao Jin, Xinbang Zhang, Xinyu Xiao, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Few-shot video object segmentation with prototype evolution
Binjie Mao, Xiyan Liu, Linsu Shi, Jiazhong Yu, Shiming Xiang |
Neural Comput. Appl. | 6 |
| 2024 | Graph Aggregating-Repelling Network: Do Not Trust All Neighbors in Heterophilic Graphs
Yuhu Wang, Jinyong Wen, Chunxia Zhang 0001, Shiming Xiang |
Neural Networks | 4 |
| 2024 | Reusable Architecture Growth for Continual Stereo MatchingabstractThe remarkable performance of recent stereo depth estimation models benefits from the successful use of convolutional neural networks to regress dense disparity. Akin to most tasks, this needs gathering training data that covers a number of heterogeneous scenes at deployment time. However, training samples are typically acquired continuously in practical applications, making the capability to learn new scenes continually even more crucial. For this purpose, we propose to perform continual stereo matching where a model is tasked to 1) continually learn new scenes, 2) overcome forgetting previously learned scenes, and 3) continuously predict disparities at inference. We achieve this goal by introducing a Reusable Architecture Growth (RAG) framework. RAG leverages task-specific neural unit search and architecture growth to learn new scenes continually in both supervised and self-supervised manners. It can maintain high reusability during growth by reusing previous units while obtaining good performance. Additionally, we present a Scene Router module to adaptively select the scene-specific architecture path at inference. Comprehensive experiments on numerous datasets show that our framework performs impressively in various weather, road, and city circumstances and surpasses the state-of-the-art methods in more challenging cross-dataset settings. Further experiments also demonstrate the adaptability of our method to unseen scenes, which can facilitate end-to-end stereo architecture learning and practical deployment. Chenghao Zhang 0003, Gaofeng Meng, Bin Fan 0001, Zhaoxiang Zhang 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | MoBoo: Memory-Boosted Vision Transformer for Class-Incremental LearningabstractContinual learning strives to acquire knowledge across sequential tasks without forgetting previously assimilated knowledge. Current state-of-the-art methodologies utilize dynamic architectural strategies to increase the network capacity for new tasks. However, these approaches often suffer from a rapid growth in the number of parameters. While some methods introduce an additional network compression stage to address this, they tend to construct complex and hyperparameter-sensitive systems. In this work, we introduce a novel solution to this challenge by proposing Memory-Boosted transformer (MoBoo), instead of conventional architecture expansion and compression. Specifically, we design a memory-augmented attention mechanism by establishing a memory bank where the “key” and “value” linear projections are stored. This memory integration prompts the model to leverage previously learned knowledge, thereby enhancing stability during training at a marginal cost. The memory bank is lightweight and can be easily managed with a straightforward queue. Moreover, to increase the model’s plasticity, we design a memory-attentive aggregator, which leverages the cross-attention mechanism to adaptively summarize the image representation from the encoder output that has historical knowledge involved. Extensive experiments on challenging benchmarks demonstrate the effectiveness of our method. For example, on ImageNet-100 under 10 tasks, our method outperforms the current state-of-the-art methods by +3.74% in average accuracy and using fewer parameters. Bolin Ni, Xing Nie, Chenghao Zhang 0003, Shixiong Xu, Xin Zhang 0093, Gaofeng Meng, Shiming Xiang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Pro-Tuning: Unified Prompt Tuning for Vision TasksabstractIn computer vision, fine-tuning is the de-facto approach to leverage pre-trained vision models to perform downstream tasks. However, deploying it in practice is quite challenging, due to adopting parameter inefficient global update and heavily relying on high-quality downstream data. Recently, prompt-based learning, which adds the task-relevant prompt to adapt the pre-trained models to downstream tasks, has drastically boosted the performance of many natural language downstream tasks. In this work, we extend this notable transfer ability benefited from prompt into vision models as an alternative to fine-tuning. To this end, we propose parameter-efficient Prompt tuning (Pro-tuning) to adapt diverse frozen pre-trained models to a wide variety of downstream vision tasks. The key to Pro-tuning is prompt-based tuning, i.e., learning task-specific vision prompts for downstream input images with the pre-trained model frozen. By only training a small number of additional parameters, Pro-tuning can generate compact and robust downstream models both for CNN-based and transformer-based network architectures. Comprehensive experiments evidence that the proposed Pro-tuning outperforms fine-tuning on a broad range of vision tasks and scenarios, including image classification (under generic objects, class imbalance, image corruption, natural adversarial examples, and out-of-distribution generalization), and dense prediction tasks such as object detection and semantic segmentation. Xing Nie, Bolin Ni, Jianlong Chang, Gaofeng Meng, Chunlei Huo, Shiming Xiang, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Active Disparity Sampling for Stereo Matching With Adjoint NetworkabstractThe sparse signals provided by external sources have been leveraged as guidance for improving dense disparity estimation. However, previous methods assume depth measurements to be randomly sampled, which restricts performance improvements due to under-sampling in challenging regions and over-sampling in well-estimated areas. In this work, we introduce an Active Disparity Sampling problem that selects suitable sampling patterns to enhance the utility of depth measurements given arbitrary sampling budgets. We achieve this goal by learning an Adjoint Network for a deep stereo model to measure its pixel-wise disparity quality. Specifically, we design a hard-soft prior supervision mechanism to provide hierarchical supervision for learning the quality map. A Bayesian optimized disparity sampling policy is further proposed to sample depth measurements with the guidance of the disparity quality. Extensive experiments on standard datasets with various stereo models demonstrate that our method is suited and effective in different stereo architectures and outperforms existing fixed and adaptive sampling methods under different sampling rates. Remarkably, the proposed method makes substantial improvements when generalized to heterogeneous unseen domains. Chenghao Zhang 0003, Gaofeng Meng, Bolin Ni, Shiming Xiang |
IEEE Trans. Image Process. | 5 |
| 2024 | Non-Maximum Suppression Guided Label Assignment for Object Detection in Crowd ScenesabstractThe detection performance in crowd scenes is limited by recalling hard objects (e.g., occluded objects). It requires that this kind of objects can be successfully detected and retained by the non-maximum suppression (NMS) while controlling false positives. The existing dynamic label assignment algorithms can help recall these objects by adaptively allocating appropriate positive samples, however, they ignore the alignment with the selecting rules of NMS. This leads to the fact that detecting objects in crowd scenes are still very sensitive to the NMS threshold setting. As a result, the existing methods can only set a low NMS threshold to avoid the excessive false positives, causing some objects failed to be recalled. And these methods also generally lack more excitation for positive samples, which hinders further facilitating the recall of hard instances in crowd scenes. This article proposes a novel dynamic label assignment strategy for object detection in crowd scenes, callednon-maximum suppression guided label assignment(NGLA), which aligns the assignment strategy with NMS process and learns more prominent positive samples. Following NMS, NGLA introduces the IoU between samples with their corresponding best samples to define positive and negative samples. To cooperate with NGLA, anNMS-aware lossis proposed to dynamically assign sample weights when supervising sample predictions, which also considers the IoU with the best sample. In addition, for better classification prediction, aregression assisted classification branchis designed to help detectors perceive the relation between the regression predictions of each sample and the corresponding best sample. Experiments demonstrate that NGLA outperforms other label assignment methods on CrowdHuman and Citypersons, and is less sensitive to the NMS threshold in crowd scenes. Hangzhi Jiang, Xin Zhang 0093, Shiming Xiang |
IEEE Trans. Multim. | 3 |
| 2024 | On the Equivalence of Linear Discriminant Analysis and Least Squares RegressionabstractStudying the relationship between linear discriminant analysis (LDA) and least squares regression (LSR) is of great theoretical and practical significance. It is well-known that the two-class LDA is equivalent to an LSR problem, and directly casting multiclass LDA as an LSR problem, however, becomes more challenging. Recent study reveals that the equivalence between multiclass LDA and LSR can be established based on a special class indicator matrix, but under a mild condition which may not hold under the scenarios with low-dimensional or oversampled data. In this article, we show that the equivalence between multiclass LDA and LSR can be established based on arbitrary linearly independent class indicator vectors and without any condition. In addition, we show that LDA is also equivalent to a constrained LSR based on the data-dependent indicator vectors. It can be concluded that under exactly the same mild condition, such two regressions are both equivalent to the null space LDA method. Illuminated by the equivalence of LDA and LSR, we propose a direct LDA classifier to replace the conventional framework of LDA plus extra classifier. Extensive experiments well validate the above theoretic analysis. Feiping Nie 0001, Hong Chen 0015, Shiming Xiang, Changshui Zhang, Shuicheng Yan, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Domain Decorrelation with Potential Energy RankingabstractMachine learning systems, especially the methods based on deep learning, enjoy great success in modern computer vision tasks under ideal experimental settings. Generally, these classic deep learning methods are built on the i.i.d. assumption, supposing the training and test data are drawn from the same distribution independently and identically. However, the aforementioned i.i.d. assumption is, in general, unavailable in the real-world scenarios, and as a result, leads to sharp performance decay of deep learning algorithms. Behind this, domain shift is one of the primary factors to be blamed. In order to tackle this problem, we propose using Potential Energy Ranking (PoER) to decouple the object feature and the domain feature in given images, promoting the learning of label-discriminative representations while filtering out the irrelevant correlations between the objects and the background. PoER employs the ranking loss in shallow layers to make features with identical category and domain labels close to each other and vice versa. This makes the neural networks aware of both objects and background characteristics, which is vital for generating domain-invariant features. Subsequently, with the stacked convolutional blocks, PoER further uses the contrastive loss to make features within the same categories distribute densely no matter domains, filtering out the domain information progressively for feature alignment. PoER reports superior performance on domain generalization benchmarks, improving the average top-1 accuracy by at least 1.20% compared to the existing methods. Moreover, we use PoER in the ECCV 2022 NICO Challenge, achieving top place with only a vanilla ResNet-18 and winning the jury award. The code has been made publicly available at: https://github.com/ForeverPs/PoER. Sen Pei, Jiaxi Sun, Shiming Xiang, Gaofeng Meng |
AAAI | 4 |
| 2023 | Robust Feature Rectification of Pretrained Vision Models for Object RecognitionabstractPretrained vision models for object recognition often suffer a dramatic performance drop with degradations unseen during training. In this work, we propose a RObust FEature Rectification module (ROFER) to improve the performance of pretrained models against degradations. Specifically, ROFER first estimates the type and intensity of the degradation that corrupts the image features. Then, it leverages a Fully Convolutional Network (FCN) to rectify the features from the degradation by pulling them back to clear features. ROFER is a general-purpose module that can address various degradations simultaneously, including blur, noise, and low contrast. Besides, it can be plugged into pretrained models seamlessly to rectify the degraded features without retraining the whole model. Furthermore, ROFER can be easily extended to address composite degradations by adopting a beam search algorithm to find the composition order. Evaluations on CIFAR-10 and Tiny-ImageNet demonstrate that the accuracy of ROFER is 5% higher than that of SOTA methods on different degradations. With respect to composite degradations, ROFER improves the accuracy of a pretrained CNN by 10% and 6% on CIFAR-10 and Tiny-ImageNet respectively. Shengchao Zhou, Gaofeng Meng, Zhaoxiang Zhang 0001, Shiming Xiang |
AAAI | 5 |
| 2023 | Bilateral Memory Consolidation for Continual LearningabstractHumans are proficient at continuously acquiring and integrating new knowledge. By contrast, deep models forget catastrophically, especially when tackling highly long task sequences. Inspired by the way our brains constantly rewrite and consolidate past recollections, we propose a novel Bilateral Memory Consolidation (BiMeCo) framework that focuses on enhancing memory interaction capabilities. Specifically, BiMeCo explicitly decouples model parameters into short-term memory module and long-term memory module, responsible for representation ability of the model and generalization over all learned tasks, respectively. BiMeCo encourages dynamic interactions between two memory modules by knowledge distillation and momentum-based updating for forming generic knowledge to prevent forgetting. The proposed BiMeCo is parameter-efficient and can be integrated into existing methods seamlessly. Extensive experiments on challenging benchmarks show that BiMeCo significantly improves the performance of existing continual learning methods. For example, combined with the state-of-the-art method CwD [55], BiMeCo brings in significant gains of around 2% to 6% while using 2x fewer parameters on CIFAR-100 under ResNet-18. Xing Nie, Shixiong Xu, Xiyan Liu, Gaofeng Meng, Chunlei Huo, Shiming Xiang |
CVPR | 6 |
| 2023 | Continual Semantic Segmentation via Scalable Contrastive Clustering and Background DiversityabstractDespite the efficacy towards static data distribution, traditional semantic segmentation methods encounter Catastrophic forgetting when tackling continually changing data streams. Another fundamental challenge is Background shift, which results from the semantic drift of the background class during continual learning steps. To extend the applicability of semantic segmentation methods, we introduce a novel, scalable segmentation architecture called ScaleSeg, designed to adapt the incremental scenarios. The architecture of ScaleSeg consists of a series of prototypes updated by online contrastive clustering. Additionally, we propose a background diversity strategy to enhance the model’s plasticity and stability, thus overcoming background shift. Comprehensive experiments and ablation studies on challenging benchmarks demonstrate that ScaleSeg surpasses previous state-of-the-art methods, particularly when dealing with extensive task sequences. Qi Yang 0015, Xing Nie, Linsu Shi, Jiazhong Yu, Shiming Xiang |
ICDM | 6 |
| 2023 | Improving the Homophily of Heterophilic Graphs for Semi-Supervised Node ClassificationabstractGraph Neural Networks (GNNs) have been applied to process the widespread graph data, including social networks and web data, etc. However, lots of GNNs can only perform well on homophilic graphs, while losing their superiority when tackling heterophilic graphs. Recent works try to use spectral theory or attention mechanism to design some more complex learning paradigms for heterophilic graphs. In this paper, we instead utilize some explored properties to construct three new graph structures of high homophily to improve the homophily of heterophilic graphs for better representation learning. Along with the original graph structure, totally four graph structures are injected into a Multi-View Graph Fusion Network (MVGFN) to learn a group of more expressive features for the semi-supervised node classification. Ablation experiments show that all three newly-constructed graph structures obtain higher homophily levels. Comparisons among several baselines indicate the superiority of our method on both homophilic and heterophilic graphs. Yuhu Wang, Shiming Xiang, Chunhong Pan |
ICME | 2 |
| 2023 | Graph Information Interaction on Feature and Structure via Cross-modal Contrastive LearningabstractThe abundant features and structure information on graphs provide a potential guarantee for learning high-quality representations without supervision. Feature attribute represents the inherent properties of nodes, while structure attribute describes their neighborhood relationship. These two types of attributes can be regarded as different modal forms of the same instance and should be consistent in identifying a member. We propose to directly regard feature and structure attributes as two separate views to embed this consistency into contrastive learning method, realizing graph information interaction on feature and structure in a cross-modal contrastive framework. Under this framework, node representations are learned in an unsupervised manner by maximizing the agreement between feature representation and structure representation. In terms of negative samples, instead of randomly sampling points from empirical distribution, a simple yet effective multi-sample mixing strategy is proposed to synthesize true negative samples with greater probability, alleviating the tricky false negative issue. Extensive experiments on multiple types of graphs demonstrate the effectiveness of the proposed method. Jinyong Wen, Yuhu Wang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
ICME | 4 |
| 2023 | Exploring Universal Principles for Graph Contrastive Learning: A Statistical PerspectiveabstractAlthough recent advances have prompted the prosperity in graph contrastive learning, the researches on universal principles for model design and desirable properties of latent representations are still inadequate. From a statistical perspective, this paper proposes two principles for guidance and constructs a general self-supervised framework for negative-free graph contrastive learning. Reformulating data augmentation as a mixture process, the first one, termed consistency principle, lays stress on exploring and mapping cross-view common information to consistent and essence-revealing representations. For the purpose of instantiation, four statistical indicators are employed to estimate and maximize the correlation between representations from various views, whose accordant variation trend during training implies the extraction of common content. With awareness of the insufficiency of a solo consistency principle, suffering from degenerated and coupled solutions, a decorrelation principle is put forward to encourage diverse and informative representations. Accordingly, two specific strategies, performing in representation space and eigen spectral space, respectively, are propounded to decouple various representation channels. Under two principles, various combinations of concrete implementations derive a family of methods. The comparison experiments with current state-of-the-arts demonstrate the effectiveness and sufficiency of two principles for high-quality graph representations. Furthermore, visual studies reveal how certain principles affect learned representations. Jinyong Wen, Shiming Xiang, Chunhong Pan |
ACM Multimedia | 2 |
| 2023 | Multi-level consistency regularization for domain adaptive object detection
Chenghao Zhang 0003, Ying Wang 0106, Shiming Xiang |
Neural Comput. Appl. | 4 |
| 2023 | Domain adaptive object detection with model-agnostic knowledge transferring
Chenghao Zhang 0003, Ying Wang 0106, Shiming Xiang |
Neural Networks | 4 |
| 2023 | Graph convolutional network with tree-guided anisotropic message passing
Yuhu Wang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
Neural Networks | 4 |
| 2023 | Patch loss: A generic multi-scale perceptual loss for single image super-resolutionabstractIn single image super-resolution (SISR), although PSNR is a key metric for signal fidelity, images with high PSNR do not necessarily render high visual quality. As a result, current perception-driven SISR methods employ perceptual metrics close to the human eye to measure the quality of the generated images. Unfortunately, the perceptual loss and adversarial loss, widely used by the perception-driven SISR methods, still underperform on these non-differentiable perceptual metrics. To this end, we propose a generic multi-scale perceptual loss, i.e., the patch loss, which can be easily plugged into off-the-shelf SISR methods to improve a broad range of perceptual metrics. Specifically, the proposed patch loss minimizes the multi-scale similarity of image patches and enhances the restoration of regions with complex textures and sharp edges via parameter-free adaptive patch-wise attention. Our proposed patch loss introduces more realistic details compared to the perceptual loss and fewer artifacts compared to the adversarial loss. Tai An, Binjie Mao, Chunlei Huo, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 5 |
| 2023 | AutoMSNet: Multi-Source Spatio-Temporal Network via Automatic Neural Architecture Search for Traffic Flow PredictionabstractRecently the research of traffic flow prediction with deep learning framework has be largely developed, whereas most current methods are still faced with the following shortcomings. For spatial feature extraction, studies have shown that both local and non-local correlations exist on traffic networks. Considering the temporal dependencies, short-term impending and longer periodic components are two most critical patterns of traffic data, which further provide different information for the prediction task. Furthermore, multi-source heterogeneous external data, which naturally holds semantic gap with traffic data, also have impact on traffic flow. To solve the above problems, this paper proposes an AutoMSNet (Multi-Source Spatio-Temporal Network via Automatic neural architecture search). The AutoMSNet is composed of an encoder-decoder structure. The encoder takes neighboring data as inputs, while the decoder captures long-term periodic patterns. Thus, different functions of two temporal features are simultaneously extracted. Moreover, a neural architecture search space is designed for spatial feature extraction. Through architecture search technique, graph convolutions with different receptive fields are automatically selected and combined to form an optimal module structure. Therefore, both local and non-local spatial features can be adaptively captured. Besides, a meta learning feature fusion strategy is proposed to integrate external data, which can alleviate the semantic gap between different data sources. Extensive experiments on three real-world traffic datasets evaluate the superiority of the proposed model. Shen Fang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Learning from the Target: Dual Prototype Network for Few Shot Semantic SegmentationabstractDue to the scarcity of annotated samples, the diversity between support set and query set becomes the main obstacle for few shot semantic segmentation. Most existing prototype-based approaches only exploit the prototype from the support feature and ignore the information from the query sample, failing to remove this obstacle.In this paper, we proposes a dual prototype network (DPNet) to dispose of few shot semantic segmentation from a new perspective. Along with the prototype extracted from the support set, we propose to build the pseudo-prototype based on foreground features in the query image. To achieve this goal, the cycle comparison module is developed to select reliable foreground features and generate the pseudo-prototype with them. Then, a prototype interaction module is utilized to integrate the information of the prototype and the pseudo-prototype based on their underlying correlation. Finally, a multi-scale fusion module is introduced to capture contextual information during the dense comparison between prototype (pseudo-prototype) and query feature. Extensive experiments conducted on two benchmarks demonstrate that our method exceeds previous state-of-the-arts with a sizable margin, verifying the effectiveness of the proposed method. Binjie Mao, Xinbang Zhang, Lingfeng Wang 0002, Qian Zhang 0009, Shiming Xiang, Chunhong Pan |
AAAI | 5 |
| 2022 | Discriminative Graph Representation Learning with Distributed SamplingabstractGraph neural networks (GNNs) have been widely used to accomplish graph classification tasks such as predicting molecular properties and classifying the labels of proteins. Discovering the latent discriminative substructures (e.g., functional groups in molecules) is a vital task to enhance the classification performance. In this paper, this task is addressed as a problem of discriminative graph representation learning. Specifically, a novel node sampling strategy is developed to achieve this goal. To this end, graph-dependent sampling vectors are first learned by a mini-network to exploit various informative substructures on graphs and sample some representative nodes, which could be regarded as performing a distributed sampling on graphs. Then, the sampled nodes are organized together topologically as a subgraph with landing probabilities of random walks. Moreover, a self-adaptive pooling ratio of nodes is obtained via feature smoothness of graphs, eliminating the trouble of manual selection of subgraph size. As a result, these treatments are equivalent to performing the difficult step of down-pooling operation on non-grid graph data. Extensive experiments and ablation studies on multiple benchmark datasets demonstrate the effectiveness and superiority of our proposed approach. Additionally, interpretability studies illustrate the ability of our model to extract discriminative substructures. Jinyong Wen, Yuhu Wang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
BIBM | 4 |
| 2022 | AME: Attention and Memory Enhancement in Hyper-Parameter OptimizationabstractTraining Deep Neural Networks (DNNs) is inherently subject to sensitive hyper-parameters and untimely feedbacks of performance evaluation. To solve these two difficulties, an efficient parallel hyper-parameter optimization model is proposed under the framework of Deep Reinforcement Learning (DRL). Technically, we develop Attention and Memory Enhancement (AME), that includes multi-head attention and memory mechanism to enhance the ability to capture both the short-term and long-term relationships between different hyper-parameter configurations, yielding an attentive sampling mechanism for searching high-performance configurations embedded into a huge search space. During the optimization of transformer-structured configuration searcher, a conceptually intuitive yet powerful strategy is applied to solve the problem of insufficient number of samples due to the untimely feedback. Experiments on three visual tasks, including image classification, object detection, semantic segmentation, demonstrate the effectiveness of AME. Nuo Xu 0006, Jianlong Chang, Xing Nie, Chunlei Huo, Shiming Xiang, Chunhong Pan |
CVPR | 5 |
| 2022 | Expanding Language-Image Pretrained Models for General Video Recognition
Bolin Ni, Houwen Peng, Songyang Zhang 0004, Gaofeng Meng, Jianlong Fu, Shiming Xiang, Haibin Ling |
ECCV (4) | 7 |
| 2022 | Spatiotemporal Contextual Consistency Network for Precipitation NowcastingabstractPrecipitation nowcasting is forecasting rainfall in the short-term conditioned by the known meteorological parameters. Recently, deep neural networks (DNNs) have shown outstanding performance in this task. But, there are several challenges imposed by the multiple meteorological elements, including the multimodal modeling, the considerable variation in scales of precipitation region, as well as the long-tailed distribution of rainfall data. To solve these problems, this paper proposes Spatiotemporal Contextual Consistency Network (SCCN) for learning from the multi meteorological elements. Architecturally, a parameter-shared multimodal fusion CNN encoder, which dynamically exchanges features between different modalities, is used to encode the multimodal meteorological data. To improve the spatial modeling of the multiple meteorological features, we compose the multi-scale filters and deconstruction convolution to modify the gate operators in ConvLSTM to propose a spatial contextual consistency ConvLSTM (SCC-ConvLSTM). Furthermore, considering the temporal consistency in rainfall, a temporal consistency module (TCM) is designed to gear to long-tailed distribution. Under this module, different long-tailed meteorological elements are calculated to encode features and residuals fused with the previous precipitation distribution in sequence. The experimental results of precipitation nowcasting demonstrate the effectiveness of our method on the ERA5 dataset and WeatherBench dataset. Xinyu Xiao, Qizhao Jin, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICDM | 4 |
| 2022 | Components Regulated Generation of Handwritten Chinese Text-lines in Arbitrary LengthabstractGenerating readable images of handwritten Chinese text-lines is very challenging due to complicated topological structures in Chinese. To address this problem, we propose a components regulated model named HCT-GAN to generate the entire lines of Chinese handwriting from text-line labels. Specifically, HCT-GAN is designed as a CGAN-based architecture that additionally integrates a Chinese text encoder (CTE), a sequence recognition module(SRM), and a spatial perception module (SPM). Compared with the one-hot embedding, CTE learns the latent content representation by reusing the structure and component embedding shared among the Chinese characters. SRM provides sequence-level constraints to the generated images. SPM can adaptively constrain the spatial correlation between the generated components, which facilitates the modeling of characters with complicated topological structures. Benefiting from such artful modeling, our model suffices to generate images of handwritten Chinese text-lines in arbitrary length. Extensive experimental results demonstrate that our model achieves state-of-the-art performance in handwritten Chinese lines generation. Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICPR | 4 |
| 2022 | AutoMF: Spatio-temporal Architecture Search for The Meteorological Forecasting TaskabstractDespite years of studies, meteorological forecasting with deep learning still faces several challenges, including the multi-modal correlation, the spatio-temporal dependency, the spatial heterogeneity, and the temporal periodicity. Though with elaborate design, manually designed networks adopted by current methods could be far from optimal in modeling the spatiotemporal dynamics of meteorological data. In this work, we propose AutoMF, i.e. a multi-level architecture search framework for the Meteorological Forecasting task. Working in a data-driven paradigm, AutoMF is capable of generating suitable convolution networks and spatio-temporal networks to capture the multi-modal correlation and the spatio-temporal dependency. Based on this framework, we develop a differential sampling-based architecture search method to optimize the architecture, and introduce the progressive search strategy to facilitate the search process. Furthermore, the spatial heterogeneity and temporal periodicity are explicitly modeled through integrating corresponding status indicators. Extensive experiments exhibit the capability of the proposed method. Xinbang Zhang, Qizhao Jin, Shiming Xiang, Chunhong Pan |
ICPR | 3 |
| 2022 | Adversarial Gradient Driven Exploration for Deep Click-Through Rate PredictionabstractExploration-Exploitation (E& E) algorithms are commonly adopted to deal with the feedback-loop issue in large-scale online recommender systems. Most of existing studies believe that high uncertainty can be a good indicator of potential reward, and thus primarily focus on the estimation of model uncertainty. We argue that such an approach overlooks the subsequent effect of exploration on model training. From the perspective of online learning, the adoption of an exploration strategy would also affect the collecting of training data, which further influences model learning. To understand the interaction between exploration and training, we design a Pseudo-Exploration module that simulates the model updating process after a certain item is explored and the corresponding feedback is received. We further show that such a process is equivalent to adding an adversarial perturbation to the model input, and thereby name our proposed approach as an the Adversarial Gradient Driven Exploration (AGE). For production deployment, we propose a dynamic gating unit to pre-determine the utility of an exploration. This enables us to utilize the limited amount of resources for exploration, and avoid wasting pageview resources on ineffective exploration. The effectiveness of AGE was firstly examined through an extensive number of ablation studies on an academic dataset. Meanwhile, AGE has also been deployed to one of the world-leading display advertising platforms, and we observe significant improvements on various top-line evaluation metrics. Kailun Wu, Weijie Bian, Zhangming Chan, Lejian Ren, Shiming Xiang, Shuguang Han, Hongbo Deng, Bo Zheng 0007 |
KDD | 5 |
| 2022 | CAN: Feature Co-Action Network for Click-Through Rate PredictionabstractFeature interaction has been recognized as an important problem in machine learning, which is also very essential for click-through rate (CTR) prediction tasks. In recent years, Deep Neural Networks (DNNs) can automatically learn implicit nonlinear interactions from original sparse features, and therefore have been widely used in industrial CTR prediction tasks. However, the implicit feature interactions learned in DNNs cannot fully retain the complete representation capacity of the original and empirical feature interactions (e.g., cartesian product) without loss. For example, a simple attempt to learn the combination of feature A and feature B < A, B > as the explicit cartesian product representation of new features can outperform previous implicit feature interaction models including factorization machine (FM)-based models and their variations. This indicates there is still a big gap between explicit and implicit feature interaction models. However, to learn all the explicit feature interaction (cartesian product) representations requires a very large sample size along with N times of original parameter space (where N is quite large in most industrial applications). In this paper, we propose a Co-Action Network (CAN) to approximate the explicit pairwise feature interactions without introducing too many additional parameters. More specifically, giving feature A and its associated feature B, their feature interaction is modeled by learning two sets of parameters: 1) the embedding of feature A, and 2) a Multi-Layer Perceptron (MLP) to represent feature B. The approximated feature interaction can be obtained by passing the embedding of feature A through the MLP network of feature B. We refer to such pairwise feature interaction as feature co-action, and such a Co-Action Network unit can provide a very powerful capacity to fitting complex feature interactions. In addition, FM can be viewed as a special case of the CAN unit when the MLP is a single layer with only one output. Experimental results on public and industrial datasets show that CAN outperforms state-of-the-art CTR models and the cartesian product method. Moreover, CAN has been deployed in the display advertisement system in Alibaba, obtaining 12% improvement on CTR and 8% on Revenue Per Mille (RPM), which is a great improvement to the business. The code for experiments in this paper is open-sourced\footnotehttps://github.com/CAN-Paper/Co-Action-Network. Weijie Bian, Kailun Wu, Lejian Ren, Qi Pi, Can Xiao, Xiang-Rong Sheng, Yong-Nan Zhu, Zhangming Chan, Na Mou, Xinchen Luo, Shiming Xiang, Guorui Zhou, Xiaoqiang Zhu, Hongbo Deng |
WSDM | 12 |
| 2022 | Urban scene based Semantical Modulation for Pedestrian Detection
Hangzhi Jiang, Shengcai Liao, Jinpeng Li 0004, Véronique Prinet, Shiming Xiang |
Neurocomputing | 5 |
| 2022 | Task-aware adaptive attention learning for few-shot semantic segmentation
Binjie Mao, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
Neurocomputing | 3 |
| 2022 | TVGCN: Time-variant graph convolutional network for traffic forecasting
Yuhu Wang, Shen Fang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
Neurocomputing | 4 |
| 2022 | Meta Graph Transformer: A Novel Framework for Spatial-Temporal Traffic Prediction
Xue Ye, Shen Fang, Chunxia Zhang 0001, Shiming Xiang |
Neurocomputing | 5 |
| 2022 | Learning adversarial point-wise domain alignment for stereo matching
Chenghao Zhang 0003, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Neurocomputing | 4 |
| 2022 | Monocular contextual constraint for stereo matching with adaptive weights assignment
Chenghao Zhang 0003, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Image Vis. Comput. | 4 |
| 2022 | HENet: Head-Level Ensemble Network for Very High Resolution Remote Sensing Images Semantic SegmentationabstractSemantic segmentation plays an important role in very high resolution (VHR) image understanding. Despite the potentials of the deep convolutional network in improving performance by end-to-end feature learning, each model has its limitations, and it is hard to discriminate complex features purely by a single model. Ensemble learning is promising for integrating the strengths of different models, however, the ensemble of deep models is challenging due to the huge amount of parameters and computation of the deep model itself as well as the difficulty in capturing complementarity between different models. To tackle these problems, a head-level ensemble network (HENet) is proposed in this letter, which reduces model complexity by sharing feature extraction networks and improves complementarity between models by novel cooperative learning (CL). Experiments on ISPRS 2-D semantic labeling benchmark demonstrate the effectiveness and advantage of the proposed method. Chunlei Huo, Nuo Xu 0005, Xin Zhang 0093, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Train in Dense and Test in Sparse: A Method for Sparse Object Detection in Aerial ImagesabstractApplications of aerial imaging, especially based on unmanned aerial vehicles (UAVs) platform, rapidly explode in recent years. Meanwhile, vision-based sensing, e.g., detection and recognition, for UAVs becomes increasingly important. Objects in aerial images are usually of tiny size, hence occupying a limited area. Terminology speaking, the images are very sparse in spatial. However, existing work in aerial object detection commonly ignores this point. Conversely, we explore the availability of such a property in improving the detection performance of aerial images. Specifically, we propose a general method, train in dense and test in sparse (TDTS), to exploit sparsity in aerial object detection: 1) in the training stage, the possible positions of object are learned by training a fully convolutional network (called prophet head) and 2) in the testing stage, prophet head identifies the possible object locations to reduce redundant computation in classification and box prediction head by sparse convolution. By extensive experiments on the VisDrone2019-Det data set, we find that the sparsity can not only help to speed up inference but also to improve accuracy. Thus, we argue that the sparsity deserves more attention. Kun Ding 0001, Guojin He, Huxiang Gu, Zisha Zhong, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | CMT: Cross Mean Teacher Unsupervised Domain Adaptation for VHR Image Semantic SegmentationabstractSemantic segmentation of remote sensing images has achieved superior results with the supervised deep learning models. However, their performance to unseen data domains could be very bad due to the domain shift between different domains. Recently, a series of unsupervised domain adaptation (UDA) methods has been developed to solve the domain shift problem in semantic segmentation. Most of them use adversarial learning to achieve global cross-domain alignment and use a self-training (ST) strategy to generate pseudo-labels for classwise alignment. However, these methods ignore the pixels that are not assigned pseudo-labels. Those pixels are mostly at the boundaries, which are vital to the final segmentation results. To solve this problem, this letter proposes a cross mean teacher (CMT) UDA method. The whole framework consists of two parts. On the one hand, the global cross-domain distribution alignment is performed, and then, reliable pseudo-labels are assigned to the target data. On the other hand, a cross teacher–student network (CTSN) is developed to effectively use those pixels with and without pseudo-labels. This network contains two student networks ($S_{1}$and$S_{2}$) and two teacher networks ($T_{1}$and$T_{2}$) for cross-consistency constraints that supervises$S_{2}$(or$S_{1}$) by the prediction results of$T_{1}$(or$T_{2}$). The cross supervision by CTSN is helpful to prevent performance bottlenecks caused by the high coupling of teacher–student network in existing methods. Extensive experiments on three different remote sensing adaptation scenes verify the effectiveness and superiority of the proposed method. Bin Fan 0001, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | MFNet: The Spatio-Temporal Network for Meteorological Forecasting With Architecture SearchabstractExploiting deep learning for the meteorological forecasting task is a challenging due to the complex spatio-temporal correlation, non-stationarity and imbalanced data distribution. Though with elaborate design, handcraft hierarchical architectures adopted by current methods could be far from optimal in sufficiently modeling the dynamics of meteorological data. For the Meteorological Forecasting task, this letter presents the MFNet, which is a spatio-temporal network with the Neural Architecture Search (NAS) technique. Working in the data-driven paradigm, our method is capable of automatically generating suitable architecture to model the spatio-temporal correlation. Moreover, the non-stationarity of meteorological data is explicitly modeled through simulating spatio-temporal variations in response to the intrinsic driven force of the meteorological state, and the Error Sensitive Regression (ESR) loss is introduced accounting for the imbalanced data distribution. Extensive experiments exhibit the capability of our method and demonstrate that deep learning is potential for serving as an operational technique for global meteorological forecasting. Xinbang Zhang, Qizhao Jin, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Subgraph-aware graph structure revision for spatial-temporal graph modeling
Yuhu Wang, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
Neural Networks | 3 |
| 2022 | Scene captioning with deep fusion of images and point clouds
Chunxia Zhang 0001, Lubin Weng, Shiming Xiang, Chunhong Pan |
Pattern Recognit. Lett. | 4 |
| 2022 | MS-Net: Multi-Source Spatio-Temporal Network for Traffic Flow PredictionabstractPredicting urban traffic flow is a challenging task, due to the complicated spatio-temporal dependencies on traffic networks. Urban traffic flow usually has both short-term neighboring and long-term periodic temporal dependencies. It is also noticed that the spatial correlations over different traffic nodes are both local and non-local. What’s more, the traffic flow is affected by various external factors. To capture the non-local spatial correlations, we propose a Dilated Attentional Graph Convolution (DAGC). The DAGC utilizes a dilated graph convolution kernel to expand the nodes’ receptive field and exploit multi-order neighborhood. Technically, the lower-order neighborhood corresponds to local spatial dependencies, while the higher-order neighborhood corresponds to non-local spatial dependencies between nodes. Based on DAGC, a Multi-Source Spatio-Temporal Network (MS-Net) is designed, which suffices to integrate long-range historical traffic data as well as multi-modal external information. MS-Net consists of four components: a spatial feature extraction module, a temporal feature fusion module, an external factors embedding module, and a multi-source data fusion module. Extensive experiments on three real traffic datasets demonstrates that the proposed model performs well on both the public transportation networks, road networks, and can handle large-scale traffic networks in particular the Beijing bus network which has more than 4,000 traffic nodes. Shen Fang, Véronique Prinet, Jianlong Chang, Michael Werman, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Decoupled Representation Learning for Character Glyph SynthesisabstractCharacter glyph synthesis is still an open challenging problem, which involves two related aspects,i.e., font style transfer and content consistency. In this paper, we propose a novel model named FontGAN, which integrates the character structure stylization, de-stylization and texture transfer into a unified framework. Specifically, we decouple character images into style representation and content representation, which offers fine-grained control of these two types of variables, thus improving the quality of the generated results. To effectively capture the style information, a style consistency module (SCM) is introduced. Technically, SCM exploits category-guided Kullback-Leibler divergence to explicitly model the style representation into different prior distributions. In this way, our model is capable of implementing transformations between multiple domains in one framework. In addition, we propose content prior module (CPM) to provide content prior for the model to guide the content encoding process and alleviates the problem of stroke deficiency during structure de-stylization. Benefiting from the idea of decoupling and regrouping, our FontGAN suffices to achieve many-to-many translation tasks for glyph structure. Experimental results demonstrate that the proposed FontGAN achieves the state-of-the-art performance in character glyph synthesis. Xiyan Liu, Gaofeng Meng, Jianlong Chang, Ruiguang Hu, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 5 |
| 2021 | Ltaf-Net: Learning Task-Aware Adaptive Features and Refining Mask for Few-Shot Semantic SegmentationabstractFew shot segmentation is a newly-developing and challenging computer vision task which is only provided with few labeled samples of the novel class. Some recent works on this problem focus more on how to design an effective comparison module but ignore how to extract the features passed to compare. In this paper we propose a novel model named LTAF-Net for few-shot segmentation. This model aims to adaptively recalibrate the extracted features which could boost the accuracy of dense comparison between support features and query features. Besides an additional prediction refinement module is designed to refine the initial mask. Meanwhile this method can apply to k-shot setting without developing a new specialized architecture and achieve competitive performance. Experiments on PASCAL-5iand FSS-1000 strongly prove the effectiveness of the proposed model. Our model outperforms the second-best method 1.4% in 1-shot and 0.76% in 5-shot respectively in PASCAL-5i. Binjie Mao, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
ICASSP | 3 |
| 2021 | Reinforcement Stacked Learning with Semantic-Associated Attention for Visual Question AnsweringabstractThe task of visual question answering (VQA) is to generate an answer for a question according to the content of an image being asked. In this process, the critical problems of effectively embedding the question feature and image feature as well as transforming the features to the prediction of answer are still faithfully unresolved. In this paper, depending on these problems, a semantic-associated attention method and a reinforcement stacked learning mechanism are proposed. Firstly, within the associations of high-level semantics, a visual spatial attention model (VSA) and a multi-semantic attention model (MSA) are proposed to extract the low-level image feature and high-level semantic feature, respectively. Furthermore, we develop a reinforcement stacked learning architecture, which splits the transformation process into multiple stages, to gradually approach the answers. At each stage, a new reinforcement learning (RL) method is introduced to directly criticize inappropriate answers to optimize the model. The extensive experiments on the VQA task show that our method can achieve state-of-the-art performance. Xinyu Xiao, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
ICASSP | 3 |
| 2021 | Knowledge Mining and Transferring for Domain Adaptive Object DetectionabstractWith the thriving of deep learning, CNN-based object detectors have made great progress in the past decade. However, the domain gap between training and testing data leads to a prominent performance degradation and thus hinders their application in the real world. To alleviate this problem, Knowledge Transfer Network (KTNet) is proposed as a new paradigm for domain adaption. Specifically, KT-Net is constructed on a base detector with intrinsic knowledge mining and relational knowledge constraints. First, we design a foreground/background classifier shared by source domain and target domain to extract the common attribute knowledge of objects in different scenarios. Second, we model the relational knowledge graph and explicitly constrain the consistency of category correlation under source domain, target domain, as well as cross-domain conditions. As a result, the detector is guided to learn object-related and domain-independent representation. Extensive experiments and visualizations confirm that transferring object-specific knowledge can yield notable performance gains. The proposed KTNet achieves state-of-the-art results on three cross-domain detection benchmarks. Chenghao Zhang 0003, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
ICCV | 4 |
| 2021 | Dual Stream Fusion Network for Multi-spectral High Resolution Remote Sensing Image Segmentation
Yiwen Shi, Chunlei Huo, Shiming Xiang, Chunhong Pan |
PRCV (2) | 5 |
| 2021 | Relational Attention with Textual Enhanced Transformer for Image Captioning
Lifei Song, Yiwen Shi, Xinyu Xiao, Chunxia Zhang 0001, Shiming Xiang |
PRCV (3) | 5 |
| 2021 | 3D-SceneCaptioner: Visual Scene Captioning Network for Three-Dimensional Point Clouds
Xianbing Pan, Shiming Xiang, Chunhong Pan |
PRCV (2) | 3 |
| 2021 | Few-Shot Learning via Feature Hallucination with Variational InferenceabstractDeep learning has achieved huge success in the field of artificial intelligence, but the performance heavily depends on labeled data. Few-shot learning aims to make a model rapidly adapt to unseen classes with few labeled samples after training on a base dataset, and this is useful for tasks lacking labeled data such as medical image processing. Considering that the core problem of few-shot learning is the lack of samples, a straightforward solution to this issue is data augmentation. This paper proposes a generative model (VI-Net) based on a cosine-classifier baseline. Specifically, we construct a framework to learn to define a generating space for each category in the latent space based on few support samples. In this way, new feature vectors can be generated to help make the decision boundary of classifier sharper during the fine-tuning process. To evaluate the effectiveness of our proposed approach, we perform comparative experiments and ablation studies on mini-ImageNet and CUB. Experimental results show that VI-Net does improve performance compared with the baseline and obtains the state-of-the-art result among other augmentation-based methods. Qinxuan Luo, Lingfeng Wang 0002, Jingguo Lv, Shiming Xiang, Chunhong Pan |
WACV | 4 |
| 2021 | DATA: Differentiable ArchiTecture Approximation With Distribution Guided SamplingabstractNeural architecture search (NAS) is inherently subject to the gap of architectures during searching and validating. To bridge this gap effectively, we develop Differentiable ArchiTecture Approximation (DATA) with Ensemble Gumbel-Softmax (EGS) estimator and Architecture Distribution Constraint (ADC) to automatically approximate architectures during searching and validating in a differentiable manner. Technically, the EGS estimator consists of a group of Gumbel-Softmax estimators, which is capable of converting probability vectors to binary codes and passing gradients reversely, reducing the estimation bias in a differentiable way. To narrow the distribution gap between sampled architectures and supernet, further, the ADC is introduced to reduce the variance of sampling during searching. Benefiting from such modeling, architecture probabilities and network weights in the NAS model can be jointly optimized with the standard back-propagation, yielding an end-to-end learning mechanism for searching deep neural architectures in an extended search space. Conclusively, in the validating process, a high-performance architecture that approaches to the learned one during searching is readily built. Extensive experiments on various tasks including image classification, few-shot learning, unsupervised clustering, semantic segmentation and language modeling strongly demonstrate that DATA is capable of discovering high-performance architectures while guaranteeing the required efficiency. Code is available at https://github.com/XinbangZhang/DATA-NAS. Xinbang Zhang, Jianlong Chang, Yiwen Guo, Gaofeng Meng, Shiming Xiang, Zhouchen Lin, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | You Only Search Once: Single Shot Neural Architecture Search via Direct Sparse OptimizationabstractRecently neural architecture search (NAS) has raised great interest in both academia and industry. However, it remains challenging because of its huge and non-continuous search space. Instead of applying evolutionary algorithm or reinforcement learning as previous works, this paper proposes a direct sparse optimization NAS (DSO-NAS) method. The motivation behind DSO-NAS is to address the task in the view of model pruning. To achieve this goal, we start from a completely connected block, and then introduce scaling factors to scale the information flow between operations. Next, sparse regularizations are imposed to prune useless connections in the architecture. Lastly, an efficient and theoretically sound optimization method is derived to solve it. Our method enjoys both advantages of differentiability and efficiency, therefore it can be directly applied to large datasets like ImageNet and tasks beyond classification. Particularly, on the CIFAR-10 dataset, DSO-NAS achieves an average test error 2.74 percent, while on the ImageNet dataset DSO-NAS achieves 25.4 percent test error under 600M FLOPs with 8 GPUs in 18 hours. As for semantic segmentation task, DSO-NAS also achieve competitive result compared with manually designed architectures on the PASCAL VOC dataset. Code is available at https://github.com/XinbangZhang/DSO-NAS. Xinbang Zhang, Zehao Huang, Naiyan Wang, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Meta-MSNet: Meta-Learning Based Multi-Source Data Fusion for Traffic Flow PredictionabstractTraffic flow prediction is a challenging task while most existing works are faced with two main problems in extracting complicated intrinsic and extrinsic features. In terms of intrinsic features, current methods don't fully exploit different functions of short-term neighboring and long-term periodic temporal patterns. As for extrinsic features, recent works mainly employ hand-crafted fusion strategies to integrate external factors but remain generalization issues. To solve these problems, we propose a meta-learning based multi-source spatio-temporal network (Meta-MSNet). The Meta-MSNet is designed with an encoder-decoder structure. The encoder captures neighboring temporal dependencies while the decoder extracts periodic features. Furthermore, two meta-learning based fusion modules are designed to integrate multi-source external data both on temporal and spatial dimensions. Experiments on three real-world traffic datasets have verified the superiority of the proposed model. Shen Fang, Xianbing Pan, Shiming Xiang, Chunhong Pan |
IEEE Signal Process. Lett. | 3 |
| 2021 | Handwritten Text Generation via Disentangled RepresentationsabstractAutomatically generating handwritten text images is a challenging task due to the diverse handwriting styles and the irregular writing in natural scenes. In this paper, we propose an effective generative model called HTG-GAN to synthesize handwritten text images from latent prior. Unlike single-character synthesis, our method is capable of generating images of sequence characters with arbitrary length, which pays more attention to the structural relationship between characters. We model the structural relationship as the style representation to avoid explicitly modeling the stroke layout. Specifically, the text image is disentangled into style representation and content representation, where the style representation is mapped into Gaussian distribution and the content representation is embedded using character index. In this way, our model can generate new handwritten text images with specified contents and various styles to perform data augmentation, thereby boosting handwritten text recognition (HTR). Experimental results show that our method achieves state-of-the-art performance in handwritten text generation. Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
IEEE Signal Process. Lett. | 3 |
| 2021 | CSLM: Convertible Short-Term and Long-Term Memory in Differential Neural ComputersabstractExternal memory-based neural networks, such as differentiable neural computers (DNCs), have recently gained importance and popularity to solve complex sequential learning tasks that pose challenges to conventional neural networks. However, a trained DNC usually has a low-memory utilization efficiency. This article introduces a variation of DNC architecture with a convertible short-term and long-term memory, named CSLM-DNC. Unlike the memory architecture of the original DNC, the new scheme of short-term and long-term memories offers different importance of memory locations for read and write, and they can be converted over time. This is mainly motivated by the human brain where short-term memory stores large amounts of noisy and unimportant information and decays rapidly, while long-term memory stores important information and lasts for a long time. The conversion of these two types of memory is allowed and is able to be learned according to their reading and writing frequency. We quantitatively and qualitatively evaluate the proposed CSLM-DNC architecture on the tasks of question answering, copy and repeat copy, showing that it can significantly improve memory efficiency and learning performance. Shiming Xiang, Bo Tang 0011 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Spatio-Temporal Graph Structure Learning for Traffic ForecastingabstractAs an indispensable part in Intelligent Traffic System (ITS), the task of traffic forecasting inherently subjects to the following three challenging aspects. First, traffic data are physically associated with road networks, and thus should be formatted as traffic graphs rather than regular grid-like tensors. Second, traffic data render strong spatial dependence, which implies that the nodes in the traffic graphs usually have complex and dynamic relationships between each other. Third, traffic data demonstrate strong temporal dependence, which is crucial for traffic time series modeling. To address these issues, we propose a novel framework named Structure Learning Convolution (SLC) that enables to extend the traditional convolutional neural network (CNN) to graph domains and learn the graph structure for traffic forecasting. Technically, SLC explicitly models the structure information into the convolutional operation. Under this framework, various non-Euclidean CNN methods can be considered as particular instances of our formulation, yielding a flexible mechanism for learning on the graph. Along this technical line, two SLC modules are proposed to capture the global and local structures respectively and they are integrated to construct an end-to-end network for traffic forecasting. Additionally, in this process, Pseudo three Dimensional convolution (P3D) networks are combined with SLC to capture the temporal dependencies in traffic data. Extensively comparative experiments on six real-world datasets demonstrate our proposed approach significantly outperforms the state-of-the-art ones. Jianlong Chang, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
AAAI | 4 |
| 2020 | AugFPN: Improving Multi-Scale Feature Learning for Object DetectionabstractCurrent state-of-the-art detectors typically exploit feature pyramid to detect objects at different scales. Among them, FPN is one of the representative works that build a feature pyramid by multi-scale features summation. However, the design defects behind prevent the multi-scale features from being fully exploited. In this paper, we begin by first analyzing the design defects of feature pyramid in FPN, and then introduce a new feature pyramid architecture named AugFPN to address these problems. Specifically, AugFPN consists of three components: Consistent Supervision, Residual Feature Augmentation, and Soft RoI Selection. AugFPN narrows the semantic gaps between features of different scales before feature fusion through Consistent Supervision. In feature fusion, ratio-invariant context information is extracted by Residual Feature Augmentation to reduce the information loss of feature map at the highest pyramid level. Finally, Soft RoI Selection is employed to learn a better RoI feature adaptively after feature fusion. By replacing FPN with AugFPN in Faster R-CNN, our models achieve 2.3 and 1.6 points higher Average Precision (AP) when using ResNet50 and MobileNet-v2 as backbone respectively. Furthermore, AugFPN improves RetinaNet by 1.6 points AP and FCOS by 0.9 points AP when using ResNet50 as backbone. Codes are available on https://github.com/Gus-Guo/AugFPN. Chaoxu Guo, Bin Fan 0001, Qian Zhang 0009, Shiming Xiang, Chunhong Pan |
CVPR | 4 |
| 2020 | Decoupled Representation Learning for Skeleton-Based Gesture RecognitionabstractSkeleton-based gesture recognition is very challenging, as the high-level information in gesture is expressed by a sequence of complexly composite motions. Previous works often learn all the motions with a single model. In this paper, we propose to decouple the gesture into hand posture variations and hand movements, which are then modeled separately. For the former, the skeleton sequence is embedded into a 3D hand posture evolution volume (HPEV) to represent fine-grained posture variations. For the latter, the shifts of hand center and fingertips are arranged as a 2D hand movement map (HMM) to capture holistic movements. To learn from the two inhomogeneous representations for gesture recognition, we propose an end-to-end two-stream network. The HPEV stream integrates both spatial layout and temporal evolution information of hand postures by a dedicated 3D CNN, while the HMM stream develops an efficient 2D CNN to extract hand movement features. Eventually, the predictions of the two streams are aggregated with high efficiency. Extensive experiments on SHREC'17 Track, DHG-14/28 and FPHA datasets demonstrate that our method is competitive with the state-of-the-art. Yongcheng Liu, Ying Wang 0008, Véronique Prinet, Shiming Xiang, Chunhong Pan |
CVPR | 5 |
| 2020 | PackDet: Packed Long-Head Object Detector
Kun Ding 0001, Guojin He, Huxiang Gu, Zisha Zhong, Shiming Xiang, Chunhong Pan |
ECCV (13) | 5 |
| 2020 | Learning Where to Focus for Efficient Video Object Detection
Zhengkai Jiang 0001, Yu Liu 0015, Ceyuan Yang, Jihao Liu, Peng Gao 0007, Qian Zhang 0009, Shiming Xiang, Chunhong Pan |
ECCV (16) | 7 |
| 2020 | Forground-Guided Vehicle Perception FrameworkabstractAs the basis of advanced visual tasks such as vehicle tracking and traffic flow analysis, vehicle detection needs to accurately predict the position and category of vehicle objects. In the past decade, deep learning based methods have made great progress. However, we also notice that some existing cases are not studied thoroughly. First, false positive on the background regions is one of the critical problems. Second, most of the previous approaches only optimize a single vehicle detection model, ignoring the relationship between different visual perception tasks. In response to the above two findings, we introduce a foreground segmentation branch for the first time, which can predict the pixel level of vehicles in advance. Furthermore, two attention modules are designed to guide the work of the detection branch. The proposed method can be easily grafted into the one-stage and two-stage detection framework. We evaluate the effectiveness of our model on LSVH, a dataset with large variations in vehicle scales, and achieve the state-of-the-art detection accuracy. Shiming Xiang, Chunhong Pan |
ICPR | 3 |
| 2020 | Deep Space Probing for Point Cloud Analysisabstract3D points distribute in a continuous 3D space irregularly, thus directly adapting 2D image convolution to 3D points is not an easy job. Previous works often artificially divide the space into regular grids, yet it could be suboptimal to learn geometry. In this paper, we propose SPCNN, namely, Space Probing Convolutional Neural Network, which naturally generalizes image CNN to deal with point clouds. The key idea of SPCNN is learning to probe the 3D space in an adaptive manner. Specifically, we define a pool of learnable convolutional weights, and let each point in the local region learn to choose a suitable convolutional weight from the pool. This is achieved by constructing a geometry guided index-mapping function that implicitly establishes a correspondence between convolutional weights and some local regions in the neighborhood (Fig. 1). In this way, the index-mapping function learns to adaptively partition nearby space for local geometry pattern recognition. With this convolution as a basic operator, SPCNN, a hierarchical architecture can be developed for effective point cloud analysis. Extensive experiments on challenging benchmarks across three tasks demonstrate that SPCNN achieves the state-of-the-art or has competitive performance. Yirong Yang, Bin Fan 0001, Yongcheng Liu, Jiyong Zhang 0001, Xin Liu 0027, Xinyu Cai, Shiming Xiang, Chunhong Pan |
ICPR | 8 |
| 2020 | Detecting Maneuvering Target Accurately Based on a Two-Phase Approach From Remote Sensing ImageryabstractManeuvering target detection in satellite images is difficult due to their small sizes, blurred appearances under various illuminations and shadows, and occlusion by trees and buildings. Recently, a fully convolutional regression network (FCRN) was proposed and achieved by the state-of-the-art performance in the Munich vehicle database. However, such a one-phase approach often makes mistakes at difficult places because of its swift glance and rejecting any second check. In this letter, a new object spatial density building net (SDBN) was designed, and a two-phase detection approach was proposed. It used the first SDBN to generate candidate regions and the second SDBN to proceed with a meticulous check on the object categories. Experiments on four maneuvering target databases, the Munich vehicle database, the Open Vehicle Database of San-Francisco (OVDS), the Overhead Imagery Research Data Set (OIRDS), and the Open Aircraft Database (OAD) show that the proposed method outperforms FCRN by an obvious margin. In addition, the accurate geometrical parameters (positions, orientations, and lengths) of all the objects were computed based on the spatial density maps, and the published experimental result of FCRN in OIRDS was pointed out and the corrected result was given. All source codes and databases are available at http://www.github.com/cxy177/SDBN. Xueyun Chen, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Deep Self-Evolution ClusteringabstractClustering is a crucial but challenging task in pattern analysis and machine learning. Existing methods often ignore the combination between representation learning and clustering. To tackle this problem, we reconsider the clustering task from its definition to develop Deep Self-Evolution Clustering (DSEC) to jointly learn representations and cluster data. For this purpose, the clustering task is recast as a binary pairwise-classification problem to estimate whether pairwise patterns are similar. Specifically, similarities between pairwise patterns are defined by the dot product between indicator features which are generated by a deep neural network (DNN). To learn informative representations for clustering, clustering constraints are imposed on the indicator features to represent specific concepts with specific representations. Since the ground-truth similarities are unavailable in clustering, an alternating iterative algorithm called Self-Evolution Clustering Training (SECT) is presented to select similar and dissimilar pairwise patterns and to train the DNN alternately. Consequently, the indicator features tend to be one-hot vectors and the patterns can be clustered by locating the largest response of the learned indicator features. Extensive experiments strongly evidence that DSEC outperforms current models on twelve popular image, text and audio datasets consistently. Jianlong Chang, Gaofeng Meng, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Local-Aggregation Graph NetworksabstractConvolutional neural networks (CNNs) provide a dramatically powerful class of models, but are subject to traditional convolution that can merely aggregate permutation-ordered and dimension-equal local inputs. It causes that CNNs are allowed to only manage signals on Euclidean or grid-like domains (e.g., images), not ones on non-Euclidean or graph domains (e.g., traffic networks). To eliminate this limitation, we develop a local-aggregation function, a sharable nonlinear operation, to aggregate permutation-unordered and dimension-unequal local inputs on non-Euclidean domains. In the context of the function approximation theory, the local-aggregation function is parameterized with a group of orthonormal polynomials in an effective and efficient manner. By replacing the traditional convolution in CNNs with the parameterized local-aggregation function, Local-Aggregation Graph Networks (LAGNs) are readily established, which enable to fit nonlinear functions without activation functions and can be expediently trained with the standard back-propagation. Extensive experiments on various datasets strongly demonstrate the effectiveness and efficiency of LAGNs, leading to superior performance on numerous pattern recognition and machine learning tasks, including text categorization, molecular activity detection, taxi flow prediction, and image classification. Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Baselines Extraction from Curved Document Images via Slope Fields RecoveryabstractBaselines estimation is a critical preprocessing step for many tasks of document image processing and analysis. The problem is very challenging due to arbitrarily complicated page layouts and various types of image quality degradations. This paper proposes a method based on slope fields recovery for curved baseline extraction from a distorted document image captured by a hand-held camera. Our method treats the curved baselines as the solution curves of an ordinary differential equation defined on a slope field. By assuming the page shape is a smooth and developable surface, we investigate a type of intrinsic geometric constraints of baselines to estimate the latent slope field. The curved baselines are finally obtained by solving an ordinary differential equation through the Euler method. Unlike the traditional text-lines based methods, our method is free from text-lines detection and segmentation. It can exploit multiple visual cues other than horizontal text-lines available in images for baselines extraction and is quite robust to document scripts, various types of image quality degradation (e.g., image distortion, blur and non-uniform illumination), large areas of non-textual objects and complex page layouts. Extensive experiments on synthetic and real-captured document images are implemented to evaluate the performance of the proposed method. Gaofeng Meng, Chunhong Pan, Shiming Xiang, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Geometric rectification of document images using adversarial gated unwarping network
Xiyan Liu, Gaofeng Meng, Bin Fan 0001, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 4 |
| 2020 | 3D PostureNet: A unified framework for skeleton-based posture recognition
Ying Wang 0008, Yongcheng Liu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. Lett. | 4 |
| 2020 | Kernel-Based Edge-Preserving Methods for Abrupt Change DetectionabstractAbrupt change detection is critical to monitor the occurrence of abnormal events from sensor data for situational awareness of complex systems. However, various disturbances and noises applied to the data observations may pose significant challenges to the robustness of many abrupt change detection methods. Recent researches have shown that bilateral filter can acquire outstanding performance on removing noises from images while preserving edge information. In this letter, we propose two improved edge-preserving memory-based cumulative sum (MB-CUSUM) methods that are able to make the abrupt change detection method more robust against noises. Our experimental studies show that the proposed methods can achieve superior performance over state-of-the-art methods to detect abrupt changes, which demonstrates the effectiveness and feasibility of their practical use. Shiming Xiang, Bo Tang 0011 |
IEEE Signal Process. Lett. | 1 |
| 2020 | Triplet Adversarial Domain Adaptation for Pixel-Level Classification of VHR Remote Sensing ImagesabstractPixel-level classification for very high resolution (VHR) images is a crucial but challenging task in remote sensing. However, since the diverse ways of satellite image acquisition and the distinct structures of various regions, the distributions of the same semantic classes among different data sets are dissimilar. Therefore, the classification model trained on one data set (source domain) may collapse, when it is directly applied to another one (target domain). To solve this problem, many adversarial-based domain adaptation methods have been proposed. However, these methods only consider the source and the target domains independently in the adversarial training, where only the target domain is explicitly contributed to narrow the gap between the distributions of both domains. Unlike previous methods, we propose a triplet adversarial domain adaptation (TriADA) method that jointly considers both domains to learn a domain-invariant classifier by a novel domain similarity discriminator. Specifically, the discriminator takes a triplet of segmentation maps as input, where two segmentation maps from the same domain are to be distinguished from the two maps from the different domains during the adversarial learning. Consequently, it explicitly considers both domains' information to narrow the distribution gap across domains. To enhance the discriminability of the classifier on the target domain, a class-aware self-training strategy, which depends on the output of the discriminator, is proposed to assign pseudo-labels with high adapted confidence on target data to retrain the classifier. Extensive experiments on several VHR pixel-level classification benchmarks demonstrate the effectiveness of our method as well as its superiority to the-state of the art. Bin Fan 0001, Hongmin Liu 0001, Chunlei Huo, Shiming Xiang, Chunhong Pan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | No-Reference Image Quality Assessment with Reinforcement Recursive List-Wise RankingabstractOpinion-unaware no-reference image quality assessment (NR-IQA) methods have received many interests recently because they do not require images with subjective scores for training. Unfortunately, it is a challenging task, and thus far no opinion-unaware methods have shown consistently better performance than the opinion-aware ones. In this paper, we propose an effective opinion-unaware NR-IQA method based on reinforcement recursive list-wise ranking. We formulate the NR-IQA as a recursive list-wise ranking problem which aims to optimize the whole quality ordering directly. During training, the recursive ranking process can be modeled as a Markov decision process (MDP). The ranking list of images can be constructed by taking a sequence of actions, and each of them refers to selecting an image for a specific position of the ranking list. Reinforcement learning is adopted to train the model parameters, in which no ground-truth quality scores or ranking lists are necessary for learning. Experimental results demonstrate the superior performance of our approach compared with existing opinion-unaware NR-IQA methods. Furthermore, our approach can compete with the most effective opinion-aware methods. It improves the state-of-the-art by over 2% on the CSIQ benchmark and outperforms most compared opinion-aware models on TID2013. Jie Gu 0002, Gaofeng Meng, Cheng Da, Shiming Xiang, Chunhong Pan |
AAAI | 4 |
| 2019 | Video Object Detection with Locally-Weighted Deformable NeighborsabstractDeep convolutional neural networks have achieved great success on various image recognition tasks. However, it is nontrivial to transfer the existing networks to video due to the fact that most of them are developed for static image. Frame-byframe processing is suboptimal because temporal information that is vital for video understanding is totally abandoned. Furthermore, frame-by-frame processing is slow and inefficient, which can hinder the practical usage. In this paper, we propose LWDN (Locally-Weighted Deformable Neighbors) for video object detection without utilizing time-consuming optical flow extraction networks. LWDN can latently align the high-level features between keyframes and keyframes or nonkeyframes. Inspired by (Zhu et al. 2017a) and (Hetang et al. 2017) who propose to aggregate features between keyframes and keyframes, we adopt brain-inspired memory mechanism to propagate and update the memory feature from keyframes to keyframes. We call this process Memory-Guided Propagation. With such a memory mechanism, the discriminative ability of features in keyframes and non-keyframes are both enhanced, which helps to improve the detection accuracy. Extensive experiments on VID dataset demonstrate that our method achieves superior performance in a speed and accuracy trade-off, i.e., 76.3% on the challenging VID dataset while maintaining 20fps in speed on Titan X GPU. Zhengkai Jiang 0001, Peng Gao 0007, Chaoxu Guo, Qian Zhang 0009, Shiming Xiang, Chunhong Pan |
AAAI | 5 |
| 2019 | What and Where the Themes Dominate in ImageabstractThe image captioning is to describe an image with natural language as human, which has benefited from the advances in deep neural network and achieved substantial progress in performance. However, the perspective of human description to scene has not been fully considered in this task recently. Actually, the human description to scene is tightly related to the endogenous knowledge and the exogenous salient objects simultaneously, which implies that the content in the description is confined to the known salient objects. Inspired by this observation, this paper proposes a novel framework, which explicitly applies the known salient objects in image captioning. Under this framework, the known salient objects are served as the themes to guide the description generation. According to the property of the known salient object, a theme is composed of two components: its endogenous concept (what) and the exogenous spatial attention feature (where). Specifically, the prediction of each word is dominated by the concept and spatial attention feature of the corresponding theme in the process of caption prediction. Moreover, we introduce a novel learning method of Distinctive Learning (DL) to get more specificity of generated captions like human descriptions. It formulates two constraints in the theme learning process to encourage distinctiveness between different images. Particularly, reinforcement learning is introduced into the framework to address the exposure bias problem between the training and the testing modes. Extensive experiments on the COCO and Flickr30K datasets achieve superior results when compared with the state-of-the-art methods. Xinyu Xiao, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
AAAI | 3 |
| 2019 | RENAS: Reinforced Evolutionary Neural Architecture SearchabstractNeural Architecture Search (NAS) is an important yet challenging task in network design due to its high computational consumption. To address this issue, we propose the Reinforced Evolutionary Neural Architecture Search (RENAS), which is an evolutionary method with reinforced mutation for NAS. Our method integrates reinforced mutation into an evolution algorithm for neural architecture exploration, in which a mutation controller is introduced to learn the effects of slight modifications and make mutation actions. The reinforced mutation controller guides the model population to evolve efficiently. Furthermore, as child models can inherit parameters from their parents during evolution, our method requires very limited computational resources. In experiments, we conduct the proposed search method on CIFAR-10 and obtain a powerful network architecture, RENASNet. This architecture achieves a competitive result on CIFAR-10. The explored network architecture is transferable to ImageNet and achieves a new state-of-the-art accuracy, i.e., 75.7% top-1 accuracy with 5.36M parameters on mobile ImageNet. We further test its performance on semantic segmentation with DeepLabv3 on the PASCAL VOC. RENASNet outperforms MobileNet-v1, MobileNet-v2 and NASNet. It achieves 75.83% mIOU without being pretrained on COCO. Yukang Chen, Gaofeng Meng, Qian Zhang 0009, Shiming Xiang, Chang Huang, Lisen Mu, Xinggang Wang |
CVPR | 4 |
| 2019 | Relation-Shape Convolutional Neural Network for Point Cloud AnalysisabstractPoint cloud analysis is very challenging, as the shape implied in irregular points is difficult to capture. In this paper, we propose RS-CNN, namely, Relation-Shape Convolutional Neural Network, which extends regular grid CNN to irregular configuration for point cloud analysis. The key to RS-CNN is learning from relation, i.e., the geometric topology constraint among points. Specifically, the convolutional weight for local point set is forced to learn a high-level relation expression from predefined geometric priors, between a sampled point from this point set and the others. In this way, an inductive local representation with explicit reasoning about the spatial layout of points can be obtained, which leads to much shape awareness and robustness. With this convolution as a basic operator, RS-CNN, a hierarchical architecture can be developed to achieve contextual shape-aware learning for point cloud analysis. Extensive experiments on challenging benchmarks across three tasks verify RS-CNN achieves the state of the arts. Yongcheng Liu, Bin Fan 0001, Shiming Xiang, Chunhong Pan |
CVPR | 3 |
| 2019 | Guiding the Flowing of Semantics: Interpretable Video Captioning via POS TagabstractXinyu Xiao, Lingfeng Wang, Bin Fan, Shinming Xiang, Chunhong Pan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xinyu Xiao, Lingfeng Wang 0002, Bin Fan 0001, Shiming Xiang, Chunhong Pan |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Progressive Sparse Local Attention for Video Object DetectionabstractTransferring image-based object detectors to the domain of videos remains a challenging problem. Previous efforts mostly exploit optical flow to propagate features across frames, aiming to achieve a good trade-off between accuracy and efficiency. However, introducing an extra model to estimate optical flow can significantly increase the overall model size. The gap between optical flow and high-level features can also hinder it from establishing spatial correspondence accurately. Instead of relying on optical flow, this paper proposes a novel module called Progressive Sparse Local Attention (PSLA), which establishes the spatial correspondence between features across frames in a local region with progressively sparser stride and uses the correspondence to propagate features. Based on PSLA, Recursive Feature Updating (RFU) and Dense Feature Transforming (DenseFT) are proposed to model temporal appearance and enrich feature representation respectively in a novel video object detection framework. Experiments on ImageNet VID show that our method achieves the best accuracy compared to existing methods with smaller model size and acceptable runtime speed. Chaoxu Guo, Bin Fan 0001, Jie Gu 0002, Qian Zhang 0009, Shiming Xiang, Véronique Prinet, Chunhong Pan |
ICCV | 5 |
| 2019 | DensePoint: Learning Densely Contextual Representation for Efficient Point Cloud ProcessingabstractPoint cloud processing is very challenging, as the diverse shapes formed by irregular points are often indistinguishable. A thorough grasp of the elusive shape requires sufficiently contextual semantic information, yet few works devote to this. Here we propose DensePoint, a general architecture to learn densely contextual representation for point cloud processing. Technically, it extends regular grid CNN to irregular point configuration by generalizing a convolution operator, which holds the permutation invariance of points, and achieves efficient inductive learning of local patterns. Architecturally, it finds inspiration from dense connection mode, to repeatedly aggregate multi-level and multi-scale semantics in a deep hierarchy. As a result, densely contextual information along with rich semantics, can be acquired by DensePoint in an organic manner, making it highly effective. Extensive experiments on challenging benchmarks across four tasks, as well as thorough model analysis, verify DensePoint achieves the state of the arts. Yongcheng Liu, Bin Fan 0001, Gaofeng Meng, Jiwen Lu, Shiming Xiang, Chunhong Pan |
ICCV | 5 |
| 2019 | GSTNet: Global Spatial-Temporal Network for Traffic Flow PredictionabstractPredicting traffic flow on traffic networks is a very challenging task, due to the complicated and dynamic spatial-temporal dependencies between different nodes on the network. The traffic flow renders two types of temporal dependencies, including short-term neighboring and long-term periodic dependencies. What's more, the spatial correlations over different nodes are both local and non-local. To capture the global dynamic spatial-temporal correlations, we propose a Global Spatial-Temporal Network (GSTNet), which consists of several layers of spatial-temporal blocks. Each block contains a multi-resolution temporal module and a global correlated spatial module in sequence, which can simultaneously extract the dynamic temporal dependencies and the global spatial correlations. Extensive experiments on the real world datasets verify the effectiveness and superiority of the proposed method on both the public transportation network and the road network. Shen Fang, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
IJCAI | 4 |
| 2019 | DATA: Differentiable ArchiTecture ApproximationabstractNeural architecture search (NAS) is inherently subject to the gap of architectures during searching and validating. To bridge this gap, we develop Differentiable ArchiTecture Approximation (DATA) with an Ensemble Gumbel-Softmax (EGS) estimator to automatically approximate architectures during searching and validating in a differentiable manner. Technically, the EGS estimator consists of a group of Gumbel-Softmax estimators, which is capable of converting probability vectors to binary codes and passing gradients from binary codes to probability vectors. Benefiting from such modeling, in searching, architecture parameters and network weights in the NAS model can be jointly optimized with the standard back-propagation, yielding an end-to-end learning mechanism for searching deep models in a large enough search space. Conclusively, during validating, a high-performance architecture that approaches to the learned one during searching is readily built. Extensive experiments on a variety of popular datasets strongly evidence that our method is capable of discovering high-performance architectures for image classification, language modeling and semantic segmentation, while guaranteeing the requisite efficiency during searching. Jianlong Chang, Xinbang Zhang, Yiwen Guo, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
NeurIPS | 5 |
| 2019 | Incremental Poisson Surface Reconstruction for Large Scale Three-Dimensional Modeling
Wei Sui, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
PRCV (3) | 4 |
| 2019 | Parameter optimization criteria guided 3D point cloud classification
Hongjun Li 0002, Weiliang Meng, Shiming Xiang, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2019 | Nonlinear Asymmetric Multi-Valued HashingabstractMost existing hashing methods resort to binary codes for large scale similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose Nonlinear Asymmetric Multi-Valued Hashing (NAMVH) supported by two distinct non-binary embeddings. Specifically, a real-valued embedding is used for representing the newly-coming query by an ideally nonlinear transformation. Besides, a multi-integer-embedding is employed for compressing the whole database, which is modeled by Binary Sparse Representation (BSR) with fixed sparsity. With these two non-binary embeddings, NAMVH preserves more precise similarities between data points and enables access to the incremental extension with database samples evolving dynamically. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the pairwise label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by a well-designed alternative optimization method. Extensive experiments on seven large scale datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency. Cheng Da, Gaofeng Meng, Shiming Xiang, Kun Ding 0001, Shibiao Xu, Qing Yang 0002, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Blind image quality assessment via learnable attention-based pooling
Jie Gu 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 3 |
| 2019 | Dense semantic embedding network for image captioning
Xinyu Xiao, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 4 |
| 2019 | Pseudo low rank video representation
Tingzhao Yu, Lingfeng Wang 0002, Chaoxu Guo, Huxiang Gu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 5 |
| 2019 | Learning graph structure via graph convolutional networks
Jianlong Chang, Gaofeng Meng, Shibiao Xu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 5 |
| 2019 | A Performance Evaluation of Local Features for Image-Based 3D ReconstructionabstractThis paper performs a comprehensive and comparative evaluation of the state-of-the-art local features for the task of image-based 3D reconstruction. The evaluated local features cover the recently developed ones by using powerful machine learning techniques and the elaborately designed handcrafted features. To obtain a comprehensive evaluation, we choose to include both float type features and binary ones. Meanwhile, two kinds of datasets have been used in this evaluation. One is a dataset of many different scene types with groundtruth 3D points, containing images of different scenes captured at fixed positions, for quantitative performance evaluation of different local features in the controlled image capturing situation. The other dataset contains Internet scale image sets of several landmarks with a lot of unrelated images, which is used for qualitative performance evaluation of different local features in the free image collection situation. Our experimental results show that binary features are competent to reconstruct scenes from controlled image sequences with only a fraction of processing time compared to using float type features. However, for the case of a large scale image set with many distracting images, float type features show a clear advantage over binary ones. Currently, the most traditional SIFT is very stable with regard to scene types in this specific task and produces very competitive reconstruction results among all the evaluated local features. Meanwhile, although the learned binary features are not as competitive as the handcrafted ones, learning float type features with CNN is promising but still requires much effort in the future. Bin Fan 0001, Qingqun Kong, Xinchao Wang, Zhiheng Wang 0001, Shiming Xiang, Chunhong Pan, Pascal Fua |
IEEE Trans. Image Process. | 5 |
| 2019 | Deep Hierarchical Encoder-Decoder Network for Image CaptioningabstractEncoder-decoder models have been widely used in image captioning, and most of them are designed via single long short term memory (LSTM). The capacity of single-layer network, whose encoder and decoder are integrated together, is limited for such a complex task of image captioning. Moreover, how to effectively increase the “vertical depth” of encoder-decoder remains to be solved. To deal with these problems, a novel deep hierarchical encoder-decoder network is proposed for image captioning, where a deep hierarchical structure is explored to separate the functions of encoder and decoder. This model is capable of efficiently exerting the representation capacity of deep networks to fuse high level semantics of vision and language in generating captions. Specifically, visual representations in top levels of abstraction are simultaneously considered, and each of these levels is associated to one LSTM. The bottom-most LSTM is applied as the encoder of textual inputs. The application of the middle layer in encoder-decoder is to enhance the decoding ability of top-most LSTM. Furthermore, depending on the introduction of semantic enhancement module of image feature and distribution combine module of text feature, variants of architectures of our model are constructed to explore the impacts and mutual interactions among the visual representation, textual representations, and the output of the middle LSTM layer. Particularly, the framework is training under a reinforcement learning method to address the exposure bias problem between the training and the testing by the policy gradient optimization. Qualitative analyses indicate the process that our model “translates” image to sentence and further visualization presents the evolution of the hidden states from different hierarchical LSTMs over time. Extensive experiments demonstrate that our model outperforms current state-of-the-art models on three benchmark datasets: Flickr8K, Flickr30K, and MSCOCO. On both image captioning and retrieval tasks, our method achieves the best results. On MSCOCO captioning Leaderboard, our method also achieves superior performance. Xinyu Xiao, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 4 |
| 2019 | Weakly Semantic Guided Action RecognitionabstractAction recognition plays a fundamental role in computer vision and video analysis. Nevertheless, extracting effective spatial-temporal features remains a challenging task. This paper proposes three simple but effective weakly semantic guided modules (SGMs) for both environment-constrained and cross-domain action recognition. The SGMs are composed of total 3-D convolution and element-wise gated operations; thus, they are efficient and easy to implement. The semantic guidance is obtained in a weakly supervised manner, in which each video clip is labeled with only an action class instead of pixel-level semantics. Benefitting from the semantic guidance, the network [called semantic guided network (SGN)] can focus on the salient parts of the video clips. Consequently, the redundant information can be reduced and the model is more robust to noise. Besides, benefitting from the intrinsic property of SGMs, SGN is totally end-to-end trainable. Quantities of experiments on both environment-constrained (e.g., Penn, HMDB-51, and UCF101) and cross-domain (e.g., ODAR) action recognition datasets demonstrate its effectiveness. Specifically, SGN gets improvements of 3.7%, 2.1%, and 5.2% for Penn, HMDB-51, and UCF-101 than the baseline ResNet3D, respectively, and SGN ranked third place in the ODAR 2017 challenge. Tingzhao Yu, Lingfeng Wang 0002, Cheng Da, Huxiang Gu, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 5 |
| 2018 | Exploiting Vector Fields for Geometric Rectification of Distorted Document Images
Gaofeng Meng, Yuanqi Su, Ying Wu 0001, Shiming Xiang, Chunhong Pan |
ECCV (16) | 4 |
| 2018 | Mgn: Multi-Glimpse Network for Action RecognitionabstractCurrent state-of-the-art action recognition approaches rely on optical flow to extract the local motion information and ignore the importance of global description of the videos. In this paper, we present a novel architecture, named Multi-Glimpse Network (MGN), to boost the performance of action recognition by combining the local and global information of the videos. Specifically, MGN makes predictions through two important modules, Local Glimpse and Global Glimpse. Local Glimpse extracts the local spatiotemporal features of different periods using temporal sampling method. Global Glimpse aggregates the extracted local features to develop global description of the videos. These two modules are complementary and indispensable. Our MGN achieves competitive results on four video action benchmarks of UCF10l, HMDB51, ODAR and Penn. Chaoxu Guo, Tingzhao Yu, Huxiang Gu, Shiming Xiang, Chunhong Pan |
ICASSP | 4 |
| 2018 | Adversarial Domain Adaptation with a Domain Similarity Discriminator for Semantic Segmentation of Urban AreasabstractExisting semantic segmentation models of urban areas have shown to perform well in a supervised setting. However, collecting lots of annotated images from each city to train such models is time-consuming or difficult. In addition, when transferring the segmentation model from the trained city (source domain) to an unseen city (target domain), the performance will largely degrade due to the domain shift. For this reason, we propose a domain adaptation method with a domain similarity discriminator to eliminate such domain shift in the framework of adversarial learning. Contrary to the single-input adversarial network, our domain similarity discriminator, which consists of a Siamese network, is able to measure the similarity of the pairwise-input data. In this way, we can use more information about the pairwise-input to measure the similarity between different distributions so as to address the problem of domain shift. Experimental results demonstrate that our approach outperforms the competing methods on three different cities. Bin Fan 0001, Shiming Xiang, Chunhong Pan |
ICIP | 3 |
| 2018 | Semantic Image Synthesis via Conditional Cycle-Generative Adversarial NetworksabstractTraditional approaches for semantic image synthesis mainly focus on text descriptions while ignoring the related structures and attributes in the original images. Therefore, some critical information, e.g., the style, backgrounds, objects shapes and pose, is missed in the generated images. In this paper, we propose a novel framework called Conditional Cycle-Generative Adversarial Network (CCGAN) to address this issue. Our model can generate photo-realistic images conditioned on the given text descriptions, while maintaining the attributes of the original images. The framework mainly consists of two coupled conditional adversarial networks, which are able to learn a desirable image mapping that can keep the structures and attributes in the images. We introduce a conditional cycle consistency loss to prevent the contradiction between two generators. This loss allows the generated images to retain most of the features of the original image, so as to improve the stability of network training. Moreover, benefiting from the mechanism of circular training, the proposed networks can learn the semantic information of the text much accurately. Experiments on Caltech-UCSD Bird dataset and Oxford-102 flower dataset demonstrate that the proposed method significantly outperforms the existing methods in terms of image details reconstruction and semantic information expression. Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICPR | 3 |
| 2018 | Kernel-Weighted Graph Convolutional Network: A Deep Learning Approach for Traffic ForecastingabstractTraffic forecasting is of great significance and has many applications in Intelligent Traffic System (ITS). In spite of many thoughtful attempts in the past decades, this task still remains far from being solved, due to the diversity, complexity and nonlinearity of traffic situations. Technically, it can be cast on the framework of regressions with spatial-template data. Typically, one may consider to employ the Convolutional Neural Network (CNN) to achieve this goal. Unfortunately, the traditional CNN is developed for grid data. By contrast, here we are facing with non-grid traffic data points that are observed spatially at locations of interest. To this end, this paper proposes a novel Kernel-Weighted Graph Convolutional Network (KW-GCN) for traffic forecasting, which learns simultaneously a group of convolutional kernels and their linear combination weights for each of the nodes in the graph. This yields a mechanism that is able to learn the features locally and exploit the structure information of traffic road-network globally. By introducing additional parameters, our KW-GCN can relax the restriction of weight sharing in classical CNN to better handle the traffic data of non-stationarity. Furthermore, it has been illustrated that the proposed linear weighting of kernels can be viewed as the low-rank decomposition of the well-known locally-connected networks, and thus it avoids over-fitting to some degree. We apply our approach to the real-world GPS data set of about 30,000 taxis in seven months in Beijing. Experiments on both taxi-flow forecasting and road-speed forecasting demonstrate that our method significantly outperforms the state-of-the-art ones. Qizhao Jin, Jianlong Chang, Shiming Xiang, Chunhong Pan |
ICPR | 4 |
| 2018 | SafeNet: Scale-normalization and Anchor-based Feature Extraction Network for Person Re-identificationabstractPerson Re-identification (ReID) is a challenging retrieval task that requires matching a person's image across non-overlapping camera views. The quality of fulfilling this task is largely determined on the robustness of the features that are used to describe the person. In this paper, we show the advantage of jointly utilizing multi-scale abstract information to learn powerful features over full body and parts. A scale normalization module is proposed to balance different scales through residual-based integration. To exploit the information hidden in non-rigid body parts, we propose an anchor-based method to capture the local contents by stacking convolutions of kernels with various aspect ratios, which focus on different spatial distributions. Finally, a well-defined framework is constructed for simultaneously learning the representations of both full body and parts. Extensive experiments conducted on current challenging large-scale person ReID datasets, including Market1501, CUHK03 and DukeMTMC, demonstrate that our proposed method achieves the state-of-the-art results. Kun Yuan 0003, Qian Zhang 0009, Chang Huang, Shiming Xiang, Chunhong Pan |
IJCAI | 4 |
| 2018 | Multi-Label Image Classification via Knowledge Distillation from Weakly-Supervised DetectionabstractMulti-label image classification is a fundamental but challenging task towards general visual understanding. Existing methods found the region-level cues (e.g., features from RoIs) can facilitate multi-label classification. Nevertheless, such methods usually require laborious object-level annotations (i.e., object labels and bounding boxes) for effective learning of the object-level visual features. In this paper, we propose a novel and efficient deep framework to boost multi-label classification by distilling knowledge from weakly-supervised detection task without bounding box annotations. Specifically, given the image-level annotations, (1) we first develop a weakly-supervised detection (WSD) model, and then (2) construct an end-to-end multi-label image classification framework augmented by a knowledge distillation module that guides the classification model by the WSD model according to the class-level predictions for the whole image and the object-level visual features for object RoIs. The WSD model is the teacher model and the classification model is the student model. After this cross-task knowledge distillation, the performance of the classification model is significantly improved and the efficiency is maintained since the WSD model can be safely discarded in the test phase. Extensive experiments on two large-scale datasets (MS-COCO and NUS-WIDE) show that our framework achieves superior performances over the state-of-the-art methods on both performance and efficiency. Yongcheng Liu, Lu Sheng, Shiming Xiang, Chunhong Pan |
ACM Multimedia | 5 |
| 2018 | Structure-Aware Convolutional Neural NetworksabstractConvolutional neural networks (CNNs) are inherently subject to invariable filters that can only aggregate local inputs with the same topological structures. It causes that CNNs are allowed to manage data with Euclidean or grid-like structures (e.g., images), not ones with non-Euclidean or graph structures (e.g., traffic networks). To broaden the reach of CNNs, we develop structure-aware convolution to eliminate the invariance, yielding a unified mechanism of dealing with both Euclidean and non-Euclidean structured data. Technically, filters in the structure-aware convolution are generalized to univariate functions, which are capable of aggregating local inputs with diverse topological structures. Since infinite parameters are required to determine a univariate function, we parameterize these filters with numbered learnable parameters in the context of the function approximation theory. By replacing the classical convolution in CNNs with the structure-aware convolution, Structure-Aware Convolutional Neural Networks (SACNNs) are readily established. Extensive experiments on eleven datasets strongly evidence that SACNNs outperform current models on various machine learning tasks, including image classification and clustering, text categorization, skeleton-based action recognition, molecular activity detection, and taxi flow prediction. Jianlong Chang, Jie Gu 0002, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
NeurIPS | 5 |
| 2018 | Deep unsupervised learning with consistent inference of latent representations
Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 4 |
| 2018 | Joint spatial-temporal attention for action recognition
Tingzhao Yu, Chaoxu Guo, Lingfeng Wang 0002, Huxiang Gu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. Lett. | 5 |
| 2018 | Deep generative video prediction
Tingzhao Yu, Lingfeng Wang 0002, Huxiang Gu, Shiming Xiang, Chunhong Pan |
Pattern Recognit. Lett. | 4 |
| 2018 | Self-Paced AutoEncoderabstractAutoencoder, which learns latent representations of samples in an unsupervised manner, has great potential in computer vision and signal processing. However, the diversity of samples makes learning a component autoencoder remaining a challenging task. This letter proposes a novel Self-Paced AutoEncoder (SPAE) for unsupervised feature extraction. The motivation behind this letter is to take samples gradually from simple to complex into consideration during training, which is similar to the mechanism of knowledge acquisition for humans. Under the unsupervised learning framework constructed on the autoencoder infrastructure, our SPAE first learns a weak autoencoder via samples with small losses and, then, elevates itself to a relatively strong autoencoder through samples with large losses. Then, the SPAE is generalized to a temporal domain, resulting to temporal SPAE (TSPAE), where the temporal information is explored and exploited to improve the performance. Typically, a TSPAE is capable of compressing temporal sequences into temporal-independent data. Experiments on the image classification and action recognition demonstrate the effectiveness of SPAE and TSPAE. Tingzhao Yu, Chaoxu Guo, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
IEEE Signal Process. Lett. | 4 |
| 2018 | Automatic Building Rooftop Extraction From Aerial Images via Hierarchical RGB-D PriorsabstractAccurate building rooftop extraction from high-resolution aerial images is of crucial importance in a wide range of applications. Owing to the varying appearance and large-scale range of scene objects, especially for building rooftops in different scales and heights, single-scale or individual prior-based extraction technique is insufficient in pursuing efficient, generic, and accurate extraction results. The trend toward integrating multiscale or several cue techniques appears to be the best way; thus, such integration is the focus of this paper. We first propose a novel salient rooftop detector integrating four correlative RGB-D priors (depth cue, uniqueness prior, shape prior, and transition surface prior) for improved rooftop extraction to address the preceding complex issues mentioned. Then, these correlative cues are computed from image layers created by our multilevel segmentation and further fused into the state-of-the-art high-order conditional random field (CRF) framework to locate the rooftop. Finally, an iterative optimization strategy is applied for high-quality solving, which can robustly handle varying appearance of building rooftops. Performance evaluations in the SZTAKI-INRIA benchmark data sets show that our method outperforms the traditional color-based algorithm and the original high-order CRF algorithm and its variants. The proposed algorithm is also evaluated and found to produce consistently satisfactory results for various large-scale, real-world data sets. Shibiao Xu, Xingjia Pan, Er Li, Baoyuan Wu, Shuhui Bu, Weiming Dong, Shiming Xiang, Xiaopeng Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2018 | Blind Image Quality Assessment via Vector Regression and Object Oriented PoolingabstractThis paper presents an effective method based on vector regression and object oriented pooling for blind image quality assessment. Unlike previous models that map the extracted features directly to a quality score, the proposed vector regression framework yields a vector of belief scores for the input image. We explore the uncertainty factors in quality assessment and design the belief scores to measure the confidences of an image to be assigned to the corresponding quality grades. Moreover, we propose an object oriented pooling strategy to further improve the performance by incorporating semantic information of image contents. According to this strategy, regions occupied by objects will be assigned more weights in the pooling phase, leading to a more accurate quality assessment. Extensive experiments on benchmark datasets demonstrate that our approach achieves state-of-the-art performance and shows a great generalization ability. Jie Gu 0002, Gaofeng Meng, Judith Redi, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 4 |
| 2018 | In Defense of Locality-Sensitive HashingabstractHashing-based semantic similarity search is becoming increasingly important for building large-scale content-based retrieval system. The state-of-the-art supervised hashing techniques use flexible two-step strategy to learn hash functions. The first step learns binary codes for training data by solving binary optimization problems with millions of variables, thus usually requiring intensive computations. Despite simplicity and efficiency, locality-sensitive hashing (LSH) has never been recognized as a good way to generate such codes due to its poor performance in traditional approximate neighbor search. We claim in this paper that the true merit of LSH lies in transforming the semantic labels to obtain the binary codes, resulting in an effective and efficient two-step hashing framework. Specifically, we developed the locality-sensitive two-step hashing (LS-TSH) that generates the binary codes through LSH rather than any complex optimization technique. Theoretically, with proper assumption, LS-TSH is actually a useful LSH scheme, so that it preserves the label-based semantic similarity and possesses sublinear query complexity for hash lookup. Experimentally, LS-TSH could obtain comparable retrieval accuracy with state of the arts with two to three orders of magnitudes faster training speed. Kun Ding 0001, Chunlei Huo, Bin Fan 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | AMVH: Asymmetric Multi-Valued hashingabstractMost existing hashing methods resort to binary codes for similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose an asymmetric multi-valued hashing method supported by two different non-binary embeddings. (1) A real-valued embedding is used for representing the newly-coming query. (2) A multi-integer-embedding is employed for compressing the whole database, which is modeled by binary sparse representation with fixed sparsity. With these two non-binary embeddings, the similarities between data points can be preserved precisely. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by alternative optimization. Extensive experiments on three multilabel datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency. Cheng Da, Shibiao Xu, Kun Ding 0001, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
CVPR | 5 |
| 2017 | Deep Adaptive Image ClusteringabstractImage clustering is a crucial but challenging task in machine learning and computer vision. Existing methods often ignore the combination between feature learning and clustering. To tackle this problem, we propose Deep Adaptive Clustering (DAC) that recasts the clustering problem into a binary pairwise-classification framework to judge whether pairs of images belong to the same clusters. In DAC, the similarities are calculated as the cosine distance between label features of images which are generated by a deep convolutional network (ConvNet). By introducing a constraint into DAC, the learned label features tend to be one-hot vectors that can be utilized for clustering images. The main challenge is that the ground-truth similarities are unknown in image clustering. We handle this issue by presenting an alternating iterative Adaptive Learning algorithm where each iteration alternately selects labeled samples and trains the ConvNet. Conclusively, images are automatically clustered based on the label features. Experimental results show that DAC achieves state-of-the-art performance on five popular datasets, e.g., yielding 97.75% clustering accuracy on MNIST, 52.18% on CIFAR-10 and 46.99% on STL-10. Jianlong Chang, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICCV | 4 |
| 2017 | Deep Networks for Degraded Document Image Binarization through Pyramid ReconstructionabstractBinarization of document images is an important processing step for document images analysis and recognition. However, this problem is quite challenging in some cases because of the quality degradation of document images, such as varying illumination, complicated backgrounds, image noises due to ink spots, water stains or document creases. In this paper, we propose a framework based on deep convolutional neural-network (DCNN) for adaptive binarization of degraded document images. The basic idea of our method is to decompose a degraded document image into a spatial pyramid structure by using DCNN, with each layer at different scale. Then the foreground image is sequentially reconstructed from these layers in a coarse-to-fine manner by using deconvolutional network. Such kind of decomposition is quite beneficial, since multi-resolution supervision information can be directly introduced into network learning. We also define several loss functions about label consistency and foregrounds smoothing to further regularize the training of the network. Experimental results demonstrate the effectiveness of the proposed method. Gaofeng Meng, Kun Yuan 0003, Ying Wu 0001, Shiming Xiang, Chunhong Pan |
ICDAR | 4 |
| 2017 | Efficient similarity learning for asymmetric hashingabstractHashing techniques with asymmetric schemes (e.g., only bi-narizing the database points) have recently attracted wide attention in the circle of image retrieval. In comparison with those methods which binarize simultaneously both of the query and database points, they not only enjoy the storage and search efficiencies, but also provide higher accuracy. Gearing to this line, this paper proposes a metric-embedded asymmetric hashing (MEAH) that learns jointly a bilinear similarity measure and binary codes of database points in an unsupervised manner. Technically, the learned similarity measure is able to bridge the gap between the binary codes and the real-valued codes, which are represented possibly with different dimensions. What is more, this measure is capable of preserving the global structure hidden in the database. Extensive experiments on two public image benchmarks demonstrate the superiority of our approach over the several state-of-the-art unsupervised hashing methods. Cheng Da, Yang Yang 0062, Chunlei Huo, Shiming Xiang, Chunhong Pan |
ICIP | 5 |
| 2017 | Context-aware cascade network for semantic labeling in VHR imageabstractSemantic labeling for the very high resolution (VHR) image of urban areas is challenging, because of many complex manmade objects with different materials and fine-structured objects located together. Under the framework of convolutional neural networks (CNNs), this paper proposes a novel end-to-end network for semantic labeling. Specifically, our network not only improves the labeling accuracy of complex manmade objects by aggregating multiple context semantics with a cascaded architecture, but also refines fine-structured objects by utilizing the low-level detail in shallow layers of CNNs with a hierarchical pyramid structure. Throughout the network, a dedicated residual correction scheme is employed to amend the latent fitting residual. As a result of these specific components, the whole model works in a global-to-local and coarse-to-fine manner. Experimental results show that our network outperforms the state-of-the-art methods on the large-scale ISPRS Vaihingen 2D Semantic Labeling Challenge dataset. Yongcheng Liu, Bin Fan 0001, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
ICIP | 5 |
| 2017 | Cascaded temporal spatial features for video action recognitionabstractExtracting spatial-temporal descriptors is a challenging task for video-based human action recognition. We decouple the 3D volume of video frames directly into a cascaded temporal spatial domain via a new convolutional architecture. The motivation behind this design is to achieve deep nonlinear feature representations with reduced network parameters. First, a 1D temporal network with shared parameters is first constructed to map the video sequences along the time axis into feature maps in temporal domain. These feature maps are then organized into channels like those of RGB image (named as Motion Image here for abbreviation), which is desired to preserve both temporal and spatial information. Second, the Motion Image is regarded as the input of the latter cascaded 2D spatial network. With the combination of the 1D temporal network and the 2D spatial network together, the size of whole network parameters is largely reduced. Benefiting from the Motion Image, our network is an end-to-end system for the task of action recognition, which can be trained with the classical algorithm of back propagation. Quantities of comparative experiments on two benchmark datasets demonstrate the effectiveness of our new architecture. Tingzhao Yu, Huxiang Gu, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
ICIP | 4 |
| 2017 | Efficient cloud detection in remote sensing images using edge-aware segmentation network and easy-to-hard training strategyabstractDetecting cloud regions in remote sensing image (RSI) is very challenging yet of great importance to meteorological forecasting and other RSI-related applications. Technically, this task is typically implemented as a pixel-level segmentation. However, traditional methods based on handcrafted or low-level cloud features often fail to achieve satisfactory performances from images with bright non-cloud and/or semitransparent cloud regions. What is more, the performances could be further degraded due to the ambiguous boundaries caused by complicated textures and non-uniform distribution of intensities. In this paper, we propose a multi-task based deep neural network for cloud detection in RSIs. Architecturally, our network is designed to combine the two tasks of cloud segmentation and cloud edge detection together to encourage a better detection near cloud boundaries, resulting in an end-to-end approach for accurate cloud detection. Accordingly, an efficient sample selection strategy is proposed to train our network in an easy-to-hard manner, in which the number of the selected samples is governed by a weight that is annealed until the entire training samples have been considered. Both visual and quantitative comparisons are conducted on RSIs collected from Google Earth. The experimental results indicate that our method can yield superior performance over the state-of-the-art methods. Kun Yuan 0003, Gaofeng Meng, Dongcai Cheng, Shiming Xiang, Chunhong Pan |
ICIP | 5 |
| 2017 | Structured binary feature extraction for hyperspectral imagery classificationabstractIn this paper, we propose a novel structured binary feature extraction method for hyperspectral image classification. To pursuit high discriminative ability and low memory cost, we resort to applying the learning to hash technique to the traditional spectral-spatial hyperspectral features. We show how the structured information among different kinds of features and different feature groups can be used to learn discriminative binary features for classification. Experiments on two standard benchmark hyperspectral data sets demonstrate the effectiveness of the proposed method. Zisha Zhong, Bin Fan 0001, Shiming Xiang, Chunhong Pan |
ICIP | 4 |
| 2017 | Active Rectification of Curved Document Images Using Structured Beams
Gaofeng Meng, Shiming Xiang, Chunhong Pan, Nanning Zheng 0001 |
Int. J. Comput. Vis. | 2 |
| 2017 | Building Regional Covariance Descriptors for Vehicle DetectionabstractWe study the question of building regional covariance descriptors (RCDs) for vehicle detection from high-resolution satellite images. A unified way is proposed to build RCD features by constant convolutional kernels in the forms of 2-D masks. Two novel formulas are designed to construct different RCD types based upon one or two convolutional masks, obtaining ten novel RCD features by four simple constant convolutional masks. Experiments show that such convolutional-mask-based RCDs outperform the previous image-derivative-based RCDs, the popular local binary patterns (LBPs), the histogram of oriented gradients (HOGs), and LBP+HOG. Furthermore, feeding to nonlinear support vector machines (SVMs) of two kernel types [L1kernel and radial basis function (RBF)], these RCDs outperform four known deep convolutional neural networks: AlexNet, GoogLeNet, CaffeNet, and LeNet, as well as their fine-tuned models by their well-trained weights of imageNet classification. Among three popular classic classifiers we have tested in the experiments, nonlinear SVMs outperform BP and Adaboost obviously, and L1kernel exceeds RBF slightly. Xueyun Chen, Ren-Xi Gong, Ling-Ling Xie, Shiming Xiang, Cheng-Lin Liu 0001, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | Automatic Road Detection and Centerline Extraction via Cascaded End-to-End Convolutional Neural NetworkabstractAccurate road detection and centerline extraction from very high resolution (VHR) remote sensing imagery are of central importance in a wide range of applications. Due to the complex backgrounds and occlusions of trees and cars, most road detection methods bring in the heterogeneous segments; besides for the centerline extraction task, most current approaches fail to extract a wonderful centerline network that appears smooth, complete, as well as single-pixel width. To address the above-mentioned complex issues, we propose a novel deep model, i.e., a cascaded end-to-end convolutional neural network (CasNet), to simultaneously cope with the road detection and centerline extraction tasks. Specifically, CasNet consists of two networks. One aims at the road detection task, whose strong representation ability is well able to tackle the complex backgrounds and occlusions of trees and cars. The other is cascaded to the former one, making full use of the feature maps produced formerly, to obtain the good centerline extraction. Finally, a thinning algorithm is proposed to obtain smooth, complete, and single-pixel width road centerline network. Extensive experiments demonstrate that CasNet outperforms the state-of-the-art methods greatly in learning quality and learning speed. That is, CasNet exceeds the comparing methods by a large margin in quantitative performance, and it is nearly 25 times faster than the comparing methods. Moreover, as another contribution, a large and challenging road centerline data set for the VHR remote sensing image will be publicly available for further studies. Ying Wang 0008, Shibiao Xu, Hongzhen Wang, Shiming Xiang, Chunhong Pan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2017 | Cross-Modal Hashing via Rank-Order PreservingabstractDue to the query effectiveness and efficiency, cross-modal similarity search based on hashing has acquired extensive attention in the multimedia community. Most existing methods do not explicitly employ the ranking information when learning hash functions, which is quite important for building practical retrieval systems. To solve this issue, this paper proposes a rank-order preserving hashing (RoPH) method with a novel regression-based rank-order preserving loss that has provable large margin property and is easy to optimize. Moreover, we jointly learn the binary codes and hash functions instead of using any relaxation trick. To solve the induced optimization problem, the alternating descent technique is adopted and each subproblem can be solved conveniently. Specifically, we show that the involved binary quadratic programming subproblem with respect to an introduced auxiliary binary variable satisfies submodularity, enabling us to use the off-the-shelf graph-cut algorithms to solve it exactly and efficiently. Extensive experiments on three benchmarks demonstrate that RoPH significantly improves the ranking quality over the state of the arts. Kun Ding 0001, Bin Fan 0001, Chunlei Huo, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 4 |
| 2016 | Data-guided random walks for fine-structured object segmentationabstractRandom walks (RW) is a popular technique for object segmentation. Apart from the satisfactory performance in various applications, its most appealing advantage is the computational efficiency. However, RW often fails to produce complete and connected results in fine-structured (FS) object segmentation. To utilize the high efficiency and overcome the drawbacks in tackling FS objects, we develop a novel approach within the RW framework. Specifically, we propose to introduce labeling preference learned from the image data into the RW model to guide the propagation of random walkers. With the help of the guidance, random walkers are more likely to propagate correctly to the FS regions, thus yielding more accurate results. Similar to RW, this approach also bears properties such as computational efficiency, closed-form solution and unique global optimum. Moreover, it has the capacities of handling disconnected objects and transferring segmentation. Comparative experimental results demonstrate that the proposed approach achieves the state-of-the-art performance in FS object segmentation, with a low requirement of runtime. Yongchao Gong, Shiming Xiang, Chunhong Pan |
ICASSP | 2 |
| 2016 | Fine-structured object segmentation via edge-guided graph cut with interaction simplificationabstractFine-structured object segmentation is a challenging problem in object segmentation community. There are mainly two difficulties that can seriously degrade the segmentation quality: 1) insufficient interactions on fine structures due to the high demand of time and manual efforts, and 2) shrinking bias that discourages long object boundaries. To address these two issues, we develop a novel method within the graph cut framework. First, the commonly used operation of scribbling or dragging bounding boxes is replaced by loosely drawing a few rectangles, thus the interaction burden is largely reduced. Second, an edge-guided graph cut model is proposed to mitigate shrinking bias. This model enforces connectivity of fine structures by adjusting the weighting between neighboring pixels. Finally, the segmentation task is formulated as an optimization problem, which can be optimized effectively and efficiently. Comparative experimental results demonstrate the effectiveness of our method. Yongchao Gong, Shiming Xiang, Lingfeng Wang 0002, Chunhong Pan |
ICASSP | 2 |
| 2016 | Efficient sea-land segmentation using seeds learning and edge directed graph cut
Dongcai Cheng, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Neurocomputing | 3 |
| 2016 | Accurate urban road centerline extraction from VHR imagery via multiscale segmentation and tensor voting
Feiyun Zhu, Shiming Xiang, Ying Wang 0008, Chunhong Pan |
Neurocomputing | 3 |
| 2016 | Road Centerline Extraction via Semisupervised Segmentation and Multidirection Nonmaximum SuppressionabstractAccurate road centerline extraction from remotely sensed images plays a significant role in road map generation and updating. In the road extraction problem, the acquisition of labeled data is time consuming and costly; thus, there are only a small amount of labeled samples in reality. In the existing centerline extraction algorithms, the thinning-based algorithms always produce small spurs that reduce the smoothness and accuracy of the road centerline; the regression-based algorithms can extract a smooth road network, but they are time consuming. To solve the aforementioned problems, we propose a novel road centerline extraction method, which is constructed based on semisupervised segmentation and multiscale filtering (MF) and multidirection nonmaximum suppression (M-NMS) (MF&M-NMS). Specifically, a semisupervised method, which explores the intrinsic structures between the labeled samples and the unlabeled ones, is introduced to obtain the segmentation result. Then, a novel MF&M-NMS-based algorithm is proposed to gain a smooth and complete road centerline network. Experimental results on a public data set demonstrate that the proposed method achieves comparable or better performances by comparing with the state-of-the-art methods. In addition, our method is nearly ten times faster than the state-of-the-art methods. Feiyun Zhu, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Fine-structured object segmentation via neighborhood propagation
Yongchao Gong, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 2 |
| 2016 | Efficient Multiple Feature Fusion With Hashing for Hyperspectral Imagery Classification: A Comparative StudyabstractDue to the complementary properties of different features, multiple feature fusion has a large potential for hyperspectral imagery classification. At the meantime, hashing is promising in representing a high-dimensional float-type feature with extremely low bit binary codes while maintaining the performance. In this paper, we study the possibility of using hashing to fuse multiple features for hyperspectral imagery classification. For this purpose, we propose a multiple feature fusion framework to evaluate the performance of using different hashing methods. For comparison and completeness, we also have an extensive comparison to five subspace-based dimension reduction methods and six fusion-based methods which are popular solutions to deal with multiple features in hyperspectral image classification. Experimental results on four benchmark hyperspectral data sets demonstrate that using hashing to fuse multiple features can achieve comparable or better performance with the traditional subspace-based dimension reduction methods and fusion-based methods. Moreover, the binary features obtained by using hashing need much less storage and are faster to compute distances with the help of machine instructions. Zisha Zhong, Bin Fan 0001, Kun Ding 0001, Haichang Li, Shiming Xiang, Chunhong Pan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2016 | Cross-Modal Retrieval via Deep and Bidirectional Representation LearningabstractCross-modal retrieval emphasizes understanding inter-modality semantic correlations, which is often achieved by designing a similarity function. Generally, one of the most important things considered by the similarity function is how to make the cross-modal similarity computable. In this paper, a deep and bidirectional representation learning model is proposed to address the issue of image-text cross-modal retrieval. Owing to the solid progress of deep learning in computer vision and natural language processing, it is reliable to extract semantic representations from both raw image and text data by using deep neural networks. Therefore, in the proposed model, two convolution-based networks are adopted to accomplish representation learning for images and texts. By passing the networks, images and texts are mapped to a common space, in which the cross-modal similarity is measured by cosine distance. Subsequently, a bidirectional network architecture is designed to capture the property of the cross-modal retrieval-the bidirectional search. Such architecture is characterized by simultaneously involving the matched and unmatched image-text pairs for training. Accordingly, a learning framework with maximum likelihood criterion is finally developed. The network parameters are optimized via backpropagation and stochastic gradient descent. A great deal of experiments are conducted to sufficiently evaluate the proposed method on three publicly released datasets: IAPRTC-12, Flickr30k, and Flickr8k. The overall results definitely show that the proposed architecture is effective and the learned representations have good semantics to achieve superior cross-modal retrieval performance. Yonghao He, Shiming Xiang, Cuicui Kang, Jian Wang 0068, Chunhong Pan |
IEEE Trans. Multim. | 2 |
| 2015 | 10, 000+ Times Accelerated Robust Subset SelectionabstractSubset selection from massive data with noised information is increasingly popular for various applications. This problem is still highly challenging as current methods are generally slow in speed and sensitive to outliers. To address the above two issues, we propose an accelerated robust subset selection (ARSS) method. Extensive experiments on ten benchmark datasets verify that our method not only outperforms state of the art methods, but also runs 10,000+ times faster than the most related method. Feiyun Zhu, Bin Fan 0001, Xinliang Zhu, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
AAAI | 5 |
| 2015 | Cross-Modal Similarity Learning: A Low Rank Bilinear FormulationabstractThe cross-media retrieval problem has received much attention in recent years due to the rapid increasing of multimedia data on the Internet. A new approach to the problem has been raised which intends to match features of different modalities directly. In this research, there are two critical issues: how to get rid of the heterogeneity between different modalities and how to match the cross-modal features of different dimensions. Recently metric learning methods show a good capability in learning a distance metric to explore the relationship between data points. However, the traditional metric learning algorithms only focus on single-modal features, which suffer difficulties in addressing the cross-modal features of different dimensions. In this paper, we propose a cross-modal similarity learning algorithm for the cross-modal feature matching. The proposed method takes a bilinear formulation, and with the nuclear-norm penalization, it achieves low-rank representation. Accordingly, the accelerated proximal gradient algorithm is successfully imported to find the optimal solution with a fast convergence rate O(1/t2). Experiments on three well known image-text cross-media retrieval databases show that the proposed method achieves the best performance compared to the state-of-the-art algorithms. Cuicui Kang, Shengcai Liao, Yonghao He, Jian Wang 0068, Wenjia Niu, Shiming Xiang, Chunhong Pan |
CIKM | 6 |
| 2015 | Extraction of Virtual Baselines from Distorted Document Images Using Curvilinear ProjectionabstractThe baselines of a document page are a set of virtual horizontal and parallel lines, to which the printed contents of document, e.g., text lines, tables or inserted photos, are aligned. Accurate baseline extraction is of great importance in the geometric correction of curved document images. In this paper, we propose an efficient method for accurate extraction of these virtual visual cues from a curved document image. Our method comes from two basic observations that the baselines of documents do not intersect with each other and that within a narrow strip, the baselines can be well approximated by linear segments. Based upon these observations, we propose a curvilinear projection based method and model the estimation of curved baselines as a constrained sequential optimization problem. A dynamic programming algorithm is then developed to efficiently solve the problem. The proposed method can extract the complete baselines through each pixel of document images in a high accuracy. It is also scripts insensitive and highly robust to image noises, non-textual objects, image resolutions and image quality degradation like blurring and non-uniform illumination. Extensive experiments on a number of captured document images demonstrate the effectiveness of the proposed method. Gaofeng Meng, Zuming Huang, Yonghong Song, Shiming Xiang, Chunhong Pan |
ICCV | 4 |
| 2015 | Segment Graph Based Image Filtering: Fast Structure-Preserving SmoothingabstractIn this paper, we design a new edge-aware structure, named segment graph, to represent the image and we further develop a novel double weighted average image filter (SGF) based on the segment graph. In our SGF, we use the tree distance on the segment graph to define the internal weight function of the filtering kernel, which enables the filter to smooth out high-contrast details and textures while preserving major image structures very well. While for the external weight function, we introduce a user specified smoothing window to balance the smoothing effects from each node of the segment graph. Moreover, we also set a threshold to adjust the edge-preserving performance. These advantages make the SGF more flexible in various applications and overcome the "halo" and "leak" problems appearing in most of the state-of-the-art approaches. Finally and importantly, we develop a linear algorithm for the implementation of our SGF, which has an O(N) time complexity for both gray-scale and high dimensional images, regardless of the kernel size and the intensity range. Typically, as one of the fastest edge-preserving filters, our CPU implementation achieves 0.15s per megapixel when performing filtering for 3-channel color images. The strength of the proposed filter is demonstrated by various applications, including stereo matching, optical flow, joint depth map upsampling, edge-preserving smoothing, edges detection, image abstraction and texture editing. Feihu Zhang, Longquan Dai, Shiming Xiang, Xiaopeng Zhang 0001 |
ICCV | 3 |
| 2015 | Super-resolution reconstruction using graph Laplacian penalizationabstractThis paper proposes to employ graph Laplacian penalization in multi-image super-resolution reconstruction. Most state-of-the-art methods use an optimization model with two items: the data fidelity item and the penalization item. However, these methods often ignore the impact of the penalization item and utilize simple formulations such as high-pass filters to fulfill the super-resolution task. As a result, they can not restore much local structural information of the high resolution image. By using graph Laplacian, the proposed method in this paper can retain more local manifold structures in the high resolution images. Based on this idea, the optimization model is constructed and the solution is presented. Comparative experiments have validated our method. The experiments have also tested our method has much faster convergence speed. Limin Shi, Bangyu Li, Shiming Xiang, Chunhong Pan |
ICIP | 4 |
| 2015 | Fine-structured object segmentation via local and nonlocal neighborhood propagationabstractIn this paper, we present a novel method for the challenging task of fine-structured (FS) object segmentation. This task is formulated as a label propagation problem on an affinity graph. To enhance the completeness and connectivity of the FS objects, we introduce a novel neighborhood system combining both local and nonlocal connections, together with a robust scheme for edge weight calculation. Additionally, region cost is incorporated into the energy function to further maintain the connectivity of fine parts where the propagation is hard to reach. An appealing advantage of the proposed method is that the energy minimization has a closed-form solution and global optimum is guaranteed. Comparative experimental results on three datasets demonstrate the effectiveness of the proposed method. Yongchao Gong, Shiming Xiang, Lingfeng Wang 0002, Chunhong Pan |
ICIP | 2 |
| 2015 | Large Scale Image Annotation via Deep Representation Learning and Tag Embedding LearningabstractIn this paper, we focus on the issue of large scale image annotation, whereas most existing methods are devised for small datasets. A novel model based on deep representation learning and tag embedding learning is proposed. Specifically, the proposed model learns an unified latent space for image visual features and tag embeddings simultaneously. Furthermore, a metric matrix is introduced to estimate the relevance scores between images and tags. Finally, an objective function modeling triplet relationships (irrelevant tag, image, relevant tag) is proposed with maximum margin pursuit. The proposed model is easy to tackle new images and tags via online learning and has a relatively low test computation complexity. Experimental results on NUS-WIDE dataset demonstrate the effectiveness of the proposed model. Yonghao He, Jian Wang 0068, Cuicui Kang, Shiming Xiang, Chunhong Pan |
ICMR | 4 |
| 2015 | Image-Text Cross-Modal Retrieval via Modality-Specific Feature LearningabstractCross-modal retrieval extends the ability of search engines to deal with the massive cross-modal data. The goal of image-text cross-modal retrieval is to search images (texts) by using text (image) queries by computing the similarities of images and texts directly. Many existing methods rely on low-level visual features and textual features for cross-modal retrieval, ignoring the characteristics existing in the raw data of different modalities. In this paper, a novel model based on modality-specific feature learning is proposed. Considering the characteristics of different modalities, the model uses two types of convolutional neural networks to map the raw data to the latent space representations for images and texts, respectively. Particularly, the convolution based network used for texts involves word embedding learning, which has been proved effective to extract meaningful textual features for text classification. In the latent space, the mapped features of images and texts form relevant and irrelevant image-text pairs, which are used by the one-vs-more learning scheme. This learning scheme can achieve ranking functionality by allowing for one relevant and more irrelevant pairs. The standard back-propagation technique is employed to update the parameters of two convolutional networks. Extensive cross-modal retrieval experiments are carried out on three challenging datasets that consist of image-document pairs or image-query click-through data from a search engine, and the results firmly demonstrate that the proposed model is much more effective. Jian Wang 0068, Yonghao He, Cuicui Kang, Shiming Xiang, Chunhong Pan |
ICMR | 4 |
| 2015 | Image Deblurring with Coupled Dictionary Learning
Shiming Xiang, Gaofeng Meng, Ying Wang 0008, Chunhong Pan, Changshui Zhang |
Int. J. Comput. Vis. | 1 |
| 2015 | Image tag-ranking via pairwise supervision based semi-supervised model
Yonghao He, Cuicui Kang, Jian Wang 0068, Shiming Xiang, Chunhong Pan |
Neurocomputing | 4 |
| 2015 | Multicluster Spatial-Spectral Unsupervised Feature Selection for Hyperspectral Image ClassificationabstractA new unsupervised spatial-spectral feature selection method for hyperspectral images has been proposed in this letter. The key idea is to select the features that better preserve the multicluster structure of the multiple spatial-spectral features. Specifically, the multicluster structure information is obtained through spectral clustering utilizing a weighted combination of the multiple features. Then, such information is preserved in a group-sparsity-based robust linear regression model. The features that contribute more in preserving the multicluster structure information are selected. Comparative experiments on two popular real hyperspectral images validate the effectiveness of the proposed method, showing higher classification accuracy. Haichang Li, Shiming Xiang, Zisha Zhong, Kun Ding 0001, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Discriminant Tensor Spectral-Spatial Feature Extraction for Hyperspectral Image ClassificationabstractWe propose to integrate spectral-spatial feature extraction and tensor discriminant analysis for hyperspectral image classification. First, we apply remarkable spectral-spatial feature extraction approaches in the hyperspectral cube to extract a feature tensor for each pixel. Then, based on class label information, local tensor discriminant analysis is used to remove redundant information for subsequent classification procedure. The approach not only extracts sufficient spectral-spatial features from original hyperspectral images but also gets better feature representation owing to tensor framework. Comparative results on two benchmarks demonstrate the effectiveness of our method. Zisha Zhong, Bin Fan 0001, Jiangyong Duan, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2015 | Robust Hyperspectral Unmixing With Correntropy-Based MetricabstractHyperspectral unmixing is one of the crucial steps for many hyperspectral applications. The problem of hyperspectral unmixing has proved to be a difficult task in unsupervised work settings where the endmembers and abundances are both unknown. In addition, this task becomes more challenging in the case that the spectral bands are degraded by noise. This paper presents a robust model for unsupervised hyperspectral unmixing. Specifically, our model is developed with the correntropy-based metric where the nonnegative constraints on both endmembers and abundances are imposed to keep physical significance. Besides, a sparsity prior is explicitly formulated to constrain the distribution of the abundances of each endmember. To solve our model, a half-quadratic optimization technique is developed to convert the original complex optimization problem into an iteratively reweighted nonnegative matrix factorization with sparsity constraints. As a result, the optimization of our model can adaptively assign small weights to noisy bands and put more emphasis on noise-free bands. In addition, with sparsity constraints, our model can naturally generate sparse abundances. Experiments on synthetic and real data demonstrate the effectiveness of our model in comparison to the related state-of-the-art unmixing models. Ying Wang 0008, Chunhong Pan, Shiming Xiang, Feiyun Zhu |
IEEE Trans. Image Process. | 3 |
| 2015 | Learning Consistent Feature Representation for Cross-Modal Multimedia RetrievalabstractThe cross-modal feature matching has gained much attention in recent years, which has many practical applications, such as the text-to-image retrieval. The most difficult problem of cross-modal matching is how to eliminate the heterogeneity between modalities. The existing methods (e.g., CCA and PLS) try to learn a common latent subspace, where the heterogeneity between two modalities is minimized so that cross-matching is possible. However, most of these methods require fully paired samples and suffer difficulties when dealing with unpaired data. Besides, utilizing the class label information has been found as a good way to reduce the semantic gap between the low-level image features and high-level document descriptions. Considering this, we propose a novel and effective supervised algorithm, which can also deal with the unpaired data. In the proposed formulation, the basis matrices of different modalities are jointly learned based on the training samples. Moreover, a local group-based priori is proposed in the formulation to make a better use of popular block based features (e.g., HOG and GIST). Extensive experiments are conducted on four public databases: Pascal VOC2007, LabelMe, Wikipedia, and NUS-WIDE. We also evaluated the proposed algorithm with unpaired data. By comparing with existing state-of-the-art algorithms, the results show that the proposed algorithm is more robust and achieves the best performance, which outperforms the second best algorithm by about 5% on both the Pascal VOC2007 and NUS-WIDE databases. Cuicui Kang, Shiming Xiang, Shengcai Liao, Changsheng Xu, Chunhong Pan |
IEEE Trans. Multim. | 2 |
| 2015 | Retargeted Least Squares Regression AlgorithmabstractThis brief presents a framework of retargeted least squares regression (ReLSR) for multicategory classification. The core idea is to directly learn the regression targets from data other than using the traditional zero-one matrix as regression targets. The learned target matrix can guarantee a large margin constraint for the requirement of correct classification for each data point. Compared with the traditional least squares regression (LSR) and a recently proposed discriminative LSR models, ReLSR is much more accurate in measuring the classification error of the regression model. Furthermore, ReLSR is a single and compact model, hence there is no need to train two-class (binary) machines that are independent of each other. The convex optimization problem of ReLSR is solved elegantly and efficiently with an alternating procedure including regression and retargeting as substeps. The experimental evaluation over a range of databases identifies the validity of our method. Xu-Yao Zhang, Lingfeng Wang 0002, Shiming Xiang, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Active Flattening of Curved Document Images via Two Structured BeamsabstractDocument images captured by a digital camera often suffer from serious geometric distortions. In this paper, we propose an active method to correct geometric distortions in a camera-captured document image. Unlike many passive rectification methods that rely on text-lines or features extracted from images, our method uses two structured beams illuminating upon the document page to recover two spatial curves. A developable surface is then interpolated to the curves by finding the correspondence between them. The developable surface is finally flattened onto a plane by solving a system of ordinary differential equations. Our method is a content independent approach and can restore a corrected document image of high accuracy with undistorted contents. Experimental results on a variety of real-captured document images demonstrate the effectiveness and efficiency of the proposed method. Gaofeng Meng, Ying Wang 0008, Shenquan Qu, Shiming Xiang, Chunhong Pan |
CVPR | 4 |
| 2014 | Image annotation via learning the image-label interrelationsabstractThe goal of image annotation is to automatically assign meaningful and content-related labels to the digital images by using machines. It is beneficial to image search and image sharing in social networks. Various methods for image annotation are proposed in last decade and they have gained much progress. However, most of them are not precise and fast enough for real-world applications. In this paper, we propose a novel and fast image annotation method via learning the image-label interrelation. The main idea of the proposed method is to predict labels by linearly propagating the label information through the image-label interrelation and the image similarities. Thus, we propose a model based on the regression between the label groundtruth and the propagated label information to learn the image-label interrelation. In addition, a label-biased regularization is integrated into our model to learn more effective and meaningful image-label interrelation. Finally, our model can be solved in closed form, therefore it achieves a fast learning process. Experimental results on three benchmark datasets demonstrate that our method shows the comparable performance with state-of-the-art methods and has faster learning time. Yonghao He, Jian Wang 0068, Shiming Xiang, Chunhong Pan |
ICIP | 3 |
| 2014 | Facade repetition extraction using block matrix based modelabstractRepetition extraction plays an important role in facade image analysis. In this paper, this task is handled within the graph cut based image segmentation framework. To model the repetitions, generalized translation symmetry (GTS) is introduced to enable aperiodic repetition layouts. More importantly, GTS is explicitly formulated in terms of matrix multiplication. That is, GTS is viewed as the product of a repetitive pattern and two block matrices. These two block matrices are employed to represent the vertical and horizontal symmetry respectively. On this basis, repetition extraction is formulated as a GTS constrained energy minimization problem. An alternatively optimization algorithm based on graph cut and dynamic programming is finally developed to solve the problem. Experimental results demonstrate the validity of our method. Hongfei Xiao, Gaofeng Meng, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
ICIP | 4 |
| 2014 | Facade Labeling via Explicit Matrix Factorization
Hongfei Xiao, Lingfeng Wang 0002, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICISP | 4 |
| 2014 | Local Label Probability Propagation for Hyperspectral Image ClassificationabstractClassification of hyper spectral images is an important issue in remote sensing image processing systems. Hyper spectral images have advantages in pixel-wise classification owing to the high spectral resolution. However, the pixel-wise classification result often introduces the salt-and-pepper appearance because of the complex noise produced by atmosphere and instrument. An effective way to overcome this phenomenon is to resort to the spatial information. This paper proposes a method to solve the above problem by using spatial similarity information. First, in order to avoid the effect of noisy pixels and mixed pixels, reliable seeds are selected in local windows according to the agreement between the central pixel and its spatial neighbors. Then, the information of the reliable seeds is propagated to their spatial neighbors by a graph Laplacian. Specifically, the graph Laplacian is designed to propagate information among spatial neighbors with close similarity relationship so that some small or long thin objects are identified. Through the seed selection and local reliable information propagation, the problem of noisy labels is solved elegantly. Experiments on three real hyper spectral data sets with different spatial resolution, spectral resolution and land covers demonstrate the effectiveness of our method. Haichang Li, Jiangyong Duan, Shiming Xiang, Lingfeng Wang 0002, Chunhong Pan |
ICPR | 3 |
| 2014 | Cross Modal Deep Model and Gaussian Process Based Model for MSR-Bing ChallengeabstractIn the MSR-Bing Image Retrieval Challenge, the contestants are required to design a system that can score the query-image pairs based on the relevance between queries and images. To address this problem, we propose a regression based cross modal deep learning model and a Gaussian Process scoring model. The regression based cross modal deep learning model takes the image features and query features as inputs respectively and outputs the relevance scores directly. The Gaussian Process scoring model regards the challenge as a ranking problem and utilizes the click (or pseudo click) information from both the training set and the development set to predict the relevance scores. The proposed models are used in different situations: matched and miss-matched queries. Experiments on the development set show the effectiveness of the proposed models. Jian Wang 0068, Cuicui Kang, Yonghao He, Shiming Xiang, Chunhong Pan |
ACM Multimedia | 4 |
| 2014 | Multifocus image fusion via focus segmentation and region reconstruction
Jiangyong Duan, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
Neurocomputing | 3 |
| 2014 | Kernel sparse representation with pixel-level and region-level local feature kernels for face recognition
Cuicui Kang, Shengcai Liao, Shiming Xiang, Chunhong Pan |
Neurocomputing | 3 |
| 2014 | Vehicle Detection in Satellite Images by Hybrid Deep Convolutional Neural NetworksabstractDetecting small objects such as vehicles in satellite images is a difficult problem. Many features (such as histogram of oriented gradient, local binary pattern, scale-invariant feature transform, etc.) have been used to improve the performance of object detection, but mostly in simple environments such as those on roads. Kembhavi et al. proposed that no satisfactory accuracy has been achieved in complex environments such as the City of San Francisco. Deep convolutional neural networks (DNNs) can learn rich features from the training data automatically and has achieved state-of-the-art performance in many image classification databases. Though the DNN has shown robustness to distortion, it only extracts features of the same scale, and hence is insufficient to tolerate large-scale variance of object. In this letter, we present a hybrid DNN (HDNN), by dividing the maps of the last convolutional layer and the max-pooling layer of DNN into multiple blocks of variable receptive field sizes or max-pooling field sizes, to enable the HDNN to extract variable-scale features. Comparative experimental results indicate that our proposed HDNN significantly outperforms the traditional DNN on vehicle detection. Xueyun Chen, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | Spectral Unmixing via Data-Guided SparsityabstractHyperspectral unmixing, the process of estimating a common set of spectral bases and their corresponding composite percentages at each pixel, is an important task for hyperspectral analysis, visualization, and understanding. From an unsupervised learning perspective, this problem is very challenging-both the spectral bases and their composite percentages are unknown, making the solution space too large. To reduce the solution space, priors. In practice, these priors would easily lead to some unsuitable solution. This is because they are achieved by applying an identical strength of constraints to all the factors, which does not hold in practice. To overcome this limitation, we propose a novel sparsity-based method by learning a data-guided map (DgMap) to describe the individual mixed level of each pixel. Through this DgMap, the l(p) (0 < p < 1) constraint is applied in an adaptive manner. Such implementation not only meets the practical situation, but also guides the spectral bases toward the pixels under highly sparse constraint. What is more, an elegant optimization scheme as well as its convergence proof have been provided in this paper. Extensive experiments on several datasets also demonstrate that the DgMap is feasible, and high quality unmixing results could be obtained by our method. Feiyun Zhu, Ying Wang 0008, Bin Fan 0001, Shiming Xiang, Gaofeng Meng, Chunhong Pan |
IEEE Trans. Image Process. | 4 |
| 2013 | Indoor frame recovering via line segments refinement and votingabstractFrame structure estimation from line segments is an important yet challenging problem in understanding indoor scenes. In practice, line segment extraction can be affected by occlusions, illumination variations, and weak object boundaries. To address this problem, an approach for frame structure recovery based on line segment refinement and voting is proposed. We refined line segments by the revising, connecting, and adding operations. We then propose an iterative voting mechanism for selecting refined line segments, where a cross ratio constraint is enforced to build crab-like models. Our algorithm outperforms state-of-the-art approaches, especially when considering complex indoor scenes. Luanzheng Guo, Lingfeng Wang 0002, Chunhong Pan, Shiming Xiang |
ICASSP | 5 |
| 2013 | Removing out-of-focus blur from similar image pairsabstractThis paper presents a new deblurring method to remove the out-of-focus blur from similar image pairs. The method is motivated by an observation that a blurred structure appearing in one image can often have its corresponding clear one in the similar clear images. Our method first extracts the patch pairs from input images by SIFT matching. Then the constraints on the patch pairs are used to estimate the blur kernel via the RANSAC algorithm. Finally, the non-blind deconvolution is adopted to restore the blurred image. The main advantage is that we can improve the deblurring results with the help of additional similar clear images in many practical applications. Our method is validated on synthetic and real images by comparing with state-of-the-art methods. Jiangyong Duan, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICASSP | 3 |
| 2013 | Robust VHR image change detection based on local features and multi-scale fusionabstractUrban change detection of Very High Resolution (VHR) remote sensing images is challenging, due to the ill-posed nature of change detection problem, the inherent nature of VHR image, the complex morphology of urban scenes, etc. To address the above difficulties, a robust approach is proposed, which is based on discriminative local features, robust distance metric and novel multi-scale fusion strategy. By integrating these components synergistically, the proposed approach is superior to the traditional approaches in capturing semantic changes and removing the false changes. Comparative experiments demonstrate the effectiveness and advantages of the proposed approach. Chunlei Huo, Shiming Xiang, Chunhong Pan |
ICASSP | 3 |
| 2013 | Active learning based automatic face segmentation for kinect videoabstractThis paper presents a novel segmentation approach for extracting faces from videos. Under an active learning framework, the segmentation is conducted automatically without human interactions. A small portion of pixels are first labeled as face or non-face. Given these labeled samples, a semi-supervised spline regression model is then applied to obtain the face region. Based on the segmentation result, new pixels are selected and labeled. These two steps perform iterately until convergence. The main novelty is that color and depth data are combined to provide the labeling information. Our approach is validated via comparisons with state-of-the-art methods on real videos captured from the commodity Kinect camera. Jixia Zhang, Shaoguo Liu, Franck Davoine, Chunhong Pan, Shiming Xiang |
ICASSP | 6 |
| 2013 | Efficient Image Dehazing with Boundary Constraint and Contextual RegularizationabstractImages captured in foggy weather conditions often suffer from bad visibility. In this paper, we propose an efficient regularization method to remove hazes from a single input image. Our method benefits much from an exploration on the inherent boundary constraint on the transmission function. This constraint, combined with a weighted L_1-norm based contextual regularization, is modeled into an optimization problem to estimate the unknown scene transmission. A quite efficient algorithm based on variable splitting is also presented to solve the problem. The proposed method requires only a few general assumptions and can restore a high-quality haze-free image with faithful colors and fine image details. Experimental results on a variety of haze images demonstrate the effectiveness and efficiency of the proposed method. Gaofeng Meng, Ying Wang 0008, Jiangyong Duan, Shiming Xiang, Chunhong Pan |
ICCV | 4 |
| 2013 | Group sparsity based semi-supervised band selection for hyperspectral imagesabstractIn this paper, we propose a novel group sparsity based semi-supervised band selection method. There are three key features in our method. First, it fulfills the band selection task by employing group sparsity on the regression coefficients in a robust linear regression for classification model, so that the selected bands hold lower classification errors. Second, the spatial smoothness prior is incorporated to preserve the similarity of spatial neighbors in band selection. Third, the objective function is efficiently optimized via an alternative iteration algorithm. Comparative results on two hyper-spectral data sets validate the effectiveness of our method, showing higher classification accuracies. Haichang Li, Ying Wang 0008, Jiangyong Duan, Shiming Xiang, Chunhong Pan |
ICIP | 4 |
| 2013 | Nonparametric Illumination Correction for Scanned Document Images via Convex HullsabstractA scanned image of an opened book page often suffers from various scanning artifacts known as scanning shading and dark borders noises. These artifacts will degrade the qualities of the scanned images and cause many problems to the subsequent process of document image analysis. In this paper, we propose an effective method to rectify these scanning artifacts. Our method comes from two observations: that the shading surface of most scanned book pages is quasi-concave and that the document contents are usually printed on a sheet of plain and bright paper. Based on these observations, a shading image can be accurately extracted via convex hulls-based image reconstruction. The proposed method proves to be surprisingly effective for image shading correction and dark borders removal. It can restore a desired shading-free image and meanwhile yield an illumination surface of high quality. More importantly, the proposed method is nonparametric and thus does not involve any user interactions or parameter fine-tuning. This would make it very appealing to nonexpert users in applications. Extensive experiments based on synthetic and real-scanned document images demonstrate the efficiency of the proposed method. Gaofeng Meng, Shiming Xiang, Nanning Zheng 0001, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Level set evolution with locally linear classification for image segmentation
Ying Wang 0008, Shiming Xiang, Chunhong Pan, Lingfeng Wang 0002, Gaofeng Meng |
Pattern Recognit. | 2 |
| 2013 | Edge-Directed Single-Image Super-Resolution Via Adaptive Gradient Magnitude Self-InterpolationabstractSuper-resolution from a single image plays an important role in many computer vision systems. However, it is still a challenging task, especially in preserving local edge structures. To construct high-resolution images while preserving the sharp edges, an effective edge-directed super-resolution method is presented in this paper. An adaptive self-interpolation algorithm is first proposed to estimate a sharp high-resolution gradient field directly from the input low-resolution image. The obtained high-resolution gradient is then regarded as a gradient constraint or an edge-preserving constraint to reconstruct the high-resolution image. Extensive results have shown both qualitatively and quantitatively that the proposed method can produce convincing super-resolution images containing complex and sharp features, as compared with the other state-of-the-art super-resolution algorithms. Lingfeng Wang 0002, Shiming Xiang, Gaofeng Meng, Huai-Yu Wu, Chunhong Pan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | A Graph-Based Classification Method for Hyperspectral ImagesabstractThe goal of this paper is to apply graph cut (GC) theory to the classification of hyperspectral remote sensing images. The task is formulated as a labeling problem on Markov random field (MRF) constructed on the image grid, and GC algorithm is employed to solve this task. In general, a large number of user interactive strikes are necessary to obtain satisfactory segmentation results. Due to the spatial variability of spectral signatures, however, hyperspectral remote sensing images often contain many tiny regions. Labeling all these tiny regions usually needs expensive human labor. To overcome this difficulty, a pixelwise fuzzy classification based on support vector machine (SVM) is first applied. As a result, only pixels with high probabilities are preserved as labeled ones. This generates a pseudouser strike map. This map is then employed for GC to evaluate the truthful likelihoods of class labels and propagate them to the MRF. To evaluate the robustness of our method, we have tested our method on both large and small training sets. Additionally, comparisons are made between the results of SVM, SVM with stacking neighboring vectors, SVM with morphological preprocessing, extraction and classification of homogeneous objects, and our method. Comparative experimental results demonstrate the validity of our method. Shiming Xiang, Chunhong Pan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2012 | Image Guided Tone Mapping with Locally Nonlinear Model
Huxiang Gu, Ying Wang 0008, Shiming Xiang, Gaofeng Meng, Chunhong Pan |
ECCV (4) | 3 |
| 2012 | Classification oriented semi-supervised band selection for hyperspectral images
Shiming Xiang, Chunhong Pan |
ICPR | 2 |
| 2012 | Kernel Homotopy based sparse representation for object classification
Cuicui Kang, Shengcai Liao, Shiming Xiang, Chunhong Pan |
ICPR | 3 |
| 2012 | Metric Rectification of Curved Document ImagesabstractIn this paper, we propose a metric rectification method to restore an image from a single camera-captured document image. The core idea is to construct an isometric image mesh by exploiting the geometry of page surface and camera. Our method uses a general cylindrical surface (GCS) to model the curved page shape. Under a few proper assumptions, the printed horizontal text lines are shown to be line convergent symmetric. This property is then used to constrain the estimation of various model parameters under perspective projection. We also introduce a paraperspective projection to approximate the nonlinear perspective projection. A set of close-form formulas is thus derived for the estimate of GCS directrix and document aspect ratio. Our method provides a straightforward framework for image metric rectification. It is insensitive to camera positions, viewing angles, and the shapes of document pages. To evaluate the proposed method, we implemented comprehensive experiments on both synthetic and real-captured images. The results demonstrate the efficiency of our method. We also carried out a comparative experiment on the public CBDAR2007 data set. The experimental results show that our method outperforms the state-of-the-art methods in terms of OCR accuracy and rectification errors. Gaofeng Meng, Chunhong Pan, Shiming Xiang, Jiangyong Duan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Image deblurring with matrix regression and gradient evolution
Shiming Xiang, Gaofeng Meng, Ying Wang 0008, Chunhong Pan, Changshui Zhang |
Pattern Recognit. | 1 |
| 2012 | Orthogonal vs. uncorrelated least squares discriminant analysis for feature extraction
Feiping Nie 0001, Shiming Xiang, Yun Liu 0021, Chenping Hou, Changshui Zhang |
Pattern Recognit. Lett. | 2 |
| 2012 | Discriminative Least Squares Regression for Multiclass Classification and Feature SelectionabstractThis paper presents a framework of discriminative least squares regression (LSR) for multiclass classification and feature selection. The core idea is to enlarge the distance between different classes under the conceptual framework of LSR. First, a technique called ε-dragging is introduced to force the regression targets of different classes moving along opposite directions such that the distances between classes can be enlarged. Then, the ε-draggings are integrated into the LSR model for multiclass classification. Our learning framework, referred to as discriminative LSR, has a compact model form, where there is no need to train two-class machines that are independent of each other. With its compact form, this model can be naturally extended for feature selection. This goal is achieved in terms of L2,1 norm of matrix, generating a sparse learning model for feature selection. The model for multiclass classification and its extension for feature selection are finally solved elegantly and efficiently. Experimental evaluation over a range of benchmark datasets indicates the validity of our method. Shiming Xiang, Feiping Nie 0001, Gaofeng Meng, Chunhong Pan, Changshui Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2011 | Image editing based on Sparse Matrix-Vector multiplicationabstractThis paper presents a unified model for image editing in terms of Sparse Matrix-Vector (SpMV) multiplication. In our framework, we cast image editing as a linear energy minimization problem and address it by solving a sparse linear system, which is able to yield a globally optimal solution. First, three classical image editing operations, including linear filtering, resizing and selecting, are reformulated in the SpMV multiplication form. The SpMV form helps us set up a straightforward mechanism to flexibly and naturally combine various image features (low-level visual features or geometrical features) and constraints together into an integrated energy minimization function under the L2norm. Then, we apply our model to implement the tasks of pan-sharpening, image cloning, image mixed editing and texture transfer, which are now popularly used in the field of digital art. Comparative experiments are reported to validate the effectiveness and efficiency of our model. Ying Wang 0008, Hongping Yan, Chunhong Pan, Shiming Xiang |
ICASSP | 4 |
| 2011 | Kernel sparse representation with local patterns for face recognitionabstractIn this paper we propose a novel kernel sparse representation classification (SRC) framework and utilize the local binary pattern (LBP) descriptor in this framework for robust face recognition. First we develop a kernel coordinate descent (KCD) algorithm for 11 minimization in the kernel space, which is based on the covariance update technique. Then we extract LBP descriptors from each image and apply two types of kernels (χ2distance based and Hamming distance based) with the proposed KCD algorithm under the SRC framework for face recognition. Experiments on both the Extended Yale B and the PIE face databases show that the proposed method is more robust against noise, occlusion, and illumination variations, even with small number of training samples. Cuicui Kang, Shengcai Liao, Shiming Xiang, Chunhong Pan |
ICIP | 3 |
| 2011 | Robust airplane detection in satellite imagesabstractAutomatic target detection in satellite images remains a challenging problem. The main difficulties lie in the cooccurrence of variations of target type, pose, and size in huge satellite image. In this paper, we propose a new airplane detection approach based on visual saliency computation and symmetry detection. The advantages are twofold. First, saliency and symmetry detection perform stably in obtaining target location and orientation information. Second, independent of target type, pose and size, saliency map and symmetry detection are computed only once. This saves a large amount of computational time but does not miss any targets. Experiments show that our method provides a promising way to detect airplanes in complex airport scenes. Shiming Xiang, Chunhong Pan |
ICIP | 2 |
| 2011 | MEAN-shift tracking algorithm with weight fusion strategyabstractIn this paper, we propose a new Mean-shift algorithm to tackle some tracking difficulties, such as background clutter and partial occlusion. First, we compare all Mean-shift-like tracking algorithms, and indicate that the main difference among them is weight calculation. Then, a new fusion strategy is proposed to unify all weight calculation methods into a framework. Based on this framework, we propose a novel weight calculation method, which takes the candidate model into consideration as well as incorporates the local background. Extensive experiments are conducted to evaluate the proposed approach. Comparative experimental results indicate that the tracking accuracy is improved as compared with the state-of-the-arts. Lingfeng Wang 0002, Chunhong Pan, Shiming Xiang |
ICIP | 3 |
| 2011 | Level set evolution with locally linear classification for image segmentationabstractThis paper presents a novel local region-based level set model for image segmentation. In each local region, we define a locally weighted least squares energy to fit a linear classification function. The local energy is then integrated over the entire image domain to form an energy functional in terms of level set function. The energy minimization is achieved by level set evolution and estimation of parameters of the locally linear function in an iterative process. By introducing the locally linear functions to separate background and foreground in local regions, our model not only ensures the accuracy of the segmentation results, but also be very robust to initialization. Experiments are reported to demonstrate the effectiveness and efficiency of our model. Ying Wang 0008, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
ICIP | 3 |
| 2011 | Interactive Image Segmentation With Multiple Linear Reconstructions in WindowsabstractThis paper proposes an algorithm for interactive image segmentation. The task is formulated as a problem of graph-based transductive classification. Specifically, given an image window, the color of each pixel in it will be reconstructed linearly with those of the remaining pixels in this window. The optimal reconstruction weights will be kept unchanged to linearly reconstruct their class labels. The label reconstruction errors are estimated in each window. These errors are further collected together to develop a learning model. Then, the class information about the user specified foreground and background pixels are integrated into a regularization framework. Under this framework, a globally optimal labeling is finally obtained. The computational complexity is analyzed, and an approach for speeding up the algorithm is presented. Comparative experimental results illustrate the validity of our algorithm. Shiming Xiang, Chunhong Pan, Feiping Nie 0001, Changshui Zhang |
IEEE Trans. Multim. | 1 |
| 2011 | Semisupervised Dimensionality Reduction and Classification Through Virtual Label RegressionabstractSemisupervised dimensionality reduction has been attracting much attention as it not only utilizes both labeled and unlabeled data simultaneously, but also works well in the situation of out-of-sample. This paper proposes an effective approach of semisupervised dimensionality reduction through label propagation and label regression. Different from previous efforts, the new approach propagates the label information from labeled to unlabeled data with a well-designed mechanism of random walks, in which outliers are effectively detected and the obtained virtual labels of unlabeled data can be well encoded in a weighted regression model. These virtual labels are thereafter regressed with a linear model to calculate the projection matrix for dimensionality reduction. By this means, when the manifold or the clustering assumption of data is satisfied, the labels of labeled data can be correctly propagated to the unlabeled data; and thus, the proposed approach utilizes the labeled and the unlabeled data more effectively than previous work. Experimental results are carried out upon several databases, and the advantage of the new approach is well demonstrated. Feiping Nie 0001, Dong Xu 0001, Xuelong Li 0001, Shiming Xiang |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2011 | Regression Reformulations of LLE and LTSA With Locally Linear TransformationabstractLocally linear embedding (LLE) and local tangent space alignment (LTSA) are two fundamental algorithms in manifold learning. Both LLE and LTSA employ linear methods to achieve their goals but with different motivations and formulations. LLE is developed by locally linear reconstructions in both high- and low-dimensional spaces, while LTSA is developed with the combinations of tangent space projections and locally linear alignments. This paper gives the regression reformulations of the LLE and LTSA algorithms in terms of locally linear transformations. The reformulations can help us to bridge them together, with which both of them can be addressed into a unified framework. Under this framework, the connections and differences between LLE and LTSA are explained. Illuminated by the connections and differences, an improved LLE algorithm is presented in this paper. Our algorithm learns the manifold in way of LLE but can significantly improve the performance. Experiments are conducted to illustrate this fact. Shiming Xiang, Feiping Nie 0001, Chunhong Pan, Changshui Zhang |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2010 | Local and Global Regressive Mapping for Manifold Learning with Out-of-Sample ExtrapolationabstractOver the past few years, a large family of manifold learning algorithms have been proposed, and applied to various applications. While designing new manifold learning algorithms has attracted much research attention, fewer research efforts have been focused on out-of-sample extrapolation of learned manifold. In this paper, we propose a novel algorithm of manifold learning. The proposed algorithm, namely Local and Global Regressive Mapping (LGRM), employs local regression models to grasp the manifold structure. We additionally impose a global regression term as regularization to learn a model for out-of-sample data extrapolation. Based on the algorithm, we propose a new manifold learning framework. Our framework can be applied to any manifold learning algorithms to simultaneously learn the low dimensional embedding of the training data and a model which provides explicit mapping of the out-of-sample data to the learned manifold. Experiments demonstrate that the proposed framework uncover the manifold structure precisely and can be freely applied to unseen data. Yi Yang 0001, Feiping Nie 0001, Shiming Xiang, Yueting Zhuang |
AAAI | 3 |
| 2010 | A general kernelization framework for learning algorithms based on kernel PCA
Changshui Zhang, Feiping Nie 0001, Shiming Xiang |
Neurocomputing | 3 |
| 2010 | A general graph-based semi-supervised learning with novel class discovery
Feiping Nie 0001, Shiming Xiang, Yun Liu 0021, Changshui Zhang |
Neural Comput. Appl. | 2 |
| 2010 | Semi-Supervised Classification via Local Spline RegressionabstractThis paper presents local spline regression for semi-supervised classification. The core idea in our approach is to introduce splines developed in Sobolev space to map the data points directly to be class labels. The spline is composed of polynomials and Green's functions. It is smooth, nonlinear, and able to interpolate the scattered data points with high accuracy. Specifically, in each neighborhood, an optimal spline is estimated via regularized least squares regression. With this spline, each of the neighboring data points is mapped to be a class label. Then, the regularized loss is evaluated and further formulated in terms of class label vector. Finally, all of the losses evaluated in local neighborhoods are accumulated together to measure the global consistency on the labeled and unlabeled data. To achieve the goal of semi-supervised classification, an objective function is constructed by combining together the global loss of the local spline regressions and the squared errors of the class labels of the labeled data. In this way, a transductive classification algorithm is developed in which a globally optimal classification can be finally obtained. In the semi-supervised learning setting, the proposed algorithm is analyzed and addressed into the Laplacian regularization framework. Comparative classification experiments on many public data sets and applications to interactive image segmentation and image matting illustrate the validity of our method. Shiming Xiang, Feiping Nie 0001, Changshui Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | TurboPixel Segmentation Using Eigen-ImagesabstractTurboPixel (TP) is a powerful tool for image over-segmentation. It is fast and can yield a lattice-like structure of superpixel regions with uniform size. This paper presents a method to learn eigen-images from the image to be segmented. Such eigen-images are used to generate the evolution speed in the TP framework. The task is formulated as a problem of pixel clustering. Specifically, for the pixels in each local window, a linear transformation is introduced to map their color vectors to be the cluster indicator vectors. The errors under all such linear transformations are estimated and summed together to obtain an objective function, from which a global optimum is finally obtained. In this process, the eigen-images are constructed. Based on these eigen-images, multidimensional image gradient operator is defined to evaluate the gradient, which is supplied to the TP algorithm to obtain the final superpixel segmentations. The computational issues are discussed, and an image pyramid is introduced to speed up the computation. Comparative experiments illustrate the effectiveness of our method. Shiming Xiang, Chunhong Pan, Feiping Nie 0001, Changshui Zhang |
IEEE Trans. Image Process. | 1 |
| 2009 | Subspace Regularization: A New Semi-supervised Learning Method
Yan-Ming Zhang 0001, Xinwen Hou, Shiming Xiang, Cheng-Lin Liu 0001 |
ECML/PKDD (2) | 3 |
| 2009 | Nonlinear dimensionality reduction with relative distance comparison
Chunxia Zhang 0001, Shiming Xiang, Feiping Nie 0001, Yangqiu Song |
Neurocomputing | 2 |
| 2009 | Embedding new data points for manifold learning via coordinate propagation
Shiming Xiang, Feiping Nie 0001, Yangqiu Song, Changshui Zhang, Chunxia Zhang 0001 |
Knowl. Inf. Syst. | 1 |
| 2009 | Soft Constraint Harmonic Energy Minimization for Transductive Learning and its Two Interpretations
Changshui Zhang, Feiping Nie 0001, Shiming Xiang, Chenping Hou |
Neural Process. Lett. | 3 |
| 2009 | Semi-supervised discriminative classification with application to tumorous tissues segmentation of MR brain images
Yangqiu Song, Changshui Zhang, Jianguo Lee, Fei Wang 0001, Shiming Xiang, Dan Zhang 0007 |
Pattern Anal. Appl. | 5 |
| 2009 | Semi-supervised orthogonal discriminant analysis via label propagation
Feiping Nie 0001, Shiming Xiang, Yangqing Jia, Changshui Zhang |
Pattern Recognit. | 2 |
| 2009 | Extracting the optimal dimensionality for local tensor discriminant analysis
Feiping Nie 0001, Shiming Xiang, Yangqiu Song, Changshui Zhang |
Pattern Recognit. | 2 |
| 2009 | Interactive Natural Image Segmentation via Spline RegressionabstractThis paper presents an interactive algorithm for segmentation of natural images. The task is formulated as a problem of spline regression, in which the spline is derived in Sobolev space and has a form of a combination of linear and Green's functions. Besides its nonlinear representation capability, one advantage of this spline in usage is that, once it has been constructed, no parameters need to be tuned to data. We define this spline on the user specified foreground and background pixels, and solve its parameters (the combination coefficients of functions) from a group of linear equations. To speed up spline construction, K-means clustering algorithm is employed to cluster the user specified pixels. By taking the cluster centers as representatives, this spline can be easily constructed. The foreground object is finally cut out from its background via spline interpolation. The computational complexity of the proposed algorithm is linear in the number of the pixels to be segmented. Experiments on diverse natural images, with comparison to existing algorithms, illustrate the validity of our method. Shiming Xiang, Feiping Nie 0001, Chunxia Zhang 0001, Changshui Zhang |
IEEE Trans. Image Process. | 1 |
| 2009 | Nonlinear Dimensionality Reduction with Local Spline EmbeddingabstractThis paper presents a new algorithm for Nonlinear Dimensionality Reduction (NLDR). Our algorithm is developed under the conceptual framework of compatible mapping. Each such mapping is a compound of a tangent space projection and a group of splines. Tangent space projection is estimated at each data point on the manifold, through which the data point itself and its neighbors are represented in tangent space with local coordinates. Splines are then constructed to guarantee that each of the local coordinates can be mapped to its own single global coordinate with respect to the underlying manifold. Thus, the compatibility between local alignments is ensured. In such a work setting, we develop an optimization framework based on reconstruction error analysis, which can yield a global optimum. The proposed algorithm is also extended to embed out of samples via spline interpolation. Experiments on toy data sets and real-world data sets illustrate the validity of our method. Shiming Xiang, Feiping Nie 0001, Changshui Zhang, Chunxia Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2008 | Trace Ratio Criterion for Feature Selection
Feiping Nie 0001, Shiming Xiang, Yangqing Jia, Changshui Zhang, Shuicheng Yan |
AAAI | 2 |
| 2008 | A unified framework for semi-supervised dimensionality reduction
Yangqiu Song, Feiping Nie 0001, Changshui Zhang, Shiming Xiang |
Pattern Recognit. | 4 |
| 2008 | Contour graph based human tracking and action sequence recognition
Shiming Xiang, Feiping Nie 0001, Yangqiu Song, Changshui Zhang |
Pattern Recognit. | 1 |
| 2008 | Learning a Mahalanobis distance metric for data clustering and classification
Shiming Xiang, Feiping Nie 0001, Changshui Zhang |
Pattern Recognit. | 1 |
| 2007 | Optimal Dimensionality Discriminant Analysis and Its Application to Image RecognitionabstractDimensionality reduction is an important issue when facing high-dimensional data. For supervised dimensionality reduction, linear discriminant analysis (LDA) is one of the most popular methods and has been successfully applied in many classification problems. However, there are several drawbacks in LDA. First, it suffers from the singularity problem, which makes it hard to preform. Second, LDA has the distribution assumption which may make it fail in applications where the distribution is more complex than Gaussian. Third, LDA can not determine the optimal dimensionality for discriminant analysis, which is an important issue but has often been neglected previously. In this paper, we propose a new algorithm and endeavor to solve all these three problems. Furthermore, we present that our method can be extended to the two-dimensional case, in which the optimal dimensionalities of the two projection matrices can be determined simultaneously. Experimental results show that our methods are effective and demonstrate much higher performance in comparison to LDA. Feiping Nie 0001, Shiming Xiang, Yangqiu Song, Changshui Zhang |
CVPR | 2 |
| 2007 | Real-time Object Classification in Video Surveillance Based on Appearance LearningabstractClassifying moving objects to semantically meaningful categories is important for automatic visual surveillance. However, this is a challenging problem due to the factors related to the limited object size, large intra-class variations of objects in a same class owing to different viewing angles and lighting, and real-time performance requirement in real-world applications. This paper describes an appearance-based method to achieve real-time and robust objects classification in diverse camera viewing angles. A new descriptor, i.e., the multi-block local binary pattern (MB-LBP), is proposed to capture the large-scale structures in object appearances. Based on MB-LBP features, an adaBoost algorithm is introduced to select a subset of discriminative features as well as construct the strong two-class classifier. To deal with the non-metric feature value of MB-LBP features, a multi-branch regression tree is developed as the weak classifiers of the boosting. Finally, the error correcting output code (ECOC) is introduced to achieve robust multi-class classification performance. Experimental results show that our approach can achieve real-time and robust object classification in diverse scenes. Lun Zhang 0001, Stan Z. Li, Xiao-Tong Yuan, Shiming Xiang |
CVPR | 4 |
| 2007 | Extracting the Optimal Dimensionality for Discriminant AnalysisabstractFor classification task, supervised dimensionality reduction is a very important method when facing with high-dimensional data. Linear discriminant analysis (LDA) is one of the most popular method for supervised dimensionality reduction. However, LDA suffers from the singularity problem, which makes it hard to work. Another problem is the determination of optimal dimensionality for discriminant analysis, which is an important issue but often been neglected previously. In this paper, we propose a new algorithm to address these two problems. Experiments show the effectiveness of our method and demonstrate much higher performance in comparison to LDA. Feiping Nie 0001, Shiming Xiang, Yangqiu Song, Changshui Zhang |
ICASSP (2) | 2 |
| 2007 | Semi-Supervised Music Genre ClassificationabstractMusic genre classification is a hot topic in pattern recognition and signal processing. Classical supervised methods need lost of labeled music data to train a classifier. In this paper, we propose a semi-supervised genre classification algorithm which is developed on several labeled music tracks and lots of unlabelled tracks. Three features are extracted from the each music track and manifold regularization method is used to design the classifier. Experiments on a large number of test music data show that semi-supervised method can improve the classification accuracy. Yangqiu Song, Changshui Zhang, Shiming Xiang |
ICASSP (2) | 3 |
| 2007 | Neighborhood MinMax Projections
Feiping Nie 0001, Shiming Xiang, Changshui Zhang |
IJCAI | 2 |
| 2007 | Interactive Visual Object Extraction Based on Belief Propagation
Shiming Xiang, Feiping Nie 0001, Changshui Zhang, Chunxia Zhang 0001 |
MMM (1) | 1 |
| 2007 | Embedding New Data Points for Manifold Learning Via Coordinate Propagation
Shiming Xiang, Feiping Nie 0001, Yangqiu Song, Changshui Zhang, Chunxia Zhang 0001 |
PAKDD | 1 |
| 2006 | Texture Image Segmentation: An Interactive Framework Based on Adaptive Features and Transductive Learning
Shiming Xiang, Feiping Nie 0001, Changshui Zhang |
ACCV (1) | 1 |
| 2006 | Exemplar-Based Human Contour Tracking
Shiming Xiang, Feiping Nie 0001, Changshui Zhang |
ACCV (1) | 1 |
| 2006 | Contour Matching Based on Belief Propagation
Shiming Xiang, Feiping Nie 0001, Changshui Zhang |
ACCV (2) | 1 |
| 2006 | Spline Embedding for Nonlinear Dimensionality Reduction
Shiming Xiang, Feiping Nie 0001, Changshui Zhang, Chunxia Zhang 0001 |
ECML | 1 |