VLDB 2026 Research / reviewers in the wild / expert
Qingshan Liu 0001
dblp:95/1247
· DBLP profile ↗
299ranked-venue papers
21as first author
118since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 153 · 9 first-author · 51 since 2021Artificial intelligence and machine learning · 140 · 11 first-author · 40 since 2021Applied, interdisciplinary, general and emerging computing · 53 · 4 first-author · 34 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | More realistic and accurate precipitation nowcasting with Conditional Rectified Flow Transformers
Yunlong Zhou, Fanfan Ji, Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Teacher Agent: A Knowledge Distillation-Free Framework for Rehearsal-Based Video Incremental Learning
Shengqin Jiang, Yaoyu Fang, Haokui Zhang, Qingshan Liu 0001, Yuankai Qi, Yang Yang 0002, Peng Wang 0023 |
Int. J. Comput. Vis. | 4 |
| 2026 | Multiple Instance Learning Framework with Masked Hard Instance Mining for Gigapixel Histopathology Image Analysis
Sheng Huang 0001, Fengtao Zhou, Bo Liu 0005, Qingshan Liu 0001 |
Int. J. Comput. Vis. | 6 |
| 2026 | Attacks in Adversarial Machine Learning: A Systematic Survey from the Lifecycle Perspective
Baoyuan Wu, Zihao Zhu 0001, Li Liu 0036, Qingshan Liu 0001, Zhaofeng He 0001, Siwei Lyu |
Int. J. Comput. Vis. | 4 |
| 2026 | Sparse-view CT image reconstruction using conditional embedding fusion diffusion model
Chenchun Zhou, Yubao Sun, Jia Liu 0034, Qingshan Liu 0001 |
Neurocomputing | 5 |
| 2026 | VLM-driven fine-grained semantic regularization for low-light image enhancement
Zixuan Sun, Chuanwei Zhou, Hui Shuai, Qingshan Liu 0001 |
Multim. Syst. | 4 |
| 2026 | Adaptive frequency collaboration for remote sensing change detection
Feng Zhou 0006, Hui Shuai, Qingshan Liu 0001, Renlong Hang |
Neural Networks | 4 |
| 2026 | Defenses in Adversarial Machine Learning: A Systematic Survey From the Lifecycle PerspectiveabstractAdversarial phenomena have been widely observed in machine learning (ML) systems, especially those using deep neural networks. These phenomena describe situations where ML systems may produce predictions that are inconsistent and incomprehensible to humans in certain specific cases. Such behavior poses a serious security threat to the practical application of ML systems. To exploit this vulnerability, several advanced attack paradigms have been developed, mainly including backdoor attacks, weight attacks, and adversarial examples. For each individual attack paradigm, various defense mechanisms have been proposed to enhance the robustness of models against the corresponding attacks. However, due to the independence and diversity of these defense paradigms, it is challenging to assess the overall robustness of an ML system against different attack paradigms. This survey aims to provide a systematic review of all existing defense paradigms from a unified lifecycle perspective. Specifically, we decompose a complete ML system into five stages: pre-training, training, post-training, deployment, and inference. We then present a clear taxonomy to categorize representative defense methods at each stage. The unified perspective and taxonomy not only help us analyze defense mechanisms but also enable us to understand the connections and differences among different defense paradigms. It inspires future research to develop more advanced and comprehensive defense strategies. Baoyuan Wu, Mingli Zhu, Meixi Zheng, Zihao Zhu 0001, Shaokui Wei, Hongrui Chen, Danni Yuan, Li Liu 0036, Qingshan Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2026 | Enhanced Spatiotemporal Consistency for Image-to-LiDAR Data PretrainingabstractLiDAR representation learning has emerged as a promising approach to reducing reliance on costly and labor-intensive human annotations. While existing methods primarily focus on spatial alignment between LiDAR and camera sensors, they often overlook the temporal dynamics critical for capturing motion and scene continuity in driving scenarios. To address this limitation, we propose SuperFlow++, a novel framework that integrates spatiotemporal cues in both pretraining and downstream tasks using consecutive LiDAR-camera pairs. SuperFlow++ introduces four key components: (1) a view consistency alignment module to unify semantic information across camera views, (2) a dense-to-sparse consistency regularization mechanism to enhance feature robustness across varying point cloud densities, (3) a flow-based contrastive learning approach that models temporal relationships for improved scene understanding, and (4) a temporal voting strategy that propagates semantic information across LiDAR scans to improve prediction consistency. Extensive evaluations on 11 heterogeneous LiDAR datasets demonstrate that SuperFlow++ outperforms state-of-the-art methods across diverse tasks and driving conditions. Furthermore, by scaling both 2D and 3D backbones during pretraining, we uncover emergent properties that provide deeper insights into developing scalable 3D foundation models. With strong generalizability and computational efficiency, SuperFlow++ establishes a new benchmark for data-efficient LiDAR-based perception in autonomous driving. Xiang Xu 0009, Lingdong Kong, Hui Shuai, Liang Pan, Kai Chen 0026, Ziwei Liu 0002, Qingshan Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | Mitigating Symptom Heterogeneity in Multimodal Depression Estimation via Level Separation and Deviation RegressionabstractMultimodal Depression Estimation (MDE) aims to infer individual depression scores by analyzing various signals, such as visual, auditory, and language signals etc. Compared to Multimodal Depression Detection (MDD) methods that only provide discrete labels, MDE can provide a more refined score evaluation. However, symptom heterogeneity leads to differences in external behaviors among patients with similar depressive states, which limits the performance of direct regression MDE methods. To address this issue, we propose a Combined Depression Level and Deviation (CDLD) method for MDE, which separates samples at different depression levels and further analyzes subtle deviations within same level to improve estimation performance. Specifically, the Multilevel Depression Separation module constructs depression levels with inherent commonalities based on psychological theories and models the ordinality of these levels, thereby separating samples with different levels of depression. Building on this, the Level-specific Deviation Regression module contrasts sample features relative to level-specific anchors, regressing the subtle depression deviation. Finally, the depression level and deviation are integrated to infer the depression score more accurately. Experiments on the DAIC-WOZ, CMDC, SEARCH, and AVEC 2014 datasets demonstrate that the proposed coarse-to-fine method effectively mitigates the impact of symptom heterogeneity on depression estimation performance, showing significant advantages in MDE tasks. The code is publicly available athttps://github.com/LIU70KG/CDLD. Chengguang Liu, Shanmin Wang, Qingshan Liu 0001, Fei Wang 0064 |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | Vision-Language-Driven Prompt Learning for Weakly Supervised Semantic SegmentationabstractThe primary challenges in image-level weakly supervised semantic segmentation (WSSS) lie in addressing the under-activation issue of target pixels and mitigating the co-occurrence phenomenon in class activation maps. In recent years, Vision-Language Models (VLM) have demonstrated exceptional performance across various vision tasks, primarily attributed to their cross-modal semantic alignment capabilities achieved through contrastive learning mechanisms. Leveraging VLM’s capability to capture fine-grained visual-textual correspondences, this paper proposes a novel Vision-Language Driven Prompt Learning (VLD-PL) framework that addresses two fundamental challenges in WSSS by establishing explicit semantic correspondences between textual descriptors and visual components, ultimately enabling efficient semantic segmentation. The VLD-PL framework consists of two core components Auxiliary Class Matching (ACM) and Background Class Filtering (BCF). The ACM module dynamically identifies semantically relevant auxiliary classes through feature alignment between image and textual embeddings, effectively enlarging target activation while mitigating co-occurrence interference by expanding semantic coverage. Simultaneously, the BCF constructs image-specific background prompts and adaptively refines background feature representations, achieving precise suppression of irrelevant background regions. These dual mechanisms synergistically address both target localization accuracy and background noise suppression, achieving state-of-the-art performance on both the PASCAL VOC 2012 and MS COCO 2014 benchmarks. Junxia Li, Yubao Sun, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | FauForensics: Boosting Audio-Visual Deepfake Detection With Facial Action UnitsabstractThe rapid evolution of generative AI has intensified the threat of realistic audio-visual deepfakes, demanding robust and generalizable detection methods. Existing solutions primarily address unimodal (e.g., audio, visual) forgeries but struggle with multimodal manipulations due to inadequate handling of heterogeneous modality features and poor cross-dataset generalization. We propose FauForensics, a novel framework leveraging biologically invariant facial action units (FAUs), which are quantitative descriptors of facial muscle activity linked to emotion physiology. They serve as forgery-resistant representations that reduce domain dependency while capturing subtle synthetic-content disruptions. In addition, unlike prior clip-level comparisons, our method computes frame-wise audio-visual similarities via a fusion module with learnable cross-modal queries, dynamically aligning lip-audio relationships and mitigating feature heterogeneity. Experiments on four publicly available datasets show state-of-the-art performance with 5.17% average cross-dataset improvement over existing methods. Jian Wang 0129, Baoyuan Wu, Li Liu 0036, Qingshan Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | An Unsupervised Image Dehazing With Scene Geometry Prior for Road Traffic ScenariosabstractDespite advances in single image dehazing, robust dehazing for real-world road traffic scenes remains challenging due to scarce paired data, traffic-specific geometry, and real-time constraints. To address this issue, we propose a novel image prior for road traffic scenes, termed scene geometry prior (SGP), which leverages depth cues derived from vanishing point (VP) to provide geometry-aware guidance and reduce reliance on paired training data. Our SGP comprises two components: a global SGP (G-SGP) that captures the global geometric distribution and a non-local SGP (NL-SGP) that corrects the errors, among obstructions belonging to the same category, in captured global distribution. Building on the proposed prior, we develop a lightweight and unsupervised road traffic image dehazing network (RTDnet). It consists of a main sub-network guided by the G-SGP to reconstruct the haze-free image, alongside two auxiliary sub-networks that leverage the NL-SGP and VP information to respectively estimate transmission map, and atmospheric light. During training, we introduce an atmospheric scattering model (ASM)-driven mutual-boost learning mechanism (ASM-ML), which is rooted in Bayesian theory and effectively integrates the strengths of different priors without mutual interference while distilling ASM-based physical knowledge into each sub-network. By coupling SGP with ASM-ML, RTDnet can be trained without paired traffic data by exploiting traffic-specific geometry, whose accurate guidance reduces the reliance on large model capacity and enables lightweight real-time deployment. Experiments demonstrate that our RTDnet surpasses state-of-the-art competitors in terms of restoration quality, efficiency, and model size. Moreover, its robust dehazing performance benefits downstream tasks operating in hazy conditions. Mingye Ju, Tianyi Lyu, Chunming He, Qingshan Liu 0001, Kai-Kuang Ma |
IEEE Trans. Image Process. | 4 |
| 2026 | Identification of Genetic Risk Factors Based on Disease Progression Derived From Modeling Longitudinal Phenotype Latent Pattern RepresentationabstractThe characteristic of neurodegenerative disorders is the progressive impairment of memory and other cognitive functions. However, these existing imaging genetic methods only use longitudinal imaging phenotypes straightforwardly, ignoring the latent pattern of the longitudinal data in the progression process. The phenotypes across multiple time-points may exhibit the latent pattern that can be used to facilitate the understanding of the progression process. Accordingly, in this paper, we explore underlying complementary information from multiple time-points and simultaneously seek the underlying latent representation. With the complementarity of multiple time-points, the latent representation depicts data more comprehensively than each individual time-point, therefore mining effective longitudinal phenotype latent pattern representation. Specifically, we first propose two latent pattern representation (LPR) for longitudinal imaging phenotypes: linear LPR (lLPR), based on linear relationships between latent representation and each time-point, and nonlinear LPR (nonlLPR), based on neural networks to deal with nonlinear relationships. Then, we calculate the imaging genetic association based on the latent pattern representation. Finally, we conduct the experiments on both synthetic and real longitudinal imaging genetic data. Related experimental results validate that our proposed approach outperforms several competing algorithms, establishes strong associations, and discovers consistent longitudinal imaging genetic biomarkers, thereby guiding disease interpretation. Meiling Wang 0001, Wei Shao 0005, Daoqiang Zhang, Qingshan Liu 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2026 | Adversarial Pruning Networks for Compact 3D Gaussian Splattingabstract3D Gaussian Splatting holds significant potential for high-quality visual scene rendering. However, the large number of Gaussian primitives it requires poses challenges in memory consumption and practical deploy. Existing methods often rely on empirical criteria to prune Gaussians, which inevitably compromises visual quality. To address this, we propose Adversarial Pruning Networks (APNet), a framework that employs adversarial learning to balances the reduction of redundant Gaussians with the preservation of visual fidelity. APNet comprises a Gaussian Learning and Pruning Network (GLPN) and a Discriminative Network. GLPN incorporates the geometric information into the learning of Gaussians and prunes these Gaussians through a data-driven mask. Meanwhile, the Discriminative Network is trained to distinguish between synthesized and real images, acting as an adversary. Through adversarial pruning, APNet significantly reduces the number of Gaussians while rendering high-quality images. Extensive experiments on the Mip-NeRF360, Tanks & Temples, and Deep Blending datasets demonstrate that APNet achieves up to a 90% reduction in the original 3DGS while maintaining high rendering quality. Hui Shuai, Yubao Sun, Qingshan Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | 3D Scenes Motion Planning and Generation with Motion Diffusion Probabilistic ModelabstractGenerating natural and realistic human motion sequences under the constraints of 3D scenes is a highly challenging task, requiring not only the precise modeling of dynamic variations in human joints but also the rigorous consideration of intricate interactions between the human body and the surrounding environment. While recent advances in deep generative models show great potential in tackling these challenges, existing methods often result in unnatural human motions and human–environment penetration during generation. In order to cope with these issues, we propose a novel approach that divides human motion generation into two stages. The first stage employs a bidirectional long short-term memory network incorporated with full-connected layers to generate motion trajectory under the input conditions including the starting and ending positions and orientations of the human model and scene feature point clouds extracted from the surrounding environment. In the second stage, we design a conditional diffusion model, guided by the trajectory generated in the first stage and the embedding of 3D scene information, to generate human motion sequences within 3D scenes. We evaluate our framework through extensive experiments on the PROX datasets, which validates its effectiveness. The results show that our method significantly outperforms existing ones in enhancing human motion naturalness and reasonableness, and reducing human penetration. Yubao Sun, Guiyu Xia, Qingshan Liu 0001, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | LiMoE: Mixture of LiDAR Representation Learners from Automotive ScenesabstractLiDAR data pretraining offers a promising approach to leveraging large-scale, readily available datasets for enhanced data utilization. However, existing methods predominantly focus on sparse voxel representation, overlooking the complementary attributes provided by other LiDAR representations. In this work, we propose ${\color{Red}\text{Li}}{\color{Green}\text{MoE}}$, a framework that integrates the Mixture of Experts (MoE) paradigm into LiDAR data representation learning to synergistically combine multiple representations, such as range images, sparse voxels, and raw points. Our approach consists of three stages: i) Image-to-LiDAR Pretraining, which transfers prior knowledge from images to point clouds across different representations; ii) Contrastive Mixture Learning (CML), which uses MoE to adaptively activate relevant attributes from each representation and distills these mixed features into a unified 3D network; iii) Semantic Mixture Supervision (SMS), which combines semantic logits from multiple representations to boost downstream segmentation performance. Extensive experiments across eleven large-scale LiDAR datasets demonstrate our effectiveness and superiority. The code has been made publicly accessible. Xiang Xu 0009, Lingdong Kong, Hui Shuai, Liang Pan, Ziwei Liu 0002, Qingshan Liu 0001 |
CVPR | 6 |
| 2025 | Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR RepresentationsabstractLiDAR representation learning aims to extract rich structural and semantic information from large-scale, readily available datasets, reducing reliance on costly human annotations. However, existing LiDAR representation strategies often overlook the inherent spatiotemporal cues in LiDAR sequences, limiting their effectiveness. In this work, we propose LiMA, a novel long-term image-to-LiDAR Memory Aggregation framework that explicitly captures longer range temporal correlations to enhance LiDAR representation learning. LiMA comprises three key components: 1) a Cross-View Aggregation module that aligns and fuses overlapping regions across neighboring camera views, constructing a more unified and redundancy-free memory bank; 2) a Long-Term Feature Propagation mechanism that efficiently aligns and integrates multi-frame image features, reinforcing temporal coherence during LiDAR representation learning; and 3) a Cross-Sequence Memory Alignment strategy that enforces consistency across driving sequences, improving generalization to unseen environments. LiMA maintains high pretraining efficiency and incurs no additional computational overhead during downstream tasks. Extensive experiments on mainstream LiDAR-based perception benchmarks demonstrate that LiMA significantly improves both LiDAR semantic segmentation and 3D object detection. We hope this work inspires more effective pretraining paradigms for autonomous driving. The code has be made publicly accessible for future research. Xiang Xu 0009, Lingdong Kong, Song Wang 0019, Chuanwei Zhou, Qingshan Liu 0001 |
ICCV | 5 |
| 2025 | MMG: Manipulation-Aware Holistic Human Motion Generation from Sparse Tracking SignalsabstractGenerating realistic avatar motion via sparse tracking signals through VR devices is essential for enhancing the immersive user experience. Human-object manipulation behaviors not only affect hand motion but also significantly impact body motion. However, existing motion generation methods for human-object interactions overlook the coordinated coupling between body and hand motions during manipulations. Due to the diversity and complexity of holistic motion (body and hand motions simultaneously) in the latent motion space, generating physically plausible and temporally consistent holistic motion in real time, via the joint constraints imposed by sparse tracking signals and manipulation content, is a major challenge in the human motion generation task. We propose the manipulation-aware holistic human motion generation method (MMG) to help resolve this issue. In MMG, first, we construct a manipulation-aware holistic human motion generation framework that serially compresses the latent motion space distribution of the body and hand to generate realistic holistic human motion with object manipulation enabled. Second, to enhance the impact of object manipulation on holistic motion generation, MMG designs a novel object manipulation representation to extract effective manipulation features. Third, MMG is trained by an elaborate progressive manipulation-guided training algorithm to improve motion generation robustness and inference performance. Compared to state-of-the-art methods, MMG achieves up to a 39% improvement in the generated holistic motion quality with a 3.55 × speedup in generation performance. In manipulation-enabled scenes, MMG generates holistic motion in real time ($\geq 24 f p s$). Compared to the state-of-the-art methods, its perceived quality is significantly improved, and the task performance of holistic motion-required VR manipulation is high-significantly improved. This paper's code is at https://github.com/XRZ-BUAA/MMG. Xuehuai Shi, Renzhi Xiao, Yilun Sheng, Xiaobai Chen, Jieming Yin, Qingshan Liu 0001 |
ISMAR | 8 |
| 2025 | Text-driven human image generation with texture and pose control
Zhedong Jin, Guiyu Xia, Paike Yang, Mengxiang Wang, Yubao Sun, Qingshan Liu 0001 |
Neurocomputing | 6 |
| 2025 | DuPt: Rehearsal-based continual learning with dual prompts
Shengqin Jiang, Daolong Zhang, Fengna Cheng, Xiaobo Lu, Qingshan Liu 0001 |
Neural Networks | 5 |
| 2025 | Riemannian manifold-based disentangled representation learning for multi-site functional connectivity analysis
Wenyang Li, Mingxia Liu 0001, Qingshan Liu 0001 |
Neural Networks | 4 |
| 2025 | Unsupervised Cross-Domain Facial Expression Recognition via Class Adaptive Self-TrainingabstractUnsupervised Cross-Domain Facial Expression Recognition (CD-FER) aims to transfer the recognition ability from annotated source domains to unlabeled target domains. Despite the advancements in CD-FER techniques based on marginal distribution matching, certain inherent properties of facial expressions, such as implicit class margins and imbalanced class distributions, still leave room for improvement in existing models. In this paper, we propose a Class-Adaptive Self-Training (CAST) model for unsupervised CD-FER. In addition to domain alignment, the CAST model leverages self-training to learn pseudo labels and dually enhance aligned representations for explicit class distinction, considering implicit class margins. Furthermore, the CAST model conducts a comprehensive analysis of the negative effects of class distributions on pseudo-label learning from perspectives of class-level representation distributions and predicted probabilities, and subsequently proposes specific solutions. By jointly matching class-level representation distributions and class distributions, the CAST model successfully alleviates conditional distribution discrepancies between domains, which is particularly pertinent for facial expression properties. Experimental results, including assessments on multiple target domains and evaluations of multiple FER models, demonstrate the effectiveness, superiority, and universality of the CAST model. Shanmin Wang, Qingshan Liu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Deep Ring-Wise Block Network for Joint Association Analysis and Alzheimer's Disease Diagnosis With InterpretabilityabstractIn the brain imaging genomic tasks, it is challenging to provide accurate prior knowledge for estimating the association between quantitative traits (QTs) extracted from neuroimaging and genetic markers like single-nucleotide polymorphisms (SNPs). The hidden structural patterns in data limit the discovery of disease-related biomarkers. To this end, we present a deep ring-wise block network (RB-Net) for association analysis and brain disease diagnosis. Specifically, we first construct a new hidden structural pattern, namely, ring-wise block pattern, that satisfies both block and ring properties within the data before the association analysis. Subsequently, a RB-Net is developed via using an auto-encoder (AE) to represent imaging genomic data. Furthermore, we design the approach for joint association learning and automated brain disease diagnosis. Additionally, the optimization scheme based on alternating update is presented for solve the built ring-wise block-perception layer model. The performance of the designed method has been experimentally assessed on the brain imaging genomic data from the Alzheimer's Disease Neuroimaging Initiative (ADNI). The results validate that the proposed approach outperforms some competing approaches, establishes strong associations, and identifies crucial regions of interest (ROIs) across different imaging phenotypes associated with genetic risk biomarkers, thereby guiding disease interpretation and diagnosis prediction. Meiling Wang 0001, Wei Shao 0005, Daoqiang Zhang, Qingshan Liu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | Text-Augmented Semantic Feature Extraction and Difference Information Learning for Remote Sensing Image Change CaptioningabstractRemote sensing image change captioning (RSICC) aims to generate sentence descriptions about land cover changes in bitemporal images. The effective acquisition of semantic-level change information is critical for this task. However, due to the effects of illumination interference, appearance similarities and scale differences between different objects, it is difficult to accurately extract change information from bitemporal images. In this article, we attempt to take advantage of the high-level semantic information inherent in text and propose a text-augmented semantic feature extraction and difference information learning model for RSICC. Specifically, we first pre-define some text prompts for each remote sensing image and use the contrastive language-image pretraining (CLIP) model to select the most suitable text descriptions for them. Then, we adopt a refined segment anything model (SAM) to learn fine-grained visual features from each image, which is further enhanced via a designed selective text-image fusion (STIF) module. After that, to extract the semantic differences between bitemporal images, we propose a text-guided difference capture (TGDC) module capable of extracting multiscale difference information under the guidance of text differences between different-time images. Finally, a transformer-based caption generator is applied to generate sentence descriptions from the extracted difference information. In order to test the performance of our proposed model, we conduct comprehensive experiments on two widely used RSICC datasets, including LEVIR-CC and Dubai-CC. The experimental results show that our proposed model is able to outperform several state-of-the-art models, which validates the effectiveness of it. The codes of our proposed model will be released at https://github.com/Richardkimyo/TACC. Renlong Hang, Jinyu Luo, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Remote Sensing Object Counting With Online Knowledge LearningabstractEfficient models for remote sensing object counting are urgently required for applications in scenarios with limited computing resources, such as drones or embedded systems. A straightforward yet powerful technique to achieve this is knowledge distillation (KD), which steers the learning of student networks by leveraging the experience of already-trained teacher networks. However, it faces a pair of challenges. First, due to its two-stage training nature, a longer training period is essential, especially as the training samples increase. Second, despite the proficiency of teacher networks in transmitting assimilated knowledge, they tend to overlook the latent insights gained during their learning process. To address these challenges, we introduce an online distillation learning method for remote sensing object counting. It builds an end-to-end training framework that seamlessly integrates two distinct networks into a unified one. It comprises a shared shallow module, a teacher branch, and a student branch. The shared module serving as the foundation for both branches is dedicated to learning some primitive information. The teacher branch utilizes prior knowledge to reduce the difficulty of learning and guides the student branch in online learning. In parallel, the student branch achieves parameter reduction and rapid inference capabilities by means of channel reduction. This design empowers the student branch not only to receive privileged insights from the teacher branch but also to tap into the latent reservoir of knowledge held by the teacher branch during the learning process. Moreover, we propose a relation-in-relation distillation (RiRD) method that allows the student branch to effectively comprehend the evolution of the relationship of intralayer teacher features among different interlayer features. Extensive experiments on two challenging datasets demonstrate the effectiveness of our method, which achieves comparable performance to state-of-the-art (SOTA) methods despite using far fewer parameters. Shengqin Jiang, Yuan Gao 0053, Fengna Cheng, Renlong Hang, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Cell-Wise Self-Optimization: Making Pretrained Model Better in Remote Sensing CountingabstractRecently, remote sensing counting has drawn a lot of attention due to its wide application requirements. However, most existing approaches tend to focus on optimizing the network backend or designing new loss functions, and overlook a foundational component, i.e., the feature extractor, which is usually based on a pre-trained model. This oversight limits the potential for further performance improvement. In this paper, we propose a cell-wise self-optimization method to enhance the feature extractor. By leveraging the powerful representation capabilities of a pre-trained model, our method further refines them for counting tasks, notably improving network performance on limited remote sensing data. Specifically, we design a lightweight cell-wise architecture optimization based on a network architecture search algorithm. It sequentially builds lightweight cells in parallel with the blocks in a pre-trained model, leveraging a newly proposed random path selection strategy for training the optimization framework. Additionally, we propose extracting high-frequency information from the blocks in a pre-trained model as self-guidance optimization to facilitate the learning of the searched cells. Experimental results on several representative datasets demonstrate that our proposed method significantly improves counting performance while substantially reducing parameters. For instance, compared with our baseline, it improves performance by 22.2% in MAE and 17.0% in MSE on the Building dataset, while reducing parameters by 61.2%. Shengqin Jiang, Qian Jie, Fengna Cheng, Haokui Zhang, Yu Liu 0029, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Input-Regulated Remote Sensing Counting With Region UnderstandingabstractRemote sensing counting aims to automatically estimate the number of objects of interest from high-resolution aerial or satellite imagery, providing critical decision-making support in areas such as urban planning, traffic monitoring, and disaster response. While most existing methods leverage pre-trained models to enhance feature generalization, their performance is often hindered by the severe scarcity of annotated remote sensing data. This limits their generalizability in complex scenarios. To address these challenges, we propose a novel remote sensing counting network that effectively captures informative signals from relatively limited annotated data. Specifically, we first introduce a graph-driven input regulator that constructs a graph structure by modeling relationships among input features, effectively capturing intrinsic contextual dependencies. This structure allows the regulator to assign adaptive pixel-level weights to network inputs, prioritizing relevant signals while mitigating the risk of overfitting to a fixed data distribution. Second, we design a dynamic region-aware module that leverages fuzzy logic to adaptively identify and enhance highly discriminative local regions. In this way, it improves the robustness of the feature representations. Extensive experiments demonstrate the effectiveness of the proposed method compared with several state-of-the-art methods. Shengqin Jiang, Haojian Long, Fengna Cheng, Yuankai Qi, Xiaobo Lu, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Multi-Granularity Aggregation Network for Remote Sensing Few-Shot SegmentationabstractFew-shot semantic segmentation (FSS) aims to segment a query image using a limited number of densely annotated support images from the same category. Most existing conventional FSS methods are tailored for coping with images from natural scenarios. Unlike natural images, remote sensing images usually have a similar background context among the support and query image pairs, and more severe intraclass inconsistency exists due to overhead shooting views. However, facing such realistic and challenging remote sensing FSS tasks, the existing methods seldom consider these intrinsic characteristics from a unified viewpoint, thus leading to inferior results. To solve the above dilemma, we propose a multi-granularity aggregation network (MGANet) to progressively capture multi-granularity discriminative information, for tackling the remote sensing FSS task. Specifically, MGANet consists of a multi-granularity similarity (MGS) module and an adaptive multiprototype aggregation (AMPA) module. To fully utilize background context, MGS extracts multi-granularity support and query feature maps from the backbone network to calculate a holistic correlation by incorporating the background information. Next, to alleviate the intraclass inconsistency of remote sensing images, AMPA decomposes the support foreground region into mainstay and auxiliary subregions by the guidance of reverse prediction on support features, thus generating three types of prototypes by masked average pooling (MAP) on these paired features and masks. Furthermore, these multiprototypes are collaboratively interacted with the query features to pursue reinforced discriminative features, relying on prototype-aware slot attention (PASA). Extensive experiments on iSAID-$5^{i}$and LoveDA-$2^{i}$demonstrate well the superiority of the proposed MGANet. The source code is available athttps://github.com/CVL-hub/MGANet/. Shi-Feng Peng, Guosen Xie, Fang Zhao 0006, Xiangbo Shu, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Spatial-Temporal Interleaved Network for Efficient Action RecognitionabstractThe decomposition of 3D convolution will considerably reduce the computing complexity of 3D convolutional neural networks, yet simple stacking restricts the performance of neural networks. To this end, we propose a spatial-temporal interleaved network for efficient action recognition. By deeply analyzing this task, it revisits the structure of 3D neural networks in action recognition from the following perspectives. To enhance the learning of robust spatial-temporal features, we initially propose an interleaved feature interaction module to comprehensively explore cross-layer features and capture the most discriminative information among them. With regards to being lightweight, a boosted parallel pseudo-3D module is introduced with the goal of circumventing a substantial number of computations from the lower to middle levels while enhancing temporal and spatial features in parallel at high levels. Furthermore, we exploit a spatial-temporal differential attention mechanism to suppress redundant features in different dimensions while reaping the benefits of nearly negligible parameters. Lastly, extensive experiments on four action recognition benchmarks are given to show the advantages and efficiency of our proposed method. Specifically, our method attains a 15.2% improvement in Top-1 accuracy compared to our baseline, a stack of full 3D convolutional layers, on the Something-Something V1 dataset while utilizing only 18.2% of the parameters. Shengqin Jiang, Haokui Zhang, Yuankai Qi, Qingshan Liu 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | FRNet: Frustum-Range Networks for Scalable LiDAR SegmentationabstractLiDAR segmentation has become a crucial component of advanced autonomous driving systems. Recent range-view LiDAR segmentation approaches show promise for real-time processing. However, they inevitably suffer from corrupted contextual information and rely heavily on post-processing techniques for prediction refinement. In this work, we propose FRNet, a simple yet powerful method aimed at restoring the contextual information of range image pixels using corresponding frustum LiDAR points. First, a frustum feature encoder module is used to extract per-point features within the frustum region, which preserves scene consistency and is critical for point-level predictions. Next, a frustum-point fusion module is introduced to update per-point features hierarchically, enabling each point to extract more surrounding information through the frustum features. Finally, a head fusion module is used to fuse features at different levels for final semantic predictions. Extensive experiments conducted on four popular LiDAR segmentation benchmarks under various task setups demonstrate the superiority of FRNet. Notably, FRNet achieves 73.3% and 82.5% mIoU scores on the testing sets of SemanticKITTI and nuScenes. While achieving competitive performance, FRNet operates 5 times faster than state-of-the-art approaches. Such high efficiency opens up new possibilities for more scalable LiDAR segmentation. The code has been made publicly available at https://github.com/Xiangxu-0103/FRNet. Xiang Xu 0009, Lingdong Kong, Hui Shuai, Qingshan Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Self-Reflection Neural Network for Class-Incremental Object CountingabstractIn crowded scenarios, achieving the counting task of dynamically evolving categories is extremely challenging. In addition to grappling with challenges such as scale variations, severe occlusion and complex backgrounds, it is imperative to mitigate the issue of catastrophic forgetting. Previous approaches have heavily relied on leveraging historical data for knowledge distillation to tackle these difficulties. However, this strategy encounters two prominent obstacles: 1) Employing the teacher network from the previous stage for distillation incurs additional computational overhead during the training stage. 2) Although knowledge distillation can facilitate effective knowledge transfer, some inaccurate predictions from the teacher network may affect the knowledge acquisition in the current stage. To overcome these issues, we introduce a novel solution: a self-reflection neural network for class-incremental object counting. First, we construct a global-aware incremental regression branch that uses stacked transformer layers as backends to capture global information, while the final regression layers dynamically expand as categories increase. Furthermore, we introduce an uncertain estimation branch that selectively isolates certain feature maps to avoid some neurons updated with excessive gradient information, thereby enhancing the network plasticity while preserving stability. The output of this branch functions as a regularization signal, steering the learning process of the incremental regression branch. To foster a more robust retention of past knowledge, we propose a self-reflection loss. It employs the rectified outputs of global-aware incremental regression branch to encourage the network to reflect upon and refine its grasp of historical knowledge, effectively averting the pitfalls of inaccurate information. Our extensive experiments validate the effectiveness of our proposed method, achieving state-of-the-art results. Shengqin Jiang, Linfei Li, Fengna Cheng, Yuankai Qi, Qingshan Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | PointAttN: You Only Need Attention for Point Cloud CompletionabstractPoint cloud completion referring to completing 3D shapes from partial 3D point clouds is a fundamental problem for 3D point cloud analysis tasks. Benefiting from the development of deep neural networks, researches on point cloud completion have made great progress in recent years. However, the explicit local region partition like kNNs involved in existing methods makes them sensitive to the density distribution of point clouds. Moreover, it serves limited receptive fields that prevent capturing features from long-range context information. To solve the problems, we leverage the cross-attention and self-attention mechanisms to design novel neural network for point cloud completion with implicit local region partition. Two basic units Geometric Details Perception (GDP) and Self-Feature Augment (SFA) are proposed to establish the structural relationships directly among points in a simple yet effective way via attention mechanism. Then based on GDP and SFA, we construct a new framework with popular encoder-decoder architecture for point cloud completion. The proposed framework, namely PointAttN, is simple, neat and effective, which can precisely capture the structural information of 3D shapes and predict complete point clouds with detailed geometry. Experimental results demonstrate that our PointAttN outperforms state-of-the-art methods on multiple challenging benchmarks. Code is available at: https://github.com/ohhhyeahhh/PointAttN Dongyan Guo, Junxia Li, Qingshan Liu 0001, Chunhua Shen |
AAAI | 5 |
| 2024 | Transcending Forgery Specificity with Latent Space Augmentation for Generalizable Deepfake DetectionabstractDeepfake detection faces a critical generalization hurdle, with performance deteriorating when there is a mismatch between the distributions of training and testing data. A broadly received explanation is the tendency of these detectors to be overfitted to forgery-specific artifacts, rather than learning features that are widely applicable across various forgeries. To address this issue, we propose a simple yet effective detector called LSDA (Latent Space Data Augmentation), which is based on a heuristic idea: representations with a wider variety of forgeries should be able to learn a more generalizable decision boundary, thereby mitigating the overfitting of method-specific features (see Fig. 1). Following this idea, we propose to enlarge the forgery space by constructing and simulating variations within and across forgery features in the latent space. This approach encompasses the acquisition of enriched, domain-specific features and the facilitation of smoother transitions between different forgery types, effectively bridging domain gaps. Our approach culminates in refining a binary classifier that leverages the distilled knowledge from the enhanced features, striving for a generalizable deepfake detector. Comprehensive experiments show that our proposed method is surprisingly effective and transcends state-of-the-art detectors across several widely used benchmarks. Zhiyuan Yan 0002, Yuhao Luo 0002, Siwei Lyu, Qingshan Liu 0001, Baoyuan Wu |
CVPR | 4 |
| 2024 | 4D Contrastive Superflows are Dense 3D Representation Learners
Xiang Xu 0009, Lingdong Kong, Hui Shuai, Liang Pan, Kai Chen 0026, Ziwei Liu 0002, Qingshan Liu 0001 |
ECCV (1) | 8 |
| 2024 | 3D human model guided pose transfer via progressive flow prediction network
Furong Ma, Guiyu Xia, Qingshan Liu 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | DiFormer: A Difference Transformer Network for Remote Sensing Change DetectionabstractChange detection (CD) is one of the most important methods for monitoring land surface changes. Recently, transformer-based models have been employed to CD. However, most of them focus on modeling the correlation within each image, while cannot well model the difference between bi-temporal images. In this paper, we propose a difference transformer network (DiFormer) to address this issue. Specifically, we propose a Token Exchange-based Difference Evaluation (TEDE) module to generate the inconsistency between the changed region and the surrounding context to highlight the difference between the bi-temporal images. In addition, to obtain semantically rich exchangeable tokens, we design a Multi-scale Semantic Perception (MSP) module, which provides assistance for difference modeling. In order to test the performance of DiFormer, we conduct qualitative and quantitative experiments on two public datasets, including LEVIR-CD and S2Looking. Experimental results show that our proposed DiFormer is able to achieve better results than several state-of-the-art models, with F1 of 92.15% on the LEVIR-CD dataset and 66.31% on the S2Looking dataset. Renlong Hang, Shanmin Wang, Qingshan Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Deep Precipitation Nowcasting With Dual Regions Displacement Information and Global Spatiotemporal Representations LearningabstractThe deep precipitation nowcasting using radar echo map prediction can mitigate the socio-economic impact of extreme precipitation events. Existing methods employ long short-term memory (LSTM) to extract rich precipitation features. However, existing methods often combine the learning and modeling of rain and nonrain regions in a single module, without clearly distinguishing their different features and motion patterns, which impairs the spatial distribution and precipitation intensity prediction of rainfall. Moreover, these LSTMs only capture local spatiotemporal features, while ignoring the global spatiotemporal features, resulting in prediction results lacking structural and strength consistency. Therefore, we propose a dual regions center displacement (DRCD) module, which separately learns and models the spatial information of rainfall and nonrainfall regions and employs this module to estimate the locations and intensity residuals of the future regions. Moreover, we also introduce a novel Global LSTM module (GLSTM) that learns the global spatiotemporal features from the sequences, which can estimate the structure and intensity of dual regions. Extensive experiments demonstrate that our method has superior or competitive performance over the state-of-the-art precipitation nowcasting methods and has the potential to be implemented as an alternative product globally. Fanfan Ji, Yunlong Zhou, Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Geometric Understanding of Discriminability and Transferability for Visual Domain AdaptationabstractTo overcome the restriction of identical distribution assumption, invariant representation learning for unsupervised domain adaptation (UDA) has made significant advances in computer vision and pattern recognition communities. In UDA scenario, the training and test data belong to different domains while the task model is learned to be invariant. Recently, empirical connections between transferability and discriminability have received increasing attention, which is the key to understand the invariant representations. However, theoretical study of these abilities and in-depth analysis of the learned feature structures are unexplored yet. In this work, we systematically analyze the essentials of transferability and discriminability from the geometric perspective. Our theoretical results provide insights into understanding the co-regularization relation and prove the possibility of learning these abilities. From methodology aspect, the abilities are formulated as geometric properties between domain/cluster subspaces (i.e., orthogonality and equivalence) and characterized as the relation between the norms/ranks of multiple matrices. Two optimization-friendly learning principles are derived, which also ensure some intuitive explanations. Moreover, a feasible range for the co-regularization parameters is deduced to balance the learning of geometric structures. Based on the theoretical results, a geometry-oriented model is proposed for enhancing the transferability and discriminability via nuclear norm optimization. Extensive experiment results validate the effectiveness of the proposed model in empirical applications, and verify that the geometric abilities can be sufficiently learned in the derived feasible range. You-Wei Luo, Chuan-Xian Ren, Xiao-Lin Xu, Qingshan Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | A Unified Object Counting Network With Object Occupation PriorabstractThe counting task, which plays a fundamental role in numerous applications (e.g., crowd counting, traffic statistics), aims to predict the number of objects with various densities. Existing object counting tasks are designed for a single object class. However, it is inevitable to encounter newly coming data with new classes in our real world. We name this scenario as evolving object counting. In this paper, we build the first evolving object counting dataset and propose a unified object counting network as the first attempt to address this task. The proposed network consists of two key components: a class-agnostic mask module and a class-incremental module. The class-agnostic mask module learns generic object occupation prior by predicting a class-agnostic binary mask (e.g., 1 denotes there exists an object at the considering position in an image and 0 otherwise). The class-incremental module is used to handle new classes and provides discriminative class guidance for density map prediction. The combined outputs of the class-agnostic mask module and image feature extractor are used to predict the final density map. When new classes arrive, we first add new neural nodes to the last regression and classification layers of the class-incremental module. Then, instead of retraining the model from scratch, we utilize knowledge distillation to help the model retain and consolidate what it has previously learned. We also employ a support sample bank to store a small number of typical training samples for each class, which are used to prevent the model from forgetting key information from old data. With this design, our model can efficiently and effectively adapt to new classes while maintaining good performance on already-seen data without large-scale retraining. Extensive experiments on the collected dataset demonstrate favorable performance. The dataset and code will be available at:https://github.com/Tanyjiang/EOCO. Shengqin Jiang, Fengna Cheng, Yuankai Qi, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Motion Compression Using Structurally Connected Neural NetworkabstractMotion compression technologies can significantly reduce the redundant information of motion data and increase the efficiency of storage and transmission. Current methods mainly utilize some ready-made universal algorithms, such as signal processing and dimensionality reduction, to model the statistical characteristics of motion data, while the individual structure of motion data is ignored. In this paper, we propose to use a deep neural network with specially designed architecture to represent motion data considering the similarity between the articulated structure of a human skeleton and the architecture of neural networks. The network parameters are then taken as the compressed data. We design a structurally connected network which just looks like a human skeleton. Within the network, only the neurons corresponding to the joints connected to each other in a human skeleton are connected. It effectively exploits the correlations between connected joints to cut down the unnecessary connections between the neurons, which leads to the significant improvement of compression efficiency. Additionally, we extract the two inherent DOFs instead of the original three DOFs of each joint by representing its movement on a sphere according to the rigidity of the articulated human skeleton. This actually achieves the theoretically lossless pre-compression with the ratio of 3:2. Extensive experiment results demonstrate the superior performances of the proposed model at the high compression ratios over other state-of-the-art methods. Guiyu Xia, Wenkai Ye, Yubao Sun, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | A Regionally Indicated Visual Grounding Network for Remote Sensing ImagesabstractVisual grounding (VG) is essential to promote the human-computer interaction in object detection tasks. Most of the current VG methods mainly focus on grounding the target objects in natural images with simple language expressions. They cannot generalize well to remote sensing images, where the target objects only cover a small fraction (e.g., 0.34%) of the whole scene and the language expression is complex. To address these challenges, we propose a regionally indicated network (RINet) for remote sensing VG in this article. Specifically, RINet first exploits DarkNet-53 and BERT to extract visual and language features, respectively. Then, these features are fed into a regional indication generator (RIG) to generate an initial indication map, which indicates the possibility of each region containing the target object. This indication map is fine-tuned by taking advantage of a high-resolution detailed feature via a comprehensive alignment module (CAM) and a correction gate (CG). In CAM, a word contribution learner is designed to evaluate the importance of each word and make it pay more attention to the words easily ignored before. The whole fine-tuning process is repeated several rounds so that the complex language information can be fully explored and the region containing the target object is located more accurately. Finally, a detection head is adopted to ground the target object. To test the performance of our proposed model, we conduct experiments on two public remote sensing datasets, including RSVG and DIOR-RSVG. The experimental results show that our proposed RINet can outperform several state-of-the-art models significantly, which validates its effectiveness. The source code of our proposed model will be released athttps://github.com/KevinDaldry/RINet. Renlong Hang, Siqi Xu, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | AANet: An Ambiguity-Aware Network for Remote-Sensing Image Change DetectionabstractRemote sensing image change detection (CD) task plays an important role in land-use survey, city construction investigation and other vital industries. Recently, deep learning has become a mainstream method for this task due to its satisfactory performance in most cases. However, it often suffers from difficulties in dealing with ambiguity regions, where pseudo-changes happen or real changes are corrupted. In this article, we propose an ambiguity-aware network (AANet) to address the aforementioned issue. Specifically, our network firstly adopts convolutional layers to learn features from dual-temporal images. After that, an ambiguity refinement module (ARM) is designed to extract the ambiguity regions and then difference features are generated based on it. Considering that the scales of different changed objects vary, a weight rearrangement module (WRM) is proposed to fuse the difference features from different layers. In order to test the performance of our proposed model, we conduct experiments on three benchmark datasets, including SYSU-CD, SVCD, and LEVIR-CD. The experimental results show that our model can outperform several state-of-the-art models on all three datasets, which validates the effectiveness of it. The source code of our proposed model will be released at https://github.com/KevinDaldry/AANet. Renlong Hang, Siqi Xu, Panli Yuan, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A Closed-Loop Circular Regression Network for 2-m Air Temperature Downscaling Over Southwestern ChinaabstractHigh-quality meteorological grid data are essential for meteorological research and applications, especially in regional scales. Statistical downscaling (SD) is an efficient method to provide more detailed information at spatial scale, and has already been implemented in many regions. In recent years, deep convolutional neural networks have exhibited promising performance in SD, effectively learning non-linear mappings from the low-resolution (LR) meteorological data to its corresponding high-resolution (HR) one. Nevertheless, existing deep-learning-based SD approaches may encounter two potential limitations. First, most of the previous deep-learning-based downscaling algorithms utilize a supervised learning framework, which necessitates the formation of data pairs consisting of HR labels and LR data for model training. However, the acquisition of meteorological data at regional scale is more challenging, making it difficult to meet the training requirements of traditional supervised-learning-based downscaling models in some cases. Second, learning the non-linear mapping between LR and HR meteorology data is typically an ill-posed issue, which means that there are infinite HR solutions for the same LR sample, making it harder to find the optimal solution within the large solution space, especially in the case of insufficient HR training labels. In this study, we propose a closed-loop circular regression network for simultaneous restoration of medium-resolution (MR) and HR 2m air temperature over Sichuan and surrounding areas, China. The model leverages the circular structure consistency to train both the downscaling and upscaling networks simultaneously. Specifically, in terms of insufficient HR labels, we introduce an additional constraint of MR supervision information to reduce the space of possible functions, forming a gradual downscaling process from LR to MR to HR data. Besides, we also establish an extra upscaling mapping from HR to MR to LR, which forms a circular consistency constraint on LR and MR data to provide additional supervision. Extensive experiments demonstrate that the proposed algorithm attains a Root Mean Square Error (RMSE) of 0.84 when utilizing 50% of the training data and 0.77 when using 75%. This performance surpasses that of many classic supervised-learning-based SD methods, even the complete supervised information is not utilized. Guangyu Liu 0002, Renlong Hang, Rui Zhang 0049, Chunxiang Shi, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | A Short-Long Term Sequence Learning Network for Precipitation NowcastingabstractPrecipitation nowcasting is a critical task for various applications, such as disaster mitigation, water resource management, traffic safety, and agricultural planning. In recent years, deep learning methods equipped with long short-term memory (LSTM) have become a mainstream method for this task. Typically, these methods take as input a radar echo sequence and focus on learning the temporal features of precipitation. However, due to the limited spatial modeling ability of the current LSTM modules, they cannot sufficiently learn the spatial–temporal features of precipitation. To address this issue, we propose in this article a short-long term sequence learning (SLTSL) network. SLTSL mainly consists of the short-term sequence learning module (SSLM) and the long-term sequence learning module (LSLM). SSLM weights and integrates the spatial distribution features of precipitation at each location of the short-term sequence by various weighted operations, where the short-term sequence is obtained by SSLM through matrix concatenation of the feature maps of four adjacent moments. LSLM integrates all feature maps into a long-term sequence through matrix fusion and then captures the temporal features of precipitation at all moments from the long-term sequence by means of multistate transitions and aggregation. In order to test the performance of the proposed network, we carry out experiments on three widely used datasets, including RadarCIKM, TAASRAD19, and RadarKNMI. The experimental results demonstrate that our proposed network can achieve superior or comparable performance to several state-of-the-art baseline methods. Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Real-Time Statistical Weather Estimation and Prediction for Tropical Cyclone Intensity in an Interpretable Manner via Causal InferenceabstractCurrently, integrating mathematical-physical (MP) knowledge into the deep learning (DL) model in an interpretable manner for tropical cyclone (TC) intensity estimation and prediction remains a challenge. In this article, we propose the statistical weather prediction for tropical cyclone intensity (SWP-TCI) model, focusing on real-time estimation and prediction of the TC intensity over the Numerical weather prediction (NWP). SWP-TCI incorporates multiple physical factors and utilizes the causal statistical fusion ensemble module (CSFEM) to integrate this physical knowledge into the model. By leveraging constraints from various physical factors, SWP-TCI can identify more distinct TC features from the satellite cloud images. CSFEM is developed using causal inference (CI), establishing the causal relationship between the TC and physical factors, thus supporting the learning of SWP-TCI and enhancing the model’s interpretability. In experiments, our results are obtained without using the historical sequence of TC intensities. The average MAE of 24-h predictions across from 2015 to 2021 is 5.9 m/s, standing in a close comparison with the official object forecast methods employed worldwide. Luhui Yue, Rui Zhang 0049, Jiamu Ding, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Estimating Tropical Cyclone Intensity Using an STIA Model From Himawari-8 Satellite Images in the Western North Pacific BasinabstractAnalyzing the temporal evolution of historical tropical cyclone (TC) structures is essential for accurate TC intensity estimation. In this article, a novel spatiotemporal interaction attention (STIA) model is proposed to estimate TC intensity using Himawari-8 data in the western North Pacific (WNP) basin. The model incorporates a spatial feature extraction module and a spatiotemporal interaction module, which leverage historical satellite images. Based on a sequence of observed satellite images, the spatial feature extraction module is expected to extract spatial features of each TC frame. After that, the spatiotemporal interaction module comprising the temporal–spatial (TS) module and the spatial–temporal (ST) module is responsible for fusing the temporal and spatial features of each frame. The experimental data are composed of Himawari-8 infrared (IR) and water vapor (WV) images from 2015 to 2020 with a time interval of 1 h. The model is trained on images from 2015 to 2018 and evaluated on images from 2019 to 2020. Ablation experiments are conducted to analyze the impact of the number of frames, and the ST and TS modules. The results demonstrate that using 18-frame inputs yields the best performance, achieving an overall root-mean-square error (RMSE) of 3.61 m/s and a mean absolute error (MAE) of 2.83 m/s. In addition, the ST and TS modules significantly contribute to enhancing the accuracy of TC intensity estimation. The performance of the STIA model already surpasses the state-of-the-art benchmarks, demonstrating its excellence in TC intensity estimation. Rui Zhang 0049, Luhui Yue, Qingshan Liu 0001, Renlong Hang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Predicting Tropical Cyclone Rapid Intensification in Western North Pacific Basin Using a TDA-RI Model From Digital Typhoon DatasetabstractThis article presents a novel temporal-differential attention rapid intensification (TDA-RI) model for predicting tropical cyclone (TC) rapid intensification (RI) in the western North Pacific (WNP) Basin. The model leverages satellite image data from the Digital Typhoon dataset, effectively capturing the characteristics of TC RI through temporal-differential images. It incorporates two key components: an intensity embedding spatial-temporal (IEST) module, and a temporal-differential normed attention (TDNA) module. The IEST module embeds intensity information into the spatial-temporal (ST) contexts of image sequences, capturing dynamic changes in intensity and image features over time and space. The TDNA module focuses on capturing changes and evolution patterns of TCs in the temporal dimension. The dataset is divided into a training dataset spanning from 1989 to 2017 and a test dataset covering 2018 to 2022. Experimental results demonstrate that the TDA-RI model exhibits excellent RI predictive capabilities, achieving ROC AUC and PR AUC scores of 0.901 and 0.465, respectively. Compared to other advanced models, the TDA-RI model shows significant advantages in terms of recall and false positive rate. Furthermore, the model can achieve satisfactory performance even without historical intensity information, highlighting its potential for real-time RI forecasting. Rui Zhang 0049, Luhui Yue, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Masked Spectral-Spatial Feature Prediction for Hyperspectral Image ClassificationabstractTransformer has emerged as a preferred method for hyperspectral (HS) image classification due to its ability to model long-range dependency. Whereas the transformer contains numerous parameters and further available labeled HS data is limited, which makes it difficult to get a well-trained transformer. Accordingly, we propose a novel HS image classification method called masked spectral–spatial feature prediction (MSSFP). It aims at helping the transformer understand the complicated spectral–spatial structures without labeled HS data, further improving the classification performance. Specifically, the input HS cube is first divided into two sequences along spectral and spatial dimensions, respectively. Then, a portion of these two sequences are masked out and we train a transformer-based encoder–decoder network to predict the hand-crafted features of masked regions. After pretraining, the encoder is fine-tuned to derive two classification results from input spectral and spatial sequences. Finally, spectral and spatial results are aggregated adaptively based on uncertainty comparison. In comparison experiments, MSSFP outperforms several state-of-the-art HS image classification methods on three benchmark datasets including Indian Pines (IP), Houston (HU), and Pavia University (PUS). Feng Zhou 0006, Guowei Yang 0002, Renlong Hang, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Compare and Focus: Multi-Scale View Aggregation for Crowd CountingabstractRecently, some state-of-the-art (SOTA) methods have designed dedicated context extractors to capture the global information that serves as a key clue for describing crowd density. A promising alternative is the transformer-based model which inherently captures long-range context dependencies. Recent related studies have made impressive progress, yet the following issues remain: (1) The size of the heads in the image is large near and small far away. The existing models fail to cope well with these variations. (2) There is an imbalance in the distribution of samples across different densities in the dataset, which leads to poor network performance on density distributions with a small number of samples. To address these issues, we propose to aggregate multi-scale views through Compare and Focus strategies. In terms of the first strategy, we mine differential hints from multi-scale view features to capture heads of varying sizes. This can effectively reduce the influence of redundant information while perceiving the subtleties of various view inputs, making it simpler to establish discriminative representations. As for the second strategy, we introduce a new activation function to formulate the Region of Interest (ROI) extraction module that enables the network to focus on relevant regions effectively. It can alleviate the extreme distribution imbalance of samples with different densities. Finally, several experiments show that our method achieves SOTA performance on four challenging datasets. Shengqin Jiang, Jialu Cai, Haokui Zhang, Yu Liu 0029, Qingshan Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Highway Visibility Level Prediction Using Geometric and Visual Features Driven Dual-Branch Fusion NetworkabstractAutomatic prediction of visibility from surveillance images can provide timely warnings for transportation management departments and drivers, which is of great significance for improving the driving safety of highways in foggy weather. The currently mainstream prediction models are based on deep networks, which mainly learn visual features as clues to predict visibility levels. However, the geometric features of highways, such as the lane lines that can be observed from surveillance images are also important clues to reflect the visibility. Therefore, we propose a dual-branch fusion network driven by both geometric and visual features to achieve robust and effective visibility prediction. Specifically, we first exploit dual branches to learn geometric features of highway and deep visual features from the foggy surveillance images, respectively. We then design a fused classification module to fuse the dual-branch features to predict the visibility level. In order to simultaneously purify features during the fusion process, it utilizes a road attention block to highlight the deep visual features corresponding to the highway road area, and a lane length estimation block to extract the feature of the length of observable lane lines. Therefore, the dual-branch features can be adaptively fused to boost prediction performance. Meanwhile, we construct a real-scene foggy image dataset, which are all gathered from the surveillance video of real highways in China. We validate the effectiveness of the proposed network on this real-scene dataset and the synthetic dataset FRIDA. The experimental results show that our method can predict visibility levels more accurately than multiple existing methods. Yubao Sun, Jihui Tang, Qingshan Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Soft Weight Pruning for Cross-Domain Few-Shot Learning With Unlabeled Target DataabstractCross-domain few-shot learning (CDFSL) has received great interest for its effectiveness in solving the problem of the shift between source and target domains in few-shot scenarios. To extract more representative features, recent CDFSL works have exploited small-scale unlabeled samples from the target domain during the feature extraction phase. Existing self-supervised CDFSL methods, however, typically fine-tune the weights of the pre-trained model without taking into account the mismatch between source and target domains. To address this shortcoming, we introduce a self-supervised soft weight pruning strategy for cross-domain few-shot classification tasks with unlabeled target data. Starting from a pre-trained network from the source domain, our approach iterates between pruning out the relatively unimportant connections of the network and reactivating the pruned connections in a joint contrastive and$L^{2}$-SPregularized training framework. By combining the soft weight pruning strategy and regularization, our method effectively restricts redundant weights while simultaneously learning crucial features for both source and target tasks. Our approach, in comparison to other methods, does not involve any additional modules in the models; however, it can still achieve remarkable performance. Our approach can be efficiently incorporated into a variety of contrastive learning methods in a plug-and-play fashion. Extensive experimental results on several benchmark datasets demonstrate that our proposed method outperforms existing representative cross-domain few-shot methods by a large margin. The code for our work can be found athttps://github.com/nuistji/swp-cdfsl. Fanfan Ji, Xiao-Tong Yuan, Qingshan Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Adaptive Activation Network for Weakly Supervised Semantic SegmentationabstractClass activation maps generated by image classifiers are widely used as priors for image-level weakly supervised semantic segmentation. However, these activation maps mainly focus on the sparse discriminative regions, which has been a bottleneck for the segmentation task. Based on our observations, the activation maps actually capture almost the entire target regions, and some regions with lower activation values are easily to be neglected. Thus, to solve the issue, we propose an adaptive activation network with two branches to recalibrate the low-confidence regions in the activation maps. Specifically, an activation enhancement branch is designed to redistribute the activation values by leveraging attention mechanism. Since multi-scale images can provide complementary information, a scale adaptation branch is paralleled to supervise the activation enhancement branch. The mutual supervision and fusion of the two branches can promote the less-discriminative parts, and deactivate the background regions. Based on them, a simple yet effective denoising module is proposed to further improve the quality of pseudo masks, which makes use of the large scale predictions of the trained segmentation network. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 benchmarks show that our method achieves state-of-the-art performance, demonstrating the effectiveness of our algorithm. Code will be made publicly available. Junxia Li, Deshuo Shi, Dongyan Guo, Qingshan Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | A Deep Learning Framework for Start-End Frame Pair-Driven Motion SynthesisabstractA start-end frame pair and a motion pattern-based motion synthesis scheme can provide more control to the synthesis process and produce content-various motion sequences. However, the data preparation for the motion training is intractable, and concatenating feature spaces of the start-end frame pair and the motion pattern lacks theoretical rationality in previous works. In this article, we propose a deep learning framework that completes automatic data preparation and learns the nonlinear mapping from start-end frame pairs to motion patterns. The proposed model consists of three modules: action detection, motion extraction, and motion synthesis networks. The action detection network extends the deep subspace learning framework to a supervised version, i.e., uses the local self-expression (LSE) of the motion data to supervise feature learning and complement the classification error. A long short-term memory (LSTM)-based network is used to efficiently extract the motion patterns to address the speed deficiency reflected in the previous optimization-based method. A motion synthesis network consists of a group of LSTM-based blocks, where each of them is to learn the nonlinear relation between the start-end frame pairs and the motion patterns of a certain joint. The superior performances in action detection accuracy, motion pattern extraction efficiency, and motion synthesis quality show the effectiveness of each module in the proposed framework. Guiyu Xia, Qingshan Liu 0001, Yubao Sun |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Unsupervised Video Object Segmentation with Online Adversarial Self-TuningabstractThe existing unsupervised video object segmentation methods depend heavily on the segmentation model trained offline on a labeled training video set, and cannot well generalize to the test videos from a different domain with possible distribution shifts. We propose to perform online fine-tuning on the pre-trained segmentation model to adapt to any ad-hoc videos at the test time. To achieve this, we design an offline semi-supervised adversarial training process, which leverages the unlabeled video frames to improve the model generalizability while aligning the features of the labeled video frames with the features of the unlabeled video frames. With the trained segmentation model, we further conduct an online self-supervised adversarial finetuning, in which a teacher model and a student model are first initialized with the pre-trained segmentation model weights, and the pseudo label produced by the teacher model is used to supervise the student model in an adversarial learning framework. Through online finetuning, the student model is progressively updated according to the emerging patterns in each test video, which significantly reduces the test-time domain gap. We integrate our offline training and online fine-tuning in a unified framework for unsupervised video object segmentation and dub our method Online Adversarial Self-Tuning (OAST). The experiments show that our method outperforms the state-of-the-arts with significant gains on the popular video object segmentation datasets. Tiankang Su, Huihui Song 0003, Dong Liu 0002, Bo Liu 0005, Qingshan Liu 0001 |
ICCV | 5 |
| 2023 | ACLM: Adaptive Compensatory Label Mining for Facial Expression Recognition
Chengguang Liu, Shanmin Wang, Hui Shuai, Qingshan Liu 0001 |
ICIG (4) | 4 |
| 2023 | Temporally Efficient Gabor Transformer for Unsupervised Video Object SegmentationabstractSpatial-temporal structural details of targets in video (e.g. varying edges, textures over time) are essential to accurate Unsupervised Video Object Segmentation (UVOS). The vanilla multi-head self-attention in the Transformer-based UVOS methods usually concentrates on learning the general low-frequency information (e.g. illumination, color), while neglecting the high-frequency texture details, leading to unsatisfying segmentation results. To address this issue, this paper presents a Temporally efficient Gabor Transformer (TGFormer) for UVOS. The TGFormer jointly models the spatial dependencies and temporal coherence intra- and inter-frames, which can fully capture the rich structural details for accurate UVOS. Concretely, we first propose an effective learnable Gabor filtering Transformer to mine the structural texture details of the object for accurate UVOS. Then, to adaptively store the redundant neighboring historical information, we present an efficient dynamic neighboring frame selection module to automatically choose the useful temporal information, which simultaneously relieves the blurry frame and reduces the computation burden. Finally, we make the UVOS model be a fully Transformer architecture, meanwhile aggregating the information from space, Gabor and time domains, yielding a strong representation with rich structure details. Extensive experiments on five mainstream UVOS benchmarks (DAVIS2016, FBMS, DAVSOD, ViSal, and MCL) demonstrate the superiority of the presented solution to sate-of-the-art methods. Jiaqing Fan, Tiankang Su, Kaihua Zhang 0001, Bo Liu 0005, Qingshan Liu 0001 |
ACM Multimedia | 5 |
| 2023 | Intellectual property protection for deep semantic segmentation models
Hongjia Ruan, Huihui Song 0003, Bo Liu 0005, Yong Cheng 0002, Qingshan Liu 0001 |
Frontiers Comput. Sci. | 5 |
| 2023 | Human pose transfer via shape-aware partial flow prediction network
Furong Ma, Guiyu Xia, Qingshan Liu 0001 |
Multim. Syst. | 3 |
| 2023 | Adaptive Multi-View and Temporal Fusing Transformer for 3D Human Pose EstimationabstractThis article proposes a unified framework dubbed Multi-view and Temporal Fusing Transformer (MTF-Transformer) to adaptively handle varying view numbers and video length without camera calibration in 3D Human Pose Estimation (HPE). It consists of Feature Extractor, Multi-view Fusing Transformer (MFT), and Temporal Fusing Transformer (TFT). Feature Extractor estimates 2D pose from each image and fuses the prediction according to the confidence. It provides pose-focused feature embedding and makes subsequent modules computationally lightweight. MFT fuses the features of a varying number of views with a novel Relative-Attention block. It adaptively measures the implicit relative relationship between each pair of views and reconstructs more informative features. TFT aggregates the features of the whole sequence and predicts 3D pose via a transformer. It adaptively deals with the video of arbitrary length and fully unitizes the temporal information. The migration of transformers enables our model to learn spatial geometry better and preserve robustness for varying application scenarios. We report quantitative and qualitative results on the Human3.6M, TotalCapture, and KTH Multiview Football II. Compared with state-of-the-art methods with camera parameters, MTF-Transformer obtains competitive results and generalizes well to dynamic capture with an arbitrary number of unseen views. Hui Shuai, Lele Wu, Qingshan Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Optimization-Based Post-Training Quantization With Bit-Split and StitchingabstractDeep neural networks have shown great promise in various domains. Meanwhile, problems including the storage and computing overheads arise along with these breakthroughs. To solve these problems, network quantization has received increasing attention due to its high efficiency and hardware-friendly property. Nonetheless, most existing quantization approaches rely on the full training dataset and the time-consuming fine-tuning process to retain accuracy. Post-training quantization does not have these problems, however, it has mainly been shown effective for 8-bit quantization. In this paper, we theoretically analyze the effect of network quantization and show that the quantization loss in the final output layer is bounded by the layer-wise activation reconstruction error. Based on this analysis, we propose an Optimization-based Post-training Quantization framework and a novel Bit-split optimization approach to achieve minimal accuracy degradation. The proposed framework is validated on a variety of computer vision tasks, including image classification, object detection, instance segmentation, with various network architectures. Specifically, we achieve near-original model performance even when quantizing FP32 models to 3-bit without fine-tuning. Peisong Wang 0001, Weihan Chen, Qiang Chen 0007, Qingshan Liu 0001, Jian Cheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Bi-RRNet: Bi-level recurrent refinement network for camouflaged object detection
Yan Liu 0004, Kaihua Zhang 0001, Yaqian Zhao, Qingshan Liu 0001 |
Pattern Recognit. | 5 |
| 2023 | Bias-Based Soft Label Learning for Facial Expression RecognitionabstractFacial Expression Recognition (FER) suffers from misrecognition due to the similarities between expressions. To address this issue, popular works replace original annotations with soft labels to reflect expression similarities. However, existing soft label learning (SLL) modules are independent of FER modules. In this article, inspired by automatic control theory, we propose a bias-based soft label learning network for FER named EC-Net. For optimizing FER and SLL modules jointly, EC-Net constitutes the closed-loop feedback between the two modules by designing a module measuring and transmitting the bias between FER module predictions and target labels. Specifically, EC-Net contains three modules: E-subNet, C-subNet, and L-Transmitter. First, E-subNet, i.e., the FER module, attempts to converge to target labels under the supervision of soft labels, acting as the executor. Then, L-Transmitter measures the bias between E-subNet predictions and target labels. It converts multiple discrete biases to the bias-based label through spectral clustering and transmits it to C-subNet. Finally, C-SubNet, i.e., the SLL module, generates soft labels from the bias-based label with a cascaded learner and progressively distinguishes similar expressions. It updates the learned soft labels for E-subNet, performing like the controller. Supervised by the bias-based soft label, E-subNet effectively reduces the dominant bias caused by similar expressions. We conduct extensive experiments on four popular benchmarks, demonstrating the effectiveness of applying closed-loop feedback in the FER task. Shanmin Wang, Hui Shuai, Chengguang Liu, Qingshan Liu 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | MSNet: Multi-Resolution Synergistic Networks for Adaptive InferenceabstractAdaptive inference with multiple networks has attracted much attention for resource-limited image classification. It assumes that a large portion of test samples can be correctly classified by small networks with fewer layers or channels, which poses a great challenge for them. In this paper, we argue that large networks have abilities to help the small ones address this challenge if fully explored. To this end, we propose a multi-resolution synergistic network (MSNet) using two different kinds of fusion modules. The first one is a cross-branch aggregation module, which aims to transfer the high-resolution features to the low-resolution ones between neighboring branches. The other one is an adaptive distillation module, whose purpose is feeding the discriminative ability of the large network to the other ones. Via these two modules, the small networks will be powerful enough to correctly classify large numbers of test samples, thus improving the classification accuracy and inference efficiency. We evaluate MSNet on three benchmark datasets: CIFAR-10, CIFAR-100, and ImageNet. Experimental results show that our network can obtain better results than several state-of-the-art networks in both anytime classification and budgeted batch classification settings. The code is available athttps://github.com/bigdata-qian/MSNet-Pytorch. Renlong Hang, Xuwei Qian, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | 3D Information Guided Motion Transfer via Sequential Image Based Human Model Refinement and Face-Attention GANabstractImage and video based human motions can be regarded as the deformation processes of person appearances, so motion transfer is usually treated as a pose guided image generation task and implemented in the 2D image plane. However, the 2D plane image generation lacks guidance of the original 3D motion information, which results in blur and shape distortions of the generated motion images. Therefore, we propose to simulate the generation process of real motion images by projecting the 3D human models, which are reconstructed from the training motion images and driven with target poses, into the 2D plane. We then take the 2D projections as the pose representations and input them into the generation model as they naturally inherit the 3D information from the original motions. Considering the unreliability on the invisible surface of the single image based human model reconstruction, we propose a sequential image based human model refinement module which exploits the complementary information between adjacent motion frames to refine the 3D human model. Furthermore, we propose a face-attention GAN model to conduct the final motion transfer, in which we use the Gaussian distribution to match the elliptical face region and design a face enhancement loss function since the faces in the generated motion images influence the performances very much. The generated motion images with reliable depth information, accurate shapes and clear faces demonstrate the effectiveness of the proposed method. Guiyu Xia, Yubao Sun, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Pose-Driven Realistic 2-D Motion SynthesisabstractA realistic 2-D motion can be treated as a deforming process of an individual appearance texture driven by a sequence of human poses. In this article, we thereby propose to transform the 2-D motion synthesis into a pose conditioned realistic motion image generation task considering the promising performance of pose estimation technology and generative adversarial nets (GANs). However, the problem is that GAN is only suitable to do the region-aligned image translation task while motion synthesis involves a large number of spatial deformations. To avoid this drawback, we design a two-step and multistream network architecture. First, we train a special GAN to generate the body segment images with given poses in step-I. Then in step-II, we input the body segment images as well as the poses into the multistream network so that it only needs to generate the textures in each aligned body region. Besides, we provide a real face as another input of the network to improve the face details of the generated motion image. The synthesized results with realism and sharp details on four training sets demonstrate the effectiveness of the proposed model. Guiyu Xia, Furong Ma, Qingshan Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2023 | Category-Level Assignment for Cross-Domain Semantic Segmentation in Remote Sensing ImagesabstractDeep learning-based semantic segmentation has made great progress in understanding very-high-resolution (VHR) remote sensing images (RSIs). However, large-scale applications are still limited. The main reason is that diverse imaging modes and geographical differences make it difficult to transfer a model trained in the source domain to the target domain. To solve this problem, unsupervised domain adaptation (UDA) for VHR RSIs has received some attention, but the accuracy of cross-domain semantic segmentation still needs to be improved. Currently, one reasonable proposal for improving accuracy is to take a close look at the category-level information. In this paper, we reveal an integer programming mechanism for modeling the category-level relationship between the source and target domains. The mechanism is based on the solution of the assignment problem, and thus, the proposed method is called category-level assignment for UDA (ClA-UDA). In ClA-UDA, a category-level assignment problem with additional constraints is defined for UDA tasks, and the solution is provided. Based on the solution, an assignment-based image-to-image transferring algorithm (AIT) is first proposed to transfer the source-domain images based on the style of the target-domain images. AIT minimizes a weighted discrepancy, and provides an analytical solution for the transfer. Two assignment-based alignment losses are then introduced to align the source and target domains based on the category-level relationship in a concise way. To validate the performance of ClA-UDA, three VHR remote sensing image datasets are employed, and six UDA tasks are designed. Extensive experiments are conducted, and the results demonstrate the superiority of ClA-UDA compared to the existing methods. Huan Ni, Qingshan Liu 0001, Haiyan Guan, Hong Tang 0002, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Geometry-Injected Image-Based Point Cloud Semantic SegmentationabstractImage-based methods have replicated the success from 2D domain to 3D point cloud semantic segmentation. However, when we directly apply 2D techniques to the projected pseudo-image, inherent differences between the point cloud and the image cause geometric distortion. This paper analyzes the geometric distortion between the point cloud and the pseudo-image, including truncation, dislocation, and hole. To ensure geometric fidelity, we propose the Geometry-injected Image-based point cloud semantic segmentation Network (GINet). We design a Cyclic Convolution to optimize the convolution operation, dealing with truncation. For dislocation and hole, we propose Dual Geometric Constraints, including Local Spatial Attention and Local Affinity Regularization, to incorporate the geometric information into semantic feature learning. Local Spatial Attention generates an attention map from the point coordinates to modulate the feature map before convolution. Local Affinity Regularization supervises the semantic similarity of pixels in the convolution kernel range. GINet rectifies the geometric distortion with these mechanisms while taking advantage of the successful 2D semantic segmentation methods. Quantitative and qualitative experiments on SemanticKITTI and SemanticPOSS demonstrate the effectiveness of GINet. Hui Shuai, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Mining Joint Intraimage and Interimage Context for Remote Sensing Change DetectionabstractRecent deep learning methods for change detection focus on excavating more discriminative context within individual images. However, due to seasonal change, noise, and so on, the appearance of objects tends to be more heterogeneous among various scenes. Consequently, the above intra-image context is inadequate to represent specific-category objects and pseudo changes would be inevitable in detection results. To deal with this issue, we propose a context aggregation network (CANet) to mine inter-image context over all training images for further enhancing intra-image context. Specifically, a Siamese network attached with temporal attention modules is served as a feature encoder to extract multi-scale temporal features from bitemporal images. Then, a context extraction module is devised to capture long-range spatial-channel context within individual images. Meanwhile, context representations of underlying categories in the scene are inferred using all training images in an unsupervised manner. Finally, these two kinds of contextual information are aggregated to one which is subsequently fed into a multi-scale fusion module to produce the detection map. CANet is compared with several state-of-the-art methods on three benchmark datasets, including the season-varying change detection (SVCD) dataset, the Sun Yat-sen University change detection (SYSU-CD) dataset, and the Learning Vision and Remote Sensing Laboratory building change detection (LEVIR-CD) dataset. It is demonstrated that our method outperforms all comparison methods in terms of F1, overall accuracy (OA), and Intersection-of-Union (IoU). The results of CANet on three datasets are available at https://github.com/NuistZF/CANet-for-change-detection and codes will be public soon. Feng Zhou 0006, Renlong Hang, Rui Zhang 0049, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Multilevel Spatial-Temporal Excited Graph Network for Skeleton-Based Action RecognitionabstractThe ability to capture joint connections in complicated motion is essential for skeleton-based action recognition. However, earlier approaches may not be able to fully explore this connection in either the spatial or temporal dimension due to fixed or single-level topological structures and insufficient temporal modeling. In this paper, we propose a novel multilevel spatial-temporal excited graph network (ML-STGNet) to address the above problems. In the spatial configuration, we decouple the learning of the human skeleton into general and individual graphs by designing a multilevel graph convolution (ML-GCN) network and a spatial data-driven excitation (SDE) module, respectively. ML-GCN leverages joint-level, part-level, and body-level graphs to comprehensively model the hierarchical relations of a human body. Based on this, SDE is further introduced to handle the diverse joint relations of different samples in a data-dependent way. This decoupling approach not only increases the flexibility of the model for graph construction but also enables the generality to adapt to various data samples. In the temporal configuration, we apply the concept of temporal difference to the human skeleton and design an efficient temporal motion excitation (TME) module to highlight the motion-sensitive features. Furthermore, a simplified multiscale temporal convolution (MS-TCN) network is introduced to enrich the expression ability of temporal features. Extensive experiments on the four popular datasets NTU-RGB+D, NTU-RGB+D 120, Kinetics Skeleton 400, and Toyota Smarthome demonstrate that ML-STGNet gains considerable improvements over the existing state of the art. Yisheng Zhu, Hui Shuai, Guangcan Liu, Qingshan Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Deep Object Co-Segmentation and Co-Saliency Detection via High-Order Spatial-Semantic Network ModulationabstractObject co-segmentation (CSG) is to segment the common objects of the same category in multiple relevant images while the co-saliency detection (CSD) aims to discover the salient and common foreground objects in a group of images. To process both tasks simultaneously, this paper presents an adaptive spatially and high-order semantically modulated deep network framework. A backbone network is first adopted to extract multi-resolution image features. With the multi-resolution features of the relevant images as input, we design an adaptive spatial modulator to learn a spatial representation that can highlight the co-object regions for each image. The adaptive spatial modulator fully captures the rich correlations of all image feature descriptors via unsupervised clustering and a graph aggregation strategy. The learned representation can well localize the common foreground object while effectively suppressing the background signals. For the high-order semantic modulator, we model it as a supervised image classification task. We propose a hierarchical high-order pooling module to learn the rich semantic features for classification use. The outputs of the two modulators manipulate the multi-resolution features by a shift-and-scale operation so that the features focus on segmenting common object regions. The proposed model is trained end-to-end without any intricate post-processing. Extensive experiments on three CSG benchmark datasets (MSRC, i-Coseg, and PASCAL-VOC) and three CSD datasets (Cosal2015, CoCA, and CoSOD3k) demonstrate the superior accuracy of the proposed method compared to state-of-the-art methods on both tasks. Kaihua Zhang 0001, Mingliang Dong, Bo Liu 0005, Dong Liu 0002, Qingshan Liu 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | Looking at Boundary: Siamese Densely Cooperative Fusion for Salient Object DetectionabstractThough deep learning-based saliency detection methods have achieved gratifying performance recently, the predicted saliency maps still suffer from the boundary challenge. From the perspective of foreground-background separation, this article attempts to extract the edge information of objects by exploiting the difference between different color channels in the RGB color space and establishes a novel multicolor contrast extraction (MCE) mechanism to improve the learning ability of exquisite boundary information of the network. To make full use of the MCE outputs and RGB colors, and well depict and capture the complementary information between them, we devise a novel Siamese densely cooperative fusion (DCF) network (SDFNet) for saliency detection, which consists of two effective components: boundary-directed feature learning (BDFL) and DCF. The BDFL provides joint learning for both MCE and RGB modalities through a Siamese network, while the DCF module is devised for complementary feature discovery, in order to effectively combine the features learned from two modalities. Experiments on five well-known benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches in terms of different evaluation metrics. We provide a detailed analysis of these results and indicate that our joint modeling of MCE and RGB colors helps to better capture the object details, especially in the object boundaries. Junxia Li, Qingshan Liu 0001, Dongyan Guo |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | ReX: An Efficient Approach to Reducing Memory Cost in Image ClassificationabstractExiting simple samples in adaptive multi-exit networks through early modules is an effective way to achieve high computational efficiency. One can observe that deployments of multi-exit architectures on resource-constrained devices are easily limited by high memory footprint of early modules. In this paper, we propose a novel approach named recurrent aggregation operator (ReX), which uses recurrent neural networks (RNNs) to effectively aggregate intra-patch features within a large receptive field to get delicate local representations, while bypassing large early activations. The resulting model, named ReXNet, can be easily extended to dynamic inference by introducing a novel consistency-based early exit criteria, which is based on the consistency of classification decisions over several modules, rather than the entropy of the prediction distribution. Extensive experiments on two benchmark datasets, i.e., Visual Wake Words, ImageNet-1k, demonstrate that our method consistently reduces the peak RAM and average latency of a wide variety of adaptive models on low-power devices. Xuwei Qian, Renlong Hang, Qingshan Liu 0001 |
AAAI | 3 |
| 2022 | Transformer-based Residual Network for Hyperspectral Snapshot Compressive ReconstructionabstractThe core problem of hyperspectral snapshot compressive imaging is to achieve high-quality reconstruction from a single snapshot measurement. Existing reconstruction methods are mainly based on convolutional networks. Inspired by the exciting success of Transformer in high-level vision tasks, we introduce Transformer into hyperspectral snapshot compressive imaging, and propose the Transformer-based residual network to learn reconstruction mapping. Specifically, the proposed network cascades multiple basic modules, each of which consists of two local-enhanced window (Lewin) Transformer blocks and two residual blocks. The advantage of this basic module is ability to exploit both local feature maps and long-range dependencies for resonstruction. The Transformer blocks treat the feature map at each pixel location as a token and capture the spatial-spectral context of each pixel within a local window through self-attention, thus effectively improving the spectral fidelity of reconstructed hyperspectral images at modest computational cost. We conduct extensive experiments on both simulation and real data. The experimental results show that the proposed network is concise and effective, achieving the best reconstruction results compared with several state-of-the-art methods. Junru Huang, Yubao Sun, Jiaxuan Wen, Qingshan Liu 0001 |
ICPR | 4 |
| 2022 | Bidirectionally Learning Dense Spatio-temporal Feature Propagation Network for Unsupervised Video Object SegmentationabstractSpatio-temporal feature representation is essential for accurate unsupervised video object segmentation, which needs an effective feature propagation paradigm for both appearance and motion features that can fully interchange information across frames. However, existing solutions mainly focus on the forward feature propagation from the preceding frame to the current one, either using the former segmentation mask or motion propagation in a frame-by-frame manner. This ignores the bi-directional temporal feature interactions (including the backward propagation from the future to the current frame) across all frames that can help to enhance the spatiotemporal feature representation for segmentation prediction. To this end, this paper presents a novel Dense Bidirectional Spatio-temporal feature propagation Network (DBSNet) to fully integrate the forward and the backward propagations across all frames. Specifically, a dense bi-ConvLSTM module is first developed to propagate the features across all frames in a forward and backward manner. This can fully capture the multi-level spatio-temporal contextual information across all frames, producing an effective feature representation that has a strong discriminative capability to tell from noisy backgrounds. Following it, a spatio-temporal Transformer refinement module is designed to further enhance the propagated features, which can effectively capture the spatio-temporal long-range dependencies among all frames. Afterwards, a Co-operative Direction-aware Graph Attention (Co-DGA) module is designed to integrate the propagated appearancemotion cues, yielding a strong spatio-temporal feature representation for segmentation mask prediction. The Co-DGA assigns proper attentional weights to neighboring points along the coordinate axis, making the segmentation model to selectively focus on the most relevant neighbors. Extensive evaluations on four mainstream challenging benchmarks including DAVIS16, FBMS, DAVSOD, and MCL demonstrate that the proposed DBSNet achieves favorable performance against state-of-the-art methods in terms of all evaluation metrics. Jiaqing Fan, Tiankang Su, Kaihua Zhang 0001, Qingshan Liu 0001 |
ACM Multimedia | 4 |
| 2022 | Waterfall-Net: Waterfall Feature Aggregation for Point Cloud Semantic Segmentation
Hui Shuai, Xiang Xu 0009, Qingshan Liu 0001 |
PRCV (3) | 3 |
| 2022 | Consistent connectome landscape mining for cross-site brain disease identification using functional MRI
Daoqiang Zhang, Jiashuang Huang, Mingxia Liu 0001, Qingshan Liu 0001 |
Medical Image Anal. | 5 |
| 2022 | CED-Net: contextual encoder-decoder network for 3D face reconstruction
Shanmin Wang, Zengqun Zhao, Xiang Xu 0009, Qingshan Liu 0001 |
Multim. Syst. | 5 |
| 2022 | Learning interlaced sparse Sinkhorn matching network for video super-resolution
Huihui Song 0003, Yutong Jin, Yong Cheng 0002, Bo Liu 0005, Dong Liu 0002, Qingshan Liu 0001 |
Pattern Recognit. | 6 |
| 2022 | Phase Space Reconstruction Driven Spatio-Temporal Feature Learning for Dynamic Facial Expression RecognitionabstractAutomatic Dynamic Facial Expression Recognition (DFER) is a challenging task, since how to effectively capture facial temporal dynamics is still an open problem. In this article, we regard variations of facial expressions as a dynamic system in accord with certain rules, and try to explore the fundamental temporal properties for recognizing dynamic expressions. Inspired by the phase space reconstruction method for time series analysis, we propose a novel network named Phase Space Reconstruction Network (PSRNet) for learning spatio-temporal features of facial expressions. First, 3D convolutional neural networks are used to extract spatial and short-term temporal features, which indicate the state of each frame and are termed as observations in the phase space. All the observations compose the trajectory of the dynamical system. Then, a data-driven across-correlation matrix is inferred to reveal the relationship of the observations. With this matrix, the phase space reconstruction module reconstructs the trajectory by aggregating the observations adaptively in the phase space. Reconstructed observations represent the gradual process of dynamic facial expressions, which is beneficial to recognize these expressions. The experiment results on three databases (Oulu, MMI, and CK+) demonstrate that the proposed PSRNet can extract more informative and representative spatio-temporal features for DFER. Moreover, the visualization of intermediate features reveals that the reconstructed features have global consistency in facial regions and the underlying evolutionary pattern of dynamic facial expression. Shanmin Wang, Hui Shuai, Qingshan Liu 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Semi-Supervised Video Object Segmentation via Learning Object-Aware Global-Local CorrespondenceabstractIn semi-supervised video object segmentation (VOS) task, temporal coherent object-level cues play a key role yet are hard to accurately model. To this end, this paper presents an object-aware global-local correspondence architecture, which enables to extract the inter-frame temporal coherent object-level features for accurate VOS. Specifically, we first generate a set of object masks by the ground-truth segmentation, and then we squeeze the current frame representation inside the object masks into a set of global object embeddings. Second, we compute the similarity between each embedding and the feature map, producing an object-aware weight for each pixel. The object-aware feature at each pixel is then constructed by summing the object embeddings weighted by their corresponding object-aware weights, which is able to capture rich object category information. Third, to establish the accurate correspondences between the inter-frame temporal coherent cues, we further design a novel global-local correspondence module to refine the temporal feature representations. Finally, we augment the object-aware features with the global-local aligned information to produce a strong spatio-temporal representation, which is essential to a more reliable pixel-wise segmentation prediction. Extensive evaluations are conducted on three popular VOS benchmarks containing Youtube-VOS, Davis2017 and Davis2016, demonstrating that the proposed method achieves favourable performance compared to the state-of-the-arts. Jiaqing Fan, Bo Liu 0005, Kaihua Zhang 0001, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Spatial Consistency Constrained GAN for Human Motion TransferabstractIn this paper, we propose a new GAN-based framework to implement video-based human motion transfer,i.e., transferring the motions from the source person to the target one with the help of pose information. Human motion transfer involves large scaled spatial deformations from pose to body image and emphasizes spatial consistency of details. However, GAN is not suitable for the region-unaligned task due to the global adversarial loss does not focus on the spatial details. Therefore, we design a two-stage Spatial Consistency Constrained GAN architecture to generate realistic target person images. Within the model, we first generate a segment map to align the regions of different body parts with a given pose in stage-I and then concatenate the pose and the segment map as condition to generate a target person image in stage-II, so that the deformation problem is avoided. Furthermore, to improve the spatial detail consistency, we propose the shape consistency loss for the segment map generation to make the model pay more attention to the shape of each body part. We also propose a pose consistency loss for the target person image generation to enforce the generated images to contain similar enough poses to the input ones. The synthesized images with clear shape and sharp details demonstrate the effectiveness of the proposed method. Furong Ma, Guiyu Xia, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Video Snapshot Compressive Imaging Using Residual Ensemble NetworkabstractVideo snapshot compressive imaging (SCI) system enables high-frame-rate imaging by projecting multiple frames into a 2D snapshot measurement during a single exposure, and the original video frames can be reconstructed by solving an optimization problem. However, existing methods usually cannot achieve a good balance between reconstruction time and reconstruction quality, which has become a major obstacle for practical application of video SCI. In order to cope with this issue, we propose a residual ensemble network to learn the explicit inverse mapping from the 2D snapshot measurement to the original video. Specifically, the proposed network aims to exploit the spatiotemporal correlations between video frames for improving reconstruction quality. The spatiotemporal correlations of video frames demonstrate multiple types, including intra-frame spatial correlation, inter-frame forward and backward temporal correlation. With the purpose of fully capturing these differentiated correlations, we design four sub-networks, namely, a pseudo-3D U-shape sub-network, two residual sub-networks, and a serial forward and backward recurrent sub-network, and further assemble these four sub-networks into an ensemble network through alternate residual links. This ensemble network can effectively fuse the predictions of each sub-network and maintain spatiotemporal consistency between video frames. We further design a compound loss function to guide the network learning, and the new video can be fast reconstructed by simply feeding its 2D snapshot measurement into the learned network. The experimental results demonstrate that our network can significantly improve the reconstruction quality while maintaining low computational cost. Yubao Sun, Xunhao Chen, Mohan Kankanhalli, Qingshan Liu 0001, Junxia Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Keyframe-Editable Real-Time Motion SynthesisabstractSince existing motion synthesis methods often lack precise controls to the synthesis process, we propose a keyframe-editable motion synthesis framework which allows users to edit the keyframes of an expected motion sequence and use the edited keyframes to control and drive the synthesis process. Specifically, a motion segment can be represented as a start-end frame pair and a group of motion patterns which record the joint angle changing processes between the start and end frames. With the pre-trained paired dictionaries relating the start-end frame space and motion pattern space, we can use the edited keyframe pair to generate its corresponding motion pattern, so the validity of keyframes is very important to the quality of synthesized motions. Thus, we use the probability to measure the motion naturalness and propose a naturalness rectification method to guarantee the validity of the edited keyframes. We also provide a joint-move interface and propose a position refinement method for the detailed adjustment of the joint positions of keyframes. Besides, we add extra naturalness constraint to the motion synthesis process to further improve the naturalness of the generated in- between frames. Extensive rectification experiments in different situations verify the effect of the proposed naturalness rectification model. The comparisons of the synthesis results with other state-of-the-art synthesis methods demonstrate the advantage of our motion synthesis model. The high efficiencies of all the algorithms make the proposed framework competent to the real-time motion synthesis. Guiyu Xia, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Self-Supervised Video Representation Learning Using Improved Instance-Wise Contrastive Learning and Deep ClusteringabstractInstance-wise contrastive learning (Instance-CL), which learns to map similar instances closer and different instances farther apart in the embedding space, has achieved considerable progress in self-supervised video representation learning. However, canonical Instance-CL does not handle properly the temporal similarities between different videos, limiting the representation capabilities of learned models. This paper presents a novel two-stage framework that combines Instance-CL and unsupervised clustering to progressively learn desirable temporal representations with high intra-class compactness. Specifically, (a) we first introduce a new consistency-preserving sampling strategy to generate positive/negative pairs. Compared to the traditional sampling methods, our sampling strategy focuses more on motion dynamics, resulting in more temporal-related feature representations. (b) To further explore the temporal similarities between videos so as to encourage intra-class compactness, we set temporal representations extracted from Instance-CL as an initializer, and iteratively use k-means clustering to generate pseudo-labels for training the encoder. We term our method as Improved Instance-CL with Deep Clustering (ICDC) and apply it to two downstream tasks, including action recognition and video retrieval. Extensive experimental results show that ICDC gains considerable improvements compared to the existing self-supervised methods. Yisheng Zhu, Hui Shuai, Guangcan Liu, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Complementarity-Aware Attention Network for Salient Object DetectionabstractIn this article, we tackle the saliency detection task from an interesting perspective: we focus both on salient regions (or foreground) detection and nonsalient regions (or background) detection instead of only the foreground and propose a novel complementarity-aware attention network. It is a unified framework with two branches, namely, positive attention module (PAM) and negative attention module (NAM), for the foreground and background detection, respectively. More specifically, the PAM exploits a position self-attention mechanism to enhance the discriminant ability of feature representation, which can detect most of the salient object regions. Meanwhile, the NAM is designed to detect the background regions, aiming to pop out the missing object parts and details in the prediction map produced by the PAM. By fusing these two attention modules together, NAM can provide complementary cues to assist PAM for precise object detection. Furthermore, in order to capture more multiscale contextual information, we introduce a bidirectional structure with multisupervision to the proposed complementarity-aware attention module for performance improvement. Experiments on five benchmark datasets show that the proposed framework achieves comparable results compared with the state-of-the-art saliency detection methods. Junxia Li, Qingshan Liu 0001, Yubao Sun |
IEEE Trans. Cybern. | 3 |
| 2022 | Hierarchical Context Network for Airborne Image SegmentationabstractMost of the recent methods focus on capturing contextual information by measuring relations (e.g., feature similarity) between each pixel and all the others for airborne image segmentation. Nevertheless, these methods have difficulty in handling confusing objects with a partially similar appearance. In this article, we attempt to simultaneously explore pixel-to-pixel (P2P) and pixel-to-object (P2O) relations to learn contextual information. For this purpose, a hierarchical context network (HCNet) is proposed. It consists of a P2P subnetwork and a P2O subnetwork. The P2P subnetwork learns the P2P relation (detail-grained context) for better preservation of the details (e.g., boundary) of the objects. Meanwhile, the P2O subnetwork models the P2O relation (semantic-grained context), aiming at improving the intraobject semantic consistency. When inferring the segmentation results, outputs of these two subnetworks are aggregated to obtain the hierarchical contextual information. Experimental results demonstrate that the proposed model achieves competitive performance on three challenging benchmarks. Feng Zhou 0006, Renlong Hang, Hui Shuai, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Cross-Modality Contrastive Learning for Hyperspectral Image ClassificationabstractDeep learning has attracted much attention in the field of hyperspectral image classification recently, due to its powerful representation and generalization abilities. Most of current deep learning models are trained in a supervised manner, which require large amounts of labeled samples to achieve state-of-the-art performance. Unfortunately, pixel-level labeling in hyperspectral imageries is difficult, time-consuming, and human-dependent. To address this issue, we propose an unsupervised feature learning model using multi-modal data, hyperspectral and LiDAR in particular. It takes advantage of the relationship between hyperspectral and LiDAR data to extract features, without using any label information. After that, we design a dual fine-tuning strategy to transfer the extracted features for hyperspectral image classification with small numbers of training samples. Such strategy is able to explore not only the semantic information but also the intrinsic structure information of training samples. In order to test the performance of our proposed model, we conduct comprehensive experiments on three hyperspectral and LiDAR datasets. Experimental results show that our proposed model can achieve better performance than several state-of-the-art deep learning models. Renlong Hang, Xuwei Qian, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Multiscale Progressive Segmentation Network for High-Resolution Remote Sensing ImageryabstractSemantic segmentation of high-resolution remote sensing imageries (HRSIs) is a critical task for a wide range of applications, such as precision agriculture and urban planning. Although convolutional neural networks (CNNs) have made great progress in accomplishing this task recently, there still exist some challenges to address, one of which is simultaneously segmenting objects with large scale variations in a HRSI. Targeting at this challenge, previous CNNs often adopt multiple convolution kernels in one layer or skip-layer connections between different layers to extract multiscale representations. However, due to the limited learning capacity of each CNN, it tends to make trade-offs in segmenting different-scale objects. This would lead to unsatisfactory segmentation results for some objects, especially the small or the large ones. In this paper, we propose a multiscale progressive segmentation network to address this issue. Instead of forcing one network to deal with all scales of objects, our network attempts to cascade three subnetworks for gradually segmenting objects with small scales, large scales, and other scales. In order to make the subnetwork focus on the specific scale objects, a scale guidance module is designed. It takes advantage of segmentation results from the preceding subnetwork to guide the feature learning of the succeeding one. Additionally, to acquire the final segmentation results, we propose a position sensitive module for adaptively combining the outputs of the three subnetworks. This module is capable of assigning combination weights of different subnetworks according to their importance. Experiments on two benchmark datasets named Vaihingen and Potsdam indicate that our proposed network can achieve considerable improvements in comparison with several state-of-the-art segmentation models. Renlong Hang, Feng Zhou 0006, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Unsupervised Spatial-Spectral Network Learning for Hyperspectral Compressive Snapshot ReconstructionabstractHyperspectral compressive imaging takes advantage of compressive sensing theory to achieve coded aperture snapshot measurement without temporal scanning, and the entire 3-D spatial–spectral data is captured by a 2-D projection during a single integration period. Its core issue is how to reconstruct the underlying hyperspectral image (HSI) using compressive sensing reconstruction algorithms. Due to the diversity in the spectral response characteristics and wavelength range of different spectral imaging devices, previous works are often inadequate to capture complex spectral variations or lack the adaptive capacity to new hyperspectral imagers. In order to address these issues, we propose an unsupervised spatial–spectral network to reconstruct HSIs only from the compressive snapshot measurement. The proposed network acts as a conditional generative model conditioned on the snapshot measurement, and it exploits the spatial–spectral attention module to capture the joint spatial–spectral correlation of HSIs. The network parameters are optimized to make sure that the network output can closely match the given snapshot measurement according to the imaging model, thus the proposed network can adapt to different imaging settings, which can inherently enhance the applicability of the network. Extensive experiments upon multiple datasets demonstrate that our network can achieve better reconstruction results than the state-of-the-art methods. Yubao Sun, Qingshan Liu 0001, Mohan Kankanhalli |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Global Tropical Cyclone Precipitation Estimation via a Multitask Convolutional Neural Network Based on HURSAT-B1 DataabstractFast and accurate global tropical cyclone (TC) precipitation estimation from satellite observations is still a challenging issue. In this article, we propose an effective model based on a multitask convolutional neural network (CNN) to estimate near-real-time global TC precipitation from HURSAT-B1 data. Our network mainly consists of three modules: the feature extraction module, the wind grade classification module, and the precipitation estimation module. The first module aims at extracting the spatial features of satellite imageries, the second module focuses on classifying the wind grades of the satellite imageries into six categories that are used to assist in estimating TC precipitation, and the third module is to estimate TC precipitation. To evaluate the effectiveness of our proposed model, we compare it with multiple linear regression (MLR) and random forest (RF) models based on integrated multisatellite retrievals for the global precipitation measurement (GPM) mission (IMERG). Besides, four typical TC events are selected to specifically analyze the temporal and spatial distribution of TC precipitation estimation. Experimental results show that the probability of detection and accuracy achieved by our proposed model are 0.68 and 0.81, while the correlation coefficient (CC) and MSE are 0.61 and 7.80, respectively. In terms of the four TC events, our proposed model obtains a more consistent and continuous spatial distribution of precipitation than MLR and RF. More importantly, our proposed model can achieve high spatiotemporal results, which has the potential to serve as an operational algorithm for global TC precipitation estimation. Mei Xue, Renlong Hang, Xiao-Tong Yuan, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Predicting Tropical Cyclogenesis Using a Deep Learning Method From Gridded Satellite and ERA5 Reanalysis Data in the Western North Pacific BasinabstractThis article proposes a deep learning model to predict tropical cyclogenesis (TCG) from gridded satellite and ERA5 reanalysis data in the western North Pacific basin. The proposed model contains two modules. First, convolutional neural network (CNN)-based deep features are extracted for each predictor, and then, the extracted features are fused with two fully connected layers to differentiate and investigate the relationship between predictors and TCG. The experimental data of this study are composed of 3232 developing tropical cluster clouds and 6657 nondeveloping ones; 90% of the collected data are utilized to train the model, and the rest are used to evaluate the trained model. Totally, nine predictors have been considered for the study, and the results show that the brightness temperature (IR), relative vorticity (Vo), and geopotential height (Z) perform better than the other predictors. A combined model with six predictors [IR, Z, RH (relative humidity), Vo, WS10 m(wind speed at the height of ten meters above the surface of the Earth), and mslp (mean sea-level pressure)] achieves the best TCG predicting performance, i.e., 97.1% of developing tropical cyclones are detected at a probability threshold of 0.13 with a false alarm rate of 20.3%. The experimental results demonstrate that the proposed method is superior to the existing methods and also indicate that the fusion of satellite and reanalysis data is a promising method to predict TCG. Rui Zhang 0049, Qingshan Liu 0001, Renlong Hang, Guangcan Liu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Limb Pose Aware Networks for Monocular 3D Pose EstimationabstractIn the task of monocular 3D pose estimation, the estimation errors of limb joints (i.e., wrist, ankle, etc) with a higher degree of freedom(DOF) are larger than that of others (i.e., hip, thorax, etc). Specifically, errors may accumulate along the physiological structure of human body parts, and trajectories of joints with higher DOF bring in higher complexity. To address this problem, we propose a limb pose aware framework, involving a kinematic constraint aware network as well as a trajectory aware temporal module, to improve the 3D prediction accuracy of limb joint positions. Two kinematic constraints named relative bone angles and absolute bone angles are introduced in this paper, the former being used for building the angular relation between adjacent bones and the latter for building the angular relation between bones and the camera plane. As a joint result of two constraints, our work suppresses errors accumulated along limbs. Furthermore, we propose a trajectory-aware network, named as Hierarchical Transformer, which takes temporal trajectories of joints as input and generates fused trajectory estimation as a result. The Hierarchical Transformer consists of Transformer Encoder blocks and aims at improving the performance of fusing temporal features. Under the effect of kinematic constraints and trajectory network, we alleviate the problem of errors accumulated along limbs and achieve promising results. Most of the off-the-shelf 2D pose estimators can be easily integrated into our framework. We perform extensive experiments on public datasets and validate the effectiveness of the framework. The ablation studies show the strength of each individual sub-module. Lele Wu, Zhenbo Yu, Yijiang Liu, Qingshan Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Local Self-Expression Subspace Learning Network for Motion Capture DataabstractDeep subspace learning is an important branch of self-supervised learning and has been a hot research topic in recent years, but current methods do not fully consider the individualities of temporal data and related tasks. In this paper, by transforming the individualities of motion capture data and segmentation task as the supervision, we propose the local self-expression subspace learning network. Specifically, considering the temporality of motion data, we use the temporal convolution module to extract temporal features. To implement the local validity of self-expression in temporal tasks, we design the local self-expression layer which only maintains the representation relations with temporally adjacent motion frames. To simulate the interpolatability of motion data in the feature space, we impose a group sparseness constraint on the local self-expression layer to impel the representations only using selected keyframes. Besides, based on the subspace assumption, we propose the subspace projection loss, which is induced from distances of each frame projected to the fitted subspaces, to penalize the potential clustering errors. The superior performances of the proposed model on the segmentation task of synthetic data and three tasks of real motion capture data demonstrate the feature learning ability of our model. Guiyu Xia, Huaijiang Sun, Yubao Sun, Qingshan Liu 0001 |
IEEE Trans. Image Process. | 6 |
| 2022 | Image Co-Saliency Detection and Instance Co-Segmentation Using Attention Graph Clustering Based Graph Convolutional NetworkabstractCo-Saliency Detection (CSD) is to explore the concurrent patterns and salient objects from a group of relevant images, while Instance Co-Segmentation (ICS) aims to identify and segment out all of these co-salient instances, generating corresponding mask for each instance. To simultaneously tackle these two tasks, we present a novel adaptive graph convolutional network with attention graph clustering (GCAGC) for CSD and ICS, termed as GCAGC-CSD and GCAGC-ICS, respectively. The GCAGC-CSD contains three key model designs: first, we develop a graph convolutional network architecture to extract multi-scale representations to characterize the intra- and inter-image consistency. Second, we propose an attention graph clustering algorithm to distinguish the salient foreground objects from common areas in an unsupervised manner. Third, we present a unified framework with encoder-decoder structure to jointly train and optimize the graph convolutional network, attention graph cluster, and CSD decoder in an end-to-end fashion. Afterwards, we design a salient instance segmentation network for GCAGC-ICS, and combine the outputs of GCAGC-CSD and the instance segmentation branch to obtain instance-aware co-segmentation masks. The proposed GCAGC-CSD and GCAGC-ICS are extensively evaluated on four CSD benchmark datasets (iCoseg, Cosal2015, COCO-SEG and CoSOD3k) and five ICS benchmark datasets (CoSOD3k, COCO-NONVOC, COCO-VOC, VOC12 and SOC), and achieve superior performance over state-of-the-arts on both tasks. Tengpeng Li, Kaihua Zhang 0001, Shiwen Shen, Bo Liu 0005, Qingshan Liu 0001, Zhu Li 0001 |
IEEE Trans. Multim. | 5 |
| 2021 | Robust Lightweight Facial Expression Recognition Network with Label Distribution TrainingabstractThis paper presents an efficiently robust facial expression recognition (FER) network, named EfficientFace, which holds much fewer parameters but more robust to the FER in the wild. Firstly, to improve the robustness of the lightweight network, a local-feature extractor and a channel-spatial modulator are designed, in which the depthwise convolution is employed. As a result, the network is aware of local and global-salient facial features. Then, considering the fact that most emotions occur as combinations, mixtures, or compounds of the basic emotions, we introduce a simple but efficient label distribution learning (LDL) method as a novel training strategy. Experiments conducted on realistic occlusion and pose variation datasets demonstrate that the proposed EfficientFace is robust under occlusion and pose variation conditions. Moreover, the proposed method achieves state-of-the-art results on RAF-DB, CAER-S, and AffectNet-7 datasets with accuracies of 88.36%, 85.87%, and 63.70%, respectively, and a comparable result on the AffectNet-8 dataset with an accuracy of 59.89%. The code is public available at https://github.com/zengqunzhao/EfficientFace. Zengqun Zhao, Qingshan Liu 0001, Feng Zhou 0006 |
AAAI | 2 |
| 2021 | DeepACG: Co-Saliency Detection via Semantic-Aware Contrast Gromov-Wasserstein DistanceabstractThe objective of co-saliency detection is to segment the co-occurring salient objects in a group of images. To address this task, we introduce a new deep network architecture via semantic-aware contrast Gromov-Wasserstein distance (DeepACG). We first adopt the Gromov-Wasserstein (GW) distance to build dense 4D correlation volumes for all pairs of image pixels within the image group. These dense correlation volumes enable the network to accurately discover the structured pair-wise pixel similarities among the common salient objects. Second, we develop a semantic-aware co-attention module (SCAM) to enhance the foreground co-saliency through predicted categorical information. Specifically, SCAM recognizes the semantic class of the foreground co-objects, and this information is then modulated to the deep representations to localize the related pixels. Third, we design a contrast edge-enhanced module (EEM) to capture richer contexts and preserve fine-grained spatial information. We validate the effectiveness of our model using three largest and most challenging benchmark datasets (Cosal2015, CoCA, and CoSOD3k). Extensive experiments have demonstrated the substantial practical merit of each module. Compared with the existing works, DeepACG shows significant improvements and achieves state-of-the-art performance. Kaihua Zhang 0001, Mingliang Dong, Bo Liu 0005, Xiao-Tong Yuan, Qingshan Liu 0001 |
CVPR | 5 |
| 2021 | Deep Transport Network for Unsupervised Video Object SegmentationabstractThe popular unsupervised video object segmentation methods fuse the RGB frame and optical flow via a two-stream network. However, they cannot handle the distracting noises in each input modality, which may vastly deteriorate the model performance. We propose to establish the correspondence between the input modalities while suppressing the distracting signals via optimal structural matching. Given a video frame, we extract the dense local features from the RGB image and optical flow, and treat them as two complex structured representations. The Wasserstein distance is then employed to compute the global optimal flows to transport the features in one modality to the other, where the magnitude of each flow measures the extent of the alignment between two local features. To plug the structural matching into a two-stream network for end-to-end training, we factorize the input cost matrix into small spatial blocks and design a differentiable long-short Sinkhorn module consisting of a long-distant Sinkhorn layer and a short-distant Sinkhorn layer. We integrate the module into a dedicated two-stream network and dub our model TransportNet. Our experiments show that aligning motion-appearance yields the state-of-the-art results on the popular video object segmentation datasets. Kaihua Zhang 0001, Zicheng Zhao, Dong Liu 0002, Qingshan Liu 0001, Bo Liu 0005 |
ICCV | 4 |
| 2021 | Former-DFER: Dynamic Facial Expression Recognition TransformerabstractThis paper proposes a dynamic facial expression recognition transformer (Former-DFER) for the in-the-wild scenario. Specifically, the proposed Former-DFER mainly consists of a convolutional spatial transformer (CS-Former) and a temporal transformer (T-Former). The CS-Former consists of five convolution blocks and N spatial encoders, which is designed to guide the network to learn occlusion and pose-robust facial features from the spatial perspective. And the temporal transformer consists of M temporal encoders, which is designed to allow the network to learn contextual facial features from the temporal perspective. The heatmaps of the leaned facial features demonstrate that the proposed Former-DFER is capable of handling the issues such as occlusion, non-frontal pose, and head motion. And the visualization of the feature distribution shows that the proposed method can learn more discriminative facial features. Moreover, our Former-DFER also achieves state-of-the-art results on the DFEW and AFEW benchmarks. Zengqun Zhao, Qingshan Liu 0001 |
ACM Multimedia | 2 |
| 2021 | Tensor LISTA: Differentiable sparse representation learning for multi-dimensional tensor
Qi Zhao 0013, Guangcan Liu, Qingshan Liu 0001 |
Neurocomputing | 3 |
| 2021 | Likelihood-constrained coupled space learning for motion synthesis
Guiyu Xia, Qingshan Liu 0001 |
Inf. Sci. | 4 |
| 2021 | Body parts relevance learning via expectation-maximization for human pose estimation
Luhui Yue, Junxia Li, Qingshan Liu 0001 |
Multim. Syst. | 3 |
| 2021 | Matrix Completion with Deterministic Sampling: Theories and MethodsabstractIn some significant applications such as data forecasting, the locations of missing entries cannot obey any non-degenerate distributions, questioning the validity of the prevalent assumption that the missing data is randomly chosen according to some probabilistic model. To break through the limits of random sampling, we explore in this paper the problem of real-valued matrix completion under the setup of deterministic sampling. We propose two conditions, isomeric condition and relative well-conditionedness, for guaranteeing an arbitrary matrix to be recoverable from a sampling of the matrix entries. It is provable that the proposed conditions are weaker than the assumption of uniform sampling and, most importantly, it is also provable that the isomeric condition is necessary for the completions of any partial matrices to be identifiable. Equipped with these new tools, we prove a collection of theorems for missing data recovery as well as convex/nonconvex matrix completion. Among other things, we study in detail a Schatten quasi-norm induced method termed isomeric dictionary pursuit (IsoDP), and we show that IsoDP exhibits some distinct behaviors absent in the traditional bilinear programs. Guangcan Liu, Qingshan Liu 0001, Xiao-Tong Yuan, Meng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Video saliency prediction using enhanced spatiotemporal alignment network
Huihui Song 0003, Kaihua Zhang 0001, Bo Liu 0005, Qingshan Liu 0001 |
Pattern Recognit. | 5 |
| 2021 | Feature Alignment and Aggregation Siamese Networks for Fast Visual TrackingabstractSiamese networks have been successfully introduced into visual tracking, which match the best candidate and a target template via a couple of networks with shared parameters. However, most Siamese network-based trackers (SNTs) are tailored to best match the canonical posture of the template and the search-region images, resulting in inferior performance when the target objects have large-scale pose variations. Besides, SNTs fail to discriminate distractors well because they only leverage high-level semantic features as target representations that cannot well tell from different targets of the same category. To address these issues, this paper presents an efficient and effective SNT that is based on feature alignment and aggregation networks. Specifically, we first design an effective feature alignment network module to calibrate the search-region image. This module results in a more reliable matching response that is robust to severe target pose variations. Then, we develop an effective shallow-level and high-level feature aggregation network module to complement the feature characteristics, making the learned feature representation not only well differentiate the target from distractors, but also robust to target appearance variations. Afterwards, we employ a channel-attention mechanism to further strengthen the discriminative capability of the aggregated feature representation. Finally, both the alignment and the aggregation modules are seamlessly integrated into the Siamese networks for robust tracking. Meanwhile, we offline learn the network parameters end-to-end without time-consuming fine-tuning. Extensive evaluations on a variety of benchmarks including VOT-2017, OTB-100, UAV123 and GOT-10k demonstrate favorable performance of our tracker against state-of-the-art ones with a speed of 60 fps. Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Hyperspectral Image Classification With Attention-Aided CNNsabstractConvolutional neural networks (CNNs) have been widely used for hyperspectral image classification. As a common process, small cubes are first cropped from the hyperspectral image and then fed into CNNs to extract spectral and spatial features. It is well known that different spectral bands and spatial positions in the cubes have different discriminative abilities. If fully explored, this prior information will help improve the learning capacity of CNNs. Along this direction, we propose an attention-aided CNN model for spectral-spatial classification of hyperspectral images. Specifically, a spectral attention subnetwork and a spatial attention subnetwork are proposed for spectral and spatial classifications, respectively. Both of them are based on the traditional CNN model and incorporate attention modules to aid networks that focus on more discriminative channels or positions. In the final classification phase, the spectral classification result and the spatial classification result are combined together via an adaptively weighted summation method. To evaluate the effectiveness of the proposed model, we conduct experiments on three standard hyperspectral data sets. The experimental results show that the proposed model can achieve superior performance compared with several state-of-the-art CNN-related models. Renlong Hang, Zhu Li 0001, Qingshan Liu 0001, Pedram Ghamisi, Shuvra S. Bhattacharyya |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Classification of Hyperspectral Images via Multitask Generative Adversarial NetworksabstractDeep learning has shown its huge potential in the field of hyperspectral image (HSI) classification. However, most of the deep learning models heavily depend on the quantity of available training samples. In this article, we propose a multitask generative adversarial network (MTGAN) to alleviate this issue by taking advantage of the rich information from unlabeled samples. Specifically, we design a generator network to simultaneously undertake two tasks: the reconstruction task and the classification task. The former task aims at reconstructing an input hyperspectral cube, including the labeled and unlabeled ones, whereas the latter task attempts to recognize the category of the cube. Meanwhile, we construct a discriminator network to discriminate the input sample coming from the real distribution or the reconstructed one. Through an adversarial learning method, the generator network will produce real-like cubes, thus indirectly improving the discrimination and generalization ability of the classification task. More importantly, in order to fully explore the useful information from shallow layers, we adopt skip-layer connections in both reconstruction and classification tasks. The proposed MTGAN model is implemented on three standard HSIs, and the experimental results show that it is able to achieve higher performance than other state-of-the-art deep learning models. Renlong Hang, Feng Zhou 0006, Qingshan Liu 0001, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Class-Guided Feature Decoupling Network for Airborne Image SegmentationabstractContextual information has been demonstrated to be helpful for airborne image segmentation. However, most of the previous works focus on the exploitation of spatially contextual information, which is difficult to segment isolated objects, mainly surrounded by uncorrelated objects. To alleviate this issue, we attempt to take advantage of the co-occurrence relations between different classes of objects in the scene. Especially, similar to other works, convolutional features are first extracted to capture the spatially contextual information. Then, a feature decoupling module is designed to encode the class co-occurrence relations into the convolutional features; thus, the most discriminative features can be decoupled. Finally, the segmentation result is inferred from the decoupled features. The whole process is integrated to form an end-to-end network, named class-guided feature decoupling network (CGFDN). Experimental results on two widely used benchmark data sets show that CGFDN obtains competitive results (>90% overall accuracy (OA) on 5-cm-resolution Potsdam and >91% OA on 9-cm-resolution Vaihingen) in comparison with several state-of-the-art models. Feng Zhou 0006, Renlong Hang, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Spectral Super-Resolution Network Guided by Intrinsic Properties of Hyperspectral ImageryabstractHyperspectral imagery (HSI) contains rich spectral information, which is beneficial to many tasks. However, acquiring HSI is difficult because of the limitations of current imaging technology. As an alternative method, spectral super-resolution aims at reconstructing HSI from its corresponding RGB image. Recently, deep learning has shown its power to this task, but most of the used networks are transferred from other domains, such as spatial super-resolution. In this paper, we attempt to design a spectral super-resolution network by taking advantage of two intrinsic properties of HSI. The first one is the spectral correlation. Based on this property, a decomposition subnetwork is designed to reconstruct HSI. The other one is the projection property, i.e., RGB image can be regarded as a three-dimensional projection of HSI. Inspired from it, a self-supervised subnetwork is constructed as a constraint to the decomposition subnetwork. These two subnetworks constitute our end-to-end super-resolution network. In order to test the effectiveness of it, we conduct experiments on three widely used HSI datasets (i.e., CAVE, NUS, and NTIRE2018). Experimental results show that our proposed network can achieve competitive reconstruction performance in comparison with several state-of-the-art networks. Renlong Hang, Qingshan Liu 0001, Zhu Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Backward Attentive Fusing Network With Local Aggregation Classifier for 3D Point Cloud Semantic SegmentationabstractIn this paper, a Backward Attentive Fusing Network with Local Aggregation Classifier (BAF-LAC) is proposed to improve the performance of 3D point cloud semantic segmentation. It consists of a Backward Attentive Fusing Encoder-Decoder (BAF-ED) to learn semantic features and a Local Aggregation Classifier (LAC) to maintain the context-awareness of points. BAF-ED narrows the semantic gap between the encoder and the decoder via fusing multi-layer encoder features with the decoder features. High-level encoder features are transformed into an attention map to modulate low-level encoder features backward. LAC adaptively enhances the intermediate features in point-wise MLPs via aggregating the features of neighboring points into the center point. It takes the place of commonly used post-processing techniques and retains context consistency into the classifier. Equipped with these modules, BAF-LAC can extract discriminative semantic features and predict smoother results. Extensive experiments on Semantic3D, SemanticKITTI, and S3DIS demonstrate that the proposed method can achieve competitive results against the state-of-the-art methods. Hui Shuai, Xiang Xu 0009, Qingshan Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Multi-Stage Feature Fusion Network for Video Super-ResolutionabstractVideo super-resolution (VSR) is to restore a photo-realistic high-resolution (HR) frame from both its corresponding low-resolution (LR) frame (reference frame) and multiple neighboring frames (supporting frames). An important step in VSR is to fuse the feature of the reference frame with the features of the supporting frames. The major issue with existing VSR methods is that the fusion is conducted in a one-stage manner, and the fused feature may deviate greatly from the visual information in the original LR reference frame. In this paper, we propose an end-to-end Multi-Stage Feature Fusion Network that fuses the temporally aligned features of the supporting frames and the spatial feature of the original reference frame at different stages of a feed-forward neural network architecture. In our network, the Temporal Alignment Branch is designed as an inter-frame temporal alignment module used to mitigate the misalignment between the supporting frames and the reference frame. Specifically, we apply the multi-scale dilated deformable convolution as the basic operation to generate temporally aligned features of the supporting frames. Afterwards, the Modulative Feature Fusion Branch, the other branch of our network accepts the temporally aligned feature map as a conditional input and modulates the feature of the reference frame at different stages of the branch backbone. This enables the feature of the reference frame to be referenced at each stage of the feature fusion process, leading to an enhanced feature from LR to HR. Experimental results on several benchmark datasets demonstrate that our proposed method can achieve state-of-the-art performance on VSR task. Huihui Song 0003, Dong Liu 0002, Bo Liu 0005, Qingshan Liu 0001, Dimitris N. Metaxas |
IEEE Trans. Image Process. | 5 |
| 2021 | Learning Deep Global Multi-Scale and Local Attention Features for Facial Expression Recognition in the WildabstractFacial expression recognition (FER) in the wild received broad concerns in which occlusion and pose variation are two key issues. This paper proposed a global multi-scale and local attention network (MA-Net) for FER in the wild. Specifically, the proposed network consists of three main components: a feature pre-extractor, a multi-scale module, and a local attention module. The feature pre-extractor is utilized to pre-extract middle-level features, the multi-scale module to fuse features with different receptive fields, which reduces the susceptibility of deeper convolution towards occlusion and variant pose, while the local attention module can guide the network to focus on local salient features, which releases the interference of occlusion and non-frontal pose problems on FER in the wild. Extensive experiments demonstrate that the proposed MA-Net achieves the state-of-the-art results on several in-the-wild FER benchmarks: CAER-S, AffectNet-7, AffectNet-8, RAFDB, and SFEW with accuracies of 88.42%, 64.53%, 60.29%, 88.40%, and 59.40% respectively. The codes and training logs are publicly available at https://github.com/zengqunzhao/MA-Net. Zengqun Zhao, Qingshan Liu 0001, Shanmin Wang |
IEEE Trans. Image Process. | 2 |
| 2021 | Parallel Connected LSTM for Matrix Sequence Prediction with Elusive CorrelationsabstractThis article is about a challenging problem called matrix sequence prediction, which is motivated from the application of taxi order prediction. Remarkably, the problem differs greatly from previous sequence prediction tasks in the sense that the time-wise correlations are quite elusive; namely, distant entries could be strongly correlated and nearby entries are unnecessarily related. Such distinct specifics make prevalent convolution-recurrence-based methods inadequate to apply. To remedy this trouble, we propose a novel architecture called Parallel Connected LSTM (PcLSTM), which integrates two new mechanisms, Multi-channel Linearized Connection (McLC) and Adaptive Parallel Unit (APU), into the framework of LSTM. Benefiting from the strengths of McLC and APU, our PcLSTM is able to handle well both the elusive correlations within each timestamp and the temporal dependencies across different timestamps, achieving state-of-the-art performance in a set of experiments demonstrated on synthetic and real-world datasets. Qi Zhao 0013, Chuqiao Chen, Guangcan Liu, Qingshan Liu 0001, Shengyong Chen |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2021 | Collaborative Local-Global Learning for Temporal Action ProposalabstractTemporal action proposal generation is an essential and challenging task in video understanding, which aims to locate the temporal intervals that likely contain the actions of interest. Although great progress has been made, the problem is still far from being well solved. In particular, prevalent methods can handle well only the local dependencies (i.e., short-term dependencies) among adjacent frames but are generally powerless in dealing with the global dependencies (i.e., long-term dependencies) between distant frames. To tackle this issue, we propose CLGNet, a novel Collaborative Local-Global Learning Network for temporal action proposal. The majority of CLGNet is an integration of Temporal Convolution Network and Bidirectional Long Short-Term Memory, in which Temporal Convolution Network is responsible for local dependencies while Bidirectional Long Short-Term Memory takes charge of handling the global dependencies. Furthermore, an attention mechanism called the background suppression module is designed to guide our model to focus more on the actions. Extensive experiments on two benchmark datasets, THUMOS’14 and ActivityNet-1.3, show that the proposed method can outperform state-of-the-art methods, demonstrating the strong capability of modeling the actions with varying temporal durations. Yisheng Zhu, Hu Han 0001, Guangcan Liu, Qingshan Liu 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2021 | Stacked U-Shape Network With Channel-Wise Attention for Salient Object DetectionabstractThis paper addresses the core issue of how to learn powerful features for saliency. We have two major observations. First, feature maps of different layers in convolutional neural networks play different roles in saliency detection. Second, different feature channels in the same layer are not of equal importance to saliency, and they often have different response to foreground or background. To address these problems, a stacked U-shape network with channel-wise attention is presented to effectively utilize these features, which mainly consists of a parallel dilated convolution (PDC) module and a multi-level attention cascaded feedback (MACF) module. More specifically, PDC aims to enlarge the receptive field without increasing the computation and effectively avoid the gridding problem. MACF is innovatively designed to adaptively select the cross-layer complementary information, and the inter-dependencies between different channel maps in the same layer can be depicted well. Finally, we adopt a multi-layer loss function to improve the commonly used binary cross entropy loss which treats all pixels equally. The extensive experiments on five saliency detection datasets demonstrate that the proposed method outperforms the state-of-the-art approaches. Junxia Li, Qingshan Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Unsupervised Network Quantization via Fixed-Point FactorizationabstractThe deep neural network (DNN) has achieved remarkable performance in a wide range of applications at the cost of huge memory and computational complexity. Fixed-point network quantization emerges as a popular acceleration and compression method but still suffers from huge performance degradation when extremely low-bit quantization is utilized. Moreover, current fixed-point quantization methods rely heavily on supervised retraining using large amounts of the labeled training data, while the labeled data are hard to obtain in the real-world applications. In this article, we propose an efficient framework, namely, fixed-point factorized network (FFN), to turn all weights into ternary values, i.e., {-1, 0, 1}. We highlight that the proposed FFN framework can achieve negligible degradation even without any supervised retraining on the labeled data. Note that the activations can be easily quantized into an 8-bit format; thus, the resulting networks only have low-bit fixed-point additions that are significantly more efficient than 32-bit floating-point multiply-accumulate operations (MACs). Extensive experiments on large-scale ImageNet classification and object detection on MS COCO show that the proposed FFN can achieve about more than 20× compression and remove most of the multiply operations with comparable accuracy. Codes are available on GitHub at https://github.com/wps712/FFN. Peisong Wang 0001, Qiang Chen 0007, Anda Cheng, Qingshan Liu 0001, Jian Cheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Nonconvex Low-Rank Kernel Sparse Subspace Learning for Keyframe Extraction and Motion SegmentationabstractBy exploiting the kernel trick, the sparse subspace model is extended to the nonlinear version with one or a combination of predefined kernels, but the high-dimensional space induced by predefined kernels is not guaranteed to be able to capture the features of the nonlinear data in theory. In this article, we propose a nonconvex low-rank learning framework in an unsupervised way to learn a kernel to replace the predefined kernel in the sparse subspace model. The learned kernel by a nonconvex relaxation of rank can better exploiting the low-rank property of nonlinear data to induce a high-dimensional Hilbert space that more closely approaches the true feature space. Furthermore, we give a global closed-form optimal solution of the nonconvex rank minimization and prove it. Considering the low-rank and sparseness characteristics of motion capture data in its feature space, we use them to verify the better representation of nonlinear data with the learned kernel via two tasks: keyframe extraction and motion segmentation. The performances on both tasks demonstrate the advantage of our model over the sparse subspace model with predefined kernels and some other related state-of-art methods. Guiyu Xia, Beijia Chen, Huaijiang Sun, Qingshan Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | 3D Tensor Auto-encoder with Application to Video CompressionabstractAuto-encoder has been widely used to compress high-dimensional data such as the images and videos. However, the traditional auto-encoder network needs to store a large number of parameters. Namely, when the input data is of dimension n , the number of parameters in an auto-encoder is in general O ( n ). In this article, we introduce a network structure called 3D Tensor Auto-Encoder (3DTAE). Unlike the traditional auto-encoder, in which a video is represented as a vector, our 3DTAE considers videos as 3D tensors to directly pass tensor objects through the network. The weights of each layer are represented by three small matrices, and thus the number of parameters in 3DTAE is just O ( n 1/3). The compact nature of 3DTAE fits well the needs of video compression. Given an ensemble of high-dimensional videos, we represent them as 3DTAE networks plus some small core tensors, and we further quantize the network parameters and the core tensors to get the final compressed data. Experimental results verify the efficiency of 3DTAE. Yang Li 0039, Guangcan Liu, Yubao Sun, Qingshan Liu 0001, Shengyong Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Deep Object Co-Segmentation via Spatial-Semantic Network ModulationabstractObject co-segmentation is to segment the shared objects in multiple relevant images, which has numerous applications in computer vision. This paper presents a spatial and semantic modulated deep network framework for object co-segmentation. A backbone network is adopted to extract multi-resolution image features. With the multi-resolution features of the relevant images as input, we design a spatial modulator to learn a mask for each image. The spatial modulator captures the correlations of image feature descriptors via unsupervised learning. The learned mask can roughly localize the shared foreground object while suppressing the background. For the semantic modulator, we model it as a supervised image classification task. We propose a hierarchical second-order pooling module to transform the image features for classification use. The outputs of the two modulators manipulate the multi-resolution features by a shift-and-scale operation so that the features focus on segmenting co-object regions. The proposed model is trained end-to-end without any intricate post-processing. Extensive experiments on four image co-segmentation benchmark datasets demonstrate the superior accuracy of the proposed method compared to state-of-the-art methods. The codes are available at http://kaihuazhang.net/. Kaihua Zhang 0001, Bo Liu 0005, Qingshan Liu 0001 |
AAAI | 4 |
| 2020 | Adaptive Graph Convolutional Network With Attention Graph Clustering for Co-Saliency DetectionabstractCo-saliency detection aims to discover the common and salient foregrounds from a group of relevant images. For this task, we present a novel adaptive graph convolutional network with attention graph clustering (GCAGC). Three major contributions have been made, and are experimentally shown to have substantial practical merits. First, we propose a graph convolutional network design to extract information cues to characterize the intra- and inter-image correspondence. Second, we develop an attention graph clustering algorithm to discriminate the common objects from all the salient foreground objects in an unsupervised fashion. Third, we present a unified framework with encoder-decoder structure to jointly train and optimize the graph convolutional network, attention graph cluster, and co-saliency detection decoder in an end-to-end manner. We evaluate our proposed GCAGC method on three co-saliency detection benchmark datasets (iCoseg, Cosal2015 and COCO-SEG). Our GCAGC method obtains significant improvements over the state-of-the-arts on most of them. Kaihua Zhang 0001, Tengpeng Li, Shiwen Shen, Bo Liu 0005, Qingshan Liu 0001 |
CVPR | 6 |
| 2020 | Learning Memory Augmented Cascading Network for Compressed Sensing of Images
Yubao Sun, Qingshan Liu 0001, Rui Huang 0001 |
ECCV (22) | 3 |
| 2020 | ProxyBNN: Learning Binarized Neural Networks via Proxy Matrices
Zitao Mo, Ke Cheng 0002, Qinghao Hu 0001, Peisong Wang 0001, Qingshan Liu 0001, Jian Cheng 0001 |
ECCV (3) | 7 |
| 2020 | Meta-learning with Network Pruning
Hongduan Tian, Bo Liu 0005, Xiao-Tong Yuan, Qingshan Liu 0001 |
ECCV (19) | 4 |
| 2020 | Prinet: A Prior Driven Spectral Super-Resolution NetworkabstractSpectral super-resolution aims to reconstruct hyperspectral images from RGB images directly. In recent years, convolutional networks have been successfully employed to this task. However, few of them take into account the specific properties of hyperspectral images. In this paper, we attempt to design a super-resolution network, named PriNET, based on two prior knowledge about hyperspectral images. The first one is spectral correlation. According to this property, we design a decomposition network to reconstruct hyperspectral images. In this network, the whole spectral bands of hyperspectral images are divided into several groups, and multiple residual networks are proposed to reconstruct them separately. The second knowledge is that the hyperspectral image should be able to generate its corresponding RGB image. Inspired from it, we design a self-supervised network to fine-tune the reconstruction results of the decomposition network. Finally, these two networks are combined together to constitute PriNET. Experimental results on two hyperspectral datasets demonstrate that the proposed PriNET can achieve better performance than several state-of-the-art networks. Renlong Hang, Zhu Li 0001, Qingshan Liu 0001, Shuvra S. Bhattacharyya |
ICME | 3 |
| 2020 | Dual Temporal Memory Network for Efficient Video Object SegmentationabstractVideo Object Segmentation (VOS) is typically formulated in a semi-supervised setting. Given the ground-truth segmentation mask on the first frame, the task of VOS is to track and segment the single or multiple objects of interests in the rest frames of the video at the pixel level. One of the fundamental challenges in VOS is how to make the most use of the temporal information to boost the performance. We present an end-to-end network which stores short- and long-term video sequence information preceding the current frame as the temporal memories to address the temporal modeling in VOS. Our network consists of two temporal sub-networks including a short-term memory sub-network and a long-term memory sub-network. The short-term memory sub-network models the fine-grained spatial-temporal interactions between local regions across neighboring frames in video via a graph-based learning framework, which can well preserve the visual consistency of local regions over time. The long-term memory sub-network models the long-range evolution of object via a Simplified-Gated Recurrent Unit (S-GRU), making the segmentation be robust against occlusions and drift errors. In our experiments, we show that our proposed method achieves a favorable and competitive performance on three frequently-used VOS datasets, including DAVIS 2016, DAVIS 2017 and Youtube-VOS in terms of both speed and accuracy. Kaihua Zhang 0001, Dong Liu 0002, Bo Liu 0005, Qingshan Liu 0001, Zhu Li 0001 |
ACM Multimedia | 5 |
| 2020 | Top-Down Fusing Multi-level Contextual Features for Salient Object Detection
Mingyuan Pan, Huihui Song 0003, Junxia Li, Kaihua Zhang 0001, Qingshan Liu 0001 |
PRCV (3) | 5 |
| 2020 | Learning lightweight Multi-Scale Feedback Residual network for single image super-resolution
Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001, Jia Liu 0034 |
Comput. Vis. Image Underst. | 4 |
| 2020 | Real-time manifold regularized context-aware correlation tracking
Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001, Wei Lian |
Frontiers Comput. Sci. | 4 |
| 2020 | Flow driven attention network for video salient object detectionabstractSalient object detection has been revolutionised by convolutional neural network (CNN) recently. However, it is hard to transfer the state‐of‐the‐art still‐image based saliency detectors to videos directly, owing to the neglect of temporal contexts between frames. In this study, the authors propose a flow‐driven attention network (FDAN) to exploit motion information for video salient object detection. FDAN consists of an appearance feature extractor, a motion‐guided attention module and a saliency map regression module. It extracts the appearance feature per frame, refines appearance feature with optical flow and infers the ultimate saliency map, respectively. Motion‐guided attention module is the core of FDAN, which extracts motion information in the form of attention. This attention mechanism is a two‐branch CNN, fusing optical flow and appearance features. In addition, a shortcut connection is applied to the attention multiplied feature map for noise suppression intensively. Experimental results show that the proposed method can achieve performance on par with the state‐of‐the‐art method flow‐guided recurrent neural encoder on challenging benchmarks of Densely Annotated Video Segmentation and Freiburg–Berkeley Motion Segmentation while being two times faster in detection. Feng Zhou 0006, Hui Shuai, Qingshan Liu 0001, Guodong Guo |
IET Image Process. | 3 |
| 2020 | Recurrent reverse attention guided residual learning for saliency object detection
Tengpeng Li, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
Neurocomputing | 4 |
| 2020 | Single image super-resolution with enhanced Laplacian pyramid network via conditional generative adversarial learning
Huihui Song 0003, Kaihua Zhang 0001, Jiaojiao Qiao, Qingshan Liu 0001 |
Neurocomputing | 5 |
| 2020 | Dual Iterative Hard ThresholdingabstractIterative Hard Thresholding (IHT) is a popular class of first-order greedy selection methods for loss minimization under cardinality constraint. The existing IHT-style algorithms, however, are proposed for minimizing the primal formulation. It is still an open issue to explore duality theory and algorithms for such a non-convex and NP-hard combinatorial optimization problem. To address this issue, we develop in this article a novel duality theory for $\ell_2$-regularized empirical risk minimization under cardinality constraint, along with an IHT-style algorithm for dual optimization. Our sparse duality theory establishes a set of sufficient and/or necessary conditions under which the original non-convex problem can be equivalently or approximately solved in a concave dual formulation. In view of this theory, we propose the Dual IHT (DIHT) algorithm as a super-gradient ascent method to solve the non-smooth dual problem with provable guarantees on primal-dual gap convergence and sparsity recovery. Numerical results confirm our theoretical predictions and demonstrate the superiority of DIHT to the state-of-the-art primal IHT-style algorithms in model estimation accuracy and computational efficiency. Xiao-Tong Yuan, Bo Liu 0005, Lezi Wang, Qingshan Liu 0001, Dimitris N. Metaxas |
J. Mach. Learn. Res. | 4 |
| 2020 | Hierarchical attentive Siamese network for real-time visual tracking
Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
Neural Comput. Appl. | 4 |
| 2020 | Learning residual refinement network with semantic context representation for real-time saliency object detection
Tengpeng Li, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
Pattern Recognit. | 4 |
| 2020 | Learning image compressed sensing with sub-pixel convolutional generative adversarial network
Yubao Sun, Qingshan Liu 0001, Guangcan Liu |
Pattern Recognit. | 3 |
| 2020 | Selecting Optimal Completion to Partial Matrix via Self-ValidationabstractIn many applications such as film recommendation, one often encounters the problem of estimating the unseen entries in a partially observed matrix, formally known as matrix completion. Over the past several decades, lots of effective methods have been established in the literature, and each method may contain several hyper-parameters. For a partial matrix, one can use those methods with certain parametric settings to obtain a large number of completions. Now, a critical question is, how to select the optimal completion from a number of candidates? This question is indeed a hard to answer, because in practice the true values of the missing entries are unknown. Thus far, the only approach for dealing with the issue is through data-validation, which is to first split the observations into two subsets, a training set and a validation set, and then choose the model that performs best on the validation set as the winner to produce the final results. Though straightforward, this approach might fall in a non-optimal model that overfits the validation set. In this work, we shall suggest a different approach called self-validation, which accounts on a special metric that can evaluate the “goodness” of a completion without using any validation data. The metric is derived from the recently established isomeric condition, measuring the identifiable degree of the completion itself. Extensive experiments demonstrate that our self-validation approach is better than the commonly used data-validation. Guangcan Liu, Wayne Zhang 0001, Qingshan Liu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2020 | Classification of Hyperspectral and LiDAR Data Using Coupled CNNsabstractIn this article, we propose an efficient and effective framework to fuse hyperspectral and light detection and ranging (LiDAR) data using two coupled convolutional neural networks (CNNs). One CNN is designed to learn spectral-spatial features from hyperspectral data, and the other one is used to capture the elevation information from LiDAR data. Both of them consist of three convolutional layers, and the last two convolutional layers are coupled together via a parameter-sharing strategy. In the fusion phase, feature-level and decision-level fusion methods are simultaneously used to integrate these heterogeneous features sufficiently. For the feature-level fusion, three different fusion strategies are evaluated, including the concatenation strategy, the maximization strategy, and the summation strategy. For the decision-level fusion, a weighted summation strategy is adopted, where the weights are determined by the classification accuracy of each output. The proposed model is evaluated on an urban data set acquired over Houston, USA, and a rural one captured over Trento, Italy. On the Houston data, our model can achieve a new record overall accuracy (OA) of 96.03%. On the Trento data, it achieves an OA of 99.12%. These results sufficiently certify the effectiveness of our proposed model. Renlong Hang, Zhu Li 0001, Pedram Ghamisi, Danfeng Hong, Guiyu Xia, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2020 | Tropical Cyclone Intensity Estimation Using Two-Branch Convolutional Neural Network From Infrared and Water Vapor ImagesabstractThis article proposes a two-branch convolutional neural network model (TCIENet) to estimate the intensity of tropical cyclone (TC) from infrared and water vapor images in the northwest Pacific basin. Three different sizes of input images are explored to train the TCIENet model, and the size of 60 × 60 pixels (radius 450 km) achieves the best performance with an overall root mean square error (RMSE) of 5.13 m/s and mean absolute error (MAE) of 4.03 m/s. TCs are divided into six categories whose RMSEs range from 4.07 to 6.05 m/s. In addition, the TCs in the year 2017 are used to analyze the correlation between the rainfall intensity from the global precipitation measurement (GPM) mission and the estimation errors of the TCIENet model. Preliminary results suggest that the model performs the best at the categories of tropical storm and super typhoon, but it degrades in performance for moderate intense categories and the weakest category of the tropical depression. The correlation coefficient between the estimation error and the rainfall intensity is 0.19. It is far from certain that the rainfall intensity accounts for the error achieved by the TCIENet model. Rui Zhang 0049, Qingshan Liu 0001, Renlong Hang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Dual-Path Attention Network for Compressed Sensing Image ReconstructionabstractAlthough deep neural network methods achieved much success in compressed sensing image reconstruction in recent years, they still have some issues, especially in preserving texture details. In this paper, we propose a new dual-path attention network for compressed sensing image reconstruction, which is composed of a structure path, a texture path and a texture attention module. Motivated by the classical paradigm of image structure-texture decomposition, the structure path aims to reconstruct the dominant structure component of the original image, and the texture path targets at recovering the remaining texture details. To better bridge the information between two paths, the texture attention module is designed to deliver the useful structure information to the texture path and predict the texture region, thereby facilitating the recovery of texture details. Two paths are optimized with a unified loss function. In the testing phase, given the measurement vector of a new image, it can be well reconstructed by carrying out the well trained dual-path attention network and integrating the outputs of the structure path and the texture path. Experimental results on the SET5, SET11 and BSD68 testing datasets demonstrate that the proposed method achieves comparable or better results compared with some state-of-the-art deep learning based methods and conventional iterative optimization based methods in terms of reconstruction quality and robustness to noise. Yubao Sun, Qingshan Liu 0001, Bo Liu 0005, Guodong Guo |
IEEE Trans. Image Process. | 3 |
| 2020 | Learning Non-Locally Regularized Compressed Sensing Network With Half-Quadratic SplittingabstractDeep learning-based Compressed Sensing (CS) reconstruction attracts much attention in recent years, due to its significant superiority of reconstruction quality. Its success is mainly attributed to the employment of a large dataset for pre-training the network to learn a reconstruction mapping. In this paper, we propose a non-locally regularized compressed sensing network for reconstructing image sequences, which can achieve high reconstruction quality without pre-training. Specifically, the proposed method attempts to learn a deep network prior for the reconstruction of an individual instance under the constraint that the network output can well match the given CS measurement. The non-local prior is designed to guide the network to capture the long-range dependencies by exploiting the self-similarities among images, and it can also make the network noise-aware. In order to deal with the compound of non-local prior and deep network prior, we construct a half-quadratic splitting based optimization method for network learning, in which the two priors are decoupled into two simple sub-problems by introducing an auxiliary variable and a quadratic fidelity constraint. Extensive experimental results demonstrate that our method is competitive to the popular methods, including sparsity prior based methods and deep learning based methods, even better than them in the cases of low measurement rates. Yubao Sun, Qingshan Liu 0001, Xiao-Tong Yuan, Guodong Guo |
IEEE Trans. Multim. | 3 |
| 2019 | Distributed Inexact Newton-type Pursuit for Non-convex Sparse LearningabstractIn this paper, we present a sample distributed greedy pursuit method for non-convex sparse learning under cardinality constraint. Given the training samples uniformly randomly partitioned across multiple machines, the proposed method alternates between local inexact sparse minimization of a Newton-type approximation and centralized global results aggregation. Theoretical analysis shows that for a general class of convex functions with Lipschitze continues Hessian, the method converges linearly with contraction factor scaling inversely to the local data size; whilst the communication complexity required to reach desirable statistical accuracy scales logarithmically with respect to the number of machines for some popular statistical learning models. For nonconvex objective functions, up to a local estimation error, our method can be shown to converge to a local stationary sparse solution with sub-linear communication complexity. Numerical results demonstrate the efficiency and accuracy of our method when applied to large-scale sparse learning tasks including deep neural nets pruning Bo Liu 0005, Xiao-Tong Yuan, Lezi Wang, Qingshan Liu 0001, Junzhou Huang, Dimitris N. Metaxas |
AISTATS | 4 |
| 2019 | Co-Saliency Detection via Mask-Guided Fully Convolutional Networks With Multi-Scale Label SmoothingabstractIn image co-saliency detection problem, one critical issue is how to model the concurrent pattern of the co-salient parts, which appears both within each image and across all the relevant images. In this paper, we propose a hierarchical image co-saliency detection framework as a coarse to fine strategy to capture this pattern. We first propose a mask-guided fully convolutional network structure to generate the initial co-saliency detection result. The mask is used for background removal and it is learned from the high-level feature response maps of the pre-trained VGG-net output. We next propose a multi-scale label smoothing model to further refine the detection result. The proposed model jointly optimizes the label smoothness of pixels and superpixels. Experiment results on three popular image co-saliency detection benchmark datasets including iCoseg, MSRC and Cosal2015 demonstrate the remarkable performance compared with the state-of-the-art methods. Kaihua Zhang 0001, Tengpeng Li, Bo Liu 0005, Qingshan Liu 0001 |
CVPR | 4 |
| 2019 | PointNet-Based Channel Attention VLAD Network
Rongrong Fan, Hui Shuai, Qingshan Liu 0001 |
PRCV (3) | 3 |
| 2019 | Local Context Embedding Neural Network for Scene Semantic Segmentation
Junxia Li, Lingzheng Dai, Qingshan Liu 0001 |
PRCV (2) | 4 |
| 2019 | Image super-resolution using conditional generative adversarial networkabstractRecently, extensive studies on a generative adversarial network (GAN) have made great progress in single image super‐resolution (SISR). However, there still exists a significant difference between the reconstructed high‐frequency and the real high‐frequency details. To address this issue, this study presents an SISR approach based on conditional GAN (SRCGAN). SRCGAN includes a generator network that generates super‐resolution (SR) images and a discriminator network that is trained to distinguish the SR images from ground‐truth high‐resolution (HR) ones. Specifically, the discriminator network uses the ground‐truth HR image as a conditional variable, which guides the network to distinguish the real images from the SR images, facilitating training a more stable generator model than GAN without this guidance. Furthermore, a residual‐learning module is introduced into the generator network to solve the issue of detail information loss in SR images. Finally, the network is trained in an end‐to‐end manner by optimizing a perceptual loss function. Extensive evaluations on four benchmark datasets including Set5, Set14, BSD100, and Urban100 demonstrate the superiority of the proposed SRCGAN over state‐of‐the‐art methods in terms of PSNR, SSIM, and visual effect. Jiaojiao Qiao, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
IET Image Process. | 5 |
| 2019 | Moving object detection via segmentation and saliency constrained RPCA
Yang Li 0039, Guangcan Liu, Qingshan Liu 0001, Yubao Sun, Shengyong Chen |
Neurocomputing | 3 |
| 2019 | Hyperspectral image classification using spectral-spatial LSTMs
Feng Zhou 0006, Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan |
Neurocomputing | 3 |
| 2019 | Retrieving Soil Moisture Over Continental U.S. via Multi-View Multi-Task LearningabstractSoil moisture (SM) is an essential variable in the hydrological cycle. Quantifying the magnitude of SM is crucial for the climate system. In this letter, we present a new model with multi-view multi-task learning (MVMTL) to estimate SM over continental U.S. Specifically, the multi-view component is used to make full use of spatial and temporal features of each grid cell (0.25° × 0.25°). Meanwhile, the multi-task component aims to capture the spatial correlations and to perform coestimations between different grid cells in the study area. To evaluate the effectiveness of MVMTL, we compare it with several retrieval methods in terms of the SM product from the European Center for Medium-Range Weather Forecasts Reanalysis Interim (ERA-Interim) and in situ SM measurements. The experimental results show that the MVMTL model can achieve higher performance than the other methods. Lingling Ge, Renlong Hang, Qingshan Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | Low-rank weighted co-saliency detection via efficient manifold ranking
Tengpeng Li, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001, Wei Lian |
Multim. Tools Appl. | 4 |
| 2019 | Dynamic Structure Embedded Online Multiple-Output Regression for Streaming DataabstractOnline multiple-output regression is an important machine learning technique for modeling, predicting, and compressing multi-dimensional correlated data streams. In this paper, we propose a novel online multiple-output regression method, called MORES, for streaming data. MORES can dynamically learn the structure of the regression coefficients to facilitate the model's continuous refinement. Considering that limited expressive ability of regression models often leading to residual errors being dependent, MORES intends to dynamically learn and leverage the structure of the residual errors to improve the prediction accuracy. Moreover, we introduce three modified covariance matrices to extract necessary information from all the seen data for training, and set different weights on samples so as to track the data streams' evolving characteristics. Furthermore, an efficient algorithm is designed to optimize the proposed objective function, and an efficient online eigenvalue decomposition algorithm is developed for the modified covariance matrix. Finally, we analyze the convergence of MORES in certain ideal condition. Experiments on two synthetic datasets and three real-world datasets validate the effectiveness and efficiency of MORES. In addition, MORES can process at least 2,000 instances per second (including training and testing) on the three real-world datasets, more than 12 times faster than the state-of-the-art online learning algorithm. Weishan Dong, Xiangfeng Wang 0001, Qingshan Liu 0001, Xin Zhang 0008 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | Joint Active Learning with Feature Selection via CUR Matrix DecompositionabstractThis paper presents an unsupervised learning approach for simultaneous sample and feature selection, which is in contrast to existing works which mainly tackle these two problems separately. In fact the two tasks are often interleaved with each other: noisy and high-dimensional features will bring adverse effect on sample selection, while informative or representative samples will be beneficial to feature selection. Specifically, we propose a framework to jointly conduct active learning and feature selection based on the CUR matrix decomposition. From the data reconstruction perspective, both the selected samples and features can best approximate the original dataset respectively, such that the selected samples characterized by the features are highly representative. In particular, our method runs in one-shot without the procedure of iterative sample selection for progressive labeling. Thus, our model is especially suitable when there are few labeled samples or even in the absence of supervision, which is a particular challenge for existing methods. As the joint learning problem is NP-hard, the proposed formulation involves a convex but non-smooth optimization problem. We solve it efficiently by an iterative algorithm, and prove its global convergence. Experimental results on publicly available datasets corroborate the efficacy of our method compared with the state-of-the-art. Xiangfeng Wang 0001, Weishan Dong, Junchi Yan, Qingshan Liu 0001, Hongyuan Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | Pan-GGF: A probabilistic method for pan-sharpening with gradient domain guided image filtering
Peixian Zhuang, Qingshan Liu 0001, Xinghao Ding |
Signal Process. | 2 |
| 2019 | Cascaded Recurrent Neural Networks for Hyperspectral Image ClassificationabstractBy considering the spectral signature as a sequence, recurrent neural networks (RNNs) have been successfully used to learn discriminative features from hyperspectral images (HSIs) recently. However, most of these models only input the whole spectral bands into RNNs directly, which may not fully explore the specific properties of HSIs. In this paper, we propose a cascaded RNN model using gated recurrent units to explore the redundant and complementary information of HSIs. It mainly consists of two RNN layers. The first RNN layer is used to eliminate redundant information between adjacent spectral bands, while the second RNN layer aims to learn the complementary information from nonadjacent spectral bands. To improve the discriminative ability of the learned features, we design two strategies for the proposed model. Besides, considering the rich spatial information contained in HSIs, we further extend the proposed model to its spectral-spatial counterpart by incorporating some convolutional layers. To test the effectiveness of our proposed models, we conduct experiments on two widely used HSIs. The experimental results show that our proposed models can achieve better results than the compared models. Renlong Hang, Qingshan Liu 0001, Danfeng Hong, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Learning-Based Sphere Nonlinear Interpolation for Motion SynthesisabstractMotion synthesis technology can produce natural and coordinated motion data without a motion capture process, which is complex and costly. Current motion synthesis methods usually provide a few interfaces to avoid the arbitrariness of the synthesis process, but this actually reduces the understandability of the synthesis process. In this paper, we propose a learning-based Sphere nonlinear interpolation (Snerp) model that can generate natural in-between motions in terms of a given start-end frame pair. Variety of the input frame pairs will enrich the diversity of the generated motions. The angle speed of natural human motion is not uniform and presents different change rules (we call them motion patterns) for different motions, so we first extract the motion patterns and then build the relation between motion pattern space and frame pair space via a paired dictionary learning process. After learning, we estimate the motion pattern according to the representation of a given start-end frame pair on the frame pair dictionary. We select several different types of start-end frame pairs from the real motion sequences as the testing data and good results of both objective and subjective evaluations on the generated motions demonstrate the superior performance of Snerp. Guiyu Xia, Huaijiang Sun, Qingshan Liu 0001, Renlong Hang |
IEEE Trans. Ind. Informatics | 3 |
| 2019 | Robust Subspace Clustering With Compressed DataabstractDimension reduction is widely regarded as an effective way for decreasing the computation, storage and communication loads of data-driven intelligent systems, leading to a growing demand for statistical methods that allow analysis (e.g., clustering) of compressed data. We therefore study in this paper a novel problem called compressive robust subspace clustering, which is to perform robust subspace clustering with the compressed data, and which is generated by projecting the original high-dimensional data onto a lower-dimensional subspace chosen at random. Given only the compressed data and sensing matrix, the proposed method, row space pursuit (RSP), recovers the authentic row space that gives correct clustering results under certain conditions. Extensive experiments show that RSP is distinctly better than the competing methods, in terms of both clustering accuracy and computational efficiency. Guangcan Liu, Zhao Zhang 0001, Qingshan Liu 0001, Hongkai Xiong |
IEEE Trans. Image Process. | 3 |
| 2019 | Parallel Attentive Correlation TrackingabstractPsychological and cognitive findings indicate that human visual perception is attentive and selective, which may process spatial and appearance selective attentions in parallel. By reflecting some aspects of these attentions, this paper presents a novel correlation filter (CF) based tracking approach, corresponding to processing a local and a semi-local background domains, respectively. In the local domain, inspired by the Gestalt principle of figure-ground segregation, we leverage an efficient Boolean map representation, which characterizes an image by a set of Boolean maps via randomly thresholding its color channels, yielding a location response map as a weighted sum of all Boolean maps. The Boolean maps capture the topological structures of target and its scene with different granularities, thereby enabling to effectively improve tracking of non-rectangular objects. Alternatively, in the semi-local domains, we introduce a novel distractor-resilient metric regularization into CF, which acts as a force to push distractors into negative space. Consequently, the unwanted boundary effects of CF can be effectively alleviated. Finally, both models associated with the local and the semi-local domains are seamlessly integrated into a Bayesian framework, and the tracked location is determined by maximizing its likelihood function. Extensive evaluations on the OTB50, OTB100, VOT2016 and VOT2017 tracking benchmarks demonstrate that the proposed method achieves favorable performance against a variety of state-of-the-art trackers with a speed of 45 fps on a single CPU. Kaihua Zhang 0001, Jiaqing Fan, Qingshan Liu 0001, Jian Yang 0003, Wei Lian |
IEEE Trans. Image Process. | 3 |
| 2018 | Joint Head Pose Estimation with Multi-task Cascaded Convolutional Networks for Face AlignmentabstractIn the past decades, face alignment has been studied widely, but it has long been impeded by the problem of pose variation. Recent studies show that pose information used as additional source of information can help address the above problem. In this paper, we adopt a multi-task cascaded CNNs based framework for simultaneous face detection, dense face alignment and fine head pose estimation. Especially, our framework exploits the inherent correlation between face alignment and fine head pose estimation to boost up landmark detection robustness in the case of various poses. Experiments show that our method not only demonstrates real-time performance for face detection, dense face alignment and fine head pose estimation, but also outperforms most state-of-the-art methods for face alignment on the challenging 300-W benchmark. Especially in the case of large pose variations, it achieves outstanding results. Zhenni Cai, Qingshan Liu 0001, Shanmin Wang, Bruce Yang |
ICPR | 2 |
| 2018 | Integrating Convolutional Neural Network and Gated Recurrent Unit for Hyperspectral Image Spectral-Spatial Classification
Feng Zhou 0006, Renlong Hang, Qingshan Liu 0001, Xiao-Tong Yuan |
PRCV (4) | 3 |
| 2018 | Fast subspace segmentation via Random Sample Probing
Yang Li 0039, Yubao Sun, Qingshan Liu 0001, Shengyong Chen |
Neurocomputing | 3 |
| 2018 | Multi-component group sparse RPCA model for motion object detection under complex dynamic background
Yubao Sun, Renlong Hang, Qingshan Liu 0001, Guangcan Liu |
Neurocomputing | 4 |
| 2018 | Spatio-temporal convolutional features with nested LSTM for facial expression recognition
Zhenbo Yu, Guangcan Liu, Qingshan Liu 0001, Jiankang Deng |
Neurocomputing | 3 |
| 2018 | Visual tracking using spatio-temporally nonlocally regularized correlation filter
Kaihua Zhang 0001, Huihui Song 0003, Qingshan Liu 0001, Wei Lian |
Pattern Recognit. | 4 |
| 2018 | Visual tracking via Boolean map representations
Kaihua Zhang 0001, Qingshan Liu 0001, Jian Yang 0003, Ming-Hsuan Yang 0001 |
Pattern Recognit. | 2 |
| 2018 | Saliency fusion via sparse and double low rank decomposition
Junxia Li, Jian Yang 0003, Chen Gong 0002, Qingshan Liu 0001 |
Pattern Recognit. Lett. | 4 |
| 2018 | Visual Tracking via Nonlocal Similarity LearningabstractEither global (e.g., intensity histograms and coefficients of sparse representation) or local (e.g., scale-invariant feature transform and histogram of oriented gradient) feature representations have been widely exploited for visual tracking. However, most of these representations describe a target appearance with a fixed spatial grid layout without considering the interactions between different grids, and hence may adversely affect their performance when the target appearance suffers from large-scale pose variations. In this paper, we learn a similarity function that considers the interactions of features in the grids not only from the same spatial positions, but also from different positions, thereby taking charge of the nonlocal information of the target appearances to effectively handle the significant appearance variations. Specifically, we explore the polynomial kernel feature map to characterize the nonlocal similarity information of all pairs of grids among the target and its background samples, and combine these feature maps as the target representations. Moveover, we learn a linear logistic regression classifier with online update to separate the target from its local background, and integrate this classifier into a particle filtering tracking framework. Extensive experimental results on the CVPR2013 tracking benchmark demonstrate the proposed approach performs favorably against some representative tracking algorithms. Qingshan Liu 0001, Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Learning Multiscale Deep Features for High-Resolution Satellite Image Scene ClassificationabstractIn this paper, we propose a multiscale deep feature learning method for high-resolution satellite image scene classification. Specifically, we first warp the original satellite image into multiple different scales. The images in each scale are employed to train a deep convolutional neural network (DCNN). However, simultaneously training multiple DCNNs is time-consuming. To address this issue, we explore DCNN with spatial pyramid pooling (SPP-net). Since different SPP-nets have the same number of parameters, which share the identical initial values, and only fine-tuning the parameters in fully connected layers ensures the effectiveness of each network, thereby greatly accelerating the training process. Then, the multiscale satellite images are fed into their corresponding SPP-nets, respectively, to extract multiscale deep features. Finally, a multiple kernel learning method is developed to automatically learn the optimal combination of such features. Experiments on two difficult data sets show that the proposed method achieves favorable performance compared with other state-of-the-art methods. Qingshan Liu 0001, Renlong Hang, Huihui Song 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Nonlinear Low-Rank Matrix Completion for Human Motion RecoveryabstractHuman motion capture data has been widely used in many areas, but it involves a complex capture process and the captured data inevitably contains missing data due to the occlusions caused by the actor's body or clothing. Motion recovery, which aims to recover the underlying complete motion sequence from its degraded observation, still remains as a challenging task due to the nonlinear structure and kinematics property embedded in motion data. Low-rank matrix completion based methods have shown promising performance in short-time-missing motion recovery problems. However, low-rank matrix completion, which is designed for linear data, lacks the theoretic guarantee when applied to the recovery of nonlinear motion data. To overcome this drawback, we propose a tailored nonlinear matrix completion model for human motion recovery. Within the model, we first learn a combined low-rank kernel via multiple kernel learning. By exploiting the learned kernel, we embed the motion data into a high dimensional Hilbert space where motion data is of desirable low-rank and we then use the low-rank matrix completion to recover motions. In addition, we add two kinematic constraints to the proposed model to preserve the kinematics property of human motion. Extensive experiment results and comparisons with five other state-of-the-art methods demonstrate the advantage of the proposed method. Guiyu Xia, Huaijiang Sun, Beijia Chen, Qingshan Liu 0001, Lei Feng 0003, Guoqing Zhang 0002, Renlong Hang |
IEEE Trans. Image Process. | 4 |
| 2018 | A Self-Paced Regularization Framework for Multilabel LearningabstractIn this brief, we propose a novel multilabel learning framework, called multilabel self-paced learning, in an attempt to incorporate the SPL scheme into the regime of multilabel learning. Specifically, we first propose a new multilabel learning formulation by introducing a self-paced function as a regularizer, so as to simultaneously prioritize label learning tasks and instances in each iteration. Considering that different multilabel learning scenarios often need different self-paced schemes during learning, we thus provide a general way to find the desired self-paced functions. To the best of our knowledge, this is the first work to study multilabel learning by jointly taking into consideration the complexities of both training instances and labels. Experimental results on four publicly available data sets suggest the effectiveness of our approach, compared with the state-of-the-art methods. Junchi Yan, Xiaoyu Zhang 0002, Qingshan Liu 0001, Hongyuan Zha |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | Reversed Spectral HashingabstractHashing is emerging as a powerful tool for building highly efficient indices in large-scale search systems. In this paper, we study spectral hashing (SH), which is a classical method of unsupervised hashing. In general, SH solves for the hash codes by minimizing an objective function that tries to preserve the similarity structure of the data given. Although computationally simple, very often SH performs unsatisfactorily and lags distinctly behind the state-of-the-art methods. We observe that the inferior performance of SH is mainly due to its imperfect formulation; that is, the optimization of the minimization problem in SH actually cannot ensure that the similarity structure of the high-dimensional data is really preserved in the low-dimensional hash code space. In this paper, we, therefore, introduce reversed SH (ReSH), which is SH with its input and output interchanged. Unlike SH, which estimates the similarity structure from the given high-dimensional data, our ReSH defines the similarities between data points according to the unknown low-dimensional hash codes. Equipped with such a reversal mechanism, ReSH can seamlessly overcome the drawback of SH. More precisely, the minimization problem in our ReSH can be optimized if and only if similar data points are mapped to adjacent hash codes, and mostly important, dissimilar data points are considerably separated from each other in the code space. Finally, we solve the minimization problem in ReSH by multilayer neural networks and obtain state-of-the-art retrieval results on three benchmark data sets. Qingshan Liu 0001, Guangcan Liu, Lai Li, Xiao-Tong Yuan, Meng Wang 0001, Wei Liu 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Self-Paced Multi-Task LearningabstractMulti-task learning is a paradigm, where multiple tasks are jointly learnt. Previous multi-task learning models usually treat all tasks and instances per task equally during learning. Inspired by the fact that humans often learn from easy concepts to hard ones in the cognitive process, in this paper, we propose a novel multi-task learning framework that attempts to learn the tasks by simultaneously taking into consideration the complexities of both tasks and instances per task. We propose a novel formulation by presenting a new task-oriented regularizer that can jointly prioritize tasks and instances.Thus it can be interpreted as a self-paced learner for multi-task learning. An efficient block coordinate descent algorithm is developed to solve the proposed objective function, and the convergence of the algorithm can be guaranteed. Experimental results on the toy and real-world datasets demonstrate the effectiveness of the proposed approach, compared to the state-of-the-arts. Junchi Yan, Weishan Dong, Qingshan Liu 0001, Hongyuan Zha |
AAAI | 5 |
| 2017 | Dual Iterative Hard Thresholding: From Non-convex Sparse Minimization to Non-smooth Concave MaximizationabstractIterative Hard Thresholding (IHT) is a class of projected gradient descent methods for optimizing sparsity-constrained minimization models, with the best known efficiency and scalability in practice. As far as we know, the existing IHT-style methods are designed for sparse minimization in primal form. It remains open to explore duality theory and algorithms in such a non-convex and NP-hard setting. In this article, we bridge the gap by establishing a duality theory for sparsity-constrained minimization with $\ell_2$-regularized objective and proposing an IHT-style algorithm for dual maximization. Our sparse duality theory provides a set of sufficient and necessary conditions under which the original NP-hard/non-convex problem can be equivalently solved in a dual space. The proposed dual IHT algorithm is a super-gradient method for maximizing the non-smooth dual objective. An interesting finding is that the sparse recovery performance of dual IHT is invariant to the Restricted Isometry Property (RIP), which is required by all the existing primal IHT without sparsity relaxation. Moreover, a stochastic variant of dual IHT is proposed for large-scale stochastic optimization. Numerical results demonstrate that dual IHT algorithms can achieve more accurate model estimation given small number of training data and have higher computational efficiency than the state-of-the-art primal IHT-style algorithms. Bo Liu 0005, Xiao-Tong Yuan, Lezi Wang, Qingshan Liu 0001, Dimitris N. Metaxas |
ICML | 4 |
| 2017 | Multicore implementation of the multi-scale adaptive deep pyramid matching model for remotely sensed image classificationabstractArtificial neural networks (ANNs) have been widely used in the analysis of remotely sensed imagery. In particular, convolutional neural networks (CNNs) are gaining more and more attention. Unlike traditional CNNs methods, where the relevant information to classify the elements of a remotely sensed image is extracted only from the last fully-connected layer, the new adaptive deep pyramid matching (ADPM) model [1] takes advantage of the features from all of the convolutional layers. This model allows the optimal fusing weights for different convolutional layers be learned from the data itself. In addition, the combination of CNNs with spatial pyramid pooling (SPP-net) to create the basic deep network allows the use of images with multiple scales, which results in better learning process thanks to the complementary information. The original ADPM method is divided in two parts: the multi-scale deep feature extraction and the ADPM core. In this paper we present a computational improvement of the ADPM core, coding a parallel-multicore version. This strategy is shown to significantly enhance performance in the analysis of remotely sensed data. Mercedes Eugenia Paoletti, Juan Mario Haut, Javier Plaza, Antonio Plaza, Qingshan Liu 0001, Renlong Hang |
IGARSS | 5 |
| 2017 | A New Theory for Matrix CompletionabstractPrevalent matrix completion theories reply on an assumption that the locations of the missing data are distributed uniformly and randomly (i.e., uniform sampling). Nevertheless, the reason for observations being missing often depends on the unseen observations themselves, and thus the missing data in practice usually occurs in a nonuniform and deterministic fashion rather than randomly. To break through the limits of random sampling, this paper introduces a new hypothesis called \emph{isomeric condition}, which is provably weaker than the assumption of uniform sampling and arguably holds even when the missing data is placed irregularly. Equipped with this new tool, we prove a series of theorems for missing data recovery and matrix completion. In particular, we prove that the exact solutions that identify the target matrix are included as critical points by the commonly used nonconvex programs. Unlike the existing theories for nonconvex matrix completion, which are built upon the same condition as convex programs, our theory shows that nonconvex programs have the potential to work with a much weaker condition. Comparing to the existing studies on nonuniform sampling, our setup is more general. Guangcan Liu, Qingshan Liu 0001, Xiao-Tong Yuan |
NIPS | 2 |
| 2017 | Face image retrieval based on shape and texture feature fusionabstractHumongous amounts of data bring various challenges to face image retrieval. This paper proposes an efficient method to solve those problems. Firstly, we use accurate facial landmark locations as shape features. Secondly, we utilise shape priors to provide discriminative texture features for convolutional neural networks. These shape and texture features are fused to make the learned representation more robust. Finally, in order to increase efficiency, a coarse-tofine search mechanism is exploited to efficiently find similar objects. Extensive experiments on the CASIAWebFace, MSRA-CFW, and LFW datasets illustrate the superiority of our method. Zongguang Lu, Jing Yang 0038, Qingshan Liu 0001 |
Comput. Vis. Media | 3 |
| 2017 | Blessing of Dimensionality: Recovering Mixture Data via Dictionary PursuitabstractThis paper studies the problem of recovering the authentic samples that lie on a union of multiple subspaces from their corrupted observations. Due to the high-dimensional and massive nature of today's data-driven community, it is arguable that the target matrix (i.e., authentic sample matrix) to recover is often low-rank. In this case, the recently established Robust Principal Component Analysis (RPCA) method already provides us a convenient way to solve the problem of recovering mixture data. However, in general, RPCA is not good enough because the incoherent condition assumed by RPCA is not so consistent with the mixture structure of multiple subspaces. Namely, when the subspace number grows, the row-coherence of data keeps heightening and, accordingly, RPCA degrades. To overcome the challenges arising from mixture data, we suggest to consider LRR in this paper. We elucidate that LRR can well handle mixture data, as long as its dictionary is configured appropriately. More precisely, we mathematically prove that LRR can weaken the dependence on the row-coherence, provided that the dictionary is well-conditioned and has a rank of not too high. In particular, if the dictionary itself is sufficiently low-rank, then the dependence on the row-coherence can be completely removed. These provide some elementary principles for dictionary learning and naturally lead to a practical algorithm for recovering mixture data. Our experiments on randomly generated matrices and real motion sequences show promising results. Guangcan Liu, Qingshan Liu 0001, Ping Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Newton-Type Greedy Selection Methods for ℓ0-Constrained MinimizationabstractWe introduce a family of Newton-type greedy selection methods for -constrained minimization problems. The basic idea is to construct a quadratic function to approximate the original objective function around the current iterate and solve the constructed quadratic program over the cardinality constraint. The next iterate is then estimated via a line search operation between the current iterate and the solution of the sparse quadratic program. This iterative procedure can be interpreted as an extension of the constrained Newton methods from convex minimization to non-convex -constrained minimization. We show that the proposed algorithms converge asymptotically and the rate of local convergence is superlinear up to certain estimation error. Our methods compare favorably against several state-of-the-art greedy selection methods when applied to sparse logistic regression and sparse support vector machines. Xiao-Tong Yuan, Qingshan Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Robust facial landmark tracking via cascade regression
Qingshan Liu 0001, Jing Yang 0038, Jiankang Deng, Kaihua Zhang 0001 |
Pattern Recognit. | 1 |
| 2017 | Adaptive Compressive Tracking via Online Vector Boosting Feature SelectionabstractRecently, the compressive tracking (CT) method has attracted much attention due to its high efficiency, but it cannot well deal with the large scale target appearance variations due to its data-independent random projection matrix that results in less discriminative features. To address this issue, in this paper, we propose an adaptive CT approach, which selects the most discriminative features to design an effective appearance model. Our method significantly improves CT in three aspects. First, the most discriminative features are selected via an online vector boosting method. Second, the object representation is updated in an effective online manner, which preserves the stable features while filtering out the noisy ones. Furthermore, a simple and effective trajectory rectification approach is adopted that can make the estimated location more accurate. Finally, a multiple scale adaptation mechanism is explored to estimate object size, which helps to relieve interference from background information. Extensive experiments on the CVPR2013 tracking benchmark and the VOT2014 challenges demonstrate the superior performance of our method. Qingshan Liu 0001, Jing Yang 0038, Kaihua Zhang 0001, Yi Wu 0001 |
IEEE Trans. Cybern. | 1 |
| 2017 | Parallel Sparse Subspace Clustering via Joint Sample and Parameter Blockwise PartitionabstractSparse subspace clustering (SSC) is a classical method to cluster data with specific subspace structure for each group. It has many desirable theoretical properties and has been shown to be effective in various applications. However, under the condition of a large-scale dataset, learning the sparse sample affinity graph is computationally expensive. To tackle the computation time cost challenge, we develop a memory-efficient parallel framework for computing SSC via an alternating direction method of multiplier (ADMM) algorithm. The proposed framework partitions the data matrix into column blocks and then decomposes the original problem into parallel multivariate Lasso regression subproblems and samplewise operations. The proposed method allows us to allocate multiple cores/machines for the processing of individual column blocks. We propose a stochastic optimization algorithm to minimize the objective function. Experimental results on real-world datasets demonstrate that the proposed blockwise ADMM framework is substantially more efficient than its matrix counterpart used by SSC, without sacrificing performance in applications. Moreover, our approach is directly applicable to parallel neighborhood selection for Gaussian graphical models structure estimation. Bo Liu 0005, Xiao-Tong Yuan, Yang Yu 0010, Qingshan Liu 0001, Dimitris N. Metaxas |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2017 | Adaptive Cascade Regression Model For Robust Face AlignmentabstractCascade regression is a popular face alignment approach, and it has achieved good performances on the wild databases. However, it depends heavily on local features in estimating reliable landmark locations and therefore suffers from corrupted images, such as images with occlusion, which often exists in real-world face images. In this paper, we present a new adaptive cascade regression model for robust face alignment. In each iteration, the shape-indexed appearance is introduced to estimate the occlusion level of each landmark, and each landmark is then weighted according to its estimated occlusion level. Also, the occlusion levels of the landmarks act as adaptive weights on the shape-indexed features to decrease the noise on the shape-indexed features. At the same time, an exemplar-based shape prior is designed to suppress the influence of local image corruption. Extensive experiments are conducted on the challenging benchmarks, and the experimental results demonstrate that the proposed method achieves better results than the state-of-the-art methods for facial landmark localization and occlusion detection. Qingshan Liu 0001, Jiankang Deng, Jing Yang 0038, Guangcan Liu, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2017 | Elastic Net Hypergraph Learning for Image Clustering and Semi-Supervised ClassificationabstractGraph model is emerging as a very effective tool for learning the complex structures and relationships hidden in data. In general, the critical purpose of graph-oriented learning algorithms is to construct an informative graph for image clustering and classification tasks. In addition to the classical K-nearest-neighbor and r-neighborhood methods for graph construction, l1-graph and its variants are emerging methods for finding the neighboring samples of a center datum, where the corresponding ingoing edge weights are simultaneously derived by the sparse reconstruction coefficients of the remaining samples. However, the pairwise links of l1-graph are not capable of capturing the high-order relationships between the center datum and its prominent data in sparse reconstruction. Meanwhile, from the perspective of variable selection, the l1norm sparse constraint, regarded as a LASSO model, tends to select only one datum from a group of data that are highly correlated and ignore the others. To simultaneously cope with these drawbacks, we propose a new elastic net hypergraph learning model, which consists of two steps. In the first step, the robust matrix elastic net model is constructed to find the canonically related samples in a somewhat greedy way, achieving the grouping effect by adding the l2penalty to the l1constraint. In the second step, hypergraph is used to represent the high order relationships between each datum and its prominent samples by regarding them as a hyperedge. Subsequently, hypergraph Laplacian matrix is constructed for further analysis. New hypergraph learning algorithms, including unsupervised clustering and multi-class semi-supervised classification, are then derived. Extensive experiments on face and handwriting databases demonstrate the effectiveness of the proposed method. Qingshan Liu 0001, Yubao Sun, Cantian Wang, Tongliang Liu, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2016 | Spatially Regularized Streaming Sensor SelectionabstractSensor selection has become an active topic aimed at energy saving, information overload prevention, and communication cost planning in sensor networks. In many real applications, often the sensors' observation regions have overlaps and thus the sensor network is inherently redundant. Therefore it is important to select proper sensors to avoid data redundancy. This paper focuses on how to incrementally select a subset of sensors in a streaming scenario to minimize information redundancy, and meanwhile meet the power consumption constraint. We propose to perform sensor selection in a multi-variate interpolation framework, such that the data sampled by the selected sensors can well predict those of the inactive sensors. Importantly, we incorporate sensors' spatial information as two regularizers, which leads to significantly better prediction performance. We also define a statistical variable to store sufficient information for incremental learning, and introduce a forgetting factor to track sensor streams' evolvement. Experiments on both synthetic and real datasets validate the effectiveness of the proposed method. Moreover, our method is over 10 times faster than the state-of-the-art sensor selection algorithm. Weishan Dong, Xiangfeng Wang 0001, Junchi Yan, Xiaobin Zhu 0001, Qingshan Liu 0001, Xin Zhang 0008 |
AAAI | 7 |
| 2016 | Decentralized Robust Subspace ClusteringabstractWe consider the problem of subspace clustering using the SSC (Sparse Subspace Clustering) approach, which has several desirable theoretical properties and has been shown to be effective in various computer vision applications.We develop a large scale distributed framework for the computation of SSC via an alternating direction method of multiplier (ADMM) algorithm. The proposed framework solves SSC in column blocks and only involves parallel multivariate Lasso regression subproblems and sample-wise operations. This appealing property allows us to allocate multiple cores/machines for the processing of individual column blocks.We evaluate our algorithm on a shared-memory architecture. Experimental results on real-world datasets confirm that the proposed block-wise ADMM framework is substantially more efficient than its matrix counterpart used by SSC,without sacrificing accuracy. Moreover, our approach is directly applicable to decentralized neighborhood selection for Gaussian graphical models structure estimation. Bo Liu 0005, Xiao-Tong Yuan, Yang Yu 0010, Qingshan Liu 0001, Dimitris N. Metaxas |
AAAI | 4 |
| 2016 | Efficient k-Support-Norm Regularized Minimization via Fully Corrective Frank-Wolfe Method
Bo Liu 0005, Xiao-Tong Yuan, Shaoting Zhang 0001, Qingshan Liu 0001, Dimitris N. Metaxas |
IJCAI | 4 |
| 2016 | Advancing Iterative Quantization Hashing Using Isotropic Prior
Lai Li, Guangcan Liu, Qingshan Liu 0001 |
MMM (2) | 3 |
| 2016 | Learning Additive Exponential Family Graphical Models via \ell_{2, 1}-norm Regularized M-EstimationabstractWe investigate a subclass of exponential family graphical models of which the sufficient statistics are defined by arbitrary additive forms. We propose two $\ell_{2,1}$-norm regularized maximum likelihood estimators to learn the model parameters from i.i.d. samples. The first one is a joint MLE estimator which estimates all the parameters simultaneously. The second one is a node-wise conditional MLE estimator which estimates the parameters for each node individually. For both estimators, statistical analysis shows that under mild conditions the extra flexibility gained by the additive exponential family models comes at almost no cost of statistical efficiency. A Monte-Carlo approximation method is developed to efficiently optimize the proposed estimators. The advantages of our estimators over Gaussian graphical models and Nonparanormal estimators are demonstrated on synthetic and real data sets. Xiao-Tong Yuan, Ping Li 0001, Tong Zhang 0001, Qingshan Liu 0001, Guangcan Liu |
NIPS | 4 |
| 2016 | Robust object tracking by online Fisher discrimination boosting feature selection
Jing Yang 0038, Kaihua Zhang 0001, Qingshan Liu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2016 | Robust visual tracking via patch based kernel correlation filters with adaptive multiple feature ensemble
Kaihua Zhang 0001, Qingshan Liu 0001 |
Neurocomputing | 3 |
| 2016 | M3 CSR: Multi-view, multi-scale and multi-component cascade shape regression
Jiankang Deng, Qingshan Liu 0001, Jing Yang 0038, Dacheng Tao |
Image Vis. Comput. | 2 |
| 2016 | A Deterministic Analysis for LRRabstractThe recently proposed low-rank representation (LRR) method has been empirically shown to be useful in various tasks such as motion segmentation, image segmentation, saliency detection and face recognition. While potentially powerful, LRR depends heavily on the configuration of its key parameter, λ. In realistic environments where the prior knowledge about data is lacking, however, it is still unknown how to choose λ in a suitable way. Even more, there is a lack of rigorous analysis about the success conditions of the method, and thus the significance of LRR is a little bit vague. In this paper we therefore establish a theoretical analysis for LRR, striving for figuring out under which conditions LRR can be successful, and deriving a moderately good estimate to the key parameter λ as well. Simulations on synthetic data points and experiments on real motion sequences verify our claims. Guangcan Liu, Huan Xu 0001, Jinhui Tang 0001, Qingshan Liu 0001, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | FaceHunter: A multi-task convolutional neural network based face detector
Jing Yang 0038, Jiankang Deng, Qingshan Liu 0001 |
Signal Process. Image Commun. | 4 |
| 2016 | Matrix-Based Discriminant Subspace Ensemble for Hyperspectral Image Spatial-Spectral Feature FusionabstractSpatial-spectral feature fusion is well acknowledged as an effective method for hyperspectral (HS) image classification. Many previous studies have been devoted to this subject. However, these methods often regard the spatial-spectral high-dimensional data as 1-D vector and then extract informative features for classification. In this paper, we propose a new HS image classification method. Specifically, matrix-based spatial-spectral feature representation is designed for each pixel to capture the local spatial contextual and the spectral information of all the bands, which can well preserve the spatial-spectral correlation. Then, matrix-based discriminant analysis is adopted to learn the discriminative feature subspace for classification. To further improve the performance of discriminative subspace, a random sampling technique is used to produce a subspace ensemble for final HS image classification. Experiments are conducted on three HS remote sensing data sets acquired by different sensors, and experimental results demonstrate the efficiency of the proposed method. Renlong Hang, Qingshan Liu 0001, Huihui Song 0002, Yubao Sun |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Dual Sparse Constrained Cascade Regression for Robust Face AlignmentabstractLocalizing facial landmarks is a fundamental step in facial image analysis. However, the problem continues to be challenging due to the large variability in expression, illumination, pose, and the existence of occlusions in real-world face images. In this paper, we present a dual sparse constrained cascade regression model for robust face alignment. Instead of using the least-squares method during the training process of regressors, sparse constraint is introduced to select robust features and compress the size of the model. Moreover, sparse shape constraint is incorporated between each cascade regression, and the explicit shape constraints are able to suppress the ambiguity in local features. To improve the model's adaptation to large pose variation, face pose is estimated by five fiducial landmarks located by deep convolutional neuron network, which is used to adaptively design the cascade regression model. To the best of our best knowledge, this is the first attempt to fuse explicit shape constraint (sparse shape constraint) and implicit context information (sparse feature selection) for robust face alignment in the framework of cascade regression. Extensive experiments on nine challenging wild data sets demonstrate the advantages of the proposed method over the state-of-the-art methods. Qingshan Liu 0001, Jiankang Deng, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2016 | Robust Visual Tracking via Convolutional Networks Without TrainingabstractDeep networks have been successfully applied to visual tracking by learning a generic representation offline from numerous training images. However, the offline training is time-consuming and the learned generic representation may be less discriminative for tracking specific objects. In this paper, we present that, even without offline training with a large amount of auxiliary data, simple two-layer convolutional networks can be powerful enough to learn robust representations for visual tracking. In the first frame, we extract a set of normalized patches from the target region as fixed filters, which integrate a series of adaptive contextual filters surrounding the target to define a set of feature maps in the subsequent frames. These maps measure similarities between each filter and useful local intensity patterns across the target, thereby encoding its local structural information. Furthermore, all the maps together form a global representation, via which the inner geometric layout of the target is also preserved. A simple soft shrinkage method that suppresses noisy values below an adaptive threshold is employed to de-noise the global representation. Our convolutional networks have a lightweight structure and perform favorably against several state-of-the-art methods on the recent tracking benchmark data set with 50 challenging videos. Kaihua Zhang 0001, Qingshan Liu 0001, Yi Wu 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Stacked Sparse Autoencoder (SSAE) for Nuclei Detection on Breast Cancer Histopathology ImagesabstractAutomated nuclear detection is a critical step for a number of computer assisted pathology related image analysis algorithms such as for automated grading of breast cancer tissue specimens. The Nottingham Histologic Score system is highly correlated with the shape and appearance of breast cancer nuclei in histopathological images. However, automated nucleus detection is complicated by 1) the large number of nuclei and the size of high resolution digitized pathology images, and 2) the variability in size, shape, appearance, and texture of the individual nuclei. Recently there has been interest in the application of "Deep Learning" strategies for classification and analysis of big image data. Histopathology, given its size and complexity, represents an excellent use case for application of deep learning strategies. In this paper, a Stacked Sparse Autoencoder (SSAE), an instance of a deep learning strategy, is presented for efficient nuclei detection on high-resolution histopathological images of breast cancer. The SSAE learns high-level features from just pixel intensities alone in order to identify distinguishing features of nuclei. A sliding window operation is applied to each image in order to represent image patches via high-level features obtained via the auto-encoder, which are then subsequently fed to a classifier which categorizes each image patch as nuclear or non-nuclear. Across a cohort of 500 histopathological images (2200 × 2200) and approximately 3500 manually segmented individual nuclei serving as the groundtruth, SSAE was shown to have an improved F-measure 84.49% and an average area under Precision-Recall curve (AveP) 78.83%. The SSAE approach also out-performed nine other state of the art nuclear detection strategies. Jun Xu 0005, Lei Xiang 0001, Qingshan Liu 0001, Hannah Gilmore, Jianzhong Wu, Jinghai Tang, Anant Madabhushi |
IEEE Trans. Medical Imaging | 3 |
| 2016 | Max-Margin-Based Discriminative Feature LearningabstractIn this brief, we propose a new max-margin-based discriminative feature learning method. In particular, we aim at learning a low-dimensional feature representation, so as to maximize the global margin of the data and make the samples from the same class as close as possible. In order to enhance the robustness to noise, we leverage a regularization term to make the transformation matrix sparse in rows. In addition, we further learn and leverage the correlations among multiple categories for assisting in learning discriminative features. The experimental results demonstrate the power of the proposed method against the related state-of-the-art methods. Qingshan Liu 0001, Weishan Dong, Xin Zhang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Efficient χ2 Kernel Linearization via Random Feature MapsabstractExplicit feature mapping is an appealing way to linearize additive kernels, such as χ2kernel for training large-scale support vector machines (SVMs). Although accurate in approximation, feature mapping could pose computational challenges in high-dimensional settings as it expands the original features to a higher dimensional space. To handle this issue in the context of χ2kernel SVMs learning, we introduce a simple yet efficient method to approximately linearize χ2kernel through random feature maps. The main idea is to use sparse random projection to reduce the dimensionality of feature maps while preserving their approximation capability to the original kernel. We provide approximation error bound for the proposed method. Furthermore, we extend our method to χ2multiple kernel SVMs learning. Extensive experiments on large-scale image classification tasks confirm that the proposed approach is able to significantly speed up the training process of the χ2kernel SVMs at almost no cost of testing accuracy. Xiao-Tong Yuan, Jiankang Deng, Qingshan Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2015 | Additive Nearest Neighbor Feature MapsabstractIn this paper, we present a concise framework to approximately construct feature maps for nonlinear additive kernels such as the Intersection, Hellinger's, and X2kernels. The core idea is to construct for each individual feature a set of anchor points and assign to every query the feature map of its nearest neighbor or the weighted combination of those of its k-nearest neighbors in the anchors. The resultant feature maps can be compactly stored by a group of nearest neighbor (binary) indication vectors along with the anchor feature maps. The approximation error of such an anchored feature mapping approach is analyzed. We evaluate the performance of our approach on large-scale nonlinear support vector machines~(SVMs) learning tasks in the context of visual object classification. Experimental results on several benchmark data sets show the superiority of our method over existing feature mapping methods in achieving reasonable trade-off between training time and testing accuracy. Xiao-Tong Yuan, Qingshan Liu 0001, Shuicheng Yan |
ICCV | 3 |
| 2015 | Hierarchical Convolutional Neural Network for Face Detection
Jing Yang 0038, Jiankang Deng, Qingshan Liu 0001 |
ICIG (2) | 4 |
| 2015 | Sparse random projection for χ2 kernel linearization: Algorithm and applications to image classification
Xiao-Tong Yuan, Qingshan Liu 0001 |
Neurocomputing | 3 |
| 2015 | Low rank driven robust facial landmark regression
Jiankang Deng, Yubao Sun, Qingshan Liu 0001, Hanqing Lu |
Neurocomputing | 3 |
| 2015 | How friends affect user behaviors? An exploration of social relation analysis for recommendation
Jian Cheng 0001, Xi Zhang 0018, Qingshan Liu 0001, Hanqing Lu |
Knowl. Based Syst. | 4 |
| 2015 | Inferring occluded features for fast object detection
Shengye Yan, Qingshan Liu 0001 |
Signal Process. | 2 |
| 2015 | Human Age Estimation Based on Locality and Ordinal InformationabstractIn this paper, we propose a novel feature selection-based method for facial age estimation. The face aging is a typical temporal process, and facial images should have certain ordinal patterns in the aging feature space. From the geometrical perspective, a facial image can be usually seen as sampled from a low-dimensional manifold embedded in the original high-dimensional feature space. Thus, we first measure the energy of each feature in preserving the underlying local structure information and the ordinal information of the facial images, respectively, and then we intend to learn a low-dimensional aging representation that can maximally preserve both kinds of information. To further improve the performance, we try to eliminate the redundant local information and ordinal information as much as possible by minimizing nonlinear correlation and rank correlation among features. Finally, we formulate all these issues into a unified optimization problem, which is similar to linear discriminant analysis in format. Since it is expensive to collect the labeled facial aging images in practice, we extend the proposed supervised method to a semi-supervised learning mode including the semi-supervised feature selection method and the semi-supervised age prediction algorithm. Extensive experiments are conducted on the FACES dataset, the Images of Groups dataset, and the FG-NET aging dataset to show the power of the proposed algorithms, compared to the state-of-the-arts. Qingshan Liu 0001, Weishan Dong, Xiaobin Zhu 0001, Jing Liu 0001, Hanqing Lu |
IEEE Trans. Cybern. | 2 |
| 2015 | A Variational Approach to Simultaneous Image Segmentation and Bias CorrectionabstractThis paper presents a novel variational approach for simultaneous estimation of bias field and segmentation of images with intensity inhomogeneity. We model intensity of inhomogeneous objects to be Gaussian distributed with different means and variances, and then introduce a sliding window to map the original image intensity onto another domain, where the intensity distribution of each object is still Gaussian but can be better separated. The means of the Gaussian distributions in the transformed domain can be adaptively estimated by multiplying the bias field with a piecewise constant signal within the sliding window. A maximum likelihood energy functional is then defined on each local region, which combines the bias field, the membership function of the object region, and the constant approximating the true signal from its corresponding object. The energy functional is then extended to the whole image domain by the Bayesian learning approach. An efficient iterative algorithm is proposed for energy minimization, via which the image segmentation and bias field correction are simultaneously achieved. Furthermore, the smoothness of the obtained optimal bias field is ensured by the normalized convolutions without extra cost. Experiments on real images demonstrated the superiority of the proposed algorithm to other state-of-the-art representative methods. Kaihua Zhang 0001, Qingshan Liu 0001, Huihui Song 0003, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2015 | Learning Multiscale Active Facial Patches for Expression AnalysisabstractIn this paper, we present a new idea to analyze facial expression by exploring some common and specific information among different expressions. Inspired by the observation that only a few facial parts are active in expression disclosure (e.g., around mouth, eye), we try to discover the common and specific patches which are important to discriminate all the expressions and only a particular expression, respectively. A two-stage multitask sparse learning (MTSL) framework is proposed to efficiently locate those discriminative patches. In the first stage MTSL, expression recognition tasks are combined to located common patches. Each of the tasks aims to find dominant patches for each expression. Secondly, two related tasks, facial expression recognition and face verification tasks, are coupled to learn specific facial patches for individual expression. The two-stage patch learning is performed on patches sampled by multiscale strategy. Extensive experiments validate the existence and significance of common and specific patches. Utilizing these learned patches, we achieve superior performances on expression recognition compared to the state-of-the-arts. Lin Zhong 0002, Qingshan Liu 0001, Peng Yang 0001, Junzhou Huang, Dimitris N. Metaxas |
IEEE Trans. Cybern. | 2 |
| 2015 | Improving the Spatial Resolution of Landsat TM/ETM+ Through Fusion With SPOT5 Images via Learning-Based Super-ResolutionabstractTo take advantage of the wide swath width of Landsat Thematic Mapper (TM)/Enhanced Thematic Mapper Plus (ETM+) images and the high spatial resolution of Système Pour l'Observation de la Terre 5 (SPOT5) images, we present a learning-based super-resolution method to fuse these two data types. The fused images are expected to be characterized by the swath width of TM/ETM+ images and the spatial resolution of SPOT5 images. To this end, we first model the imaging process from a SPOT image to a TM/ETM+ image at their corresponding bands, by building an image degradation model via blurring and downsampling operations. With this degradation model, we can generate a simulated Landsat image from each SPOT5 image, thereby avoiding the requirement for geometric coregistration for the two input images. Then, band by band, image fusion can be implemented in two stages: 1) learning a dictionary pair representing the high- and low-resolution details from the given SPOT5 and the simulated TM/ETM+ images; 2) super-resolving the input Landsat images based on the dictionary pair and a sparse coding algorithm. It is noteworthy that the proposed method can also deal with the conventional spatial and spectral fusion of TM/ETM+ and SPOT5 images by using the learned dictionary pairs. To examine the performance of the proposed method of fusing the swath width of TM/ETM+ and the spatial resolution of SPOT5, we illustrate the fusion results on the actual TM images and compare with several classic pansharpening methods by assuming that the corresponding SPOT5 panchromatic image exists. Furthermore, we implement the classification experiments on both actual images and fusion results to demonstrate the benefits of the proposed method for further classification applications. Huihui Song 0003, Bo Huang 0001, Qingshan Liu 0001, Kaihua Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2015 | Ordinal Distance Metric Learning for Image RankingabstractRecently, distance metric learning (DML) has attracted much attention in image retrieval, but most previous methods only work for image classification and clustering tasks. In this brief, we focus on designing ordinal DML algorithms for image ranking tasks, by which the rank levels among the images can be well measured. We first present a linear ordinal Mahalanobis DML model that tries to preserve both the local geometry information and the ordinal relationship of the data. Then, we develop a nonlinear DML method by kernelizing the above model, considering of real-world image data with nonlinear structures. To further improve the ranking performance, we finally derive a multiple kernel DML approach inspired by the idea of multiple-kernel learning that performs different kernel operators on different kinds of image features. Extensive experiments on four benchmarks demonstrate the power of the proposed algorithms against some related state-of-the-art methods. Qingshan Liu 0001, Jing Liu 0001, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | Newton Greedy Pursuit: A Quadratic Approximation Method for Sparsity-Constrained OptimizationabstractFirst-order greedy selection algorithms have been widely applied to sparsity-constrained optimization. The main theme of this type of methods is to evaluate the function gradient in the previous iteration to update the non-zero entries and their values in the next iteration. In contrast, relatively less effort has been made to study the second-order greedy selection method additionally utilizing the Hessian information. Inspired by the classic constrained Newton method, we propose in this paper the NewTon Greedy Pursuit (NTGP) method to approximately minimizes a twice differentiable function over sparsity constraint. At each iteration, NTGP constructs a second-order Taylor expansion to approximate the cost function, and estimates the next iterate as the solution of the constructed quadratic model over sparsity constraint. Parameter estimation error and convergence property of NTGP are analyzed. The superiority of NTGP to several representative first-order greedy selection methods is demonstrated in synthetic and real sparse logistic regression tasks. Xiao-Tong Yuan, Qingshan Liu 0001 |
CVPR | 2 |
| 2014 | Fast Visual Tracking via Dense Spatio-temporal Context Learning
Kaihua Zhang 0001, Lei Zhang 0006, Qingshan Liu 0001, David Zhang 0001, Ming-Hsuan Yang 0001 |
ECCV (5) | 3 |
| 2014 | Multiple-Output Regression with High-Order Structure Information
Qingshan Liu 0001, Fan Jing Meng, Weishan Dong, Yu Wang 0021, Jingmin Xu |
ICPR | 3 |
| 2014 | Discriminative Context Models for Collective Activity RecognitionabstractContext information has been widely studied for recognizing collective activities. Most existing works assume that all individuals in a single image share the same activity label. However, in many cases, multiple activities can be coexisted and serve as the context for each other in real-world scenarios. Based on this observation, we propose a novel approach to model both the intra-class and inter-class behavior interactions among persons in the scenario. By introducing the intra-class and inter-class context descriptors, we propose a unified discriminative model to jointly capture the individual appearance information and the context patterns around the focal person in a max-margin framework. Finally, a greedy forward search method is utilized to optimally label the activities in the testing scene. Experimental results demonstrate the superiority of our approach in activity recognition. Chaoyang Zhao, Jinqiao Wang, Xiao Bai 0001, Qingshan Liu 0001, Hanqing Lu |
ICPR | 5 |
| 2014 | Immune network-based swarm intelligence and its application to unmanned aerial vehicle (UAV) swarm coordination
Liguo Weng, Qingshan Liu 0001, Min Xia 0002, Yongduan Song 0001 |
Neurocomputing | 2 |
| 2014 | Learning the object location, scale and view for image categorization with adapted classifier
Shengye Yan, Xinxing Xu, Qingshan Liu 0001 |
Inf. Sci. | 3 |
| 2014 | Learning Discriminative Dictionary for Group Sparse RepresentationabstractIn recent years, sparse representation has been widely used in object recognition applications. How to learn the dictionary is a key issue to sparse representation. A popular method is to use l1 norm as the sparsity measurement of representation coefficients for dictionary learning. However, the l1 norm treats each atom in the dictionary independently, so the learned dictionary cannot well capture the multisubspaces structural information of the data. In addition, the learned subdictionary for each class usually shares some common atoms, which weakens the discriminative ability of the reconstruction error of each subdictionary. This paper presents a new dictionary learning model to improve sparse representation for image classification, which targets at learning a class-specific subdictionary for each class and a common subdictionary shared by all classes. The model is composed of a discriminative fidelity, a weighted group sparse constraint, and a subdictionary incoherence term. The discriminative fidelity encourages each class-specific subdictionary to sparsely represent the samples in the corresponding class. The weighted group sparse constraint term aims at capturing the structural information of the data. The subdictionary incoherence term is to make all subdictionaries independent as much as possible. Because the common subdictionary represents features shared by all classes, we only use the reconstruction error of each class-specific subdictionary for classification. Extensive experiments are conducted on several public image databases, and the experimental results demonstrate the power of the proposed method, compared with the state-of-the-arts. Yubao Sun, Qingshan Liu 0001, Jinhui Tang 0001, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2013 | A Weighted One Class Collaborative Filtering with Content Topic Features
Jian Cheng 0001, Xi Zhang 0018, Qingshan Liu 0001, Hanqing Lu |
MMM (2) | 4 |
| 2013 | Video event description in scene context
Changbo Hu, Qingshan Liu 0001, Jake K. Aggarwal |
Neurocomputing | 3 |
| 2013 | Learning for scalable multimedia representation
Qingshan Liu 0001, Yueting Zhuang |
Neurocomputing | 1 |
| 2013 | Recognizing expressions from face and body gesture by temporal normalized motion and appearance features
Shizhi Chen, Yingli Tian, Qingshan Liu 0001, Dimitris N. Metaxas |
Image Vis. Comput. | 3 |
| 2013 | Ordinal regularized manifold feature extraction for image ranking
Qingshan Liu 0001, Jing Liu 0001, Hanqing Lu |
Signal Process. | 2 |
| 2012 | Learning ordinal discriminative features for age estimationabstractIn this paper, we present a new method for facial age estimation based on ordinal discriminative feature learning. Considering the temporally ordinal and continuous characteristic of aging process, the proposed method not only aims at preserving the local manifold structure of facial images, but also it wants to keep the ordinal information among aging faces. Moreover, we try to remove redundant information from both the locality information and ordinal information as much as possible by minimizing nonlinear correlation and rank correlation. Finally, we formulate these two issues into a unified optimization problem of feature selection and present an efficient solution. The experiments are conducted on the public available Images of Groups dataset and the FG-NET dataset, and the experimental results demonstrate the power of the proposed method against the state-of-the-art methods. Qingshan Liu 0001, Jing Liu 0001, Hanqing Lu |
CVPR | 2 |
| 2012 | Learning active facial patches for expression analysisabstractIn this paper, we present a new idea to analyze facial expression by exploring some common and specific information among different expressions. Inspired by the observation that only a few facial parts are active in expression disclosure (e.g., around mouth, eye), we try to discover the common and specific patches which are important to discriminate all the expressions and only a particular expression, respectively. A two-stage multi-task sparse learning (MTSL) framework is proposed to efficiently locate those discriminative patches. In the first stage MTSL, expression recognition tasks, each of which aims to find dominant patches for each expression, are combined to located common patches. Second, two related tasks, facial expression recognition and face verification tasks, are coupled to learn specific facial patches for individual expression. Extensive experiments validate the existence and significance of common and specific patches. Utilizing these learned patches, we achieve superior performances on expression recognition compared to the state-of-the-arts. Lin Zhong 0002, Qingshan Liu 0001, Peng Yang 0001, Bo Liu 0005, Junzhou Huang, Dimitris N. Metaxas |
CVPR | 2 |
| 2012 | Learning distance metric regression for facial age estimation
Qingshan Liu 0001, Jing Liu 0001, Hanqing Lu |
ICPR | 2 |
| 2012 | Ordinal preserving projection: a novel dimensionality reduction method for image rankingabstractLearning to rank has been demonstrated as a powerful tool for image ranking, but the issue of the "curse of dimensionality" is a key challenge of learning a ranking model from a large image database. This paper proposes a novel dimensionality reduction algorithm named ordinal preserving projection (OPP) for learning to rank. We first define two matrices, which work in the row direction and column direction respectively. The two matrices aim at leveraging the global structure of the data set and ordinal information of the observations. By maximizing the corresponding objective functions, we can obtain two optimal projection matrices mapping original data points into low-dimensional subspace, in which both global structure and ordinal information can be preserved. The experiments are conducted on the public available MSRA-MM image data set and "Web Queries" image data set, and the experimental results demonstrate the effectiveness of the proposed method. Jing Liu 0001, Yan Liu 0004, Changsheng Xu, Qingshan Liu 0001, Hanqing Lu |
ICMR | 5 |
| 2012 | Temporal Spectral Residual for fast salient motion detection
Xinyi Cui, Qingshan Liu 0001, Shaoting Zhang 0001, Fei Yang 0001, Dimitris N. Metaxas |
Neurocomputing | 2 |
| 2012 | Special Section on Object and Event Classification in Large-Scale Video CollectionsabstractThe nine papers in this special section on object and event classification in large-scale video collections can be categorized into four themes: video indexing, concept detection, video summarization, and event recognition. Changsheng Xu, Alan Hanjalic, Shuicheng Yan, Qingshan Liu 0001, Alan F. Smeaton |
IEEE Trans. Multim. | 4 |
| 2011 | A Framework for the Recognition of Nonmanual Markers in Segmented Sequences of American Sign LanguageabstractDespite the fact that there is critical grammatical information expressed through facial expressions and head gestures, most research in the field of sign language recognition has primarily focused on the manual component of signing.We propose a novel framework for robust tracking and analysis of non-manual behaviours, with an application to sign language recognition.The novelty of our method is threefold.First, we propose a dynamic feature representation.Instead of using only the features available in the current frame (e.g., head pose), we additionally aggregate and encode the feature values in neighbouring frames to better encode the dynamics of expressions and gestures (e.g., head shakes).Second, we use Multiple Instance Learning [12] to handle feature misalignment resulting from drifting of the face tracker and partial occlusions.Third, we utilize a discriminative Hidden Markov Support Vector Machine (HMSVM) [1] to learn finer temporal dependencies between the features of interest.We apply our signerindependent framework to segmented recognition of five classes of grammatical constructions conveyed through facial expressions and head gestures: wh-questions, negation, conditional/when clauses, yes/no questions and topics, and show improvement over previous methods. Nicholas Michael, Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas, Carol Neidle |
BMVC | 3 |
| 2011 | Abnormal detection using interaction energy potentialsabstractA new method is proposed to detect abnormal behaviors in human group activities. This approach effectively models group activities based on social behavior analysis. Different from previous work that uses independent local features, our method explores the relationships between the current behavior state of a subject and its actions. An interaction energy potential function is proposed to represent the current behavior state of a subject, and velocity is used as its actions. Our method does not depend on human detection or segmentation, so it is robust to detection errors. Instead, tracked spatio-temporal interest points are able to provide a good estimation of modeling group interaction. SVM is used to find abnormal events. We evaluate our algorithm in two datasets: UMN and BEHAVE. Experimental results show its promising performance against the state-of-art methods. Xinyi Cui, Qingshan Liu 0001, Mingchen Gao, Dimitris N. Metaxas |
CVPR | 2 |
| 2011 | Segment and recognize expression phase by fusion of motion area and neutral divergence featuresabstractAn expression can be approximated by a sequence of temporal segments called neutral, onset, offset and apex. However, it is not easy to accurately detect such temporal segments only based on facial features. Some researchers try to temporally segment expression phases with the help of body gesture analysis. The problem of this approach is that the expression temporal phases from face and gesture channels are not synchronized. Additionally, most previous work adopted facial key points tracking or body tracking to extract motion information, which is unreliable in practice due to illumination variations and occlusions. In this paper, we present a novel algorithm to overcome the above issues, in which two simple and robust features are designed to describe face and gesture information, i.e., motion area and neutral divergence features. Both features do not depend on motion tracking, and they can be easily calculated too. Moreover, it is different from previous work in that we integrate face and body gesture together in modeling the temporal dynamics through a single channel of sensorial source, so it avoids the unsynchronized issue between face and gesture channels. Extensive experimental results demonstrate the effectiveness of the proposed algorithm. Shizhi Chen, Yingli Tian, Qingshan Liu 0001, Dimitris N. Metaxas |
FG | 3 |
| 2011 | Content quality based image retrieval with multiple instance boost rankingabstractMost previous works treated image retrieval as a classification problem or a similarity measurement problem. In this paper, we propose a new idea for image retrieval, in which we regard image retrieval as a ranking issue by evaluating image content quality. Based on the content preference between the images, the image pairs are organized to build the data set for rank learning. Because image content generally is disclosed by image patches with meaningful objects, each image is looked as one bag, and the regions inside are the corresponding instances. In order to save the computation cost, the instances in the image are the rectangle regions and the integral histogram is applied to speed up histogram feature extraction. Due to the feature dimension is high, we propose a boost-based multiple instance learning for image retrieval. Based on different assumptions in multiple instance setting, Mean, Max and TopK ranking models are developed with Boost learning. Experiments on the real-world images from Flickr, Pisca, and Google shows that the power of the proposed method. Peng Yang 0001, Qingshan Liu 0001, Lin Zhong 0002, Dimitris N. Metaxas |
ACM Multimedia | 3 |
| 2011 | Dynamic soft encoded patterns for facial event analysis
Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas |
Comput. Vis. Image Underst. | 2 |
| 2011 | Unsupervised Image Categorization by Hypergraph PartitionabstractWe present a framework for unsupervised image categorization in which images containing specific objects are taken as vertices in a hypergraph and the task of image clustering is formulated as the problem of hypergraph partition. First, a novel method is proposed to select the region of interest (ROI) of each image, and then hyperedges are constructed based on shape and appearance features extracted from the ROIs. Each vertex (image) and its k-nearest neighbors (based on shape or appearance descriptors) form two kinds of hyperedges. The weight of a hyperedge is computed as the sum of the pairwise affinities within the hyperedge. Through all of the hyperedges, not only the local grouping relationships among the images are described, but also the merits of the shape and appearance characteristics are integrated together to enhance the clustering performance. Finally, a generalized spectral clustering technique is used to solve the hypergraph partition problem. We compare the proposed method to several methods and its effectiveness is demonstrated by extensive experiments on three image databases. Yuchi Huang, Qingshan Liu 0001, Fengjun Lv, Yihong Gong, Dimitris N. Metaxas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Hypergraph with sampling for image retrieval
Qingshan Liu 0001, Yuchi Huang, Dimitris N. Metaxas |
Pattern Recognit. | 1 |
| 2011 | Identifying Regional Cardiac Abnormalities From Myocardial Strains Using Nontracking-Based Strain Estimation and Spatio-Temporal Tensor AnalysisabstractMyocardial strain is a critical indicator of many cardiac diseases and dysfunctions. The goal of this paper is to extract and use the myocardial strain pattern from tagged magnetic resonance imaging (MRI) to identify and localize regional abnormal cardiac function in human subjects. In order to extract the myocardial strains from the tagged images, we developed a novel nontracking-based strain estimation method for tagged MRI. This method is based on the direct extraction of tag deformation, and therefore avoids some limitations of conventional displacement or tracking-based strain estimators. Based on the extracted spatio-temporal strain patterns, we have also developed a novel tensor-based classification framework that better conserves the spatio-temporal structure of the myocardial strain pattern than conventional vector-based classification algorithms. In addition, the tensor-based projection function keeps more of the information of the original feature space, so that abnormal tensors in the subspace can be back-projected to reveal the regional cardiac abnormality in a more physically meaningful way. We have tested our novel methods on 41 human image sequences, and achieved a classification rate of 87.80%. The regional abnormalities recovered from our algorithm agree well with the patient's pathology and clinical image interpretation, and provide a promising avenue for regional cardiac function analysis. Qingshan Liu 0001, Dimitris N. Metaxas, Leon Axel |
IEEE Trans. Medical Imaging | 2 |
| 2011 | A Component-Based Framework for Generalized Face AlignmentabstractThis paper presents a component-based deformable model for generalized face alignment, in which a novel bistage statistical model is proposed to account for both local and global shape characteristics. Instead of using statistical analysis on the entire shape, we build separate Gaussian models for shape components to preserve more detailed local shape deformations. In each model of components, a Markov network is integrated to provide simple geometry constraints for our search strategy. In order to make a better description of the nonlinear interrelationships over shape components, the Gaussian process latent variable model is adopted to obtain enough control of shape variations. In addition, we adopt an illumination-robust feature to lead the local fitting of every shape point when light conditions change dramatically. To further boost the accuracy and efficiency of our component-based algorithm, an efficient subwindow search technique is adopted to detect components and to provide better initializations for shape components. Based on this approach, our system can generate accurate shape alignment results not only for images with exaggerated expressions and slight shading variation but also for images with occlusion and heavy shadows, which are rarely reported in previous work. Yuchi Huang, Qingshan Liu 0001, Dimitris N. Metaxas |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2010 | Image retrieval via probabilistic hypergraph rankingabstractIn this paper, we propose a new transductive learning framework for image retrieval, in which images are taken as vertices in a weighted hypergraph and the task of image search is formulated as the problem of hypergraph ranking. Based on the similarity matrix computed from various feature descriptors, we take each image as a `centroid' vertex and form a hyperedge by a centroid and its k-nearest neighbors. To further exploit the correlation information among images, we propose a probabilistic hypergraph, which assigns each vertex vito a hyperedge ejin a probabilistic way. In the incidence structure of a probabilistic hypergraph, we describe both the higher order grouping information and the affinity relationship between vertices within each hy-peredge. After feedback images are provided, our retrieval system ranks image labels by a transductive inference approach, which tends to assign the same label to vertices that share many incidental hyperedges, with the constraints that predicted labels of feedback images should be similar to their initial labels. We compare the proposed method to several other methods and its effectiveness is demonstrated by extensive experiments on Corel5K, the Scene dataset and Caltech 101. Yuchi Huang, Qingshan Liu 0001, Shaoting Zhang 0001, Dimitris N. Metaxas |
CVPR | 2 |
| 2010 | Exploring facial expressions with compositional featuresabstractMost previous work focuses on how to learn discriminating appearance features over all the face without considering the fact that each facial expression is physically composed of some relative action units (AU). However, the definition of AU is an ambiguous semantic description in Facial Action Coding System (FACS), so it makes accurate AU detection very difficult. In this paper, we adopt a scheme of compromise to avoid AU detection, and try to interpret facial expression by learning some compositional appearance features around AU areas. We first divided face image into local patches according to the locations of AUs, and then we extract local appearance features from each patch. A minimum error based optimization strategy is adopted to build compositional features based on local appearance features, and this process embedded into Boosting learning structure. Experiments on the Cohn-Kanada database show that the proposed method has a promising performance and the built compositional features are basically consistent to FACS. Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas |
CVPR | 2 |
| 2010 | Lesion-Specific Coronary Artery Calcium Quantification for Predicting Cardiac Event with Multiple Instance Support Vector Machines
Qingshan Liu 0001, Idean Marvasty, Sarah Rinehart, Szilard Voros, Dimitris N. Metaxas |
MICCAI (1) | 1 |
| 2010 | Lennard-Jones force field for geometric active contour
Zhenglong Li 0001, Qingshan Liu 0001, Hanqing Lu, Dimitris N. Metaxas |
Signal Process. | 2 |
| 2009 | Video object segmentation by hypergraph cutabstractIn this paper, we present a new framework of video object segmentation, in which we formulate the task of extracting prominent objects from a scene as the problem of hypergraph cut. We initially over-segment each frame in the sequence, and take the over-segmented image patches as the vertices in the graph. Different from the traditional pairwise graph structure, we build a novel graph structure, hypergraph, to represent the complex spatio-temporal neighborhood relationship among the patches. We assign each patch with several attributes that are computed from the optical flow and the appearance-based motion profile, and the vertices with the same attribute value is connected by a hyperedge. Through all the hyperedges, not only the complex non-pairwise relationships between the patches are described, but also their merits are integrated together organically. The task of video object segmentation is equivalent to the hypergraph partition, which can be solved by the hypergraph cut algorithm. The effectiveness of the proposed method is demonstrated by extensive experiments on nature scenes. Yuchi Huang, Qingshan Liu 0001, Dimitris N. Metaxas |
CVPR | 2 |
| 2009 | RankBoost with l1 regularization for facial expression recognition and intensity estimationabstractMost previous facial expression analysis works only focused on expression recognition. In this paper, we propose a novel framework of facial expression analysis based on the ranking model. Different from previous works, it not only can do facial expression recognition, but also can estimate the intensity of facial expression, which is very important to further understand human emotion. Although it is hard to label expression intensity quantitatively, the ordinal relationship in temporal domain is actually a good relative measurement. Based on this observation, we convert the problem of intensity estimation to a ranking problem, which is modeled by the RankBoost. The output ranking score can be directly used for intensity estimation, and we also extend the ranking function for expression recognition. To further improve the performance, we propose to introduce l 1 based regularization into the Rankboost. Experiments on the Cohn-Kanade database show that the proposed method has a promising performance compared to the state-of-the-art. Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas |
ICCV | 2 |
| 2009 | Expanded bag of words representation for object classificationabstractCurrently, the bag of visual words (BOW) representation has received wide applications in object categorization. However, the BOW representation ignores the dependency relationship among visual words, which could provide informative knowledge to understand an image. In this paper, we first design a simple method to discover this dependency through computing the spatial correlation between visual words in overlapped local patches. Obtaining the dependency relationship, we further propose a novel update strategy to modify the BOW representation. The modification is motivated by the idea of Query Expansion applied successfully in text retrieval. We implement our approach on challenging PASCAL 2006 database, and the experimental results show its improved performance against the BOW representation. Tinglin Liu, Jing Liu 0001, Qingshan Liu 0001, Hanqing Lu |
ICIP | 3 |
| 2009 | A variational multi-view learning framework and its application to image segmentationabstractThe paper presents a novel multi-view learning framework based on variational inference. We formulate the framework as a graph representation in form of graph factorization: the graph comprises of factor graphs, which are used to describe internal states of views. Each view is modeled with a Gaussian mixture model. The proposed framework has three main advantages (1) less constraint assumed on data, (2) effective utilization of unlabeled data, and (3) automatic data structure inferring: proper data structure can be inferred in only one round. The experiments on image segmentation demonstrate its effectiveness. Zhenglong Li 0001, Qingshan Liu 0001, Hanqing Lu |
ICME | 2 |
| 2009 | Temporal spectral residual: fast motion saliency detectionabstractSaliency detection has attracted much attention in recent years. It aims at locating semantic regions in images for further image understanding. In this paper, we address the issue of motion saliency detection for video content analysis. Inspired by the idea of Spectral Residual for image saliency detection, we propose a new method Temporal Spectral Residual on video slices along X-T and Y-T planes, which can automatically separate foreground motion objects from backgrounds, also with the help of threshold selection and voting schemes. Different from conventional background modeling methods with complex mathematical model, the proposed method is only based on Fourier spectrum analysis, so it is simple and fast. The power of our proposed method is demonstrated in the experiments of four typical videos with different dynamic background. Xinyi Cui, Qingshan Liu 0001, Dimitris N. Metaxas |
ACM Multimedia | 2 |
| 2009 | Introduction to computer vision and image understanding the special issue on video analysis
Qingshan Liu 0001, Xuelong Li 0001, Ahmed M. Elgammal, Xian-Sheng Hua 0001, Dong Xu 0001, Dacheng Tao |
Comput. Vis. Image Underst. | 1 |
| 2009 | Image annotation via graph learning
Jing Liu 0001, Mingjing Li, Qingshan Liu 0001, Hanqing Lu, Songde Ma |
Pattern Recognit. | 3 |
| 2009 | Introduction to the special issue on Video-based Object and Event Analysis
Shiguang Shan, Qingshan Liu 0001, Dacheng Tao, Dong Xu 0001, Shuicheng Yan, Xuelong Li 0001 |
Pattern Recognit. Lett. | 2 |
| 2009 | Boosting encoded dynamic features for facial expression recognition
Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas |
Pattern Recognit. Lett. | 2 |
| 2008 | Facial expression recognition using encoded dynamic featuresabstractIn this paper, we propose a novel framework for video-based facial expression recognition, which can handle the data with various time resolution including a single frame. We first use the haar-like features to represent facial appearance, due to their simplicity and effectiveness. Then we perform K-Means clustering on the facial appearance features to explore the intrinsic temporal patterns of each expression. Based on the temporal pattern models, we further map the facial appearance variations into dynamic binary patterns. Finally, boosting learning is performed to construct the expression classifiers. Compared to previous work, the dynamic binary patterns encode the intrinsic dynamics of expression, and our method makes no assumption on the time resolution of the data. Extensive experiments carried on the Cohn-Kanade database show the promising performance of the proposed method. Peng Yang 0001, Qingshan Liu 0001, Xinyi Cui, Dimitris N. Metaxas |
CVPR | 2 |
| 2008 | Similarity Features for Facial Event Analysis
Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas |
ECCV (1) | 2 |
| 2008 | Human-centered image navigation on mobile devicesabstractHow to navigate large images on mobile device is an open problem due to small display screen. Currently region of interest (ROI) image compressing is a popular approach. However, since the semantics of image is diversified by different users, it is hard to extract general ROI. In this paper, we propose a human-centered image navigation method for mobile users with simple human interaction. Different from previous work, we first extract Local Saliency Map (LSM) of an image according to personalized requirement, which can reduce the semantic ambiguity of the image for different users. Based on LSM, we can detect the ROIs of the image, which fully satisfies the different interests of the different users. Experiments and user studies show an encouraging performance of the proposed method. Cunxun Zang, Qingshan Liu 0001, Jian Cheng 0001, Hanqing Lu |
ICME | 2 |
| 2008 | A variational inference based approach for image segmentationabstractIn this paper, we present a variational Bayes (VB) approach for image segmentation. First, image is modeled by a mixture model, and then with the techniques of factor analyzer, the underlying structure of image content is inferred automatically. Different from the traditional EM algorithm that seriously suffers from component number selection, the proposed method can accurately infer the underlying image structure including suitable component number without usual sub- or over-segmentation problem. To overcome the problem of local optimization, a component split strategy is adopted in inference optimization process. Extensive experiments on various images validate the proposed method. Zhenglong Li 0001, Qingshan Liu 0001, Jian Cheng 0001, Hanqing Lu |
ICPR | 2 |
| 2008 | Lennard-Jones force field for Geometric Active ContourabstractThis paper presents a new Geometric Active Contour (GAC) model based on Lennard-Jones (L-J) force field, which is inspired by the theory of intermolecular interaction. It is different from gradient based GAC models in that the proposed model does not rely on any pre-computed edge map and is directly computed from image data. Moreover, it can integrate various information including grayscale, color and texture etc. The proposed L-J force field has two different characteristics controlled by a switch parameter c. In the case of c = 0, the force vector flows will bi-directionally converge to boundaries, and it will obtain a morphological dilation-like effect with c ≠ 0. We test the proposed method on various images, and the experimental results are very promising. Zhenglong Li 0001, Qingshan Liu 0001, Hanqing Lu, Dimitris N. Metaxas |
ICPR | 2 |
| 2008 | A Multilevel Contextual Approach to Change Detection for very high Resolution ImagesabstractA multilevel contextual approach is proposed in this paper for change detection of VHR images. By representing the change features in a hierarchical contextual manner, the changes are detected level-by-level. By taking advantages of SVMs, the ambiguity of changes is mitigated and the optimal changes are detected peculiar to the specific user. Compared to the traditional methods, the proposed approach is more accurate, more robust and faster. Experiments demonstrate the effectiveness and advantages of the proposed approach. Chunlei Huo, Zhixin Zhou, Hanqing Lu, Jian Cheng 0001, Qingshan Liu 0001 |
IGARSS (4) | 6 |
| 2008 | Urban Change Detection based on Local Features and Multiscale FusionabstractA multiscale approach is presented in this paper for urban change detection of VHR images. The proposed approach detects the changes at different scales by local-region-based approach, which consists of local region extraction, local region description and local region comparison. To combine the changes at different scales and improve the accuracy, multiscale fusion strategy is applied to local-region-based change detection. Experimental results obtained on Quickbird images confirm the effectiveness of the proposed approach. Chunlei Huo, Zhixin Zhou, Qingshan Liu 0001, Jian Cheng 0001, Hanqing Lu |
IGARSS (3) | 3 |
| 2008 | Identifying Regional Cardiac Abnormalities from Myocardial Strains Using Spatio-temporal Tensor Analysis
Qingshan Liu 0001, Dimitris N. Metaxas, Leon Axel |
MICCAI (1) | 2 |
| 2008 | A new extension of kernel feature and its application for visual recognition
Qingshan Liu 0001, Hongliang Jin, Xiaoou Tang, Hanqing Lu, Songde Ma |
Neurocomputing | 1 |
| 2008 | A Multimodal Scheme for Program Segmentation and Representation in Broadcast Video StreamsabstractWith the advance of digital video recording and playback systems, the request for efficiently managing recorded TV video programs is evident so that users can readily locate and browse their favorite programs. In this paper, we propose a multimodal scheme to segment and represent TV video streams. The scheme aims to recover the temporal and structural characteristics of TV programs with visual, auditory, and textual information. In terms of visual cues, we develop a novel concept named program-oriented informative images (POIM) to identify the candidate points correlated with the boundaries of individual programs. For audio cues, a multiscale Kullback-Leibler (K-L) distance is proposed to locate audio scene changes (ASC), and accordingly ASC is aligned with video scene changes to represent candidate boundaries of programs. In addition, latent semantic analysis (LSA) is adopted to calculate the textual content similarity (TCS) between shots to model the inter-program similarity and intra-program dissimilarity in terms of speech content. Finally, we fuse the multimodal features of POIM, ASC, and TCS to detect the boundaries of programs including individual commercials (spots). Towards effective program guide and attracting content browsing, we propose a multimodal representation of individual programs by using POIM images, key frames, and textual keywords in a summarization manner. Extensive experiments are carried out over an open benchmarking dataset TRECVID 2005 corpus and promising results have been achieved. Compared with the electronic program guide (EPG), our solution provides a more generic approach to determine the exact boundaries of diverse TV programs even including dramatic spots. Jinqiao Wang, Ling-Yu Duan, Qingshan Liu 0001, Hanqing Lu, Jesse S. Jin |
IEEE Trans. Multim. | 3 |
| 2007 | Image Segmentation Using Co-EM Strategy
Zhenglong Li 0001, Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu |
ACCV (2) | 3 |
| 2007 | Face Mis-alignment Analysis by Multiple-Instance Subspace
Qingshan Liu 0001, Dimitris N. Metaxas |
ACCV (2) | 2 |
| 2007 | Topology-Preserved Diffusion Distance for Histogram ComparisonabstractIn most previous works, histograms are simply treated as n-dimensional arrays or even reshaped into vectors when measuring the distances between them. However many histograms have their intrinsic topologies, such as HSV histogram (cone), shape context (polar), orientation histogram (circle). The topologies are important for so-called cross-bin distance, because they determine the similarities between histogram bins, and influence the crossbin distances between histograms. In this paper, we proposed the topologypreserved diffusion distance to take the topology into account. This method extracts the distance by measuring the heat diffusion process defined on the topology of the histogram. Moreover, a fast implementation with time complexity O(N) is developed. Experiments on image retrieval and interest point matching show the effectiveness and efficiency of the proposed method. 1 Wang Yan, Qiqi Wang 0001, Qingshan Liu 0001, Hanqing Lu, Songde Ma |
BMVC | 3 |
| 2007 | Boosting Coded Dynamic Features for Facial Action Units and Facial Expression RecognitionabstractIt is well known that how to extract dynamical features is a key issue for video based face analysis. In this paper, we present a novel approach of facial action units (AU) and expression recognition based on coded dynamical features. In order to capture the dynamical characteristics of facial events, we design the dynamical haar-like features to represent the temporal variations of facial events. Inspired by the binary pattern coding, we further encode the dynamic haar-like features into binary pattern features, which are useful to construct weak classifiers for boosting learning. Finally the Adaboost is performed to learn a set of discriminating coded dynamic features for facial active units and expression recognition. Experiments on the CMU expression database and our own facial AU database show its encouraging performance. Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas |
CVPR | 2 |
| 2007 | A Component Based Deformable Model for Generalized Face AlignmentabstractThis paper presents a component based deformable model for generalized face alignment, in which a novel bi-stage statistical framework is proposed to account for both local and global shape characteristics. Instead of using statistical analysis on the entire shape as in previous alignment work, we build separate Gaussian models for shape components to preserve more detailed local shape deformations. In each model of components the Markov Network is integrated to provide simple geometry constraints for our search strategy. In order to make a better description of the nonlinear interrelationships over the shape components, the Gaussian process latent variable model is adopted to obtain enough control of full range shape variations. Furthermore, we propose an illumination-robust feature to lead the local fitting of every shape point when light conditions change dramatically. Based on this approach, our system can generate optimal shape for images with exaggerated expressions and under variable illumination, as evidenced by extensive experimentation. Yuchi Huang, Qingshan Liu 0001, Dimitris N. Metaxas |
ICCV | 2 |
| 2007 | Image Annotation Refinement using NSC-Based Word CorrelationabstractImage annotation refinement is crucial to improve the performance of automatic image annotation, in which the estimation of word correlation is a key issue. Typically, the word co-occurrence information may be utilized to estimate the word correlation. However, this approach is not accurate enough because it equally treats any word pair co-occurring in the training data and cannot extract synonymy relationship effectively. In this paper, a novel method is developed to estimate the word correlation based on the improved nearest spanning chains (NSC). It can extract more informative and reasonable relations among keywords. Obtaining the enhanced word correlation, a word-based graph is constructed, which is used to re-rank the candidate annotations for an untagged image. Experiments conducted on the typical Corel dataset demonstrate the effectiveness of the proposed method. Jing Liu 0001, Mingjing Li, Qingshan Liu 0001, Hanqing Lu, Songde Ma |
ICME | 3 |
| 2007 | Robust Commercial Retrieval in Video StreamsabstractTV commercial video is a kind of informative medium. To fast and robustly index and retrieve commercial videos is of interest to commercial monitor, copyright protection, and commercial management, we propose a coarse-to-fine scheme to robustly retrieve commercial videos. Different from previous work using clip or key frames-based matching, our scheme has incorporated the commercial production knowledge to search the candidate commercial positions. Color and ordinal features are extracted for locating the exact commercial positions with dynamic time warping distance. Comparison experiments were carried out over TRECVID 2006 news videos and some videos from Chinese channels. Our scheme has achieved promising simulation results. Jinqiao Wang, Ling-Yu Duan, Qingshan Liu 0001, Hanqing Lu, Jesse S. Jin |
ICME | 3 |
| 2007 | Facial Expression Recognition using Encoded Dynamic FeaturesabstractIn this paper, we propose a new approach of facial expression recognition. In order to capture the temporal characteristic of facial expressions, we design dynamic haar-like features to represent the facial images, and code them into binary patterns for the further analysis. Based on the encoded features, Adaboost is employed to learn the combination of optimal discriminant features to construct the classifier. The experiments carried on the CMU database show the promising performance of the proposed method. Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas |
ICME | 2 |
| 2007 | A New Multimedia Message Customizing Framework for mobile DevicesabstractIn this paper, we present a novel framework to customize multimedia messages for mobile users. The goal is to generate a video message from a series of pictures. The framework includes visual attention view detection, image grouping, image ranking, and slideshow generation. Considering the limitation of mobile device, we use a simple color feature based attention model to detect interesting regions of the images. We group the images, and rank them based on the attention view similarities. Finally a human perception based slideshow is designed to keep the mobile users' eye on attention regions efficiently. In addition, a short music is selected to match the video message. Extensive experiments and user studies show the promising performance of the proposed system. Cunxun Zang, Qingshan Liu 0001, Hanqing Lu, Kongqiao Wang |
ICME | 2 |
| 2007 | Registration of Lung Tissue Between Fluoroscope and CT Images: Determination of Beam Gating Parameters in Radiotherapy
Sukmoon Chang, Jinghao Zhou, Qingshan Liu 0001, Dimitris N. Metaxas, Bruce G. Haffty, Sung N. Kim, Salma J. Jabbour, Ning J. Yue |
MICCAI (1) | 3 |
| 2007 | Automatic TV Logo Detection, Tracking and Removal in Broadcast Video
Jinqiao Wang, Qingshan Liu 0001, Ling-Yu Duan, Hanqing Lu, Changsheng Xu |
MMM (2) | 2 |
| 2007 | An improved variable-size block-matching algorithm
Qingshan Liu 0001, Hanqing Lu |
Multim. Tools Appl. | 2 |
| 2006 | A Geometric Contour Framework with Vector Field Support
Zhenglong Li 0001, Qingshan Liu 0001, Hanqing Lu |
ACCV (2) | 2 |
| 2006 | Boosting Multi-gabor Subspaces for Face Recognition
Qingshan Liu 0001, Hongliang Jin, Xiaoou Tang, Hanqing Lu, Songde Ma |
ACCV (1) | 1 |
| 2006 | Automatic Moving Object Segmentation with Accurate Boundaries
Qingshan Liu 0001, Hanqing Lu |
ACCV (1) | 3 |
| 2006 | Fast Global Motion Estimation Via Iterative Least-Square Method
Qingshan Liu 0001, Hanqing Lu |
ACCV (2) | 3 |
| 2006 | Multiple Similarities Based Kernel Subspace Learning for Image Classification
Wang Yan, Qingshan Liu 0001, Hanqing Lu, Songde Ma |
ACCV (2) | 2 |
| 2006 | Web Image Mining Based on Modeling Concept-Sensitive Salient RegionsabstractIn this paper, we propose a probabilistic model for Web image mining, which is based on concept-sensitive salient regions without human intervene. Our goal is to achieve a middle-level understanding of image semantics to bridge the semantic gap existing in the field of image mining and retrieval. With the help of a popular search engine, semantically relevant images are collected, and concept-sensitive salient regions are extracted automatically based on an attention model. Then the semantic concept model is learned from the joint distribution of all salient regions with Gaussian mixture model and expectation-maximization algorithm. In addition, by incorporating semantically irrelevant un-salient regions as negative samples, the discriminative power of the solution is further enhanced. Experiments demonstrate the encouraging performance of the proposed method Jing Liu 0001, Qingshan Liu 0001, Jinqiao Wang, Hanqing Lu, Songde Ma |
ICME | 2 |
| 2006 | Fast Progressive Model Refinement Global Motion Estimation Algorithm with PredictionabstractGlobal motion estimation (GME) is an important part in the object-based applications. In this paper, a fast progressive model refinement (FPMR) GME algorithm is proposed. It can select the appropriate motion model according to the complexity of the camera motion. Two techniques are used to accelerate the procedure of FPMR. The first is an outlier prediction based feature point selection method. It can predict outliers from that of the last frame and therefore can effectively remove the influence of outliers on parameter calculation. The second is an intermediate-level model prediction method, which is used to fast the model selection and the parameter calculation procedure. Experiments show that the proposed algorithm is above two times faster than that of the feature-based fast and robust GME technique Qingshan Liu 0001, Hanqing Lu |
ICME | 3 |
| 2006 | Dynamic Similarity Kernel for Visual Recognition
Wang Yan, Qingshan Liu 0001, Hanqing Lu, Songde Ma |
KES (2) | 2 |
| 2006 | Ensemble learning for independent component analysis
Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001 |
Pattern Recognit. | 2 |
| 2006 | Face recognition using kernel scatter-difference-based discriminant analysisabstractThere are two fundamental problems with the Fisher linear discriminant analysis for face recognition. One is the singularity problem of the within-class scatter matrix due to small training sample size. The other is that it cannot efficiently describe complex nonlinear variations of face images because of its linear property. In this letter, a kernel scatter-difference-based discriminant analysis is proposed to overcome these two problems. We first use the nonlinear kernel trick to map the input data into an implicit feature space F. Then a scatter-difference-based discriminant rule is defined to analyze the data in F. The proposed method can not only produce nonlinear discriminant features but also avoid the singularity problem of the within-class scatter matrix. Extensive experiments show encouraging recognition performance of the new algorithm. Qingshan Liu 0001, Xiaoou Tang, Hanqing Lu, Songde Ma |
IEEE Trans. Neural Networks | 1 |
| 2005 | A Nonlinear Approach for Face Sketch Synthesis and RecognitionabstractMost face recognition systems focus on photo-based face recognition. In this paper, we present a face recognition system based on face sketches. The proposed system contains two elements: pseudo-sketch synthesis and sketch recognition. The pseudo-sketch generation method is based on local linear preserving of geometry between photo and sketch images, which is inspired by the idea of locally linear embedding. The nonlinear discriminate analysis is used to recognize the probe sketch from the synthesized pseudo-sketches. Experimental results on over 600 photo-sketch pairs show that the performance of the proposed method is encouraging. Qingshan Liu 0001, Xiaoou Tang, Hongliang Jin, Hanqing Lu, Songde Ma |
CVPR (1) | 1 |
| 2005 | Learning Local Descriptors for Face DetectionabstractIn this paper, we propose a realtime face detection approach based on local structure and texture of the objects in gray-level images. Our strategy is to map the local spatial structures and image textures of face class into binary patterns, and use these binary patterns as local descriptors. Boosting based face detector is constructed using these local descriptors, and cascade scheme is employed to further improve the efficiency of the face detector. Compared to the existing face detection approaches, our proposed method has two advantages: (1) it is robust to illumination changes to some extend, for the features use the information of local relationship instead of the original gray values; (2) the computational cost is very low, both in training procedure and evaluation step. The experimental results show that the proposed method can meet the demand of realtime applications with a satisfied detection performance. Hongliang Jin, Qingshan Liu 0001, Xiaoou Tang, Hanqing Lu |
ICME | 2 |
| 2005 | AIRE: an ambient interactive and responsive environment for mobile image managementabstractThis paper proposes an ambient interactive and responsive environment (AIRE) to improve user's image browsing experiences on the small-form-factor devices. This solution is characterized as two aspects: (1) establishing an ambient interactive communication across various devices; and (2) designing a distributed interface that crosses various devices to overcome the display constraint in mobile devices. In the AIRE system, a two-level image browsing scheme is designed to meet users' various image browsing needs. Zhigang Hua, Xiang-Jun Wang, Xing Xie 0001, Qingshan Liu 0001, Hanqing Lu, Wei-Ying Ma |
Mobile HCI | 4 |
| 2005 | Semantic knowledge extraction and annotation for web imagesabstractNowadays, images have become widely available on the World Wide Web (WWW). It's essential to develop effective ways for managing and retrieving such abundant images. Advantageously, compared to the traditional images where very little information is provided, the web images contain plentiful context data. This paper introduces a system that can automatically acquire semantic knowledge for web image annotation. By using a page layout analysis method that can precisely assign context to web images, we developed efficient algorithms to extract semantic knowledge for web images, such as description, people, temporal and geographic information. To validate the practicality and efficiency of this system, we applied it to about 6,500 images crawled from Web. Experiments demonstrated that our approach could achieve satisfactory results. Zhigang Hua, Xiang-Jun Wang, Qingshan Liu 0001, Hanqing Lu |
ACM Multimedia | 3 |
| 2005 | Providing on-demand sports video to mobile devicesabstractThis paper introduces a system for providing on-demand sports video to mobile devices, which has two main contributions. First, we construct an infrastructure for extracting and delivering the highlights instead of the whole sport videos to mobile clients, which can significantly reduce the bandwidth consumption. Second, we design an advanced UI for the mobile clients to effectively browse and interact with the video highlights. To validate the practicality and effectiveness of this system, we conduct the experiments on several real soccer videos. The results demonstrated that more than 65% of bandwidth consumption could be reduced. Moreover, the initial user study results show that the mobile users could interact effectively with the interface to seek or navigate sports videos. Qingshan Liu 0001, Zhigang Hua, Cunxun Zang, Xiaofeng Tong, Hanqing Lu |
ACM Multimedia | 1 |
| 2005 | Highlight ranking for sports video browsingabstractSports video has been extensively studied for its wide viewer-ship and tremendous commercial potentials. Many studies focused on highlight extraction for summarizing a lengthy video. In this paper, we present an advanced highlight analysis system for sports video browsing, in which highlight evaluation and ranking are concerned besides highlight detection. First, we use replay detection to efficiently localize the highlights. Then incorporating with domain-specific knowledge, we adopt several significant cues to evaluate the importance degree of the highlights with support vector regression. Finally, the highlights are ranked with descending sort according to their importance value. The ranking results can provide a hierarchical video browsing and customized content delivery scheme. Initial experimental results on soccer videos show an encouraging performance comparing with human subjective evaluation. Xiaofeng Tong, Qingshan Liu 0001, Yifan Zhang 0001, Hanqing Lu |
ACM Multimedia | 2 |
| 2005 | Supervised kernel locality preserving projections for face recognition
Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001 |
Neurocomputing | 2 |
| 2004 | Face detection using improved LBP under Bayesian frameworkabstractIn this paper, we present a novel face detection approach using improved local binary patterns (ILBP) as facial representation. ILBP feature is an improvement of LBP feature that considers both local shape and texture information instead of raw grayscale information and it is robust to illumination variation. We model the face and non-face class using multivariable Gaussian model and classify them under Bayesian framework. Extensive experiments show that the proposed method has an encouraging performance. Hongliang Jin, Qingshan Liu 0001, Hanqing Lu, Xiaofeng Tong |
ICIG | 2 |
| 2004 | Robust color-based trackingabstractColor as a distinct feature is widely used for object representation and tracking. However, color-based tracking is often influenced by clutter background and illumination variation. This paper presents a robust color-based tracking method, in which robust color feature is extracted for constructing the observation model under the modified particle filter tracking framework. The object is represented by its dominant color, and the weighted histogram with spatial information of the dominant color is used to optimize object models. In the particle filter framework, an extended iterated likelihood weighting scheme is employed to utilize more valuable particles. The experimental results show it is a real-time robust tracker, and it can obtain more than 30 fps with 2.4 G CPU and 512 MRAM. Qingshan Liu 0001, Hanqing Lu |
ICIG | 2 |
| 2004 | Replay detection in broadcasting sports videoabstractAn effective replay detection method for broadcasting sports videos is proposed in this paper. The algorithm is composed of logos detection and replay recognition. Firstly, a logo-template is automatically extracted from logo-transition at the beginning of a video. It is further used to detect the other logos in the same video. Then, with taking logos as boundaries, the video can be divided into replay and non-replay segments. To identify replay segments, an SVM classifier with mid-level features is utilized. Experiments demonstrate the effectiveness of this method. Xiaofeng Tong, Hanqing Lu, Qingshan Liu 0001, Hongliang Jin |
ICIG | 3 |
| 2004 | A supervised nonlinear local embedding for face recognitionabstractMany recent works demonstrated that subspace analysis is a good method for face recognition. How to find the subspace is a key issue. In this paper, a supervised nonlinear local embedding (SNLE) method is proposed to construct a subspace for face recognition, in which we combine the idea of nonlinear kernel mapping and preserving local geometric relations of the samples belonging to same class. SNLE can not only gain a perfect approximation of the nonlinear face manifold, but also enhance within-class local information. Moreover, it is also equivalent to solving a generalized eigenvalue problem in mathematics. Our experiments are performed on two benchmarks, and experimental results show that the proposed method has an encouraging performance. Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001 |
ICIP | 2 |
| 2004 | Semantic units based events detection in soccer videos
Xiaofeng Tong, Qingshan Liu 0001, Hanqing Lu |
ICIP | 2 |
| 2004 | Feature space analysis using low-order tensor voting
Hanqing Lu, Qingshan Liu 0001 |
ICIP | 3 |
| 2004 | A three-layer event detection framework and its application in soccer videoabstractA three-layer event detection scheme is proposed and applied to shoot and card event detection in soccer videos. At the lowest layer low-level features including color texture, edge and motion are considered. At the middle layer a concept of semantic unit is presented to bridge the semantic gap. The semantic unit is a sequence of consecutive frames tagged a special semantic cue. At the highest layer, taking the semantic units as observations, a Bayesian inference-based probabilistic framework is used to reason on the presence of interesting events. The experimental results of shoot and red/yellow card event detection demonstrate that the scheme is effective and the semantic units extraction is reliable. Xiao-Feng Tong, Hanqing Lu, Qingshan Liu 0001 |
ICME | 3 |
| 2004 | Random Independent Subspace for Face Recognition
Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001 |
KES | 2 |
| 2004 | Improving ICA Performance for Modeling Image Appearance with the Kernel Trick
Qingshan Liu 0001, Jian Cheng 0001, Hanqing Lu, Songde Ma |
KES | 1 |
| 2004 | Improving kernel Fisher discriminant analysis for face recognitionabstractThis work is a continuation and extension of our previous research where kernel Fisher discriminant analysis (KFDA), a combination of the kernel trick with Fisher linear discriminant analysis (FLDA), was introduced to represent facial features for face recognition. This work makes three main contributions to further improving the performance of KFDA. First, a new kernel function, called the cosine kernel, is proposed to increase the discriminating capability of the original polynomial kernel function. Second, a geometry-based feature vector selection scheme is adopted to reduce the computational complexity of KFDA. Third, a variant of the nearest feature line classifier is employed to enhance the recognition performance further as it can produce virtual samples to make up for the shortage of training samples. Experiments have been carried out on a mixed database with 125 persons and 970 images and they demonstrate the effectiveness of the improvements. Qingshan Liu 0001, Hanqing Lu, Songde Ma |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Kernel-Based Nonlinear Discriminant Analysis for Face Recognition
Qingshan Liu 0001, Rui Huang 0001, Hanqing Lu, Songde Ma |
J. Comput. Sci. Technol. | 1 |
| 2002 | Head Tracking Using Shapes and Adaptive Color Histograms
Qingshan Liu 0001, Songde Ma, Hanqing Lu |
J. Comput. Sci. Technol. | 1 |