VLDB 2026 Research / reviewers in the wild / expert
Qiang Wu 0001
dblp:87/2533-1
· DBLP profile ↗
200ranked-venue papers
4as first author
79since 2021 · last 2026
0000-0001-5641-2483ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 135 · 3 first-author · 47 since 2021Artificial intelligence and machine learning · 63 · 30 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 since 2021Computer networks · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Security and privacy · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PromptHG: Prompt-Enhanced Heterogeneous Graph for Personalized News Recommendation
Hai-Dang Kieu, Delvin Ce Zhang, Qiang Wu 0001, Min Xu 0001, Dung D. Le |
ECIR (1) | 4 |
| 2026 | MEGG: replay via maximally extreme GGscore in incremental learning for neural recommendation modelsabstractAbstract Neural collaborative filtering (NCF)-based recommendation models have been widely adopted in practical recommender systems due to their effectiveness. However, these models are typically developed under the static deep learning paradigm, where training is conducted on fixed datasets with the implicit assumption of a static data distribution. This approach is ill-suited for dynamic environments, such as those encountered in real-world platforms, where user preferences and collaborative filtering patterns evolve continuously. To address this limitation, incremental learning-a paradigm designed to integrate new knowledge while preserving previously learned information-emerges as a promising alternative. Despite its potential, the direct application of conventional incremental learning methods, which are prevalent in domains like computer vision and natural language processing, is hindered by unique challenges in recommender systems. These include the distinct task paradigm, data complexity, and sparsity issues. Moreover, existing incremental learning approaches tailored for neural recommendation models remain scarce and often suffer from limited generalizability. To bridge this gap, we propose an innovative experience replay-based incremental learning framework specifically designed for neural recommendation models, termed Replay Samples with Maximally Extreme GGscore (MEGG). At the core of MEGG is a novel metric, the GGscore, which quantifies the influence of individual samples on model training. By selectively replaying samples with the most extreme GGscores, our method effectively mitigates catastrophic forgetting, thereby maintaining high predictive performance over time. A key advantage of MEGG lies in its data-centric nature, which renders it agnostic to the underlying model architecture. This ensures broad applicability across various neural recommendation models and seamless integration with existing incremental learning frameworks to further enhance performance. Extensive experiments conducted on three neural recommendation models across four benchmark datasets demonstrate the superior effectiveness of MEGG compared to state-of-the-art methods. Furthermore, additional evaluations highlight its scalability, efficiency, and robustness. The implementation of MEGG will be made publicly available upon acceptance. Yunxiao Shi, Shuo Yang 0006, Haimin Zhang 0001, Li Wang 0064, Yongze Wang, Qiang Wu 0001, Min Xu 0001 |
Data Min. Knowl. Discov. | 6 |
| 2026 | Wrinkles in time: Multi-scale patching and super-resolution for efficient time series forecasting
Yuwei Chen 0007, Wenjing Jia, Qiang Wu 0001 |
Neurocomputing | 3 |
| 2026 | FedPCL-CDR: A federated prototype-based contrastive learning framework for privacy-preserving cross-domain recommendation
Li Wang 0064, Qiang Wu 0001, Min Xu 0001 |
Neural Networks | 2 |
| 2026 | GaitADIB: Adversarial disentangled information bottleneck network for unseen-view gait recognition
Hanyue Du, Xianye Ben, Zunxiao Xu, Lei Chen 0095, Qiang Wu 0001 |
Pattern Recognit. | 6 |
| 2026 | Tensor completion via Tucker decomposition with Correlated Total Variation regularization on factor matrices
Min Wang 0022, Zhuying Chen, Qiang Wu 0001 |
Signal Process. | 3 |
| 2026 | Weakly Supervised Composed Object Re-Identification With Large ModelsabstractExisting object re-identification (re-ID) and composed image retrieval (CIR) methods capture different aspects of real-world retrieval requirements; re-ID preserves identity but cannot specify desired appearance changes, whereas CIR supports attribute-guided retrieval but does not enforce identity consistency. To bridge this gap, we introduce composed object re-identification (CORI), a new task that requires the retrieved target to simultaneously satisfy identity preservation and text-guided attribute modification. This problem is fundamentally different from existing re-ID and CIR settings and has not been explicitly studied before. To make CORI feasible without costly manual annotation, we propose a weakly supervised framework that leverages large language models (LLMs) and visual question answering (VQA) models to automatically generate reference-to-target descriptions using ID labels alone. We further develop the first baseline model tailored for CORI, which jointly learns multimodal composition and identity-aware matching through shared-weight image encoders, a text encoder, and a compositor module optimized by contrastive, ID, and triplet losses. We also establish four CORI benchmark datasets covering person and vehicle retrieval. Experiments show that the proposed method consistently outperforms representative baselines adapted from existing CIR and re-ID methods for the newly introduced CORI setting, improving Rank@1 by 2.1% and 0.8% on RAP and Celeb-reID-light, and by 9.9% and 9.5% on VeRi-776 and VRIC, respectively. Jie Lu 0001, Qiang Wu 0001, Guangquan Zhang 0001 |
IEEE Trans. Cybern. | 3 |
| 2026 | SeeGait: Synergistic Co-Evolving Representations for Multimodal Gait Recognition via Hierarchical Multi-Stage FusionabstractGait recognition offers non-contact, long-distance identification but struggles with robustness against covariates like clothing variations, carrying conditions, and viewpoint changes. Existing methods predominantly rely on single modalities (e.g., silhouettes or skeletons) or employ shallow multimodal fusion, such as simple concatenation, which treats modalities as independent and static, failing to exploit their complementary strengths, shape cues from silhouettes and structural kinematics from skele-tons. To address these limitations, we introduce the Synergistic co-evolving representations (See) principle, enabling modalities to iteratively interact, guide, and refine each other across semantic hierarchies, fostering a unified, robust identity representation resilient to complex environments. This is realized through SeeGait, a novel multimodal framework featuring hierarchical multi-stage fusion. At its core, the Bidirectional Hierarchical Cross-Attention Synergy Module (BiHCASM) employs adaptive cross-modal attention to dynamically align and reweight features bidirectionally, allowing structural insights to enhance appearance focus and vice versa. Complementing this, the Hierarchical Spatiotemporal Transformer Encoder (HSTE) captures long-range skeleton dynamics, overcoming GCN limitations, while the Hierarchical Convolutional Silhouette Encoder (HCSE) extracts multi-scale silhouette pyramids for rich shape priors. Finally, a Holistic Feature Aggregation (HFA) strategy consolidates features from all stages for deep supervision, ensuring comprehensive optimization. By promoting mutual refinement, SeeGait mitigates covariate disruptions through enhanced complementarity, yielding superior discriminability. Extensive experiments show state-of-the-art performance, with 97.1% average Rank-1 accuracy on CASIA-B, and top results on CCPG and SUSTech1K. Hanyue Du, Xianye Ben, Xiankai Lu, Zunxiao Xu, Qingshuo Gao, Qiang Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | A Self-Supervised Diffusion Model With Edge Prior for Unpaired LDCT DenoisingabstractLow-dose computed tomography (LDCT) reduces health risks from radiation exposure but introduces imaging noise and artifacts. While numerous studies have employed deep learning for LDCT image denoising, the field continues to face significant challenges. Recent advancements have seen diffusion models applied to overcome issues of over-smoothness and unstable training inherent in prior deep learning approaches. However, the diffusion models face challenges in direct practical applications due to the extensive sampling steps, significant inference time required, and the need for hard-to-obtain paired data during training. To address these difficulties, this paper introduces a self-supervised diffusion model with edge prior for unpaired LDCT denoising. This method enables denoising within a lower-dimensional space, reducing computational complexity. Our proposed approach enhances denoised image clarity by applying prior edge constraints to compressed encodings; it employs a noise-conditioned encoding strategy to facilitate self-supervised image training, enabling the method to be applicable to unpaired CT data; and it utilizes compressed LDCT encoding as intermediate sampling results during the inference process, thereby accelerating sampling and reducing the time required for inference, making the method more real-time capable. Extensive validation across multiple datasets demonstrates that our method achieves competitive performance against state-of-the-art approaches in terms of peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and perceptual quality (LPIPS), while maintaining a practically acceptable inference time. Zhen Zhang 0057, Huizhen Zhang, Shaohua Zheng, Liqin Huang, Qiang Wu 0001, Xiahai Zhuang, Mingdian Yu |
IEEE J. Biomed. Health Informatics | 7 |
| 2026 | Classifier Enhancement Using Extended Context and Domain Experts for Semantic SegmentationabstractPrevalent semantic segmentation methods generally adopt a vanilla classifier to categorize each pixel into specific classes. Although such a classifier learns global information from the training data, this information is represented by a set of fixed parameters (weights and biases). However, each image has a different class distribution, which prevents the classifier from addressing the unique characteristics of individual images. At the dataset level, class imbalance leads to segmentation results being biased towards majority classes, limiting the model's effectiveness in identifying and segmenting minority class regions. In this paper, we propose an Extended Context-Aware Classifier (ECAC) that dynamically adjusts the classifier using global (dataset-level) and local (image-level) contextual information. Specifically, we leverage a memory bank to learn dataset-level contextual information of each class, incorporating the class-specific contextual information from the current image to improve the classifier for precise pixel labeling. Additionally, a teacher-student network paradigm is adopted, where the domain expert (teacher network) dynamically adjusts contextual information with ground truth and transfers knowledge to the student network. Comprehensive experiments illustrate that the proposed ECAC can achieve state-of-the-art performance across several datasets, including ADE20K, COCO-Stuff10K, and Pascal-Context. Huadong Tang, Youpeng Zhao 0002, Min Xu 0001, Jun Wang 0001, Qiang Wu 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | A Multi-Modal Prompt-Tuning Framework for Non-Overlapping Multi-Domain RecommendationabstractCross-domain recommendation (CDR) aims to enhance recommendation accuracy in data-sparse domains by transferring knowledge from data-rich domains. Most existing CDR methods conduct knowledge transfer based on overlapping users or items to address the user cold-start problems, including few-shot (i.e., users with sparse interactions) and zero-shot (i.e., users with no interactions) scenarios. However, in real-world scenarios, such overlap is often sparse or non-existent, limiting the effectiveness of these approaches. To overcome this challenge, we propose a novelMulti-modalPrompt-tuningFramework forNon-overlappingMulti-DomainRecommendation (MPF-NMDR). MPF-NMDR transfers knowledge across non-overlapping domains, enhancing recommendation performance in both few-shot and zero-shot scenarios. Specifically, we first pre-train the MPF-NMDR framework on data from all domains to capture users' generalized cross-domain preferences, which are learned through the generalized multi-modal interest mining module. We then conduct prompt-tuning with domain, user, and item prompts in the target domain to capture distinctions among various domains, users, and items. In this process, only the prompt parameters are fine-tuned, while all other parameters remain frozen, enabling the model to capture the distinctions among domains, users, and items while preserving the cross-domain knowledge. Extensive experiments on Amazon and Douban review datasets validate the superior performance of MPF-NMDR compared to SOTA baselines. We release our code athttps://github.com/Lili1013/MPF_NMDR. Li Wang 0064, Shoujin Wang, Qiang Wu 0001, Min Xu 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | Beyond KAN: Introducing KarSein for Adaptive High-Order Feature Interaction Modeling in CTR PredictionabstractModeling high-order feature interactions is crucial for Click-Through Rate (CTR) prediction, yet traditional approaches typically predefine a maximum interaction order and exhaustively enumerate feature combinations up to that order. This paradigm depends heavily on prior domain knowledge to delimit the interaction space and incurs substantial computational overhead. As a result, conventional CTR models face a persistent tension between enriching representations with complex high-order interactions and keeping computation tractable. To address this dual challenge, this study introduces the Kolmogorov–Arnold Represented Sparse Efficient Interaction Network (KarSein). Drawing inspiration from the learnable activation mechanism in the Kolmogorov–Arnold Network (KAN), KarSein leverages this mechanism to adaptively transform low-order basic features into high-order feature interactions, offering a novel approach to feature interaction modeling. KarSein extends the capabilities of KAN by introducing a more efficient architecture that significantly reduces computational costs while accommodating 2D embedding vectors as feature inputs. Furthermore, it overcomes the limitation of KAN’s its inability to spontaneously capture multiplicative relationships among features. Extensive experiments highlight the superiority of KarSein, demonstrating its ability to surpass not only the vanilla implementation of KAN in CTR prediction tasks but also other baseline methods. Remarkably, KarSein achieves exceptional predictive accuracy while maintaining a highly compact parameter size and minimal computational overhead. Moreover, KarSein retains the key advantages of KAN, such as strong interpretability and structural sparsity. As the first systematic adaptation of KAN to CTR prediction, KarSein offers a practical, parameter-efficient, and interpretable alternative for modeling complex feature interactions in large-scale recommendation systems. Yunxiao Shi, Wujiang Xu, Haimin Zhang 0001, Qiang Wu 0001, Min Xu 0001 |
ACM Trans. Inf. Syst. | 4 |
| 2026 | PECC: Position Encoding Coordinate Classification System Design for Human Pose EstimationabstractCoordinate classification is an efficient approach to 2-D human pose estimation (HPE), treating keypoint predictions as sub-pixel bins along horizontal and vertical axes, thereby avoiding the computationally intensive upsampling process required in traditional heatmap-based methods. In this article, we introduce the Position Encoding Coordinate Classification (PECC) system, which enhances coordinate classification by embedding position information directly into keypoint feature representations through a novel position encoding mechanism. We further design a tailored attention mechanism, Filtering Amplified Attention (FAA), optimized for coordinate classification. FAA provides finer relative positional information, improving the system’s ability to model relationships between keypoints and enhancing coordinate localization accuracy. Our method maintains the efficiency of coordinate classification by utilizing 1-D vectors, significantly reducing model parameters and computational cost. Additionally, the incorporation of positional encoding enhances the system’s ability to effectively model and exploit spatial information within a coordinate-classification-based pose estimation framework. Extensive experiments on mainstream datasets demonstrate that PECC achieves superior accuracy and robustness in 2-D HPE, advancing the state-of-the-art in this domain. Tao Zhang 0010, Qiang Wu 0001, Yeh-Cheng Chen, Naixue Xiong |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2026 | Volume Feature Aware View-Epipolar Transformers for Generalizable NeRFabstractGeneralizable NeRF synthesizes novel views of unseen scenes without per-scene training. The view-epipolar transformer has become popular in this field for its ability to produce high-quality views. Existing methods with this architecture rely on the assumption that texture consistency across views can identify object surfaces, with such identification crucial for determining where to reconstruct texture. However, this assumption is not always valid, as different surface positions may share similar texture features, creating ambiguity in surface identification. To handle this ambiguity, this paper introduces 3D volume features into the view-epipolar transformer. These features contain geometric information, which will be a supplement to texture features. By incorporating both texture and geometric cues in consistency measurement, our method mitigates the ambiguity in surface detection. This leads to more accurate surfaces and thus better novel view synthesis. Additionally, we propose a decoupled decoder where volume and texture features are used for density and color prediction respectively. In this way, the two properties can be better predicted without mutual interference. Experiments show improved results over existing transformer-based methods on both real-world and synthetic datasets. Ping An 0001, Xinpeng Huang, Qiang Wu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | MeRino: Entropy-Driven Design for Generative Language Models on IoT DevicesabstractGenerative Large Language Models (LLMs) stand as a revolutionary advancement in the modern era of artificial intelligence (AI). However, scaling down LLMs for resource-constrained hardware, such as Internet-of-Things (IoT) devices requires non-trivial efforts and domain knowledge. In this paper, we propose a novel information-entropy framework for designing mobile-friendly generative language models. The whole design procedure involves solving a mathematical programming (MP) problem, which can be done on the CPU within minutes, making it nearly zero-cost. We evaluate our designed models, termed MeRino, across fourteen NLP downstream tasks, showing their competitive performance against the state-of-the-art autoregressive transformer models under the mobile setting. Notably, MeRino achieves similar or better performance on both language modeling and zero-shot learning tasks, compared to the 350M parameter OPT while being 4.9x faster on NVIDIA Jetson Nano with 5.5x reduction in model size. Youpeng Zhao 0002, Huadong Tang, Qiang Wu 0001, Jun Wang 0001 |
AAAI | 4 |
| 2025 | Multilingual Model Enhancement Framework using a Human-Centered Approach for Arabic Spam DetectionabstractArabic spam detection remains a critical challenge in cybersecurity, due to the complexity of language and inadequate resources compared to those available for English. This research introduces a human-centered framework for Arabic spam classification, integrating behavioral insights from phishing vulnerability studies with advanced machine learning models. Building on our previous work for student phishing awareness and behavioral patterns, we have developed customized translation workflows and enhanced state-of-the-art detection techniques through the integration of human factors. Our enhanced models demonstrate a significant improvement in classification accuracy and a reduction in false positive rates. The results indicate that incorporating human perceptual elements not only bolsters technical performance but also enhances the real-world effectiveness of Arabic spam detection systems. This approach effectively bridges the gap between technical capability and practical deployment, providing a more robust solution for Arabic-language cybersecurity applications. Saleh Alqahtani, Priyadarsi Nanda, Qiang Wu 0001, Raddad Faqihi, Bashair Alrashed |
AICCSA | 3 |
| 2025 | Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute RecognitionabstractPedestrian attribute recognition (PAR) seeks to predict multiple semantic attributes associated with a specific pedestrian. There are two types of approaches for PAR: unimodal framework and bimodal framework. The former one is to seek a robust visual feature. However, the lack of exploiting semantic feature of linguistic modality is the main concern. The latter one utilizes prompt learning techniques to integrate linguistic data. However, static prompt templates and simple bimodal concatenation cannot to capture the extensive intra-class attribute variability and support active modalities collaboration. In this paper, we propose an Enhanced Visual-Semantic Interaction with Tailored Prompts (EVSITP) framework for PAR. We present an Image-Conditional Dual-Prompt Initialization Module (IDIM) to adaptively generate context-sensitive prompts from visual inputs. Subsequently, a Prompt Enhanced and Regularization Module (PERM) is proposed to strengthen linguistic information from IDIM. We further design a Bimodal Mutual Interaction Module (BMIM) to ensure bidirectional modalities communication. In addition, existing PAR datasets are collected over a short period in limited scenarios, which do not align with real-world scenarios. Therefore, we annotate a long-term person re-identification dataset to create a new PAR dataset, Celeb-PAR. Experiments on several challenging PAR datasets show that our method outperforms state-of-the-art approaches. Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001 |
CVPR | 6 |
| 2025 | ANNs-SaDE: A Machine-Learning-Based Design Automation Framework for Microwave Branch-Line CouplersabstractThe traditional method for designing branch-line couplers involves a trial-and-error optimization process that requires multiple design iterations through electromagnetic (EM) simulations. Thus, it is extremely time consuming and labor intensive. In this paper, a novel machine-learning-based framework is proposed to tackle this issue. It integrates artificial neural networks with a self-adaptive differential evolution algorithm (ANNs-SaDE). This framework enables the self-adaptive design of various types of microwave branch-line couplers by precisely optimizing essential electrical properties, such as coupling factor, isolation, and phase difference between output ports. The effectiveness of the ANNs-SaDE framework is demonstrated by the designs of folded single-stage branch-line couplers and multi-stage wideband branch-line couplers. Qiang Wu 0001, Li Yang 0011, Roberto Gómez-García, Xi Zhu 0001 |
ISCAS | 3 |
| 2025 | CSFRNet: Integrating Clothing Status Awareness for Long-Term Person Re-identification
Yan Huang 0008, Yan Huang 0023, Zhang Zhang 0001, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Causal disentanglement for regulating social influence bias in social recommendation
Li Wang 0064, Min Xu 0001, Quangui Zhang, Yunxiao Shi, Qiang Wu 0001 |
Neurocomputing | 5 |
| 2025 | Rethinking attention mechanism for enhanced pedestrian attribute recognitionabstractPedestrian Attribute Recognition (PAR) plays a crucial role in various computer vision applications, demanding precise and reliable identification of attributes from pedestrian images. Traditional PAR methods, though effective in leveraging attention mechanisms, often suffer from the lack of direct supervision on attention, leading to potential overfitting and misallocation. This paper introduces a novel and model-agnostic approach, Attention-Aware Regularization (AAR), which rethinks the attention mechanism by integrating causal reasoning to provide direct supervision of attention maps. AAR employs perturbation techniques and a unique optimization objective to assess and refine attention quality, encouraging the model to prioritize attribute-specific regions. Our method demonstrates significant improvement in PAR performance by mitigating the effects of incorrect attention and fostering a more effective attention mechanism. Experiments on standard datasets showcase the superiority of our approach over existing methods, setting a new benchmark for attention-driven PAR models. Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001 |
Neurocomputing | 6 |
| 2025 | Codebook prior-guided hybrid attention dehazing networkabstractTransformers have been widely used in image dehazing tasks due to their powerful self-attention mechanism for capturing long-range dependencies. However, directly applying Transformers often leads to coarse details during image reconstruction, especially in complex real-world hazy scenarios. To address this problem, we propose a novel Hybrid Attention Encoder (HAE). Specifically, a channel-attention-based convolution block is integrated into the Swin-Transformer architecture. This design enhances the local features at each position through an overlapping block-wise spatial attention mechanism while leveraging the advantages of channel attention in global information processing to strengthen the network’s representation capability. Moreover, to adapt to various complex hazy environments, a high-quality codebook prior encapsulating the color and texture knowledge of high-resolution clear scenes is introduced. We also propose a more flexible Binary Matching Mechanism (BMM) to better align the codebook prior with the network, further unlocking the potential of the model. Extensive experiments demonstrate that our method consistently outperforms the second-best methods by a margin of 8% to 19% across multiple metrics on the RTTS and URHI datasets. The source code has been released at https://github.com/HanyuZheng25/HADehzeNet . Liqin Huang, Hanyu Zheng, Zhipeng Su, Qiang Wu 0001 |
Image Vis. Comput. | 5 |
| 2025 | Camera-aware Embedding Refinement for unsupervised person re-identification
Yimin Liu 0001, Meibin Qi, Yongle Zhang 0001, Wenbo Xu 0004, Qiang Wu 0001 |
Knowl. Based Syst. | 5 |
| 2025 | High-order diversity feature learning for pedestrian attribute recognitionabstractPedestrian attribute recognition (PAR) involves accurately identifying multiple attributes present in pedestrian images. There are two main approaches for PAR: part-based method and attention-based method. The former relies on existing segmentation or region detection methods to localize body parts and learn corresponding attribute-specific feature from the corresponding regions, where the performance heavily depends on the accuracy of body region localization. The latter adopts the embedded attention modules or transformer attention to exploit detailed feature. However, it can focus on certain body regions but often provide coarse attention, failing to capture fine-grained details, the learned feature may also be interfered with by irrelevant information. Meanwhile, these methods overlook the global contextual information. This work argues for replacing coarse attention with detailed attention and integrating it with global contextual feature from ViT to jointly represent attribute-specific regions. To tackle this issue, we propose a High-order Diversity Feature Learning (HDFL) method for PAR based on ViT. We utilize a polynomial predictor to design an Attribute-specific Detailed Feature Exploration (ADFE) module, which can construct the high-order statistics and gain more fine-grained feature. Our ADFE module is a parameter-friendly method that provides flexibility in deciding its utilization during the inference phase. A Soft-redundancy Perception Loss (SPLoss) is proposed to adaptively measure the redundancy between feature of different orders, which can promote diverse characterization of features. Experiments on several PAR datasets show that our method achieves a new state-of-the-art (SOTA) performance. On the most challenging PA100K dataset, our method outperforms previous SOTA by 1.69% and achieves the highest mA of 84.92%. Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001 |
Neural Networks | 6 |
| 2025 | Learning Comprehensive Representation via Selective Activation and Dual-Level Orthogonality for Pedestrian Attribute RecognitionabstractMulti-label Pedestrian Attribute Recognition (PAR) involves identifying a series of semantic attributes in person images. Existing PAR solutions typically rely on CNN as the backbone network to extract pedestrian features. Unfortunately, CNNs process only one adjacent region at a time, resulting in the disappearance of long-range relations between different attribute-specific regions. To address this limitation, we adopt the Vision Transformer (ViT) instead of CNN as the backbone for PAR, aiming to build long-range relations and extract more robust features. However, PAR suffers from an inherent attribute imbalance issue, causing ViT to naturally focus more on attributes that appear frequently in the training set and ignore some pedestrian attributes that appear less. The native features extracted by ViT are not able to tolerate the imbalance attribute distribution issue. To tackle this issue, we propose a novel component and a dual-level loss: the Selective Feature Activation Method (SFAM), the Orthogonal Feature Activation Loss (OFALoss), and Orthogonal Weight Regularization Loss (OWRLoss). SFAM smartly suppresses the more informative attribute-specific features, thus compelling the PAR model to pay greater attention to attribute-specific regions that are often overlooked. The proposed OFALoss enforces an orthogonal constraint on the original feature extracted by ViT and the suppressed features from SFAM, promoting the comprehensiveness of feature representation in each attribute-specific region. Furthermore, OWRLoss is employed for decreasing correlations among entries of the last shared classification layer, which can alleviate the highly correlated of weight vectors caused by non-uniform distribution. This can prevent excessive mutual interference among different attributes during attribute recognition. Our model-agnostic approach is plug-and-play, requiring no additional training parameters in the training process. We conduct experiments on several benchmark PAR datasets, including PETA, PA100K, RAPv1, and RAPv2, demonstrating the effectiveness of our method. Specifically, our method outperforms existing state-of-the-art approaches. Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001, Jianqiang Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Consistent Image Inpainting With Pre-Perception and Cross-Perception Collaborative ProcessesabstractIt has been proven that introducing multiple guidance sources boosts image inpainting performance. However, existing methods primarily focus on local relationships and neglect the holistic interplay between guidance and texture information. Moreover, they lack an effective feedback mechanism to adaptively update the guidance process as corrupted texture information is progressively restored, potentially resulting in inconsistent inpainting. To tackle this issue, we propose a novel scheme aligned with pre-perception and cross-perception collaborative processes in human drawing. To mimic the pre-perception process, we introduce a pre-perceptual transformer block that captures long-range contextual dependencies and activates meaningful information to individually optimize image structures, semantic layouts, and textures, thereby effectively controlling their respective generation. To mimic the cross-perception collaborative process, we propose a cyclic cross-perceptual interaction to maintain consistency across the entire image regarding structure, layout, and texture while progressively refining their details. This interaction accounts for the global attention relationship between texture and other guidance sources (including image structure and semantic layout) to enhance image texture, alongside integrating a dedicated feedback mechanism to update guidance information. The proposed components are alternately deployed in three-branch decoders of the new scheme from rough to fine-grained levels to achieve these two iterative processes of human drawing. Experimental results prove the superiority of the proposed scheme over state-of-the-art methods across three datasets. Yongle Zhang 0001, Yimin Liu 0001, Hao Fan 0004, Ruotong Hu, Jian Zhang 0002, Qiang Wu 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | Federated User Preference Modeling for Privacy-Preserving Cross-Domain RecommendationabstractCross-domain recommendation (CDR) aims to address the data-sparsity problem by transferring knowledge across domains. Existing CDR methods generally assume that the user-item interaction data is shareable between domains, which leads to privacy leakage. Recently, some privacy-preserving CDR (PPCDR) models have been proposed to solve this problem. However, they primarily transfer simple representations learned only from user-item interaction histories, overlooking other useful side information, leading to inaccurate user preferences. Additionally, they transfer differentially private user-item interaction matrices or embeddings across domains to protect privacy. However, these methods offer limited privacy protection, as attackers may exploit external information to infer the original data. To address these challenges, we propose a novel Federated User Preference Modeling (FUPM) framework. In FUPM, first, a novel comprehensive preference exploration module is proposed to learn users' comprehensive preferences from both interaction data and additional data including review texts and potentially positive items. Next, a private preference transfer module is designed to first learn differentially private local and global prototypes, and then privately transfer the global prototypes using a federated learning strategy. These prototypes are generalized representations of user groups, making it difficult for attackers to infer individual information. Extensive experiments on four CDR tasks conducted on the Amazon and Douban datasets validate the superiority of FUPM over SOTA baselines. Li Wang 0064, Shoujin Wang, Quangui Zhang, Qiang Wu 0001, Min Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Hierarchical Multi-Prototype Discrimination: Boosting Support-Query Matching for Few-Shot SegmentationabstractFew-shot segmentation (FSS) aims at training a model on base classes with sufficient annotations and then tasking the model with predicting a binary mask to identify novel class pixels with limited labeled images. Mainstream FSS methods adopt a support-query matching paradigm that activates target regions of the query image according to their similarity with a single support class prototype. However, this prototype vector is inclined to overfit the support images, leading to potential under-matching in latent query object regions and incorrect mismatches with base class features in the query image. To address these issues, this study reformulates conventional single foreground prototype matching to a multi-prototype matching paradigm. In this paradigm, query features exhibiting high confidence with non-target prototypes will be categorized as background. Specifically, the target query features are drawn closer to the novel class prototype through a Masked Cross-Image Encoding (MCE) module and a Semantic Multi-prototype Matching (SMM) module is incorporated to collaboratively filter unexpected base class regions on multi-scale features. Furthermore, we devise an adaptive class activation map, termed target-aware class activation map (TCAM) to preserve semantically coherent regions that might be inadvertently suppressed under pixel-wise matching guidance. Experimental results on PASCAL-5$^{i}$and COCO-20$^{i}$datasets demonstrate the advantage of the proposed novel modules, with the holistic approach outperforming compared state-of-the-art methods. Wenbo Xu 0004, Huaxi Huang, Yongshun Gong, Litao Yu, Qiang Wu 0001, Jian Zhang 0002 |
IEEE Trans. Multim. | 5 |
| 2025 | FRFCNet: Feature Refinement and Flexible Concatenation for Object DetectionabstractThe state-of-the-art YOLO detection algorithms still suffer from the issue of redundant extraction of similar features during feature propagation, and the simplistic stacking approach of connecting different features limits the flexibility of feature fusion. We propose a new feature recombination mechanism involving refining feature extraction and flexible concatenation. It includes the HFConv (Hybrid Flexibility Convolution) module, the MFD (Multivariate Flexibility Downsampling) module, and the DFSPP (Deformable and Flexible Spatial Pyramid Pooling) module. Specifically, the HFConv module employs feature refinement and flexible connection strategies to optimize feature representation and reduce redundancy in a dynamic way, acquiring diverse feature information from local and surrounding regions. The MFD module leverages multiple downsampling methods to address the issue of feature redundancy that may arise from a single downsampling method, thereby enhancing feature diversity. The DFSPP module learns an offset corresponding to the pooling kernel size, allowing for the extraction of the most critical information in a dynamic manner. By incorporating these modules into the YOLO architecture, we develop a more robust network called FRFCNet, and the experimental results show a notable 4.1% and 2.8% improvement in AP values on the VOC2012 and COCO2017 datasets, respectively, compared to the baseline (YOLOV7-Tiny-SiLu), outperforming current one-stage detectors. Tao Zhang 0010, Xiangjian He, Qiang Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | EMS: A Large-Scale Eye Movement Dataset, Benchmark, and New Model for Schizophrenia RecognitionabstractSchizophrenia (SZ) is a common and disabling mental illness, and most patients encounter cognitive deficits. The eye-tracking technology has been increasingly used to characterize cognitive deficits for its reasonable time and economic costs. However, there is no large-scale and publicly available eye movement dataset and benchmark for SZ recognition. To address these issues, we release a large-scale Eye Movement dataset for SZ recognition (EMS), which consists of eye movement data from 104 schizophrenics and 104 healthy controls (HCs) based on the free-viewing paradigm with 100 stimuli. We also conduct the first comprehensive benchmark, which has been absent for a long time in this field, to compare the related 13 psychosis recognition methods using six metrics. Besides, we propose a novel mean-shift-based network (MSNet) for eye movement-based SZ recognition, which elaborately combines the mean shift algorithm with convolution to extract the cluster center as the subject feature. In MSNet, first, a stimulus feature branch (SFB) is adopted to enhance each stimulus feature with similar information from all stimulus features, and then, the cluster center branch (CCB) is utilized to generate the cluster center as subject feature and update it by the mean shift vector. The performance of our MSNet is superior to prior contenders, thus, it can act as a powerful baseline to advance subsequent study. To pave the road in this research field, the EMS dataset, the benchmark results, and the code of MSNet are publicly available at https://github.com/YingjieSong1/EMS. Zhi Liu 0003, Gongyang Li, Qiang Wu 0001, Dan Zeng 0001, Lihua Xu, Tianhong Zhang, Jijun Wang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Attribute-Guided Pedestrian Retrieval: Bridging Person Re-ID with Internal Attribute VariabilityabstractIn various domains such as surveillance and smart retail, pedestrian retrieval, centering on person re-identification (Re-ID), plays a pivotal role. Existing Re-ID methodologies often overlook subtle internal attribute variations, which are crucial for accurately identifying individuals with changing appearances. In response, our paper introduces the Attribute-Guided Pedestrian Retrieval (AGPR) task, focusing on integrating specified attributes with query images to refine retrieval results. Although there has been progress in attribute-driven image retrieval, there remains a notable gap in effectively blending robust Re-ID models with intra-class attribute variations. To bridge this gap, we present the Attribute-Guided Transformer-based Pedestrian Retrieval (ATPR) framework. ATPR adeptly merges global ID recognition with local attribute learning, ensuring a co-hesive linkage between the two. Furthermore, to effectively handle the complexity of attribute interconnectivity, ATPR organizes attributes into distinct groups and applies both inter-group correlation and intra-group decorrelation regularizations. Our extensive experiments on a newly estab-lished benchmark using the RAP dataset [32] demonstrate the effectiveness of ATPR within the AGPR paradigm. Yan Huang 0023, Zhang Zhang 0001, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001 |
CVPR | 3 |
| 2024 | Enhancing Retrieval and Managing Retrieval: A Four-Module Synergy for Improved Quality and Efficiency in RAG SystemsabstractRetrieval-augmented generation (RAG) techniques leverage the in-context learning capabilities of large language models (LLMs) to produce more accurate and relevant responses. Originating from the simple ‘retrieve-then-read’ approach, the RAG framework has evolved into a highly flexible and modular paradigm. A critical component, the Query Rewriter module, enhances knowledge retrieval by generating a search-friendly query. This method aligns input questions more closely with the knowledge base. Our research identifies opportunities to enhance the Query Rewriter module to Query Rewriter+ by generating multiple queries to overcome the Information Plateaus associated with a single query and by rewriting questions to eliminate Ambiguity, thereby clarifying the underlying intent. We also find that current RAG systems exhibit issues with Irrelevant Knowledge; to overcome this, we propose the Knowledge Filter. These two modules are both based on the instruction-tuned Gemma-2B model, which together enhance response quality. The final identified issue is Redundant Retrieval; we introduce the Memory Knowledge Reservoir and the Retriever Trigger to solve this. The former supports the dynamic expansion of the RAG system’s knowledge base in a parameter-free manner, while the latter optimizes the cost for accessing external knowledge, thereby improving resource utilization and response efficiency. These four RAG modules synergistically improve the response quality and efficiency of the RAG system. The effectiveness of these modules has been validated through experiments and ablation studies across six common QA datasets. The source code can be accessed at https://github.com/Ancientshi/ERM4. Yunxiao Shi, Xing Zi, Zijing Shi, Haimin Zhang 0001, Qiang Wu 0001, Min Xu 0001 |
ECAI | 5 |
| 2024 | Task Consistent Prototype Learning for Incremental Few-Shot Semantic Segmentation
Wenbo Xu 0004, Yang Wang 0002, Qiang Wu 0001, Jian Zhang 0002 |
ICPR (23) | 5 |
| 2024 | Fine-scale deep learning model for time series forecastingabstractAbstract Time series data, characterized by large volumes and wide-ranging applications, requires accurate predictions of future values based on historical data. Recent advancements in deep learning models, particularly in the field of time series forecasting, have shown promising results by leveraging neural networks to capture complex patterns and dependencies. However, existing models often overlook the influence of short-term cyclical patterns in the time series, leading to a lag in capturing changes and accurately tracking fluctuations in forecast data. To overcome this limitation, this paper introduces a new method that utilizes an interpolation technique to create a fine-scaled representation of the cyclical pattern, thereby alleviating the impact of the irregularity in the cyclical component and hence enhancing prediction accuracy. The proposed method is presented along with evaluation metrics and loss functions suitable for time series forecasting. Experiment results on benchmark datasets demonstrate the effectiveness of the proposed approach in effectively capturing cyclical patterns and improving prediction accuracy. Yuwei Chen 0007, Wenjing Jia, Qiang Wu 0001 |
Appl. Intell. | 3 |
| 2024 | Unsupervised Point Cloud Representation Learning by Clustering and Neural RenderingabstractAbstract Data augmentation has contributed to the rapid advancement of unsupervised learning on 3D point clouds. However, we argue that data augmentation is not ideal, as it requires a careful application-dependent selection of the types of augmentations to be performed, thus potentially biasing the information learned by the network during self-training. Moreover, several unsupervised methods only focus on uni-modal information, thus potentially introducing challenges in the case of sparse and textureless point clouds. To address these issues, we propose an augmentation-free unsupervised approach for point clouds, named CluRender, to learn transferable point-level features by leveraging uni-modal information for soft clustering and cross-modal information for neural rendering. Soft clustering enables self-training through a pseudo-label prediction task, where the affiliation of points to their clusters is used as a proxy under the constraint that these pseudo-labels divide the point cloud into approximate equal partitions. This allows us to formulate a clustering loss to minimize the standard cross-entropy between pseudo and predicted labels. Neural rendering generates photorealistic renderings from various viewpoints to transfer photometric cues from 2D images to the features. The consistency between rendered and real images is then measured to form a fitting loss, combined with the cross-entropy loss to self-train networks. Experiments on downstream applications, including 3D object detection, semantic segmentation, classification, part segmentation, and few-shot learning, demonstrate the effectiveness of our framework in outperforming state-of-the-art techniques. Guofeng Mei, Cristiano Saltori, Elisa Ricci 0001, Nicu Sebe, Qiang Wu 0001, Jian Zhang 0002, Fabio Poiesi |
Int. J. Comput. Vis. | 5 |
| 2024 | CAA: Class-Aware Affinity calculation add-on for semantic segmentationabstractLeveraging contextual dependencies is a commonly used technique to enhance the performance of image segmentation. However, existing solutions do not effectively catch the class-level association between the pixels along the boundary across the objects of the different classes but focus more on the local pixel-to-pixel relation. This work proposes a Class-Aware Affinity module (CAA) that considers both pixel-to-pixel relation and pixel-to-class association. We try to argue that the pixel-to-pixel relations still catch the relation (e.g. similarity, attention, or affiliation) on the local texture level. At the same time, it should also consider the association between the pixel and the class context produced by the given image. Pixel-to-class association can best reveal the co-occurrent dependency on the semantic level between the given pixels and their nearby context. Such pixel-to-class association combined with the pixel-to-pixel relations aggregating the local texture information will best mitigate the confusion caused in the boundary regions across the objects of the different classes. Moreover, the proposed framework can serve as a generic add-on to be integrated with the existing image segmentation solution to boost the current performance. Equipped with CAA, we achieve promising performance against the existing work with 54.59% mIoU on ADE20K, 49.96% mIoU on COCO-Stuff10k, and 64.38% mIoU on Pascal-Context. Huadong Tang, Youpeng Zhao 0002, Chaofan Du, Min Xu 0001, Qiang Wu 0001 |
Knowl. Based Syst. | 5 |
| 2024 | A privacy-preserving framework with multi-modal data for cross-domain recommendationabstractCross-domain recommendation (CDR) aims to enhance the recommendation accuracy in a target domain with sparse data by leveraging rich information in a source domain, thereby addressing the data-sparsity problem. Some existing CDR methods highlight the advantages of extracting domain-common and domain-specific features to learn comprehensive user and item representations. However, these methods cannot effectively disentangle these components, as they often rely on simple user-item historical interaction information (such as ratings, clicks, and browsing), neglecting the rich multi-modal features. In addition, they do not protect user-sensitive data from potential leakage during knowledge transfer between domains. To address these challenges, we propose a P rivacy- P reserving Framework with M ulti- M odal Data for C ross- D omain R ecommendation, called P2M2-CDR. Specifically, we first design a multi-modal disentangled encoder that utilizes multi-modal information to disentangle more informative domain-common and domain-specific embeddings. Furthermore, we introduce a privacy-preserving decoder to mitigate user privacy leakage during knowledge transfer. Local differential privacy (LDP) is used to obfuscate disentangled embeddings before the inter-domain exchange, thereby enhancing privacy protection. To ensure both consistency and differentiation among these obfuscated disentangled embeddings, we incorporate contrastive learning-based domain-inter and domain-intra losses. Extensive experiments conducted on six CDR tasks from two real-world datasets demonstrate that P2M2-CDR outperforms other state-of-the-art single- and cross-domain baselines. The code is available at https://github.com/Lili1013/P2M2-CDR . Li Wang 0064, Lei Sang 0001, Quangui Zhang, Qiang Wu 0001, Min Xu 0001 |
Knowl. Based Syst. | 4 |
| 2024 | Customized meta-dataset for automatic classifier accuracy evaluation
Yan Huang 0023, Zhang Zhang 0001, Yan Huang 0008, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001 |
Pattern Recognit. | 4 |
| 2024 | Light Field Salient Object Detection With Sparse Views via Complementary and Discriminative Interaction Networkabstract4D light field data record the scene from multiple views, thus implicitly providing beneficial depth cue for salient object detection in challenging scenes. Existing light field salient object detection (LF SOD) methods usually use a large number of views to improve the detection accuracy. However, using so many views for LF SOD brings difficulties to its practical applications. Considering that adjacent views in a light field are actually with very similar contents, in this work, we propose defining a more efficient pattern of input views, i. e., key sparse views, and design a network to effectively explore the depth cue from sparse views for LF SOD. Specifically, we firstly introduce a low rank-based statistical analysis to the existing LF SOD datasets, which allows us to conclude a fixed yet universal pattern for our key sparse views, including the number and positions of views. These views maintain the sufficient depth cue, but greatly lower the number of views to be captured and processed, facilitating practical applications. Then, we propose an effective solution with a key Complementary and Discriminative Interaction Module (CDIM) for LF SOD from key sparse views, named CDINet. The CDINet follows a two-stream structure to extract the depth cue from the light field stream (i. e., sparse views) and the appearance cue from the RGB stream (i. e., center view), generating features and initial saliency maps for each stream. The CDIM is tailored for inter-stream interaction of both these features and saliency maps, using the depth cue to complement the missing salient regions in RGB stream and discriminate the background distraction, to enhance the final saliency map further. Extensive experiments on three LF multi-view datasets demonstrate that our CDINet not only outperforms the state-of-the-art 2D methods, but also achieves competitive performance as compared with the state-of-the-art 3D and 4D methods. The code and results of our method are available athttps://github.com/GilbertRC/LFSOD-CDINet. Gongyang Li, Ping An 0001, Zhi Liu 0003, Xinpeng Huang, Qiang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | GaitDAN: Cross-View Gait Recognition via Adversarial Domain AdaptationabstractView change causes significant differences in the gait appearance. Consequently, recognizing gait in cross-view scenarios is highly challenging. Most recent approaches either convert the gait from the original view to the target view before recognition is carried out or extract the gait feature irrelevant to the camera view through either brute force learning or decouple learning. However, these approaches have many constraints, such as the difficulty of handling unknown camera views. This work treats the view-change issue as a domain-change issue and proposes to tackle this problem through adversarial domain adaptation. This way, gait information from different views is regarded as the data from different sub-domains. The proposed approach focuses on adapting the gait feature differences caused by such sub-domain change and, at the same time, maintaining sufficient discriminability across the different people. For this purpose, a Hierarchical Feature Aggregation (HFA) strategy is proposed for discriminative feature extraction. By incorporating HFA, the feature extractor can well aggregate the spatial-temporal feature across the various stages of the network and thereby comprehensive gait features can be obtained. Then, an Adversarial View-change Elimination (AVE) module equipped with a set of explicit models for recognizing the different gait viewpoints is proposed. Through the adversarial learning process, AVE would not be able to identify the gait viewpoint in the end, given the gait features generated by the feature extractor. That is, the adversarial domain adaptation mitigates the view change factor, and discriminative gait features that are compatible with all sub-domains are effectively extracted. Extensive experiments on three of the most popular public datasets, CASIA-B, OULP, and OUMVLP richly demonstrate the effectiveness of our approach. Tianhuan Huang, Xianye Ben, Chen Gong 0002, Wenzheng Xu, Qiang Wu 0001, Hongchao Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Improving Consistency of Proxy-Level Contrastive Learning for Unsupervised Person Re-IdentificationabstractRecently, contrastive learning-based unsupervised person re-identification (Re-ID) methods have garnered significant attention due to their effectiveness. These methods rely on predicted pseudo-labels to construct contrastive pairs, optimizing the network gradually. Some methods also utilize camera labels to explore intra-camera and inter-camera contrastive relations, achieving state-of-the-art results. However, these methods fail to address the issue of inconsistency in proxy-level contrastive learning, which arises from variations in the distribution of instances belonging to the same proxy. Specifically, they are sensitive to the distribution of instances in a mini-batch used for contrastive pair construction, and uncertainty or noise in the data distribution can lead to turbulence in the contrastive loss, degrading the effectiveness of contrastive learning. In this work, we first propose a dual-branch contrastive learning (DBCL) framework. The framework comprises a dual-branch structure with an identity discrimination branch and a camera view awareness branch. These branches are mutually trained to produce a jointly optimized model with both high person identification accuracy and cross-camera robustness. Moreover, to mitigate the proxy-level contrastive inconsistency issue in the camera view awareness branch, we design intra-camera and inter-camera consistent contrastive losses. Our DBCL has been extensively evaluated on several person Re-ID datasets and has demonstrated superior performance compared to state-of-the-art methods. Notably, on the challenging MSMT17 dataset with complex scenes, our method achieved an mAP of 45.3% and Rank-1 accuracy of 75.3%. Yimin Liu 0001, Meibin Qi, Yongle Zhang 0001, Qiang Wu 0001, Jingjing Wu 0001, Shuo Zhuang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Enhancing Person Re-Identification Performance Through In Vivo LearningabstractThis research investigates the potential of in vivo learning to enhance visual representation learning for image-based person re-identification (re-ID). Compared to traditional self-supervised learning (which require external data), the introduced in vivo learning utilizes supervisory labels generated from pedestrian images to improve re-ID accuracy without relying on external data sources. Three carefully designed in vivo learning tasks, leveraging statistical regularities within images, are proposed without the need for laborious manual annotations. These tasks enable feature extractors to learn more comprehensive and discriminative person representations by jointly modeling various aspects of human biological structure information, contributing to enhanced re-ID performance. Notably, the method seamlessly integrates with existing re-ID frameworks, requiring minimal modifications and no additional data beyond the existing training set. Extensive experiments on diverse datasets, including Market1501, CUHK03-NP, Celeb-reID, Celeb-reid-light, PRCC, and LTCC, demonstrate substantial enhancements in rank-1 precision compared to state-of-the-art methods. Yan Huang 0008, Yan Huang 0023, Zhang Zhang 0001, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Meta Clothing Status Calibration for Long-Term Person Re-IdentificationabstractRecent studies have seen significant advancements in the field of long-term person re-identification (LT-reID) through the use of clothing-irrelevant or insensitive features. This work takes the field a step further by addressing a previously unexplored issue, the Clothing Status Distribution Shift (CSDS). CSDS refers to the differing ratios of samples with clothing changes to those without clothing changes between the training and test sets, leading to a decline in LT-reID performance. We establish a connection between the performance of LT-reID and CSDS, and argue that addressing CSDS can improve LT-reID performance. To that end, we propose a novel framework called Meta Clothing Status Calibration (MCSC), which uses meta-learning to optimize the LT-reID model. Specifically, MCSC simulates CSDS between meta-train and meta-test with meta-optimization objectives, optimizing the LT-reID model and making it robust to CSDS. This framework is designed to prevent overfitting and improve the generalization ability of the LT-reID model in the presence of CSDS. Comprehensive evaluations on seven datasets demonstrate that the proposed MCSC framework effectively handles CSDS and improves current state-of-the-art LT-reID methods on several LT-reID benchmarks. Yan Huang 0023, Qiang Wu 0001, Zhang Zhang 0001, Caifeng Shan, Yan Huang 0008, Yi Zhong 0002, Liang Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | PCL: Point Contrast and Labeling for Weakly Supervised Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation is a fundamental task in 3D scene understanding and has recently achieved remarkable progress. The success of existing approaches is attributed to recent advanced deep networks for point clouds and the availability of a large amount of labeled training data. However, creating such fully annotated training datasets for supervised point cloud semantic segmentation methods is a time-consuming and labor-intensive process, which increases the difficulty of extending supervised approaches to new application scenarios. To alleviate the data-hungry nature of deep learning, we propose PCL, the point contrast and labeling framework for weakly supervised point cloud semantic segmentation with small percentages of point-level annotations. The core idea of this method is to exploit contrastive learning to help learn a larger number of discriminative feature representations with limited annotations. By introducing two types of contrastive relationships, cross-sample point contrast and low-level similarity-based point contrast, our proposed framework can directly regularize the learned feature space, considering not only the low-level similarity within each point cloud but also the discriminative semantics within and across point clouds on both labeled and unlabeled points via pseudo labels. In addition, we propose a pseudo label refinery module to generate robust and reliable pseudo labels online, reducing the negative impact of incorrect pseudo labels. Our method achieves state-of-the-art performance on a diverse set of label-efficient semantic segmentation tasks. Anan Du, Tianfei Zhou, Shuchao Pang, Qiang Wu 0001, Jian Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | Mutual Dual-Task Generator With Adaptive Attention Fusion for Image InpaintingabstractImage segmentation can reveal the semantic structure information in an image, which is helpful guidance information for image inpainting. Notably, it can help mitigate the artifacts on the boundaries of different semantic regions during the inpainting process. Existing semantic guidance-based image inpainting provides one-way guidance from the semantic segmentation task to the image inpainting task. There is no feedback from the inpainting results to adjust the guidance process, which causes inferior performance. To tackle this issue, this work proposes mutual dual-task generators to establish the interaction between image segmentation and image inpainting tasks. Thus, semantic segmentation guides image inpainting and also receives feedback from image inpainting. These two processes interact with each other and progressively improve the inpainting quality. The mutual dual-task generator consists of a shared encoder and mutual decoders with the bidirectional Cross-domain Feature DeNormalization (CFDN) module inside, which hierarchically models the Segmentation-guided image Texture (ST) generation and Texture-guided semantic Segmentation (TS) generation. At the end of mutual decoders, an Adaptive Attention Fusion (AAF) module is proposed to augment the texture and semantic class affinity between pixels, further refining the inpainted results. Experimental results demonstrate that the proposed mutual dual-task generator pipeline achieves superior inpainting performances over the state of the arts on three public datasets. Yongle Zhang 0001, Yimin Liu 0001, Ruotong Hu, Qiang Wu 0001, Jian Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2023 | Unsupervised Deep Probabilistic Approach for Partial Point Cloud RegistrationabstractDeep point cloud registration methods face challenges to partial overlaps and rely on labeled data. To address these issues, we propose UDPReg, an unsupervised deep probabilistic registration framework for point clouds with partial overlaps. Specifically, we first adopt a network to learn posterior probability distributions of Gaussian mixture models (GMMs) from point clouds. To handle partial point cloud registration, we apply the Sinkhorn algorithm to predict the distribution-level correspondences under the constraint of the mixing weights of GMMs. To enable unsupervised learning, we design three distribution consistency-based losses: self-consistency, cross-consistency, and local contrastive. The self-consistency loss is formulated by encouraging GMMs in Euclidean and feature spaces to share identical posterior distributions. The cross-consistency loss derives from the fact that the points of two partially overlapping point clouds belonging to the same clusters share the cluster centroids. The cross-consistency loss allows the network to flexibly learn a transformation-invariant posterior distribution of two aligned point clouds. The local contrastive loss facilitates the network to extract discriminative local features. Our UDPReg achieves competitive performance on the 3DMatch/3DLoMatch and ModelNet/ModelLoNet benchmarks. Guofeng Mei, Hao Tang 0005, Xiaoshui Huang, Weijie Wang 0002, Juan Liu 0006, Jian Zhang 0002, Luc Van Gool, Qiang Wu 0001 |
CVPR | 8 |
| 2023 | Class-Aware Contextual Information for Semantic SegmentationabstractExploring spatial contextual information is a well-adopted approach to achieving better semantic segmentation performance. However, most existing methods neglect the class association between the neighboring pixels. In this paper, we propose a CACINet, which consists of a Semantic Affinity Module (SAM) and a Class Association Module (CAM), to generate class-aware contextual information among pixels on a fine-grained level. SAM analyzes the affiliation of any two given pixels belonging to the same or different class. It produces intra-class and inter-class pixel contextual information. CAM classifies the image into different class regions globally and then it encodes the pixel based on the degree of affiliation of the pixels with each class in the image. In this way, it augments the class affiliation of the pixels into the corresponding context calculation. Comprehensive experiments demonstrate that the proposed method achieves competitive performance on two semantic segmentation benchmarks: ADE20K and PASCAL-Context. Huadong Tang, Youpeng Zhao 0002, Yingying Jiang 0001, Zhuoxin Gan, Qiang Wu 0001 |
ICASSP | 5 |
| 2023 | Parameter-Efficient Vision Transformer with Linear AttentionabstractRecent advances in vision transformers (ViTs) have achieved outstanding performance in visual recognition tasks, including image classification and detection. ViTs can learn global representations with their self-attention mechanism, but they are usually heavy-weight and unsuitable for resource-constrained devices. In this paper, we propose a novel linear feature attention (LFA) module to reduce computation costs for vision transformers and combine efficient mobile CNN modules to form a parameter-efficient and high-performance CNN-ViT hybrid model, called LightFormer, which can serve as a general-purpose backbone to learn both global and local representation. Comprehensive experiments demonstrate that LightFormer achieves competitive performance across different visual recognition tasks. On the ImageNet-1K dataset, LightFormer achieves top-1 accuracy of 78.5% with 5.5 million parameters. Our model also performs well when transferred to object detection and semantic segmentation tasks. On the MS COCO dataset, LightFormer attains mAP of 33.2 within the YOLOv3 framework, and on the Cityscapes dataset, with only a simple all-MLP decoder, LightFormer achieves mIoU of 78.5 and FPS of 15.3, surpassing state-of-the-art lightweight segmentation networks. Youpeng Zhao 0002, Huadong Tang, Yingying Jiang 0001, Yong A, Qiang Wu 0001, Jun Wang 0001 |
ICIP | 5 |
| 2023 | Camera Proxy based Contrastive Learning with Hard Sampling for Unsupervised Person Re-identificationabstractBecause of the advantages of dealing with large-scale unlabelled data, unsupervised learning has recently attracted more attention for person re-identification. Particularly, the combination of the unsupervised learning paradigm with contrastive learning shows promising efficiency in network optimization. This work adopts the successful camera-aware contrastive learning approach and further explores its capability on the camera proxy level to improve the data pair consistency. Thus, it is more robust to the camera change, which still challenges the unsupervised person re-identification. This work proposed a Camera Proxy-based Contrastive Learning framework, which explicitly considers inter-camera scenario and intra-camera scenario. Moreover, this work is motivated by the strategy of selecting a hard negative sample in triplet loss learning and further extends it to contrastive learning for both negative and positive pair creation on the camera proxy level. Extensive experiments demonstrate the superiority of the proposed framework over state-of-the-art approaches on purely unsupervised re-identification. Yimin Liu 0001, Meibin Qi, Qiang Wu 0001, Yanfang Yang, Xiaohong Li 0002, Jian Zhang 0002 |
ICME | 3 |
| 2023 | Masked Cross-image Encoding for Few-shot SegmentationabstractFew-shot segmentation (FSS) is a dense prediction task that aims to infer the pixel-wise labels of unseen classes using only a limited number of annotated images. The key challenge in FSS is to classify the labels of query pixels using class prototypes learned from the few labeled support exemplars. Prior approaches to FSS have typically focused on learning class-wise descriptors independently from support images, thereby ignoring the rich contextual information and mutual dependencies among support-query features. To address this limitation, we propose a joint learning method termed Masked Cross-Image Encoding (MCE), which is designed to capture common visual properties that describe object details and to learn bidirectional inter-image dependencies that enhance feature interaction. MCE is more than a visual representation enrichment module; it also considers cross-image mutual dependencies and implicit guidance. Experiments on FSS benchmarks PASCAL-5iand COCO-20idemonstrate the advanced meta-learning ability of the proposed method. Wenbo Xu 0004, Huaxi Huang, Litao Yu, Qiang Wu 0001, Jian Zhang 0002 |
ICME | 5 |
| 2023 | Automated Flock Density and Activity Recognition for Welfare Monitoring on Commercial Egg FarmsabstractMonitoring poultry behaviour provides the opportunity to aid egg production and animal welfare. With the current development in machine learning and computer vision, automated content analysis has become a practical way for low-cost and continuous monitoring of animal behaviours. In this demo, we will show a simple yet effective flock monitoring system based on computer vision and machine learning techniques for egg farmers that allows them to reduce labour yet improve performance. This demo shows that it is possible to auto-analyse flock activities thereby providing early warning of welfare issues, by applying object detection, tracking and crowd-counting techniques. Summaries of individual bird activity and their distribution are closely related to the flock behaviour, which in turn reflects the welfare status. Specifically, the density and movement patterns of birds provide reliable information on the welfare status of the flock. For example, the real-time monitoring of density and movement can give early warnings of pile-ups. To observe these and other important flock activities, we developed a low-cost and easy-use system based on recent computer vision techniques to auto-estimate the density and movement of birds on commercial egg farms. Litao Yu, Wenbo Xu 0004, Qiang Wu 0001, Jian Zhang 0002 |
MMSP | 3 |
| 2023 | Guest Editorial: Learning from limited annotations for computer vision tasksabstractThe past decade has witnessed remarkable achievements in computer vision, owing to the fast development of deep learning. With the advancement of computing power and deep learning algorithms, we can process and apply millions or even hundreds of millions of large-scale data to train robust and advanced deep learning models. In spite of the impressive success, current deep learning methods tend to rely on massive annotated training data and lack the capability of learning from limited exemplars. However, constructing a million-scale annotated dataset like ImageNet is time-consuming, labour-intensive and even infeasible in many applications. In certain fields, very limited annotated examples can be gathered due to various reasons such as privacy or ethical issues. Consequently, one of the pressing challenges in computer vision is to develop approaches that are capable of learning from limited annotated data. The purpose of this Special Issue is to collect high-quality articles on learning from limited annotations for computer vision tasks (e.g. image classification, object detection, semantic segmentation, instance segmentation and many others), publish new ideas, theories, solutions and insights on this topic and showcase their applications. In this Special Issue we received 29 papers, all of which underwent peer review. Of the 29 originally submitted papers, 9 have been accepted. The nine accepted papers can be clustered into two main categories: theoretical and applications. The papers that fall into the first category are by Liu et al., Li et al. and He et al. The second category of papers offers a direct solution to various computer vision tasks. These papers are by Ma et al., Wu et al., Rao et al., Sun et al., Hou et al. and Gong et al. A brief presentation of each of the papers in this Special Issue follows. Liu et al. present a Gaussianisation prototypical classifier (GPC) for few-shot classification which mainly focuses on solving the issue of prototype bias. GPC consists of handling the features with the Gaussianisation operation and estimating a reliable prototype using the maximum a posteriori method using base class features as prior information. The proposed method is simple yet effective, which does not use any extra labelled data or knowledge. Moreover, it's also a one-step prototype rectification method, which does not resort any complex continuous optimisation. The ablation study shows that GPC can benefit from features pretrained only with CE loss or jointly trained with self-supervised loss. The results demonstrate that the proposed method outperforms related work and other state-of-the-art methods. Li et al. present a novel hyperspectral unmixing method named ‘Global centralised and Structured discriminative Nonnegative Matrix Factorisation (GSNMF)’. The proposed GSNMF offers several distinct advantages over the traditional unmixing techniques. Constructed on the foundation of the manifold regularisation techniques, GSNMF captures the intrinsic structural information by using the local affinity and distant repulsion constraints concurrently. With the structured discriminative information, local affinity constraint ensures that similar elements share similar estimated abundances, while the distant repulsion constraint ensures that dissimilar elements have different abundances. All experiments and analyses have demonstrated that the proposed GSNMF exhibits a remarkable performance compared to the other methods. He et al. present a taxonomy of existing algorithms in the task of makeup transfer. Evaluation methods are proposed, existing methods are analysed and existing datasets are reviewed. Finally the current problems in the field of makeup transfer are discussed, and the trend of future research is analysed. Ma et al. present a dense transformer framework for person re-identification tasks. This paper introduces densely connected class tokens to connect any two layers implicitly. The framework, Denseformer, outperforms other vision transformer models on four widely used benchmarks, namely Market-1501, DukeMTMC-reID, MSMT17 and Occluded-Duke datasets with only a small amount of extra calculation cost. According to the visualisation results, the proposed Denseformer pays more attention to the main parts of human bodies, obtaining discriminative global features. The Denseformer is a general improvement on ViT and works well on other tasks that use ViT as a backbone according to the promising results. Wu et al. present a new homology-continuous-based makeup transformation method, which can be roughly divided into two network branches: the age compensation branch and the makeup transformation branch. Specifically, in the age compensation branch, based on the same source continuity the authors designed a new encoding module which can map the face vector into the corresponding high-dimensional vector space and realise the compensation for age by adjusting the vector direction. In the makeup transformation branch, this work designed a multi-style encoder to handle different types of makeup, such as Chinese Japanese Korean makeup etc. In addition, the proposed network structure is a two-pass encoder-decoder architecture which has good parallelism and can achieve better results with training and inference on GPU. Rao et al. present a novel end-to-end architecture for point completion by using a stack-style folding network called the SSFN. Due to the fact that the output shape code cannot completely represent semantically, they propose a Stack-Style Folding module that transforms the bottleneck output into the style code analogous to StyleGAN. Experiments on ShapeNet and KITTI datasets indicate that the proposed SSFN architecture achieves a decent visual quality and metric performance. Sun et al. present a method for a zero-shot temporal event localisation (ZSTEL) that leverage large-scale video and language models, for example, CLIP. They solve the two key problems for ZSTEL: (1) how to find the relevant region where the event is likely to occur, (2) how to determine event duration after the relevant region is found. They propose the query-guided optimisation for local frame relevance. Relying on the query-to-frame relationship, this method can find the most relevant local frame region where the event is most likely to occur, guided by a constructed objective. The experimental results on the two standard benchmark datasets, Charades-STA and ActivityCaptions have shown the effectiveness of the proposed approach. Hou et al. present a cutting-edge few-shot detection method for logo images. To avoid the misclassification between the base and novel classes, they add an extra classification head. They also apply the convolutional layer into regression heads to improve the accuracy of location by using the limited training data. Considering the characteristics of logo images, they add balanced feature pyramid with Deformable RoI Pooling and unfreeze region proposal network in the fine-tuning stage. The extensive comparative experimentation and ablation studies illustrate the advantage of the proposed method and the effectiveness of every component in the model. Gong et al. present a method for object detection with a long-tail distribution that includes a dual-balanced network and balanced classification loss. This work investigates how the long-tailed distribution impacts the sub-networks in the general two-stage object detection framework Faster-RCNN and finds that unbalanced proposal sampling and unbalanced classification logic deteriorate the performance of the model in terms of AP. They propose the balanced region proposal network and balanced the classification network to address the above issues. Experiments on the LVIS-v0.5 dataset demonstrate that the framework improves the performance of AP without sacrificing too much from the performance of head categories in long-tail distribution. All of the papers selected for this Special Issue show that the field of learning from limited annotations for computer vision tasks is steadily moving forward. The possibility of a weakly supervised learning paradigm will remain a source of inspiration for new techniques in the years to come. Firstly, we wish to express our thanks to Ph.D. students at Nanjing University of Science and Technology for their continuous assistance throughout this process. Also, we wish to express our gratitude to all the contributors who submitted novel scientific results in this special issue and to the anonymous reviewers, whose expert work allowed the realisation of this endeavor. We aspire that this effort should contribute to the further development of DL and increase the concern of the scientific and technological community in the respective area. Last, we should not omit to express our appreciation to the journal's Editors-in-Chief and the Editorial Office for their support throughout this venture. Data sharing is not applicable to this article as no new data were created or analysed in this study. Yazhou Yao is a professor at the School of Computer Science and Engineering and Nanjing University of Science and Technology. With the support of the China Scholarship Council, he received his Ph.D. degree in Computer Science, University of Technology Sydney, Australia at 2018. From July 2018 to July 2019, he worked as a Research Scientist at the Inception Institute of Artificial Intelligence, Abu Dhabi, UAE. His research interests include multimedia processing and machine learning. Wenguan Wang is currently a ZJU100 Young Professor at Zhejiang University. He received his Ph.D. degree from Beijing Institute of Technology in 2018. From 2016 to 2018, he was a joint Ph.D. candidate at the University of California, Los Angeles. From 2018 to 2019, he was a senior scientist at the Inception Institute of Artificial Intelligence, UAE. From 2020 to 2022, he worked as a postdoc researcher at ETH Zurich, Switzerland. After that, he worked as a lecturer and ARC DECRA Fellow at the University of Technology Sydney. His current research interests include computer vision, image processing and deep learning. Qiang Wu received the BEng and MEng degrees in electronic engineering from the Harbin Institute of Technology, Harbin, China, in 1996 and 1998, respectively, and the Ph.D. degree in computing science from the University of Technology Sydney, Sydney, Australia, in 2004. He is currently an Associate Professor and a Core Member of the Global Big Data Technologies Centre, University of Technology Sydney. He has published more than 70 refereed papers, including those published in prestigious journals and top international conferences. His major research interests include computer vision, image processing, pattern recognition, machine learning and multimedia processing. He has served as the chair and/or a Programme Committee Member for a number of international conferences. Dongfang Liu is an Assistant Professor in the Department of Computer Engineering at the Rochester Institute of Technology (RIT). He earned his Ph.D. degree from Purdue University. Dr. Dongfang Liu's research focus on embodied AI and creates general AI solutions to address significant societal challenges. His ongoing work consists of: (1) developing attention-guided perception models that behave like a human's perpetual capacity; and (2) developing structured and human-centred recognition systems that comprehend the surrounding visual world. His publication portfolio includes papers from major conferences in the artificial intelligence and robotics fields, such as CVPR, ECCV, ICCV, ICLR, NIPS, ICML, AAAI, IJCAI, ACL, EMNLP, WWW, WACV, IROS etc. He currently serves on the senior programme committee for AAAI and IJCAI and as an associate editor for IEEE Transactions on Circuits and Systems for Video Technology (TCSVT). Jin Zheng received the BS and MS degrees from Liaoning Technical University, in 2001 and 2004, respectively, and the Ph.D. degree from the School of Computer Science and Engineering, Beihang University, in 2009. She joined the School of Computer Science and Engineering, Beihang University, in 2009. In 2014, she visited Harvard University, MA, USA, as a Visiting Scholar for 1 year. Her current research interests include object detection, tracking and recognition, among other similar interests. Yazhou Yao, Wenguan Wang, Qiang Wu 0001, Dongfang Liu |
IET Comput. Vis. | 3 |
| 2023 | Improving Disentangled Representation Learning for Gait Recognition Using Group SupervisionabstractIn decades, gait has been gathering extensive interest for the advantage that it can be measured from a distance without physical contact. However, for image/video-based gait recognition, its performance can be remarkably influenced by exterior factors, such as viewing angles and clothing changes. Thus, in this paper, a group-supervised disentangled representation learning network is proposed for gait recognition to extract features invariant to these factors. First, sequences are explicitly disentangled into pose, gait, appearance, and view features through a generic encoder-decoder framework. To ensure the feature adaptability and independency, a disentanglement swap module is specifically adopted during our encode-decoder process through a series of swap operations based on the feature attributes. Following the feature disentanglement, a disentanglement aggregation module is also specially proposed for pose, gait, and appearance features to enhance their effectiveness. Finally, the enhanced three features are concatenated together for gait recognition. Relevant experiments certify that compared with other disentangled representation learning-based gait recognition methods, our proposed method enables to obtain a more excellent recognition result, despite fewer gait frames being utilized. Lingxiang Yao, Worapan Kusakunniran, Peng Zhang 0057, Qiang Wu 0001, Jian Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2022 | Data Augmentation-free Unsupervised Learning for 3D Point Cloud Understanding
Guofeng Mei, Cristiano Saltori, Fabio Poiesi, Jian Zhang 0002, Elisa Ricci 0001, Nicu Sebe, Qiang Wu 0001 |
BMVC | 7 |
| 2022 | Unsupervised Point Cloud Pre-Training Via Contrasting and ClusteringabstractThe annotation for large-scale point clouds is still time-consuming and unavailable for many complex real-world tasks. Point cloud pre-training is a promising direction to auto-extract features without labeled data. Therefore, this paper proposes a general unsupervised approach, named ConClu for point cloud pre-training by jointly performing contrasting and clustering. Specifically, the contrasting is formulated by maximizing the similarity feature vectors produced by encoders fed with two augmentations of the same point cloud. The clustering simultaneously clusters the data while enforcing consistency between cluster assignments produced different augmentations. Experimental evaluations on downstream applications outperform state-of-the-art techniques, which demonstrates the effectiveness of our framework. Guofeng Mei, Xiaoshui Huang, Juan Liu 0006, Jian Zhang 0002, Qiang Wu 0001 |
ICIP | 5 |
| 2022 | Partial Point Cloud Registration Via Soft SegmentationabstractMost existing correspondence-free registration methods suffer from performance degradation in partial overlapped point clouds. To solve the partial overlapped point cloud registration, this paper proposes, SegReg, a soft Segmentation-based correspondence-free Registration approach. Specifically, we first softly segment both source and target point clouds into a discrete number of geometric partitions, respectively. Then registration is achieved through iteratively using the IC-LK algorithm to minimize the distance between the feature descriptors of the corresponded partitions. Extensive experiments on synthetic synthetic dataset ModelNet40 and real dataset 7Scene show that the proposed method achieves state-of-the-art performance. Guofeng Mei, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001 |
ICIP | 4 |
| 2022 | Overlap-Guided Coarse-to-Fine Correspondence Prediction for Point Cloud RegistrationabstractEstablishing reliable correspondences between a pair of point clouds is essential for registration with partial overlaps. However, existing correspondence estimation works usually struggle to distinguish the points in overlap and non-overlap regions. This paper thus proposes an Overlap-guided Coarse-to-Fine Network, named OCFNet, which first establishes correspondences at a coarse level and then refines them at a point level. Specifically, at the coarse level, our model first aggregates two point clouds into smaller sets of super-points with associated features and overlap scores, followed by establishing coarse-level correspondences between the two sets of super-points under the guidance of overlap scores. On the fine stage, a decoder recovers the raw points while jointly learning the associated features and overlap scores. Coarse-level proposals are then expanded to patches, and point-level correspondences are sequentially refined from the corresponding patches. We conducted comprehensive experiments on 3DMatch, 3DLoMatch, and KITTI benchmarks to show the effectiveness of the proposed method. [code] Guofeng Mei, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001 |
ICME | 4 |
| 2022 | Blockchain-Enabled Fish Provenance and Quality Tracking SystemabstractAccurate assessment of fish quality is difficult in practice due to the lack of trusted fish provenance and quality tracking information. Working with Sydney Fish Market (SFM), we develop a Blockchain-enabled fish provenance and quality tracking (BeFAQT) system. A multilayer Blockchain architecture based on attribute-based encryption (ABE) is proposed to tackle the privacy issue caused by applying Blockchain to secure supply chain data and achieve trusted and confidential data sharing among parties in fish supply chains. An Internet-of-Things (IoT) chain saves encrypted fish provenance and quality tracking data, and an ABE chain is specifically designed for the access control to the data in the IoT chain. Latest IoT and artificial intelligence (AI) technologies, including NarrowBand-IoT, image processing, and biosensing, are developed for fish origin proof, supply chain tracking, and objective fish quality assessment. As proven by field trials with SFM and a local fish supply chain, the BeFAQT is able to provide trusted and comprehensive fish provenance and quality tracking information in real time. Xu Wang 0004, Guangsheng Yu, Ren Ping Liu 0001, Jian Zhang 0002, Qiang Wu 0001, Steven W. Su, Ying He 0011, Zongjian Zhang, Litao Yu, Taoping Liu, Wentian Zhang, Peter Loneragan, Eryk Dutkiewicz, Erik Poole, Nick Paton |
IEEE Internet Things J. | 5 |
| 2022 | Enhanced Spatial-Temporal Salience for Cross-View Gait RecognitionabstractGait recognition can be used in person identification and re-identification by itself or in conjunction with other biometrics. Although gait has both spatial and temporal attributes, and it has been observed that decoupling spatial feature and temporal feature can better exploit the gait feature on the fine-grained level. However, the spatial-temporal correlations of gait video signals are also lost in the decoupling process. Direct 3D convolution approaches can retain such correlations, but they also introduce unnecessary interferences. Instead of common 3D convolution solutions, this paper proposes an integration of decoupling process into a 3D convolution framework for cross-view gait recognition. In particular, a novel block consisting of a Parallel-insight Convolution layer integrated with a Spatial-Temporal Dual-Attention (STDA) unit is proposed as the basic block for global spatial-temporal information extraction. Under the guidance of the STDA unit, this block can well integrate spatial-temporal information extracted by two decoupled models and at the same time retain the spatial-temporal correlations. In addition, a Multi-Scale Salient Feature Extractor is proposed to further exploit the fine-grained features through context awareness extension of part-based features and adaptively aggregating the spatial features. Extensive experiments on three popular gait datasets, namely CASIA-B, OULP and OUMVLP, demonstrate that the proposed method outperforms state-of-the-art methods. Tianhuan Huang, Xianye Ben, Chen Gong 0002, Baochang Zhang 0001, Qiang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | TOAN: Target-Oriented Alignment Network for Fine-Grained Image Categorization With Few Labeled SamplesabstractIn this paper, we study the fine-grained categorization problem under the few-shot setting, i.e., each fine-grained class only contains a few labeled examples, termed Fine-Grained Few-Shot classification (FGFS). The core predicament in FGFS is the high intra-class variance yet low inter-class fluctuations in the dataset. In traditional fine-grained classification, the high intra-class variance can be somewhat relieved by conducting the supervised training on the abundant labeled samples. However, with few labeled examples, it is hard for the FGFS model to learn a robust class representation with the significantly higher intra-class variance. Moreover, the inter- and intra-class variance are closely related. The significant intra-class variance in FGFS often aggravates the low inter-class variance issue. To address the above challenges, we propose a Target-Oriented Alignment Network (TOAN) to tackle the FGFS problem from both intra- and inter-class perspective. To reduce the intra-class variance, we propose a target-oriented matching mechanism to reformulate the spatial features of each support image to match the query ones in the embedding space. To enhance the inter-class discrimination, we devise discriminative fine-grained features by integrating local compositional concept representations with the global second-order pooling. We conducted extensive experiments on four public datasets for fine-grained categorization, and the results show the proposed TOAN obtains the state-of-the-art. Huaxi Huang, Junjie Zhang 0002, Litao Yu, Jian Zhang 0002, Qiang Wu 0001, Chang Xu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Collaborative Feature Learning for Gait Recognition Under Cloth ChangesabstractSince gait can be utilized to identify individuals from a far distance without their interaction and coordination, recently many gait recognition methods have been proposed. However, due to a real-world scenario of clothing changes, a degradation occurs for most of these methods. Thus in this paper, a more efficient gait recognition method is proposed to address the problem of clothing variances. First, part-based gait features are formulated from two different perspectives,i.e., the separated body parts that are more robust to clothing changes and the estimated human skeleton key-point regions. It is reasonable to formulate such features for cloth-changing gait recognition, because these two perspectives are both less vulnerable to clothing changes. Given that each feature has its own advantages and disadvantages, a more efficient gait feature is generated in this paper by assembling these two features together. Moreover, since local features are more discriminative than global features, in this paper more attention is focused on the local short-range features. Also, unlike most methods, in our method we treat the estimated key-point features as a set of word embeddings, and a transformer encoder is specifically used to learn the dependence of each correlative key-points. The robustness and effectiveness of our proposed method are certified by experiments on CASIA Gait Dataset B, and it has achieved the state-of-the-art performance on this dataset. Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Alleviating Modality Bias Training for Infrared-Visible Person Re-IdentificationabstractThe task of infrared-visible person re-identification (IV-reID) is to recognize people across two modalities (i.e., RGB and IR). Existing cutting-edge approaches normally use a pair of images that have the same IDs (i.e., ID-tied cross-modality image pairs) and input them into an ImageNet-trained ResNet50. The ResNet50 backbone model can learn shared features across modalities to tolerate modality discrepancies between RGB and IR. This work will unveil a Modality Bias Training (MBT) problem that is less discussed in IV-reID, which will demonstrate that MBT significantly compromises the performance of IV-reID. Due to MBT, IR information can be overwhelmed by RGB information during training when the ResNet50 model is pretrained based on a large amount of RGB images from ImageNet. Thus, the trained models are more inclined to RGB information. Accordingly, the cross-modality generalization ability of the model is also compromised. To tackle this issue, we present a Dual-level Learning Strategy (DLS) that 1) enforces the focus of the network on ID-exclusive (rather than ID-tied) labels of cross-modality image pairs to mitigate the problem of MBT and 2) introduces third modality data that contain both RGB and IR information to further prevent the information from the IR modality from being overwhelmed during training. Our third modality images are generated by a generative adversarial network. A dynamic ID-exclusive Smooth (dIDeS) label is proposed for the generated third modality data. In experiments, comprehensive experiments are carried out to demonstrate the success of DLS in tackling the MBT issue exposed in IV-reID. Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Peng Zhang 0057, Zhaoxiang Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Dual Attention on Pyramid Feature Maps for Image CaptioningabstractGenerating natural sentences from images is a fundamental learning task for visual-semantic understanding in multimedia. In this paper, we propose to apply dual attention on pyramid image feature maps to fully explore the visual-semantic correlations and improve the quality of generated sentences. Specifically, with the full consideration of the contextual information provided by the hidden state of the RNN controller, the pyramid attention can better localize the visually indicative and semantically consistent regions in images. On the other hand, the contextual information can help re-calibrate the importance of feature components by learning the channel-wise dependencies, to improve the discriminative power of visual features for better content description. We conducted comprehensive experiments on three well-known datasets: Flickr8K, Flickr30 K and MS COCO, which achieved impressive results in generating descriptive and smooth natural sentences from images. Using either convolution visual features or more informative bottom-up attention features, the composite model can boost the performance of image-to-sentence translation, with a limited computational resource overhead. The proposed pyramid attention and dual attention methods are highly modular, which can be inserted into various image captioning modules to further improve the performance. Litao Yu, Jian Zhang 0002, Qiang Wu 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Exploring Pairwise Relationships Adaptively From Linguistic Context in Image CaptioningabstractFor image captioning, recent works start to focus on exploring visual relationships for generating high-quality interactive words (i.e. verbs and prepositions). However, many existing works only focus on semantic level by analysing the feature similarity between objects in the visual domain but ignore the linguistic context included in the caption decoder. When captioning is being carried out, the entity words can be inferred based on visual information of objects. The interactive words representing the relationships between entity words can only be inferred based on high-level language meaning generated in the process of captioning decoding. Such high-level language meaning is called linguistic context, which refers to the relational context between words or phrases in the caption sentences. The linguistic context can be used as strong guidance to explore related visual relationships between different objects effectively. To achieve this, we propose a novel context-adaptive attention module that is strongly driven by the linguistic context from the caption decoder. In this module, a novel design of visual relationship attention is proposed based on a bilinear self-attention model to explore related visual relationships and encode more discriminative features under the linguistic context. To achieve the adaptive process of attending to related visual relationships for generating interactive words or related visual objects for entity words, an attention modulator is integrated as an attention channel controller responding to the changing linguistic context of the caption decoder dynamically. Experimented on MSCOCO dataset, our model achieves promising performances compared with all counterpart models that explore visual relationships. Zongjian Zhang, Qiang Wu 0001, Yang Wang 0002, Fang Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | MIG-Net: Multi-Scale Network Alternatively Guided by Intensity and Gradient Features for Depth Map Super-ResolutionabstractThe studies of previous decades have shown that the quality of depth maps can be significantly lifted by introducing the guidance from intensity images describing the same scenes. With the rising of deep convolutional neural network, the performance of guided depth map super-resolution is further improved. The variants always consider deep structure, optimized gradient flow and feature reusing. Nevertheless, it is difficult to obtain sufficient and appropriate guidance from intensity features without any prior. In fact, features in the gradient domain, e.g., edges, present strong correlations between the intensity image and the corresponding depth map. Therefore, the guidance in the gradient domain can be more efficiently explored. In this paper, the depth features are iteratively upsampled by 2×. In each upsampling stage, the low-quality depth features and the corresponding gradient features are iteratively refined by the guidance from the intensity features via two parallel streams. Then, to make full use of depth features in the image and gradient domains, the depth features and gradient features are alternatively complemented with each other. Compared with state-of-the-art counterparts, the sufficient experimental results show improvements according to the objective and subjective assessments. The code is available athttps://github.com/Yifan-Zuo/MIG-net-gradient_guided_depth_enhancement. Yifan Zuo 0001, Yuming Fang 0001, Xiaoshui Huang, Xiwu Shang, Qiang Wu 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | Recognizing Gaits Across Walking and Running SpeedsabstractFor decades, very few methods were proposed for cross-mode (i.e., walking vs. running) gait recognition. Thus, it remains largely unexplored regarding how to recognize persons by the way they walk and run. Existing cross-mode methods handle the walking-versus-running problem in two ways, either by exploring the generic mapping relation between walking and running modes or by extracting gait features which are non-/less vulnerable to the changes across these two modes. However, for the first approach, a mapping relation fit for one person may not be applicable to another person. There is no generic mapping relation given that walking and running are two highly self-related motions. The second approach does not give more attention to the disparity between walking and running modes, since mode labels are not involved in their feature learning processes. Distinct from these existing cross-mode methods, in our method, mode labels are used in the feature learning process, and a mode-invariant gait descriptor is hybridized for cross-mode gait recognition to handle this walking-versus-running problem. Further research is organized in this article to investigate the disparity between walking and running. Running is different from walking not only in the speed variances but also, more significantly, in prominent gesture/motion changes. According to these rationales, in our proposed method, we give more attention to the differences between walking and running modes, and a robust gait descriptor is developed to hybridize the mode-invariant spatial and temporal features. Two multi-task learning-based networks are proposed in this method to explore these mode-invariant features. Spatial features describe the body parts non-/less affected by mode changes, and temporal features depict the instinct motion relation of each person. Mode labels are also adopted in the training phase to guide the network to give more attention to the disparity across walking and running modes. In addition, relevant experiments on OU-ISIR Treadmill Dataset A have affirmed the effectiveness and feasibility of the proposed method. A state-of-the-art result can be achieved by our proposed method on this dataset. Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | PTN: A Poisson Transfer Network for Semi-supervised Few-shot LearningabstractThe predicament in semi-supervised few-shot learning (SSFSL) is to maximize the value of the extra unlabeled data to boost the few-shot learner. In this paper, we propose a Poisson Transfer Network (PTN) to mine the unlabeled information for SSFSL from two aspects. First, the Poisson Merriman–Bence–Osher (MBO) model builds a bridge for the communications between labeled and unlabeled examples. This model serves as a more stable and informative classifier than traditional graph-based SSFSL methods in the message-passing process of the labels. Second, the extra unlabeled samples are employed to transfer the knowledge from base classes to novel classes through contrastive learning. Specifically, we force the augmented positive pairs close while push the negative ones distant. Our contrastive transfer scheme implicitly learns the novel-class embeddings to alleviate the over-fitting problem on the few labeled data. Thus, we can mitigate the degeneration of embedding generality in novel classes. Extensive experiments indicate that PTN outperforms the state-of-the-art few-shot and SSFSL models on miniImageNet and tieredImageNet benchmark datasets. Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Qiang Wu 0001, Chang Xu 0002 |
AAAI | 4 |
| 2021 | Multi-Models Fusion for Light Field Angular Super-ResolutionabstractLight field (LF) imaging has received increasing attention due to its richer interpretation of the scene. However, an inherent spatial-angular trade-off exists in LF that prevents LF from practical applications. Consequently, how to break such a trade-off has become one of the main challenges in sparsely sampled LF reconstruction. LF super-resolution (SR) can provide an opportunity to solve this issue, but most methods exploit only one form of LF, thereby leading to much loss of information. We believe that different LF forms can compensate each other to obtain higher gains via fusion strategy. In this paper, therefore, we propose a multi-models fusion for LF SR in angular domain. Cascading models which are trained by different LF forms can fully exploit rich LF information. Experimental results demonstrate that our method is effective and achieves a comparable result against state-of-the-art techniques. Fengyin Cao, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Qiang Wu 0001 |
ICASSP | 5 |
| 2021 | Clothing Status Awareness for Long-Term Person Re-IdentificationabstractLong-Term person re-identification (LT-reID) exposes extreme challenges because of the longer time gaps between two recording footages where a person is likely to change clothing. There are two types of approaches for LT-reID: biometrics-based approach and data adaptation based approach. The former one is to seek clothing irrelevant biometric features. However, seeking high quality biometric feature is the main concern. The latter one adopts fine-tuning strategy by using data with significant clothing change. However, the performance is compromised when it is applied to cases without clothing change. This work argues that these approaches in fact are not aware of clothing status (i.e., change or no-change) of a pedestrian. Instead, they blindly assume all footages of a pedestrian have different clothes. To tackle this issue, a Regularization via Clothing Status Awareness Network (RCSANet) is proposed to regularize descriptions of a pedestrian by embedding the clothing status awareness. Consequently, the description can be enhanced to maintain the best ID discriminative feature while improving its robustness to real-world LT-reID where both clothing-change case and no-clothing-change case exist. Experiments show that RCSANet performs reasonably well on three LT-reID datasets. Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Zhaoxiang Zhang 0001 |
ICCV | 2 |
| 2021 | Unsupervised Domain Adaptation with Background Shift Mitigating for Person Re-Identification
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 2 |
| 2021 | Multilocation Human Activity Recognition via MIMO-OFDM-Based Wireless Networks: An IoT-Inspired Device-Free Sensing ApproachabstractDevice-free sensing (DFS) is an emerging technology that empowers wireless communication systems with the ability for not only data communication but also smart sensing. By taking advantage of machine-learning technologies, DFS transforms traditional wireless communication networks into intelligent context-aware networks and will open the doors for a myriad of promising 6G-enabled Internet of Things (IoT) applications, ranging from smart home to smart buildings. Although significant progress has been made for human activity recognition at a single location by leveraging this technology, performance at multiple locations has not been fully explored. As far as multilocation activity sensing is concerned, the performance is compromised along with the change of locations and labor-intensive annotation works caused by multilocation. To tackle this issue, an activity decomposition network (ActNet) is presented to decompose the activity information directly from input samples by using the training data from different locations together. Instead of dealing with different locations separately, our ActNet can assemble data from different locations together for training to mitigate the data limitation issue caused by a single location. To achieve this, a multiple-input–multiple-output (MIMO)-orthogonal frequency-division multiplexing (OFDM) technology-based prototype system is utilized to collect data samples at 24 different locations in a cluttered office environment. Especially, for each location, only ten samples of each activity are used for training. Experiments demonstrate that the average classification accuracy is 94.6% across all locations with ensured robustness produced by our method. Yi Zhong 0002, Ju Wang 0008, Siliang Wu, Ting Jiang 0008, Yan Huang 0023, Qiang Wu 0001 |
IEEE Internet Things J. | 6 |
| 2021 | Exploring region relationships implicitly: Image captioning with visual relationship attention
Zongjian Zhang, Qiang Wu 0001, Yang Wang 0002, Fang Chen 0001 |
Image Vis. Comput. | 2 |
| 2021 | Beyond modality alignment: Learning part-level representation for visible-infrared person re-identification
Peng Zhang 0057, Qiang Wu 0001, Xunxiang Yao, Jingsong Xu |
Image Vis. Comput. | 2 |
| 2021 | Robust gait recognition using hybrid descriptors based on Skeleton Gait Energy Image
Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang, Wankou Yang |
Pattern Recognit. Lett. | 3 |
| 2021 | Learning from EPI-Volume-Stack for Light Field image angular super-resolution
Deyang Liu, Qiang Wu 0001, Yan Huang 0023, Xinpeng Huang, Ping An 0001 |
Signal Process. Image Commun. | 2 |
| 2021 | Low-Rank Pairwise Alignment Bilinear Network For Few-Shot Fine-Grained Image ClassificationabstractDeep neural networks have demonstrated advanced abilities on various visual classification tasks, which heavily rely on the large-scale training samples with annotated ground-truth. However, it is unrealistic always to require such annotation in real-world applications. Recently, Few-Shot learning (FS), as an attempt to address the shortage of training samples, has made significant progress in generic classification tasks. Nonetheless, it is still challenging for current FS models to distinguish the subtle differences between fine-grained categories given limited training data. To filling the classification gap, in this paper, we address the Few-Shot Fine-Grained (FSFG) classification problem, which focuses on tackling the fine-grained classification under the challenging few-shot learning setting. A novel low-rank pairwise bilinear pooling operation is proposed to capture the nuanced differences between the support and query images for learning an effective distance metric. Moreover, a feature alignment layer is designed to match the support image features with query ones before the comparison. We name the proposed model Low-Rank Pairwise Alignment Bilinear Network (LRPABN), which is trained in an end-to-end fashion. Comprehensive experimental results on four widely used fine-grained classification data sets demonstrate that our LRPABN model achieves the superior performances compared to state-of-the-art methods. Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Jingsong Xu, Qiang Wu 0001 |
IEEE Trans. Multim. | 5 |
| 2021 | Weighted Adaptive Image Super-Resolution Scheme Based on Local Fractal Feature and Image RoughnessabstractImage super-resolution aims to reconstruct a high-resolution image from the known low-resolution version. During this process, it should keep the degree of image roughness non-decreasing, which reflects various texture features and appearance. However, this point is not well addressed in the current work. This work argues that reducing roughness during image super-resolution is the key reason causing various problems such as artificial texture and/or edge blur. In this work, keeping the image roughness non-decreasing during super-resolution is being well investigated for the first time to our best knowledge. Image super-resolution is cast as an optimization problem to keep image roughness non-decreasing. In order to tackle this problem, the image super-resolution is approached based on the theory of fractal, where adaptive fractal interpolation function is proposed. In this way, the rational fractal interpolation model is adaptive to every local region. Thus, the roughness of every image region can be best maintained while super-resolution is carried out through fractal interpolation. In this work, the image roughness is reflected by the fractal dimension, which is a key element affecting the construction of fractal interpolation model. That is, the image roughness is measurable using fractal dimension. Mathematically, the overall image super-resolution process can be converted into a fractal interpolation optimization problem where the local fractal dimension is maintained. Although adaptive super-resolution on image segments may best maintain image roughness using the proposed method, it still generates unnecessary block artifacts. To tackle this problem, this work proposes a fine-grained pixel-wise fractal function. Our extensive experimental results demonstrate that the proposed method achieves encouraging performance with the state-of-the-art super-resolution algorithms. Xunxiang Yao, Qiang Wu 0001, Peng Zhang 0057, Fangxun Bao |
IEEE Trans. Multim. | 2 |
| 2021 | Learning Spatial-Temporal Representations Over Walking Tracklet for Long-Term Person Re-Identification in the WildabstractLong-term person re-identification (re-ID) aims to build identity correspondence of the Target Subject of Interest (TSI) exposed under surveillance cameras over a long time interval. Compared to the conventional short-term re-ID studied by most existing works, it suffers an additional problem: significant dressing change observed with time lapsing. Unfortunately, this variation in long-term person re-ID case contradicts the assumption of prior short-term re-ID approaches, and thus causes significant difficulties if conventional short-term re-ID methods are applied. To address the problem, this paper proposes to learn hybrid feature representation via a two-stream network named SpTSkM, including a spatial-temporal stream and a skeleton motion stream. The former performs directly on image sequences, which tends to learn identity-related spatial-temporal patterns such as body geometric structure and body movement. The latter operates on normalized 3D skeletons by adapting graph convolutional network, which tends to learn pure motion patterns from skeleton sequences. Both streams extract fine-grained level time-gap stable information that is robust to appearance changes in long-term re-ID and meanwhile maintains sufficient discriminability to differentiate different people. The final matching metric is obtained by mixing information of the two streams in a score-level fusion strategy. In addition, we collect a Cloth-Varying vIDeo re-ID (CVID-reID) dataset particularly for long-term re-ID. It contains video tracklets of celebrities posted on the Internet. These videos are snapshots under extremely different scenarios that include highly dynamic background, diverse camera views and abundant cloth variations on each TSI. These factors cause CVID-reID more complicated and closer to practice. Our experiments demonstrate the difficulty of long-term person re-ID and also validate the effectiveness of the proposed SpTSkM, showing the best performance. Peng Zhang 0057, Jingsong Xu, Qiang Wu 0001, Yan Huang 0023, Xianye Ben |
IEEE Trans. Multim. | 3 |
| 2021 | Dual-Stream Guided-Learning via a Priori Optimization for Person Re-identificationabstractThe task of person re-identification (re-ID) is to find the same pedestrian across non-overlapping camera views. Generally, the performance of person re-ID can be affected by background clutter. However, existing segmentation algorithms cannot obtain perfect foreground masks to cover the background information clearly. In addition, if the background is completely removed, some discriminative ID-related cues (i.e., backpack or companion) may be lost. In this article, we design a dual-stream network consisting of a Provider Stream (P-Stream) and a Receiver Stream (R-Stream). The R-Stream performs an a priori optimization operation on foreground information. The P-Stream acts as a pusher to guide the R-Stream to concentrate on foreground information and some useful ID-related cues in the background. The proposed dual-stream network can make full use of the a priori optimization and guided-learning strategy to learn encouraging foreground information and some useful ID-related information in the background. Our method achieves Rank-1 accuracy of 95.4% on Market-1501, 89.0% on DukeMTMC-reID, 78.9% on CUHK03 (labeled), and 75.4% on CUHK03 (detected), outperforming state-of-the-art methods. Junyi Wu 0001, Yan Huang 0023, Qiang Wu 0001, Jianqiang Zhao, Liqin Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Exploring Long-Short-Term Context For Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation attracts numerous attention following the success of the point-based convolution neural network. Due to the ambiguity of the point-based feature, many methods study on integrating contextual information to solve the ambiguous problem. However, the extracted context is severely limited to the small input blocks. Few prior works exploit contextual information beyond the blocks to capture long-range dependencies. To address this limitation, we propose a novel long-short-term context framework, which adopts a long-short-term feature bank to exploit both the local context within each block and the long-range context beyond the current task block. The proposed framework is flexible and easy to be combined with existing models, thereby enables existing models to capture the larger range context. Extensive experiments demonstrate that the proposed model achieves improved segmentation performance, and augmenting existing models with a long-short-term feature bank consistently increases the performance. Anan Du, Shuchao Pang, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001 |
ICIP | 5 |
| 2020 | Part-based Collaborative Spatio-temporal Feature Learning for Cloth-changing Gait RecognitionabstractIn decades many gait recognition methods have been proposed using different techniques. However, due to a real-world scenario of clothing variations, a reduction of the recognition rate occurs for most of these methods. Thus in this paper, a part-based spatio-temporal feature learning method is proposed to tackle the problem of clothing variations for gait recognition. First, based on the anatomical properties, human bodies are segmented into two regions, which are affected and unaffected by clothing variations. A learning network is particularly proposed in this paper to grasp principal spatio-temporal features from those unaffected regions. Different from most part-based methods with spatial or temporal features solely being utilized, in our method these two features are associated in a more collaborative manner. Snapshots are created for each gait sequence from the H-W and T-W views. Stable spatial information is embedded in the H-W view and adequate temporal information is embedded in the T-W view. An inherent relationship exists between these two views. Thus, a collaborative spatio-temporal feature will be hybridized by concatenating these correlative spatial and temporal information. The robustness and efficiency of our proposed method are validated by experiments on CASIA Gait Dataset B and OU-ISIR Treadmill Gait Dataset B. Our proposed method can both achieve the state-of-the-art results on these two databases. Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Jingsong Xu |
ICPR | 3 |
| 2020 | Field-wise Learning for Multi-field Categorical DataabstractWe propose a new method for learning with multi-field categorical data. Multi-field categorical data are usually collected over many heterogeneous groups. These groups can reflect in the categories under a field. The existing methods try to learn a universal model that fits all data, which is challenging and inevitably results in learning a complex model. In contrast, we propose a field-wise learning method leveraging the natural structure of data to learn simple yet efficient one-to-one field-focused models with appropriate constraints. In doing this, the models can be fitted to each category and thus can better capture the underlying differences in data. We present a model that utilizes linear models with variance and low-rank constraints, to help it generalize better and reduce the number of parameters. The model is also interpretable in a field-wise manner. As the dimensionality of multi-field categorical data can be very high, the models applied to such data are mostly over-parameterized. Our theoretical analysis can potentially explain the effect of over-parametrization on the generalization of our model. It also supports the variance constraints in the learning objective. The experiment results on two large-scale datasets show the superior performance of our model, the trend of the generalization error bound, and the interpretability of learning outcomes. Our code is available at https://github.com/lzb5600/Field-wise-Learning. Zhibin Li 0002, Jian Zhang 0002, Yongshun Gong, Yazhou Yao, Qiang Wu 0001 |
NeurIPS | 5 |
| 2020 | Automatic Sheep Counting by Multi-object TrackingabstractAnimal counting is a highly skilled yet tedious task in livestock transportation and trading. To effectively free up the human labour and provide accurate counts for sheep loading/unloading, we develop an auto sheep counting system based on multi-object detection, tracking and extrapolation techniques. Our system has demonstrated more than 99.9% accuracy with sheep moving freely in a race under optimal visual conditions. Jingsong Xu, Litao Yu, Jian Zhang 0002, Qiang Wu 0001 |
VCIP | 4 |
| 2020 | A Vision Based Fish Processing SystemabstractThe digital fish provenance and quality tracking system is essential for the seafood supply chain. As a part of this system, we develop a vision-based fish processing system to automatically perform fish freshness estimation, size measurement and species classification. Under the constrained illumination environment, our system is able to auto-process the fish selection, thus greatly reduce the human labour and bring trust and efficiency to the seafood supply chain from catch to market. Zongjian Zhang, Litao Yu, Jian Zhang 0002, Qiang Wu 0001 |
VCIP | 4 |
| 2020 | Error sensitivity model based on spatial and temporal features
Dezhi Bo, Qiang Wu 0001, Ping An 0001 |
Multim. Tools Appl. | 4 |
| 2020 | Light Field Compression Using Global Multiplane Representation and Two-Step PredictionabstractDue to its spatio-angular structure, light field image allows for a wealth of post-processing techniques like digital refocusing and depth estimation. In order to compress the data of the two domains, the current proposal intends to embed the disparity-based view synthesis method into the decoder. However, predicting each view separately or in local groups means bringing more computational burden to the decoder and destroying the light field structure. Since disparity contains the relationship between all light rays in the light field, the proposed solution is to predict a disparity-based global representation as the first step. In the second step, all the views can be predicted easily based on this representation. In this letter, we use the recently proposed multiplane as the form of this global representation. The experimental results show the effectiveness of the proposed solution, and the better RD performance compared to other schemes especially under low bitrates. Ping An 0001, Xinpeng Huang, Chao Yang 0021, Deyang Liu, Qiang Wu 0001 |
IEEE Signal Process. Lett. | 6 |
| 2020 | Generated Data With Sparse Regularized Multi-Pseudo Label for Person Re-IdentificationabstractRecently, Generative Adversarial Network (GAN) has been adopted to improve person re-identification (person re-ID) performance through data augmentation. However, directly leveraging generated data to train a re-ID model may easily lead to over-fitting issue on these extra data and decrease the generalisability of model to learn true ID-related features from real data. Inspired by the previous approach which assigns multi-pseudo labels on the generated data to reduce the risk of over-fitting, we propose to take sparse regularization into consideration. We attempt to further improve the performance of current re-ID models by using the unlabeled generated data. The proposed Sparse Regularized Multi-Pseudo Label (SRMpL) can effectively prevent the over-fitting issue when some larger weights are assigned to the generated data. Our experiments are carried out on two publicly available person re-ID datasets (e.g., Market-1501 and DukeMTMC-reID). Compared with existing unlabeled generated data re-ID solutions, our approach achieves competitive performance. Two classical re-ID models are used to verify our sparse regularization label on generated data, i.e., an ID-embedding network and a two-stream network. Liqin Huang, Junyi Wu 0001, Yan Huang 0023, Qiang Wu 0001, Jingsong Xu |
IEEE Signal Process. Lett. | 5 |
| 2020 | Coupled Bilinear Discriminant Projection for Cross-View Gait RecognitionabstractA problem that hinders good performance of general gait recognition systems is that the appearance features of gaits are more affected-prone by views than identities, especially when the walking direction of the probe gait is different from the register gait. This problem cannot be solved by traditional projection learning methods because these methods can learn only one projection matrix, and thus for the same subject, it cannot transfer cross-view gait features into similar ones. This paper presents an innovative method to overcome this problem by aligning gait energy images (GEIs) across views with the coupled bilinear discriminant projection (CBDP). Specifically, the CBDP generates the aligned gait matrix features for two views with two sets of bilinear transformation matrices, so that the original GEIs' spatial structure information can be preserved. By iteratively maximizing the ratio of inter-class distance metric to intra-class distance metric, the CBDP can learn the optimal matrix subspace where the GEIs across views are aligned in both horizontal and vertical coordinates. Therefore, the CBDP is also able to avoid the under-sample problem. We also theoretically prove that the upper and lower bounds of the objective function sequence of the CBDP are both monotonically increasing, so the convergence of the CBDP is demonstrated. In the terms of accuracy, the comparative experiments on the CASIA (B) and OU-ISIR gait databases show that our method is superior to the state-of-the-art cross-view gait recognition methods. More impressively, encouraging performance is obtained by our method even in matching a lateral-view gait with a frontal-view gait. Xianye Ben, Chen Gong 0002, Peng Zhang 0057, Qiang Wu 0001, Weixiao Meng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Beyond Scalar Neuron: Adopting Vector-Neuron Capsules for Long-Term Person Re-IdentificationabstractCurrent person re-identification (re-ID) works mainly focus on the short-term scenario where a person is less likely to change clothes. However, in the long-term re-ID scenario, a person has a great chance to change clothes. A sophisticated re-ID system should take such changes into account. To facilitate the study of long-term re-ID, this paper introduces a large-scale re-ID dataset called “Celeb-reID” to the community. Unlike previous datasets, the same person can change clothes in the proposed Celeb-reID dataset. Images of Celeb-reID are acquired from the Internet using street snap-shots of celebrities. There is a total of 1,052 IDs with 34,186 images making Celeb-reID being the largest long-term re-ID dataset so far. To tackle the challenge of cloth changes, we propose to use vector-neuron (VN) capsules instead of the traditional scalar neurons (SN) to design our network. Compared with SN, one extra-dimensional information in VN can perceive cloth changes of the same person. We introduce a well-designed ReIDCaps network and integrate capsules to deal with the person re-ID task. Soft Embedding Attention (SEA) and Feature Sparse Representation (FSR) mechanisms are adopted in our network for performance boosting. Experiments are conducted on the proposed long-term re-ID dataset and two common short-term re-ID datasets. Comprehensive analyses are given to demonstrate the challenge exposed in our datasets. Experimental results show that our ReIDCaps can outperform existing state-of-the-art methods by a large margin in the long-term scenario.The new dataset and code will be released to facilitate future researches. Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Yi Zhong 0002, Peng Zhang 0057, Zhaoxiang Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Top-Push Constrained Modality-Adaptive Dictionary Learning for Cross-Modality Person Re-IdentificationabstractPerson re-identification aims to match person captured by multiple non-overlapping cameras that mainly mean standard RGB cameras. In contemporary surveillance, cameras of different modalities such as infrared cameras and depth cameras are introduced because of their unique advantages in poor illumination scenarios. However, re-identifying the persons across such cameras of different modalities is extremely difficult and, unfortunately, seldom discussed. It is mainly caused by extremely different appearances of the person shown under such different camera modalities. In this paper, we tackle this challenging cross-modality people re-identification through a top-push constrained modality-adaptive dictionary learning. The proposed model asymmetrically projects the heterogeneous features from dissimilar modalities onto a common space. In this way, the modality-specific bias is mitigated. Thus, the heterogeneous data can be simultaneously enforced by a shared dictionary in a canonical space. Moreover, a top-push ranking graph regularization is embedded in the proposed model to improve the discriminability, which efficiently further boosts the matching accuracy. In order to implement the proposed model, an iterative process is developed in this paper to optimize these two processes jointly. Extensive experiments on the benchmark SYSU-MM01 and BIWI RGBD-ID person re-identification datasets show promising results which outperform state-of-the-art methods. Peng Zhang 0057, Jingsong Xu, Qiang Wu 0001, Yan Huang 0023, Jian Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Depth Map Enhancement by Revisiting Multi-Scale Intensity Guidance Within Coarse-to-Fine StagesabstractBeing different from the most methods of guided depth map enhancement based on deep convolutional neural network which focus on increasing the depth of networks, this paper is to improve the effectiveness of intensity guidance when the network goes deep. Overall, the proposed network upsamples the low-resolution depth maps from coarse to fine. Within each refinement stage of certain-scale depth features, the current-scale and all coarse-scales of the guidance features are revisited by dense connection. Therefore, the multi-scale guidance is efficiently maintained as the propagation of features. Furthermore, the proposed network maintains the intensity features in the high-resolution domain from which the multi-scale guidance is directly extracted. This design further improves the quality of intensity guidance. In addition, the shallow depth features upsampled via transposed convolution layer are directly transferred to the final depth features for reconstruction, which is called global residual learning in feature domain. Similarly, the global residual learning in pixel domain learns the difference between the depth ground truth and the coarsely upsampled depth map. Also, the local residual learning is to maintain the low frequency within each refinement stage and progressively recover the high frequency. The proposed method is tested for noise-free and noisy cases which compares against 16 state-of-the-art methods. Our experimental results show the improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Yuming Fang 0001, Yong Yang 0001, Xiwu Shang, Qiang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Multi-Scale Frequency Reconstruction for Guided Depth Map Super-Resolution via Deep Residual NetworkabstractThe depth maps obtained by the consumer-level sensors are always noisy in the low-resolution (LR) domain. Existing methods for the guided depth super-resolution, which are based on the pre-defined local and global models, perform well in general cases (e.g., joint bilateral filter and Markov random field). However, such model-based methods may fail to describe the potential relationship between RGB-D image pairs. To solve this problem, this paper proposes a data-driven approach based on the deep convolutional neural network with global and local residual learning. It progressively upsamples the LR depth map guided by the high-resolution intensity image in multiple scales. A global residual learning is adopted to learn the difference between the ground truth and the coarsely upsampled depth map, and the local residual learning is introduced in each scale-dependent reconstruction sub-network. This scheme can restore the depth structure from coarse to fine via multi-scale frequency synthesis. In addition, batch normalization layers are used to improve the performance of depth map denoising. Our method is evaluated in noise-free and noisy cases. A comprehensive comparison against 17 state-of-the-art methods is carried out. The experimental results show that the proposed method has faster convergence speed as well as improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Qiang Wu 0001, Yuming Fang 0001, Ping An 0001, Liqin Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | SBSGAN: Suppression of Inter-Domain Background Shift for Person Re-IdentificationabstractCross-domain person re-identification (re-ID) is challenging due to the bias between training and testing domains. We observe that if backgrounds in the training and testing datasets are very different, it dramatically introduces difficulties to extract robust pedestrian features, and thus compromises the cross-domain person re-ID performance. In this paper, we formulate such problems as a background shift problem. A Suppression of Background Shift Generative Adversarial Network (SBSGAN) is proposed to generate images with suppressed backgrounds. Unlike simply removing backgrounds using binary masks, SBSGAN allows the generator to decide whether pixels should be preserved or suppressed to reduce segmentation errors caused by noisy foreground masks. Additionally, we take ID-related cues, such as vehicles and companions into consideration. With high-quality generated images, a Densely Associated 2-Stream (DA-2S) network is introduced with Inter Stream Densely Connection (ISDC) modules to strengthen the complementarity of the generated data and ID-related cues. The experiments show that the proposed method achieves competitive performance on three re-ID datasets, i.e., Market-1501, DukeMTMC-reID, and CUHK03, under the cross-domain person re-ID scenario. Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002 |
ICCV | 2 |
| 2019 | Bi-Level Masked Multi-scale CNN-RNN Networks for Short Text RepresentationabstractRepresenting short text is becoming extremely important for a variety of valuable applications. However, representing short text is critical yet challenging because it involves lots of informal words and typos (i.e. the noise problem) but only few vocabularies in each text (i.e. the sparsity problem). Most of existing work on representing short text relies on noise recognition and sparsity expansion. However, the noises in short text are with various forms and changing fast, but, most of the current methods may fail to adaptively recognize the noise. Also, it is hard to explicitly expand a sparse text to a high-quality dense text. In this paper, we tackle the noise and sparsity problems in short text representation by learning multi-grain noise-tolerant patterns and then embedding the most significant patterns in a text as its representation. To achieve this goal, we propose a bi-level multi-scale masked CNN-RNN network to embed the most significant multi-grain noise-tolerant relations among words and characters in a text into a dense vector space. Comprehensive experiments on five large real-world data sets demonstrate our method significantly outperforms the state-of-the-art competitors. Qian Li 0006, Qiang Wu 0001, Chengzhang Zhu, Jian Zhang 0002 |
ICDAR | 2 |
| 2019 | Kpsnet: Keypoint Detection and Feature Extraction for Point Cloud RegistrationabstractThis paper presents the KPSNet, a KeyPoint Siamese Network to simultaneously learn task-desirable keypoint detector and feature extractor. The keypoint detector is optimized to predict a score vector, which signifies the probability of each candidate being a keypoint. The feature extractor is optimized to learn robust features of keypoints by exploiting the correspondence between the keypoints generated from two inputs, respectively. For training, the KPSNet does not require to manually annotate keypoints and local patches pairwise. Instead, we design an alignment module to establish the correspondence between the two inputs and generate positive and negative samples on-the-fly. Therefore, our method can be easily extended to new scenes. We test the proposed method on the open-source benchmark and experiments show the validity of our method. Anan Du, Xiaoshui Huang, Jian Zhang 0002, Lingxiang Yao, Qiang Wu 0001 |
ICIP | 5 |
| 2019 | Fast Registration for Cross-Source Point Clouds by using Weak Regional Affinity and Pixel-Wise RefinementabstractMany types of 3D acquisition sensors have emerged in recent years and point cloud has been widely used in many areas. Accurate and fast registration of cross-source 3D point clouds from different sensors is an emerged research problem in computer vision. This problem is extremely challenging because cross-source point clouds contain a mixture of various variances, such as density, partial overlap, large noise and outliers, viewpoint changing. In this paper, an algorithm is proposed to align cross-source point clouds with both high accuracy and high efficiency. There are two main contributions: firstly, two components, the weak region affinity and pixel-wise refinement, are proposed to maintain the global and local information of 3D point clouds. Then, these two components are integrated into an iterative tensor-based registration algorithm to solve the cross-source point cloud registration problem. We conduct experiments on a synthetic cross-source benchmark dataset and real cross-source datasets. Comparison with six state-of-the-art methods, the proposed method obtains both higher efficiency and accuracy. Xiaoshui Huang, Lixin Fan, Qiang Wu 0001, Jian Zhang 0002, Chun Yuan 0003 |
ICME | 3 |
| 2019 | Compare More Nuanced: Pairwise Alignment Bilinear Network for Few-Shot Fine-Grained LearningabstractThe recognition ability of human beings is developed in a progressive way. Usually, children learn to discriminate various objects from coarse to fine-grained with limited supervision. Inspired by this learning process, we propose a simple yet effective model for the Few-Shot Fine-Grained (FSFG) recognition, which tries to tackle the challenging fine-grained recognition task using meta-learning. The proposed method, named Pairwise Alignment Bilinear Network (PABN), is an end-to-end deep neural network. Unlike traditional deep bilinear networks for fine-grained classification, which adopt the self-bilinear pooling to capture the subtle features of images, the proposed model uses a novel pairwise bilinear pooling to compare the nuanced differences between base images and query images for learning a deep distance metric. In order to match base image features with query image features, we design feature alignment losses before the proposed pairwise bilinear pooling. Experiment results on four fine-grained classification datasets and one generic few-shot dataset demonstrate that the proposed model outperforms both the state-of-the-art few-shot fine-grained and general few-shot methods. Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Qiang Wu 0001, Jingsong Xu |
ICME | 4 |
| 2019 | Celebrities-ReID: A Benchmark for Clothes Variation in Long-Term Person Re-IdentificationabstractThis paper considers person re-identification (re-ID) in the case of long-time gap (i.e., long-term re-ID) that concentrates on the challenge of clothes variation of each person. We introduce a new dataset, named Celebrities-reID to handle that challenge. Compared with current datasets, the proposed Celebrities-reID dataset is featured in two aspects. First, it contains 590 persons with 10,842 images, and each person does not wear the same clothing twice, making it the largest clothes variation person re-ID dataset to date. Second, a comprehensive evaluation using state of the arts is carried out to verify the feasibility and new challenge exposed by this dataset. In addition, we propose a benchmark approach to the dataset where a two-step fine-tuning strategy on human body parts is introduced to tackle the challenge of clothes variation. In experiments, we evaluate the feasibility and quality of the proposed Celebrities-reID dataset. The experimental results demonstrate that the proposed benchmark approach is not only able to best tackle clothes variation shown in our dataset but also achieves competitive performance on a widely used person re-ID dataset Market1501, which further proves the reliability of the proposed benchmark approach. Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002 |
IJCNN | 2 |
| 2019 | An Inferable Representation Learning for Fraud Review Detection with Cold-start ProblemabstractFraud review significantly damages the business reputation and also customers' trust to certain products. It has become a serious problem existing on the current social media. Various efforts have been put in to tackle such problems. However, in the case of cold-start where a review is posted by a new user who just pops up on the social media, common fraud detection methods may fail because most of them are heavily depended on the information about the user's historical behavior and its social relation to other users, yet such information is lacking in the cold-start case. This paper presents a novel Joint-bEhavior-and-Social-relaTion-infERable (JESTER) embedding method to leverage the user reviewing behavior and social relations for cold-start fraud review detection. JESTER embeds the deep characteristics of existing user behavior and social relations of users and items in an inferable user-item-review-rating representation space where the representation of a new user can be efficiently inferred by a closed-form solution and reflects the user's most probable behavior and social relations. Thus, a cold-start fraud review can be effectively detected accordingly. Our experiments show JESTER (i) performs significantly better in detecting fraud reviews on four real-life social media data sets, and (ii) effectively infers new user representation in the cold-start problem, compared to three state-of-the-art and two baseline competitors. Qian Li 0006, Qiang Wu 0001, Chengzhang Zhu, Jian Zhang 0002 |
IJCNN | 2 |
| 2019 | Visual Relationship Attention for Image CaptioningabstractVisual attention mechanisms have been broadly used by image captioning models to attend to related visual information dynamically, allowing fine-grained image understanding and reasoning. However, they are only designed to discover the region-level alignment between visual features and the language feature. The exploration of higher-level visual relationship information between image regions, which is rarely researched in recent works, is beyond their capabilities. To fill this gap, we propose a novel visual relationship attention model based on the parallel attention mechanism under the learnt spatial constraints. It can extract relationship information from visual regions and language and then achieve the relationship-level alignment between them. Using combined visual relationship attention and visual region attention to attend to related visual relationships and regions respectively, our image captioning model can achieve state-of-the-art performances on the MSCOCO dataset. Both quantitative analysis and qualitative analysis demonstrate that our novel visual relationship attention model can capture related visual relationship and further improve the caption quality. Zongjian Zhang, Yang Wang 0002, Qiang Wu 0001, Fang Chen 0001 |
IJCNN | 3 |
| 2019 | VT-GAN: View Transformation GAN for Gait Recognition Across ViewsabstractRecognizing gaits without human cooperation is of importance in surveillance and forensics because of the benefits that gait is unique and collected remotely. However, change of camera view angle severely degrades the performance of gait recognition. To address the problem, previous methods usually learn mappings for each pair of views which incurs abundant independently built models. In this paper, we proposed a View Transformation Generative Adversarial Networks (VT-GAN) to achieve view transformation of gaits across two arbitrary views using only one uniform model. In specific, we generated gaits in target view conditioned on input images from any views and the corresponding target view indicator. In addition to the classical discriminator in GAN which makes the generated images look realistic, a view classifier is imposed. This controls the consistency of generated images and conditioned target view indicator and ensures to generate gaits in the specified target view. On the other hand, retaining identity information while performing view transformation is another challenge. To solve the issue, an identity distilling module with triplet loss is integrated, which constrains the generated images inheriting identity information from inputs and yields discriminative feature embeddings. The proposed VT-GAN generates visually promising gaits and achieves promising performances for cross-view gait recognition, which exhibits great effectiveness of the proposed VT-GAN. Peng Zhang 0057, Qiang Wu 0001, Jingsong Xu |
IJCNN | 2 |
| 2019 | VN-GAN: Identity-preserved Variation Normalizing GAN for Gait RecognitionabstractGait is recognized as a unique biometric characteristic to identify a walking person remotely across surveillance networks. However, the performance of gait recognition severely suffers challenges from view angle diversity. To address the problem, an identity-preserved Variation Normalizing Generative Adversarial Network (VN-GAN) is proposed for learning purely identity-related representations. It adopts a coarse-to-fine manner which firstly generates initial coarse images by normalizing view to an identical one and then refines the coarse images by injecting identity-related information. In specific, Siamese structure with discriminators for both camera view angles and human identities is utilized to achieve variation normalization and identity preservation of two stages, respectively. In addition to discriminators, reconstruction loss and identity-preserving loss are integrated, which forces the generated images to be the same in view and to be discriminative in identity. This ensures to generate identity-related images in an identical view of good visual effect for gait recognition. Extensive experiments on benchmark datasets demonstrate that the proposed VN-GAN can generate visually interpretable results and achieve promising performance for gait recognition. Peng Zhang 0057, Qiang Wu 0001, Jingsong Xu |
IJCNN | 2 |
| 2019 | Sample Adaptive Multiple Kernel Learning for Failure Prediction of Railway PointsabstractRailway points are among the key components of railway infrastructure. As a part of signal equipment, points control the routes of trains at railway junctions, having a significant impact on the reliability, capacity, and punctuality of rail transport. Meanwhile, they are also one of the most fragile parts in railway systems. Points failures cause a large portion of railway incidents. Traditionally, maintenance of points is based on a fixed time interval or raised after the equipment failures. Instead, it would be of great value if we could forecast points' failures and take action beforehand, minimising any negative effect. To date, most of the existing prediction methods are either lab-based or relying on specially installed sensors which makes them infeasible for large-scale implementation. Besides, they often use data from only one source. We, therefore, explore a new way that integrates multi-source data which are ready to hand to fulfil this task. We conducted our case study based on Sydney Trains rail network which is an extensive network of passenger and freight railways. Unfortunately, the real-world data are usually incomplete due to various reasons, e.g., faults in the database, operational errors or transmission faults. Besides, railway points differ in their locations, types and some other properties, which means it is hard to use a unified model to predict their failures. Aiming at this challenging task, we firstly constructed a dataset from multiple sources and selected key features with the help of domain experts. In this paper, we formulate our prediction task as a multiple kernel learning problem with missing kernels. We present a robust multiple kernel learning algorithm for predicting points failures. Our model takes into account the missing pattern of data as well as the inherent variance on different sets of railway points. Extensive experiments demonstrate the superiority of our algorithm compared with other state-of-the-art methods. Zhibin Li 0002, Jian Zhang 0002, Qiang Wu 0001, Yongshun Gong, Jinfeng Yi, Christina Kirsch |
KDD | 3 |
| 2019 | Unsupervised User Behavior Representation for Fraud Review Detection with Cold-Start Problem
Qian Li 0006, Qiang Wu 0001, Chengzhang Zhu, Jian Zhang 0002 |
PAKDD (1) | 2 |
| 2019 | Modified Baseline for Light Field StitchingabstractIn traditional 2D image stitching, the baseline method usually means global homography via Direct Linear Transformation (DLT) on inliers. In this paper, a modified baseline method for light field (LF) stitching is proposed to stitch two LFs. The depth map and the center sub-aperture image (SAI) are used to filter the feature points of the entire LF. The global 4D homography is then calculated by DLT to align all SAIs corresponding to the same angular domain coordinates of two LFs. Finally, the improved Markov Random Field (MRF) energy considering the global LF is used to find the seam of 2D SAIs instead of computational 4D graph cut. Experimental results show that the proposed method can effectively stitch the 4D LFs, and preserve the consistency of the angular and spatial domains of the stitched LF compared with implementing 2D image stitching to the corresponding SAIs. Moreover, the method proposed in this paper can easily extend all advanced 2D image stitching methods to 4D LF, so that the acquired LF can have larger field of view and wider applications. Ping An 0001, Xinpeng Huang, Chunli Meng, Qiang Wu 0001 |
VCIP | 5 |
| 2019 | Improving Person Re-Identification Performance Using Body Mask Via Cross-Learning StrategyabstractThe task of person re-identification (re-id) is to find the same pedestrian across non-overlapping cameras. Normally, the performance of person re-id can be affected by background clutters. However, existing segmentation algorithms are hard to obtain perfect foreground person images. To effectively leverage the body (foreground) cue, and in the meantime pay attention to discriminative information in the background (e.g., companion or vehicle), we propose to use a cross-learning strategy to take both foreground and other discriminative information into account. In addition, since currently existing foreground segmentation result always involves noise, we use Label Smoothing Regularization (LSR) to strengthen the generalization capability during our learning process. In experiments, we pick up two state-of-the-art person re-id methods to verify the effectiveness of our proposed cross-learning strategy. Our experiments are carried out on two publicly available person re-id datasets. Obvious performance improvements can be observed on both datasets. Junyi Wu 0001, Lingxiang Yao, Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Liqin Huang |
VCIP | 5 |
| 2019 | Cost-Effective Foliage Penetration Human Detection Under Severe Weather Conditions Based on Auto-Encoder/Decoder Neural NetworkabstractMilitary surveillance events and rescue activities are vital missions for the Internet-of-Things. To this end, foliage penetration for human detection plays an important role. However, although the feasibility of that mission has been validated, we observe that it still cannot perform promisingly under severe weather conditions, such as rainy, foggy, and snowy days. Therefore, in this paper, experiments are conducted under severe weather conditions based on a proposed deep learning approach. We present an auto-encoder/decoder (Auto-ED) deep neural network that can learn the deep representation and conduct classification task concurrently. Since the property of cost-effective, the device-free sensing techniques are used to address human detection in our case. As we pursue the signal-based mission, two components are involved in the proposed Auto-ED approach. First, an encoder is utilized that encode signal-based inputs into higher dimensional tensors by fractionally strided convolution operations. Then, a decoder is leveraged with convolution operations to extract deep representations and learn the classifier simultaneously. To verify the effectiveness of the proposed approach, we compare it with several machine learning approaches under different weather conditions. Also, a simulation experiment is conducted by adding additive white Gaussian noise to the original target signals with different signal to noise ratios. Experimental results demonstrate that the proposed approach can best tackle the challenge of human detection under severe weather conditions in the high-clutter foliage environment, which indicates its potential application values in the near future. Yan Huang 0023, Yi Zhong 0002, Qiang Wu 0001, Eryk Dutkiewicz, Ting Jiang 0008 |
IEEE Internet Things J. | 3 |
| 2019 | Adaptive rational fractal interpolation function for image super-resolution via local fractal analysis
Xunxiang Yao, Qiang Wu 0001, Peng Zhang 0057, Fangxun Bao |
Image Vis. Comput. | 2 |
| 2019 | Heritage image annotation via collective knowledge
Junjie Zhang 0002, Qi Wu 0001, Jian Zhang 0002, Chunhua Shen, Jianfeng Lu 0003, Qiang Wu 0001 |
Pattern Recognit. | 6 |
| 2019 | Pyramid-Structured Depth MAP Super-Resolution Based on Deep Dense-Residual NetworkabstractAlthough deep convolutional neural networks (DCNN) show significant improvement for single depth map (SD) super-resolution (SR) over the traditional counterparts, most SDSR DCNNs do not reuse the hierarchical features for depth map SR resulting in blurred high-resolution (HR) depth maps. They always stack convolutional layers to make network deeper and wider. In addition, most SDSR networks generate HR depth maps at a single level, which is not suitable for large up-sampling factors. To solve these problems, we present pyramid-structured depth map super-resolution based on deep dense-residual network. Specially, our networks are made up of dense residual blocks that use densely connected layers and residual learning to model the mapping between high-frequency residuals and low-resolution (LR) depth map. Furthermore, based on the pyramid structure, our network can progressively generate depth maps of various levels by taking advantages of features from different levels. The proposed network adopts a deep supervision scheme to reduce the difficulty of model training and further improve the performance. The proposed method is evaluated on Middlebury datasets which shows improved performance compared with 6 state-of-the-art methods. Liqin Huang, Jianjia Zhang, Yifan Zuo 0001, Qiang Wu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2019 | Coupled Patch Alignment for Matching Cross-View GaitsabstractGait recognition has attracted growing attention in recent years as the gait of humans has a strong discriminative ability even under low resolution at a distance. Unfortunately, the performance of gait recognition can be largely affected by view change. To address this problem, we propose a Coupled Patch Alignment (CPA) algorithm that effectively matches a pair of gaits across different views. To realize CPA, we first build a certain amount of patches, and each of them is made up of a sample as well as its intra-class and inter-class nearest-neighbors. Then we design an objective function for each patch to balance the cross-view intra-class compactness and the cross-view inter-class separability. Finally, all the local independent patches are combined to render a unified objective function. Theoretically, we show that the proposed CPA has a close relationship with Canonical Correlation Analysis (CCA). Algorithmically, we extend CPA to "Multi-dimensional Patch Alignment" (MPA) that can handle an arbitrary number of views. Comprehensive experiments on CASIA(B), USF and OU-ISIR gait databases firmly demonstrate the effectiveness of our methods over other existing popular methods in terms of cross-view gait recognition. Xianye Ben, Chen Gong 0002, Peng Zhang 0057, Xitong Jia, Qiang Wu 0001, Weixiao Meng 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Multi-Pseudo Regularized Label for Generated Data in Person Re-IdentificationabstractSufficient training data normally is required to train deeply learned models. However, due to the expensive manual process for labelling large number of images (i.e., annotation), the amount of available training data (i.e., real data) is always limited. To produce more data for training a deep network, Generative Adversarial Network (GAN) can be used to generate artificial sample data (i.e., generated data). However, the generated data usually does not have annotation labels. To solve this problem, in this paper, we propose a virtual label called Multi-pseudo Regularized Label (MpRL) and assign it to the generated data. With MpRL, the generated data will be used as the supplementary of real training data to train a deep neural network in a semi-supervised learning fashion. To build the corresponding relationship between the real data and generated data, MpRL assigns each generated data a proper virtual label which reflects the likelihood of the affiliation of the generated data to predefined training classes in the real data domain. Unlike the traditional label which usually is a single integral number, the virtual label proposed in this work is a set of weight-based values each individual of which is a number in (0,1] called multi-pseudo label and reflects the degree of relation between each generated data to every pre-defined class of real data. A comprehensive evaluation is carried out by adopting two state-of-the-art convolutional neural networks (CNNs) in our experiments to verify the effectiveness of MpRL. Experiments demonstrate that by assigning MpRL to generated data, we can further improve the person re-ID performance on five re-ID datasets, i.e., Market-1501, DukeMTMC-reID, CUHK03, VIPeR, and CUHK01. The proposed method obtains +6.29%, +6.30%, +5.58%, +5.84%, and +3.48% improvements in rank-1 accuracy over a strong CNN baseline on the five datasets respectively, and outperforms state-of-the-art methods. Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Zhedong Zheng, Zhaoxiang Zhang 0001, Jian Zhang 0002 |
IEEE Trans. Image Process. | 3 |
| 2019 | A Computational Model for Stereoscopic Visual Saliency PredictionabstractDepth information plays an important role in human vision as it provides additional cues that distinguish objects from their backgrounds. This paper explores depth information for analyzing stereoscopic saliency and presents a computational model that predicts stereoscopic visual saliency based on three aspects of human vision: 1) the pop-out effect; 2) comfort zones; and 3) background effects. Through an analysis of these three phenomena, we find that most of the stereoscopic saliency region can be explained. Our model comprises three modules, each describing one aspect of saliency distribution, and a control function that can be used to adjust the three models independently. The relationship between the three models is not mutually exclusive. One, two, or three phenomena may appear in one image. Therefore, to accurately determine which phenomena the image conforms to, we have devised a selection strategy that chooses the appropriate combination of models based on the content of the image. Our approach is implemented within a framework based on the multifeature analysis. The framework considers surrounding regions, color/depth contrast, and points of interest. The selection strategy can improve the performance of the framework. A series of experiments on two recent eye-tracking datasets shows that our proposed method outperforms several state-of-the-art saliency models. Hao Cheng 0003, Jian Zhang 0002, Qiang Wu 0001, Ping An 0001 |
IEEE Trans. Multim. | 3 |
| 2019 | High-Quality Image Captioning With Fine-Grained and Semantic-Guided Visual AttentionabstractThe soft-attention mechanism is regarded as one of the representative methods for image captioning. Based on the end-to-end convolutional neural network (CNN)-long short term memory (LSTM) framework, the soft-attention mechanism attempts to link the semantic representation in text (i.e., captioning) with relevant visual information in the image for the first time. Motivated by this approach, several state-of-the-art attention methods are proposed. However, due to the constraints of CNN architecture, the given image is only segmented to the fixed-resolution grid at a coarse level. The visual feature extracted from each grid indiscriminately fuses all inside objects and/or their portions. There is no semantic link between grid cells. In addition, the large area “stuff” (e.g., the sky or a beach) cannot be represented using the current methods. To address these problems, this paper proposes a new model based on the fully convolutional network (FCN)-LSTM framework, which can generate an attention map at a fine-grained grid-wise resolution. Moreover, the visual feature of each grid cell is contributed only by the principal object. By adopting the grid-wise labels (i.e., semantic segmentation), the visual representations of different grid cells are correlated to each other. With the ability to attend to large area “stuff,” our method can further summarize an additional semantic context from semantic labels. This method can provide comprehensive context information to the language LSTM decoder. In this way, a mechanism of fine-grained and semantic-guided visual attention is created, which can accurately link the relevant visual information with each semantic meaning inside the text. Demonstrated by three experiments including both qualitative and quantitative analyses, our model can generate captions of high quality, specifically high levels of accuracy, completeness, and diversity. Moreover, our model significantly outperforms all other methods that use VGG-based CNN encoders without fine-tuning. Zongjian Zhang, Qiang Wu 0001, Yang Wang 0002, Fang Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Fine-Grained and Semantic-Guided Visual Attention for Image CaptioningabstractSoft-attention is regarded as one of the representative methods for image captioning. Based on the end-to-end CNN-LSTM framework, it tries to link the relevant visual information on the image with the semantic representation in the text (i.e. captioning) for the first time. In recent years, there are several state-of-the-art methods published, which are motivated by this approach and include more elegant fine-tune operation. However, due to the constraints of CNN architecture, the given image is only segmented to fixed-resolution grid at a coarse level. The overall visual feature created for each grid cell indiscriminately fuses all inside objects and/or their portions. There is no semantic link among grid cells, although an object may be segmented into different grid cells. In addition, the large-area stuff (e.g. sky and beach) cannot be represented in the current methods. To tackle the problems above, this paper proposes a new model based on the FCN-LSTM framework which can segment the input image into a fine-grained grid. Moreover, the visual feature representing each grid cell is contributed only by the principal object or its portion in the corresponding cell. By adopting the pixel-wise labels (i.e. semantic segmentation), the visual representations of different grid cells are correlated to each other. In this way, a mechanism of fine-grained and semantic-guided visual attention is created, which can better link the relevant visual information with each semantic meaning inside the text through LSTM. Without using the elegant fine-tune, the comprehensive experiments show promising performance consistently across different evaluation metrics. Zongjian Zhang, Qiang Wu 0001, Yang Wang 0002, Fang Chen 0001 |
WACV | 2 |
| 2018 | Long-Term Person Re-identification Using True Motion from VideosabstractMost person re-identification approaches and benchmarks assume that pedestrians go across the surveillance network without significant appearance changes in a brief period, which explicitly restricts person re-identification to a short-term event and incurs inter-sample similarity measurement by appearance matching. However, pedestrians are likely to reappear in the surveillance network after a long-time interval (long-term) and change their wearing in many real-world scenarios. These scenarios inevitably cause appearances between subjects more ambiguous and indistinguishable. In this paper we consider these scenarios and propose a unified feature representation based on true motion cues from videos named FIne moTion encoDing (FITD). Our hypothesis is that people keep constant motion patterns under non-distraction walking condition. Therefore, the motion characteristics are more reliable than static appearance feature to describe a walking person. Particularly, we extract motion patterns hierarchically by encoding trajectory-aligned descriptors with Fisher vectors in a spatial-aligned pyramid. To verify benefits of the proposed FITD, we collect a new dataset typically for the long-term situations. Extensive experiments demonstrate the merits of our FITD especially for the long-term scenarios. Peng Zhang 0057, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002 |
WACV | 2 |
| 2018 | A survey: facial micro-expression recognition
Madhumita A. Takalkar, Min Xu 0001, Qiang Wu 0001, Zenon Chaczko |
Multim. Tools Appl. | 3 |
| 2018 | Single Image Dehazing Based on Dark Channel Prior and Energy MinimizationabstractHazy images have limited visibility and low contrast. The degradation is expressed by transmission map, which is one of the most important estimates of single image dehazing. Transmission map estimation is an underconstraint problem, and lots of priors have been proposed. Among them, the dark channel prior is widely recognized. However, traditional methods have not fully exploited its power due to improper assumptions or operations, which cause unwanted artifacts. The postrefinement algorithms employed to remove these artifacts in turn undermine the merits of the prior. In this letter, a novel method for estimating transmission map by energy minimization is proposed to solve this problem. The energy function combines the dark channel prior with piecewise smoothness. The method is compared to the state-of-the-art methods and shows outstanding performance. Mingzhu Zhu, Bingwei He, Qiang Wu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2018 | A Coarse-to-Fine Algorithm for Matching and Registration in 3D Cross-Source Point CloudsabstractWe propose an efficient method to deal with the matching and registration problem found in cross-source point clouds captured by different types of sensors. This task is especially challenging due to the presence of density variation, scale difference, a large proportion of noise and outliers, missing data, and viewpoint variation. The proposed method has two stages: in the coarse matching stage, we use the ensemble of shape functions descriptor to select potential K regions from the candidate point clouds for the target. In the fine stage, we propose a scale embedded generative Gaussian mixture models registration method to refine the results from the coarse matching stage. Following the fine stage, both the best region and accurate camera pose relationships between the candidates and target are found. We conduct experiments in which we apply the method to two applications: one is 3D object detection and localization in street-view outdoor (LiDAR/VSFM) cross-source point clouds and the other is 3D scene matching and registration in indoor (KinectFusion/VSFM) cross-source point clouds. The experiment results show that the proposed method performs well when compared with the existing methods. It also shows that the proposed method is robust under various sensing techniques, such as LiDAR, Kinect, and RGB camera. Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001, Lixin Fan, Chun Yuan 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Explicit Edge Inconsistency Evaluation Model for Color-Guided Depth Map EnhancementabstractColor-guided depth enhancement is used to refine depth maps according to the assumption that the depth edges and the color edges at the corresponding locations are consistent. In methods on such low-level vision tasks, the Markov random field (MRF), including its variants, is one of the major approaches that have dominated this area for several years. However, the assumption above is not always true. To tackle the problem, the state-of-the-art solutions are to adjust the weighting coefficient inside the smoothness term of the MRF model. These methods lack an explicit evaluation model to quantitatively measure the inconsistency between the depth edge map and the color edge map, so they cannot adaptively control the efforts of the guidance from the color image for depth enhancement, leading to various defects such as texture-copy artifacts and blurring depth edges. In this paper, we propose a quantitative measurement on such inconsistency and explicitly embed it into the smoothness term. The proposed method demonstrates promising experimental results compared with the benchmark and state-of-the-art methods on the Middlebury ToF-Mark, and NYU data sets. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | GII Representation-Based Cross-View Gait Recognition by Discriminative Projection With List-Wise ConstraintsabstractRemote person identification by gait is one of the most important topics in the field of computer vision and pattern recognition. However, gait recognition suffers severely from the appearance variance caused by the view change. It is very common that gait recognition has a high performance when the view is fixed but the performance will have a sharp decrease when the view variance becomes significant. Existing approaches have tried all kinds of strategies like tensor analysis or view transform models to slow down the trend of performance decrease but still have potential for further improvement. In this paper, a discriminative projection with list-wise constraints (DPLC) is proposed to deal with view variance in cross-view gait recognition, which has been further refined by introducing a rectification term to automatically capture the principal discriminative information. The DPLC with rectification (DPLCR) embeds list-wise relative similarity measurement among intraclass and inner-class individuals, which can learn a more discriminative and robust projection. Based on the original DPLCR, we have introduced the kernel trick to exploit nonlinear cross-view correlations and extended DPLCR to deal with the problem of multiview gait recognition. Moreover, a simple yet efficient gait representation, namely gait individuality image (GII), based on gait energy image is proposed, which could better capture the discriminative information for cross view gait recognition. Experiments have been conducted in the CASIA-B database and the experimental results demonstrate the outstanding performance of both the DPLCR framework and the new GII representation. It is shown that the DPLCR-based cross-view gait recognition has outperformed the-state-of-the-art approaches in almost all cases under large view variance. The combination of the GII representation and the DPLCR has further enhanced the performance to be a new benchmark for cross-view gait recognition. Zhaoxiang Zhang 0001, Jiaxin Chen 0002, Qiang Wu 0001, Ling Shao 0001 |
IEEE Trans. Cybern. | 3 |
| 2018 | Depth Super-Resolution on RGB-D Video Sequences With Large Displacement 3D MotionabstractTo enhance the resolution and accuracy of depth data, some video-based depth super-resolution methods have been proposed which utilizes its neighboring depth images in the temporal domain. They often consist of two main stages: motion compensation of temporally neighboring depth images and fusion of compensated depth images. However, large displacement 3D motion often leads to compensation error, and the compensation error is further introduced into the fusion. A video-based depth super-resolution method with novel motion compensation and fusion approaches is proposed in this paper. We claim that, 3D Nearest Neighboring Field (NNF) is a better choice than using positions with true motion displacement for depth enhancements. To handle large displacement 3D motion, the compensation stage utilized 3D NNF instead of true motion used in previous methods. Next, the fusion approach is modeled as a regression problem to predict the super-resolution result efficiently for each depth image by using its compensated depth images. A new deep convolutional neural network architecture is designed for fusion, which is able to employ a large amount of video data for learning the complicated regression function. We comprehensively evaluate our method on various RGB-D video sequences to show its superior performance. Yucheng Wang 0003, Jian Zhang 0002, Zicheng Liu 0001, Qiang Wu 0001, Zhengyou Zhang, Yunde Jia |
IEEE Trans. Image Process. | 4 |
| 2018 | Minimum Spanning Forest With Embedded Edge Inconsistency Measurement Model for Guided Depth Map EnhancementabstractGuided depth map enhancement based on Markov Random Field (MRF) normally assumes edge consistency between the color image and the corresponding depth map. Under this assumption, the low-quality depth edges can be refined according to the guidance from the high-quality color image. However, such consistency is not always true, which leads to texture-copying artifacts and blurring depth edges. In addition, the previous MRF-based models always calculate the guidance affinities in the regularization term via a non-structural scheme which ignores the local structure on the depth map. In this paper, a novel MRF-based method is proposed. It computes these affinities via the distance between pixels in a space consisting of the Minimum Spanning Trees (Forest) to better preserve depth edges. Furthermore, inside each Minimum Spanning Tree, the weights of edges are computed based on explicit edge inconsistency measurement model, which significantly mitigates texture-copying artifacts. To further tolerate the effects caused by noise and better preserve depth edges, a bandwidth adaption scheme is proposed. Our method is evaluated for depth map super-resolution and depth map completion problems on synthetic and real datasets including Middlebury, ToF-Mark and NYU. A comprehensive comparison against 16 state-of-the-art methods is carried out. Both qualitative and quantitative evaluation present the improved performances. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Minimum spanning forest with embedded edge inconsistency measurement for color-guided depth map upsamplingabstractColor-guided depth map up-sampling, such as Markov-Random-Field-based (MRF-based) methods, is a popular depth map enhancement solution, which normally assumes edge consistency between color image and corresponding depth map. It calculates the coefficients of smoothness term in MRF according to such assumption. However, such consistency is not always true which leads to texture-copying artifacts and blurring depth edges. In this paper, we propose a novel coefficient computing scheme for smoothness term in MRF which is based on the distance between pixels in the Minimum Spanning Trees (Forest) to better preserve depth edges. The explicit edge inconsistency measurement is embedded into weights of edges in Minimum Spanning Trees, which significantly mitigates texture-copying artifacts. The proposed method is evaluated on Middlebury datasets and ToF-Mark datasets which demonstrates improved results compared with state-of-the-art methods. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
ICME | 2 |
| 2017 | Variable Bandwidth Weighting for Texture Copy Artifact Suppression in Guided Depth UpsamplingabstractIn this paper, we mathematically analyze one of the most challenging issues in color image-guided depth upsampling: the texture copy artifacts. The optimal guidance weights denoted by balanced weights are proposed to best suppress texture copy artifacts. To both suppress texture copy artifacts and preserve depth discontinuities, a new general weighting scheme called variable bandwidth weighting is proposed. The variable bandwidth weighting scheme is able to adjust guidance weights according to the local depth smoothness. A new concept called relative smoothness is proposed for measuring the local depth smoothness. Given this quantitative smoothness measurement, the proposed weighting scheme can adaptively adjust the bandwidth for calculating the guidance weights in the existing methods. As we use the computationally efficient balanced weights instead of the guidance weights of a large bandwidth, the proposed method can speed up the upsampling process for about $2\times \sim 5\times $ when compared with the original upsampling methods. Experimental results show the effectiveness and efficiency of the proposed method in suppressing texture copy artifacts, preserving the depth discontinuities and reducing the computational cost at the same time. Wei Liu 0044, Jie Yang 0002, Qiang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | A Systematic Approach for Cross-Source Point Cloud Registration by Preserving Macro and Micro StructuresabstractWe propose a systematic approach for registering cross-source point clouds that come from different kinds of sensors. This task is especially challenging due to the presence of significant missing data, large variations in point density, scale difference, large proportion of noise, and outliers. The robustness of the method is attributed to the extraction of macro and micro structures. Macro structure is the overall structure that maintains similar geometric layout in cross-source point clouds. Micro structure is the element (e.g., local segment) being used to build the macro structure. We use graph to organize these structures and convert the registration into graph matching. With a novel proposed descriptor, we conduct the graph matching in a discriminative feature space. The graph matching problem is solved by an improved graph matching solution, which considers global geometrical constraints. Robust cross source registration results are obtained by incorporating graph matching outcome with RANSAC and ICP refinements. Compared with eight state-of-the-art registration algorithms, the proposed method invariably outperforms on Pisa Cathedral and other challenging cases. In order to compare quantitatively, we propose two challenging cross-source data sets and conduct comparative experiments on more than 27 cases, and the results show we obtain much better performance than other methods. The proposed method also shows high accuracy in same-source data sets. Xiaoshui Huang, Jian Zhang 0002, Lixin Fan, Qiang Wu 0001, Chun Yuan 0003 |
IEEE Trans. Image Process. | 4 |
| 2017 | Robust Color Guided Depth Map RestorationabstractOne of the most challenging issues in color guided depth map restoration is the inconsistency between color edges in guidance color images and depth discontinuities on depth maps. This makes the restored depth map suffer from texture copy artifacts and blurring depth discontinuities. To handle this problem, most state-of-the-art methods design complex guidance weight based on guidance color images and heuristically make use of the bicubic interpolation of the input depth map. In this paper, we show that using bicubic interpolated depth map can blur depth discontinuities when the upsampling factor is large and the input depth map contains large holes and heavy noise. In contrast, we propose a robust optimization framework for color guided depth map restoration. By adopting a robust penalty function to model the smoothness term of our model, we show that the proposed method is robust against the inconsistency between color edges and depth discontinuities even when we use simple guidance weight. To the best of our knowledge, we are the first to solve this problem with a principled mathematical formulation rather than previous heuristic weighting schemes. The proposed robust method performs well in suppressing texture copy artifacts. Moreover, it can better preserve sharp depth discontinuities than previous heuristic weighting schemes. Through comprehensive experiments on both simulated data and real data, we show promising performance of the proposed method. Wei Liu 0044, Jie Yang 0002, Qiang Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Robust weighted least squares for guided depth upsamplingabstractIn this paper, we propose a new guided depth upsampling method denoted as Robust Weighted Least Squares (RWLS). Our work is inspired by the connection between the Weighted Least Squares (WLS) and the Auto Regressive (AR) model. By adopting a new robust penalty function to model the smoothness of the proposed model, we show that the proposed method performs much better in preserving sharp depth discontinuities than previous work. Through both mathematical analysis and experimental results, we show that our method has promising performance on handling the inconsistency between the guidance image and the depth map in both preserving sharp depth discontinuities and suppressing the texture copy artifacts. Wei Liu 0044, Jie Yang 0002, Qiang Wu 0001 |
ICIP | 4 |
| 2016 | Explicit measurement on depth-color inconsistency for depth completionabstractColor-guided depth completion is to refine depth map through structure light sensing by filling missing depth structure and de-nosing. It is based on the assumption that depth discontinuity and color edge at the corresponding location are consistent. Among all proposed methods, MRF-based method including its variants is one of major approaches. However, the assumption above is not always true, which causes texture-copy and depth discontinuity blurring artifacts. The state-of-the-art solutions usually are to modify the weighting inside smoothness term of MRF model. Because there is no any method explicitly considering the inconsistency occurring between depth discontinuity and the corresponding color edge, they cannot adaptively control the effect of guidance from color image when completing depth map. In this paper, we propose quantitative measurement on such inconsistency and explicitly embed it into weighting value of smoothness term. The proposed method is evaluated on NYU Kinect datasets and demonstrates promising results. Yifan Zuo 0001, Qiang Wu 0001, Ping An 0001, Jian Zhang 0002 |
ICIP | 2 |
| 2016 | Explicit modeling on depth-color inconsistency for color-guided depth up-samplingabstractColor-guided depth up-sampling is to enhance the resolution of depth map according to the assumption that the depth discontinuity and color image edge at the corresponding location are consistent. Through all methods reported, MRF including its variants is one of major approaches, which has dominated in this area for several years. However, the assumption above is not always true. Solution usually is to adjust the weighting inside smoothness term in MRF model. But there is no any method explicitly considering the inconsistency occurring between depth discontinuity and the corresponding color edge. In this paper, we propose quantitative measurement on such inconsistency and explicitly embed it into weighting value of smoothness term. Such solution has not been reported in the literature. The improved depth up-sampling based on the proposed method is evaluated on Middlebury datasets and ToFMark datasets and demonstrate promising results. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
ICME | 2 |
| 2016 | Handling Occlusion and Large Displacement Through Improved RGB-D Scene Flow EstimationabstractThe accuracy of scene flow is restricted by several challenges such as occlusion and large displacement motion. When occlusion happens, the positions inside the occluded regions lose their corresponding counterparts in preceding and succeeding frames. Large displacement motion will increase the complexity of motion modeling and computation. Moreover, occlusion and large displacement motion are highly related problems in scene flow estimation, e.g., large displacement motion often leads to considerably occluded regions in the scene. An improved dense scene flow method based on red-green-blue-depth (RGB-D) data is proposed in this paper. To handle occlusion, we model the occlusion status for each point in our problem formulation, and jointly estimate the scene flow and occluded regions. To deal with large displacement motion, we employ an over-parameterized scene flow representation to model both the rotation and translation components of the scene flow, since large displacement motion cannot be well approximated using translational motion only. Furthermore, we employ a two-stage optimization procedure for this overparameterized scene flow representation. In the first stage, we propose a new RGB-D PatchMatch method, which is mainly applied in the RGB-D image space to reduce the computational complexity introduced by the large displacement motion. According to the quantitative evaluation based on the Middlebury data set, our method outperforms other published methods. The improved performance is also comprehensively confirmed on the real data acquired by Kinect sensor. Yucheng Wang 0003, Jian Zhang 0002, Zicheng Liu 0001, Qiang Wu 0001, Philip A. Chou, Zhengyou Zhang, Yunde Jia |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Small target detection using an optimization-based filterabstractSmall target detection is a critical problem in the Infrared Search And Track (IRST) system. Although it has been studied for years, there are some challenges remained, e.g. cloud edges and horizontal lines are likely to cause false alarms. This paper proposes a novel method using an optimization-based filter to detect infrared small target in heavy clutter. First, we design a certain pixel area as active area. Second, a weighted quadratic cost function is performed in the active area. Finally, a filter based on statistics of active area is derived from the cost function. Our method could preserve heterogeneous area, meanwhile, remove target region. Experimental results show our method achieves satisfied performance in heavy clutter. Keren Fu, Tao Zhou 0002, Jie Yang 0002, Qiang Wu 0001, Xiangjian He |
ICASSP | 5 |
| 2015 | Enhancing person re-identification by integrating gait biometric
Zheng Liu 0014, Zhaoxiang Zhang 0001, Qiang Wu 0001, Yunhong Wang 0001 |
Neurocomputing | 3 |
| 2015 | Fast and robust head detection with arbitrary pose and occlusion
Tao Zhang 0010, Wenjing Jia, Qiang Wu 0001, Jie Yang 0002, Xiangjian He |
Multim. Tools Appl. | 4 |
| 2015 | An MRF-Based Depth Upsampling: Upsample the Depth Map With Its Own PropertyabstractIn this letter, we propose a novel method for upsampling the noisy low resolution depth map with the guidance of the companion color image. The problem is modeled with an Markov Random Field (MRF)-based optimization framework. The novelty relies on the smoothness term that is modeled with an exponential function as the error norm. By using this novel error norm, our method can take the property of the depth map into account. Depth discontinuity cues are not only obtained from the color image but also the depth map itself. Our method has much better performance in preserving sharp depth discontinuities and suppressing the texture copy artifacts. Experimental results show that our method outperforms state-of-art solutions in both visual quality and accuracy. Wei Liu 0044, Shaoyong Jia, Penglin Li, Jie Yang 0002, Qiang Wu 0001 |
IEEE Signal Process. Lett. | 6 |
| 2015 | Local N-Ary Pattern and Its Extension for Texture ClassificationabstractTexture image classification is important in computer vision research. To effectively capture texture patterns, a distinctive feature such as a local binary pattern (LBP) is needed. An LBP is robust against monotonic and gray-scale variations and it computes quickly. Its robustness and speed advantage have made it popular in various texture analysis applications. However, an LBP is sensitive to noise, particularly smooth weak illumination gradients in near-uniform regions. To mitigate the effect of noise and increase distinctiveness, a local ternary pattern (LTP) is proposed. Compared with a binary coding LBP, an LTP adopts ternary coding. As a result, an LTP can better tolerate noise and is significantly more distinctive. These advantages of an LTP effectively improve its classification accuracy. However, the potential of ternary coding is not fully explored in LTPs because a ternary pattern is split into a pair of binary patterns. In this paper, to fully explore the distinctiveness in the local pattern, the feature extraction process is formulated as an integer decomposition problem, which is a generalized version of the Bachet de Meziriac weight problem (BMWP). Following this generalization, a local n-ary pattern (LNP) is proposed, for which the LBP is a special case parametrized under n = 2. The LTP is not a special case of the LNP. Both LBP and LTP are used as benchmark methods to evaluate LNPs performance due to their well-recognized success. In addition, a rotation-invariant and uniform LNP is also proposed and compared with a rotation-invariant and uniform LBP. The proposed LNP achieves significantly improved texture classification accuracy compared with the LBP and also demonstrates considerable improvement over the LTP. Sheng Wang 0003, Qiang Wu 0001, Xiangjian He, Jie Yang 0002, Yi Wang 0037 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Violent video detection based on MoSIFT feature and sparse codingabstractTo detect violence in a video, a common video description method is to apply local spatio-temporal description on the query video. Then, the low-level description is further summarized onto the high-level feature based on Bag-of-Words (BoW) model. However, traditional spatio-temporal descriptors are not discriminative enough. Moreover, BoW model roughly assigns each feature vector to only one visual word, therefore inevitably causing quantization error. To tackle the constrains, this paper employs Motion SIFT (MoSIFT) algorithm to extract the low-level description of a query video. To eliminate the feature noise, Kernel Density Estimation (KDE) is exploited for feature selection on the MoSIFT descriptor. In order to obtain the highly discriminative video feature, this paper adopts sparse coding scheme to further process the selected MoSIFTs. Encouraging experimental results are obtained based on two challenging datasets which record both crowded scenes and non-crowded scenes. Chen Gong 0002, Jie Yang 0002, Qiang Wu 0001, Lixiu Yao |
ICASSP | 4 |
| 2014 | Street view cross-sourced point cloud matching and registrationabstractObject registration has been widely discussed with the development of various range sensing technologies. In most work, however, the point clouds of reference and target are generated by the same technology, such as a Kinect range camera, LiDAR sensor, or Structure from Motion technique. Cases in which reference and target point clouds are generated by different technologies are rarely discussed. Due to the significant differences across various point cloud data in terms of point cloud density, sensing noise, scale, occlusion etc., object registration between such different point clouds becomes extremely difficult. In this study, we address for the first time an even more challenging case in which the differently-sourced point clouds are acquired from a real street view. One is generated on the basis of an image sequence through the SfM process, and the other is produced directly by the LiDAR system. We propose a two-stage matching and registration algorithm to achieve object registration between these two different point clouds. The experiments are based on real building object point cloud data and demonstrate the effectiveness and efficiency of the proposed solution. The newly proposed solution can be further developed to contribute to several related applications, such as Location Based Service. Furong Peng, Qiang Wu 0001, Lixin Fan, Jian Zhang 0002, Yu You, Jianfeng Lu 0003, Jing-Yu Yang 0001 |
ICIP | 2 |
| 2014 | Shape Preserving RGB-D Depth Map Restoration
Wei Liu 0044, Haoyang Xue, Yun Gu, Jie Yang 0002, Qiang Wu 0001, Zhenhong Jia |
ICONIP (3) | 5 |
| 2014 | Semi-supervised classification with pairwise constraints
Chen Gong 0002, Keren Fu, Qiang Wu 0001, Enmei Tu, Jie Yang 0002 |
Neurocomputing | 3 |
| 2014 | Exploiting Universum data in AdaBoost using gradient descent
Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang |
Image Vis. Comput. | 2 |
| 2014 | An efficient color quantization based on generic roughness measure
Xiaodong Yue 0002, Duoqian Miao 0001, Longbing Cao, Qiang Wu 0001, Yufei Chen 0002 |
Pattern Recognit. | 4 |
| 2014 | Boosting Separability in Semisupervised Learning for Object ClassificationabstractBoosting algorithms, especially AdaBoost, have attracted great attention in computer vision. In the early version of boosting algorithms, the weak classifier selection and the strong classifier learning are linked together. It has been demonstrated that decoupling of these two processes can provide more flexibility for training a better classifier. In these studies, linear discriminant analysis (LDA) has been adopted to select weak classifiers independently based on class separability rather than a training error that occurs normally in AdaBoost. It is observed that LDA is successful only if a large number of labeled training samples is available. However, a large-scale labeled training set is not always available in many computer vision applications such as object classification. To tackle this problem, this paper proposes semisupervised subspace learning combined with a boosting framework for object classification, through which unlabeled data can participate in the boosting training to compensate for the lack of enough labeled data. With the proposed framework, this paper develops three various approaches that utilize unlabeled data in different ways. According to the experiments on several public image data sets, the proposed methods achieve superior performance over AdaBoost and existing semisupervised algorithms. Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Fumin Shen, Zhenmin Tang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | PageRank Tracker: From Ranking to TrackingabstractVideo object tracking is widely used in many real-world applications, and it has been extensively studied for over two decades. However, tracking robustness is still an issue in most existing methods, due to the difficulties with adaptation to environmental or target changes. In order to improve adaptability, this paper formulates the tracking process as a ranking problem, and the PageRank algorithm, which is a well-known webpage ranking algorithm used by Google, is applied. Labeled and unlabeled samples in tracking application are analogous to query webpages and the webpages to be ranked, respectively. Therefore, determining the target is equivalent to finding the unlabeled sample that is the most associated with existing labeled set. We modify the conventional PageRank algorithm in three aspects for tracking application, including graph construction, PageRank vector acquisition and target filtering. Our simulations with the use of various challenging public-domain video sequences reveal that the proposed PageRank tracker outperforms mean-shift tracker, co-tracker, semiboosting and beyond semiboosting trackers in terms of accuracy, robustness and stability. Chen Gong 0002, Keren Fu, Artur Loza, Qiang Wu 0001, Jie Yang 0002 |
IEEE Trans. Cybern. | 4 |
| 2014 | Recognizing Gaits Across Views Through Correlated Motion Co-ClusteringabstractHuman gait is an important biometric feature, which can be used to identify a person remotely. However, view change can cause significant difficulties for gait recognition because it will alter available visual features for matching substantially. Moreover, it is observed that different parts of gait will be affected differently by view change. By exploring relations between two gaits from two different views, it is also observed that a part of gait in one view is more related to a typical part than any other parts of gait in another view. A new method proposed in this paper considers such variance of correlations between gaits across views that is not explicitly analyzed in the other existing methods. In our method, a novel motion co-clustering is carried out to partition the most related parts of gaits from different views into the same group. In this way, relationships between gaits from different views will be more precisely described based on multiple groups of the motion co-clustering instead of a single correlation descriptor. Inside each group, a linear correlation between gait information across views is further maximized through canonical correlation analysis (CCA). Consequently, gait information in one view can be projected onto another view through a linear approximation under the trained CCA subspaces. In the end, a similarity between gaits originally recorded from different views can be measured under the approximately same view. Comprehensive experiments based on widely adopted gait databases have shown that our method outperforms the state-of-the-art. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li, Liang Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Generalized local N-ary patterns for texture classificationabstractLocal Binary Pattern (LBP) has been well recognised and widely used in various texture analysis applications of computer vision and image processing. It integrates properties of texture structural and statistical texture analysis. LBP is invariant to monotonic gray-scale variations and has also extensions to rotation invariant texture analysis. In recent years, various improvements have been achieved based on LBP. One of extensive developments was replacing binary representation with ternary representation and proposed Local Ternary Pattern (LTP). This paper further generalises the local pattern representation by formulating it as a generalised weight problem of Bachet de Meziriac and proposes Local N-ary Pattern (LNP). The encouraging performance is achieved based on three benchmark datasets when compared with its predecessors. Sheng Wang 0003, Xiangjian He, Qiang Wu 0001, Jie Yang 0002 |
AVSS | 3 |
| 2013 | Training boosting-like algorithms with semi-supervised subspace learningabstractBoosting algorithms have attracted great attention since the first real-time face detector by Viola & Jones through feature selection and strong classifier learning simultaneously. On the other hand, researchers have proposed to decouple such two procedures to improve the performance of Boosting algorithms. Motivated by this, we propose a boosting-like algorithm framework by embedding semi-supervised subspace learning methods. It selects weak classifiers based on class-separability. Combination weights of selected weak classifiers can be obtained by subspace learning. Three typical algorithms are proposed under this framework and evaluated on public data sets. As shown by our experimental results, the proposed methods obtain superior performances over their supervised counterparts and AdaBoost. Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Fumin Shen, Zhenmin Tang |
ICIP | 2 |
| 2013 | Attribute-based learning for large scale object classificationabstractScalability to large numbers of classes is an important challenge for multi-class classification. It can often be computationally infeasible at test phase when class prediction is performed by using every possible classifier trained for each individual class. This paper proposes an attribute-based learning method to overcome this limitation. First is to define attributes and their associations with object classes automatically and simultaneously. Such associations are learned based on greedy strategy under certain conditions. Second is to learn a classifier for each attribute instead of each class. Then, these trained classifiers are used to predict classes based on their attribute representations. The proposed method also allows trade-off between test-time complexity (which grows linearly with the number of attributes) and accuracy. Experiments based on Animals-with-Attributes and ILSVRC2010 datasets have shown that the performance of our method is promising when compared with the state-of-the-art. Worapan Kusakunniran, Shin'ichi Satoh 0001, Jian Zhang 0002, Qiang Wu 0001 |
ICME | 4 |
| 2013 | Multi-view urban scene reconstruction in non-uniform volumeabstractThis paper presents a new fully automatic approach for multi-view urban scene reconstruction. Our algorithm is based on the Manhattan-World assumption, which can provide compact models while preserving fidelity of synthetic architectures. Starting from a dense point cloud, we extract its main axes by global optimization, and construct a nonuniform volume based on them. A graph model is created from volume facets rather than voxels. Appropriate edge weights are defined to ensure the validity and quality of the surface reconstruction. Compared with the common pointcloud- to-model methods, the proposed methodology exploits image information to unveil the real structures of holes in the point cloud. Experiments demonstrate the encouraging performance of the algorithm. Runchao Mao, Qiang Wu 0001, Yu Qiao 0003, Li Bai 0001, Jie Yang 0002 |
ICMV | 2 |
| 2013 | Violence detection based on histogram of optical flow orientationabstractIn this paper, we propose a novel approach for violence detection and localization in a public scene. Currently, violence detection is considerably under-researched compared with the common action recognition. Although existing methods can detect the presence of violence in a video, they cannot precisely locate the regions in the scene where violence is happening. This paper will tackle the challenge and propose a novel method to locate the violence location in the scene, which is important for public surveillance. The Gaussian Mixed Model is extended into the optical flow domain in order to detect candidate violence regions. In each region, a new descriptor, Histogram of Optical Flow Orientation (HOFO), is proposed to measure the spatial-temporal features. A linear SVM is trained based on the descriptor. The performance of the method is demonstrated on the publicly available data sets, BEHAVE and CAVIAR. Tao Zhang 0010, Jie Yang 0002, Qiang Wu 0001, Li Bai 0001, Lixiu Yao |
ICMV | 4 |
| 2013 | MIL-SKDE: Multiple-instance learning with supervised kernel density estimation
Ruo Du, Qiang Wu 0001, Xiangjian He, Jie Yang 0002 |
Signal Process. | 2 |
| 2013 | A New View-Invariant Feature for Cross-View Gait RecognitionabstractHuman gait is an important biometric feature which is able to identify a person remotely. However, change of view causes significant difficulties for recognizing gaits. This paper proposes a new framework to construct a new view-invariant feature for cross-view gait recognition. Our view-normalization process is performed in the input layer (i.e., on gait silhouettes) to normalize gaits from arbitrary views. That is, each sequence of gait silhouettes recorded from a certain view is transformed onto the common canonical view by using corresponding domain transformation obtained through invariant low-rank textures (TILTs). Then, an improved scheme of procrustes shape analysis (PSA) is proposed and applied on a sequence of the normalized gait silhouettes to extract a novel view-invariant gait feature based on procrustes mean shape (PMS) and consecutively measure a gait similarity based on procrustes distance (PD). Comprehensive experiments were carried out on widely adopted gait databases. It has been shown that the performance of the proposed method is promising when compared with other existing methods in the literature. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2012 | On splitting dataset: Boosting Locally Adaptive Regression Kernels for car localizationabstractIn this paper, we study the impact of learning an Adaboost classifier with small sample set (i.e., with fewer training examples). In particular, we make use of car localization as an underlying application, because car localization can be widely used to various real world applications. In order to evaluate the performance of Adaboost learning with a few examples, we simply apply Adaboost learning to a recently proposed feature descriptor - Locally Adaptive Regression Kernel (LARK). As a type of state-of-the-art feature descriptor, LARK is robust against illumination changes and noises. More importantly, we use LARK because its spatial property is also favorable for our purpose (i.e., each patch in the LARK descriptor corresponds to one unique pixel in the original image). In addition to learning a detector from the entire training dataset, we also split the original training dataset into several sub-groups and then we train one detector for each sub-group. We compare those features associated using the detector of each sub-group with that of the detector learnt with the entire training dataset and propose improvements based on the comparison results. Our experimental results indicate that the Adaboost learning is only successful on a small dataset when those learnt features simultaneously satisfy two conditions that: 1. features are learnt from the Region of Interest (ROI), and 2. features are sufficiently far away from each other. Sheng Wang 0003, Qiang Wu 0001, Xiangjian He, Min Xu 0001 |
ICARCV | 2 |
| 2012 | Object Detection Based on Co-occurrence GMuLBP FeaturesabstractImage co-occurrence has shown great powers on object classification because it captures the characteristic of individual features and spatial relationship between them simultaneously. For example, Co-occurrence Histogram of Oriented Gradients (CoHOG) has achieved great success on human detection task. However, the gradient orientation in CoHOG is sensitive to noise. In addition, CoHOG does not take gradient magnitude into account which is a key component to reinforce the feature detection. In this paper, we propose a new LBP feature detector based image co-occurrence. Building on uniform Local Binary Patterns, the new feature detector detects Co-occurrence Orientation through Gradient Magnitude calculation. It is known as CoGMuLBP. An extension version of the GoGMuLBP is also presented. The experimental results on the UIUC car data set show that the proposed features outperform state-of-the-art methods. Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang |
ICME | 2 |
| 2012 | Multiscale roughness measure for color image segmentation
Xiaodong Yue 0002, Duoqian Miao 0001, L. B. Cao, Qiang Wu 0001 |
Inf. Sci. | 5 |
| 2012 | Cross-view and multi-view gait recognitions based on view transformation model using multi-layer perceptron
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
Pattern Recognit. Lett. | 2 |
| 2012 | Directional high-pass filter for blurry image analysis
Jie Yang 0002, Qiang Wu 0001, Xiangjian He |
Signal Process. Image Commun. | 3 |
| 2012 | Fast and Accurate Human Detection Using a Cascade of Boosted MS-LBP FeaturesabstractIn this letter, a new scheme for generating local binary patterns (LBP) is presented. This Modified Symmetric LBP (MS-LBP) feature takes advantage of LBP and gradient features. It is then applied into a boosted cascade framework for human detection. By combining MS-LBP with Haar-like feature into the boosted framework, the performances of heterogeneous features based detectors are evaluated for the best trade-off between accuracy and speed. Two feature training schemes, namely Single AdaBoost Training Scheme (SATS) and Dual AdaBoost Training Scheme (DATS) are proposed and compared. On the top of AdaBoost, two multidimensional feature projection methods are described. A comprehensive experiment is presented. Apart from obtaining higher detection accuracy, the detection speed based on DATS is 17 times faster than HOG method. Jingsong Xu, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang |
IEEE Signal Process. Lett. | 2 |
| 2012 | Gait Recognition Under Various Viewing Angles Based on Correlated Motion RegressionabstractIt is well recognized that gait is an important biometric feature to identify a person at a distance, e.g., in video surveillance application. However, in reality, change of viewing angle causes significant challenge for gait recognition. A novel approach using regression-based view transformation model (VTM) is proposed to address this challenge. Gait features from across views can be normalized into a common view using learned VTM(s). In principle, a VTM is used to transform gait feature from one viewing angle (source) into another viewing angle (target). It consists of multiple regression processes to explore correlated walking motions, which are encoded in gait features, between source and target views. In the learning processes, sparse regression based on the elastic net is adopted as the regression function, which is free from the problem of overfitting and results in more stable regression models for VTM construction. Based on widely adopted gait database, experimental results show that the proposed method significantly improves upon existing VTM-based methods and outperforms most other baseline methods reported in the literature. Several practical scenarios of applying the proposed method for gait recognition under various views are also discussed in this paper. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Gait Recognition Across Various Walking Speeds Using Higher Order Shape Configuration Based on a Differential Composition ModelabstractGait has been known as an effective biometric feature to identify a person at a distance. However, variation of walking speeds may lead to significant changes to human walking patterns. It causes many difficulties for gait recognition. A comprehensive analysis has been carried out in this paper to identify such effects. Based on the analysis, Procrustes shape analysis is adopted for gait signature description and relevant similarity measurement. To tackle the challenges raised by speed change, this paper proposes a higher order shape configuration for gait shape description, which deliberately conserves discriminative information in the gait signatures and is still able to tolerate the varying walking speed. Instead of simply measuring the similarity between two gaits by treating them as two unified objects, a differential composition model (DCM) is constructed. The DCM differentiates the different effects caused by walking speed changes on various human body parts. In the meantime, it also balances well the different discriminabilities of each body part on the overall gait similarity measurements. In this model, the Fisher discriminant ratio is adopted to calculate weights for each body part. Comprehensive experiments based on widely adopted gait databases demonstrate that our proposed method is efficient for cross-speed gait recognition and outperforms other state-of-the-art methods. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2011 | Pairwise Shape configuration-based PSA for gait recognition under small viewing angle changeabstractTwo main components of Procrustes Shape Analysis (PSA) are adopted and adapted specifically to address gait recognition under small viewing angle change: 1) Procrustes Mean Shape (PMS) for gait signature description; 2) Procrustes Distance (PD) for similarity measurement. Pairwise Shape Configuration (PSC) is proposed as a shape descriptor in place of existing Centroid Shape Configuration (CSC) in conventional PSA. PSC can better tolerate shape change caused by viewing angle change than CSC. Small variation of viewing angle makes large impact only on global gait appearance. Without major impact on local spatio-temporal motion, PSC which effectively embeds local shape information can generate robust view-invariant gait feature. To enhance gait recognition performance, a novel boundary re-sampling process is proposed. It provides only necessary re-sampled points to PSC description. In the meantime, it efficiently solves problems of boundary point correspondence, boundary normalization and boundary smoothness. This re-sampling process adopts prior knowledge of body pose structure. Comprehensive experiment is carried out on the CASIA gait database. The proposed method is shown to significantly improve performance of gait recognition under small viewing angle change without additional requirements of supervised learning, known viewing angle and multi-camera system, when compared with other methods in literatures. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
AVSS | 2 |
| 2011 | An effective document image deblurring algorithmabstractDeblurring camera-based document image is an important task in digital document processing, since it can improve both the accuracy of optical character recognition systems and the visual quality of document images. Traditional deblurring algorithms have been proposed to work for natural-scene images. However the natural-scene images are not consistent with document images. In this paper, the distinct characteristics of document images are investigated. We propose a content-aware prior for document image deblurring. It is based on document image foreground segmentation. Besides, an upper-bound constraint combined with total variation based method is proposed to suppress the rings in the deblurred image. Comparing with the traditional general purpose deblurring methods, the proposed deblurring algorithm can produce more pleasing results on document images. Encouraging experimental results demonstrate the efficacy of the proposed method. Xiangjian He, Jie Yang 0002, Qiang Wu 0001 |
CVPR | 4 |
| 2011 | Speed-invariant gait recognition based on Procrustes Shape Analysis using higher-order shape configurationabstractWalking speed change is considered a typical challenge hindering reliable human gait recognition. This paper proposes a novel method to extract speed-invariant gait feature based on Procrustes Shape Analysis (PSA). Two major components of PSA, i.e., Procrustes Mean Shape (PMS) and Procrustes Distance (PD), are adopted and adapted specifically for the purpose of speed-invariant gait recognition. One of our major contributions in this work is that, instead of using conventional Centroid Shape Configuration (CSC) which is not suitable to describe individual gait when body shape changes particularly due to change of walking speed, we propose a new descriptor named Higher-order derivative Shape Configuration (HSC) which can generate robust speed-invariant gait feature. From the first order to the higher order, derivative shape configuration contains gait shape information of different levels. Intuitively, the higher order of derivative is able to describe gait with shape change caused by the larger change of walking speed. Encouraging experimental results show that our proposed method is efficient for speed-invariant gait recognition and evidently outperforms other existing methods in the literatures. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
ICIP | 2 |
| 2011 | Learning Global and Local Features for License Plate Detection
Sheng Wang 0003, Wenjing Jia, Qiang Wu 0001, Xiangjian He, Jie Yang 0002 |
ICONIP (3) | 3 |
| 2011 | Facial Expression Recognition on Hexagonal Structure Using LBP-Based Histogram Variances
Xiangjian He, Ruo Du, Wenjing Jia, Qiang Wu 0001, Wei-Chang Yeh 0001 |
MMM (2) | 5 |
| 2011 | SKRWM based descriptor for pedestrian detection in thermal imagesabstractPedestrian detection in a thermal image is a difficult task due to intrinsic challenges:1) low image resolution, 2) thermal noising, 3) polarity changes, 4) lack of color, texture or depth information. To address these challenges, we propose a novel mid-level feature descriptor for pedestrian detection in thermal domain, which combines pixel-level Steering Kernel Regression Weights Matrix (SKRWM) with their corresponding covariances. SKRWM can properly capture the local structure of pixels, while the covariance computation can further provide the correlation of low level feature. This mid-level feature descriptor not only captures the pixel-level data difference and spatial differences of local structure, but also explores the correlations among low-level features. In the case of human detection, the proposed mid-level feature descriptor can discriminatively distinguish pedestrian from complexity. For testing the performance of proposed feature descriptor, a popular classifier framework based on Principal Component Analysis (PCA) and Support Vector Machine (SVM) is also built. Overall, our experimental results show that proposed approach has overcome the problems caused by background subtraction in [1] while attains comparable detection accuracy compared to the state-of-the-arts. Qiang Wu 0001, Jian Zhang 0002, Glenn Geers |
MMSP | 2 |
| 2011 | More on Weak Feature: Self-correlate Histogram Distances
Sheng Wang 0003, Qiang Wu 0001, Xiangjian He, Wenjing Jia |
PSIVT (1) | 2 |
| 2010 | Canny Edge Detection Using Bilateral Filter on Real Hexagonal Structure
Xiangjian He, Daming Wei, Kin-Man Lam 0001, Wenjing Jia, Qiang Wu 0001 |
ACIVS (1) | 7 |
| 2010 | Support vector regression for multi-view gait recognition based on local motion feature selectionabstractGait is a well recognized biometric feature that is used to identify a human at a distance. However, in real environment, appearance changes of individuals due to viewing angle changes cause many difficulties for gait recognition. This paper re-formulates this problem as a regression problem. A novel solution is proposed to create a View Transformation Model (VTM) from the different point of view using Support Vector Regression (SVR). To facilitate the process of regression, a new method is proposed to seek local Region of Interest (ROI) under one viewing angle for predicting the corresponding motion information under another viewing angle. Thus, the well constructed VTM is able to transfer gait information under one viewing angle into another viewing angle. This proposal can achieve view-independent gait recognition. It normalizes gait features under various viewing angles into a common viewing angle before similarity measurement is carried out. The extensive experimental results based on widely adopted benchmark dataset demonstrate that the proposed algorithm can achieve significantly better performance than the existing methods in literature. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
CVPR | 2 |
| 2010 | ECCH: A novel color coocurrence histogramabstractIn this paper, a novel color cooccurrence histogram method, named eCCH which stands for color cooccurrence histogram at edge points, is proposed to describe the spatial-color joint distribution of images. Unlike all existing ideas, we only investigate the color distribution of pixels located at the two sides of edge points on gradient direction lines. When measuring the similarity of two eCCHs, the Gaussian weighted histogram intersection method is adopted, where both identical and similar color pairs are considered to compensate color variations. Comparative experimental results demonstrate the performance of the proposed eCCH in terms of robustness to color variance and small computational complexity. Wenjing Jia, Xiangjian He, Qiang Wu 0001 |
ICASSP | 3 |
| 2010 | Motion blur detection based on lowest directional high-frequency energyabstractMotion blur detection and the relevant blurring parameter estimation are important for many computer vision tasks. The contribution of this paper is in two folds. First, we propose a closed-form solution for motion direction estimation on blurred image. Secondly, a novel method is proposed for motion blurred region detection. The proposed direction estimation is based on measurement of lowest directional high-frequency energy. Compared with traditional methods, it will improve accuracy with less computational cost. Moreover, the proposed motion blurred region detection can efficiently estimate blurred regions without Point Spread Function estimation. Encouraging results are shown by experiments. Jie Yang 0002, Qiang Wu 0001 |
ICIP | 3 |
| 2010 | Multi-view Gait Recognition Based on Motion Regression Using Multilayer PerceptronabstractIt has been shown that gait is an efficient biometric feature for identifying a person at a distance. However, it is a challenging problem to obtain reliable gait feature when viewing angle changes because the body appearance can be different under the various viewing angles. In this paper, the problem above is formulated as a regression problem where a novel View Transformation Model (VTM) is constructed by adopting Multilayer Perceptron (MLP) as regression tool. It smoothly estimates gait feature under an unknown viewing angle based on motion information in a well selected Region of Interest (ROI) under other existing viewing angles. Thus, this proposal can normalize gait features under various viewing angles into a common viewing angle before gait similarity measurement is carried out. Encouraging experimental results have been obtained based on widely adopted benchmark database. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
ICPR | 2 |
| 2010 | Context-aware fusion: A case study on fusion of gait and face for human identification in video
Xin Geng 0001, Kate Smith-Miles, Liang Wang 0001, Ming Li 0010, Qiang Wu 0001 |
Pattern Recognit. | 5 |
| 2009 | Automatic Gait Recognition Using Weighted Binary Pattern on VideoabstractHuman identification by recognizing the spontaneous gait recorded in real-world setting is a tough and not yet fully resolved problem in biometrics research. Several issues have contributed to the difficulties of this task. They include various poses, different clothes, moderate to large changes of normal walking manner due to carrying diverse goods when walking, and the uncertainty of the environments where the people are walking. In order to achieve a better gait recognition, this paper proposes a new method based on Weighted Binary Pattern (WBP). WBP first constructs binary pattern from a sequence of aligned silhouettes. Then, adaptive weighting technique is applied to discriminate significances of the bits in gait signatures. Being compared with most of existing methods in the literatures, this method can better deal with gait frequency, local spatial-temporal human pose features, and global body shape statistics. The proposed method is validated on several well known benchmark databases. The extensive and encouraging experimental results show that the proposed algorithm achieves high accuracy, but with low complexity and computational time. Worapan Kusakunniran, Qiang Wu 0001, Hongdong Li, Jian Zhang 0002 |
AVSS | 2 |
| 2009 | Facial expression recognition using histogram variances facesabstractIn human's expression recognition, the representation of expression features is essential for the recognition accuracy. In this work we propose a novel approach for extracting expression dynamic features from facial expression videos. Rather than utilising statistical models e.g. Hidden Markov Model (HMM), our approach integrates expression dynamic features into a static image, the Histogram Variances Face (HVF), by fusing histogram variances among the frames in a video. The HVFs can be automatically obtained from videos with different frame rates and immune to illumination interference. In our experiments, for the videos picturing the same facial expression, e.g., surprise, happy and sadness etc., their corresponding HVFs are similar, even though the performers and frame rates are different. Therefore the static facial recognition approaches can be utilised for the dynamic expression recognition. We have applied this approach on the well-known Cohn-Kanade AU-Coded Facial Expression database then classified HVFs using PCA and Support Vector Machine (SVMs), and found the accuracy of HVFs classification is very encouraging. Ruo Du, Qiang Wu 0001, Xiangjian He, Wenjing Jia, Daming Wei |
WACV | 2 |
| 2009 | Editorial
Liang Wang 0001, Qiang Wu 0001, Ming Li 0010, Jordi Gonzàlez 0001, Xin Geng 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2009 | Face recognition using message passing based clustering method
Chunhua Du, Jie Yang 0002, Qiang Wu 0001, Tianhao Zhang 0002 |
J. Vis. Commun. Image Represent. | 3 |
| 2009 | Image/video-based pattern analysis and HCI applications
Liang Wang 0001, Qiang Wu 0001, Hanzi Wang, Xin Geng 0001, Ming Li 0010 |
Pattern Recognit. Lett. | 2 |
| 2008 | Using dynamic programming to match human behavior sequencesabstractThis paper proposed a new approach for recognition and matching the human behavior sequence. Each human behavior sequence is represented by its key postures to greatly reduce the computation time. Normalization is applied to all the behavior sequences key postures for matching. A dynamic time warping (DTW) algorithm is used to perform the alignment of two time series. Experiments are carried out on an open human behavior database and exciting results have been obtained. Yan Chen 0020, Qiang Wu 0001, Xiangjian He |
ICARCV | 2 |
| 2008 | An approach of canny edge detection with virtual hexagonal image structureabstractEdge detection plays an important role in the areas of image processing, multimedia and computer vision. Gradient-based edge detection is a straightforward method to identify the edge points in the original grey-level image. It is intuitive that, in the human vision system, the edge points always appear where the gradient magnitude assumes a maximum. Hexagonal structure is an image structure alternative to traditional square image structure. The geometrical arrangement of pixels on a hexagonal structure can be described as a collection of hexagonal pixels. Because all the existing hardware for capturing image and for displaying image are produced based on square structure, an approach that uses bilinear interpolation and tri-linear interpolation is applied for conversion between square and hexagonal structures. Based on this approach, an edge detection method is proposed. This method performs Gaussian filtering to suppress image noise and computes gradients on the hexagonal structure. The pixel edge strengths on the square structure are then estimated before Canny' edge detector is applied to determine the final edge map. The experimental results show that the proposed method improves the edge detection accuracy and efficiency. Xiangjian He, Wenjing Jia, Qiang Wu 0001 |
ICARCV | 3 |
| 2008 | Extracting key postures in a human action video sequenceabstractHuman key posture extraction from videos will benefit video storage, video retrieval, human action recognition, human behaviour understanding and so on. This paper presents an approach to select key postures from human action sequences using 2D information. There are two steps in the proposed method. Information measurement which is a kind of global feature of a frame is used to roughly find key posture candidates. Then, a body skeleton feature which is a kind of local feature is applied to select final key postures from the candidates obtained in the first step. The experiments show that the proposed method is efficient. Yan Chen 0020, Qiang Wu 0001, Xiangjian He, Chunhua Du, Jie Yang 0002 |
MMSP | 2 |
| 2008 | Segmentation of characters on car license platesabstractLicense plate recognition usually contains three steps, namely license plate detection/localization, character segmentation and character recognition. When reading characters on a license plate one by one after license plate detection step, it is crucial to accurately segment the characters. The segmentation step may be affected by many factors such as license plate boundaries (frames). The recognition accuracy will be significantly reduced if the characters are not properly segmented. This paper presents an efficient algorithm for character segmentation on a license plate. The algorithm follows the step that detects the license plates using an AdaBoost algorithm. It is based on an efficient and accurate skew and slant correction of license plates, and works together with boundary (frame) removal of license plates. The algorithm is efficient and can be applied in real-time applications. The experiments are performed to show the accuracy of segmentation. Xiangjian He, Lihong Zheng, Qiang Wu 0001, Wenjing Jia, Bijan Samali, Marimuthu Palaniswami |
MMSP | 3 |
| 2008 | Pedestrian detection using hybrid statistical featureabstractA novel approach for walking people detection is proposed in this paper, which is inspired by the idea of gait energy image (GEI). Unlike most of common human detection methods where usually a trained detector scans a single image and then generates a detection result, the proposed method detects people on a sequence of silhouettes which contain both appearance characteristics and motion characteristics. Thus, our method is more robust. Encouraging experimental results are obtained based on CASIA gait database and the additional non-human objects data. Qiang Wu 0001, Chunhua Du, Jie Yang 0002, Xiangjian He, Yan Chen 0020 |
MMSP | 1 |
| 2008 | Adaptive Fusion of Gait and Face for Human Identification in VideoabstractMost work on multi-biometric fusion is based on static fusion rules which cannot respond to the changes of the environment and the individual users. This paper proposes adaptive multi-biometric fusion, which dynamically adjusts the fusion rules to suit the real-time external conditions. As a typical example, the adaptive fusion of gait and face in video is studied. Two factors that may affect the relationship between gait and face in the fusion are considered, i.e., the view angle and the subject-to-camera distance. Together they determine the way gait and face are fused at an arbitrary time. Experimental results show that the adaptive fusion performs significantly better than not only single biometric traits, but also those widely adopted static fusion rules including SUM, PRODUCT, MIN, and MAX. Xin Geng 0001, Liang Wang 0001, Ming Li 0010, Qiang Wu 0001, Kate Smith-Miles |
WACV | 4 |
| 2007 | Parallel Edge Detection on a Virtual Hexagonal Structure
Xiangjian He, Wenjing Jia, Qiang Wu 0001, Tom Hintz |
GPC | 3 |
| 2007 | Local Binary Patterns for Human Detection on Hexagonal StructureabstractLocal binary pattern (LBP) was designed and has been widely used for efficient texture classification. LBP provides a simple and effective way to represent texture patterns. Uniform LBPs play an important role for LBP-based pattern/object recognition as they include majority of LBPs. On the other hand, Human detection based on Mahalanobis distance map (MDM) recognizes appearance of human based on geometrical structure. Each MDM shows a clear texture pattern that can be classified using LBPs. In this paper, we compute LBPs of MDMs on a hexagonal structure. The circular pixel arrangement in hexagonal structure results in higher accuracy for LBP representation than on square structure. Chi-square as a measure is used for human detection based on uniform LBPs obtained. We show that our method using LBPs built on MDMs has a higher human detection rate and a lower false positive rate compared to the method merely based on MDMs. We will also show using experimental results that LBPs on hexagonal structure lead to more robust human classification. Xiangjian He, Yan Chen 0020, Qiang Wu 0001, Wenjing Jia |
ISM | 4 |
| 2006 | Symmetric Color Ratio in Spiral Architecture
Wenjing Jia, Huaifeng Zhang, Xiangjian He, Qiang Wu 0001 |
ACCV (2) | 4 |
| 2006 | Estimation of Internal and External Parameters for Camera Calibration Using 1D PatternabstractCamera calibration is to estimate the intrinsic and extrinsic parameters of a camera. Most of object-based calibration methods used 3D or 2D pattern. A novel and more flexible 1D object-based calibration was introduced only a couple of years ago, but merely for estimation of intrinsic parameters. The estimation of extrinsic papers is essential when multiple cameras are involved for simultaneously taking images from different view angles and when the knowledge of relative locations between the cameras is required. Though it is relatively simple using 2D or 3D calibration pattern, the estimation of extrinsic parameters is not obvious using 1D pattern. In this paper, we will perform a 1D camera calibration involving both intrinsic and extrinsic parameters. Xiangjian He, Huaifeng Zhang, Namho Hur, Jinwoong Kim, Qiang Wu 0001, Taeone Kim |
AVSS | 5 |
| 2006 | A Comparison on Histogram Based Image Matching MethodsabstractUsing colour histogram as a stable representation over change in view has been widely used for object recognition. In this paper, three newly proposed histogram-based methods are compared with other three popular methods, including conventional histogram intersection (HI) method, Wong and Cheung's merged palette histogram matching (MPHM) method, and Gevers' colour ratio gradient (CRG) method. These methods are tested on vehicle number plate images for number plate classification. Experimental results disclose that, the CRG method is the best choice in terms of speed, and the GWHI method can give the best classification results. Overall, the CECH method produces the best performance when both speed and classification performance are concerned. Wenjing Jia, Huaifeng Zhang, Xiangjian He, Qiang Wu 0001 |
AVSS | 4 |
| 2006 | Car Plate Detection Using Cascaded Tree-Style Learner Based on Hybrid Object FeaturesabstractCar plate detection is a key component in automatic license plate recognition system. This paper adopts an enhanced cascaded tree style learner framework for car plate detection using the hybrid object features including the simple statistical features and Harr-like features. The statistical features are useful for simplifying the process on cascade classifier. The cascaded tree-style detector design will further reduce the false alarm and the false dismissal while retaining a high detection ratio. The experimental results obtained by the proposed algorithm exhibit the encouraging performance. Qiang Wu 0001, Huaifeng Zhang, Wenjing Jia, Xiangjian He, Jie Yang 0002, Tom Hintz |
AVSS | 1 |
| 2006 | Uniformly Partitioning Images on Virtual Hexagonal StructureabstractHexagonal structure is different from the traditional square structure for image representation. The geometrical arrangement of pixels on hexagonal structure can be described in terms of a hexagonal grid. Uniformly separating image into seven similar copies with a smaller scale has commonly been used for parallel and accurate image processing on hexagonal structure. However, all the existing hardware for capturing image and for displaying image are produced based on square architecture. It has become a serious problem affecting the advanced research based on hexagonal structure. Furthermore, the current techniques used for uniform separation of images on hexagonal structure do not coincide with the rectangular shape of images. This has been an obstacle in the use of hexagonal structure for image processing. In this paper, we briefly review a newly developed virtual hexagonal structure that is scalable. Based on this virtual structure, algorithms for uniform image separation are presented. The virtual hexagonal structure retains image resolution during the process of image separation, and does not introduce distortion. Furthermore, images can be smoothly and easily transferred between the traditional square structure and the hexagonal structure while the image shape is kept in rectangle Xiangjian He, Huaqing Wang, Namho Hur, Wenjing Jia, Qiang Wu 0001, Jinwoong Kim, Tom Hintz |
ICARCV | 5 |
| 2006 | Learning-Based Number Recognition on Spiral ArchitectureabstractIn this paper, a number recognition algorithm is proposed on spiral architecture, a hexagonal image structure. This algorithm employs RULES-3 inductive learning method to recognize numbers. The algorithm starts from a collection of samples of numbers from number plates. Edge maps of the samples are then detected based on spiral architecture. A set of rules are extracted using these samples by RULES-3. The rules describe the frequencies of 9 different edge masks appearing in the samples. Each mask is a cluster of 7 hexagonal pixels. In order to recognize a number plate, all numbers are tested one by one using the extracted rules. The number recognition is achieved by counting the frequencies of the 9 masks. In this paper, a comparison between results based on rectangular structure and the results based on spiral architecture is given. From the experimental results, we can make the conclusion that Spiral Architecture is better than rectangular structure for inductive learning-based number recognition Lihong Zheng, Xiangjian He, Qiang Wu 0001, Tom Hintz |
ICARCV | 3 |
| 2006 | A New Approach for SA-Based Fractal Image CompressionabstractSpiral Architecture based fractal image compression is proposed in this paper. Perceptually, a new definition of range block and domain block is presented on such enhanced image structure. Compared with the common square image architecture, spiral architecture provides higher fidelity to fractal image compression, which is demonstrated by the experimental results. Huaqing Wang, Qiang Wu 0001, Xiangjian He, Tom Hintz |
ICIP | 2 |
| 2006 | Image Matching Using Colour Edge Cooccurrence HistogramsabstractIn this paper, a novel colour edge cooccurrence histogram (CECH) method is proposed to match images by measuring similarities between their CECH histograms. Unlike the previous colour edge cooccurrence histogram proposed by Crandall and Luo (2004 ) we only investigate those pixels which are located at the two sides of edge points in their gradient direction lines and at a distance away from the edge points. When measuring similarities between two CECH histograms, a newly proposed Gaussian weighted histogram intersection (GWHI) method is extended for this purpose. Both identical colour pairs and similar colour pairs are taken into account in our algorithm, and the weights are decided by the larger distance between two colour pairs involved in matching. The proposed algorithm is tested for matching vehicle number plate images captured under various illumination conditions. Experimental results demonstrate that the proposed algorithm can be used to compare images in real-time, and is robust to illumination variations and insensitive to the model images selected. Wenjing Jia, Huaifeng Zhang, Xiangjian He, Qiang Wu 0001 |
SMC | 4 |
| 2006 | A Fast Algorithm for License Plate Detection in Various ConditionsabstractThis paper proposes a fast algorithm detecting license plates in various conditions. There are three main contributions in this paper. The first contribution is that we define a new vertical edge map, with which the license plate detection algorithm is extremely fast. The second contribution is that we construct a cascade classifier which is composed of two kinds of classifiers. The classifiers based on statistical features decrease the complexity of the system. They are followed by the classifiers based on Haar-features, which make it possible to detect license plate in various conditions. Our algorithm is robust to the variance of the illumination, view angle, the position, size and color of the license plates when working in complex environment. The third contribution is that we experimentally analyze the relations of the scaling factor with detection rate and processing time. On the basis of the analysis, we select the optimal scaling factor in our algorithm. In the experiments, both high detection rate (with low false positive rate) and high speed are achieved when the algorithm is used to detect license plates in various complex conditions. Huaifeng Zhang, Wenjing Jia, Xiangjian He, Qiang Wu 0001 |
SMC | 4 |
| 2006 | Real-Time License Plate Detection Under Various Conditions
Huaifeng Zhang, Wenjing Jia, Xiangjian He, Qiang Wu 0001 |
UIC | 4 |
| 2005 | Bi-Lateral Filtering Based Edge Detection on Hexagonal ArchitectureabstractEdge detection plays an important role in image processing but is still an open problem. This paper presents a novel edge detection method based on bi-lateral filtering which achieves better performance than single Gaussian filtering. In this form of filtering, both spatial closeness and intensity similarity of pixels are considered in order to preserve important visual cues provided by edges and reduce the sharpness of transitions in intensity values as well. In addition, the edge detection method proposed in this paper is achieved on hexagonally sampled images. Due to the compact and circular nature of the hexagonal lattice, a better quality edge map is obtained on hexagonal architecture than common edge detection on square architecture. Experimental results using our proposed method in this paper exhibit encouraging performance. Qiang Wu 0001, Xiangjian He, Tom Hintz |
ICASSP (2) | 1 |
| 2005 | Modified Color Ratio GradientabstractColor ratio gradient is an efficient method used for color image retrieval and object recognition, which is shown to be illumination-independent and geometry-insensitive when tested on scenery images. However, color ratio gradient produces unsatisfied matching result while dealing with relatively uniform objects without rich color texture. In addition, performance of color ratio gradient degenerates while processing unsaturated color image objects. In this paper, a scheme with modified color ratio gradient is presented, which addresses the two problems above. Experimental results using the proposed method in this paper exhibit more robust performance Huaifeng Zhang, Wenjing Jia, Xiangjian He, Qiang Wu 0001 |
MMSP | 4 |
| 2003 | Complete Image Partitioning on Spiral Architecture
Qiang Wu 0001, Xiangjian He, Tom Hintz, Yuhuang Ye |
ISPA | 1 |
| 2001 | A Skeleton Algorithm on Clusters for Image Edge DetectionabstractImage edge detection in computer vision and image processing is a process which detects one kind of significant feature in an image that appears as large delta values in intensities. In this paper, a parallel algorithmic skeleton for edge detection is proposed based on the Spiral Architecture and the Gaussian multi-scale theory. UNIX-based network programming mechanisms in C are used for the implementation on a cluster of Sun-workstations. Our work provides an efficient algorithm for edge detection and is robust to noise. Xiangjian He, Tom Hintz, Qiang Wu 0001 |
IPDPS | 3 |