VLDB 2026 Research / reviewers in the wild / expert
Sheng Huang 0001
dblp:56/6585-1
· DBLP profile ↗
86ranked-venue papers
17as first author
54since 2021 · last 2026
0000-0001-5610-0826ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 7 first-author · 23 since 2021Artificial intelligence and machine learning · 33 · 7 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Computer networks · 2Systems, architecture and hardware · 1Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse-Scale Transformer with Bidirectional Awareness for Time Series ForecastingabstractTime series forecasting (TSF) plays a crucial role in many real-world applications, such as weather prediction and economic planning. While Transformer-based models have shown strong capabilities in modeling long-range dependencies, effectively capturing the multi-scale temporal dynamics inherent in time series remains a major challenge. Existing methods often adopt time-windows of varying sizes, which may introduce noisy or irrelevant representations when mismatched with the underlying temporal patterns, potentially leading to overfitting. In this paper, we propose Sparse-Scale Transformer (SSformer) with Bidirectional Awareness for Time Series Forecasting to enhance the multi-scale modeling for time series. Specifically, we propose a novel Sparse-Scale Convolution (SSC) block that imposes sparsity on scales to obtain the informative representations by evaluating the intra-scale segment similarity of time series, and utilizes scale-specific convolutions to extract local patterns. Furthermore, we design a Bidirectional-Scale Interaction (BSI) block to explicitly model scale correlations in both coarse-to-fine and fine-to-coarse directions. Finally, scale predictions are ensembled to fully exploit the complementary forecasting capabilities across scales. Extensive experiments on various real-world datasets demonstrate that SSformer achieves state-of-the-art performance with superior efficiency. Ying Liu 0096, Bo Liu 0005, Sheng Huang 0001, Wenbo Hu 0001, Meng Wang 0001, Richang Hong |
AAAI | 3 |
| 2026 | Multiple Instance Learning Framework with Masked Hard Instance Mining for Gigapixel Histopathology Image Analysis
Sheng Huang 0001, Fengtao Zhou, Bo Liu 0005, Qingshan Liu 0001 |
Int. J. Comput. Vis. | 2 |
| 2026 | Noisy label effects for out-of-distribution detection in single-positive multi-label settings
Yi Zhang 0113, Xiaohong Zhang 0002, Dan Yang 0001, Sheng Huang 0001 |
Pattern Anal. Appl. | 5 |
| 2026 | HDG-CLIP: Hierarchical Dual-Granularity Vision-Semantic Alignment for Open-Vocabulary Multi-Label Image ClassificationabstractOpen-Vocabulary Multi-Label Image Classification (OV-MLIC) is an emerging task in computer vision aimed at recognizing unseen categories in real-world scenarios, leveraging Vision and Language Pre-training (VLP) models like CLIP. However, existing methods overlook the impact of category coupling and scale variation on cross-category knowledge transfer, thereby restricting performance on unseen categories. To address this issue, we propose a novel OV-MLIC method called Hierarchical Dual-Granularity Alignment-CLIP (HDG-CLIP), which emphasizes the complementary characteristics of different modalities and introduces a sample-category matching mechanism. Specifically, to address the category coupling issue, we construct semantic category prototypes to enhance cross-category knowledge transfer. Through the interaction between visual embeddings and category prototypes, we decouple category-specific information from mixed visual features and leverage the visual context of samples to learn category-level visual features. For mitigating the scale variation issue, we build a sample-category dual-granularity matching mechanism based on the difference in capture capability of different modalities across scales, thereby improving the object localization accuracy from a multi-dimensional perspective. Extensive experimental results show that HDG-CLIP exhibits state-of-art performance over existing methods on both the NUS-WIDE and the Open-Images datasets. Our code is available at https://github.com/wakihy/HDG-CLIP. Beiyan Liu, Sheng Huang 0001, Bo Liu 0005, Xiaobin Huang, Nankun Mu, Richang Hong, Meng Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Scan-Invariant Mamba With Differentiated Sequence Contrastive Learning in Computational PathologyabstractMultiple instance learning (MIL) is a commonly used paradigm for histopathological analysis due to the ultra-high resolution and coarse-grained labels of Whole Slide Images (WSIs). Recent studies apply Mamba architecture to WSI classification by modeling MIL as long-sequence tasks, but a key discrepancy remains: Mamba's output is sensitive to scanning modes, whereas MIL requires scan-invariant predictions. To address this problem, we propose Scan-invariant Mamba with Differentiated Sequence Contrastive Learning (SMDC-MIL), a novel Mamba-based MIL approach enabling bag-level feature learning independent of input modes. Our method mitigates scanning-mode impacts and adapts Mamba to learn the bag discrimination features that are independent of the input mode via two innovations: 1) a differentiated sequence generation mechanism that employs instance rearrangement, augmentation, and masking to simulate real-world scanning variations by maximizing differences in sequence order, length, and composition from the same WSI; and 2) a differentiated sequence contrastive learning architecture that enforces consistent bag-level representations and predictions across diverse sequences using the same Mamba model, guiding it to prioritize scan-invariant discriminative features. Experimental results on 4 computational pathology tasks and 10 datasets demonstrate that our SMDC-MIL achieves state-of-the-art performance compared to other methods. The corresponding code is available at https://github.com/LianYueZ/SMDCMIL.git. Sheng Huang 0001, Xin Zhang 0131, Bo Liu 0005, Fengtao Zhou, Kang Li 0004, Hao Chen 0011, Meng Wang 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Semantic Guided Dual-Branch Co-inference for Few-Shot 3D Point Cloud Classification
Sheng Huang 0001, Jiexuan Yan, Xin Zhang 0131, Nankun Mu |
PRCV (5) | 2 |
| 2025 | DDFM: A Damage Detail Fusion Model Based on MambaOut for Pavement Damage Classification
Shizheng Zhang, Meng Zeng, Min Huang 0014, Sheng Huang 0001 |
PRCV (12) | 5 |
| 2025 | Learning complementary visual information for few-shot food recognition by Regional Erasure and Reactivation
Yi Zhang 0113, Luwen Huangfu, Lili Balazs, Sheng Huang 0001 |
Expert Syst. Appl. | 5 |
| 2025 | Rethinking the sample relations for few-shot classification
Guowei Yin, Sheng Huang 0001, Luwen Huangfu, Yi Zhang 0113, Xiaohong Zhang 0002 |
Image Vis. Comput. | 2 |
| 2025 | Enhanced prototype network with gated point recyclable feature mining for few-shot 3D point cloud classification
Sheng Huang 0001, Luwen Huangfu, Ma Rui, Bo Liu 0005 |
Knowl. Based Syst. | 2 |
| 2025 | Adaptive learning of instance representatives in dual spaces for medical image classification
Sheng Huang 0001, Yi Zhang 0113, Xiaoxian Zhang, Chen Liu 0026, Xiahong Zhang |
Neural Comput. Appl. | 2 |
| 2025 | Dual-View Alignment Learning With Hierarchical-Prompt for Class-Imbalance Multi-Label Image ClassificationabstractReal-world datasets often exhibit class imbalance across multiple categories, manifesting as long-tailed distributions and few-shot scenarios. This is especially challenging in Class-Imbalanced Multi-Label Image Classification (CI-MLIC) tasks, where data imbalance and multi-object recognition present significant obstacles. To address these challenges, we propose a novel method termed Dual-View Alignment Learning with Hierarchical Prompt (HP-DVAL), which leverages multi-modal knowledge from vision-language pretrained (VLP) models to mitigate the class-imbalance problem in multi-label settings. Specifically, HP-DVAL employs dual-view alignment learning to transfer the powerful feature representation capabilities from VLP models by extracting complementary features for accurate image-text alignment. To better adapt VLP models for CI-MLIC tasks, we introduce a hierarchical prompt-tuning strategy that utilizes global and local prompts to learn task-specific and context-related prior knowledge. Additionally, we design a semantic consistency loss during prompt tuning to prevent learned prompts from deviating from general knowledge embedded in VLP models. The effectiveness of our approach is validated on two CI-MLIC benchmarks: MS-COCO and VOC2007. Extensive experimental results demonstrate the superiority of our method over SOTA approaches, achieving mAP improvements of 10.0% and 5.2% on the long-tailed multi-label image classification task, and 6.8% and 2.9% on the multi-label few-shot image classification task. Sheng Huang 0001, Jiexuan Yan, Beiyan Liu, Bo Liu 0005, Richang Hong |
IEEE Trans. Image Process. | 1 |
| 2025 | Learning the Difference of Few-Shot Food Data Using Multivariate Knowledge-Guided Variational AutoencoderabstractRecent advancements in food image recognition have underscored its importance in dietary monitoring, which promotes a healthy lifestyle and aids in the prevention of diseases such as diabetes and obesity. While mainstream food recognition methods excel in scenarios with large-scale annotated datasets, they falter in few-shot regimes where data is limited. This paper addresses this challenge by introducing a variational generative method, the Multivariate Knowledge-guided Variational AutoEncoder (MK-VAE), for few-shot food recognition. MK-VAE leverages handcrafted features and semantic embeddings as multivariate prior knowledge to strengthen feature learning and feature generation in different phases. Specifically, we design a lightweight and flexible feature distillation module that distills handcrafted features to enhance the feature learning network for capturing the salient visual information in few-shot samples. During the feature generation phase, we utilize a variational autoencoder to learn the difference distribution of food data and explicitly boost the latent representation with category-level semantic embeddings to pull homogeneous features closer together while pushing inhomogeneous features apart. Experimental results demonstrate that our proposed MK-VAE significantly outperforms state-of-the-art few-shot food recognition methods in both 5-way 1-shot and 5-way 5-shot settings on three widely-used benchmark datasets: Food-101, VIREO Food-172, and UECFood-256. Yi Zhang 0113, Sheng Huang 0001, Mingjian Hong, Dan Yang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Feature Noise Resilient for QoS Prediction With Probabilistic Deep SupervisionabstractAccurate Quality of Service (QoS) prediction is essential for enhancing user satisfaction in web recommendation systems, yet existing prediction models often overlook feature noise, focusing predominantly on label noise. In this paper, we present the Probabilistic Deep Supervision Network (PDS-Net), a robust framework designed to effectively identify and mitigate feature noise, thereby improving QoS prediction accuracy. PDS-Net operates with adual-branch architecture: the main branch utilizes a decoder network to learn a Gaussian-based prior distribution from known features, while the second branch derives a posterior distribution based on true labels. A key innovation of PDS-Net is its condition-based noise recognition loss function, which enables precise identification of noisy features in objects (users or services). Once noisy features are identified, PDS-Net refines the feature's prior distribution, aligning it with the posterior distribution, and propagates this adjusted distribution to intermediate layers, effectively reducing noise interference. Extensive experiments conducted on two real-world QoS datasets demonstrate that PDS-Net consistently outperforms existing models, achieving an average improvement of 8.91% in MAE on Dataset D1 and 8.32% on Dataset D2 compared to the state-of-the-art. These results highlight PDS-Net's ability to accurately capture complex user-service relationships and handle feature noise, underscoring its robustness and versatility across diverse QoS prediction environments. Xiaohong Zhang 0002, Ze Shi Li, Sheng Huang 0001, Meng Yan 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | Learning Feature Exploration and Selection With Handcrafted Features for Few-Shot LearningabstractInterest in few-shot learning (FSL) has grown recently, but the value of feature learning, which bridges the gap between base and novel classes, remains largely understudied. The limited availability of labeled samples for each class poses a major challenge. To tackle this, we propose a simple yet effective approach called deep discriminative handcrafted feature regression (DDHFR) to explore intrinsic information and select improved discriminative features in few-shot data by mining knowledge from classical handcrafted features. To explore intrinsic information, we design several deep handcrafted feature regression (DHFR) modules and plugged them separately into different layers of the backbone to use feature engineering knowledge for feature learning optimization at different granularities. To achieve discriminative feature selection, we incorporate an auxiliary classifier (AC) into each DHFR module to enhance the acquisition of discriminative information. Furthermore, we employed self-distillation to boost ability of ACs ot be classified. Experimental results in three backbones on three datasets show that DDHFR can generally improve the performance of existing FSL methods. On average, it improves the recognition accuracy by 1.16% in two common few-shot settings. Yi Zhang 0113, Sheng Huang 0001, Luwen Huangfu, Daniel Dajun Zeng |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | Data Distribution Distilled Generative Model for Generalized Zero-Shot RecognitionabstractIn the realm of Zero-Shot Learning (ZSL), we address biases in Generalized Zero-Shot Learning (GZSL) models, which favor seen data. To counter this, we introduce an end-to-end generative GZSL framework called D3GZSL. This framework respects seen and synthesized unseen data as in-distribution and out-of-distribution data, respectively, for a more balanced model. D3GZSL comprises two core modules: in-distribution dual space distillation (ID2SD) and out-of-distribution batch distillation (O2DBD). ID2SD aligns teacher-student outcomes in embedding and label spaces, enhancing learning coherence. O2DBD introduces low-dimensional out-of-distribution representations per batch sample, capturing shared structures between seen and un seen categories. Our approach demonstrates its effectiveness across established GZSL benchmarks, seamlessly integrating into mainstream generative frameworks. Extensive experiments consistently showcase that D3GZSL elevates the performance of existing generative GZSL methods, under scoring its potential to refine zero-shot learning practices. The code is available at: https://github.com/PJBQ/D3GZSL.git Mingjian Hong, Luwen Huangfu, Sheng Huang 0001 |
AAAI | 4 |
| 2024 | Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyabstractMultiple instance learning (MIL) is the most widely used framework in computational pathology, encompassing sub-typing, diagnosis, prognosis, and more. However, the ex-isting MIL paradigm typically requires an offline instance feature extractor, such as a pre-trained ResNet or a foun-dation model. This approach lacks the capability for feature fine-tuning within the specific downstream tasks, limiting its adaptability and performance. To address this issue, we propose a Re-embedded Regional Transformer (R2T) for re-embedding the instance features online, which captures fine-grained local features and establishes connections across different regions. Unlike existing works that focus on pre-training powerful feature extractor or designing sophisticated instance aggregator, R2T is tailored to re-embed instance features online. It serves as a portable module that can seamlessly integrate into mainstream MIL models. Extensive experimental results on common computational pathology tasks validate that: 1) feature re-embedding improves the performance of MIL models based on ResNet-50 features to the level of foundation model features, and further enhances the performance of foundation model features; 2) the R2T can introduce more signifi-cant performance improvements to various MIL models; 3) R2T-MIL, as an R2T-enhanced AB-MIL, outperforms other latest methods by a large margin. The code is available at: https://github.com/DearCaat/RRT-MIL. Fengtao Zhou, Sheng Huang 0001, Yi Zhang 0113, Bo Liu 0005 |
CVPR | 3 |
| 2024 | SAM-MIL: A Spatial Contextual Aware Multiple Instance Learning Approach for Whole Slide Image ClassificationabstractMultiple Instance Learning (MIL) represents the predominant framework in Whole Slide Image (WSI) classification, covering aspects such as sub-typing, diagnosis, and beyond. Current MIL models predominantly rely on instance-level features derived from pretrained models such as ResNet. These models segment each WSI into independent patches and extract features from these local patches, leading to a significant loss of global spatial context and restricting the model's focus to merely local features. To address this issue, we propose a novel MIL framework, named SAM-MIL, that emphasizes spatial contextual awareness and explicitly incorporates spatial context by extracting comprehensive, image-level information. The Segment Anything Model (SAM) represents a pioneering visual segmentation foundational model that can capture segmentation features without the need for additional fine-tuning, rendering it an outstanding tool for extracting spatial context directly from raw WSIs. Our approach includes the design of group feature extraction based on spatial context and a SAM-Guided Group Masking strategy to mitigate class imbalance issues. We implement a dynamic mask ratio for different segmentation categories and supplement these with representative group features of categories. Moreover, SAM-MIL divides instances to generate additional pseudo-bags, thereby augmenting the training set, and introduces consistency of spatial context across pseudo-bags to further enhance the model's performance. Experimental results on the CAMELYON-16 and TCGA Lung Cancer datasets demonstrate that our proposed SAM-MIL model outperforms existing mainstream methods in WSIs classification. Our open-source implementation code is is available at https://github.com/FangHeng/SAM-MIL. Sheng Huang 0001, Luwen Huangfu, Bo Liu 0005 |
ACM Multimedia | 2 |
| 2024 | Category-Prompt Refined Feature Learning for Long-Tailed Multi-Label Image ClassificationabstractReal-world data consistently exhibits a long-tailed distribution, often spanning multiple categories. This complexity underscores the challenge of content comprehension, particularly in scenarios requiring Long-Tailed Multi-Label image Classification (LTMLC). In such contexts, imbalanced data distribution and multi-object recognition pose significant hurdles. To address this issue, we propose a novel and effective approach for LTMLC, termed Category-Prompt Refined Feature Learning (CPRFL), utilizing semantic correlations between different categories and decoupling category-specific visual representations for each category. Specifically, CPRFL initializes category-prompts from the pretrained CLIP's embeddings and decouples category-specific visual representations through interaction with visual features, thereby facilitating the establishment of semantic correlations between the head and tail classes. To mitigate the visual-semantic domain bias, we design a progressive Dual-Path Back-Propagation mechanism to refine the prompts by progressively incorporating context-related visual information into prompts. Simultaneously, the refinement process facilitates the progressive purification of the category-specific visual representations under the guidance of the refined prompts. Furthermore, taking into account the negative-positive sample imbalance, we adopt the Asymmetric Loss as our optimization objective to suppress negative samples across all classes and potentially enhance the head-to-tail recognition performance. We validate the effectiveness of our method on two LTMLC benchmarks and extensive experiments demonstrate the superiority of our work over baselines.The code is available at https://github.com/jiexuanyan/CPRFL. Jiexuan Yan, Sheng Huang 0001, Nankun Mu, Luwen Huangfu, Bo Liu 0005 |
ACM Multimedia | 2 |
| 2024 | Text kernel expansion for real-time scene text detection
Sheng Huang 0001, Bo Liu 0005 |
Pattern Anal. Appl. | 2 |
| 2024 | Mirrored EAST: An Efficient Detector for Automatic Vehicle Identification Number Detection in the WildabstractVehicle identification number (VIN) is a unique serial number used to identify individual vehicles across various applications. The first crucial step in automatically collecting VINs is to accurately localize the VIN area. In this article, we present a novel VIN detection approach called Mirrored EAST (MEAST) based on an efficient and accurate scene text (EAST) Detection framework. MEAST learns to exploit the spatial consistency between an image and its mirrored version to improve localization performance, and employs a lighter but more discriminative backbone network to improve its applicability in mobile scenarios. To evaluate the VIN detection performance, we constructed a large-scale VIN image dataset named CQU-VD20 K, consisting of 20 000 VIN images in real scenarios. Based on this dataset, we have conducted a comprehensive empirical study of VIN detection. The results demonstrate the superiority of MEAST over other methods in VIN detection. Additionally, we also conducted extended experiments on a license plate dataset named CCPD-Rotate, which confirms the effectiveness of our approach in other industrial inspection tasks. Guowei Yin, Sheng Huang 0001, Jin Xie 0005, Dan Yang 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Self-Supervised Adversarial Learning for Domain Adaptation of Pavement Distress ClassificationabstractPavement distress classification is crucial for the maintenance of highways. Although many methods for classifying pavement distress are available, they all assume that training and testing datasets are drawn from the same distribution. When we introduce a new unlabeled dataset with a different distribution, the performance of existing methods decreases considerably due to domain shift, motivating us to look beyond the supervised setting to utilize unlabeled datasets directly in training a model. Therefore, we develop a novel unsupervised domain adaptation (UDA) framework, namely, the Self-supervised Adversarial Network (SSAN) for the first time in this study to conduct multi-category pavement distress classification on an unlabeled target domain. In particular, SSAN leverages adversarial domain adaptation (ADA) thoughts to align the features of different domains. However, distress typically occupies a small Section of high-resolution pavement images. Consequently, aligning features directly is unreasonable because the aligning procedure is still dominated by background features instead of foreground features, which are the most useful information for classification. Therefore, we design a pretext module, called Self-supervised Learning for the Target domain (SLT), to mine foreground information. To validate our method, we use two challenging pavement crack datasets, namely, the Chonqing University Bituminous Pavement Disease Detection (CQU-BPDD) and the Chongqing University Bituminous Pavement Multi-label Disease Detection (CQU-BPMDD) datasets. Moreover, extensive experiments demonstrate that SSAN outperforms state-of-the-art UDA methods. Yanwen Wu, Mingjian Hong, Sheng Huang 0001, Yongxin Ge |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Semi-Identical Twins Variational AutoEncoder for Few-Shot LearningabstractData augmentation is a popular way for few-shot learning (FSL). It generates more samples as supplements and then transforms the FSL task into a common supervised learning problem for a solution. However, most data-augmentation-based FSL approaches only consider the prior visual knowledge for feature generation, thereby leading to low diversity and poor quality of generated data. In this study, we attempt to address this issue by incorporating both prior visual and prior semantic knowledge to condition the feature generation process. Inspired by some genetic characteristics of semi-identical twins, a novel multimodal generative FSL approach was developed named semi-identical twins variational autoencoder (STVAE) to better exploit the complementarity of these modality information by considering the multimodal conditional feature generation process as a process that semi-identical twins are born and collaborate to simulate their father. STVAE conducts feature synthesis by pairing two conditional variational autoencoders (CVAEs) with the same seed but different modality conditions. Subsequently, the generated features of two CVAEs are considered as semi-identical twins and adaptively combined to yield the final feature, which is considered as their fake father. STVAE requires that the final feature can be converted back into its paired conditions while ensuring these conditions remain consistent with the original in both representation and function. Moreover, STVAE is able to work in the partial modality-absence case due to the adaptive linear feature combination strategy. STVAE essentially provides a novel idea to exploit the complementarity of different modality prior information inspired by genetics in FSL. Extensive experimental results demonstrate that our work achieves promising performances in comparison to the recent state-of-the-art approaches, as well as validate its effectiveness on FSL under various modality settings. Yi Zhang 0113, Sheng Huang 0001, Xi Peng 0005, Dan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Multiple Instance Learning Framework with Masked Hard Instance Mining for Whole Slide Image ClassificationabstractThe whole slide image (WSI) classification is often formulated as a multiple instance learning (MIL) problem. Since the positive tissue is only a small fraction of the gigapixel WSI, existing MIL methods intuitively focus on identifying salient instances via attention mechanisms. However, this leads to a bias towards easy-to-classify instances while neglecting hard-to-classify instances. Some literature has revealed that hard examples are beneficial for modeling a discriminative boundary accurately. By applying such an idea at the instance level, we elaborate a novel MIL framework with masked hard instance mining (MHIM-MIL), which uses a Siamese structure (Teacher-Student) with a consistency constraint to explore the potential hard instances. With several instance masking strategies based on attention scores, MHIM-MIL employs a momentum teacher to implicitly mine hard instances for training the student model, which can be any attention-based MIL model. This counter-intuitive strategy essentially enables the student to learn a better discriminating boundary. Moreover, the student is used to update the teacher with an exponential moving average (EMA), which in turn identifies new hard instances for subsequent training iterations and stabilizes the optimization. Experimental results on the CAMELYON-16 and TCGA Lung Cancer datasets demonstrate that MHIM-MIL outperforms other latest methods in terms of performance and training cost. The code is available at: https://github.com/DearCaat/MHIM-MIL. Sheng Huang 0001, Xiaoxian Zhang, Fengtao Zhou, Yi Zhang 0113, Bo Liu 0005 |
ICCV | 2 |
| 2023 | ASDFL: An adaptive super-pixel discriminative feature-selective learning for vehicle matchingabstractAbstract There are a large number of cameras in modern transportation system that capture numerous vehicle images continuously. Therefore, automatic analysis of these vehicle images is helpful for traffic flow management, criminal investigations and vehicle inspections. Vehicle matching, which aims to determine whether two input images depict an identical vehicle, is one of the core tasks in vehicle analysis. Recent relevant studies have focused on local feature extraction instead of global extraction, since local details can provide crucial cues to distinguish between cars. However, these methods do not select local features; that is, they do not assign weights to local features. Therefore, in this research, we systematically study the vehicle matching task, and present a novel annotation‐free local‐based deep learning method called Adaptive super‐pixel discriminative feature‐selective learning (ASDFL) to address this issue. In ASDFL, vehicle images are segmented into clusters of super‐pixels of similar size by considering the location and colour similarities of pixels without using any component‐level annotation. These super‐pixels are deemed to be the virtual components of vehicles. Moreover, a convolutional neural network is used to extract the deep features of these virtual components. Thereafter, an instance‐specific mask generation module driven by the extracted global features is enhanced to produce a mask to select the most distinctive virtual components of each vehicle image pair in the feature space. Finally, the vehicle matching task is accomplished by classifying the selected virtual component features of each imaged vehicle pair. Extensive experiments on two popular vehicle identification benchmarks demonstrate that our method is 1.57% and 0.8% more accurate than the previous baselines in a vehicle matching task on the VeRi and VehicleID datasets, respectively, which demonstrates the effectiveness of our method. Rong Qin 0001, Huanhuan Lv, Yi Zhang 0113, Luwen Huangfu, Sheng Huang 0001 |
Expert Syst. J. Knowl. Eng. | 5 |
| 2023 | Anchor-based discriminative dual distribution calibration for transductive zero-shot learning
Yi Zhang 0113, Sheng Huang 0001, Xiaohong Zhang 0002, Dan Yang 0001 |
Image Vis. Comput. | 2 |
| 2023 | MSTIL: Multi-cue Shape-aware Transferable Imbalance Learning for effective graphic API recommendationabstractApplication Programming Interface (API) recommendation based on graphs is a valuable task in the fields of data visualization and software engineering. However, this task was previously undefined until a recently published paper coining the task as Plot2API and utilizing a deep learning-based method named SPGNN. Compared to general image classification methods, this dedicated approach uses semantic parsing to exploit deep features and yields better performance. However, its performance declines sharply in unbalanced datasets, thus limiting its generalizability. To address this issue, we propose a method named Multi-cue Shape and software engineering-aware Transferable Imbalance Learning (MSTIL), consisting of three major components: Cross-Language Shape-Aware Plot Transfer Learning (CLSAPTL), Cross-Language API Semantic Similarity-based Data Augmentation (CLASSDA), and Imbalance Plot2API Learning (IPL). Motivated by the hierarchical classification of the graphs, CLSAPTL guides the model to learn the graphs’ class hierarchy and thereby enabling the model to learn more transferable visual features. Given that a graph can be associated with multiple APIs and motivated by the fact that many APIs that exert similar functions in different languages have semantically similar names, CLASSDA leverages the samples of APIs with semantically similar names to assist in feature learning. Finally, inspired by the essence of softmax cross entropy loss, IPL alleviates the imbalances between positive and negative samples during training. We conduct our experiments on two public datasets. Extensive experimental results shows that MSTIL improves the performance of classic CNNs along with the state-of-the-art method, demonstrating its effectiveness. Specifically, MSTIL has an average relative mAP improvement of 12.94% across the models on all datasets. Rong Qin 0001, Zeyu Wang 0001, Sheng Huang 0001, Luwen Huangfu |
J. Syst. Softw. | 3 |
| 2023 | GSAL: Geometric structure adversarial learning for robust medical image segmentationabstractAutomatic medical image segmentation plays a crucial role in clinical diagnosis and treatment. However, it is still a challenging task due to the complex interior characteristics ( e.g. , inconsistent intensity, low contrast, texture heterogeneity) and ambiguous external boundary structures. In this paper, we introduce a novel geometric structure learning mechanism (GSLM) to overcome the limitations of existing segmentation models that lack learning ”focus, path, and difficulty.” The geometric structure in this mechanism is jointly characterized by the skeleton-like structure extracted by the mask distance transform (MDT) and the boundary structure extracted by the mask distance inverse transform (MDIT). Among them, the skeleton-like and boundary pay attention to the trend of interior characteristics consistency and external structure continuity, respectively. With this idea, we design GSAL, a novel end-to-end geometric structure adversarial learning for robust medical image segmentation. GSAL has four components: a geometric structure generator, which yields the geometric structure to learn the most discriminative features that preserve interior characteristics consistency and external boundary structure continuity, skeleton-like and boundary structure discriminators , which enhance and correct the characterization of internal and external geometry to mutually promote the capture of global contextual dependencies, and a geometric structure fusion sub-network, which fuses the two complementary and refined skeleton-like and boundary structures to generate the high-quality segmentation results. The proposed approach has been successfully applied to three different challenging medical image segmentation tasks , including polyp segmentation , COVID-19 lung infection segmentation, and lung nodule segmentation. Extensive experimental results demonstrate that the proposed GSAL achieves favorably against most state-of-the-art methods under different evaluation metrics . The code is available at: https://github.com/DLWK/GSAL . Kun Wang 0021, Xiaohong Zhang 0002, Sheng Huang 0001, Dan Yang 0001 |
Pattern Recognit. | 5 |
| 2023 | Adaptively Weighted k-Tuple Metric Network for Kinship VerificationabstractFacial image-based kinship verification is a rapidly growing field in computer vision and biometrics. The key to determining whether a pair of facial images has a kin relation is to train a model that can enlarge the margin between the faces that have no kin relation while reducing the distance between faces that have a kin relation. Most existing approaches primarily exploit duplet (i.e., two input samples without cross pair) or triplet (i.e., single negative pair for each positive pair with low-order cross pair) information, omitting discriminative features from multiple negative pairs. These approaches suffer from weak generalizability, resulting in unsatisfactory performance. Inspired by human visual systems that incorporate both low-order and high-order cross-pair information from local and global perspectives, we propose to leverage high-order cross-pair features and develop a novel end-to-end deep learning model called the adaptively weighted k -tuple metric network (AW k -TMN). Our main contributions are three-fold. First, a novel cross-pair metric learning loss based on k -tuplet loss is introduced. It naturally captures both the low-order and high-order discriminative features from multiple negative pairs. Second, an adaptively weighted scheme is formulated to better highlight hard negative examples among multiple negative pairs, leading to enhanced performance. Third, the model utilizes multiple levels of convolutional features and jointly optimizes feature and metric learning to further exploit the low-order and high-order representational power. Extensive experimental results on three popular kinship verification datasets demonstrate the effectiveness of our proposed AW k -TMN approach compared with several state-of-the-art approaches. The source codes and models are released.1. Sheng Huang 0001, Jingkai Lin, Luwen Huangfu, Junlin Hu 0001, Daniel Dajun Zeng |
IEEE Trans. Cybern. | 1 |
| 2023 | Weakly Supervised Patch Label Inference Networks for Efficient Pavement Distress Detection and Recognition in the WildabstractAutomatic image-based pavement distress detection and recognition are vital for pavement maintenance and management. However, existing deep learning-based methods largely omit the specific characteristics of pavement images, such as high image resolution and low distress area ratio, and are not end-to-end trainable. In this paper, we present a series of simple yet effective end-to-end deep learning approaches named Weakly Supervised Patch Label Inference Networks (WSPLIN) for efficiently addressing these tasks under various application settings. WSPLIN transforms the fully supervised pavement image classification problem into a weakly supervised pavement patch classification problem for solutions. Specifically, WSPLIN first divides the pavement image under different scales into patches with different collection strategies and then employs a Patch Label Inference Network (PLIN) to infer the labels of these patches to fully exploit the resolution and scale information. Notably, we design a patch label sparsity constraint based on the prior knowledge of distress distribution and leverage the Comprehensive Decision Network (CDN) to guide the training of PLIN in a weakly supervised way. Therefore, the patch labels produced by PLIN provide interpretable intermediate information, such as the rough location and the type of distress. We evaluate our method on a large-scale bituminous pavement distress dataset named CQU-BPDD and the augmented Crack500 (Crack500-PDD) dataset, which is a newly constructed pavement distress detection dataset augmented from the Crack500. Extensive results demonstrate the superiority of our method over baselines in both performance and efficiency. The source codes of WSPLIN are released onhttps://github.com/DearCaat/wsplin. Sheng Huang 0001, Guixin Huang, Luwen Huangfu, Dan Yang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Deep Domain Adaptation for Pavement Crack DetectionabstractDeep learning-based pavement cracks detection methods often require large-scale labels with detailed crack location information to learn accurate predictions. In practice, however, crack locations are very difficult to be manually annotated due to various visual patterns of pavement crack. In this paper, we propose a Deep Domain Adaptation-based Crack Detection Network (DDACDN), which learns domain invariant features by taking advantage of the source domain knowledge to predict the multi-category crack location information in the target domain, where only image-level labels are available. Specifically, DDACDN first extracts crack features from both the source and target domain by a two-branch weights-shared backbone network. And in an effort to achieve the cross-domain adaptation, an intermediate domain is constructed by aggregating the three-scale features from the feature space of each domain to adapt the crack features from the source domain to the target domain. Finally, the network involves the knowledge of both domains and is trained to recognize and localize pavement cracks. To facilitate accurate training and validation for domain adaptation, we use two challenging pavement crack datasets CQU-BPDD and RDD2020. Furthermore, we construct a new large-scale Bituminous Pavement Multi-label Disease Dataset named CQU-BPMDD, which contains 38994 high-resolution pavement disease images to further evaluate the robustness of our model. Extensive experiments demonstrate that DDACDN outperforms state-of-the-art pavement crack detection methods in predicting the crack location on the target domain. Chunhua Yang 0003, Sheng Huang 0001, Zhimin Ruan, Yongxin Ge |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Dual Space Multiple Instance Representative Learning for Medical Image Classification
Xiaoxian Zhang, Sheng Huang 0001, Yi Zhang 0113, Xiaohong Zhang 0002, Mingchen Gao, Chen Liu 0026 |
BMVC | 2 |
| 2022 | Kernel Inversed Pyramidal Resizing Network for Efficient Pavement Distress Recognition
Rong Qin 0001, Luwen Huangfu, Devon Hood, James Ma, Sheng Huang 0001 |
ICONIP (6) | 5 |
| 2022 | Boosting Multi-Label Image Classification with Complementary Parallel Self-DistillationabstractMulti-Label Image Classification (MLIC) appro-aches usually exploit label correlations to achieve good performance. However, emphasizing correlation like co-occurrence may overlook discriminative features and lead to model overfitting. In this study, we propose a generic framework named Parallel Self-Distillation (PSD) for boosting MLIC models. PSD decomposes the original MLIC task into several simpler MLIC sub-tasks via two elaborated complementary task decomposition strategies named Co-occurrence Graph Partition (CGP) and Dis-occurrence Graph Partition (DGP). Then, the MLIC models of fewer categories are trained with these sub-tasks in parallel for respectively learning the joint patterns and the category-specific patterns of labels. Finally, knowledge distillation is leveraged to learn a compact global ensemble of full categories with these learned patterns for reconciling the label correlation exploitation and model overfitting. Extensive results on MS-COCO and NUS-WIDE datasets demonstrate that our framework can be easily plugged into many MLIC approaches and improve performances of recent state-of-the-art approaches. The source code is released at https://github.com/Robbie-Xu/CPSD. Jiazhi Xu, Sheng Huang 0001, Fengtao Zhou, Luwen Huangfu, Daniel Dajun Zeng, Bo Liu 0005 |
IJCAI | 2 |
| 2022 | PicT: A Slim Weakly Supervised Vision Transformer for Pavement Distress ClassificationabstractAutomatic pavement distress classification facilitates improving the efficiency of pavement maintenance and reducing the cost of labor and resources. A recently influential branch of this task divides the pavement image into patches and infers the patch labels for addressing these issues from the perspective of multi-instance learning. However, these methods neglect the correlation between patches and suffer from a low efficiency in the model optimization and inference. As a representative approach of vision Transformer, Swin Transformer is able to address both of these issues. It first provides a succinct and efficient framework for encoding the divided patches as visual tokens, then employs self-attention to model their relations. Built upon Swin Transformer, we present a novel vision Transformer named Pavement Image Classification Transformer (PicT) for pavement distress classification. In order to better exploit the discriminative information of pavement images at the patch level, the Patch Labeling Teacher is proposed to leverage a teacher model to dynamically generate pseudo labels of patches from image labels during each iteration, and guides the model to learn the discriminative features of patches via patch label inference in a weakly supervised manner. The broad classification head of Swin Transformer may dilute the discriminative features of distressed patches in the feature aggregation step due to the small distressed area ratio of the pavement image. To overcome this drawback, we present a Patch Refiner to cluster patches into different groups and only select the highest distress-risk group to yield a slim head for the final image classification. We evaluate our method on a large-scale bituminous pavement distress dataset named CQU-BPDD. Extensive results demonstrate the superiority of our method over baselines and also show that PicT outperforms the second-best performed model by a large margin of +2.4% in [email protected] on detection task, +3.9% in F1 on recognition task, and 1.8x throughput, while enjoying 7x faster training speed using the same computing resources. Our codes and models have been released on https://github.com/DearCaat/PicT. Sheng Huang 0001, Xiaoxian Zhang, Luwen Huangfu |
ACM Multimedia | 2 |
| 2022 | PixelSeg: Pixel-by-Pixel Stochastic Semantic Segmentation for Ambiguous Medical ImagesabstractSemantic segmentation tasks often have multiple output hypotheses for a single input image. Particularly in medical images, these ambiguities arise from unclear object boundaries or differences in physicians' annotation. Learning the distribution of annotations and automatically giving multiple plausible predictions is useful to assist physicians in their decision-making. In this paper, we propose a semantic segmentation framework, PixelSeg, for modelling aleatoric uncertainty in segmentation maps and generating multiple plausible hypotheses. Unlike existing works, PixelSeg accomplishes the semantic segmentation task by sampling the segmentation maps pixel by pixel, which is achieved by the PixelCNN layers used to capture the conditional distribution between pixels. We propose (1) a hierarchical architecture to model high-resolution segmentation maps more flexibly, (2) a fast autoregressive sampling algorithm to improve sampling efficiency by 96.2, and (3) a resampling module to further improve predictions' quality and diversity. In addition, we demonstrate the great advantages of PixelSeg in the novel area of interactive uncertainty segmentation, which is beyond the capabilities of existing models. Extensive experiments and state-of-the-art results on the LIDC-IDRI and BraTS 2017 datasets demonstrate the effectiveness of our proposed model. Xiaohong Zhang 0002, Sheng Huang 0001, Kun Wang 0021 |
ACM Multimedia | 3 |
| 2022 | A Probabilistic Model for Controlling Diversity and Accuracy of Ambiguous Medical Image SegmentationabstractMedical image segmentation tasks often have more than one plausible annotation for a given input image due to its inherent ambiguity. Generating multiple plausible predictions for a single image is of interest for medical critical applications. Many methods estimate the distribution of the annotation space by developing probabilistic models to generate multiple hypotheses. However, these methods aim to improve the diversity of predictions at the expense of the more important accuracy. In this paper, we propose a novel probabilistic segmentation model, called Joint Probabilistic U-net, which successfully achieves flexible control over the two abstract conceptions of diversity and accuracy. Specifically, we (i) model the joint distribution of images and annotations to learn a latent space, which is used to decouple diversity and accuracy, and (ii) transform the Gaussian distribution in the latent space to a complex distribution to improve model's expressiveness. In addition, we explore two strategies for preventing the latent space collapse, which are effective in improving the model's performance on datasets with limited annotation. We demonstrate the effectiveness of the proposed model on two medical image datasets, i.e. LIDC-IDRI and ISBI 2016, and achieved state-of-the-art results on several metrics. Xiaohong Zhang 0002, Sheng Huang 0001, Kun Wang 0021 |
ACM Multimedia | 3 |
| 2022 | Adversarial Bidirectional Feature Generation for Generalized Zero-Shot Learning Under Unreliable Semantics
Guowei Yin, Yi Zhang 0113, Sheng Huang 0001 |
PRCV (2) | 4 |
| 2022 | BCL-FL: A Data Augmentation Approach with Between-Class Learning for Fault LocalizationabstractAutomated fault localization (FL) techniques collect runtime information as input data and then analyze input data to identify the relationship between program statements and failures. They usually take advantages of the statistics of the input data to develop a suspiciousness evaluation methodology (e.g., spectrum-based formulas and deep neural network models) by exploring the underlying correlation rooted in the input data. Thus, the quality of input data is critical for FL. In the actual process of development, developers seek to generate adequate test cases for testing the function or the robustness of a subject program. However, regarding a fault, most test cases are passed test cases and a very few ones are failed test cases since a very small portion of inputs in input domain will lead to a program failure. It means that FL usually faces a problem of imbalanced data, and this problem has been proven to pose an adverse effect on FL effectiveness. To address this problem, we propose BCL-FL: a data augmentation approach based on between-class learning, which produces new synthesized failed test samples by mixing two classes of real test cases (i.e., a passed test case and a failed one) with a random ratio. Specifically, BCL-FL uses the characteristics of real failed test cases to design a data synthesis formula suitable for failed test samples, which can make the synthesized failed test samples closer to real test cases. Since the synthesized data is different from real data, we ingeniously assign a continuous value between 0 and 1 to label the synthesized sample according to the mixing ratio of original labels. We take the synthesized failed test samples and the original test cases as the balanced input data for FL techniques to address the imbalanced data problem. To evaluate the effectiveness of BCL-FL, we conduct large-scale experiments on 287 faulty versions of eight large-sized programs (from ManyBugs and Defects4J) using six state-of-the-art FL approaches. The experimental results show that BCL-FL significantly improves the effectiveness of existing FL techniques, e.g., BCL-FL improves the CNN-FL approach in Top-1, Top-5, and Top-10 by 150%, 136.36%, and 193.1%, respectively. Yan Lei 0005, Huan Xie 0002, Sheng Huang 0001, Meng Yan 0001, Zhou Xu 0003 |
SANER | 4 |
| 2022 | Multi-label out-of-distribution detection via exploiting sparsity and co-occurrence of labels
Lei Wang 0062, Sheng Huang 0001, Luwen Huangfu, Bo Liu 0005, Xiaohong Zhang 0002 |
Image Vis. Comput. | 2 |
| 2022 | EANet: Iterative edge attention network for medical image segmentation
Kun Wang 0021, Xiaohong Zhang 0002, Sheng Huang 0001, Dan Yang 0001 |
Pattern Recognit. | 5 |
| 2022 | Multi-Label Image Classification via Category Prototype Compositional LearningabstractReal-world images are often compositions of multiple objects with different categories, scales, poses and locations. Adding nonexistent objects to an image (composing) or removing existent objects from an image (decomposing) leads to higher discrepancy in appearance, which reveals an important but long-neglected compositional nature of multi-label images. In light of this observation, we propose a novel end-to-end compositional learning framework named Category Prototype Compositional Learning (CPCL) to model such compositional nature for multi-label image classification. In CPCL, each image is represented by a collection of category-related features used to eliminate the negative effects from location information. Then, a compositional learning module is introduced to compose and decompose the category-related features with their corresponding category prototypes, which are derived from the semantic representations of categories. If the image has the given object, the output after composing should be closer to the original input than the output after decomposing. Contrarily, if the image does not have the given object, the output after decomposing should be closer to the original input than the output after composing. We introduce the Transformed Appearance Distance (TAD) to measure the appearance change between the composed and decomposed features relative to the category-related features with respect to each category. Finally, multi-label image classification is accomplished by performing a TAD-based metric learning. Experimental results on three multi-label image classification benchmarks,i.e., NUS-WIDE, MS-COCO and VOC 2007, validate the effectiveness and superiority of our work in comparison with the state-of-the-arts. The source codes of our model have been released onhttps://github.com/ZFT-CQU/CPCL. Fengtao Zhou, Sheng Huang 0001, Bo Liu 0005, Dan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | An Iteratively Optimized Patch Label Inference Network for Automatic Pavement Distress DetectionabstractWe present a novel deep learning framework named the Iteratively Optimized Patch Label Inference Network (IOPLIN) for automatically detecting various pavement distresses that are not solely limited to specific ones, such as cracks and potholes. IOPLIN can be iteratively trained with only the image label via the Expectation-Maximization Inspired Patch Label Distillation (EMIPLD) strategy, and accomplish this task well by inferring the labels of patches from the pavement images. IOPLIN enjoys many desirable properties over the state-of-the-art single branch CNN models such as GoogLeNet and EfficientNet. It is able to handle images in different resolutions, and sufficiently utilize image information particularly for the high-resolution ones, since IOPLIN extracts the visual features from unrevised image patches instead of the resized entire image. Moreover, it can roughly localize the pavement distress without using any prior localization information in the training phase. In order to better evaluate the effectiveness of our method in practice, we construct a large-scale Bituminous Pavement Disease Detection dataset named CQU-BPDD consisting of 60,059 high-resolution pavement images, which are acquired from different areas at different times. Extensive results on this dataset demonstrate the superiority of IOPLIN over the state-of-the-art image classification approaches in automatic pavement distress detection. The source codes of IOPLIN are released onhttps://github.com/DearCaat/ioplin, and the CQU-BPDD dataset is able to be accessed onhttps://dearcaat.github.io/CQU-BPDD/. Sheng Huang 0001, Qiming Zhao, Luwen Huangfu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Low-resolution assisted three-stream network for person re-identification
Jiahong Xie, Yongxin Ge, Junyin Zhang, Sheng Huang 0001, Feiyu Chen 0002, Hongxing Wang 0001 |
Vis. Comput. | 4 |
| 2021 | Deep Semantic Dictionary Learning for Multi-label Image ClassificationabstractCompared with single-label image classification, multi-label image classification is more practical and challenging. Some recent studies attempted to leverage the semantic information of categories for improving multi-label image classification performance. However, these semantic-based methods only take semantic information as type of complements for visual representation without further exploitation. In this paper, we present an innovative path towards the solution of the multi-label image classification which considers it as a dictionary learning task. A novel end-to-end model named Deep Semantic Dictionary Learning (DSDL) is designed. In DSDL, an auto-encoder is applied to generate the semantic dictionary from class-level semantics and then such dictionary is utilized for representing the visual features extracted by Convolutional Neural Network (CNN) with label embeddings. The DSDL provides a simple but elegant way to exploit and reconcile the label, semantic and visual spaces simultaneously via conducting the dictionary learning among them. Moreover, inspired by iterative optimization of traditional dictionary learning, we further devise a novel training strategy named Alternately Parameters Update Strategy (APUS) for optimizing DSDL, which alternately optimizes the representation coefficients and the semantic dictionary in forward and backward propagation. Extensive experimental results on three popular benchmarks demonstrate that our method achieves promising performances in comparison with the state-of-the-arts. Our codes and models have been released. Fengtao Zhou, Sheng Huang 0001 |
AAAI | 2 |
| 2021 | DFDM: A Deep Feature Decoupling Module for Lung Nodule SegmentationabstractIn this paper, we propose a novel feature decoupling method to tackle two critical problems in the lung nodule segmentation task: (i) ambiguity of nodule boundary leads to the imprecise segmentation boundary and (ii) the high false positive rate of segmentation result. Our motivation is that an accurate segmentation network needs explicitly modeling the nodule boundary and texture information, and suppressing the noise information. To do so, a novel Deep Feature Decoupling Module (DFDM) is proposed to decouple the nodule boundary, noise, and texture information from the original feature maps. The decoupled boundary and texture information is used to benefit the segmentation, and the noise information is removed from the input features to reduce the false positive rate. The proposed DFDM consists of three parallel branches, including Boundary Sensitive Branch (BSB), Noise Removal Branch (NRB), and Texture Preserving Branch (TPB) to decouple the mentioned three information, respectively. In particular, we design our BSB with a novel architecture to effectively capture the boundary information of lung nodules. We apply the proposed DFDM to the U-Net architecture and achieve convincing segmentation results on the LIDC–IDRI dataset. Code and models are available at https://github.com/chinichenw/DFDM. Wei Chen 0090, Qiuli Wang 0001, Sheng Huang 0001, Xiaohong Zhang 0002, Yucong Li, Chen Liu 0026 |
ICASSP | 3 |
| 2021 | Weakly Supervised Patch Label Inference Network with Image Pyramid for Pavement Diseases Recognition in the WildabstractAutomatic pavement disease recognition is vital for pavement maintenance and management. In this paper, we present an end-to-end deep learning approach named Weakly Super-vised Patch Label Inference Network with Image Pyramid (WSPLIN-IP) for recognizing various types of pavement diseases that are not just limited to the specific ones, such as crack and pothole. WSPLIN-IP first divides the pavement image into patches with an image pyramid for fully exploiting the resolution and scale information. Then, a Patch Label Inference Network (PLIN) is employed for inferring the labels of these patches constrained with a patch label sparsity loss. Finally, the patch labels are fed into a Comprehensive Decision Network (CDN) for disease recognition. Since only the image label is available during whole training, the training of PLIN is conducted in a weakly supervised way under the guidance of CDN and the trained PLIN can provide the interpretable intermediate information. We evaluate our method on a large-scale Bituminous Pavement Disease Dataset named CQU-BPDD whose samples are acquired in the real world. Extensive results demonstrate the superiority of our method over baselines. Guixin Huang, Sheng Huang 0001, Luwen Huangfu, Dan Yang 0001 |
ICASSP | 2 |
| 2021 | Generally Boosting Few-Shot Learning with HandCrafted FeaturesabstractExisting Few-Shot Learning (FSL) methods predominantly focus on developing different types of sophisticated models to extract the transferable prior knowledge for recognizing novel classes, while they almost pay less attention to the feature learning part in FSL which often simply leverage some well-known CNN as the feature learner. However, feature is the core medium for encoding such transferable knowledge. Feature learning is easy to be trapped in the over-fitting particularly in the scarcity of the training data, and thereby degenerates the performances of FSL. The handcrafted features, such as Histogram of Oriented Gradient (HOG) and Local Binary Pattern (LBP), have no requirement on the amount of training data, and used to perform quite well in many small-scale data scenarios, since their extractions involve no learning process, and are mainly based on the empirically observed and summarized prior feature engineering knowledge. In this paper, we intend to develop a general and simple approach for generally boosting FSL via exploiting such prior knowledge in the feature learning phase. To this end, we introduce two novel handcrafted feature regression modules, namely HOG and LBP regression, to the feature learning parts of deep learning-based FSL models. These two modules are separately plugged into the different convolutional layers of backbone based on the characteristics of the corresponding handcrafted features to guide the backbone optimization from different feature granularity, and also ensure that the learned feature can encode the handcrafted feature knowledge which improves the generalization ability of feature and alleviate the over-fitting of the models. Three recent state-of-the-art FSL approaches are leveraged for examining the effectiveness of our method. Extensive experiments on miniImageNet, CIFAR-FS and FC100 datasets show that the performances of all these FSL approaches are well boosted via applying our method on all three datasets. Our codes and models have been released. Yi Zhang 0113, Sheng Huang 0001, Fengtao Zhou |
ACM Multimedia | 2 |
| 2021 | Pulmonary Nodule Classification of CT Images with Attribute Self-guided Graph Convolutional V-Shape Networks
Kun Wang 0021, Xiaohong Zhang 0002, Sheng Huang 0001 |
PRICAI (1) | 4 |
| 2021 | Plot2API: Recommending Graphic API from Plot via Semantic Parsing Guided Neural NetworkabstractPlot-based Graphic API recommendation (Plot2API) is an unstudied but meaningful issue, which has several important applications in the context of software engineering and data visualization, such as the plotting guidance of the beginner, graphic API correlation analysis, and code conversion for plotting. Plot2API is a very challenging task, since each plot is often associated with multiple APIs and the appearances of the graphics drawn by the same API can be extremely varied due to the different settings of the parameters. Additionally, the samples of different APIs also suffer from extremely imbalanced.Considering the lack of technologies in Plot2API, we present a novel deep multi-task learning approach named Semantic Parsing Guided Neural Network (SPGNN) which translates the Plot2API issue as a multi-label image classification and an image semantic parsing tasks for the solution. In SPGNN, the recently advanced Convolutional Neural Network (CNN) named EfficientNet is employed as the backbone network for API recommendation. Meanwhile, a semantic parsing module is complemented to exploit the semantic relevant visual information in feature learning and eliminate the appearance-relevant visual information which may confuse the visual-information-based API recommendation. Moreover, the recent data augmentation technique named random erasing is also applied for alleviating the imbalance of API categories.We collect plots with the graphic APIs used to drawn them from Stack Overflow, and release three new Plot2API datasets corresponding to the graphic APIs of R and Python programming languages for evaluating the effectiveness of Plot2API techniques. Extensive experimental results not only demonstrate the superiority of our method over the recent deep learning baselines but also show the practicability of our method in the recommendation of graphic APIs. Zeyu Wang 0001, Sheng Huang 0001, Zhongxin Liu 0002, Meng Yan 0001, Xin Xia 0001, Bei Wang 0010, Dan Yang 0001 |
SANER | 2 |
| 2021 | Deep feature enhancing and selecting network for weakly supervised temporal action localization
Jiaruo Yu, Yongxin Ge, Xiaolei Qin, Sheng Huang 0001, Feiyu Chen 0002 |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | Discriminative deep semi-nonnegative matrix factorization network with similarity maximization for unsupervised feature learning
Feiyu Chen 0002, Yongxin Ge, Sheng Huang 0001, Xiaohong Zhang 0002, Dan Yang 0001 |
Pattern Recognit. Lett. | 4 |
| 2021 | Realistic Lung Nodule Synthesis With Multi-Target Co-Guided Adversarial MechanismabstractThe important cues for a realistic lung nodule synthesis include the diversity in shape and background, controllability of semantic feature levels, and overall CT image quality. To incorporate these cues as the multiple learning targets, we introduce the Multi-Target Co-Guided Adversarial Mechanism, which utilizes the foreground and background mask to guide nodule shape and lung tissues, takes advantage of the CT lung and mediastinal window as the guidance of spiculation and texture control, respectively. Further, we propose a Multi-Target Co-Guided Synthesizing Network with a joint loss function to realize the co-guidance of image generation and semantic feature learning. The proposed network contains a Mask-Guided Generative Adversarial Sub-Network (MGGAN) and a Window-Guided Semantic Learning Sub-Network (WGSLN). The MGGAN generates the initial synthesis using the mask combined with the foreground and background masks, guiding the generation of nodule shape and background tissues. Meanwhile, the WGSLN controls the semantic features and refines the synthesis quality by transforming the initial synthesis into the CT lung and mediastinal window, and performing the spiculation and texture learning simultaneously. We validated our method using the quantitative analysis of authenticity under the Fréchet Inception Score, and the results show its state-of-the-art performance. We also evaluated our method as a data augmentation method to predict malignancy level on the LIDC-IDRI database, and the results show that the accuracy of VGG-16 is improved by 5.6%. The experimental results confirm the effectiveness of the proposed method. Qiuli Wang 0001, Xiaohong Zhang 0002, Mingchen Gao, Sheng Huang 0001, Jian Wang 0135, Jiuquan Zhang, Dan Yang 0001, Chen Liu 0026 |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Erratum to "Realistic Lung Nodule Synthesis With Multi-Target Co-Guided Adversarial Mechanism"
Qiuli Wang 0001, Xiaohong Zhang 0002, Mingchen Gao, Sheng Huang 0001, Jian Wang 0135, Jiuquan Zhang, Dan Yang 0001, Chen Liu 0026 |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Knowledge-Guided And Hyper-Attention Aware Joint Network For Benign-Malignant Lung Nodule ClassificationabstractAccurate identification and early diagnosis of malignant lung nodules are crucial for improving the survival rate of patients with lung cancer. Deep learning methods have recently been proven success in computer-aided diagnostic tasks. However, to the best of our knowledge, the features of tissues and vessels will disturb the model resulting in inaccurate classification of the nodules. To reduce the interference and capture crucial contextual information from different channels in a more efficient way, we introduce a Hyper-Attention Mechanism(HAM) that can be easily integrated into convolutional neural networks(CNNs). Moreover, without incorporating prior-domain knowledge, traditional methods lack interpretability, which is difficult to understand and utilize them in the clinic by radiologists. Based on this, we propose a novel Knowledge-Guided model to predict malignant pulmonary nodules from chest CT data, which inject external medical knowledge into CNNs to guide the training process. We evaluate the proposed model on the LIDC-IDRI dataset and demonstrate its effectiveness by achieving comparable state-of-the-art performance. Weixin Xu 0002, Kun Wang 0021, Jingkai Lin, Sheng Huang 0001, Xiaohong Zhang 0002 |
ICIP | 5 |
| 2020 | Robust Bidirectional Generative Network For Generalized Zero-Shot LearningabstractIn this work, we propose a novel generative approach named Robust Bidirectional Generative Network (RBGN) based on Conditional Generative Adversarial Network (CGAN) for Generalized Zero-shot Learning (GZSL). RBGN employs the adversarial attack to train a more rigorous discriminator, thus enhancing the generalizability and robustness of the feature generator under minimax strategy. Moreover, RBGN decodes the generated visual features back to their semantic representations to further improve the representational ability of generated visual features and alleviate the hubness problem. The experimental results of GZSL on four datasets, i.e. CUB, SUN, AWA1, AWA2, demonstrate that our model achieves competitive performance compared to state-of-the-art approaches and owns better generalizability to the unseen classes over conventional generative GZSL models. Further robustness analysis also validates the strong robustness of our model to the different types of semantic disturbance. Sheng Huang 0001, Luwen Huangfu, Feiyu Chen 0002, Yongxin Ge |
ICME | 2 |
| 2020 | Corner detection using the point-to-centroid distance techniqueabstractCorners, highly important local features of images and corner finding, play a crucial role in computer vision and image processing, such as object tracking and vehicle detection. Proposing effective and efficient corner detectors is the aim of corner detection. In this study, the authors first present a new measure of corner sharpness termed as the point‐to‐centroid distance (PCD) and then examine its behaviours, which display beneficial characteristics that help distinguish corners from non‐corners. Based on PCD behaviours, the authors propose a novel corner detector. Extensive experimental results demonstrate that the PCD technique is effective and simultaneously efficient for corner detection compared with six other contour‐based corner detectors in terms of two commonly used evaluation metrics – average repeatability and localisation error. Shizheng Zhang, Luwen Huangfu, Zhifeng Zhang 0002, Sheng Huang 0001, Heng Wang 0004 |
IET Image Process. | 4 |
| 2020 | DTMMN: Deep transfer multi-metric network for RGB-D action recognition
Xiaolei Qin, Yongxin Ge, Jinyuan Feng, Dan Yang 0001, Feiyu Chen 0002, Sheng Huang 0001 |
Neurocomputing | 6 |
| 2020 | Class-Prototype Discriminative Network for Generalized Zero-Shot LearningabstractWe present a novel end-to-end deep metric learning model named Class-Prototype Discriminative Network (CPDN) for Generalized Zero-Shot Learning (GZSL). It consists of a generative network for producing the visual prototype of each class by feeding its semantic representation, and a metric network for measuring the similarities between the sample and the generated class-prototypes to accomplish the classification. In CPDN, a query sample intends to posses a higher similarity with its homogenous class-prototypes while the lower similarities with the inhomogenous ones, and the class-prototypes also intend to be distinguished with each other through the metric network. Moreover, a discriminative version of Relation Network (RN) named Discriminative Relation Network (DRN) is presented by incorporating the aforementioned idea into the conventional RN model for further achieving the complementation CDPN and RN in metric learning. Extensive experimental results on standard benchmarks demonstrate that our proposed approaches consistently outperform RN, and achieve the competitive performances compared with the state-of-the-arts in GZSL. Sheng Huang 0001, Jingkai Lin, Luwen Huangfu |
IEEE Signal Process. Lett. | 1 |
| 2019 | KGZNet: Knowledge-Guided Deep Zoom Neural Networks for Thoracic Disease ClassificationabstractThis paper aims to automatically diagnose thoracic diseases in Chest X-ray(CXR) images using deep neural net-works(DNN). However, the existing approaches generally use the global CXR images as input for training purposes. This strategy is low-efficiency, coarse, and might introduce many unnecessary noises. We believe that the deep learning, which is inherently an algebraic computation system, is not the most efficient way to acquire highly sophisticated human knowledge, for example those thoracic diseases are typically limited within the lung regions and interdependence between lesion location. In this paper, we address the above problem by proposing to explore how external medical knowledge can be injected into DNN to guide its training process. We design four feature extraction modules to construct a knowledge-guided deep zoom neural network(KGZNet), which can gradually make full use of the most medical discriminative feature information(from coarse to fine) of global, lung regions, and lesion regions. Specifically, we first learn global branch using global images. Second, learn the lung region branch using lung region images, which are identified and cropped by the Lung Region Generator(LRG-1). Then, guided by the attention heat map generated from the lung region branch learning, we inference a mask to crop a medical discriminative lesion region from the lung region images by the Lesion Region Generator(LRG-2). The lesion region images are used for training a lesion branch. Lastly, the obtained medical discriminative features knowledge are fused by the feature fusion model for disease classification. We have evaluated the proposed method on the NIH ChestX-ray 14 dataset and achieves the average AUC of 0.878, and the experiment results demonstrate the superiority and effectiveness of the proposed method, compared to other state-of-the-art methods. Kun Wang 0021, Xiaohong Zhang 0002, Sheng Huang 0001 |
BIBM | 3 |
| 2019 | Automatic Detection of Pneumonia in Chest X-Ray Images Using Cooperative Convolutional Neural Networks
Kun Wang 0021, Xiaohong Zhang 0002, Sheng Huang 0001, Feiyu Chen 0002 |
PRCV (2) | 3 |
| 2019 | Fine Grain Lung Nodule Diagnosis Based on CT Using 3D Convolutional Neural Network
Qiuli Wang 0001, Sheng Huang 0001, Chen Liu 0026, Xiaohong Zhang 0002, Dan Yang 0001 |
PRCV (2) | 3 |
| 2019 | Corner detection based on tangent-to-point distance accumulation technique
Shizheng Zhang, Sheng Huang 0001, Zhifeng Zhang 0002, Heng Wang 0004, Junxia Ma |
Multim. Tools Appl. | 2 |
| 2019 | Discriminative Probabilistic Latent Semantic Analysis with Application to Single Sample Face Recognition
Daoxiang Zhou, Dan Yang 0001, Xiaohong Zhang 0002, Sheng Huang 0001, Shu Feng |
Neural Process. Lett. | 4 |
| 2018 | Residual Inception: A New Module Combining Modified Residual with Inception to Improve Network PerformanceabstractResiduals and inception are two commonly used module that makes the network deeper and wider to achieve better performance. And the combination of these two modules which is usually referred to as inception-resnet can get a better result. In this paper, we propose a new type of combination to give full play to the role of residuals and inception, making network learning more abundant features. The new proposed module is called Residual Inception (RI) which enjoys the same width as the inception module in GoogLeNet. In RI, each parallel cascade structure is replaced by a densely block or a modified residual block for gaining a better performance and a lower computational cost. Finally, we evaluate our proposed network on three highly competitive datasets and the results demonstrate its superiority in comparison with the state-of-the-art. Xingpeng Zhang, Sheng Huang 0001, Xiaohong Zhang 0002, Qiuli Wang 0001, Dan Yang 0001 |
ICIP | 2 |
| 2018 | Joint Deep Learning for RGB-D Action RecognitionabstractRecent approaches in RGB-based and depth-based human action recognition achieved outstanding performance respectively, which demonstrate the effectiveness of RGB and depth modalities for action classification, however it is infrequent to consider them both. Currently, available multimodal-based methods of action recognition suffer from some limitations, including non-end-to-end training, violent fusion and inefficiency. In this paper, we propose a novel joint deep learning (JDL) model which is capable of: 1) jointly optimizing the object of classification and feature extraction through a novel end-to-end two-stream deep learning model, 2) refining common-specific features via introducing the constraint of similarity loss in high-level, and 3) using 2D convolution kernel instead of 3D convolution kernel during feature extraction for gaining the efficiency. The experiments on two challenging datasets show the promising performance of our architecture. Xiaolei Qin, Yongxin Ge, Liuwei Zhan, Guangrui Li 0003, Sheng Huang 0001, Hongxing Wang 0001, Feiyu Chen 0002 |
VCIP | 5 |
| 2018 | Improved hypergraph regularized Nonnegative Matrix Factorization with sparse representationabstractAs a commonly used data representation technique, Nonnegative Matrix Factorization (NMF) has received extensive attentions in the pattern recognition and machine learning communities over decades, since its working mechanism is in accordance with the way how the human brain recognizes objects. Inspired by the remarkable successes of manifold learning, more and more researchers attempt to incorporate the manifold learning into NMF for finding a compact representation ,which uncovers the hidden semantics and respects the intrinsic geometric structure simultaneously. Graph regularized Nonnegative Matrix Factorization (GNMF) is one of the representative approaches in this category. The core of such approach is the graph, since a good graph can accurately reveal the relations of samples which benefits the data geometric structure depiction. In this paper, we leverage the sparse representation to construct a sparse hypergraph for better capturing the manifold structure of data, and then impose the sparse hypergraph as a regularization to the NMF framework to present a novel GNMF algorithm called Sparse Hypergraph regularized Nonnegative Matrix Factorization (SHNMF). Since the sparse hypergraph inherits the merits of both the sparse representation and the hypergraph model, SHNMF enjoys more robustness and can better exploit the high-order discriminant manifold information for data representation . We apply our work to address the image clustering issue for evaluation. The experimental results on five popular image databases show the promising performances of the proposed approach in comparison with the state-of-the-art NMF algorithms. Sheng Huang 0001, Hongxing Wang 0001, Yongxin Ge, Luwen Huangfu, Xiaohong Zhang 0002, Dan Yang 0001 |
Pattern Recognit. Lett. | 1 |
| 2018 | Background Modeling by Stability of Adaptive Features in Complex ScenesabstractThe single-feature-based background model often fails in complex scenes, since a pixel is better described by several features, which highlight different characteristics of it. Therefore, the multi-feature-based background model has drawn much attention recently. In this paper, we propose a novel multi-feature-based background model, named stability of adaptive feature (SoAF) model, which utilizes the stabilities of different features in a pixel to adaptively weigh the contributions of these features for foreground detection. We do this mainly due to the fact that the features of pixels in the background are often more stable. In SoAF, a pixel is described by several features and each of these features is depicted by a unimodal model that offers an initial label of the target pixel. Then, we measure the stability of each feature by its histogram statistics over a time sequence and use them as weights to assemble the aforementioned unimodal models to yield the final label. The experiments on some standard benchmarks, which contain the complex scenes, demonstrate that the proposed approach achieves promising performance in comparison with some state-of-the-art approaches. Dan Yang 0001, Chenqiu Zhao, Xiaohong Zhang 0002, Sheng Huang 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Robust face alignment with cascaded coarse-to-fine auto-encoder networkabstractIn this paper, we present a novel face alignment method using a two-level cascaded auto-encoder networks (2-LCAN). In our framework, the first level auto-encoder networks generate rough facial landmarks locations by taking detected face images with low-resolution as inputs. The second level autoencoder networks are constructed by cascading several sub stacked auto-encoder networks (SSAN) in a coarse-to-fine manner. Each SSAN extracts SIFT features and local pixels features around current landmark positions, then fuses them together to further refine landmarks of different facial components with higher image resolutions. Finally, experimental results on LFPW and HELEN datasets demonstrate that our proposed method is significantly superior to the compared approaches both in accuracy and robustness. Yongxin Ge, Mingjian Hong, Sheng Huang 0001, Dan Yang 0001 |
ICIP | 4 |
| 2017 | On the effect of hyperedge weights on hypergraph learning
Sheng Huang 0001, Ahmed M. Elgammal, Dan Yang 0001 |
Image Vis. Comput. | 1 |
| 2017 | Robust corner detection using the eigenvector-based angle estimator
Shizheng Zhang, Dan Yang 0001, Sheng Huang 0001, Xiaohong Zhang 0002, Liyun Tu, Zemin Ren |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Joint Local Regressors Learning for Face Alignment
Yongxin Ge, Mingjian Hong, Sheng Huang 0001, Dan Yang 0001 |
Neurocomputing | 4 |
| 2016 | Discriminant Hyper-Laplacian Projections and its scalable extension for dimensionality reduction
Sheng Huang 0001, Dan Yang 0001, Yongxin Ge, Xiaohong Zhang 0002 |
Neurocomputing | 1 |
| 2016 | Collaborative Graph Embedding: A Simple Way to Generally Enhance Subspace Learning AlgorithmsabstractCollaborative representation (CR), known as an effective way to address the signal representation (regression) problem, has achieved remarkable success in visual classification. According to our theoretical analysis, the subspace learning issue can also be deemed as a signal representation problem. Therefore, we extend the graph embedding (GE) framework as a CR model to improve the discriminating power of the subspace learning algorithm. The new GE framework, which is named collaborative GE (CGE) framework, enjoys many desirable properties of CR. From theoretical analysis, CGE is robust to the noise and has the same computational complexity as GE. From experimental analysis, CGE can generally enhance the subspace learning algorithms and a reasonable regularization parameter can be inferred from its intrinsic graph. Several state-of-the-art subspace learning algorithms are plugged into our framework to produce their collaborative versions. Meanwhile, by exploring the intrinsic relation among GE methods, we present a new collaborative method named collaborative class-scattering locality preserving projections (CCSLPPs). The results of extensive experiments on ORL, AR, Scene15, Caltech256, LFW-A, and OU-ISIR-A databases demonstrate that the collaborative versions consistently outperform their original algorithms with a remarkable improvement and CCSLPP gets the best performance compared with all used methods. Sheng Huang 0001, Yu Yang 0010, Dan Yang 0001, Ahmed M. Elgammal |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Learning Hypergraph-regularized Attribute PredictorsabstractWe present a novel attribute learning framework named Hypergraph-based Attribute Predictor (HAP). In HAP, a hypergraph is leveraged to depict the attribute relations in the data. Then the attribute prediction problem is casted as a regularized hypergraph cut problem, in which a collection of attribute projections is jointly learnt from the feature space to a hypergraph embedding space aligned with the attributes. The learned projections directly act as attribute classifiers (linear and kernelized). This formulation leads to a very efficient approach. By considering our model as a multi-graph cut task, our framework can flexibly incorporate other available information, in particular class label. We apply our approach to attribute prediction, Zero-shot and N-shot learning tasks. The results on AWA, USAA and CUB databases demonstrate the value of our methods in comparison with the state-of-the-art approaches. Sheng Huang 0001, Ahmed M. Elgammal, Dan Yang 0001 |
CVPR | 1 |
| 2015 | Weather classification with deep convolutional neural networksabstractIn this paper, we study weather classification from images using Convolutional Neural Networks (CNNs). Our approach outperforms the state of the art by a huge margin in the weather classification task. Our approach achieves 82.2% normalized classification accuracy instead of 53.1% for the state of the art (i.e., 54.8% relative improvement). We also studied the behavior of all the layers of the Convolutional Neural Networks, we adopted, and interesting findings are discussed. Sheng Huang 0001, Ahmed M. Elgammal |
ICIP | 2 |
| 2015 | A load balancing multi-path routing scheme based on effective voids for optical burst switching networks
Sheng Huang 0001, Yunshui Zhang, Liqin Sun, Keping Long |
Sci. China Inf. Sci. | 1 |
| 2015 | Combined supervised information with PCA via discriminative component selection
Sheng Huang 0001, Dan Yang 0001, Yongxin Ge, Xiaohong Zhang 0002 |
Inf. Process. Lett. | 1 |
| 2015 | Graph regularized linear discriminant analysis and its generalization
Sheng Huang 0001, Dan Yang 0001, Xiaohong Zhang 0002 |
Pattern Anal. Appl. | 1 |
| 2015 | Class specific sparse representation for classification
Sheng Huang 0001, Yu Yang 0010, Dan Yang 0001, Luwen Huangfu, Xiaohong Zhang 0002 |
Signal Process. | 1 |
| 2015 | Cross-Speed Gait Recognition Using Speed-Invariant Gait Templates and Globality-Locality Preserving ProjectionsabstractWe present a novel manifold-based approach for cross-speed gait recognition. In our approach, the walking action is considered as residing on a manifold, in the feature space, that is homomorphic to a unit circle. We employ thin plate spline (TPS) kernel-based radial basis function (RBF) interpolation to fit such manifold. TPS kernel-based RBF interpolation separates the learned coefficients into an affine component and a nonaffine component, which, respectively, encodes the dynamic and static characteristics of the gait manifold. We introduce the use of the nonaffine component as a cross-speed gait representation, and denote it speed invariant gait template (SIGT). We also propose an enhanced locality preserving projections (LPP) algorithm named globality LPP (GLPP) for reducing the dimension of SIGT. In GLPP, the graph Laplacians of intrasubject part and intersubjects part are separately constructed, and then to combine as a new graph Laplacian. Finally, a manifold learning-based classifier named normalized hypergraph classifier is employed for classification. Experimental results on two gait databases demonstrate the effectiveness of our proposed approach in comparison with the state-of-the-art gait recognition methods. Sheng Huang 0001, Ahmed M. Elgammal, Jiwen Lu, Dan Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2014 | Improving non-negative matrix factorization via ranking its basesabstractAs a considerable technique in image processing and computer vision, Nonnegative Matrix Factorization (NMF) generates its bases by iteratively multiplicative update with two initial random nonnegative matrices W and H, that leads to the randomness of the bases selection. For this reason, the potentials of NMF algorithms are not completely exploited. To address this issue, we present a novel framework which uses the feature selection techniques to evaluate and rank the bases of the NMF algorithms to enhance the NMF algorithms. We adopted the well known Fisher criterion and Least Reconstruction Error criterion, which is proposed by us, as two instances to show how that works successfully under our framework. Moreover, in order to avoid the hard combinatorial optimization issue in ranking procedure, a de-correlation constraint can be optionally imposed to the NMF algorithms for giving a better approximation to the global optimum of the NMF projections. We evaluate our works in face recognition, object recognition and image reconstruction on ORL and ETH-80 databases and the results demonstrate the enhancement of the state-of-the-art NMF under our framework. Sheng Huang 0001, Ahmed M. Elgammal, Dan Yang 0001 |
ICIP | 1 |
| 2013 | Learning Speed Invariant Gait Template via Thin Plate Spline Kernel Manifold FittingabstractWe present a novel approach for cross-speed gait recognition. In our approach, the cyclic walking action is considered as residing on a manifold which is homeomorphic to a unit circle in the gait space. Thin Plate Spline (TPS) kernel-based Radial Basis Function (RBF) interpolation is used to fit the walking manifold for each gait sequence. The sub-ject related kernel mapping coefficients are learned for representing the gait. According to the property of TPS, the coefficients can be naturally separated as an affine component and a non-affine component. The affine component is the style factor corresponding to the deformation of the homeomorphic manifold caused by the walking action, while the non-affine component is the shape factor, invariant to the walking speed. We denote this non-affine component as Speed Invariant Gait Template (SIGT) and use it as cross-speed gait feature. To address the curse of dimensionality issue and speed up the recognition, we use Globality Locality Preserving Projections (GLPP) to reduce the dimensions of SIGTs. Two walking speeds related gait databases are employed for evaluating our pro-posed method. The experimental results demonstrate the superiority of our method over the state-of-the-art. 1 Sheng Huang 0001, Ahmed M. Elgammal, Dan Yang 0001 |
BMVC | 1 |
| 2009 | VB-Rescheduling: An Efficient Data Channel Rescheduling Algorithm Based on Virtual Burst for OBS NetworksabstractIn optical burst switching (OBS) networks, the data channel scheduling algorithm is one of the most important issues, which have a great impact on network performances. Currently, there are various data channel scheduling algorithms. Among them, the rescheduling algorithm is more attractive because it could adaptively reallocate the data channels even when they have been occupied by some data bursts (DB), and release some channel resource for the latter DB in most situations. However when the traffic load is heavy, it is not effective any more, and would worsen network performance. Therefore, this paper proposes a new rescheduling algorithm, namely VB-Rescheduling algorithm. According to the state of the data channels, it reschedules data blocks on demand by three granularities (i.e., virtual burst, child-burst cluster and normal burst). Compared with other rescheduling algorithms, it has some advantages as follows. Firstly, it could keep the same sequence of the arriving data bursts at a node as the corresponding control packets. Secondly, it is more flexible to reschedule data blocks. Finally, simulation results show that it can greatly improve OBS network performance in terms of the overall packet loss probability and the link utilization, compared with traditional OBS rescheduling algorithm (whose rescheduling granularity is normal burst) and the native virtual burst scheduling scheme. Keping Long, Fenfen Dong, Sheng Huang 0001, Xiaolin Duan |
ICC | 5 |
| 2006 | The SLA-Compatible Fault Management Model for Differentiated Fault Recovery
Keping Long, Sheng Huang 0001, Yujun Kuang |
HPCC | 3 |
| 2006 | An Adaptive Parameter Deflection Routing to Resolve Contentions in OBS Networks
Keping Long, Sheng Huang 0001, Qianbin Chen, Ruyan Wang |
Networking | 3 |