VLDB 2026 Research / reviewers in the wild / expert
Chuanfei Hu
dblp:272/6122
· DBLP profile ↗
29ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0003-1669-9429ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A fine-grained information fusion and inference method for out-of-distribution detection in fault diagnosis
Guoliang Wu, Xinde Li, Fir Dunkin, Chuanfei Hu, Heqing Li, Zhentong Zhang, Kaixuan Wu, Erfeng Liu |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Cross-view encoder with pseudo label for incomplete multi-view clustering in postoperative liver diagnosis
Xinde Li, Chuanfei Hu |
Pattern Recognit. | 3 |
| 2026 | TranSpike: Pixel-wise frequency reconstruction and spike interaction for remote photoplethysmography
Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Chuanfei Hu, Shuo Chen 0003, Jian Yang 0003 |
Pattern Recognit. | 4 |
| 2026 | PipeCLIP: Defect-Conditioned and Cross-Focus-Driven Vision-Language Model for Video-Based Sewer Defect InspectionabstractIn recent years, vision-language models (VLMs), such as CLIP, have excelled in the visual domain due to the availability of vast paired image-text data. However, directly adapting CLIP to the sewer defect inspection achieves a poor performance, since the vocabularies of defect category are highly “unfamiliar” to CLIP, resulting in the weak representations among the defect categories in the text embedding space. Besides, the substantial differences of pipe characteristics are challenging to align the multiple pairs of visual and text features across multi-focus segments. We propose PipeCLIP for adapting the text information to the video-based multi-label sewer defect classification, which is the first method to integrate a CLIP-based model into the sewer defect inspection. First, expert prior descriptions (EPD) are introduced to differentiate the distinctions between defect category in the text embedding space. Second, defect-attribute coupling (DAC) prompt is proposed to strengthen the coupling relationship between the sewer defects and pipe multi-attributes. The two prompts proposed above together with the traditional category prompt, collectively constitute defect-conditioned text prompt (DecTP) for CLIP. Then, cross-focus temporal (CFT) module is designed to integrate feature information from different focal length, strengthening the visual-text alignment across multi-focus segments. Extensive experiments are conducted on the public benchmark, in which the superiority of PipeCLIP is demonstrated compared with the state-of-the-art methods. Code is available at: https://anonymous.4open.science/r/PipeCLIP-0925. Chenyang Zhao 0009, Chuanfei Hu, Zhenzhong Cao, Yinuo Song, Jingtai Liu |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | Trustworthy Driver State Perception via Contextual Interaction-Driven Evidential Vision-Language Fusion in Vehicular Cyber-Physical SystemsabstractA vision-driven driver monitoring system plays a vital role of vehicular cyber-physical systems (VCPS) to guarantee the driving safety. Recent advances focus on modeling a deep learning-based method to realize the driver monitoring system, which benefits from the powerful capability of data-driven feature extraction. Although the acceptable performances of driver state monitoring methods are achieved, there is still a gap between the emerged techniques and actual application scenarios. First, the human-centric visual appearances are not involved comprehensively to represent the driver states, resulting in ignoring the contextual interaction of behaviors. Second, the inherent uncertainty of driver situation is not considered, while the unreliable samples would lead to the untrustworthy results. In this paper, we focus on a vision-based driver state monitoring method, where a trustworthy driver state perception (TDSP) is proposed via human-centric contextual interaction-driven evidential vision-language fusion in VCPS. Specifically, a vision-language model-based architecture is first modified in temporal dimension to represent the visual human-centric contextual interactions, while a vision-language consistency loss is designed to mitigate the gap between visual and textual representations. Then, an evidence-based learning method is introduced to jointly conduct the classification and uncertainty estimation for driver states. Furthermore, to model the human-centric contextual interactions towards the evidence-based paradigm comprehensively, Dempster-Shafer theory-based combination rule is introduced to fuse the visual and textual representations. Extensive experiments are conducted on two public benchmarks, where the superiority of TDSP is demonstrated compared with the state-of-the-art methods. The superior performance of TDSP to recognize dangerous states are 85.41% and 83.12% in terms of accuracy and F1 score, which outperforms the state-of-the-art methods by 4.68% and 3.99%. Moreover, we validate the reliability of TDSP against the noisy data for VCPS. The code will be public at https://github.com/w64228013/TDSP. Chuanfei Hu, Xinde Li, Jianxin Pang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | AVQACL: A Novel Benchmark for Audio-Visual Question Answering Continual LearningabstractIn this paper, a novel benchmark for audio-visual question answering continual learning (AVQACL) is introduced, aiming to study fine-grained scene understanding and spatial-temporal reasoning in videos under a continual learning setting. To facilitate this multimodal continual leaning task, we create two audio-visual question answering continual learning datasets, named Split-AVQA and Split-MUSIC-AVQA based on the AVQA and MUSIC-AVQA datasets, respectively. The experimental results suggest that the model exhibits limited cognitive and reasoning abilities and experiences catastrophic forgetting when processing three modalities simultaneously in a continuous data stream. To address above challenges, we propose a novel continual learning method that incorporates question-guided cross-modal information fusion (QCIF) to focus on question-relevant details for improved feature representation and task-specific knowledge distillation with spatial-temporal feature constraints (TKD-STFC) to preserve the spatial-temporal reasoning knowledge acquired from previous dynamic scenarios. Furthermore, a question semantic consistency constraint (QSCC) is employed to ensure that the model maintains a consistent understanding of question semantics across tasks throughout the continual learning process. Extensive experimental results on Split-AVQA and Split-MUSIC-AVQA datasets illustrate that our method achieves state-of-the-art audio-visual question answering continual learning performance. The code is available at https://github.com/kx-wu/AVQACL. Kaixuan Wu, Xinde Li, Chuanfei Hu, Guoliang Wu |
CVPR | 4 |
| 2025 | UCFN: Uncertainty-aware cross-granularity fusion network for visual intention understanding
Xinde Li, Chuanfei Hu |
Neurocomputing | 3 |
| 2025 | Stochastic human motion prediction using a quantized conditional diffusion model
Biaozhang Huang, Xinde Li, Chuanfei Hu, Heqing Li |
Knowl. Based Syst. | 3 |
| 2025 | Trusted Video-Based Sewer Inspection via Support Clip-Based Pareto-Optimal Evidential NetworkabstractAn automatic vision-based sewer inspection plays a vital role of sewage system in a modern city. Existing methods have utilized evidential deep learning to construct trusted models. Although the acceptable performance has been achieved in sewer defect classification, the fine-grained information of sewer defects in videos is ignored. Meanwhile, the trade-off between multi-label classification and uncertainty estimation remains challenging. In this paper, support clip-based pareto-optimal evidential network (POEN) is proposed for trusted video-based sewer inspection. Specifically, support clip module (SCM) is designed to capture the fine-grained visual representation of defects from local scale segments. Then, evidential deep learning is introduced to quantify the uncertainty for out-of-distribution detection. Furthermore, Pareto-optimal weighting scheme (PWS) is designed to solve the common trade-off dilemma in multi-task learning. Extensive experiments are conducted on VideoPipe, in which the superiority of POEN is demonstrated compared with the state-of-the-art methods. Chenyang Zhao 0009, Chuanfei Hu, Hang Shao 0001, Fir Dunkin, Yongxiong Wang |
IEEE Signal Process. Lett. | 2 |
| 2025 | MgCNL: A Sample Separation Approach via Multi-Granularity Balls for Fault Diagnosis With the Interference of Noisy LabelsabstractThe fault diagnosis based on supervised learning has achieved remarkable results in the intelligent manufacturing, making it an important guarantee for long-term safe and stable operation in modern industry. However, the accuracy heavily relies on high-quality annotation labels, which are expensive to obtain, limiting the diagnosis models applicability in many scenarios. Although obtaining automatically annotated samples from annotators is a promising solution, the generated dataset is always containing incorrect labels (noisy labels), due to perceptual limitations, resulting in low or even invalid the accuracy of model. With the goal of handling this challenge, a diagnostic approach based on multi-granularity information fusion to combat noisy labels, called MgCNL, is proposed, to train the model with high-accuracy, without knowing the specific noise ratio. Specifically, inspired by granular-ball computing, a confidence evaluation method of labels is designed, so that samples with high confidence labels can be selected from dataset with noisy labels for supervised learning, thus avoiding the negative impact of incorrect labels on model performance. Finally, the efficacy was demonstrated on three datasets using different backbones: MgCNL successfully reduced the adverse impact of noisy labels, achieving significantly better results than other advanced methods in various noisy scenarios, which offers a competitive model training strategy for practitioners in intelligent manufacturing or industrial fault diagnosis who are hampered by the costs associated with sample labeling. Note to Practitioners—In modern industry, the cost of manual/expert annotation for high-quality data is is prohibitively expensive, and the data annotated by automatic annotators often contains noisy labels that seriously damages the accuracy of models, which makes many data-driven diagnosis models constrained by training data and difficult to put into practice, posing an urgent challenge to the automation and intelligence of the manufacturing industry. To address this challenge, this article proposed a robust training strategy called MgCNL, aimed at offsetting the negative impact of noisy labels, in the hope that automatic annotation strategy with lower cost can be more widely applied in model training tasks for industrial practice. MgCNL, based on multi-granularity information, can effectively select high-confidence samples from datasets for supervised learning, even under unknown proportions of noise labels, thus reducing the misleading impact of noisy labels on diagnostic models. As a result, MgCNL possesses the ability to robustly train high-accuracy diagnostic models in data with noisy labels, thus enabling automatic annotators to replace experts in dataset construction as a more economical and efficient potential technical approach. Meanwhile, MgCNL also brings value to datasets with uncertain labels, making them applicable without the need to invest significant human resources to verify label reliability. Fir Dunkin, Xinde Li, Heqing Li, Guoliang Wu, Chuanfei Hu, Shuzhi Sam Ge |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | ASD: Towards Attribute Spatial Decomposition for Prior-Free Facial Attribute RecognitionabstractRepresenting the spatial properties of facial attributes is a vital challenge for facial attribute recognition (FAR). Recent advances have achieved the reliable performances for FAR, benefiting from the description of spatial properties via extra prior information. However, the extra prior information might not be always available, resulting in the restricted application scenario of the prior-based methods. Meanwhile, the spatial ambiguity of facial attributes caused by inherent spatial diversities of facial parts is ignored. To address these issues, we propose a prior-free method for attribute spatial decomposition (ASD), mitigating the spatial ambiguity of facial attributes. The attribute components could be formally described in terms of the spatial locations without any extra prior information. Experimental results demonstrate the superiority of ASD compared with state-of-the-art prior-based methods on both CelebA and LFWA. Chuanfei Hu, Hang Shao 0001, Bo Dong 0001, Zhe Wang 0027, Yongxiong Wang |
ICME | 1 |
| 2024 | A Novel Framework for Structure Descriptors-Guided Hand-drawn Floor Plan ReconstructionabstractIn the absence of a pre-built indoor map, robot navigation suffers from the limitations of sensors and environments, resulting in decreased efficiency in performing ad-hoc tasks. Given that blueprints are difficult to obtain, an intuitive method is to provide robots with prior knowledge via hand-drawn floor plans. However, due to the inability of robots to directly comprehend hand-drawn styles, the applicability of this method is limited. In this paper, we present a novel framework for hand-drawn floor plan reconstruction that can recognize abstract hand-drawn elements and standardize the reconstruction of hand-drawn floor plans, thereby providing robots with valuable global map information. Specifically, we design a new series of structure descriptors as reconstruction components and employ a deep learning-based model for recognition. Then the standardized results are obtained through the proposed floor plan reconstruction algorithm. To verify the effectiveness of the framework, we conduct experiments on electronic and paper hand-drawn floor plans. Compared with other state-of-the-art methods, our proposed method achieves superior reconstruction results. This work expands the application scenarios for indoor robots, enabling them to quickly comprehend the semantics of complex scenes, thereby enhancing the competitiveness in downstream tasks. Zhentong Zhang, Xinde Li, Chuanfei Hu, Fir Dunkin |
IROS | 4 |
| 2024 | Like draws to like: A Multi-granularity Ball-Intra Fusion approach for fault diagnosis models to resists misleading by noisy labels
Fir Dunkin, Xinde Li, Chuanfei Hu, Guoliang Wu, Heqing Li, Zhentong Zhang |
Adv. Eng. Informatics | 3 |
| 2024 | Empowering intelligent manufacturing with edge computing: A portable diagnosis and distance localization approach for bearing faults
Hairui Fang, Jialin An, Jingyu Bai, Jiawei Xiang, Wenjie Bai, Siyuan Fan, Chuanfei Hu, Fir Dunkin |
Adv. Eng. Informatics | 11 |
| 2024 | Trustworthy multi-phase liver tumor segmentation via evidence-based uncertainty
Chuanfei Hu, Tianyi Xia, Quchen Zou, Yuancheng Wang, Shenghong Ju, Xinde Li |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Semi-supervised human action recognition via dual-stream cross-fusion and class-aware memory bank
Biaozhang Huang, Shaojiang Wang, Chuanfei Hu, Xinde Li |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | TMFF: Trustworthy Multi-Focus Fusion Framework for Multi-Label Sewer Defect Classification in Sewer Inspection VideosabstractAn automatic vision-based sewer inspection plays a vital role of sewage system in a modern city. Recent advances focus on modeling a deep learning-based method to realize the sewer inspection system, benefiting from the capability of data-driven feature extraction. Although the acceptable performances of sewer defect classification are achieved, there is still a gap between the emerged methods and actual application scenarios. The first issue is that the multi-focus complementarity is ignored to represent the sewer defect, resulting in capturing the multi-scale information of sewer defect inefficiently. Second, the inherent uncertainty of sewer defect is not considered, while the serious unknown sewer defect categories would be missed, resulting in the untrustworthy sewer inspection. In this paper, we focus on quick-view (QV)-based sewer inspection, while a trustworthy multi-focus fusion framework (TMFF) is proposed, jointly combining multi-label classification and uncertainty estimation. Specifically, focal segment module (FSM) is designed based on optical flow to split the QV sewer video into long-focus and short-focus segments, where the multi-focus segments can be modeled to represent the multi-scale information of sewer defect. Then, evidential deep learning (EDL) is introduced to quantify the uncertainty, while joint expert scheme (JES) is designed to aggregate the expert opinions of multi-focus segments. Moreover, evidential disambiguating strategy (EDS) is proposed to alleviate the ambiguity of uncertainty estimation. Extensive experiments are conducted on VideoPipe, in which the superiority of TMFF is demonstrated compared with the state-of-the-art methods. Furthermore, we validate the potential capability of TMFF against the unknown cases of sewer defects. Chuanfei Hu, Chenyang Zhao 0009, Hang Shao 0001, Jin Deng, Yongxiong Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | TranPhys: Spatiotemporal Masked Transformer Steered Remote Photoplethysmography EstimationabstractSubtle variations are invisible to the naked eyes in human physiological signals can reflect important biological and health indicators. Although numerous computer vision methods have been proposed to recover and magnify these changes, most of them either only focus on identifying and recognizing explicit features such as shapes and textures, or are weak in long-term temporal modeling and spatiotemporal interactive perception of implicit biometrics. Therefore, it is difficult for them to robustly overcome various disturbances that affect detection performance. To address these issues, this paper presents TranPhys, a novel remote photoplethysmography (rPPG) network for facial video-based heart rate estimation. Specifically, first, we argue that facial subregions vary over time due to their biological personalities. So we split the input face video into multiple spatiotemporal tubes, build the 3D vision transformer with encoders and decoders to adequately model the high-dimensional representations of the respective regulars in each subregion, and globally coordinate their feedback on the cardiac pulsing waveform. Second, we design the temporal pooling attention to more finely mine the subtle changes hidden in the skin color over time and their long-term contextual rhythm cues. Third, we leverage the self-supervised masked autoencoding paradigm to overcome redundancy to enhance the robustness of our model, and construct the targeted spatiotemporal sampling maps instead of raw input sequences as the pretrained constraint labels to fully inspire self-supervision. We train, validate, and practice our TranPhys on multiple public datasets to demonstrate that our method achieves the competitive performance in remote heart rate estimation. Hang Shao 0001, Lei Luo 0001, Jianjun Qian, Shuo Chen 0003, Chuanfei Hu, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | ESUAV-NI: Endogenous Security Framework for UAV Perception System Based on Neural ImmunityabstractUnmanned aerial vehicles (UAVs) represent an essential component of advanced intelligent equipment that can be used as an aerial perception system by installing various sensors such as vision, hearing, touch, taste, and smell to achieve intelligently integrated perception of environments. However, these perception system with environmental information may be threatened by various internal and external attacks, causing a great challenge to the security of the UAV. The original security system relied on an expert knowledge base to prevent attacks, but the weaknesses of lacking proactivity and flexibility are gradually exposed. The strong resistance and survivability of biological systems can be used to fill this capability gap and provide new ideas for the security of the UAV perception system. Therefore, an endogenous security framework (ESUAV-NI) based on the neural system and immune system is proposed in this article. Through breeding artificial intelligence (AI) vaccines and distributed neural hierarchical control, we achieve the security protection for the UAV perception system. Moreover, we evaluated the AI vaccine breeding approach in the ESUAV-NI by conducting extensive experiments on internal threats and external aerial imagery camouflage data, respectively. The results show that the proposed approach has a superior performance for the UAV perception system. Heqing Li, Xinde Li, Zhentong Zhang, Chuanfei Hu, Fir Dunkin, Shuzhi Sam Ge |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Towards Trustworthy Multi-Label Sewer Defect Classification via Evidential Deep LearningabstractAn automatic vision-based sewer inspection plays a key role of sewage system in a modern city. Recent advances focus on utilizing deep learning model to realize the sewer inspection system, benefiting from the capability of data-driven feature representation. However, the inherent uncertainty of sewer defects is ignored, resulting in the missed detection of serious unknown sewer defect categories. In this paper, we propose a trustworthy multi-label sewer defect classification (TMSDC) method, which can quantify the uncertainty of sewer defect prediction via evidential deep learning. Meanwhile, a novel expert base rate assignment (EBRA) is proposed to introduce the expert knowledge for describing reliable evidences in practical situations. Experimental results demonstrate the effectiveness of TMSDC and the superior capability of uncertainty estimation is achieved on the latest public benchmark. Chenyang Zhao 0009, Chuanfei Hu, Hang Shao 0001, Zhe Wang 0027, Yongxiong Wang |
ICASSP | 2 |
| 2023 | A two-branch deep learning with spatial and pose constraints for social group detection
Xinde Li, Chuanfei Hu, Jin Deng, Weijie Sheng 0001, Lianli Zhu |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Hyperbolic embedding steered spatiotemporal graph convolutional network for video-based remote heart rate estimation
Hang Shao 0001, Lei Luo 0001, Shuo Chen 0003, Chuanfei Hu, Jian Yang 0003 |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | TNTC: Two-Stream Network with Transformer-Based Complementarity for Gait-Based Emotion RecognitionabstractRecognizing the human emotion automatically from visual characteristics plays a vital role in many intelligent applications. Recently, gait-based emotion recognition, especially gait skeletons-based characteristic, has attracted much attention, while many available methods have been proposed gradually. The popular pipeline is to first extract affective features from joint skeletons, and then aggregate the skeleton joint and affective features as the feature vector for classifying the emotion. However, the aggregation procedure of these emerged methods might be rigid, resulting in insufficiently exploiting the complementary relationship between skeleton joint and affective features. Meanwhile, the long range dependencies in both spatial and temporal domains of the gait sequence are scarcely considered. To address these issues, we propose a novel two-stream network with transformer-based complementarity, termed as TNTC. Skeleton joint and affective features are encoded into two individual images as the inputs of two streams, respectively. A new transformer-based complementarity module (TCM) is proposed to bridge the complementarity between two streams hierarchically via capturing long range dependencies. Experimental results demonstrate that TNTC outperforms state-of-the-art methods on the latest dataset in terms of accuracy. Chuanfei Hu, Weijie Sheng 0001, Bo Dong 0001, Xinde Li |
ICASSP | 1 |
| 2022 | An Anomaly Detection Method Based on Self-Supervised Learning with Soft Label Assignment for Defect Visual InspectionabstractRecently, local-editing-based transformations are introduced in anomaly detection for defect visual inspection, which construct a pretext task with the paradigm of self-supervised learning. However, supervised information of local-editing-based transformation may be incorrect when invalid trans-formation occurs in the pretext task. The reason is that the conventional method to generate labels ignores the differences of images between before and after the transformation. To address this issue, we propose soft label assignment (SLA) to construct soft labels via measuring the similarity between the original and transformed images. Meanwhile, a novel self-supervised learning-based anomaly detection method is proposed for defect visual inspection, which exploits local-editing-based transformation with SLA as a pretext classification task. A convolutional neural network (CNN) is trained to extract deep features of ambiguity and irregularity by the pretext classification task. In the main task, an anomaly detection is modeled via the deep representations to estimate defects regarded as anomalies. Experimental results demonstrate the effect of SLA, and the proposed method achieves superior performance than state-of-the-art methods in terms of the receiver operating characteristic curve (AUC-ROC). Chuanfei Hu, Yongxiong Wang |
ICASSP | 1 |
| 2021 | A Semantic-Enhanced Method Based On Deep SVDD for Pixel-Wise Anomaly DetectionabstractDetecting the anomalous information in multimedia is valuable to many computer vision applications. Recently, many pixel-wise methods modeling by deep learning model have been presented, which can be divided in reconstruction-based and distance-based methods. However, reconstruction-based methods suffer from the low precision of pixel reconstructions. Distance-based methods extract the hierarchical features by a pre-trained model, in order to estimate the anomalies by distances between normal and anomalous features. Nevertheless, multi-level features are ignored in these methods, and semantic information is not considered which is important to enhance the description of anomalies. To over-come the problems, we propose a novel semantic-enhanced anomaly detection method based on deep Support Vector Data Description (SVDD). A new semantic correlation module (SCB) is introduced to enhance the semantic information of the feature representations by cosine similarity. Mean-while, the multi-level architecture is utilized to estimate the final pixel-wise anomaly score. Experimental results demonstrate the proposed method outperforms state-of-the-art methods on MVTec and STC dataset. Chuanfei Hu, Hang Shao 0001 |
ICME | 1 |
| 2021 | BCNet: Bidirectional collaboration network for edge-guided salient object detection
Bo Dong 0001, Chuanfei Hu, Keren Fu, Geng Chen 0001 |
Neurocomputing | 3 |
| 2020 | Salient Object Detection with Boundary InformationabstractHow to distinguish the low-contrast area near boundaries is a basic challenge in salient object detection. Most of recent state-of-the-art methods can achieve a good performance but still can't work well near boundaries. In this paper, we propose a novel network based on multi-level feature fusion with boundary information to solve this problem. Our model includes two separate decoding sub-networks, one is object sub-network to detect salient objects and another is boundary sub-network which outputs error maps to get boundary information by boundary maps. Moreover, we design a connection and fusion module to exchange and fuse information of objects and boundaries. In addition, to balance the two subnetworks, the optimal weight of loss function is obtained by experiments. The experimental results show that our model can distinguish the low-contrast area near boundaries well by boundary information and achieves the state-of-the-art performance on five common datasets. Yongxiong Wang, Chuanfei Hu, Hang Shao 0001 |
ICME | 3 |
| 2020 | A saliency-guided clothing attribute recognition method by fusing salient prior informationabstractConvolutional neural network (CNN) based clothing attribute recognition has been applied in many clothing-related applications, such as recommendation system and clothes retrieval. However, the outperformance of existing recognition approaches is limited on account of the requirement of manual spatial prior information, such as landmarks and bounding box. In this paper, we innovatively transfer existing instance-irrelative knowledge to our proposed method where salient prior information is employed to assist the attribute prediction to avoid such laborsome annotations. Concretely, we first propose a saliency-guided method composed of salient object detection network (SOD-N) and clothing attribute recognition network (CAR-N). SOD-N provides the saliency map as the prior information guiding CAR-N to focus on the valid region. Furthermore, a new learnable fusion module is designed in CAR-N to aggregate the high-level features from the deep salient and clothing features. That can strengthen the ability of CAR-N to represent the correlative clothing attributes, resulting in further improving the final performance. The effectiveness of our method is demonstrated by the experimental results on clothing- related datasets, which improves among 1% to 10% mAP on each attributes category and almost 5% mean AP. Chuanfei Hu, Dong Bo |
SMC | 1 |
| 2020 | A CNN Model for Herb Identification Based on Part Priority Attention MechanismabstractAutomated herb identification plays an important role in protecting and investigating the herbs for botanists, which has been widely applied in the field of cosmetic, medical and food industry areas. Traditionally, due to the complicated background and various herb patterns, herb discriminative feature extractions is a hard work. And some existing methods may not be applicable when the herbs exist in actual wild environment. Therefore, how to locate the valid herb regions and extract the effective features is an open issue. In this paper, a novel CNN model is proposed for herb identification by using the part-information perception module (PPM) and species classification module (SCM). A new attention mechanism, namely part priority attention mechanism (PPAM), is proposed by training PPM independently with herb part labels. It should be pointed out that the proposed PPAM can guide the model to focus on the position of herb parts and suppress the irrelevant noisy regions. Moreover, depthwise separable convolution and label smoothing technique are introduced to decrease the model complexity and regularize the impact of mistake labels. In the experiment part, a large-scale herb dataset is constructed, which consists many kinds of challenging herb species in the wild environment. Additionally, these images contains both species labels and part labels. Experimental results demonstrate that our model achieves obviously improvement in term of accuracy and model size compared with other classical deep learning models. Zhanquan Sun 0001, Engang Tian, Chuanfei Hu, Hui Zong |
SMC | 4 |