EDBT 2026 Demo / reviewers in the wild / expert
Huanjie Tao
dblp:220/1301
· DBLP profile ↗
31ranked-venue papers
21as first author
25since 2021 · last 2026
0000-0002-1453-2468ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 7 first-author · 8 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accuracy evaluation of classification models on partially labeled datasets
Huanjie Tao, Wu Gao |
Expert Syst. Appl. | 1 |
| 2026 | DAAM: semi-supervised continual semantic segmentation based on distribution alignment and affinity matching
Huanjie Tao |
Expert Syst. Appl. | 2 |
| 2026 | MRKD-PBCL: Multi-level region-wise knowledge distillation and prototype balanced contrastive learning for class incremental semantic segmentation
Huanjie Tao, Qiuhan Zhang, Benran Li |
Neurocomputing | 2 |
| 2026 | Towards robust incomplete multimodal open-set domain generalization with uncertain missing modalities
Xin Chen 0121, Huanjie Tao, Benran Li |
Knowl. Based Syst. | 2 |
| 2026 | Unifying Modality and Scale: Visual Mamba for Feature Fusion in RGB-X Crowd CountingabstractCrowd counting has long been a crucial topic in the domains of computer vision and video surveillance. In particular, with the widespread use of thermal or depth cameras, RGB-X crowd counting has emerged as a prominent research focus. Although depth or infrared images provide complementary information, the core challenge remains in effectively unifying heterogeneous cross-modality and cross-scale information to form a comprehensive representation of crowd distributions in complex scenes. To address this problem, we propose a novel Mamba-based framework, termed UMS-VMamba-CC for multi-modal (RGB-X) crowd counting. Specifically, we design the cross-modality disentanglement fusion visual Mamba (CMDF-VMamba) that uses self-supervised learning to decompose modality-invariant and modality-specific features in spatial-frequency domains, followed by multi-modal feature aggregation via the gating mechanism. For cross-scale fusion, we design the cross-scale pyramid fusion visual Mamba (CSPF-VMamba), which adopts the bi-directional pyramid structure to incorporate low-level features into high-level representations during downsampling, while upsampling high-level features and aggregating them with low-level features through state space contextual modeling. Comprehensive experiments on multiple mainstream datasets demonstrate that the UMS-VMamba-CC framework achieves competitive performance for RGB-X crowd counting. Yaocong Hu, Mengbo Jia, Pindeng Wang, Wenbo Zhu 0002, Huanjie Tao, Tianming Ni, Teng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | A weakly supervised pavement crack segmentation based on adversarial learning and transformers
Yvon Apedo, Huanjie Tao |
Multim. Syst. | 2 |
| 2025 | Enhanced Multiview attention network with random interpolation resize for few-shot surface defect detection
Penghao Li, Huanjie Tao, Yishi Deng |
Multim. Syst. | 2 |
| 2025 | A content-style control network with style contrastive learning for underwater image enhancement
Zhenguang Wang, Huanjie Tao, Yishi Deng |
Multim. Syst. | 2 |
| 2025 | An end-to-end repair-based joint training framework for weakly supervised pavement crack segmentation
Huanjie Tao, Qianyue Duan, Zhenwu Hu, Yishi Deng |
Multim. Tools Appl. | 2 |
| 2025 | Residual Quotient Learning for Zero-Reference Low-Light Image EnhancementabstractRecently, neural networks have become the dominant approach to low-light image enhancement (LLIE), with at least one-third of them adopting a Retinex-related architecture. However, through in-depth analysis, we contend that this most widely accepted LLIE structure is suboptimal, particularly when addressing the non-uniform illumination commonly observed in natural images. In this paper, we present a novel variant learning framework, termed residual quotient learning, to substantially alleviate this issue. Instead of following the existing Retinex-related decomposition-enhancement-reconstruction process, our basic idea is to explicitly reformulate the light enhancement task as adaptively predicting the latent quotient with reference to the original low-light input using a residual learning fashion. By leveraging the proposed residual quotient learning, we develop a lightweight yet effective network called ResQ-Net. This network features enhanced non-uniform illumination modeling capabilities, making it more suitable for real-world LLIE tasks. Moreover, due to its well-designed structure and reference-free loss function, ResQ-Net is flexible in training as it allows for zero-reference optimization, which further enhances the generalization and adaptability of our entire framework. Extensive experiments on various benchmark datasets demonstrate the merits and effectiveness of the proposed residual quotient learning, and our trained ResQ-Net outperforms state-of-the-art methods both qualitatively and quantitatively. Furthermore, a practical application in dark face detection is explored, and the preliminary results confirm the potential and feasibility of our method in real-world scenarios. Linfeng Fei, Huanjie Tao, Yaocong Hu, Wei Zhou 0042, Jiun Tian Hoe, Weipeng Hu, Yap-Peng Tan |
IEEE Trans. Image Process. | 3 |
| 2024 | A label-relevance multi-direction interaction network with enhanced deformable convolution for forest smoke recognition
Huanjie Tao |
Expert Syst. Appl. | 1 |
| 2024 | Smoke Recognition in Satellite Imagery via an Attention Pyramid Network With Bidirectional Multilevel Multigranularity Feature Aggregation and Gated FusionabstractMingyuan Ren, Xiuwen Fu, Pasquale Pace, Gianluca Aloi, and Giancarlo FortinoRecognizing smoke in satellite imagery is a critical approach in an Internet of Things (IoT) system for monitoring forest fires. However, the task remains challenging due to false alarms of smoke-like occurrences caused by complex land cover types, and missing detections caused by the diversity of fire smoke. Some reasons are that existing methods overlook attention granularity, neglect all-layer-based fusion of low-level features with high-level semantic information, and fail to address interferences arising from fusing different kinds of features. To solve these issues, this paper presents an attention pyramid network with bidirectional multi-level multi-granularity feature aggregation and gated fusion for smoke recognition. First, to guide the model sequentially extract multi-granularity smoke attention clues for complementary smoke perception, we design an attention-guided feature pyramid module by concatenating residual blocks and attention pyramid blocks. Second, to leverage both low-level fine-grained and high-level semantic features in all network layers, we design a bidirectional feature aggregation module using multi-level multi-granularity feature blocks. Finally, to selectively integrate the features with different resolutions and semantic levels to effectively achieve feature complementarity and avoid feature mutual interference, we design a gated feature fusion module using gated feature fusion blocks. The experimental results demonstrate that our model achieves an accuracy of 98.33% on the USTC-SmokeRS dataset. Additionally, on the E-USTC-SmokeRS dataset, our model achieves a detection rate of 94.92%, a false alarm rate of 3.00%, and an F1-score of 0.9553. These results surpass the performance of existing satellite-imagery-based smoke recognition methods. Huanjie Tao |
IEEE Internet Things J. | 1 |
| 2024 | A Spatial-Channel Feature-Enriched Module Based on Multicontext Statistics AttentionabstractConvolutional neural networks (CNNs) have demonstrated remarkable performance in various computer vision tasks, such as image classification, semantic segmentation, and object detection. However, learning discriminative and generalizable features that can overcome both intra-class variations and inter-class ambiguity remains a challenging problem. To address this issue, we propose a novel module called the spatial-channel feature-enriched module (SCFEM), which can be easily integrated into existing CNN architectures. Specifically, we propose a refined multi-scale residual block (RMRB) to dynamically produce channel-wise weights via a shared aggregation gate to selectively fuse multi-scale features. Moreover, a multi-context statistics attention block (MSAB) is proposed to explore both the semantic dependency between channels and the long-range spatial contextual dependency between pixels via multi-context channel attention (MCA) and statistic spatial attention (SSA). MCA captures the local context information of each channel by considering the associated neighbor channels. SSA captures complex long-range dependencies and discriminative but subtle differences among pixels by leveraging high-order statistics. Experimental results demonstrate the superiority of our SCFEM over state-of-the-art methods on multiple computer vision tasks. Huanjie Tao, Qianyue Duan |
IEEE Internet Things J. | 1 |
| 2024 | Hierarchical and progressive learning with key point sensitive loss for sonar image classification
Xin Chen 0121, Huanjie Tao, Yishi Deng |
Multim. Syst. | 2 |
| 2024 | Hierarchical attention network with progressive feature fusion for facial expression recognition
Huanjie Tao, Qianyue Duan |
Neural Networks | 1 |
| 2024 | Weakly-Supervised Pavement Surface Crack Segmentation Based on Dual Separation and Domain GeneralizationabstractAutomatic pavement surface crack segmentation is crucial for efficient and cost-effective road maintenance. Despite fully-supervised crack segmentation methods have achieved significant success, the laborious task of pixel-level annotation hampers their widespread applicability. To address this issue, this paper presents a weakly-supervised pavement surface crack segmentation method based on Dual Separation and Domain Generalization (DSDGNet). Firstly, a crack image formulation model (CIFM) is developed by separating the crack image into a background component and a crack component. Additionally, we treat the crack component as a linear fusion of the pavement texture component and the crack mask. Secondly, a local-to-global learning method (L2G-L) is proposed to learn complete crack via local learning based on a random cropping and pasting algorithm. This idea stems from the observation that the crack component can be separated into several local regions, akin to the local regions found in hand-drawn crack components. Thirdly, A progressive interaction training algorithm (PIT) is crafted to train the image generation model by leveraging both generated and real images, thereby narrowing the divide between generated and authentic crack images. Finally, realistic and diverse crack images, along with their crack masks, are generated to facilitate the training of fully-supervised segmentation models. A generalizable loss is proposed to enhance the model generalization ability by combining reconstruction, segmentation, and domain adversarial losses. Extensive experiments on six public pavement crack datasets show the effectiveness and superiority of DSDGNet in weakly-supervised methods. Huanjie Tao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | An adaptive frame selection network with enhanced dilated convolution for video smoke recognition
Huanjie Tao, Qianyue Duan |
Expert Syst. Appl. | 1 |
| 2023 | A gated multi-hierarchical feature fusion network for recognizing steel plate surface defects
Huanjie Tao, Minghao Lu, Zhenwu Hu, Jianfeng An |
Multim. Syst. | 1 |
| 2023 | An improved interaction-and-aggregation network for person re-identification
Huanjie Tao, Wenjie Bao, Qianyue Duan, Zhenwu Hu, Jianfeng An |
Multim. Tools Appl. | 1 |
| 2023 | Controllable smoke image generation network based on smoke imaging principle
Huanjie Tao, Jing Wang 0162, Zhouxin Xin |
Multim. Tools Appl. | 1 |
| 2023 | Learning discriminative feature representation with pixel-level supervision for forest smoke recognition
Huanjie Tao, Qianyue Duan, Minghao Lu, Zhenwu Hu |
Pattern Recognit. | 1 |
| 2023 | An Adaptive Interference Removal Framework for Video Person Re-IdentificationabstractVideo person re-identification (V-ReID) can leverage rich spatial-temporal information embedded in sequence data to achieve better accuracy. However, it is vulnerable to the interference of inaccurate and redundant noisy frames in each sequence as well as the background clutter and person-irrelevant pixels in each frame. To solve the above issues, this paper presents an adaptive interference removal framework (IRF) to learn discriminative feature representations by removing various interference. Our IRF mainly consists of two modules including an attention-guided adaptive interference frame removal module (IFRM) and an attention-guided adaptive interference pixel removal module (IPRM). IFRM and IPRM are designed to locate task-relevant keyframes and key pixels, respectively. IFRM adopts the attention mechanism to predict frame-wise scores to characterize the contribution of each frame to the final identification task. IPRM collaboratively utilizes camera identity classification loss, person identity classification loss, target attention loss, and person mask adversarial loss for extracting pure pedestrian representations. A progressive mask augmentation strategy is designed to restrain the data distribution of the generated person masks to further guide model training. Extensive experiments demonstrate that our models outperform state-of-the-art accuracy on seven person ReID datasets. Huanjie Tao, Qianyue Duan, Jianfeng An |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | CENet: A Channel-Enhanced Spatiotemporal Network With Sufficient Supervision Information for Recognizing Industrial Smoke EmissionsabstractVision-based industrial smoke emission recognition technology can identify smoke emissions and provide a visual evidence for humans to pursue environmental justice. However, the existing methods still face the issues of low detection rates (DRs) and high false alarm rates (FARs) due to the insufficient supervision information and limited feature representation capability. To solve these issues, this article presents a channel-enhanced spatiotemporal network (CENet) with sufficient supervision information for recognizing industrial smoke emissions. First, to provide sufficient supervision information for learning discriminative feature representation, we propose a new loss function by collaboratively taking the binary category, pixel-level smoke density, and background information as supervision information and use them in the network final layer as well as the network middle layers to guide model training. Second, to solve the deficiencies of max/average pooling and convolution operations in feature extraction, we propose channel-enhanced modules, including channel-enhanced pooling (CEPool), channel-enhanced convolution (CEConv), and channel-enhanced upsampling (CEUpsample) to learn high-response values of bright features as well as the low-response values with discriminative features. The channel-enhanced modules selectively enhance the learned features with large amount of information and suppress those useless features. Third, we propose a two-stage method based on spatiotemporal information extraction module (SIEM) and smoke recognition module (SRM), which are designed to learn spatiotemporal information between input frames and that between smoke density frames, respectively. Extensive experiments show that CENet achieves the best performance among the existing smoke recognition methods. Huanjie Tao, Jing Wang 0162, Zhouxin Xin |
IEEE Internet Things J. | 1 |
| 2022 | Attention-Aggregated Attribute-Aware Network With Redundancy Reduction Convolution for Video-Based Industrial Smoke Emission RecognitionabstractExisting video-based industrial smoke emission recognition methods face the issues of low detection rates and high false alarm rates. An important reason is that they only consider binary category information as supervision information and ignore video attribute information, which provides important supplementary in improving model performances. To solve it, we propose an attention-aggregated attribute-aware network (AANet). First, to effectively guide the model for discriminative feature learning, a video attribute information decoding module is proposed to increase supervision information by designing attribute vector construction and attribute information decoding methods. Second, to learn discriminative feature representations, some attentions are designed to aggregate spatiotemporal and context information based on ConvLSTM, global feature extraction, and cascaded pyramid attention. Final, the redundancy reduction convolution is proposed to reduce redundant channels by channelwise weights considering matrix elements summation and information spatial distribution characterized by information entropy. Extensive experiments show that AANet significantly outperforms existing methods Huanjie Tao, Minghao Lu, Zhenwu Hu, Zhouxin Xin, Jing Wang 0162 |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Learning Discriminative Feature Representation for Estimating Smoke Density of Smoky Vehicle RearabstractEstimating the smoke density from a single image of the smoky vehicle rear is significant for smoke level estimation and smoky vehicle recognition. However, this is a highly ill-posed problem. To solve it, this paper presents a novel smoke density estimation network (SDENet). First, to reduce the susceptibility of deeper convolutions towards the smoke scale variant and enhance feature diversity, we propose the attention multi-scale encoding blocks based on multi-scale blocks and the convolutional block attention. The multi-scale block exploits the multi-scale features at a granular level within a single basic block. The attention makes our model pay more attention to specific regions which are beneficial to smoke density estimation. Second, the extracted features for smoke recognition may contain background interference information due to the smoke translucency, so we propose semantic-guided feature selection blocks to progressively select smoke-relevant features and suppress background interference features from the encoded feature via global high-level semantic information for learning more discriminative features. Finally, to make the learned features more adaptive to smoke feature resolution and visual appearance, we design attention gate decoding blocks to fuse different features via gate blocks, which enhance the features at spatial locations and channel locations where the features are essential for smoke density estimation. Extensive experiments on smoke density estimation show that our model achieves the best performance among existing methods. Huanjie Tao, Qianyue Duan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Detecting smoky vehicles from traffic surveillance videos based on dynamic features
Huanjie Tao |
Appl. Intell. | 1 |
| 2020 | A three-stage framework for smoky vehicle detection in traffic surveillance videos
Huanjie Tao, Xiaobo Lu |
Inf. Sci. | 1 |
| 2020 | Smoke Vehicle Detection Based on Spatiotemporal Bag-Of-Features and Professional Convolutional Neural NetworkabstractExisting smoke vehicle detection methods are vulnerable to false alarms. To solve this issue, this paper presents two automatic smoke vehicle detection methods based on spatiotemporal bag-of-features (S-BoF) and professional convolutional neural network (P-CNN). In the first method, we propose the S-BoF model to characterize the key regions detected by the visual background extractor (ViBe) algorithm. The S-BoF model contains three groups of features, including color moments on three orthogonal planes (CM-TOP), completed robust local binary pattern on three orthogonal planes (CRLBP-TOP), and histogram of oriented gradient on three orthogonal planes (HOG-TOP). The extracted features are fed to the support vector machine (SVM) and classify the key regions to smoke regions or non-smoke regions to further detect smoke vehicles. In the second method, we propose the P-CNN model to extract more robust and complementary spatiotemporal features by designing three professional models to analyze different kinds of features in the key region sequence on three orthogonal planes. The three professional models, including color CNN (CCNN), texture CNN (TCNN), and gradient CNN (GCNN), are based on three independent CNN128 models with different inputs. The experimental results show that the proposed methods achieve higher detection rates and lower false alarm rates than existing smoke detection methods. Huanjie Tao, Xiaobo Lu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Smoke vehicle detection based on robust codebook model and robust volume local binary count patterns
Huanjie Tao, Xiaobo Lu |
Image Vis. Comput. | 1 |
| 2018 | Correction of micro-CT image geometric artefacts based on markerabstractSmall geometric misalignments of micro computed tomography (CT) system will cause geometric artefacts in the reconstructed image. A new correction method of geometrical artefacts based on marker and non‐linear optimisation model is proposed. In this method, the simple balls marker and the measured objects are scanned simultaneously, and the geometric parameters of the micro‐CT system are precisely estimated by solving the non‐linear optimisation model which is based on the scanning data. The geometric artefacts caused by geometric parameters are corrected and the authors can reconstruct the image correctly. In addition to estimating geometric parameters for the traditional scanning mode, the proposed method can also be used for the limited angle CT scanning and the half detector CT scanning. Simulated experiments and real experiments verify that the correction method effectively decrease the geometric artefacts of micro‐CT images. Huanjie Tao, Xiaobo Lu |
IET Image Process. | 1 |
| 2018 | Smoky vehicle detection based on multi-feature fusion and ensemble neural networks
Huanjie Tao, Xiaobo Lu |
Multim. Tools Appl. | 1 |