VLDB 2026 Research / reviewers in the wild / expert
Yichao Cao
dblp:160/6077
· DBLP profile ↗
37ranked-venue papers
9as first author
36since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 6 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud Trackingabstract3D single object tracking (SOT) in LiDAR point clouds is a critical task in computer vision and autonomous driving. Despite great success having been achieved, the inherent sparsity of point clouds introduces a dual-redundancy challenge that limits existing trackers: (1) vast spatial redundancy from background noise impairs accuracy, and (2) informational redundancy within the foreground hinders efficiency. To tackle these issues, we propose CompTrack, a novel end-to-end framework that systematically eliminates both forms of redundancy in point clouds. First, CompTrack incorporates a Spatial Foreground Predictor (SFP) module to filter out irrelevant background noise based on information entropy, addressing spatial redundancy. Subsequently, its core is an Information Bottleneck-guided Dynamic Token Compression (IB-DTC) module that eliminates the informational redundancy within the foreground. Theoretically grounded in low-rank approximation, this module leverages an online SVD analysis to adaptively compress the redundant foreground into a compact and highly informative set of proxy tokens. Extensive experiments on KITTI, nuScenes and Waymo datasets demonstrate that CompTrack achieves top-performing tracking performance with superior efficiency, running at a real-time 90 FPS on a single RTX 3090 GPU. Sifan Zhou, Yichao Cao, Jiahao Nie 0001, Yuqian Fu, Xiaobo Lu, Shuo Wang 0030 |
AAAI | 2 |
| 2026 | Refining the granularity of smoke representation: SAM-powered density-aware progressive smoke segmentation framework
Yichao Cao, Xuanpeng Li, Xiaolin Meng, Xiaobo Lu |
Pattern Recognit. | 1 |
| 2026 | Incremental mixture of experts: Continual learning for object detection in forestry scenarios
Ximeng Cheng, Qiaonan Zhu, Shukun Jia, Yichao Cao, Xiaobo Lu |
Pattern Recognit. | 4 |
| 2026 | Tracking by detection and query: An efficient end-to-end framework for multi-object tracking
Shukun Jia, Yichao Cao, Xin Lu 0007, Xiaobo Lu |
Pattern Recognit. | 3 |
| 2025 | Perturbating, Tuning, and Collaborating: Harnessing Vision Foundation Models for Single Domain Generalization on Medical ImagingabstractSingle Domain Generalization (SDG) is critical in medical imaging applications. Recently, Vision Foundation Models (VFMs) have spearheaded a trend in AI development due to their robust generalizability and versatility. This work aims to fully explore the generalization capabilities of VFMs alongside the domain-specific expertise of specialized models, thoroughly investigating the boundaries of their respective capabilities, thereby collaboratively addressing SDG challenges within medical imaging. We propose a framework for Collaborative reasoning between Specialized and Universal models for Single Domain Generalization (CollaSU-SDG) in medical imaging. Specifically, we first design a model-aware perturbation injection method from the perspective of single-source domain data, enabling differentiated and adaptive perturbation injection for two different scales of models. Then, a domain expansion adapter is designed for the VFM to adapt to the augmented single-source domain medical data. Lastly, we introduce an adaptive hierarchical transfer and dynamic dense prompting method that facilitate collaborative reasoning between the specialized and universal models, eliminating the need for explicit prompts. Through these designs, CollaSU-SDG fully leverages the strengths of both specialized and universal models, achieving robust out-of-distribution generalization capabilities on single-source domain data. Experimental results demonstrate that CollaSU-SDG significantly advances the state-of-the-art performance across a wide range of medical datasets. All the code will be publicly available. Yichao Cao, YingYing Zhang, Xiu Su, Haogang Zhu |
AAAI | 2 |
| 2025 | A Novel Image-Graph Heterogeneous Fusion Framework for Static IR Drop PredictionabstractIR drop analysis is crucial for ensuring the reliability and performance of integrated circuits (ICs) but poses computational challenges as the IC designs grow larger, especially for ultra deep-submicron VLSI designs. Deep learnings (DL) as the efficiency-promising solutions, mainly employ various CNN-based networks to achieve image-to-image IR drop predictions. However, they neglect and lose the power delivery network (PDN) global spatial features and cell instance topological information. This paper proposes a novel image-graph heterogeneous fusion framework (IGHF), which integrates the effectiveness and complementarity of dual branches (CNN and GNN) for higher prediction performance. In the CNN-based Power ScaleFusion Unet branch, the proposed long-range and local-detail encoder (LLE) integrates seamlessly with the hierarchical and adjacent compensation group (HACG) module. This design facilitates effective multi-scale global-to-local spatial power feature extraction within the PDN and enables adaptive high-to-low-level feature fusion and compensation in the decoder. Moreover, a cell voltage aware (CVA) module in the GNN branch is designed to adaptively aggregate PDN topological features of heterogeneous neighbors of different orders. Comparative experiments demonstrate that the proposed IGHF achieves significant accuracy improvements, outperforming the state-of-the-art MAUNet and widely-used IREDGe methods by considerable margins of 24.6% and 55.0% reduction in prediction error, while the prediction maps possess higher structural fidelity. Transfer experiments indicate that IGHF with transfer learning can improve the accuracy in real circuits with the few-shot real circuit test cases. Dan Niu, Dekang Zhang, Yichao Cao, Zhou Jin 0001, Chao Wang 0120, Yichao Dong, Changyin Sun 0001 |
DAC | 3 |
| 2025 | Debiased Prototype Evolving for Point Cloud Domain Adaptation via 3D Foundation ModelsabstractDomain adaptation in point cloud data is essential for improving downstream tasks in autonomous driving, robotics, and 3D modeling. 3D Foundation models, driven by scaling laws, have significantly advanced point cloud applications by embedding rich semantic knowledge of geometric structures. However, a significant gap remains between their broad zero-shot generalization capabilities and the specialized requirements of domain adaptation tasks. Furthermore, pre-training can induce a model bias towards samples that resemble the pre-training dataset. To bridge this gap, we propose an Evolving Alignment strategy to apply large-scale 3D foundation models to domain adaptation in 3D point clouds, named EvoAlign3D. Specifically, we implement a joint domain alignment strategy to align the foundation model’s feature space with a transferable feature space across the source and target domains. Meanwhile, we propose a debiased prototype evolving method, which refines class-level prototypes and optimizes pseudo-label consistency, progressively mitigating model biases and enhancing the transferability of discriminative features for better cross-domain generalization. Our method significantly improves domain adaptation classification performance on the PointDA dataset, which spans both synthetic and real-world data domains. All the code and pre-trained weights will be publicly available. Yichao Cao, Xuanpeng Li |
ICASSP | 2 |
| 2025 | A Geometry-Material Aware Point Cloud Transformer for Large-scale Unstructured Thermal Analysis in 2.5D ICsabstractThermal management in large-scale unstructured 2.5D ICs faces the challenges due to the integration of complex geometries and heterogeneous materials. Existing deep learning (DL) methods urgently require a memory-efficient and high-fidelity unstructured representation method for multiscale complex ICs to simultaneously model macroscopic components and microscopic structure. Moreover, it further needs to achieve multiscale geometric thermal feature capture and thermal distribution difference adaptation among heterogeneous materials. Combining a multiscale unstructured point-cloud representation, this paper introduces Therm-PCT, a geometry-material aware point-cloud transformer framework to achieve high-accuracy thermal and its gradient prediction. Therm-PCT incorporates three key modules: adaptive multipath-coupled diffusion (AMD), a wavelet-based fine-grained recovery (WFR), and a thermal-aware Mixture-of-Material-Experts (TA-MoME) adapter. AMD adaptively learns heat diffusion path interaction with serialization-gate-based attention. Furthermore, the WFR module recovers fine-grained thermal gradients through high-frequency wavelet domain enhancement, and the TA-MoME adapter adapts to heterogeneous material by dynamically routing material-specific experts. Experiments demonstrate that the Thermal-PCT’s accuracy performance metric improvements are substantial, outperforming the newly proposed method FSA-Heat, by considerable margins of 78.03%, 84.00%, 67.61%, and 78.25% in 80 K-scale point clouds. It also achieves a 147× speed-up compared to the commercial software COMSOL. Additionally, Therm-PCT shows the potential of zero-shot generalization up to 0.4 M-scale points (5.7× than training scale) and robust performance on unseen geometric shapes. Dekang Zhang, Dan Niu, Yichao Cao, Yichao Dong, Zhenya Zhou, Zhou Jin 0001 |
ICCAD | 3 |
| 2025 | CounterPC: Counterfactual Feature Realignment for Unsupervised Domain Adaptation on Point Clouds
Yichao Cao, Xiu Su, Dan Niu, Xuanpeng Li |
ICCV | 2 |
| 2025 | TinyMIG: Transferring Generalization from Vision Foundation Models to Single-Domain Medical ImagingabstractMedical imaging faces significant challenges in single-domain generalization (SDG) due to the diversity of imaging devices and the variability among data collection centers. To address these challenges, we propose \textbf{TinyMIG}, a framework designed to transfer generalization capabilities from vision foundation models to medical imaging SDG. TinyMIG aims to enable lightweight specialized models to mimic the strong generalization capabilities of foundation models in terms of both global feature distribution and local fine-grained details during training. Specifically, for global feature distribution, we propose a Global Distribution Consistency Learning strategy that mimics the prior distributions of the foundation model layer by layer. For local fine-grained details, we further design a Localized Representation Alignment method, which promotes semantic alignment and generalization distillation between the specialized model and the foundation model. These mechanisms collectively enable the specialized model to achieve robust performance in diverse medical imaging scenarios. Extensive experiments on large-scale benchmarks demonstrate that TinyMIG, with extremely low computational cost, significantly outperforms state-of-the-art models, showcasing its superior SDG capabilities. All the code and model weights will be publicly available. Hongyan Xu 0002, Yichao Cao, Xiu Su, Tianfa Li, Shan An, Haogang Zhu |
ICML | 3 |
| 2025 | FocusTrack: One-Stage Focus-and-Suppress Framework for 3D Point Cloud Object Tracking
Sifan Zhou, Jiahao Nie 0001, Yichao Cao, Xiaobo Lu |
ACM Multimedia | 4 |
| 2025 | SmokeAgent: Multimodal agent for fine-grained smoke event analysis in large-scale wild environments
Yichao Cao, Xuanpeng Li, Xiaobo Lu |
Pattern Recognit. | 1 |
| 2025 | Improved Specific Emitter Identification Based on Margin Disparity Discrepancy in Varying Modulation ScenariosabstractIn Specific Emitter Identification (SEI), transmitters are typically distinguished through Radio Frequency Fingerprint (RFF) features. However, modulation schemes can be deliberately coupled to confound RFF information. This paper addresses modulation variation as a Domain Adaptation (DA) problem and proposes an SEI framework based on Margin Disparity Discrepancy (MDD) to enhance robustness in modulation-varying scenarios. Specifically, we first establish a theoretical tight upper bound for the discrepancy between modulation domains using MDD theory. Then, we design an adversarial network to align variable features to shorten the discrepancy between modulations. Finally, we experimented with complex modulated signals including digital and analog modulation. Numerical results indicate that our approach achieves an average improvement of over 20% in accuracy compared to classical SEI methods and outperforms traditional DA techniques. Yezhuo Zhang, Zinan Zhou, Yichao Cao, Xuanpeng Li |
IEEE Signal Process. Lett. | 3 |
| 2025 | Variational Feature Imitation Conditioned on Visual Descriptions for Few-Shot Fine-Grained RecognitionabstractIn few-shot fine-grained recognition (FS-FGR) tasks, the main challenge is to distinguish novel categories with high intra-class variations and low inter-class differences given scarce training data. Existing studies explore discriminative features through a compact network to avoid overfitting, while they achieve marginal performance gain owing to the limited representation capability. Motivated by the significant progress of the vision foundation model, we introduce it to describe visual attributes and boost the performance of the compact feature extractor. A few-shot fine-grained recognition method with Variational Feature Imitation Conditioned on Visual Descriptions, VFI-CVD for short, has been proposed in this paper. It simultaneously exploits the pre-trained knowledge from a vision foundation model and the expert knowledge mined by a feature extractor. Specifically, the intra-class variations shared across object categories are encoded into a common distribution thus we can augment features by sampling latent variables. To enhance the learning of intra-class variations, a condition exchange strategy (CES) is put forward to interact the knowledge between samples through feature cross-imitation. In the inference stage, the learned knowledge is further integrated through the joint prediction of visual descriptions and cross-imitated features. Comprehensive experimental results on four fine-grained benchmark datasets show that the proposed VFI-CVD achieves state-of-the-art performance, e.g., 90.37% under the 5-way 1-shot setting on CUB-200-2011. It surpasses existing methods by a large margin, especially in the challenging 30-way recognition tasks and cross-domain evaluation. The source code is publicly available:https://github.com/Lx-zjwf/VFI-CVD. Xin Lu 0007, Yixuan Pan, Yichao Cao, Xin Zhou 0030, Xiaobo Lu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | The Devil is in the Frequency: Constrained and Adaptive Fine-Grained Domain Perturbation for Robust Medical SegmentationabstractDomain generalization (DG) in medical image analysis is critical for achieving consistent and reliable diagnostics across diverse healthcare systems. However, domain shifts resulting from variations in imaging protocols, devices, and practices hinder accurate anatomical identification. While data augmentation shows promise, it struggles to generate diverse samples that bridge domain gaps and often distorts invariant anatomical features, compromising diagnostic integrity. This paper introduces the Adaptive Dual-Space Spectral Perturbation (AdaDSP) framework to address these issues at both broad and fine-grained levels. At the broad level, AdaDSP injects learnable spectral perturbations into input images and intermediate feature maps, significantly enhancing the diversity of the training data. At the fine-grained level, we propose a Fine-Grained Spectral Perturbation module that utilizes two lightweight attention mechanisms to capture sensitive frequency bands that hinder generalization. By injecting multivariate Gaussian noise within a mini-batch, this module better modulates the distribution of frequencies and accomplishes adaptive perturbation of sensitive frequency bands. Furthermore, we introduce a Universal Triple-stage Semantic Constraint Framework to encourage the networks to learn domain-invariant representations while retaining the discriminabtive capacity. Extensive experiments show that our method outperforms state-of-the-art benchmarks, with improvements of 2.40% and 2.99% in two notable medical imaging tasks, respectively. Yichao Cao, Haogang Zhu |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Image Anomaly Detection Based on Controllable Self-AugmentationabstractBased on data synthesis, anomaly detection (AD) methods often rely on external data for data synthesis. However, most external abnormal data exhibits strong randomness, which may lead to a reduced range of diversity among the synthesized data. In order to achieve a broader diversity in data synthesis, it is necessary to not only have highly diverse data but also to incorporate low-diversity noise data. To enhance the diversity range of the synthesized data, this study proposes a diversity measurement assisted by image self-representation: measuring the distance between noise data and normal data and quantitatively synthesizing diversified data by selecting diverse noise data for synthesis, namely, Diversified Synthesis (DS). Diversified Synthesis introduces patch measurement and a controllable enhancement module to establish controllable diversified enhanced data. The contribution of this study lies in proposing a novel diversified synthesis method, which achieves a broader diversity synthesis through the introduction of image self-representation-assisted diversity measurement and quantitative synthesis. Furthermore, through the self-enhancement data augmentation method, the use of image intrinsic features for enhancement achieves diversity and multi-scale characteristics in the synthesized data, thereby improving the training performance of the discriminative model. This provides an effective optimization solution for comprehensive anomaly detection methods. Liujie Hua, Yichao Cao, Yitian Long, Shan You, Xiu Su, Yueyi Luo, Chang Xu 0002 |
IJCNN | 2 |
| 2024 | Universal Frequency Domain Perturbation for Single-Source Domain GeneralizationabstractIn this work, we introduce a novel approach to single-source domain generalization (SDG) in medical imaging, focusing on overcoming the challenge of style variation in out-of-distribution (OOD) domains without requiring domain labels or additional generative models. We propose a Universal Frequency Perturbation framework for SDG termed as UniFreqSDG, that performs hierarchical feature-level frequency domain perturbations, facilitating the model's ability to handle diverse OOD styles. Specifically, we design a learnable spectral perturbation module that adaptively learns the frequency distribution range of samples, allowing for precise low-frequency (LF) perturbation. This adaptive approach not only generates stylistically diverse samples but also preserves domain-invariant anatomical features without the need for manual hyperparameter tuning. Then, the frequency features before and after perturbation are decoupled and recombined through the Content Preservation Reconstruction operation, effectively preventing the loss of discriminative content information. Furthermore, we introduce the Active Domain-variance Inducement Loss to encourage effective perturbation in the frequency domain while ensuring the sufficient decoupling of domain-invariant and domain-style features. Extensive experiments demonstrate that UniFreqSDG increases the dice score by an average of 7.47% (from 77.98% to 85.45%) on the fundus dataset and 4.99% (from 71.42% to 76.73%) on the prostate dataset compared to the state-of-the-art approaches. Yichao Cao, Xiu Su, Haogang Zhu |
ACM Multimedia | 2 |
| 2024 | Distilling object detectors with efficient logit mimicking and mask-guided feature imitation
Xin Lu 0007, Yichao Cao, Shikun Chen, Xin Zhou 0030, Xiaobo Lu |
Expert Syst. Appl. | 2 |
| 2024 | Multi-temporal dependency handling in video smoke recognition: A holistic approach spanning spatial, short-term, and long-term perspectives
Qifan Xue, Yichao Cao, Xuanpeng Li, Weigong Zhang |
Expert Syst. Appl. | 3 |
| 2024 | Towards better small object detection in UAV scenes: Aggregating more object-oriented information
Chenyue Yang, Yichao Cao, Xiaobo Lu |
Pattern Recognit. Lett. | 2 |
| 2024 | CEDR: Contrastive Embedding Distribution Refinement for 3D point cloud representation
Yichao Cao, Qifan Xue, Shuai Jin, Xuanpeng Li, Weigong Zhang |
Signal Process. Image Commun. | 2 |
| 2023 | Coarse2Fine: Local Consistency Aware Re-prediction for Weakly Supervised Object LocalizationabstractWeakly supervised object localization aims to localize objects of interest by using only image-level labels. Existing methods generally segment activation map by threshold to obtain mask and generate bounding box. However, the activation map is locally inconsistent, i.e., similar neighboring pixels of the same object are not equally activated, which leads to the blurred boundary issue: the localization result is sensitive to the threshold, and the mask obtained directly from the activation map loses the fine contours of the object, making it difficult to obtain a tight bounding box. In this paper, we introduce the Local Consistency Aware Re-prediction (LCAR) framework, which aims to recover the complete fine object mask from locally inconsistent activation map and hence obtain a tight bounding box. To this end, we propose the self-guided re-prediction module (SGRM), which employs a novel superpixel aggregation network to replace the post-processing of threshold segmentation. In order to derive more reliable pseudo label from the activation map to supervise the SGRM, we further design an affinity refinement module (ARM) that utilizes the original image feature to better align the activation map with the image appearance, and design a self-distillation CAM (SD-CAM) to alleviate the locator dependence on saliency. Experiments demonstrate that our LCAR outperforms the state-of-the-art on both the CUB-200-2011 and ILSVRC datasets, achieving 95.89% and 70.72% of GT-Know localization accuracy, respectively. Yixuan Pan, Yichao Cao, Chongjin Chen, Xiaobo Lu |
AAAI | 3 |
| 2023 | Re-mine, Learn and Reason: Exploring the Cross-modal Semantic Correlations for Language-guided HOI detectionabstractHuman-Object Interaction (HOI) detection is a challenging computer vision task that requires visual models to address the complex interactive relationship between humans and objects and predicttriplets. Despite the challenges posed by the numerous interaction combinations, they also offer opportunities for multi-modal learning of visual texts. In this paper, we present a systematic and unified framework (RmLR) that enhances HOI detection by incorporating structured text knowledge. Firstly, we qualitatively and quantitatively analyze the loss of interaction information in the two-stage HOI detector and propose a re-mining strategy to generate more comprehensive visual representation. Secondly, we design more fine-grained sentence- and word-level alignment and knowledge transfer strategies to effectively address the many-to-many matching problem between multiple interactions and multiple texts. These strategies alleviate the matching confusion problem that arises when multiple interactions occur simultaneously, thereby improving the effectiveness of the alignment process. Finally, HOI reasoning by visual features augmented with textual knowledge substantially improves the understanding of interactions. Experimental results illustrate the effectiveness of our approach, where state-of-the-art performance is achieved on public benchmarks. Yichao Cao, Qingfei Tang, Xiu Su, Shan You, Xiaobo Lu, Chang Xu 0002 |
ICCV | 1 |
| 2023 | Attributes Grouping and Mining Hashing for Fine-Grained Image RetrievalabstractIn recent years, hashing methods have been popular in the large-scale media search for low storage and strong representation capabilities. To describe objects with similar overall appearance but subtle differences, more and more studies focus on hashing-based fine-grained image retrieval. Existing hashing networks usually generate both local and global features through attention guidance on the same deep activation tensor, which limits the diversity of feature representations. To handle this limitation, we substitute convolutional descriptors for attention-guided features and propose an Attributes Grouping and Mining Hashing (AGMH), which groups and embeds the category-specific visual attributes in multiple descriptors to generate a comprehensive feature representation for efficient fine-grained image retrieval. Specifically, an Attention Dispersion Loss (ADL) is designed to force the descriptors to attend to various local regions and capture diverse subtle details. Moreover, we propose a Stepwise Interactive External Attention (SIEA) to mine critical attributes in each descriptor and construct correlations between fine-grained attributes and objects. The attention mechanism is dedicated to learning discrete attributes, which will not cost additional computations in hash codes generation. Finally, the compact binary codes are learned by preserving pairwise similarities. Experimental results demonstrate that AGMH consistently yields the best performance against state-of-the-art methods on fine-grained benchmark datasets. Xin Lu 0007, Shikun Chen, Yichao Cao, Xin Zhou 0030, Xiaobo Lu |
ACM Multimedia | 3 |
| 2023 | Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation ModelsabstractHuman-object interaction (HOI) detection aims to comprehend the intricate relationships between humans and objects, predicting <human, action, object> triplets, and serving as the foundation for numerous computer vision tasks. The complexity and diversity of human-object interactions in the real world, however, pose significant challenges for both annotation and recognition, particularly in recognizing interactions within an open world context. This study explores the universal interaction recognition in an open-world setting through the use of Vision-Language (VL) foundation models and large language models (LLMs). The proposed method is dubbed as UniHOI. We conduct a deep analysis of the three hierarchical features inherent in visual HOI detectors and propose a method for high-level relation extraction aimed at VL foundation models, which we call HO prompt-based learning. Our design includes an HO Prompt-guided Decoder (HOPD), facilitates the association of high-level relation representations in the foundation model with various HO pairs within the image. Furthermore, we utilize a LLM (i.e. GPT) for interaction interpretation, generating a richer linguistic understanding for complex HOIs. For open-category interaction recognition, our method supports either of two input types: interaction phrase or interpretive sentence. Our efficient architecture design and learning methods effectively unleash the potential of the VL foundation models and LLMs, allowing UniHOI to surpass all existing methods with a substantial margin, under both supervised and zero-shot settings. The code and pre-trained weights will be made publicly available. Yichao Cao, Qingfei Tang, Xiu Su, Shan You, Xiaobo Lu, Chang Xu 0002 |
NeurIPS | 1 |
| 2023 | IMDet: Injecting more supervision to CenterNet-like object detection
Shukun Jia, Yichao Cao, Xiaobo Lu |
Expert Syst. Appl. | 3 |
| 2022 | CNN-Transformer Hybrid Architecture for Early Fire Detection
Chenyue Yang, Yixuan Pan, Yichao Cao, Xiaobo Lu |
ICANN (4) | 3 |
| 2022 | Searching for Better Spatio-temporal Alignment in Few-Shot Action RecognitionabstractSpatio-Temporal feature matching and alignment are essential for few-shot action recognition as they determine the coherence and effectiveness of the temporal patterns. Nevertheless, this process could be not reliable, especially when dealing with complex video scenarios. In this paper, we propose to improve the performance of matching and alignment from the end-to-end design of models. Our solution comes at two-folds. First, we encourage to enhance the extracted Spatio-Temporal representations from few-shot videos in the perspective of architectures. With this aim, we propose a specialized transformer search method for videos, thus the spatial and temporal attention can be well-organized and optimized for stronger feature representations. Second, we also design an efficient non-parametric spatio-temporal prototype alignment strategy to better handle the high variability of motion. In particular, a query-specific class prototype will be generated for each query sample and category, which can better match query sequences against all support sequences. By doing so, our method SST enjoys significant superiority over the benchmark UCF101 and HMDB51 datasets. For example, with no pretraining, our method achieves 17.1\% Top-1 accuracy improvement than the baseline TRX on UCF101 5-way 1-shot setting but with only 3x fewer FLOPs. Yichao Cao, Xiu Su, Qingfei Tang, Shan You, Xiaobo Lu, Chang Xu 0002 |
NeurIPS | 1 |
| 2022 | ED-DRAP: Encoder-Decoder Deep Residual Attention Prediction Network for Radar EchoesabstractPrecipitation nowcasting is quite important and fundamental. It underlies various public services ranging from rainstorm warnings to flight safety. In order to further improve the prediction accuracy for the spatiotemporal sequence forecasting problem, we propose an encoder–decoder deep residual attention prediction network, which adaptively rescales the multiscale sequence- and spatial-wise features and achieves very deep trainable residual prediction by integrating global residual learning and local deep residual sequence and spatial attention blocks (RSSABs). Experiments in a real-world radar echo map dataset of South China show that compared with the ingenious PredRNN++, TrajGRU methods, and newly proposed Unet-based methods, our ED-DRAP network performs better on the precipitation nowcasting metrics, as well as occupies small GPU memory. Hongshu Che, Dan Niu, Zengliang Zang, Yichao Cao, Xisong Chen |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | STCNet: spatiotemporal cross network for industrial smoke detection
Yichao Cao, Qingfei Tang, Xiaobo Lu |
Multim. Tools Appl. | 1 |
| 2022 | QuasiVSD: efficient dual-frame smoke detection
Yichao Cao, Qingfei Tang, Shaosheng Xu, Xiaobo Lu |
Neural Comput. Appl. | 1 |
| 2022 | EFFNet: Enhanced Feature Foreground Network for Video Smoke Source Prediction and DetectionabstractSmoke detection in video is a challenging task because of the irregular shape of smoke, its complex motion state, which is affected by temperature, wind and other external factors, and background disturbances. Pixel-based foreground modeling method is a crucial step in many smoke detection systems and can be applied to efficiently focus on a certain object or a specific region to detect movements or anomalies. In video analysis, it is a natural idea to move the focus from the pixel-level foreground to the feature-level foreground. In this paper, the feature foreground is generated by the middle layer of a convolutional neural network (CNN) to guide the temporal modeling process for smoke objects. A novel temporal module called the Feature Foreground Module (FFM) is proposed to boost learning of a smoke temporal representation. Consider the problem of smoke analysis in video, we present a novel unifying approach, named an enhanced feature foreground network (EFFNet), that performs both smoke source prediction and detection. Efficient branch networks are designed in EFFNet, to predict the source mask and bounding boxes of smoke plumes in video. To the best of our knowledge, this is the first paper to study the source of smoke using deep learning methods. Finally, experiments on a realistic smoke dataset and a public dataset show that EFFNet method performs much better than do previous state-of-the-art methods. Yichao Cao, Qingfei Tang, Xuehui Wu, Xiaobo Lu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Combining the Convolution and Transformer for Classification of Smoke-Like Scenes in Remote Sensing ImagesabstractRemote sensing (RS) images are used in a wide range of tasks. In the fire detection field, smoke in RS images is considered as an indicator of wildfires. However, smoke-like scenes, e.g., cloud, in RS images increase the difficulty of smoke recognition. Convolutional neural networks (CNNs) have greatly promoted the development of image processing. CNNs are good at capturing local features; however, their ability to capture global features is relatively weak. Recently, the transformer deep learning model has shown strong potential in vision tasks. The transformer model utilizes self-attention modules to extract global features but may lose local details. Recognition of smoke in RS images depends strongly on the combination of both local and global features. Thus, this article proposes the transformer enhanced convolutional network (TECN) to classify RS smoke-like scenes. The proposed hybrid TECN model exploits the advantages of the CNN and transformer techniques at the same time. In TECN, the feature merge and intelligent aggregation modules are used to promote conversion and aggregation between CNN feature maps and transformer patch embeddings. Experiments are conducted on the USTC_SmokeRS dataset, which is developed for the classification of RS smoke-like scenes. The experimental results demonstrate that the proposed TECN achieves a competitive accuracy of 98.39% on this dataset. Shikun Chen, Yichao Cao, Xiaobo Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Online human action recognition with spatial and temporal skeleton features using a distributed camera networkabstractOnline action recognition is an important task for human-centered intelligent services. However, it remains a highly challenging problem due to the high varieties and uncertainties of spatial and temporal scales of human actions. In this paper, the following core ideas are proposed to deal with the online action recognition problem. First, we combine spatial and temporal skeleton features to represent human actions, which include not only geometrical features, but also multiscale motion features, such that both spatial and temporal information of the actions are covered. We use an efficient one-dimensional convolutional neural network to fuse spatial and temporal features and train them for action recognition. Second, we propose a group sampling method to combine the previous action frames and current action frames, which are based on the hypothesis that the neighboring frames are largely redundant, and the sampling mechanism ensures that the long-term contextual information is also considered. Third, the skeletons from multiview cameras are fused in a distributed manner, which can improve the human pose accuracy in the case of occlusions. Finally, we propose a Restful style based client-server service architecture to deploy the proposed online action recognition module on the remote server as a public service, such that camera networks for online action recognition can benefit from this architecture due to the limited onboard computational resources. We evaluated our model on the data sets of JHMDB and UT-Kinect, which achieved highly promising accuracy levels of 80.1% and 96.9%, respectively. Our online experiments show that our memory group sampling mechanism is far superior to the traditional sliding window. Yichao Cao, Guohui Tian, Ze Ji |
Int. J. Intell. Syst. | 3 |
| 2021 | Global2Salient: Self-adaptive feature aggregation for remote sensing smoke detection
Shikun Chen, Yichao Cao, Xiaoqiang Feng, Xiaobo Lu |
Neurocomputing | 2 |
| 2021 | Patchwise dictionary learning for video forest fire smoke detection in wavelet domain
Xuehui Wu, Yichao Cao, Xiaobo Lu, Henry Leung 0001 |
Neural Comput. Appl. | 2 |
| 2019 | Learning spatial-temporal representation for smoke vehicle detection
Yichao Cao, Xiaobo Lu |
Multim. Tools Appl. | 1 |