VLDB 2026 Research / reviewers in the wild / expert
Bing Cao 0002
dblp:59/4329-2
· DBLP profile ↗
44ranked-venue papers
16as first author
39since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 14 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 8 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reconcile Gradient Modulation for Harmony Multimodal LearningabstractMultimodal learning frequently faces two coupled challenges: modality imbalance, where dominant modalities suppress others during training, and modality conflict, where opposing gradient directions hinder optimization. Existing methods typically address these issues in isolation, yet they are intrinsically correlated and most fundamentally reflected in the gradient space—severe imbalance may obscure conflicts, while suppressing conflict may homogenize features and worsen imbalance, affecting fusion performance. To jointly address this coupled challenge, we propose Reconcile Gradient Modulation (RGM), a unified framework that adaptively adjusts gradient magnitude and direction for harmony multimodal learning. The core of RGM is SynOrth Grad, which minimizes Dirichlet energy to perform minimal-gradient surgery. It enhances cooperation synergy when modalities are aligned and enforces orthogonality to preserve uniqueness in conflict situations, thus promoting stable and balanced learning. To guide this modulation, we propose Cumulative Gradient Energy (CGE) as a convergence-guaranteed measure of modality-wise progress, and construct a Balance-nonConflict Plane (BCP) for real-time diagnosis and control of training dynamics. Experiments on diverse benchmarks validate our effectiveness and generalizability, consistently outperforming counterparts that are designed to handle multimodal imbalance or conflict independently. Xiyuan Gao, Bing Cao 0002, Baoquan Gong, Pengfei Zhu 0001 |
AAAI | 2 |
| 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image FusionabstractImage fusion aims to integrate comprehensive information from images acquired through multiple sources. However, images captured by diverse sensors often encounter various degradations that can negatively affect fusion quality. Traditional fusion methods generally treat image enhancement and fusion as separate processes, overlooking the inherent correlation between them; notably, the dominant regions in one modality of a fused image often indicate areas where the other modality might benefit from enhancement. Inspired by this observation, we introduce the concept of dominant regions for image enhancement and present a Dynamic Relative EnhAnceMent framework for Image Fusion (Dream-IF). This framework quantifies the relative dominance of each modality across different layers and leverages this information to facilitate reciprocal cross-modal enhancement. By integrating the relative dominance derived from image fusion, our approach supports not only image restoration but also a broader range of image enhancement applications. Furthermore, we employ prompt-based encoding to capture degradation-specific details, which dynamically steer the restoration process and promote coordinated enhancement in both multi-modal image fusion and image enhancement scenarios. Extensive experimental results demonstrate that Dream-IF consistently outperforms its counterparts. Xingxin Xu, Bing Cao 0002, Dongdong Li 0004, Qinghua Hu, Pengfei Zhu 0001 |
AAAI | 2 |
| 2026 | Bi-directional Self-Registration for misaligned infrared-visible image fusion
Timing Li, Bing Cao 0002, Qinghua Hu, Pengfei Zhu 0001 |
Pattern Recognit. | 2 |
| 2026 | Hyperbolic Cycle Alignment for Infrared-Visible Image FusionabstractAccurate alignment is a fundamental prerequisite for multi-modal image fusion, yet aligning multi-modal remains challenging due to nonlinear geometric distortions and substantial appearance discrepancies. This work presents the Hyperbolic Cycle Alignment Network (Hy-CycleAlign), a geometry-aware cyclic alignment framework formulated in hyperbolic space. By embedding multi-level representations into a negatively curved manifold, Hy-CycleAlign departs from conventional Euclidean-space paradigms and provides enhanced sensitivity to spatial perturbations, enabling more reliable modeling of cross-modal correspondences. The framework integrates a dual-path cyclic structure that enforces bidirectional deformation consistency and prevents accumulated alignment drift. In addition, a hyperbolic hierarchy contrastive alignment module jointly constrains semantic and structural representations within a unified hyperbolic embedding domain, promoting coherent alignment across global and local geometric scales. From a theoretical standpoint, we derive the sensitivity properties of the Poincaré model and show that its metric inherently amplifies positional variations, thereby strengthening the discriminability of subtle cross-modal misalignments compared with Euclidean geometry. Extensive experiments on diverse misaligned multi-modal datasets show that Hy-CycleAlign achieves strong overall results in alignment accuracy, structural fidelity, and downstream fusion quality. These results validate the effectiveness of hyperbolic geometric modeling for robust multi-modal image alignment. Timing Li, Bing Cao 0002, Jiahe Feng, Haifang Cao, Qinghua Hu, Pengfei Zhu 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Asymmetric Reinforcing Against Multi-Modal Representation BiasabstractThe strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the challenge of dynamic modality contributions, the dominance of different modalities may change with the environments, leading to suboptimal performance in multimodal learning. Current methods mainly enhance weak modalities to balance multimodal representation bias, which inevitably optimizes from a partialmodality perspective, easily leading to performance descending for dominant modalities. To address this problem, we propose an Asymmetric Reinforcing method against Multimodal representation bias (ARM). Our ARM dynamically reinforces the weak modalities while maintaining the ability to represent dominant modalities through conditional mutual information. Moreover, we provide an in-depth analysis that optimizing certain modalities could cause information loss and prevent leveraging the full advantages of multimodal data. By exploring the dominance and narrowing the contribution gaps between modalities, we have significantly improved the performance of multimodal learning, making notable progress in mitigating imbalanced multimodal learning. Xiyuan Gao, Bing Cao 0002, Pengfei Zhu 0001, Nannan Wang 0001, Qinghua Hu |
AAAI | 2 |
| 2025 | Unknown Text Learning for Clip-Based Few-Shot Open-Set Recognition
Qilong Wang 0001, Bing Cao 0002, Qinghua Hu, Yahong Han |
ICCV | 3 |
| 2025 | Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale BenchmarkabstractThe dynamic imbalance of the fore-background is a major challenge in video object counting, which is usually caused by the sparsity of target objects. This remains understudied in existing works and often leads to severe under-/over-prediction errors. To tackle this issue in video object counting, we propose a density-embedded Efficient Masked Autoencoder Counting (E-MAC) framework in this paper. To empower the model’s representation ability on density regression, we develop a new Density-Embedded Masked mOdeling (DEMO) method, which first takes the density map as an auxiliary modality to perform multimodal self-representation learning for image and density map. Although DEMO contributes to effective cross-modal regression guidance, it also brings in redundant background information, making it difficult to focus on the foreground regions. To handle this dilemma, we propose an efficient spatial adaptive masking derived from density maps to boost efficiency. Meanwhile, we employ an optical flow-based temporal collaborative fusion strategy to effectively capture the dynamic variations across frames, aligning features to derive multi-frame density residuals. The counting
accuracy of the current frame is boosted by harnessing the information from adjacent frames. In addition, considering that most existing datasets are limited to human-centric scenarios, we propose a large video bird counting dataset, DroneBird, in natural scenarios for migratory bird protection. Extensive experiments on three crowd datasets and our DroneBird validate our superiority against
the counterparts. The code and dataset are available. Bing Cao 0002, Quanhao Lu, Jiekang Feng, Qilong Wang 0001, Pengfei Zhu 0001, Qinghua Hu |
ICLR | 1 |
| 2025 | Multimodal Negative LearningabstractMultimodal learning systems often encounter challenges related to modality imbalance, where a dominant modality may overshadow others, thereby hindering the learning of weak modalities. Conventional approaches often force weak modalities to align with dominant ones in "Learning to be (the same)" (Positive Learning), which risks suppressing the unique information inherent in the weak modalities. To address this challenge, we offer a new learning paradigm: "Learning Not to be" (Negative Learning). Instead of enhancing weak modalities’ target-class predictions, the dominant modalities dynamically guide the weak modality to suppress non-target classes. This stabilizes the decision space and preserves modality-specific information, allowing weak modalities to preserve unique information without being over-aligned. We proceed to reveal the multimodal learning from a robustness perspective and theoretically derive the Multimodal Negative Learning (MNL) framework, which introduces a dynamic guidance mechanism tailored for negative learning. Our method provably tightens the robustness lower bound of multimodal learning by increasing the Unimodal Confidence Margin (UCoM) and reduces the empirical error of weak modalities, particularly under noisy and imbalanced scenarios. Extensive experiments across multiple benchmarks demonstrate the effectiveness and generalizability of our approach against the competing methods. The code will be available at: https://github.com/BaoquanGong/Multimodal-Negative-Learning.git Baoquan Gong, Xiyuan Gao, Pengfei Zhu 0001, Qinghua Hu, Bing Cao 0002 |
NeurIPS | 5 |
| 2025 | Unknown Support Prototype Set for Open Set Recognition
Guosong Jiang, Pengfei Zhu 0001, Bing Cao 0002, Qinghua Hu |
Int. J. Comput. Vis. | 3 |
| 2025 | Geometry-semantic aware for monocular 3D Semantic Scene Completion
Zonghao Lu, Bing Cao 0002, Shuyin Xia, Qinghua Hu |
Pattern Recognit. | 2 |
| 2025 | RTF: Recursive TransFusion for Multi-Modal Image SynthesisabstractMulti-modal image synthesis is crucial for obtaining complete modalities due to the imaging restrictions in reality. Current methods, primarily CNN-based models, find it challenging to extract global representations because of local inductive bias, leading to synthetic structure deformation or color distortion. Despite the significant global representation ability of transformer in capturing long-range dependencies, its huge parameter size requires considerable training data. Multi-modal synthesis solely based on one of the two structures makes it hard to extract comprehensive information from each modality with limited data. To tackle this dilemma, we propose a simple yet effective Recursive TransFusion (RTF) framework for multi-modal image synthesis. Specifically, we develop a TransFusion unit to integrate local knowledge extracted from the individual modality by connecting a CNN-based local representation block (LRB) and a transformer-based global fusion block (GFB) via a feature translating gate (FTG). Considering the numerous parameters introduced by the transformer, we further unfold a TransFusion unit with recursive constraint repeatedly, forming recursive TransFusion (RTF), which progressively extracts multi-modal information at different depths. Our RTF remarkably reduces network parameters while maintaining superior performance. Extensive experiments validate our superiority against the competing methods on multiple benchmarks. The source code will be available at https://github.com/guoliangq/RTF. Bing Cao 0002, Guoliang Qi, Pengfei Zhu 0001, Qinghua Hu, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | CharacterFactory: Sampling Consistent Characters With GANs for Diffusion ModelsabstractRecent advances in text-to-image models have opened new frontiers in human-centric generation. However, these models cannot be directly employed to generate images with consistent newly coined identities. In this work, we propose CharacterFactory, a framework that allows sampling new characters with consistent identities in the latent space of GANs for diffusion models. More specifically, we consider the word embeddings of celeb names as ground truths for the identity-consistent generation task and train a GAN model to learn the mapping from a latent space to the celeb embedding space. In addition, we design a context-consistent loss to ensure that the generated identity embeddings can produce identity-consistent images in various contexts. Remarkably, the whole model only takes 10 minutes for training, and can sample infinite characters end-to-end during inference. Extensive experiments demonstrate excellent performance of the proposed CharacterFactory on character creation in terms of identity consistency and editability. Furthermore, the generated characters can be seamlessly combined with the off-the-shelf image/video/3D diffusion models. We believe that the proposed CharacterFactory is an important step for identity-consistent character generation. Code and Gradio demo are available at: https://qinghew.github.io/CharacterFactory/. Baolu Li 0001, Xiaomin Li 0001, Bing Cao 0002, Liqian Ma, Huchuan Lu, Xu Jia 0012 |
IEEE Trans. Image Process. | 4 |
| 2025 | RD-OpenMax: Rethinking OpenMax for Robust Realistic Open-Set RecognitionabstractOpen-set recognition (OSR) toward a practical open-world setting has attracted increasing research attention in recent years. However, existing OSR settings are either too idealized or focus on specific scenes such as long-tailed distribution and few-shot samples, which fail to capture the complexity of real-world scenarios. In this article, we propose a realistic OSR (ROSR) setting that covers a diverse range of challenging and real-world scenarios, including fine-grained cases with strong semantic correlation and a large number of species, few-shot samples, long-tailed sample distribution, dynamic inputs (e.g., images, spatio-temporal, and multimodal signals) and cross-domain adaptation. In particular, we rethink the simple and basic OpenMax for the ROSR setting and introduce a novel method, regularized discriminative OpenMax (RD-OpenMax), to handle the challenges in the ROSR setting. RD-OpenMax improves upon the basic OpenMax approach by introducing a covariance attention-based covariance pooling (CACP) module as a global aggregation step before the deep architecture's classifier. This module explores rich statistical information on features and provides discriminative distance scores for OpenMax. To address the instability of extreme value theory (EVT) estimation due to insufficient training samples under few-shot and long-tailed scenarios, we propose a regularized EVT (REVT) method based on Monte Carlo sampling to recalibrate the distribution of distance scores. As such, our RD-OpenMax performs a REVT model of distance scores generated by discriminative CACP representations to distinguish known classes and recognize unknown ones effectively and robustly. Extensive experiments are conducted on more than ten visual benchmarks across several scenarios, and the empirical comparisons show that the ROSR setting challenges existing state-of-the-art OSR approaches. Moreover, our RD-OpenMax clearly outperforms its counterparts under the ROSR setting while performing favorably against state-of-the-arts under the traditional OSR setting. Xiaojie Yin, Bing Cao 0002, Qinghua Hu, Qilong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Bi-directional Adapter for Multimodal TrackingabstractDue to the rapid development of computer vision, single-modal (RGB) object tracking has made significant progress in recent years. Considering the limitation of single imaging sensor, multi-modal images (RGB, infrared, etc.) are introduced to compensate for this deficiency for all-weather object tracking in complex environments. However, as acquiring sufficient multi-modal tracking data is hard while the dominant modality changes with the open environment, most existing techniques fail to extract multi-modal complementary information dynamically, yielding unsatisfactory tracking performance. To handle this problem, we propose a novel multi-modal visual prompt tracking model based on a universal bi-directional adapter, cross-prompting multiple modalities mutually. Our model consists of a universal bi-directional adapter and multiple modality-specific transformer encoder branches with sharing parameters. The encoders extract features of each modality separately by using a frozen, pre-trained foundation model. We develop a simple but effective light feature adapter to transfer modality-specific information from one modality to another, performing visual feature prompt fusion in an adaptive manner. With adding fewer (0.32M) trainable parameters, our model achieves superior tracking performance in comparison with both the full fine-tuning methods and the prompt learning-based methods. Our code is available: https://github.com/SparkTempest/BAT. Bing Cao 0002, Junliang Guo, Pengfei Zhu 0001, Qinghua Hu |
AAAI | 1 |
| 2024 | ID-like Prompt Learning for Few-Shot Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection methods often exploit auxiliary outliers to train model identifying OOD samples, especially discovering challenging outliers from auxiliary outliers dataset to improve OOD detection. However, they may still face limitations in effectively distinguishing between the most challenging OOD samples that are much like in-distribution (ID) data, i.e., ID-like samples. To this end, we propose a novel OOD detection framework that discovers ID-like outliers using CLIP [32]from the vicinity space of the ID samples, thus helping to identify these most challenging OOD samples. Then a prompt learning framework is proposed that utilizes the identified ID-like outliers to further leverage the capabilities of CLIP for OOD detection. Benefiting from the powerful CLIP, we only need a small number of ID samples to learn the prompts of the model without exposing other auxiliary outlier datasets. By focusing on the most challenging ID-like OOD samples and elegantly exploiting the capabilities of CLIP, our method achieves superior few-shot learning performance on various real-world image datasets (e.g., in 4-shot OOD detection on the ImageNet-1k dataset, our method reduces the average FPR95 by 12.16% and improves the average AUROC by 2.76%, compared to state-of-the-art methods). Code is available at https://github.com/ycfate/ID-like Yichen Bai, Zongbo Han, Bing Cao 0002, Xiaoheng Jiang, Qinghua Hu, Changqing Zhang 0002 |
CVPR | 3 |
| 2024 | Task-Customized Mixture of Adapters for General Image FusionabstractGeneral image fusion aims at integrating important in-formation from multi-source images. However, due to the significant cross-task gap, the respective fusion mechanism varies considerably in practice, resulting in limited performance across subtasks. To handle this problem, we pro-pose a novel task-customized mixture of adapters (TC-MoA) for general image fusion, adaptively prompting various fusion tasks in a unified model. We borrow the insight from the mixture of experts (MoE), taking the experts as effi-cient tuning adapters to prompt a pre-trained foundation model. These adapters are shared across different tasks and constrained by mutual information regularization, ensuring compatibility with different tasks while complementarity for multi-source images. The task-specific routing networks customize these adapters to extract task-specific information from different sources with dynamic dominant inten-sity, performing adaptive visual feature prompt fusion. No-tably, our TC-MoA controls the dominant intensity bias for different fusion tasks, successfully unifying multiple fusion tasks in a single model. Extensive experiments show that TC-MoA outperforms the competing approaches in learning commonalities while retaining compatibility for gen-eral image fusion (multi-modal, multi-exposure, and multi-focus), and also demonstrating striking controllability on more generalization experiments. The code is available at https://github.com/YangSun22/TC-MoA. Pengfei Zhu 0001, Bing Cao 0002, Qinghua Hu |
CVPR | 3 |
| 2024 | Visible and Clear: Finding Tiny Objects in Difference Map
Bing Cao 0002, Haiyu Yao, Pengfei Zhu 0001, Qinghua Hu |
ECCV (17) | 1 |
| 2024 | Predictive Dynamic FusionabstractMultimodal fusion is crucial in joint decision-making systems for rendering holistic judgments. Since multimodal data changes in open environments, dynamic fusion has emerged and achieved remarkable progress in numerous applications. However, most existing dynamic multimodal fusion methods lack theoretical guarantees and easily fall into suboptimal problems, yielding unreliability and instability. To address this issue, we propose a Predictive Dynamic Fusion (PDF) framework for multimodal learning. We proceed to reveal the multimodal fusion from a generalization perspective and theoretically derive the predictable Collaborative Belief (Co-Belief) with Mono- and Holo-Confidence, which provably reduces the upper bound of generalization error. Accordingly, we further propose a relative calibration strategy to calibrate the predicted Co-Belief for potential uncertainty. Extensive experiments on multiple benchmarks confirm our superiority. Our code is available at https://github.com/Yinan-Xia/PDF. Bing Cao 0002, Yinan Xia, Changqing Zhang 0002, Qinghua Hu |
ICML | 1 |
| 2024 | Dynamic Brightness Adaptation for Robust Multi-modal Image Fusion
Yiming Sun 0003, Bing Cao 0002, Pengfei Zhu 0001, Qinghua Hu |
IJCAI | 2 |
| 2024 | Exploring What to Share for Multimodal Image SynthesisabstractMultimodal image synthesis has emerged as an effective solution to the challenge of modality missing. Most existing methods extract specific features from each modality, which suffers from insufficient common representation across modalities. A fundamental solution to capture multimodal common features is to adopt a shared feature extractor. However, it struggles to preserve the modality-specific features. To handle this dilemma, we explore a dynamic "what to share" mechanism to extract modality-common features while simultaneously preserving modality-specific information. Specifically, we propose a shared-specific feature modeling (SSFM) framework for multimodal image synthesis. We first develop a parameter selective sharing (PSS) mechanism to determine which layers in the feature extractor should share parameters across modalities. The PSS computes the gradient-based modality conflict score (MC-Score) for each layer, separating the high-conflict layers while sharing the low-conflict layers for individual modality. Additionally, we propose a multimodal interactive attention (MIA) module to adaptively fuse multimodal complementary information from each modality. Experimental results show that our approach outperforms the state-of-the-art methods in quantitative and qualitative comparisons. Guoliang Qi, Bing Cao 0002, Qinghua Hu |
IJCNN | 2 |
| 2024 | Conditional Controllable Image FusionabstractImage fusion aims to integrate complementary information from multiple input images acquired through various sources to synthesize a new fused image. Existing methods usually employ distinct constraint designs tailored to specific scenes, forming fixed fusion paradigms. However, this data-driven fusion approach is challenging to deploy in varying scenarios, especially in rapidly changing environments. To address this issue, we propose a conditional controllable fusion (CCF) framework for general image fusion tasks without specific training. Due to the dynamic differences of different samples, our CCF employs specific fusion constraints for each individual in practice. Given the powerful generative capabilities of the denoising diffusion model, we first inject the specific constraints into the pre-trained DDPM as adaptive fusion conditions. The appropriate conditions are dynamically selected to ensure the fusion process remains responsive to the specific requirements in each reverse diffusion stage. Thus, CCF enables conditionally calibrating the fused images step by step. Extensive experiments validate our effectiveness in general fusion tasks across diverse scenarios against the competing methods without additional training. The code is publicly available. Bing Cao 0002, Xingxin Xu, Pengfei Zhu 0001, Qilong Wang 0001, Qinghua Hu |
NeurIPS | 1 |
| 2024 | Test-Time Dynamic Image FusionabstractThe inherent challenge of image fusion lies in capturing the correlation of multi-source images and comprehensively integrating effective information from different sources. Most existing techniques fail to perform dynamic image fusion while notably lacking theoretical guarantees, leading to potential deployment risks in this field. Is it possible to conduct dynamic image fusion with a clear theoretical justification? In this paper, we give our solution from a generalization perspective. We proceed to reveal the generalized form of image fusion and derive a new test-time dynamic image fusion paradigm. It provably reduces the upper bound of generalization error. Specifically, we decompose the fused image into multiple components corresponding to its source data. The decomposed components represent the effective information from the source data, thus the gap between them reflects the \textit{Relative Dominability} (RD) of the uni-source data in constructing the fusion image. Theoretically, we prove that the key to reducing generalization error hinges on the negative correlation between the RD-based fusion weight and the uni-source reconstruction loss. Intuitively, RD dynamically highlights the dominant regions of each source and can be naturally converted to the corresponding fusion weight, achieving robust results. Extensive experiments and discussions with in-depth analysis on multiple benchmarks confirm our findings and superiority. Our code is available at https://github.com/Yinan-Xia/TTD. Bing Cao 0002, Yinan Xia, Changqing Zhang 0002, Qinghua Hu |
NeurIPS | 1 |
| 2024 | Deep Attention-Guided Spatial-Spectral Network for Hyperspectral Image UnmixingabstractDeep-learning (DL)-based methods have been increasingly used in hyperspectral unmixing (HU), especially the recent trend of unsupervised autoencoder (AE) networks, which have achieved excellent performances. Although some existing unmixing methods take spatial information into account, the utilization of spatial structure is not sufficient and effective. In this letter, we present a deep attention-guided spatial–spectral network for hyperspectral image (HSI) unmixing called DASS-Net, which adopts a parallel dual-stream structure. We design a neighborhood spatial attention (NSA) module, where the abundance features of the central pixel are dynamically weighted by the coarse-grained features of the neighborhood pixels. In addition, a dual-gated mechanism is introduced to further integrate and express the spatial and spectral information. Experimental results show that the proposed DASS-Net performs particularly well in endmember extraction and outperforms all compared methods. Lin Qi 0004, Mengyi Yue, Feng Gao 0005, Bing Cao 0002, Junyu Dong, Xinbo Gao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | LiDAR-Camera Continuous Fusion in Voxelized Grid for Semantic Scene CompletionabstractSemantic Scene Completion (SSC) requires a comprehensive perception of both the geometry and semantics across the entire 3D scene. In the domain of autonomous driving, the majority of existing SSC methods rely on single-modal images (e.g., MonoScene, TPVformer) or point clouds (e.g., S3CNet, JS3C-Net), without taking into account the complementary information from bimodal sources. In this work, we propose an Image and Point Cloud continuous fusion in Voxel Network (IPVoxelNet) to address SSC within the voxelized space. IPVoxelNet represents images and point clouds within a unified voxelized space and utilizes the Image and Point Cloud Fusion (IPF) layers for continuous fusion of bimodal features. Specifically, IPVoxelNet utilizes pixel-to-voxel reprojection to map pixels into 3D space, leveraging the dense semantics of images. Unordered point clouds are represented in voxel space through regularization. IPVoxelNet independently learns the geometry and semantics of each modality. Additionally, we propose cross-modal knowledge distillation to transfer geometric information from point clouds to images. We validate our model on the challenging SemanticKITTI and nuScenes-Occupancy datasets, achieving state-of-the-art results across multiple classes. IPVoxelNet demonstrates competitive performance in both geometry (SC IoU) and semantics (mIoU). Zonghao Lu, Bing Cao 0002, Qinghua Hu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Autoencoder-Based Collaborative Attention GAN for Multi-Modal Image SynthesisabstractMulti-modal images are required in a wide range of practical scenarios, from clinical diagnosis to public security. However, certain modalities may be incomplete or unavailable because of the restricted imaging conditions, which commonly leads to decision bias in many real-world applications. Despite the significant advancement of existing image synthesis techniques, learning complementary information from multi-modal inputs remains challenging. To address this problem, we propose an autoencoder-based collaborative attention generative adversarial network (ACA-GAN) that uses available multi-modal images to generate the missing ones. The collaborative attention mechanism deploys a single-modal attention module and a multi-modal attention module to effectively extract complementary information from multiple available modalities. Considering the significant modal gap, we further developed an autoencoder network to extract the self-representation of target modality, guiding the generative model to fuse target-specific information from multiple modalities. This considerably improves cross-modal consistency with the desired modality, thereby greatly enhancing the image synthesis performance. Quantitative and qualitative comparisons for various multi-modal image synthesis tasks highlight the superiority of our approach over several prior methods by demonstrating more precise and realistic results. Bing Cao 0002, Haifang Cao, Pengfei Zhu 0001, Changqing Zhang 0002, Qinghua Hu |
IEEE Trans. Multim. | 1 |
| 2024 | Multi-View Knowledge Ensemble With Frequency Consistency for Cross-Domain Face TranslationabstractCross-domain face translation aims to transfer face images from one domain to another. It can be widely used in practical applications, such as photos/sketches in law enforcement, photos/drawings in digital entertainment, and near-infrared (NIR)/visible (VIS) images in security access control. Restricted by limited cross-domain face image pairs, the existing methods usually yield structural deformation or identity ambiguity, which leads to poor perceptual appearance. To address this challenge, we propose a multi-view knowledge (structural knowledge and identity knowledge) ensemble framework with frequency consistency (MvKE-FC) for cross-domain face translation. Due to the structural consistency of facial components, the multi-view knowledge learned from large-scale data can be appropriately transferred to limited cross-domain image pairs and significantly improve the generative performance. To better fuse multi-view knowledge, we further design an attention-based knowledge aggregation module that integrates useful information, and we also develop a frequency-consistent (FC) loss that constrains the generated images in the frequency domain. The designed FC loss consists of a multidirection Prewitt (mPrewitt) loss for high-frequency consistency and a Gaussian blur loss for low-frequency consistency. Furthermore, our FC loss can be flexibly applied to other generative models to enhance their overall performance. Extensive experiments on multiple cross-domain face datasets demonstrate the superiority of our method over state-of-the-art methods both qualitatively and quantitatively. Bing Cao 0002, Pengfei Zhu 0001, Qinghua Hu, Dongwei Ren, Wangmeng Zuo, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Multi-Task Credible Pseudo-Label Learning for Semi-Supervised Crowd CountingabstractAs a widely used semi-supervised learning strategy, self-training generates pseudo-labels to alleviate the labor-intensive and time-consuming annotation problems in crowd counting while boosting the model performance with limited labeled data and massive unlabeled data. However, the noise in the pseudo-labels of the density maps greatly hinders the performance of semi-supervised crowd counting. Although auxiliary tasks, e.g., binary segmentation, are utilized to help improve the feature representation learning ability, they are isolated from the main task, i.e., density map regression and the multi-task relationships are totally ignored. To address the above issues, we develop a multi-task credible pseudo-label learning (MTCP) framework for crowd counting, consisting of three multi-task branches, i.e., density regression as the main task, and binary segmentation and confidence prediction as the auxiliary tasks. Multi-task learning is conducted on the labeled data by sharing the same feature extractor for all three tasks and taking multi-task relations into account. To reduce epistemic uncertainty, the labeled data are further expanded, by trimming the labeled data according to the predicted confidence map for low-confidence regions, which can be regarded as an effective data augmentation strategy. For unlabeled data, compared with the existing works that only use the pseudo-labels of binary segmentation, we generate credible pseudo-labels of density maps directly, which can reduce the noise in pseudo-labels and therefore decrease aleatoric uncertainty. Extensive comparisons on four crowd-counting datasets demonstrate the superiority of our proposed model over the competing methods. The code is available at: https://github.com/ljq2000/MTCP. Pengfei Zhu 0001, Jingqing Li, Bing Cao 0002, Qinghua Hu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image FusionabstractInfrared and visible image fusion aims to integrate comprehensive information from multiple sources to achieve superior performances on various practical tasks, such as detection, over that of a single modality. However, most existing methods directly combined the texture details and object contrast of different modalities, ignoring the dynamic changes in reality, which diminishes the visible texture in good lighting conditions and the infrared contrast in low lighting conditions. To fill this gap, we propose a dynamic image fusion framework with a multi-modal gated mixture of local-to-global experts, termed MoE-Fusion, to dynamically extract effective and comprehensive information from the respective modalities. Our model consists of a Mixture of Local Experts (MoLE) and a Mixture of Global Experts (MoGE) guided by a multi-modal gate. The MoLE performs specialized learning of multi-modal local features, prompting the fused images to retain the local information in a sample-adaptive manner, while the MoGE focuses on the global information that complements the fused image with overall texture detail and contrast. Extensive experiments show that our MoE-Fusion outperforms state-of-the-art methods in preserving multi-modal image texture and contrast through the local-to-global dynamic learning paradigm, and also achieves superior performance on detection tasks. Our code is available: https://github.com/SunYM2020/MoE-Fusion. Bing Cao 0002, Yiming Sun 0003, Pengfei Zhu 0001, Qinghua Hu |
ICCV | 1 |
| 2023 | AutoEncoder-Driven Multimodal Collaborative Learning for Medical Image Synthesis
Bing Cao 0002, Zhiwei Bi, Qinghua Hu, Han Zhang 0002, Nannan Wang 0001, Xinbo Gao 0001, Dinggang Shen |
Int. J. Comput. Vis. | 1 |
| 2023 | Cross-Drone Transformer Network for Robust Single Object TrackingabstractDrones have been widely used in a variety of applications, e.g., aerial photography and military security, because of their high maneuverability and broad views compared with fixed cameras. Multi-drone tracking systems can provide rich information about targets by collecting complementary video clips from different views, especially when targets are occluded or disappear in some views. However, it is challenging to handle cross-drone information interaction and multi-drone information fusion in multi-drone visual tracking. Recently, Transformer has shown significant advantages in automatically modeling the correlation between templates and search regions for visual tracking. To leverage its potential in multi-drone tracking, we propose a novel cross-drone Transformer network (TransMDOT) for visual object tracking tasks. The self-attention mechanism is used to automatically capture the correlation between multiple templates and the corresponding search region to achieve multi-drone feature fusion. During the tracking process, a cross-drone mapping mechanism is proposed by using the surrounding information of the drone with promising tracking status as reference, assisting drones that lost targets to re-calibrate, which implements real-time cross-drone information interaction. As the existing multi-drone evaluation metrics only consider spatial information while ignore temporal information, we further present a system perception index (SPFI) that combines both temporal and spatial information to evaluate the tracking status of multiple drones. Experiments on the MDOT dataset prove that TransMDOT greatly surpasses the state-of-the-art methods in both single-drone performance and multi-drone system fusion performance. Our code will be available onhttps://github.com/cgjacklin/transmdot. Pengfei Zhu 0001, Bing Cao 0002, Xing Wang 0002, Qinghua Hu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Semantic-shape Adaptive Feature Modulation for Semantic Image SynthesisabstractRecent years have witnessed substantial progress in se-mantic image synthesis, it is still challenging in synthesizing photo-realistic images with rich details. Most previ-ous methods focus on exploiting the given semantic map, which just captures an object-level layout for an image. Obviously, a fine-grained part-level semantic layout will benefit object details generation, and it can be roughly in-ferred from an object's shape. In order to exploit the part-level layouts, we propose a Shape-aware Position Descrip-tor (SPD) to describe each pixel's positional feature, where object shape is explicitly encoded into the SP D feature. Fur-thermore, a Semantic-shape Adaptive Feature Modulation (SAFM) block is proposed to combine the given semantic map and our positional features to produce adaptively mod-ulated features. Extensive experiments demonstrate that the proposed SPD and SAFM significantly improve the gener-ation of objects with rich details. Moreover, our method performs favorably against the SOTA methods in terms of quantitative and qualitative evaluation. The source code and model are available at SAFM. Zhengyao Lv, Xiaoming Li 0002, Zhenxing Niu, Bing Cao 0002, Wangmeng Zuo |
CVPR | 4 |
| 2022 | DetFusion: A Detection-driven Infrared and Visible Image Fusion NetworkabstractInfrared and visible image fusion aims to utilize the complementary information between the two modalities to synthesize a new image containing richer information. Most existing works have focused on how to better fuse the pixel-level details from both modalities in terms of contrast and texture, yet ignoring the fact that the significance of image fusion is to better serve downstream tasks. For object detection tasks, object-related information in images is often more valuable than focusing on the pixel-level details of images alone. To fill this gap, we propose a detection-driven infrared and visible image fusion network, termed DetFusion, which utilizes object-related information learned in the object detection networks to guide multimodal image fusion. We cascade the image fusion network with the detection networks of both modalities and use the detection loss of the fused images to provide guidance on task-related information for the optimization of the image fusion network. Considering that the object locations provide a priori information for image fusion, we propose an object-aware content loss that motivates the fusion model to better learn the pixel-level information in infrared and visible images. Moreover, we design a shared attention module to motivate the fusion network to learn object-specific information from the object detection networks. Extensive experiments show that our DetFusion outperforms state-of-the-art methods in maintaining pixel intensity distribution and preserving texture details. More notably, the performance comparison with state-of-the-art image fusion methods in task-driven evaluation also demonstrates the superiority of the proposed method. Our code will be available: https://github.com/SunYM2020/DetFusion. Yiming Sun 0003, Bing Cao 0002, Pengfei Zhu 0001, Qinghua Hu |
ACM Multimedia | 2 |
| 2022 | Learning Self-supervised Low-Rank Network for Single-Stage Weakly and Semi-supervised Semantic Segmentation
Junwen Pan, Pengfei Zhu 0001, Kaihua Zhang 0001, Bing Cao 0002, Yu Wang 0106, Dingwen Zhang, Junwei Han 0001, Qinghua Hu |
Int. J. Comput. Vis. | 4 |
| 2022 | Common feature learning for brain tumor MRI synthesis by context-aware generative adversarial network
Pu Huang 0001, Dengwang Li, Zhicheng Jiao, Dongming Wei, Bing Cao 0002, Zhanhao Mo, Qian Wang 0001, Han Zhang 0002, Dinggang Shen |
Medical Image Anal. | 5 |
| 2022 | Face photo-sketch synthesis via full-scale identity supervision
Bing Cao 0002, Nannan Wang 0001, Jie Li 0001, Qinghua Hu, Xinbo Gao 0001 |
Pattern Recognit. | 1 |
| 2022 | Curiosity-Driven Class-Incremental Learning via Adaptive Sample SelectionabstractModern artificial intelligence systems require class-incremental learning while suffering from catastrophic forgetting in many real-world applications. Due to the missing knowledge of past data, performance substantially degrades. Recent methods often used knowledge distillation and bias correction to avoid catastrophic forgetting caused by cognitive bias. However, since these methods mainly learn all samples indiscriminately, the model is hard to learn what it truly needs from the data stream to balance the new and old knowledge, leading to inevitable forgetting. Instead of considering each sample indiscriminately, the model should learn from its curious samples automatically. To tackle this problem, we propose a curiosity-driven class-incremental learning approach via adaptive sample selection for learning a more generalized model with fewer ineffective updates. Specifically, our method quantifies the model’s curiosity in each sample by two properties: uncertainty and novelty. Our model learns informative samples selectively during training utilizing the proposed uncertainty property, which benefits the classification decision boundary. In light of the imbalanced data, a novelty property is used to selectively optimize the model by employing dissimilar samples, endowing it with more robustness and less cognitive bias. Our method successfully reduces catastrophic forgetting and can be flexibly incorporated with other techniques. Extensive experiments and in-depth analysis on the CIFAR-100, Tiny-ImageNet and Caltech-101 datasets show that our approach outperforms competing methods for class-incremental learning in terms of preventing catastrophic forgetting. Qinghua Hu, Yucong Gao, Bing Cao 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Drone-Based RGB-Infrared Cross-Modality Vehicle Detection Via Uncertainty-Aware LearningabstractDrone-based vehicle detection aims at detecting vehicle locations and categories in aerial images. It empowers smart city traffic management and disaster relief. Researchers have made a great deal of effort in this area and achieved considerable progress. However, because of the paucity of data under extreme conditions, drone-based vehicle detection remains a challenge when objects are difficult to distinguish, particularly in low-light conditions. To fill this gap, we constructed a large-scale drone-based RGB-infrared vehicle detection dataset called DroneVehicle, which contains 28, 439 RGB-infrared image pairs covering urban roads, residential areas, parking lots, and other scenarios from day to night. Cross-modal images provide complementary information for vehicle detection, but also introduce redundant information. To handle this dilemma, we further propose an uncertainty-aware cross-modality vehicle detection (UA-CMDet) framework to improve detection performance in complex environments. Specifically, we design an uncertainty-aware module using cross-modal intersection over union and illumination estimation to quantify the uncertainty of each object. Our method takes uncertainty as a weight to boost model learning more effectively while reducing bias caused by high-uncertainty objects. For more robust cross-modal integration, we further perform illumination-aware non-maximum suppression during inference. Extensive experiments on our DroneVehicle and two challenging RGB-infrared object detection datasets demonstrated the advanced flexibility and superior performance of UA-CMDet over competing methods. Our code and DroneVehicle will be available:https://github.com/VisDrone/DroneVehicle. Yiming Sun 0003, Bing Cao 0002, Pengfei Zhu 0001, Qinghua Hu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Learning a Prototype Discriminator With RBF for Multimodal Image SynthesisabstractMultimodal image synthesis has emerged as a viable solution to the modality missing challenge. Most existing approaches employ softmax-based classifiers to provide modal constraints for the generated models. These methods, however, focus on learning to distinguish inter-domain differences while failing to build intra-domain compactness, resulting in inferior synthetic results. To provide sufficient domain-specific constraint, we hereby introduce a novel prototype discriminator for generative adversarial network (PT-GAN) to effectively estimate the missing or noisy modalities. Different from most previous works, we introduce the Radial Basis Function (RBF) network, endowing the discriminator with domain-specific prototypes, to improve the optimization of generative model. Since the prototype learning extracts more discriminative representation of each domain, and emphasizes intra-domain compactness, it reduces the sensitivity of discriminator to pixel changes in generated images. To address this dilemma, we further propose a reconstructive regularization term which connects the discriminator with the generator, thus enhancing its pixel detectability. To this end, the proposed PT-GAN provides not only consistent domain-specific constraints, but also reasonable uncertainty estimation of generated images with the RBF distance. Experimental results show that our method outperforms the state-of-the-art techniques. The source code will be available at: https://github.com/zhiweibi/PT-GAN. Zhiwei Bi, Bing Cao 0002, Wangmeng Zuo, Qinghua Hu |
IEEE Trans. Image Process. | 2 |
| 2021 | Deep Bayesian Hashing With Center Prior for Multi-Modal Neuroimage RetrievalabstractMulti-modal neuroimage retrieval has greatly facilitated the efficiency and accuracy of decision making in clinical practice by providing physicians with previous cases (with visually similar neuroimages) and corresponding treatment records. However, existing methods for image retrieval usually fail when applied directly to multi-modal neuroimage databases, since neuroimages generally have smaller inter-class variation and larger inter-modal discrepancy compared to natural images. To this end, we propose a deep Bayesian hash learning framework, called CenterHash, which can map multi-modal data into a shared Hamming space and learn discriminative hash codes from imbalanced multi-modal neuroimages. The key idea to tackle the small inter-class variation and large inter-modal discrepancy is to learn a common center representation for similar neuroimages from different modalities and encourage hash codes to be explicitly close to their corresponding center representations. Specifically, we measure the similarity between hash codes and their corresponding center representations and treat it as a center prior in the proposed Bayesian learning framework. A weighted contrastive likelihood loss function is also developed to facilitate hash learning from imbalanced neuroimage pairs. Comprehensive empirical evidence shows that our method can generate effective hash codes and yield state-of-the-art performance in cross-modal retrieval on three multi-modal neuroimage datasets. Erkun Yang, Mingxia Liu 0001, Dongren Yao, Bing Cao 0002, Chunfeng Lian, Pew-Thian Yap, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Auto-GAN: Self-Supervised Collaborative Learning for Medical Image SynthesisabstractIn various clinical scenarios, medical image is crucial in disease diagnosis and treatment. Different modalities of medical images provide complementary information and jointly helps doctors to make accurate clinical decision. However, due to clinical and practical restrictions, certain imaging modalities may be unavailable nor complete. To impute missing data with adequate clinical accuracy, here we propose a framework called self-supervised collaborative learning to synthesize missing modality for medical images. The proposed method comprehensively utilize all available information correlated to the target modality from multi-source-modality images to generate any missing modality in a single model. Different from the existing methods, we introduce an auto-encoder network as a novel, self-supervised constraint, which provides target-modality-specific information to guide generator training. In addition, we design a modality mask vector as the target modality label. With experiments on multiple medical image databases, we demonstrate a great generalization ability as well as specialty of our method compared with other state-of-the-arts. Bing Cao 0002, Han Zhang 0002, Nannan Wang 0001, Xinbo Gao 0001, Dinggang Shen |
AAAI | 1 |
| 2020 | Deep Disentangled Hashing with Momentum Triplets for Neuroimage Search
Erkun Yang, Dongren Yao, Bing Cao 0002, Pew-Thian Yap, Dinggang Shen, Mingxia Liu 0001 |
MICCAI (1) | 3 |
| 2019 | Multi-Margin based Decorrelation Learning for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) refers to matching face images acquired from different domains with wide applications in security scenarios. However, HFR is still a challenging problem due to the significant cross-domain discrepancy and the lacking of sufficient training data in different domains. This paper presents a deep neural network approach namely Multi-Margin based Decorrelation Learning (MMDL) to extract decorrelation representations in a hyperspherical space for cross-domain face images. The proposed framework can be divided into two components: heterogeneous representation network and decorrelation representation learning. First, we employ a large scale of accessible visual face images to train heterogeneous representation network. The decorrelation layer projects the output of the first component into decorrelation latent subspace and obtain decorrelation representation. In addition, we design a multi-margin loss (MML), which consists of tetradmargin loss (TML) and heterogeneous angular margin loss (HAML), to constrain the proposed framework. Experimental results on two challenging heterogeneous face databases show that our approach achieves superior performance on both verification and recognition tasks, comparing with state-of-the-art methods. Bing Cao 0002, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Zhifeng Li 0001 |
IJCAI | 1 |
| 2019 | Data Augmentation-Based Joint Learning for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) is the process of matching face images captured from different sources. HFR plays an important role in security scenarios. However, HFR remains a challenging problem due to the considerable discrepancies (i.e., shape, style, and color) between cross-modality images. Conventional HFR methods utilize only the information involved in heterogeneous face images, which is not effective because of the substantial differences between heterogeneous face images. To better address this issue, this paper proposes a data augmentation-based joint learning (DA-JL) approach. The proposed method mutually transforms the cross-modality differences by incorporating synthesized images into the learning process. The aggregated data augments the intraclass scale, which provides more discriminative information. However, this method also reduces the interclass diversity (i.e., discriminative information). We develop the DA-JL model to balance this dilemma. Finally, we obtain the similarity score between heterogeneous face image pairs through the log-likelihood ratio. Extensive experiments on a viewed sketch database, forensic sketch database, near-infrared image database, thermal-infrared image database, low-resolution photo database, and image with occlusion database illustrate that the proposed method achieves superior performance in comparison with the state-of-the-art methods. Bing Cao 0002, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Asymmetric Joint Learning for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) refers to matching a probe face image taken from one modality to face images acquired from another modality. It plays an important role in security scenarios. However, HFR is still a challenging problem due to great discrepancies between cross-modality images. This paper proposes an asymmetric joint learning (AJL) approach to handle this issue. The proposed method transforms the cross-modality differences mutually by incorporating the synthesized images into the learning process which provides more discriminative information. Although the aggregated data would augment the scale of intra-classes, it also reduces the diversity (i.e. discriminative information) for inter-classes. Then, we develop the AJL model to balance this dilemma. Finally, we could obtain the similarity score between two heterogeneous face images through the log-likelihood ratio. Extensive experiments on viewed sketch database, forensic sketch database and near infrared image database illustrate that the proposed AJL-HFR method achieve superior performance in comparison to state-of-the-art methods. Bing Cao 0002, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001 |
AAAI | 1 |