Cong Bai

dblp:02/11432 · DBLP profile ↗
← Back
101ranked-venue papers
11as first author
81since 2021 · last 2026
0000-0002-6177-3862ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 60 · 7 first-author · 44 since 2021Artificial intelligence and machine learning · 34 · 4 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 15 since 2021Databases, data management, data science and information retrieval · 13 · 2 first-author · 9 since 2021Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text Retrieval
abstract
Remote sensing (RS) image–text retrieval faces significant challenges in real-world datasets due to the presence of Pseudo-Matched Pairs (PMPs), semantically mismatched or weakly aligned image–text pairs, which hinder the learning of reliable cross-modal alignments. To address this issue, we propose a novel retrieval framework that leverages Cross-Modal Gated Attention and a Positive–Negative Awareness Attention mechanism to mitigate the impact of such noisy associations. The gated module dynamically regulates cross-modal information flow, while the awareness mechanism explicitly distinguishes informative (positive) cues from misleading (negative) ones during alignment learning. Extensive experiments on three benchmark RS datasets, i.e., RSICD, RSITMD, and RS5M, demonstrate that our method consistently achieves state-of-the-art performance, highlighting its robustness and effectiveness in handling real-world mismatches and PMPs in RS image–text retrieval tasks.
Pengxiang Ouyang, Zheng Wang 0059, Cong Bai
AAAI4
2026 KFTD: Koopman-Fourier Time-Differentiable Network for Continuous Ocean Spatiotemporal Forecasting
abstract
Accurate oceanic forecasting is critical for climate monitoring and disaster early-warning. However, ocean spatiotemporal forecasting encounters the double challenges of modeling complex dynamical systems and ensuring computational efficiency. We present Koopman–Fourier Time-Differentiable (KFTD) Network, a time-continuous two-stage paradigm that decouples interpolation from prediction to achieve efficient and scalable spatiotemporal modeling. We map complex nonlinear dynamics into the Koopman linear space and exploit Fourier analysis to enable continuous-time interpolation at arbitrary sub-steps. A lightweight residual network consumes the high-fidelity intermediate states to yield the final forecast. Unlike diffusion models, KFTD eliminates multi-step noise sampling and directly evolves the system in continuous time, yielding a 4× computational speed-up. We further introduce a D-PP Loss that supports arbitrary PDE constraints in an end-to-end manner, breaking the physical-consistency bottleneck of pure data-driven approaches. Empirical results on four ocean datasets confirm that our continuous-time framework reduces MSE by an average of 5.6% (up to 12.7% for SST) and improves efficiency over MCVD by 76.25%.
Qinghui Chen, Hailong Liu 0007, Jinglin Zhang 0001, Cong Bai
KDD (1)5
2026 KAN-FIF: Spline-Parameterized Lightweight Physics-based Tropical Cyclone Estimation on Meteorological Satellite
Jiakang Shen, Qinghui Chen, Runtong Wang, Chenrui Xu, Jinglin Zhang 0001, Cong Bai, Feng Zhang 0041
KDD (1)6
2026 Short-Length Hashing via Bit-Level Semantic Representation for Image-Text Retrieval
abstract
The explosive growth of multi-modal data in the era of big data has significantly heightened the urgent need for efficient cross-modal retrieval methods. While traditional real-valued retrieval approaches struggle with high storage costs and slow query speeds, hashing techniques have emerged as a promising solution by mapping high-dimensional data into compact binary codes. Among various hashing paradigms, short-length hashing offers superior advantages in terms of retrieval speed and storage efficiency, making it particularly suitable for resource-constrained edge devices and large-scale real-time applications. However, existing short-length hashing methods typically suffer from weak classification boundaries and significant information loss due to the extremely limited capacity of the hash bits. Most state-of-the-art methods treat the hash code as a holistic vector, failing to maximize the distinctiveness of individual bits. To tackle these challenges effectively, this paper proposes a novel method termed Bit-Level Semantic Representation Hashing (BLSRH). First, by establishing a Bit Semantic Learning Network (BSLN), we enhance the representational and discriminative capabilities of each bit in short-length hash codes independently. Additionally, a contrastive learning mechanism is introduced between bits to improve semantic consistency across modalities and reduce semantic redundancy among bits. Furthermore, to preserve the manifold structure of the original data, an adapted global similarity-preserving method is designed. Finally, a modal alignment loss based on soft-constraint is proposed to bridge the heterogeneity gap and reduce quantization errors, which replaces strict discrete constraints with flexible symbolic constraints. Comprehensive experimental results on three benchmark datasets demonstrate that BLSRH significantly outperforms state-of-the-art baselines, including recent transformer-based approaches, particularly in low-bit scenarios.
Siyang Zhang, Cong Bai
ICMR4
2026 3D-MolGL: A multimodal framework for integrating 3D molecular graphs into language models
Huizhi Li, Dagang Li 0001, Jinglin Zhang 0001, Yuhui Zheng, Cong Bai
Expert Syst. Appl.5
2026 A Novel Dataset and Lightweight Distillation Baseline for Highlight Transparent Object Detection
Gang Li 0005, Qinghui Chen, Qunshu Zhang, Jin Wan, Maomao Xiong, Cong Bai, Dagang Li 0001, Wenyin Zhang, Jinglin Zhang 0004, Shengyong Chen
Int. J. Comput. Vis.8
2026 GraphVSum:graph guided multimodal video summarization
Zhengqi Zhao, Cong Bai, Pengyi Hao
Multim. Syst.2
2026 Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks, Challenges and Baselines
abstract
Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detection remains challenging due to two fundamental limitations: (i) the scarcity of large-scale industrial datasets that cover diverse defect categories across multiple domains, and (ii) the reliance on manual prompts (points, boxes, masks) that introduce subjective noise and lack text-visual interaction for fine-grained understanding. To address these challenges, we introduce a Large-Scale Multi-Modal Industrial Open-Closed benchmark (MMIOC-1 M) containing over one million samples across 14 super-categories, 29 industrial scenes, and 351 defect subcategories. To our knowledge, MMIOC-1 M is the first unified largest benchmark supporting both open-vocabulary and closed-set industrial detection, providing valuable pre-training data for LVLMs in industrial scenarios. Furthermore, we propose a Refined Text-Visual Prompt Network (RTVPNet) that incorporates three key innovations: (1) an expert-assisted domain projection mechanism that enables rapid adaptation of general vision models to industrial domains, (2) an energy-based sparse sampling strategy that automatically generates refined visual prompts without manual intervention, and (3) a bidirectional text-visual interaction module that enhances cross-modal semantic alignment and understanding. Extensive experiments demonstrate that RTVPNet achieves state-of-the-art performance on MMIOC-1 M, LVIS, and COCO benchmarks while maintaining computational efficiency.
Jinglin Zhang 0001, Qinghui Chen, Gang Li 0005, Da Chen 0002, Shuainan Jing, Dagang Li 0001, Cong Liu 0012, Cong Bai, Shengyong Chen
IEEE Trans. Pattern Anal. Mach. Intell.10
2026 Boundary mutual information hashing for cross-modal retrieval
Cong Bai
Pattern Recognit.3
2026 Refocal Loss in Transformer for Long-Tailed Multi-Granularity Cataract Classification
abstract
Different cataract types and various severities usually require different countermeasures. For automatic cataract diagnosis, existing cataract classification methods group cataracts into common types, such as nuclear cataract, cortical cataract, and posterior subcapsular cataract, while existing cataract grading works aim to achieve fine-grained evaluation of the severity of the most common types of cataract. The severity assessment differs among various types of cataracts. Existing work is limited in predicting various cataract types at different granularity levels. In order to improve diagnostic efficiency, our study explores this matter in the context of multi-granularity cataract classification. Firstly, a large-scale dataset called Multi-Granularity Long-Tailed Cataract is collected. Secondly, an end-to-end training network is proposed, in which the Transformer is investigated for the extraction of multi-granularity cataract features. What is more, considering the imbalanced cataract data with the long-tailed distribution, the Refocal loss is proposed to rebalance the loss contribution of different classes by enhancing the reciprocal value of the effective number of samples. Compared with state-of-the-art methods, the experiments conducted on the multi-granularity cataract classification dataset demonstrate that the proposed model achieves the highest Precision of 78.22%, F1-score of 68.35%, Kappa of 64.38% and MCC of 64.49%, indicating that the proposed framework is promising in offering physicians reliable quantitative evaluations for multi-granularity cataract classification, which can help guide appropriate treatment decisions before the patient's cataracts worsen.
Yan Wang 0132, Hongdi Sun, Cong Bai
IEEE J. Biomed. Health Informatics6
2026 GVLTrack: Global Vision-Language Tracking with Multi-Stage Modal Fusion
abstract
In general, local Visual-Language (VL) trackers search targets around the previous bounding box by initial VL annotations. However, there is an inherent contradiction between the local searching perspective of the tracker and the orientation descriptions in language conducted under the global perspective. Furthermore, most methods only fuse modality information in a single stage, which tends to an insufficient relation modeling. To address these issues, we propose a Global Vision-Language Tracker (GVLTrack) with multi-stage modal fusion. First, it tracks the target in the entire image instead of local tracking based on previous results to resolve the above contradiction. Second, GVLTrack incorporates three modal interaction modules: Consistent Relationship Modeling (CRM), VL-Guided Query (VLQ) Initialization, and Recurrent Cross-Modal Decoder (RC-Decoder) to fuse vision-language modality and refine the bounding box progressively comprehensively. We conduct extensive experiments on several benchmarks and achieve competitive performance, demonstrating the effectiveness of our approach. The code will be made publicly available as soon as it is accepted.
Sixian Chan 0001, Cong Bai, Xiaoqin Zhang 0002
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Dust-Mamba: An Efficient Dust Storm Detection Network with Multiple Data Sources
abstract
Accurate detection of dust storms is challenging due to complex meteorological interactions. With the development of deep learning, deep neural networks have been increasingly applied to dust storm detection, offering better learning and generalization capabilities compared to traditional physical modeling. However, existing methods face some limitations, leading to performance bottlenecks in dust storm detection. From the task perspective, existing research focuses on occurrence detection while neglecting intensity detection. From the data perspective, existing research fails to explore the utilization of multi-source data. From the model perspective, most models are built on convolutional neural networks, which have an inherent limitation in capturing long-range dependencies. To address the challenges mentioned, this study proposes Dust-Mamba. To the best of our knowledge, this study is the first attempt to accomplish both the occurrence and intensity detection of dust storms with advanced deep learning technology. In Dust-Mamba, multi-source data is introduced to provide a comprehensive perspective, Mamba and attention are applied to boost feature selection while maintaining long-range modeling capability. Additionally, this study proposes Structure Sharing Transfer Learning Strategies for intensity detection, which further enhances the performance of Dust-Mamba with minimal time cost. As shown by experiments, Dust-Mamba achieves Dice scores of 0.963 for occurrence detection and 0.560 for intensity detection, surpassing several baseline models. In conclusion, this study offers valuable baselines for dust storm detection, with significant reference value and promising application potential.
Cong Bai, Zhonghao Lin, Jinglin Zhang 0001, Shengyong Chen
AAAI1
2025 Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline
abstract
Large Visual Language Models (LVLMs) have achieved remarkable success in vision tasks. However, the significant differences between industrial and natural scenes make applying LVLMs challenging. Existing LVLMs rely on user-provided prompts to segment objects. This often leads to suboptimal performance due to the inclusion of irrelevant pixels. In addition, the scarcity of data also makes the application of LVLMs in industrial scenarios remain unexplored. To fill this gap, this paper proposes an open industrial dataset and a Refined Text-Visual Prompt (RTVP) for zero-shot industrial defect detection. First, this paper constructs the Multi-Modal Industrial Open Dataset (MMIO) containing 80K+ samples. MMIO contains diverse industrial categories, including 6 super categories and 18 subcategories. MMIO is the first large-scale multi-scenes pre-training dataset for industrial zero-shot learning, and provides valuable training data for open models in future industrial scenarios. Based on MMIO, this paper provides a RTVP specifically for industrial zero-shot tasks. RTVP has two significant advantages: First, this paper designs an expert-guided large model domain adaptation mechanism and designs an industrial zero-shot method based on Mobile-SAM, which enhances the generalization ability of large models in industrial scenarios. Second, RTVP automatically generates visual prompts directly from images and considers text-visual prompt interactions ignored by previous LVLM, improving visual and textual content understanding. RTVP achieves SOTA with 42.2% and 24.7% AP in zero-shot and closed scenes of MMIO.
Qinghui Chen, Maomao Xiong, Shijiao Ding, Zhanzhi Su, Xinjie Yao, Cong Bai
AAAI8
2025 TC-Diffuser: Bi-Condition Multi-Modal Diffusion for Tropical Cyclone Forecasting
abstract
Tropical cyclones (TCs) are complex weather systems with strong winds and heavy rainfall, causing substantial loss of life and property. Therefore, accurate TC forecasting is crucial for the effective prevention of disasters caused by TCs. TC forecasting can be regarded as a spatio-temporal prediction problem. It has been proven that using multi-modal data can effectively introduce atmospheric information to achieve better prediction results and higher interpretability. But it also introduces inevitably introduces noise into the prediction process. The diffusion model's unique noise modeling capability can reduce prediction noise when using multi-modal datasets. However, adapting it to TC forecasting has two main challenges: how to extract valuable information from multi-modal data, and how to utilize them to guide the generation process. For the first challenge, while recent methods can predict multiple TC attributes using multi-modal data, they often overlook the interdependence of multiple attributes and the semantic gap between modalities. Considering the interdependence of attributes, we propose two condition generators that capture the commonalities and characteristics of TC attributes, extracting spatio-temporal and environmental features and incorporating expert knowledge. To reduce the semantic gap between multi-modal data, we introduce the PGSA-LSTM module to map primary and auxiliary modalities. For the second challenge, we propose a novel Bi-condition diffusion model that sequentially processes conditions from the characteristics to commonalities of attributes, thereby expanding the guidance information that the diffusion model can accept. Our results surpass state-of-the-art deep learning models and outperform the numerical weather prediction model used by the China Central Meteorological Observatory. TC-Diffuser shows high generalizability across global ocean areas, strong robustness in handling missing data, and higher computational efficiency.
Shiqi Zhang 0009, Pan Mu, Cong Bai
AAAI5
2025 NeighborRetr: Balancing Hub Centrality in Cross-Modal Retrieval
abstract
Cross-modal retrieval aims to bridge the semantic gap between different modalities, such as visual and textual data, enabling accurate retrieval across them. Despite significant advancements with models like CLIP that align cross-modal representations, a persistent challenge remains: the hubness problem, where a small subset of samples (hubs) dominate as nearest neighbors, leading to biased representations and degraded retrieval accuracy. Existing methods often mitigate hubness through post-hoc normalization techniques, relying on prior data distributions that may not be practical in real-world scenarios. In this paper, we directly mitigate hubness during training and introduce NeighborRetr, a novel method that effectively balances the learning of hubs and adaptively adjusts the relations of various kinds of neighbors. Our approach not only mitigates the hubness problem but also enhances retrieval performance, achieving state-of-the-art results on multiple cross-modal retrieval benchmarks. Furthermore, Neighbor-Retr demonstrates robust generalization to new domains with substantial distribution shifts, highlighting its effectiveness in real-world applications. We make our code publicly available at: https://github.com/NeighborRetr.
Zengrong Lin, Zheng Wang 0059, Tianwen Qian, Pan Mu, Sixian Chan 0001, Cong Bai
CVPR6
2025 LS-Mamba: Spatial-Spectral Mamba for Multispectral Cloud Image Semantic Segmentation
abstract
Cloud semantic segmentation, which assigns semantic labels to each pixel in multispectral images, plays a critical role in weather analysis and climate studies. Despite recent advancements in deep learning and the emergence of the Mamba architecture, existing methods for cloud segmentation continue to face significant challenges. In particular, current approaches often fall short in effectively modeling the complex relationships among spectral channels, which can lead to ambiguous representations and result in misclassification, especially of spectrally similar cloud types. Additionally, while Mamba excels at long-range modeling, it often overlooks local 2D structural dependencies, resulting in inaccuracies for clouds with complex spatial distributions. To address these challenges, we propose a novel cloud semantic classification model based on Spatial-Spectral Mamba. We design a spectral Mamba block (SpeMamba) to capture intricate intraspectral relationships to improve discrimination between confused cloud types, and also design a spatial Mamba block to model local-global dependencies through local-global scanning, preserving fine-grained spatial structures while maintaining global features. The proposed method is evaluated on the Himawari-8 image dataset, and the experimental results demonstrate the effectiveness of the proposed method, achieving the new state-of-the-art performance. Codes are available at https://github.com/Zjut-MultimediaPlus/LS-Mamba.
Zhiying Hu 0005, Cong Bai
ECAI3
2025 PiCNet: Physics-infused Convolution Network for Radar-Based Precipitation Nowcasting
abstract
Meteorological disasters, especially extreme precipitation, cause significant socioeconomic damage, highlighting the need for effective quantitative precipitation nowcasting. Existing methods, often data-driven and resource-intensive, struggle to capture the underlying physical laws of meteorology. This paper introduces a simple yet effective model using an advection simulator to learn precipitation’s physical dynamics, making the predictions more interpretable. Our model also incorporates a physics-guided module to enhance sensitivity to high-intensity rainfall, improving rainfall prediction accuracy. Experiments on the KNMI radar echo dataset demonstrate that our model outperforms state-of-the-art methods, offering better insights into physics-infused precipitation nowcasting.
Zheng Wang 0059, Hanyi Zhang, Cong Bai
ICASSP3
2025 Prompt-UIE: A Unified Prompt-Driven Framework for Underwater Image Enhancement
abstract
The complex and diverse underwater environment causes various types of degradation in underwater images. However, most existing methods focus on single underwater datasets, where the similarities in degradation limit the model’s exploration of different degradation characteristics. To address this challenge, we developed a new unified model for underwater image enhancement based on prompt learning, called PromptUIE, focusing on the common attributes of underwater images. Prompt-UIE is designed to adapt the pre-trained specific model to various underwater conditions without the need of multiple datasets. It builds upon a specific model and integrates a visual prompt module along with a reverse transmission map (RTM) guided loss function. First, the visual prompt module based on background light guides the specific model to perform enhancements based on different water types through a carefully designed visual prompt strategy. Next, the RTM guided loss function improves the model’s ability to handle non-uniform degradation. Experiments on both full-reference and no-reference datasets demonstrate the effectiveness and robustness of our method. The code is available at https://github.com/Zjut-MultimediaPlus/Prompt-UIE.
Yanling Zhang, Linxuan Luo, Pan Mu, Cong Bai
ICASSP4
2025 TCP-Diffusion: A Multi-modal Diffusion Model for Global Tropical Cyclone Precipitation Forecasting with Change Awareness
abstract
Deep learning methods have made significant progress in regular rainfall forecasting, yet the more hazardous tropical cyclone (TC) rainfall has not received the same attention. While regular rainfall models can offer valuable insights for designing TC rainfall forecasting models, most existing methods suffer from cumulative errors and lack physical consistency. Additionally, these methods overlook the importance of meteorological factors in TC rainfall and their integration with the numerical weather prediction (NWP) model. To address these issues, we propose Tropical Cyclone Precipitation Diffusion (TCP-Diffusion), a multi-modal model for forecasting of TC precipitation given an existing TC in any location globally. It forecasts rainfall around the TC center for the next 12 hours at 3 hourly resolution based on past rainfall observations and multi-modal environmental variables. Adjacent residual prediction (ARP) changes the training target from the absolute rainfall value to the rainfall trend and gives our model the capability of rainfall change awareness, reducing cumulative errors and ensuring physical consistency. Considering the influence of TC-related meteorological factors and the useful information from NWP model forecasts, we propose a multi-model framework with specialized encoders to extract richer information from environmental variables and results provided by NWP models. The results of extensive experiments show that our method outperforms other DL methods and the NWP method from the European Centre for Medium-Range Weather Forecasts (ECMWF).
Pan Mu, Cong Bai, Peter AG Watson
ICML3
2025 From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion
abstract
Accurate near-real-time precipitation retrieval has been enhanced by satellite-based technologies.However, infrared-based algorithms have low accuracy due to weak relations with surface precipitation, whereas passive microwave and radar-based methods are more accurate but limited in range.This challenge motivates the Precipitation Retrieval Expansion (PRE) task, which aims to enable accurate, infrared-based full-disc precipitation retrievals beyond the scanning swath.We introduce Multimodal Knowledge Expansion, a two-stage pipeline with the proposed PRE-Net model.In the Swath-Distilling stage, PRE-Net transfers knowledge from a multimodal data integration model to an infrared-based model within the scanning swath via Coordinated Masking and Wavelet Enhancement (CoMWE).In the Full-Disc Adaptation stage, Self-MaskTune refines predictions across the full disc by balancing multimodal and full-disc infrared knowledge.Experiments on the introduced PRE benchmark demonstrate that PRE-Net significantly advanced precipitation retrieval performance, outperforming leading products like PERSIANN-CCS, PDIR, and IMERG.The code will be available at https://github.com/Zjut-MultimediaPlus/PRE-Net.
Zheng Wang 0059, Kai Ying, Bin Xu 0017, Chunjiao Wang, Cong Bai
KDD (2)5
2025 Event-Driven Hybrid and Cross-Stage Guide for Video Corpus Moment Retrieval
Zheng Wang 0059, Zengrong Lin, Cong Bai
ICMR4
2025 Physics-Coupled Frequency Dynamic Adaptation Network for Domain Generalized Underwater Object Detection
Linxuan Luo, Pan Mu, Cong Bai
ACM Multimedia3
2025 Amodal-KAN: The First Look at Kolmogorov-Arnold Network for Amodal Instance Segmentation
abstract
Amodal instance segmentation has emerged as a critical task in the field of segmentation, facilitating the understanding of complex real-world scenes. Although numerous innovative designs and improvements have been introduced by incorporating techniques such as transformers, existing networks are still constrained to linear pattern modeling and struggle to effectively capture complex nonlinear relationships. Motivated by the strong performance of Kolmogorov–Arnold Networks (KANs), which redefine the learning paradigm by employing stacks of nonlinear, learnable activation functions derived from the Kolmogorov–Arnold representation theorem, we aim to address these limitations. Specifically, in this paper, we explore the underutilized potential of KANs to improve amodal instance segmentation architectures. We study existing amodal instance segmentation pipelines and integrate KANs to enhance amodal feature representations. Additionally, we investigate replacing traditional predictors in amodal instance segmentation with KAN-based predictors. Experiments on amodal instance segmentation datasets demonstrates the superiority of our proposed KANet.
Cong Bai
MMAsia3
2025 IDOL: Meeting Diverse Distribution Shifts with Prior Physics for Tropical Cyclone Multi-Task Estimation
abstract
Tropical Cyclone (TC) estimation aims to accurately estimate various TC attributes in real time. However, distribution shifts arising from the complex and dynamic nature of TC environmental fields, such as varying geographical conditions and seasonal changes, present significant challenges to reliable estimation. Most existing methods rely on multi-modal fusion for feature extraction but overlook the intrinsic distribution of feature representations, leading to poor generalization under out-of-distribution (OOD) scenarios. To address this, we propose an effective Identity Distribution-Oriented Physical Invariant Learning framework (IDOL), which imposes identity-oriented constraints to regulate the feature space under the guidance of prior physical knowledge, thereby dealing distribution variability with physical invariance. Specifically, the proposed IDOL employs the wind field model and dark correlation knowledge of TC to model task-shared and task-specific identity tokens. These tokens capture task dependencies and intrinsic physical invariances of TC, enabling robust estimation of TC wind speed, pressure, inner-core, and outer-core size under distribution shifts. Extensive experiments conducted on multiple datasets and tasks demonstrate the outperformance of the proposed IDOL, verifying that imposing identity-oriented constraints based on prior physical knowledge can effectively mitigates diverse distribution shifts in TC estimation.
Hanting Yan, Pan Mu, Shiqi Zhang 0009, Yuchao Zhu, Cong Bai
NeurIPS6
2025 Lightweight Multi-Stage Aggregation Transformer for robust medical image segmentation
Xiaoyan Wang 0007, Yating Zhu, Dongyan Guo, Pan Mu, Ming Xia 0005, Cong Bai, Zhongzhao Teng, Shengyong Chen
Medical Image Anal.8
2025 FocTrack: Focus attention for visual tracking
Sixian Chan 0001, Zhenchao Shi, Cong Bai, Shengyong Chen
Pattern Recognit.4
2025 DBFA-TSNet: A Three-Stage Building Extraction Network Based on Dual-Branch Fusion and Adaptive Enhancement
Jianan Chen 0002, Sixian Chan 0001, Cong Bai
IEEE Trans. Geosci. Remote. Sens.4
2025 Sparse Information Perception Network for Remote Sensing Cross-Modal Retrieval
abstract
In recent years, research on remote sensing (RS) cross-modal retrieval (RSCMR) has gained significant attention, yet the inherent sparsity of RS images has posed challenges in aligning them with corresponding text, leading to reduced retrieval performance. To address this, we propose the sparse information perception network (SIPN), an instance-level dynamic network that adapts to sparse information in RS images and text. Our approach incorporates a dynamic feature fusion module (DFFM) for initializing the network, dynamic training adjustments using random processes and pretrained policy networks, and feature alignment via contrastive loss. Additionally, we enhance feature selection and fusion capabilities through a multilevel attention module (MLAM), which includes cross-attention (CA) and a multilayer perceptron (MLP), with triplet loss for further alignment. Our method demonstrates superior performance in RSCMR on the RSICD and RSITMD datasets.
Pengxiang Ouyang, Cong Bai
IEEE Trans. Geosci. Remote. Sens.3
2025 Precipitation Retrieval Integrating Multiple Satellite Observations: A Dataset and a Framework
abstract
Multimodal satellite observations have been widely used for precipitation retrieval. Numerous retrieval algorithms and precipitation products have been developed based on these data. However, the integrated retrieval of multimodal data remains challenging due to the modality heterogeneity caused by different data characteristics reflecting precipitation patterns. To address these issues, effectively integrating multimodal data is crucial. We propose a framework named Precipitation Retrieval Integrating Multiple Satellite Observations And Geographical Information (PRMG), consisting of two networks for precipitation identification and estimation respectively. Specifically, PRMG includes a Multi-Branch Fusion (MBF) module for integrating three types of satellite observations: infrared (IR), passive microwave (PMW), and spaceborne precipitation radar (PR), and a geographical information correction (GIC) module to incorporate the geographical information to calibrate precipitation features. We also collect a new precipitation retrieval dataset for multimodal precipitation retrieval, called Precipitation-MG, which includes satellite observations, corresponding geographical information, and precipitation products. Extensive experiments on Precipitation-MG demonstrate the effectiveness of the multimodal fusion method and the geographic correction method. The retrieval performance of PRMG achieves significant improvements compared to the Global Precipitation Measurement currently in operation, i.e., (GPM) Level-2 DPR and GMI Combined (2B-CMB) product. The source code and dataset are publicly available at https://github.com/Zjut-MultimediaPlus/PRMG.
Zheng Wang 0059, Boxian He, Chunjiao Wang, Bin Xu 0017, Cong Bai
IEEE Trans. Geosci. Remote. Sens.5
2024 Phy-CoCo: Physical Constraint-Based Correlation Learning for Tropical Cyclone Intensity and Size Estimation
abstract
Tropical Cyclone (TC) estimation aims to estimate various attributes of TC in real-time to alleviate and prevent disasters caused by violent TCs. As artificial intelligence technology advances, various deep learning-based multi-task estimation approaches have been proposed. However, most of them only focus on extracting common features of tasks, disregarding potential negative transfer and task interactions between different tasks. This paper is thus motivated to propose a Physical Constraint-based Correlation (Phy-CoCo) learning framework from the perspective of Multi-Task Learning (MTL). Specifically, for task-specific feature learning, we introduce Correlation Modeling (CoM) based on Centrally Expanded Pooling (CEP). Furthermore, for cross-task interaction, we propose a Multi-Domain Recurrent Convolution (MDRC) module to incorporate physical constraints into MTL. These physical constraints enable the transformation of different task features by simulating the physical relations among different attributes of TC. Lastly, in combination with a task-shared network that leverages the hybrid fusion of multi-modal data, our MTL framework accurately estimates various TC attributes. Extensive experiments conducted on our constructed dataset demonstrate that the proposed Phy-CoCo outperforms previous methods in TC estimation in terms of estimation error, verifying the potential of the physics-incorporated MTL model.
Hanting Yan, Pan Mu, Cong Bai
ECAI5
2024 Distinguishing Visually Similar Images: Triplet Contrastive Learning Framework for Image-text Retrieval
abstract
In recent years, contrastive learning techniques, particularly InfoNCE loss, have propelled advancements in image-text alignment. However, aligning indistinguishable image-text pairs remains a challenge for conventional methods, often overlooking semantically similar content. To overcome these limitations, we introduce the Triplet Contrast Learning Framework (TCLF). Accompanying TCLF is the Stricter Noise Contrast Estimation (SNCE) loss, designed to minimize mutual information between positive and negative image-text pairs. Additionally, we propose the intra-modal mutual information (IMI) loss to encourage a uniform distribution of image features and enhance discriminative capacity for similar images. SNCE aligns two enhanced image versions, and collectively, SNCE and IMI empower TCLF for effective learning with challenging image-text pairs. Experimental evaluations on MSCOCO and Flickr30K datasets demonstrate superior performance compared to state-of-the-art methods. Ablation studies confirm the efficacy of SNCE and IMI in overcoming identified challenges.
Pengxiang Ouyang, Jianan Chen 0002, Zheng Wang 0059, Cong Bai
ICME5
2024 Live on the Hump: Self Knowledge Distillation via Virtual Teacher-Students Mutual Learning
abstract
For solving the limitations of the current self knowledge distillation including never fully utilizing the knowledge of shallow exits and neglecting the impact of auxiliary exits' structure on the performance of network, a novel self knowledge distillation framework via virtual teacher-students mutual learning named LOTH is proposed in this paper. A knowledgeable virtual teacher is constructed from the rich feature maps of each exit to help the learning of each exit. Meanwhile, the logit knowledges of each exit are incorporated to guide the learning of the virtual teacher. They learn mutually through the well-designed loss in LOTH. Moreover, two kinds of auxiliary building blocks are designed to balance the efficiency and effectiveness of network. Extensive experiments with diverse backbones on CIFAR-100 and Tiny-ImageNet validate the effectiveness of LOTH, which realizes superior performance with less resource by the comparison with the state-of-the-art distillation methods. The code of LOTH is available on Github https://github.com/cloak-s/LOTH.
Shuang Wang 0016, Pengyi Hao, Fuli Wu, Cong Bai
ACM Multimedia4
2024 Intent-Aware Graph-Level Embedding Learning Based Recommendation
Pengyi Hao, Si-Hao Liu, Cong Bai
J. Comput. Sci. Technol.3
2024 Multi-scale feature correspondence and restriction mechanism for visible X-ray baggage re-Identification
Sixian Chan 0001, Jiaao Cui, Yonggan Wu, Cong Bai
Multim. Syst.5
2024 An efficient and real-time steel surface defect detection method based on single-stage detection algorithm
Qiqi Miao, Suqiang Li, Sixian Chan 0001, Jie Hu 0041, Cong Bai
Multim. Tools Appl.7
2024 Direction-Oriented Visual-Semantic Embedding Model for Remote Sensing Image-Text Retrieval
abstract
Image-text retrieval has developed rapidly in recent years. However, it is still a challenge in remote sensing due to visual-semantic imbalance, which leads to incorrect matching of non-semantic visual and textual features. To solve this problem, we propose a novel Direction-Oriented Visual-semantic Embedding Model (DOVE) to mine the relationship between vision and language. Our highlight is to conduct visual and textual representations in latent space, directing them as close as possible to a redundancy-free regional visual representation. Concretely, a Regional-Oriented Attention Module (ROAM) adaptively adjusts the distance between the final visual and textual embeddings in the latent semantic space, oriented by regional visual features. Meanwhile, a lightweight Digging Text Genome Assistant (DTGA) is designed to expand the range of tractable textual representation and enhance global word-level semantic connections using less attention operations. Ultimately, we exploit a global visual-semantic constraint to reduce single visual dependency and serve as an external constraint for the final visual and textual representations. The effectiveness and superiority of our method are verified by extensive experiments including parameter evaluation, quantitative comparison, ablation studies and visual analysis, on two benchmark datasets, RSICD and RSITMD.
Jiancheng Pan, Cong Bai
IEEE Trans. Geosci. Remote. Sens.3
2024 Diverse-Feature Collaborative Progressive Learning for Visible-Infrared Person Re-Identification
abstract
Visible–infrared person reidentification (VI-ReID) aims to search for pedestrian identities in different spectra. The major challenge is the modality differences between infrared and visible images for the VI-ReID task. Existing approaches try to design networks based on a single-stage training strategy to extract features. However, they often excessively rely on a particular feature, such as modality-specific features or modality-independent features, and overlook the significance of the diverse features obtained by combining them. To address this problem, we propose a diverse-feature collaborative progressive learning network (DCPLNet) for VI-ReID in this article. With the benefit of diverse information, our DCPLNet can effectively learn informative representations for reducing the modality differences. Specifically, we propose a novel three-stage progressive learning strategy (t-PLS) to progressively learn diverse features. For the proposed t-PLS, we design a contour feature enhancement module to mine human contour features and raise a perceptual contour feature loss for supervised feature extraction. Finally, we advance a batch adaptation module to establish feature links between samples. Extensive experiments on SYSU-MM01, RegDB, and LLCM datasets demonstrate that our proposed model performs better than most state-of-the-art methods.
Sixian Chan 0001, Weihao Meng, Cong Bai, Jie Hu 0041, Shenyong Chen
IEEE Trans. Ind. Informatics3
2024 Graph Convolutional Network Discrete Hashing for Cross-Modal Retrieval
abstract
With the rapid development of deep neural networks, cross-modal hashing has made great progress. However, the information of different types of data is asymmetrical, that is to say, if the resolution of an image is high enough, it can reproduce almost 100% of the real-world scenes. However, text usually carries personal emotion and it is not objective enough, so we generally think that the information of image will be much richer than text. Although most of the existing methods unify the semantic feature extraction and hash function learning modules for end-to-end learning, they ignore this issue and do not use information-rich modalities to support information-poor modalities, leading to suboptimal results, although they unify the semantic feature extraction and hash function learning modules for end-to-end learning. Furthermore, previous methods learn hash functions in a relaxed way that causes nontrivial quantization losses. To address these issues, we propose a new method called graph convolutional network (GCN) discrete hashing. This method uses a GCN to bridge the information gap between different types of data. The GCN can represent each label as word embedding, with the embedding regarded as a set of interdependent object classifiers. From these classifiers, we can obtain predicted labels to enhance feature representations across modalities. In addition, we use an efficient discrete optimization strategy to learn the discrete binary codes without relaxation. Extensive experiments conducted on three commonly used datasets demonstrate that our proposed method graph convolutional network-based discrete hashing (GCDH) outperforms the current state-of-the-art cross-modal hashing methods.
Cong Bai, Jinglin Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Self-Supervised Enhancement for Named Entity Disambiguation via Multimodal Graph Convolution
abstract
Named entity disambiguation (NED) finds the specific meaning of an entity mention in a particular context and links it to a target entity. With the emergence of multimedia, the modalities of content on the Internet have become more diverse, which poses difficulties for traditional NED, and the vast amounts of information make it impossible to manually label every kind of ambiguous data to train a practical NED model. In response to this situation, we present MMGraph, which uses multimodal graph convolution to aggregate visual and contextual language information for accurate entity disambiguation for short texts, and a self-supervised simple triplet network (SimTri) that can learn useful representations in multimodal unlabeled data to enhance the effectiveness of NED models. We evaluated these approaches on a new dataset, MMFi, which contains multimodal supervised data and large amounts of unlabeled data. Our experiments confirm the state-of-the-art performance of MMGraph on two widely used benchmarks and MMFi. SimTri further improves the performance of NED methods. The dataset and code are available at https://github.com/LanceZPF/NNED_MMGraph.
Kaining Ying, Zhenhua Wang 0003, Dongyan Guo, Cong Bai
IEEE Trans. Neural Networks Learn. Syst.5
2024 Auxiliary Feature Fusion and Noise Suppression for HOI Detection
abstract
In recent years, one-stage HOI (Human–Object Interaction) detection methods tend to divide the original task into multiple sub-tasks by using a multi-branch network structure. However, there is no sufficient attention to information communication between these branches. The inference approach in the cascaded structure is singular, while fully parallel methods will disrupt the associations between different pieces of information. Besides, noise interference may occur during the fusion of different features and thus affect the detection performance. To address these issues, this article proposes a one-stage three-branch parallel HOI detection method, which treats HOI as three separate sub-tasks (human detection, object detection, and interaction detection) and leverages three distinct reasoning relationships to generate richer relational information. Firstly , an auxiliary feature fusion (AFF) module is introduced, which integrates features originally extracted independently to form fused features enriched with supplementary information. This approach strengthens communication between branches in the network while handling the three sub-tasks concurrently, thereby facilitating the exchange of more contextual information. Secondly , to mitigate noise interference generated during the fusion process, a fusion noise suppression (FNS) module is introduced, which effectively suppresses noise and enhances the model’s performance in interaction detection tasks. Finally , experiments are conducted on two major benchmark datasets, and experimental results show that our HOI detection method is superior to previous methods. Also, ablation studies confirm the effectiveness of all the components in our proposed method.
Sixian Chan 0001, Xianpeng Zeng, Jie Hu 0041, Cong Bai
ACM Trans. Multim. Comput. Commun. Appl.5
2023 MGTCF: Multi-Generator Tropical Cyclone Forecasting with Heterogeneous Meteorological Data
abstract
Accurate forecasting of tropical cyclone (TC) plays a critical role in the prevention and defense of TC disasters. We must explore a more accurate method for TC prediction. Deep learning methods are increasingly being implemented to make TC prediction more accurate. However, most existing methods lack a generic framework for adapting heterogeneous meteorological data and do not focus on the importance of the environment. Therefore, we propose a Multi-Generator Tropical Cyclone Forecasting model (MGTCF), a generic, extensible, multi-modal TC prediction model with the key modules of Generator Chooser Network (GC-Net) and Environment Net (Env-Net). The proposed method can utilize heterogeneous meteorologic data efficiently and mine environmental factors. In addition, the Multi-generator with Generator Chooser Net is proposed to tackle the drawbacks of single-generator TC prediction methods: the prediction of undesired out-of-distribution samples and the problems stemming from insufficient learning ability. To prove the effectiveness of MGTCF, we conduct extensive experiments on the China Meteorological Administration Tropical Cyclone Best Track Dataset. MGTCF obtains better performance compared with other deep learning methods and outperforms the official prediction method of the China Central Meteorological Observatory in most indexes.
Cong Bai, Sixian Chan 0001, Yuquan Wu
AAAI2
2023 Multi-Stage Aggregation Transformer for Medical Image Segmentation
abstract
Capturing rich multi-scale features is essential for resolving complex variations in medical image segmentation. In this paper, we explore how to fully utilize the advantages of Convolutional neural networks (CNN) and Transformer, and propose a novel multi-stage aggregation architecture named MA-Transformer for accurate segmentation of medical images with large variations and blurs. Specifically, an encoder module is introduced in each stage, which is a dual-branch structure parallelly combining Transformers and convolutions. By such design, the self-attention can provide a global context for CNN to extract multi-resolution complementary features stage by stage, thus the feature representations are gradually enhanced with local details and contextual information. Multi-scale semantic features are then combined with skip connections in the decoder to produce the final result. Extensive experiments on public medical imaging datasets demonstrate our superior segmentation performance, compared to the state-of-the-art CNN-based, Transformer-based approaches and CNN-Transformer combined approaches. Code will be made publicly available.
Xiaoyan Wang 0007, Minghan Shao, Dongyan Guo, Ming Xia 0005, Cong Bai
ICASSP7
2023 Towards General and Fast Video Derain via Knowledge Distillation
abstract
As a common natural weather condition, rain can obscure video frames and thus affect the performance of the visual system, so video derain receives a lot of attention. In natural environments, rain has a wide variety of streak types, which increases the difficulty of the rain removal task. In this paper, we propose a Rain Review-based General video derain Network via knowledge distillation (named RRGNet) that handles different rain streak types with one pre-training weight. Specifically, we design a frame grouping-based encoder-decoder network that makes full use of the temporal information of the video. Further, we use the old task model to guide the current model in learning new rain streak types while avoiding forgetting. To consolidate the network’s ability to derain, we design a rain review module to play back data from old tasks for the current model. The experimental results show that our developed general method achieves the best results in terms of running speed and derain effect.
Defang Cai, Pan Mu, Sixian Chan 0001, Zhanpeng Shao, Cong Bai
ICME5
2023 Visible-Xray Cross-Modality Package Re-Identification
abstract
During the package inspection, once prohibited articles are checked under the X-ray, the inspector needs to find the corresponding package in time for further confirmation. When the number of passengers increases, this process is time-consuming. Recently, prohibited articles detection as a detection task has attracted much attention, but few studies have focused on Visible-Xray package re-identification (VX-ReID) task. In this paper, we mainly explore the VX-ReID task. Firstly, we establish the first VX-ReID dataset RX01, which includes 55883 Visible and 29174 X-ray package images. Furthermore, we introduce a baseline model that includes a cross-modality channel attention module (CMCA) and momentum contrast mAP (MoCoAP). CMCA is used to enhance channels that contain modality-invariant information. MoCoAP is a differentiable mAP approximation strategy that directly optimizes the retrieve performance of the model. By combining these two strategies, we achieve competitive performance on the RX01 and SYSU-MM01 datasets. Code will be released at https://github.com/cjjjao/VX-ReID.
Sixian Chan 0001, Jiaao Cui, Yonggan Wu, Cong Bai
ICME5
2023 Histogram-guided Video Colorization Structure with Spatial-Temporal Connection
abstract
Video colorization, aiming at obtaining colorful and plausible results from grayish frames, has aroused a lot of interest recently. Nevertheless, how to maintain temporal consistency while keeping the quality of colorized results remains challenging. To tackle the above problems, we present a Histogram-guided Video Colorization with Spatial-Temporal connection structure (named ST-HVC). To fully exploit the chroma and motion information, the joint flow and histogram module is tailored to integrate the histogram and flow features. To manage the blurred and artifact, we design a combination scheme attending to temporal detail and flow feature combination. We further recombine the histogram, flow and sharpness features via a U-shape network. Extensive comparisons are conducted with several state-of-the-art image and video-based methods, demonstrating that the developed method achieves excellent performance both quantitatively and qualitatively in two video datasets.
Zheyuan Liu 0009, Pan Mu, Hanning Xu, Cong Bai
ICME4
2023 Transmission and Color-guided Network for Underwater Image Enhancement
abstract
In recent years, with the continuous development of the marine industry, underwater image enhancement has attracted plenty of attention. Unfortunately, the propagation of light in water will be absorbed by water bodies and scattered by suspended particles, resulting in color deviation and low contrast. To solve these two problems, we propose an Adaptive Transmission and Dynamic Color guided network (named ATDCnet) for underwater image enhancement. In particular, to exploit the knowledge of physics, we design an Adaptive Transmission-directed Module (ATM) to better guide the network. To deal with the color deviation problem, we design a Dynamic Color-guided Module (DCM) to post-process the enhanced image color. Further, we design an Encoder-Decoder-based (EDC) structure with attention and a multistage feature fusion mechanism to perform color restoration and contrast enhancement simultaneously. Extensive experiments demonstrate the state-of-the-art performance of the ATDCnet on multiple benchmark datasets.
Pan Mu, Haotian Qian, Cong Bai
ICME4
2023 SGPT: The Secondary Path Guides the Primary Path in Transformers for HOI Detection
abstract
HOI detection is essential for human-computer interaction, especially in behavior detection and robot manipulation. Existing mainstream transformer methods of HOI detection are focused on single-stream detection only, e.g.,$image \rightarrow HOI(\mathcal{P}_{1})$, or$image \rightarrow HO\rightarrow I(\mathcal{P}_{2})$. Both paths have their own characteristics of concern, so we propose a novel method, using the Secondary path$(\mathcal{P}_{2})$Guides the Primary path$(\mathcal{P}_{1})$in Transformers (SGPT). SGPT contains two core modules: the Dual-Path Consistency (DPC) module and the Instance Interaction Attention (IIA) module. DPC keeps human, object and interaction consistent on the dual-path and lets$\mathcal{P}_{2}$guide$\mathcal{P}_{1}$to learn more meaningful features. IIA fuses human and object to enhance interaction in$\mathcal{P}_{2}$, which allows instance to constrain interaction. Our proposed dual-path are employed during training, and only the$\mathcal{P}_{1}$path is used for inference. Hence, SGPT improves generalization without increasing model capacity in HICO-DET and V-COCO datasets compared to the state-of-the-arts. The code of this work is available at https://github.com/visualVk/sgpt.git.
Sixian Chan 0001, Weixiang Wang, Zhanpeng Shao, Cong Bai
ICRA4
2023 Reducing Semantic Confusion: Scene-aware Aggregation Network for Remote Sensing Cross-modal Retrieval
abstract
Recently, remote sensing cross-modal retrieval has received incredible attention from researchers. However, the unique nature of remote-sensing images leads to many semantic confusion zones in the semantic space, which greatly affects retrieval performance. We propose a novel scene-aware aggregation network (SWAN) to reduce semantic confusion by improving scene perception capability. In visual representation, a visual multiscale fusion module (VMSF) is presented to fuse visual features with different scales as a visual representation backbone. Meanwhile, a scene fine-grained sensing module (SFGS) is proposed to establish the associations of salient features at different granularity. A scene-aware visual aggregation representation is formed by the visual information generated by these two modules. In textual representation, a textual coarse-grained enhancement module (TCGE) is designed to enhance the semantics of text and to align visual information. Furthermore, as the diversity and differentiation of remote sensing scenes weaken the understanding of scenes, a new metric, namely, scene recall is proposed to measure the perception of scenes by evaluating scene-level retrieval performance, which can also verify the effectiveness of our approach in reducing semantic confusion. By performance comparisons, ablation studies and visualization analysis, we validated the effectiveness and superiority of our approach on two datasets, RSICD and RSITMD. The source code is available at https://github.com/kinshingpoon/SWAN-pytorch.
Jiancheng Pan, Cong Bai
ICMR3
2023 Little Strokes Fell Great Oaks: Boosting the Hierarchical Features for Multi-exposure Image Fusion
abstract
In recent years, deep learning networks have made remarkable strides in the domain of multi-exposure image fusion. Nonetheless, prevailing approaches often involve directly feeding over-exposed and under-exposed images into the network, which leads to the under-utilization of inherent information present in the source images. Additionally, unsupervised techniques predominantly employ rudimentary weighted summation for color channel processing, culminating in an overall desaturated final image tone. To partially mitigate these issues, this study proposes a gamma correction module specifically designed to fully leverage latent information embedded within source images. Furthermore, a modified transformer block, embracing self-attention mechanisms, is introduced to optimize the fusion process. Ultimately, a novel color enhancement algorithm is presented to augment image saturation while preserving intricate details. The source code is available at https://github.com/ZhiyingDu/BHFMEF.
Pan Mu, Zhiying Du, Jinyuan Liu 0001, Cong Bai
ACM Multimedia4
2023 A Generalized Physical-knowledge-guided Dynamic Model for Underwater Image Enhancement
abstract
Underwater images often suffer from color distortion and low contrast resulting in various image types, due to the scattering and absorption of light by water. While it is difficult to obtain high-quality paired training samples with a generalized model. To tackle these challenges, we design a Generalized Underwater image enhancement method via a Physical-knowledge-guided Dynamic Model (short for GUPDM). In particular, to cover complex underwater scenes, this study changes the global atmosphere light and the transmission to simulate various underwater image types through the formation model. We then design an Atmosphere-based Dynamic Structure (ADS) and Transmission-guided Dynamic Structure (TDS) that use dynamic convolutions to adaptively extract prior information from underwater images and generate parameters for Prior-based Multi-scale Structure (PMS). These two modules enable the network to select appropriate parameters for various water types adaptively. Besides, the multi-scale feature extraction module in PMS uses convolution blocks with different kernel sizes and obtains weights for each feature map via channel attention block. The source code will be available at https://github.com/shiningZZ/GUPDM
Pan Mu, Hanning Xu, Zheyuan Liu 0009, Zheng Wang 0059, Sixian Chan 0001, Cong Bai
ACM Multimedia6
2023 A Prior Instruction Representation Framework for Remote Sensing Image-text Retrieval
abstract
This paper presents a prior instruction representation framework (PIR) for remote sensing image-text retrieval, aimed at remote sensing vision-language understanding tasks to solve the semantic noise problem. Our highlight is the proposal of a paradigm that draws on prior knowledge to instruct adaptive learning of vision and text representations. Concretely, two progressive attention encoder (PAE) structures, Spatial-PAE and Temporal-PAE, are proposed to perform long-range dependency modeling to enhance key feature representation. In vision representation, Vision Instruction Representation (VIR) based on Spatial-PAE exploits the prior-guided knowledge of the remote sensing scene recognition by building a belief matrix to select key features for reducing the impact of semantic noise. In text representation, Language Cycle Attention (LCA) based on Temporal-PAE uses the previous time step to cyclically activate the current time step to enhance text representation capability. A cluster-wise affiliation loss is proposed to constrain the inter-classes and to reduce the semantic confusion zones in the common subspace. Comprehensive experiments demonstrate that using prior knowledge instruction could enhance vision and text representations and could outperform the state-of-the-art methods on two benchmark datasets, RSICD and RSITMD. Codes are available at https://github.com/Zjut-MultimediaPlus/PIR-pytorch.
Jiancheng Pan, Cong Bai
ACM Multimedia3
2023 Meta-relationship for course recommendation in MOOCs
Pengyi Hao, Cong Bai
Multim. Syst.3
2023 Triple-level relationship enhanced transformer for image captioning
Anqi Zheng, Shiqi Zheng, Cong Bai, Deng Chen
Multim. Syst.3
2023 Mine Diversified Contents of Multispectral Cloud Images Along With Geographical Information for Multilabel Classification
abstract
Multispectral multilabel cloud image classification (MSMLCIC) aims to predict a set of labels presented in a multispectral (MS) cloud image, which usually contains more than one cloud type or weather system. However, the exploration of diversified contents reflected by multiple bands of MS image is limited and the consideration of geographical information (time and location information) is insufficient. To cope with the abovementioned problems, this work proposes the multispectral cloud image multilabel classifier with group feature extractor and geo-queries (MS-GoGo). With a group feature extractor, different bands of MS images are processed separately according to the content they reflected, and a group of distinctive yet complementary image features are generated. Geo-queries are responsible for implicitly embedding different labels with time and location information to probe the corresponding similar semantic ingredients. Due to the coarse classification of the existing dataset, a new dataset named LSCIDMR-V2 is generated with fine-grained cloud-type annotation and multichannel data. The experiment shows that, using the group feature extractor and geo-queries, the popular used metric subset accuracy is improved from 40.06 to 42.87 and 44.35, respectively. The proposed method achieves the mean average precision of 82.40, outperforming state-of-the-art methods.
Dongxiaoyuan Zhao, Jinglin Zhang 0001, Cong Bai
IEEE Trans. Geosci. Remote. Sens.4
2023 Community aware graph embedding learning for item recommendation
Pengyi Hao, Zhaojie Qian, Shuang Wang 0016, Cong Bai
World Wide Web (WWW)4
2022 ISDA: Position-Aware Instance Segmentation with Deformable Attention
abstract
Most instance segmentation models are not end-to-end trainable due to either the incorporation of proposal estimation (RPN) as a pre-processing or non-maximum suppression (NMS) as a post-processing. Here we propose a novel end-to-end instance segmentation method termed ISDA. It reshapes the task into predicting a set of object masks, which are generated via traditional convolution operation with learned position-aware kernels and features of objects. Such kernels and features are learned by leveraging a deformable attention network with multi-scale representation. Thanks to the introduced set-prediction mechanism, the proposed method is NMS-free. Empirically, ISDA outperforms Mask R-CNN (the strong baseline) by 2.6 points on MS-COCO, and achieves leading performance compared with recent models. Code will be available soon.
Kaining Ying, Zhenhua Wang 0003, Cong Bai
ICASSP3
2022 Intra-Modal Constraint Loss for Image-Text Retrieval
abstract
Cross-modal retrieval has drawn much attention in both computer vision and natural language processing domains. With the development of convolutional and recurrent neural networks, the bottleneck of retrieval across image-text modalities is no longer the extraction of image and text features but an efficient loss function learning in embedding space. Many loss functions try to closer pairwise features from heterogeneous modalities. This paper proposes a method for learning joint embedding of images and texts using an intra-modal constraint loss function to reduce the violation of negative pairs from the same homogeneous modality. Experimental results show that our approach outperforms state-of-the-art bi-directional image-text retrieval methods on Flickr30K and Microsoft COCO datasets. Our code is publicly available1
Jianan Chen 0002, Lu Zhang 0037, Cong Bai, Kidiyo Kpalma
ICIP4
2022 MMINR: Multi-frame-to-Multi-frame Inference with Noise Resistance for Precipitation Nowcasting with Radar
abstract
Precipitation nowcasting based on radar echo maps is essential in meteorological research. Recently, Convolutional RNNs based methods dominate this field, but they cannot be solved by parallel computation resulting in longer inference time. FCN based methods adopt a multi-frame-to-single-frame inference (MSI) strategy to avoid this problem. They feedback into the model again to predict the next time step to get multi-frame nowcasting results in the prediction phase, which will lead to the accumulation of prediction errors. In addition, precipitation noise is a crucial factor contributing to high prediction errors because of its unpredictability. To address this problem, we propose a novel Multi-frame-to-Multi-frame Inference (MMI) model with Noise Resistance (NR) named MMINR. It avoids error accumulation and resists precipitation noiseś negative effect in parallel computation. NR contains a Noise Dropout Module (NDM) and a Semantic Restore Module (SRM). NDM deliberately dropout noise simple yet efficient, and SRM supplements semantic information of features to alleviate the problem of semantic information mistakenly lost by NDM. Experimental results demonstrate that MMINR can attain competitive scores compared with other SOTAs. The ablation experiments show that the proposed NDM and SRM can solve the aforementioned problems.
Cong Bai
ICPR2
2022 Structural and Temporal Learning for Dropout Prediction in MOOCs
Tianxing Han, Pengyi Hao, Cong Bai
KSEM (2)3
2022 Structure-Inferred Bi-level Model for Underwater Image Enhancement
abstract
Very recently, with the development of underwater robots, underwater image enhancement arising growing interests in the computer vision community. However, owing to light being scattered and absorbed while it traveling in water, underwater captured images often suffer from color cast and low visibility. Existing methods depend on specific prior knowledge and training data to enhance underwater images in the absence of structure information, which results in poor and unnatural performance. To this end, we propose a Structural-Inferred Bi-level Model (SIBM) that incorporates different modalities of knowledge (i.e., semantic domain, gradient-domain, and pixel domain) hierarchically enhancing underwater images. In particular, by introducing a semantic mask, we individually optimize the forehand branch that avoids unnecessary interference arising from the background region. We design a gradient-based high-frequency branch to exploit gradient-space guidance for preserving texture structures. Moreover, we construct a pixel-based branch by feeding semantic and gradient information to enhance underwater images. To exploit different modalities, we introduce a hyper-parameter optimization scheme to fuse the above domain information. Experimental results illustrate that the developed method not only outperforms the previous methods in quantitative scores but also generalizes well on real-world underwater datasets. Source code is available at \hrefhttps://github.com/IntegralCoCo/SIBM https://github.com/IntegralCoCo/SIBM.
Pan Mu, Haotian Qian, Cong Bai
ACM Multimedia3
2022 Exploring Implicit and Explicit Relations with the Dual Relation-Aware Network for Image Captioning
Zhiwei Zha, Cong Bai
MMM (2)3
2022 Edge-preserving Image Smoothing via Counting-weighted Total Variation
abstract
We present a new counting-weighted total variation measure, which captures the consistency of gradient directions in local regions to distinguish edges and details. A novel optimization framework with the proposed counting-weighted total variation in the l1regularization term is then developed to realize the edge-preserving image smoothing. In order to solve the optimization problem, we adopt an iteratively re-weighted least square based algorithm. Experimental results demonstrate that the proposed method is capable of completing edge-preserving image smoothing, while avoiding blurring and over-sharpening the edge. The proposed method can also be used for the image detail enhancement without involving halos or gradient reversal artifacts, while achieve better quality scores in the comparison with other enhancement methods.
Jiachao Dang, Yi Liu 0004, Wenjing Shuai, Cong Bai, Shishun Tian
MMSP4
2022 Deep Correlation based Concept Recommendation for MOOCs
abstract
The current course recommendation in massive open online courses (MOOCs) usually ignores students' interests in some certain type of knowledge concepts, resulting in low completion of most courses.Therefore, it requires a concept recommendation to help students accurately choose courses in MOOCs.In this paper, we propose Deep Correlation based Concept Recommendation (DCCR) for MOOCs.It gathers the interactive information obtained by different entities through meta-paths in MOOCs and extracts the semantic information of concepts.To deeply capture the correlation information among users, a multi-relation graph is built to generate the correlation features which aggregates the abundant information under different meta-paths.Then through the graph convolutional neural networks, entity embeddings of users and knowledge concepts are generated.Additionally, a concatenation-based fusion function is designed to get the final joint representations reasonably.By verifying on two public datasets, experiments show that DCCR outperforms the state-of-the-art methods.
Shengyu Mao, Pengyi Hao, Cong Bai
SEKE3
2022 Automatic and accurate segmentation of peripherally inserted central catheter (PICC) from chest X-rays using multi-stage attention-guided learning
Xiaoyan Wang 0007, Ye Sheng, Chenglu Zhu, Cong Bai, Ming Xia 0005, Zhanpeng Shao, Ruiyi Zhao, Zhenjie Liu
Neurocomputing6
2022 An Effective Federated Learning Verification Strategy and Its Applications for Fault Diagnosis in Industrial IoT Systems
abstract
Due to the diverse equipment and uneven load distribution in industrial environments, data regarding faults are often unbalanced. Moreover, data and models from clients may become contaminated or damaged, affecting diagnostic performance. To overcome these problems, this study proposes a stacking model for diagnosing interturn short circuit (ITSC) faults in permanent magnet synchronous motors (PMSMs). Federated learning (FL) is used to train the model to increase data security and overcome data islanding in distributed scenarios. Moreover, an improved verification strategy was adopted to select appropriate client models in each round to update the FL global model. We created a secondary server-side data set to validate the client weightings. The data set contains clean sample data for all ITSC fault categories. By calculating the fault diagnosis accuracy of the global model on the auxiliary data set, the model eliminates low-quality clients with uneven fault distributions. The improved particle swarm optimization (PSO) is used to optimize the weight coefficients of clients involved in aggregation, improving the robustness of the aggregation strategy under a joint learning system. In evaluation experiments, compared with the federated average (FedAvg) model, the proposed dynamic verification model exhibited the better diagnostic accuracy in situations of data imbalance, incurred lower communication costs, and prevented local oscillations in the model.
Yuanjiang Li, Kai Zhu 0005, Cong Bai, Jinglin Zhang 0001
IEEE Internet Things J.4
2022 CMRDF: A Real-Time Food Alerting System Based on Multimodal Data
abstract
A healthy diet is a major concern for everyone, especially for those with specific diseases, such as diabetes. Meanwhile, with the rapid development of new technologies, it is feasible for us to detect the deep latent relationship between daily meals and wellbeing. Advanced Internet of Things devices, such as smart bracelets and wearable cameras, make it possible for people to know how food is related to health at any time. However, it is still arduous for individuals to memorize all the health information and utilize them to regulate their diet. To deal with such problems, we propose a novel system called cross-modal retrieval on diabetogenic food (CMRDF) which realizes a real-time dietary notice based on multimodal data captured from wearable devices. In this system, we propose a new graph-based cross-modal retrieval method named graph correlation analysis with ranking loss that finds the latent information in multimodal data. We use graph convolutional networks to dig the deep latent information in modalities and represent the data in finer granularity. It uses visual and physiological information to estimate whether the food that a user tries to obtain is diabetogenic or not, and feeds back the reasons in detail. Extensive experiments on the MSCOCO data set and the new proposed multimodal diabetogenic food database real-life diabetogenic show that the proposed cross-modal retrieval method outperforms state-of-the-art methods and CMRDF can achieve reliable results on preventing diabetic patients from inappropriate food.
Cong Bai, Shengyong Chen
IEEE Internet Things J.2
2022 Rainformer: Features Extraction Balanced Network for Radar-Based Precipitation Nowcasting
abstract
Precipitation nowcasting is one of the fundamental challenges in natural hazard research. High-intensity rainfall, especially the rainstorm, will lead to the enormous loss of people’s property. Existing methods usually utilize convolution operation to extract rainfall features and increase the network depth to expand the receptive field to obtain fake global features. Although this scheme is simple, only local rainfall features can be extracted leading to insensitivity to high-intensity rainfall. This letter proposes a novel precipitation nowcasting framework named Rainformer, in which, two practical components are proposed: the global features extraction unit and the gate fusion unit (GFU). The former provides robust global features learning ability depending on the window-based multi-head self-attention (W-MSA) mechanism, while the latter provides a balanced fusion of local and global features. Rainformer has a simple yet efficient architecture and significantly improves the accuracy of rainfall prediction, especially on high-intensity rainfall. It offers a potential solution for real-world applications. The experimental results show that Rainformer outperforms seven state of the arts methods on the benchmark database and provides more insights into the high-intensity rainfall prediction task.
Cong Bai, Jinglin Zhang 0001, Shengyong Chen
IEEE Geosci. Remote. Sens. Lett.1
2022 SSA-Net: Spatial self-attention network for COVID-19 pneumonia infection segmentation with semi-supervised few-shot learning
Xiaoyan Wang 0007, Yiwen Yuan, Dongyan Guo, Ming Xia 0005, Zhenhua Wang 0003, Cong Bai, Shengyong Chen
Medical Image Anal.8
2022 Unsupervised adversarial image retrieval
Ling Huang 0003, Cong Bai, Yijuan Lu, Shaobo Zhang 0005, Shengyong Chen
Multim. Syst.2
2022 Res2-UNeXt: a novel deep learning framework for few-shot cell image segmentation
Sixian Chan 0001, Cong Bai, Shengyong Chen
Multim. Tools Appl.3
2022 Online multiple object tracking using joint detection and embedding network
Sixian Chan 0001, Yangwei Jia, Xiaolong Zhou 0001, Cong Bai, Shengyong Chen, Xiaoqin Zhang 0002
Pattern Recognit.4
2022 LSCIDMR: Large-Scale Satellite Cloud Image Database for Meteorological Research
abstract
People can infer the weather from clouds. Various weather phenomena are linked inextricably to clouds, which can be observed by meteorological satellites. Thus, cloud images obtained by meteorological satellites can be used to identify different weather phenomena to provide meteorological status and future projections. How to classify and recognize cloud images automatically, especially with deep learning, is an interesting topic. Generally speaking, large-scale training data are essential for deep learning. However, there is no such cloud images database to date. Thus, we propose a large-scale cloud image database for meteorological research (LSCIDMR). To the best of our knowledge, it is the first publicly available satellite cloud image benchmark database for meteorological research, in which weather systems are linked directly with the cloud images. LSCIDMR contains 104 390 high-resolution images, covering 11 classes with two different annotation methods: 1) single-label annotation and 2) multiple-label annotation, called LSCIDMR-S and LSCIDMR-M, respectively. The labels are annotated manually, and we obtain a total of 414 221 multiple labels and 40 625 single labels. Several representative deep learning methods are evaluated on the proposed LSCIDMR, and the results can serve as useful baselines for future research. Furthermore, experimental results demonstrate that it is possible to learn effective deep learning models from a sufficiently large image database for the cloud image classification.
Cong Bai, Minjing Zhang, Jinglin Zhang 0001, Jianwei Zheng 0001, Shengyong Chen
IEEE Trans. Cybern.1
2022 Siamese Implicit Region Proposal Network With Compound Attention for Visual Tracking
abstract
Recently, siamese-based trackers have achieved significant successes. However, those trackers are restricted by the difficulty of learning consistent feature representation with the object. To address the above challenge, this paper proposes a novel siamese implicit region proposal network with compound attention for visual tracking. First, an implicit region proposal (IRP) module is designed by combining a novel pixel-wise correlation method. This module can aggregate feature information of different regions that are similar to the pre-defined anchor boxes in Region Proposal Network. To this end, the adaptive feature receptive fields then can be obtained by linear fusion of features from different regions. Second, a compound attention module including a channel and non-local attention is raised to assist the IRP module to perform a better perception of the scale and shape of the object. The channel attention is applied for mining the discriminative information of the object to handle the background clutters of the template, while non-local attention is trained to aggregate the contextual information to learn the semantic range of the object. Finally, experimental results demonstrate that the proposed tracker achieves state-of-the-art performance on six challenging benchmark tests, including VOT-2018, VOT-2019, OTB-100, GOT-10k, LaSOT, and TrackingNet. Further, our obtained results demonstrate that the proposed approach can be run at an average speed of 72 FPS in real time.
Sixian Chan 0001, Xiaolong Zhou 0001, Cong Bai, Xiaoqin Zhang 0002
IEEE Trans. Image Process.4
2021 Multi-scale Hierarchical Transformer structure for 3D medical image segmentation
abstract
Transformers have demonstrated great potential in computer vision tasks which is benefited from its ability of long-range dependency modeling. However, its core material multi-head self-attention (MHSA) has high computational and spacial complexity. Besides, tokens embedded with 1D position sequence are unable to represent inductive bias of locality and neighbor contextual information, which are essential for high-resolution downstream vision tasks. Thus, transformer is hard to use in image segmentation tasks, especially in 3D medical images. In this paper, we propose a novel multi-scale hierarchical framework (MSHT) that efficiently incorporates the local modeling ability of CNN and long-range modeling ability of transformer for accurate 3D medical image segmentation. Unlike many prior Transformer-based solutions, the proposed MSHT first adopts parallel CNN and transformer as encoder block to extract the global and local feature representations. As the core component for our MSHT, a position embedding method for 3D medical image (3D IPE) is proposed to generate relational and absolute position encoding for tokens fed in transformer block. We conduct an extensive evaluation on both Lits2017 dataset and Kits2019 dataset. The results indicate that our MSHT leads to a substantial performance over other CNN-based methods on 3D image segmentation.
Xiaoyan Wang 0007, Bangze Zhang, Cong Bai, Ming Xia 0005, Peiliang Sun
BIBM5
2021 Classmates Enhanced Diversity-Self-Attention Network for Dropout Prediction in MOOCs
Dongen Wu, Pengyi Hao, Tianxing Han, Cong Bai
ICONIP (4)5
2021 Community Enhanced Course Concept Recommendation in MOOCs with Multiple Entities
Binglong Ye, Shengyu Mao, Pengyi Hao, Wei Chen 0001, Cong Bai
KSEM5
2021 Tropical Cyclones Tracking Based on Satellite Cloud Images: Database and Comprehensive Study
Sixian Chan 0001, Cong Bai
MMM (2)3
2021 Box Regression-Guided Anchor-free for Robust Visual Tracking
abstract
The Siamese tracker-based approach has achieved significant success in recent years. However, these approaches do not consider the different requirements for input feature in classification and regression branches. The regression branch needs feature information slightly larger than the object region, while the classification branch needs to avoid classification failure caused by the introduction of background information. In this paper, we present a novel Box Regression-Guided Anchor-free for Robust Visual Tracking. Firstly, a scale-aware regression module is designed to satisfy the feature requirements of the regression branch, which can capture feature information of various scales. Secondly, regression-guided classification module is applied to aligning the feature between the regression result and correlation feature, thereby avoiding the introduction of background information to classification branch. In addition, the new correlation operation is introduced to gain more superb correlation feature. Comparsion experimental exhibits that the proposed tracker achieves promising results in five challenging benchmark tests, including GOT-10K, OTB-2015, VOT-2018, VOT-2019 and TrackingNet, and run at an average speed of 60 FPS in real-time.
Sixian Chan 0001, Xiaolong Zhou 0001, Cong Bai, Hua Gao, Shengyong Chen
SMC4
2021 PaI-Net: A modified U-Net of reducing semantic gap for surgical instrument segmentation
abstract
Abstract Tracking the instruments in a surgical scene is an essential task in minimally invasive surgery. However, due to the unpredictability of scenes, automatically segmenting the instruments is very challenging. In this paper, a novel method named parallel inception network (PaI‐Net) is proposed, in which an attention parallel module (APM) and an output fusion module (OFM) are integrated with U‐Net to improve the segmentation ability. Specially, APM utilizes multi‐scale convolution kernels and global average pooling operations to extract semantic information and global context information of different scales, while OFM combines the feature maps of the decoder part to aggregate the abundant boundary information of shallow layers and the rich semantic information of deep layers together, which achieve a significant improvement in generating segmentation masks. Finally, the evaluation of proposed method on robotic instruments segmentation task from Medical Image Computing and Computer Assisted Intervention Society (MICCAI) and retinal image segmentation task from International Symposium on Biomedical Imaging (ISBI) show that our model has achieved advanced performance on multi‐scale semantic segmentation and is superior to the current state‐of‐the‐art models.
Xiaoyan Wang 0007, Xingyu Zhong, Cong Bai, Ruiyi Zhao, Ming Xia 0005
IET Image Process.4
2021 Hyperspectral Image Classification Using Mixed Convolutions and Covariance Pooling
abstract
Recently, convolution neural network (CNN)-based hyperspectral image (HSI) classification has enjoyed high popularity due to its appealing performance. However, using 2-D or 3-D convolution in a standalone mode may be suboptimal in real applications. On the one hand, the 2-D convolution overlooks the spectral information in extracting feature maps. On the other hand, the 3-D convolution suffers from heavy computation in practice and seems to perform poorly in scenarios having analogous textures along with consecutive spectral bands. To solve these problems, we propose a mixed CNN with covariance pooling for HSI classification. Specifically, our network architecture starts with spectral-spatial 3-D convolutions that followed by a spatial 2-D convolution. Through this mixture operation, we fuse the feature maps generated by 3-D convolutions along the spectral bands for providing complementary information and reducing the dimension of channels. In addition, the covariance pooling technique is adopted to fully extract the second-order information from spectral-spatial feature maps. Motivated by the channel-wise attention mechanism, we further propose two principal component analysis (PCA)-involved strategies, channel-wise shift and channel-wise weighting, to highlight the importance of different spectral bands and recalibrate channel-wise feature response, which can effectively improve the classification accuracy and stability, especially in the case of limited sample size. To verify the effectiveness of the proposed model, we conduct classification experiments on three well-known HSI data sets, Indian Pines, University of Pavia, and Salinas Scene. The experimental results show that our proposal, although with less parameters, achieves better accuracy than other state-of-the-art methods.
Jianwei Zheng 0001, Yuchao Feng, Cong Bai, Jinglin Zhang 0003
IEEE Trans. Geosci. Remote. Sens.3
2021 Unsupervised Adversarial Instance-Level Image Retrieval
abstract
With the wide use of visual sensors in the Internet of Things (IoT) in the past decades, huge amounts of images are captured in people's daily lives, which poses challenges to traditional deep-learning-based image retrieval frameworks. Most such frameworks need a large amount of annotated training data, which are expensive. Moreover, machines still lack human intelligence, as illustrated by the fact that they pay less attention to the interesting regions that humans generally focus on when searching for images. Hence, this paper proposes a novel unsupervised framework that focuses on the instance object in the image and integrates human intelligence into the deep-learning-based image retrieval. This framework is called adversarial instance-level image retrieval (AILIR). We incorporate adversarial training and an attention mechanism into this framework that considers human intelligence with artificial intelligence. The generator and discriminator are redesigned to guarantee that the generator retrieves similar images while the discriminator selects unmatched images and creates an adversarial reward for the generator. A minimax game is conducted by the adversarial reward retrieval mechanism until the discriminator is unable to judge whether the image sequence retrieved matches the query. Comparison and ablation experiments on four benchmark datasets prove that the proposed adversarial training framework indeed improves instance retrieval and outperforms the state-of-the-art methods focused on instance retrieval.
Cong Bai, Jinglin Zhang 0003, Ling Huang 0003, Lu Zhang 0037
IEEE Trans. Multim.1
2020 Co-Saliency Detection Using Collaborative Feature Extraction And High-To-Low Feature Integration
abstract
Co-saliency detection, as a developing research branch of saliency detection, devotes to identify the common salient objects in a group of related images. The major challenge of co-saliency detection is how to effectively represent features considering both intra-image and inter-image information. In this paper, we propose a co-saliency detection model using collaborative feature extraction and high-to-low feature integration. We first feed the target image and its co-images into the Individual Feature Extraction Module (IFEM) to produce multi-level individual features. Then, to capture the collaborative inter-image information, the Collaborative Feature Extraction Module (CFEM) is applied to all highest-level individual features, generating the collaborative feature. Finally, we build a High-to-low Feature Integration Module (HFIM), which integrates the collaborative feature and multi-level individual features of the target image, to enrich the collaborative feature with individual intra-image information. Extensive experiments on two public datasets demonstrate that the proposed model achieves the state-of-the-art performance.
Jingru Ren, Zhi Liu 0003, Gongyang Li, Xiaofei Zhou 0003, Cong Bai, Guangling Sun
ICME5
2020 RWMF: A Real-World Multimodal Foodlog Database
abstract
With the increasing health concerns on diet, it's worthwhile to develop an intelligent assistant that can help users eat healthier. Such an assistant can automatically give personal advice for the user's diet and generate health report about eating on a regular basis. To boost the research on such diet assistant, we establish a real-world foodlog database using various methods such as filter, cluster and graph convolutional network. This database is built based on real-world lifelog and medical data, which is named as Real-World Multimodal Foodlog (RWMF). It contains 7500 multimodal pairs, and each pair consists of a food image paired with a line of personal biometrics data (such as Blood Glucose) and a textual food description of food composition paired with a line of food nutrition data. In this paper, we present the detailed procedures for setting up the database. We evaluate the performance of RWMF using different food image classification and cross-modal retrieval approaches. We also test the performance of multimodal fusion on RWMF through ablation experiments. The experimental results show that the RWMF database is quite challenging and can be widely used to evaluate the performance of food analysis methods based on multimodal data.
Cong Bai, Kaining Ying, Lixin Huang
ICPR2
2020 Deep Adversarial Discrete Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing has received widespread attentions on cross-modal retrieval task due to its superior retrieval efficiency and low storage cost. However, most existing cross-modal hashing methods learn binary codes directly from multimedia data, which cannot fully utilize the semantic knowledge of the data. Furthermore, they cannot learn the ranking based similarity relevance of data points with multi-label. And they usually use a relax constraint of hash code which causes non-negligible quantization loss in the optimization. In this paper, a hashing method called Deep Adversarial Discrete Hashing (DADH) is proposed to address these issues for cross-modal retrieval. The proposed method uses adversarial training to learn features across modalities and ensure the distribution consistency of feature representations across modalities. We also introduce a weighted cosine triplet constraint which can make full use of semantic knowledge from the multi-label to ensure the precise ranking relevance of item pairs. In addition, we use a discrete hashing strategy to learn the discrete binary codes without relaxation, by which the semantic knowledge from label in the hash codes can be preserved while the quantization loss can be minimized. Ablation experiments and comparison experiments on two cross-modal databases show that the proposed DADH improves the performance and outperforms several state-of-the-art hashing methods for cross-modal retrieval.
Cong Bai, Jinglin Zhang 0001, Shengyong Chen
ICMR1
2020 Overlap classification mechanism for skeletal bone age assessment
abstract
The bone development is a continuous process, however, discrete labels are usually used to represent bone ages. This inevitably causes a semantic gap between actual situation and label representation scope. In this paper, we present a novel method named as overlap classification network to narrow the semantic gap in bone age assessment. In the proposed network, discrete bone age labels (such as 0-228 month) are considered as a sequence that is used to generate a series of subsequences. Then the proposed network makes use of the overlapping information between adjacent subsequences and output several bone age ranges at the same time for one case. The overlapping part of these age ranges is considered as the final predicted bone age. The proposed method without any preprocessing can achieve a much smaller mean absolute error compared with state-of-the-art methods on a public dataset.
Pengyi Hao, Xuhang Xie, Tianxing Han, Cong Bai
MMAsia4
2020 Instance Image Retrieval with Generative Adversarial Training
Cong Bai, Ling Huang 0003, Yu-Gang Jiang 0001, Shengyong Chen
MMM (1)2
2020 Co-saliency detection via integration of multi-layer convolutional features and inter-image propagation
Jingru Ren, Zhi Liu 0003, Xiaofei Zhou 0003, Cong Bai, Guangling Sun
Neurocomputing4
2020 Cross-domain representation learning by domain-migration generative adversarial network for sketch based image retrieval
Cong Bai, Jian Chen 0009, Pengyi Hao, Shengyong Chen
J. Vis. Commun. Image Represent.1
2019 End-to-End Panoptic Segmentation with Pixel-Level Non-Overlapping Embedding
abstract
Recent panoptic segmentation even instance segmentation methods usually rely on the region-based method or highly-specialized combination with heuristics module, followed by post-processing techniques. While most of the recent methods neglect low-fill rate linear objects and cannot recognize pixels located in bounding box margins. We propose a branched, end-to-end trainable multi-task architecture focusing on pixel-level grouping problems for panoptic segmentation. The embedding branch regress pixels into an embedding space, so that pixels from the same group are at close range while those from different groups have a specified margin. Every pixel can be considered in an image without overlapping. And semantic branch produces best seed scores with labels as clustering center. The further-embedding branch disentangles each pixel in pixel embedding space. Thus, we are able to segment both thing and stuff classes, and explain all the pixels in the image. We obtain state-of-the-art results on Pascal VOC2012 and Cityscapes.
Qieshi Zhang, Jun Cheng 0002, Cong Bai, Pengyi Hao
ICME4
2019 Embodied One-Shot Video Recognition: Learning from Actions of a Virtual Embodied Agent
abstract
One-shot learning aims to recognize novel target classes from few examples by transferring knowledge from source classes, under a general assumption that the source and target classes are semantically related but not exactly the same. Based on this assumption, recent work has focused on image-based one-shot learning, while little work has addressed video-based one shot learning. One of the challenges lies in that it is difficult to maintain the disjoint-class assumption for videos, since video clips of target classes may potentially appear in the videos of source classes. To address this issue, we introduce a novel setting, termed as embodied agents based one-shot learning, which leverages synthetic videos produced in a virtual environment to understand realistic videos of target classes. In this setting, we further propose two types of learning tasks: embodied one-shot video domain adaptation and embodied one-shot video transfer recognition. These tasks serve as a testbed for evaluating video related one-shot learning tasks. In addition, we propose a general video segment augmentation method, which significantly facilitates a variety of one-shot learning tasks. Experimental results validate the soundness of our setting and learning tasks, and also show the effectiveness of our augmentation approach to video recognition in the small-sample size regime.
Yuqian Fu, Chengrong Wang, Yanwei Fu 0001, Yu-Xiong Wang, Cong Bai, Xiangyang Xue 0001, Yu-Gang Jiang 0001
ACM Multimedia5
2019 Session details: Poster Session
abstract
No abstract available.
Cong Bai
MMAsia1
2019 Video Summarization based on Sparse Subspace Clustering with Automatically Estimated Number of Clusters
abstract
Advancements in technology resulted in a sharp growth in the number of digital cameras at people's disposal all across the world. Consequently, the huge storage space consumed by the videos from these devices on video repositories make the job of video processing and analysis to be time-consuming. Furthermore, this also slows down the video browsing and retrieval. Video summarization plays a very crucial role in solving these issues. Despite the number of video summarization approaches proposed up to the present time, the goal is to take a long video and generate a video summary in form of a short video skim without losing the meaning or the message transmitted by the original lengthy video. This is done by selecting the important frames called key-frames. The approach proposed by this work performs automatic summarization of digital videos based on detected objects' deep features. To this end, we apply sparse subspace clustering with an automatically estimated number of clusters to the objects' deep features. The summary generated from our scheme will store the meta-data for each short video inferred from the clustering results. In this paper, we also suggest a new video dataset for video summarization. We evaluate the performance of our work using the TVSum dataset and our video summarization dataset.
Pengyi Hao, Edwin Manhando, Taotao Ye, Cong Bai
MMAsia4
2019 Supervised learning based discrete hashing for image retrieval
Cong Bai, Jinglin Zhang 0001, Zhi Liu 0003, Shengyong Chen
Pattern Recognit.2
2019 Weighted Mixed-Norm Regularized Regression for Robust Face Identification
abstract
Face identification (FI) via regression-based classification has been extensively studied during the recent years. Most vector-based methods achieve appealing performance in handing the noncontiguous pixelwise noises, while some matrix-based regression methods show great potential in dealing with contiguous imagewise noises. However, there is a lack of consideration of the mixture noises case, where both contiguous and noncontiguous noises are jointly contained. In this paper, we propose a weighted mixed-norm regression (WMNR) method to cope with the mixture image corruption. WMNR reveals certain essential characteristics of FI problems and bridges the vector- and matrix-based methods. Particularly, WMNR provides two advantages for both theoretical analysis and practical implementation. First, it generalizes possible distributions of the residuals into a unified feature weighted loss function. Second, it constrains the residual image as low-rank structure that can be quantified with general nonconvex functions and a weight factor. Moreover, a new reweighted alternating direction method of multipliers algorithm is derived for the proposed WMNR model. The algorithm exhibits great computational efficiency since it divides the original optimization problem into certain subproblems with analytical solution or can be implemented in a parallel manner. Extensive experiments on several public face databases demonstrate the advantages of WMNR over the state-of-the-art regression-based approaches. More specifically, the WMNR achieves an appealing tradeoff between identification accuracy and computational efficiency. Compared with the pure vector-based methods, our approach achieves more than 10% performance improvement and saves more than 70% of runtime, especially in severe corruption scenarios. Compared with the pure matrix-based methods, although it requires slightly more computation time, the performance benefits are even larger; up to 20% improvement can be obtained.
Jianwei Zheng 0001, Kechen Lou, Xi Yang 0006, Cong Bai, Jinhui Tang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2018 Optimization of deep convolutional neural network for large scale image retrieval
Cong Bai, Ling Huang 0003, Jianwei Zheng 0001, Shengyong Chen
Neurocomputing1
2018 Saliency-based multi-feature modeling for semantic image retrieval
Cong Bai, Jianan Chen 0002, Ling Huang 0003, Kidiyo Kpalma, Shengyong Chen
J. Vis. Commun. Image Represent.1
2018 Saliency integration driven by similar images
Jingru Ren, Zhi Liu 0003, Xiaofei Zhou 0003, Guangling Sun, Cong Bai
J. Vis. Commun. Image Represent.5
2017 Salient Object Segmentation via Effective Integration of Saliency and Objectness
abstract
This paper proposes an effective salient object segmentation method via the graph-based integration of saliency and objectness. Based on the superpixel segmentation result of the input image, a graph is built to represent superpixels using regular vertex, background seed vertex with the addition of a terminal vertex. The edge weights on the graph are defined by integrating the difference of appearance, saliency, and objectness between superpixels. Then, the object probability of each superpixel is measured by finding the shortest path from the corresponding vertex to the terminal vertex on the graph, and the resultant object probability map can generally better highlight salient objects and suppress background regions compared to both saliency map and objectness map. Finally, the object probability map is used to initialize salient object and background, and effectively incorporated into the framework of graph cut to obtain the final salient object segmentation result. Extensive experimental results on three public benchmark datasets show that the proposed method consistently improves the salient object segmentation performance and outperforms the state-of-the-art salient object segmentation methods. Furthermore, experimental results also demonstrate that the proposed graph-based integration method is more effective than other fusion schemes and robust to saliency maps generated using various saliency models.
Linwei Ye, Zhi Liu 0003, Liquan Shen, Cong Bai, Yang Wang 0003
IEEE Trans. Multim.5
2015 K-means based histogram using multiresolution feature vectors for color texture database retrieval
Cong Bai, Jinglin Zhang 0001, Zhi Liu 0003, Wanlei Zhao
Multim. Tools Appl.1
2014 Online Glocal Transfer for Automatic Figure-Ground Segmentation
abstract
This paper addresses the problem of automatic figure-ground segmentation, which aims at automatically segmenting out all foreground objects from background. The underlying idea of this approach is to transfer segmentation masks of globally and locally (glocally) similar exemplars into the query image. For this purpose, we propose a novel high-level image representation method named as object-oriented descriptor. Using this descriptor, a set of exemplar images glocally similar to the query image is retrieved. Then, using over-segmented regions of these retrieved exemplars, a discriminative classifier is learned on-the-fly and subsequently used to predict foreground probability for the query image. Finally, the optimal segmentation is obtained by combining the online prediction with typical energy optimization of Markov random field. The proposed approach has been extensively evaluated on three datasets, including Pascal VOC 2010, VOC 2011 segmentation challenges, and iCoseg dataset. Experiments show that the proposed approach outperforms state-of-the-art methods and has the potential to segment large-scale images containing unknown objects, which never appear in the exemplar images.
Wenbin Zou, Cong Bai, Kidiyo Kpalma, Joseph Ronsin
IEEE Trans. Image Process.2
2013 Multi-object tracking using sparse representation
abstract
Recently sparse representation has been successfully applied to single object tracking by observing the reconstruction error of candidate object with sparse representation. In practice, sparse representation also shows competitive performance on multi-class classification, and thus is potential for multi-object tracking. In this paper we explore this technique for on-line multi-object tracking through a simple tracking-by-detection scheme, with background subtraction for object detection and sparse representation for object recognition. Final experiments demonstrate that the proposed approach only combining color histogram and 2-dimensional coordinates as features, achieves favorable performance over state-of-the-art work in persistent identity tracking.
Weizhi Lu, Cong Bai, Kidiyo Kpalma, Joseph Ronsin
ICASSP2