EDBT 2026 Demo / reviewers in the wild / expert
Shichao Kan
dblp:234/2854 · also Shi-Chao Kan
· DBLP profile ↗
52ranked-venue papers
10as first author
46since 2021 · last 2026
0000-0003-0097-6196ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 7 first-author · 23 since 2021Artificial intelligence and machine learning · 22 · 4 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAPL: Enhancing Visual Prompt Encoding for Robust Open-Set Blood Cell Detection
Wenzhuo Xu, Jianfeng Liu 0001, Shichao Kan, Yixiong Liang |
ICIC (29) | 4 |
| 2026 | Leveraging artificial intelligence in advance care planning: A scoping review
Minghui Tan, Zhao Ni, Shichao Kan, Paul Macharia, Jinfeng Ding |
Artif. Intell. Medicine | 4 |
| 2026 | Multi-level contrastive learning with graph convolutional network for multi-view clustering
Jie Wang 0067, Haiwei Deng, Shichao Kan |
Expert Syst. Appl. | 3 |
| 2026 | Query-guided predicate decoupling and prototype approximation learning for scene graph generation
Shichao Kan, Yue Zhang 0065, Yi-Gang Cen, Wanru Xu, Yi Jin 0001, Yidong Li |
Expert Syst. Appl. | 2 |
| 2026 | Integrating spatial features and dynamically learned temporal features via contrastive learning for video temporal grounding in LLM
Peifu Wang, Yixiong Liang, Yi-Gang Cen, Jin Liu 0012, Shichao Kan |
Image Vis. Comput. | 7 |
| 2026 | Progressively multi-scale feature fusion for semantic segmentation
Shichao Kan, Yi-Gang Cen, Qi Cao 0002, Yansen Huang, Ming Zeng 0012 |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | Modality-incomplete Federated Learning via personalized prompt generation and alignment
Shichao Kan |
Pattern Recognit. | 4 |
| 2026 | Reasoning elicitation and multi-granularity contrastive learning for text-rich image understanding in large vision-language models
Jiazhi Xia, Bingchuan Jiang, Shichao Kan |
Pattern Recognit. | 4 |
| 2026 | Vision-Semantics-Label: A New Two-Step Paradigm for Action Recognition With Large Language ModelabstractIn recent years, the rapid advancement of multi-modal large language models has propelled the development of video-based conversation models. Due to their exceptional video understanding capabilities, there is often an expectation that these models can handle all video-related tasks, including action recognition. However, because action recognition datasets typically lack semantic information, limiting the performance of dialogue models. Additionally, as these dialogue models are designed for video understanding, they frequently overlook critical information required for action recognition—continuous motion—in their model architecture and training dataset configurations. To address these challenges, we first propose a novel two-step mapping framework based on large language models, termed “Vision-Semantics-Label” mapping, to better adapt video-based large language models for action recognition. In the first step, we proposed a visual-skeletal collaborative learning large language model (VS-LLM), which utilizes human keypoints to compensate for the missing motion details without increasing the input token length of the large language model. In the second step, we designed two mapping methods: verb noun match (VN-Match) and all text match (ALL-Match), which can effectively extract relevant action descriptions from the text. Finally, we construct semantic action recognition datasets to ensure that the training data inherently contains action details, enabling the model to better achieve action recognition. We evaluate our approach on five benchmark datasets, demonstrating the state-of-the-art performance of large language models in action recognition. The source code and dataset are publicly available at https://github.com/xiaoyu92568/VS-LLM. Wanru Xu, Shichao Kan, Linna Zhang, Yi Jin 0001, Yi-Gang Cen, Yidong Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | How Does the Smoothness Approximation Method Facilitate Generalization for Federated Adversarial Learning?abstractFederated Adversarial Learning (FAL) is a robust framework for resisting adversarial attacks on federated learning. Although some FAL studies have developed efficient algorithms, they primarily focus on convergence performance and overlook generalization. Generalization is crucial for evaluating algorithm performance on unseen data. However, generalization analysis is more challenging due to non-smooth adversarial loss functions. A common approach to addressing this issue is to leverage smoothness approximation. In this paper, we develop algorithm stability measures to evaluate the generalization performance of two popular FAL algorithms: Vanilla FAL (VFAL) and Slack FAL (SFAL), using three different smooth approximation methods: 1) Surrogate Smoothness Approximation (SSA), (2) Randomized Smoothness Approximation (RSA), and (3) Over-Parameterized Smoothness Approximation (OPSA). Based on our in-depth analysis, we answer how to properly set the smoothness approximation method to mitigate generalization error in FAL. Moreover, we identify RSA as the most effective generalization error reduction method. In highly data-heterogeneous scenarios, we also recommend employing SFAL to mitigate the deterioration of generalization performance caused by heterogeneity. Based on our theoretical results, we provide insights to help develop more efficient FAL algorithms, such as designing new metrics and dynamic aggregation rules to mitigate heterogeneity. Wenjun Ding, Ying An, Lixing Chen, Shichao Kan |
AAAI | 4 |
| 2025 | Multimodal Foundation Model Adaptation with Clinical Knowledge Guidance for IDH GenotypingabstractAccurately predicting isocitrate dehydrogenase (IDH) mutations is crucial for glioma diagnosis, but the limited availability of multimodal MRI restricts the generalization of existing methods. Fine-tuning foundation models is a common solution, yet their lack of domain-specific knowledge often impairs optimal performance. Clinical studies show that both the imagemodal T2-FLAIR mismatch sign knowledge and text-modal demographic information are useful for IDH genotyping. Thus, we propose a novel network that integrates multimodal clinical knowledge to guide the fine-tuning of the multimodal foundation model M3D for glioma IDH genotyping on multimodal MRI. To fully utilize multimodal knowledge, we first extract multigranularity T2-FLAIR mismatch features from different layers via a Mixture-of-Experts (MoE) pool-based mismatch adapter (incorporating an MoE-based pooling mechanism and a spatialchannel mismatch attention module). Meanwhile, demographic information is converted into text prompts and encoded by the M3D text encoder to generate demographic-based text embeddings. In the encoder, T2-FLAIR mismatch features are integrated at the end of each ViT Block to introduce imagingspecific knowledge, while text features are further processed via a cross-modal text-image attention fusion adapter to enhance representation learning in the joint feature space. We evaluated the approach on an internal dataset (from 3 public datasets, 871 patients) and an independent external dataset (501 patients). It achieved 93.25 % accuracy on the internal dataset and 86.03 % on the external dataset, with only$\mathbf{2. 6 7 M}$trainable parameters, outperforming 7 existing state-of-the-art IDH genotyping methods. Hulin Kuang, Yingxu Chen, Jin Liu 0012, Jie Wang 0067, Shichao Kan |
BIBM | 6 |
| 2025 | GMReg: Group Mamba Correlation Based Pyramid Network with Edge Enhancement for Medical Image RegistrationabstractDeformable image registration is fundamental in medical image analysis. Existing pyramid-based deep learning methods suffer from coarse deformation decomposition, poor inter-level transitions, and error accumulation. To address these, we propose GMReg, an unsupervised pyramid network for medical image registration based on Mamba correlation. GMReg introduces intra-level multi-scale decomposition: each pyramid level splits deformation fields into subfields with different receptive fields via grouping, using Group Mamba correlation layers for feature matching/fusion, and convolutional prediction for sub-fields. A channel attention-based context fusion module enhances inter-group interaction, while a multi-scale edge enhancement module guides subfield fusion to improve boundary sensitivity. Experiments on two public brain MRI datasets (LPBA40, Mind-Boggle) show GMReg significantly outperforms state-of-the-art methods in registration accuracy. Additionally, results on the FIRE dataset demonstrate that GMReg also holds potential for the fundus image registration task. Hulin Kuang, Guangheng Wu, Jin Liu 0012, Shichao Kan, Jie Wang 0067 |
BIBM | 5 |
| 2025 | EPCPE: A Real-time End-to-End Pipeline for RGB-based Category-level 6D Pose EstimationabstractRGB-based category-level 6D pose estimation methods have faced significant challenges in achieving real-time performance, primarily due to the design of two-stage pipeline. To address this issue, we propose a novel end-to-end pipeline named EPCPE. In detail, we first extract implicit rotation features via Large Visual Model (LVM), and then adaptively obtain pose-specific features with a fined-tuned Lite Feature Extractor. Finally, we introduce a novel Pose Decoder with two parallel branches, enabling simultaneous 6D pose estimation and 2D object detection. We also propose a novel rotation loss function to further enhance the performance. Extensive experiments on the CAMERA25 and REAL275 datasets demonstrate that our pipeline is concise, achieves state-of-the-art (SOTA) and real-time performance. Xiaofeng Fan, Shichao Kan, Yixiong Liang |
ICASSP | 3 |
| 2025 | Scene Graph Generation with Large Vision-Language Model and Its ApplicationsabstractScene graph generation (SGG) is pivotal for acquiring valuable knowledge in visual scene understanding, making it crucial for tasks such as visual question answering and visual reasoning. In recent times, multimodal large language models (MLLMs) have demonstrated remarkable proficiency in object grounding and recognition. Nevertheless, constructing the scene graph directly poses a challenging task for MLLMs due to the intricate nature of predicting relationships between objects and the action states of objects. Simultaneously, reliance solely on the large language model (LLM) for achieving region understanding in complex scenes makes MLLMs susceptible to hallucinations, leading to potential errors in determining the coordinates of objects. To tackle these challenges, we introduce a large vision-language model (LVLM) within the Shikra framework for scene graph generation. Our approach involves the creation of an instruction-following SGG dataset for the fine-tuning of the LVLM. After SGG, we recognize that the scene graph generated by the LVLM can assist the LLM in answering visual questions. Thus, we evaluate LVLM-based question-and-answering models by leveraging the scene graph as a rationale, introducing a concept termed Scene Graph Chain of Thought (SGCoT). The proposed SGG method is rigorously evaluated through both quantitative and qualitative experiments on closed-set and open-set SGG tasks, affirming its effectiveness. Moreover, using the scene graph generated by fine-tuning LVLM as a rationale in the chain of thoughts results in a competitive performance on several vision-language compositional benchmarks. Wei-Xin Chen, Yong-Yong Chen, Shichao Kan |
ICME | 3 |
| 2025 | Noise-Guided Predicate Representation Extraction and Diffusion-Enhanced Discretization for Scene Graph GenerationabstractScene Graph Generation (SGG) is a fundamental task in visual understanding, aimed at providing more precise local detail comprehension for downstream applications. Existing SGG methods often overlook the diversity of predicate representations and the consistency among similar predicates when dealing with long-tail distributions. As a result, the model's decision layer fails to effectively capture details from the tail end, leading to biased predictions. To address this, we propose a Noise-Guided Predicate Representation Extraction and Diffusion-Enhanced Discretization (NoDIS) method. On the one hand, expanding the predicate representation space enhances the model's ability to learn both common and rare predicates, thus reducing prediction bias caused by data scarcity. We propose a conditional diffusion model to reconstructs features and increase the diversity of representations for same category predicates. On the other hand, independent predicate representations in the decision phase increase the learning complexity of the decision layer, making accurate predictions more challenging. To address this issue, we introduce a discretization mapper that learns consistent representations among similar predicates, reducing the learning difficulty and decision ambiguity in the decision layer. To validate the effectiveness of our method, we integrate NoDIS with various SGG baseline models and conduct experiments on multiple datasets. The results consistently demonstrate superior performance. Shichao Kan, Fanghui Zhang, Wanru Xu, Yue Zhang 0065, Yi-Gang Cen |
ICML | 2 |
| 2025 | Dynamic Residual Encoding with Slide-Level Contrastive Learning for End-to-End Whole Slide Image RepresentationabstractWhole Slide Image (WSI) representation is critical for cancer subtyping, cancer recognition and mutation prediction.Training an end-to-end WSI representation model poses significant challenges, as a standard gigapixel slide can contain tens of thousands of image tiles, making it difficult to compute gradients of all tiles in a single mini-batch due to current GPU limitations. To address this challenge, we propose a method of dynamic residual encoding with slide-level contrastive learning (DRE-SLCL) for end-to-end WSI representation. Our approach utilizes a memory bank to store the features of tiles across all WSIs in the dataset. During training, a mini-batch usually contains multiple WSIs. For each WSI in the batch, a subset of tiles is randomly sampled and their features are computed using a tile encoder. Then, additional tile features from the same WSI are selected from the memory bank. The representation of each individual WSI is generated using a residual encoding technique that incorporates both the sampled features and those retrieved from the memory bank. Finally, the slide-level contrastive loss is computed based on the representations and histopathology reports ofthe WSIs within the mini-batch. Experiments conducted over cancer subtyping, cancer recognition, and mutation prediction tasks proved the effectiveness of the proposed DRE-SLCL method. Te Gao, Zhihong Shi, Yixiong Liang, Ruiqing Zheng, Hulin Kuang, Min Zeng 0004, Shichao Kan |
ACM Multimedia | 9 |
| 2025 | Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental LearningabstractMultimodal Biomedical Image Incremental Learning (MBIIL) is essential for handling diverse tasks and modalities in the biomedical domain, as training separate models for each modality or task significantly increases inference costs. Existing incremental learning methods focus on task expansion within a single modality, whereas MBIIL seeks to train a unified model incrementally across modalities. The MBIIL faces two challenges: I) How to preserve previously learned knowledge during incremental updates? II) How to effectively leverage knowledge acquired from existing modalities to support new modalities? To address these challenges, we propose MSLoRA-CR, a method that fine-tunes Modality-Specific LoRA modules while incorporating Contrastive Regularization to enhance intra-modality knowledge sharing and promote inter-modality knowledge differentiation. Our approach builds upon a large vision-language model (LVLM), keeping the pretrained model frozen while incrementally adapting new LoRA modules for each modality or task. Experiments on the incremental learning of biomedical images demonstrate that MSLoRA-CR outperforms both the state-of-the-art (SOTA) approach of training separate models for each modality and the general incremental learning method (incrementally fine-tuning LoRA). Specifically, MSLoRA-CR achieves a 1.88% improvement in overall performance compared to unconstrained incremental learning methods while maintaining computational efficiency. Our code is publicly available at https://github.com/VentusAislant/MSLoRA_CR. Yixiong Liang, Hulin Kuang, Yi-Gang Cen, Min Zeng 0004, Shichao Kan |
ACM Multimedia | 8 |
| 2025 | 2OMe-LM: predicting 2′-O-methylation sites in human RNA using a pre-trained RNA language modelabstractMOTIVATION: 2'-O-methylation (2OMe) is a common post-transcriptional modification in RNA that plays a crucial role in regulating gene expression and is implicated in various biological processes and diseases. Computational methods offer an efficient alternative to the time-consuming and costly experimental identification of 2OMe sites. Recent advancements in RNA pre-trained language models have revolutionized RNA bioinformatics. However, there remains a gap in their application specifically for predicting 2OMe sites. RESULTS: In the study, we propose a novel deep learning framework, 2OMe-LM, for predicting 2OMe sites in RNA. 2OMe-LM integrates RNA sequence features derived from RNA pre-trained language models with those obtained from the word2vec technique. Then, 2OMe-LM employs fully connected layers and a bidirectional long short-term memory network to process the two types of features separately, followed by a feature fusion module for the final prediction. Additionally, an attention block is incorporated to provide the interpretability of the prediction results. The results demonstrate that 2OMe-LM significantly outperforms existing state-of-the-art predictors, with features from RNA pre-trained language models proving to be critical. Motif analysis further demonstrates 2OMe-LM's potential for discovering 2OMe-related motifs. AVAILABILITY AND IMPLEMENTATION: The 2OMe-LM web server is available at https://csuligroup.com:9200/2OMe-LM. The source code can be obtained from https://github.com/CSUBioGroup/2OMe-LM. Qianpei Liu, Min Zeng 0004, Chengqian Lu, Shichao Kan, Fei Guo 0001, Min Li 0007 |
Bioinform. | 5 |
| 2025 | Feature Transformation Reconstruction (FTR) Network for Unsupervised Anomaly DetectionabstractThe goal of the feature reconstruction network based on an autoencoder in the training phase is to force the network to reconstruct the input features well. The network tends to learn shortcuts of “identity mapping,” which leads to the network outputting abnormal features as they are in the inference phase. As such, the abnormal features based on reconstruction error cannot be distinguished from normal features, significantly limiting the detection performance of such methods. To address this issue, we propose a feature transformation reconstruction (FTR) network, which can avoid the identity mapping problem. Specifically, we use a normalizing flow model as a feature transformation (FT) network to transform input features into other forms. The training goal of the feature reconstruction (FR) network is no longer to reconstruct the input features but to reconstruct the transformed features, effectively avoiding the shortcut of learning the “identity map.” Furthermore, this paper proposes a masked convolutional attention (MCA) module, which randomly masks the input features in the training phase and reconstructs the input features in a self‐supervised manner. In the testing phase, the MCA can effectively suppress the excessive reconstruction of abnormal features and further improve anomaly detection performance. FTR achieves the scores of the area under the receiver operating characteristic curve (AUROC) at 99.5% and 97.8% on the MVTec AD and BTAD datasets, respectively, outperforming other state‐of‐the‐art methods. Moreover, FTR is faster than the existing methods, with a high speed of 137 frames per second (FPS) on a 3080ti GPU. Linna Zhang, Lanyao Zhang, Qi Cao 0002, Shichao Kan, Yi-Gang Cen, Fugui Zhang, Yansen Huang |
Int. J. Intell. Syst. | 4 |
| 2025 | Cross-scene visual context parsing with large vision-language modelabstractRelation analysis is crucial for image-based applications such as visual reasoning and visual question answering . Current relation analysis such as scene graph generation (SGG) only focuses on building relationships among objects within a single image. However, in real-world applications, relationships among objects across multiple images, as seen in video understanding , may hold greater significance as they can capture global information. This is still a challenging and unexplored task. In this paper, we aim to explore the technique of Cross-Scene Visual Context Parsing (CS-VCP) using a large vision-language model. To achieve this, we first introduce a cross-scene dataset comprising 10,000 pairs of cross-scene visual instruction data, with each instruction describing the common knowledge of a pair of cross-scene images. We then propose a Cross-Scene Visual Symbiotic Linkage (CS-VSL) model to understand both cross-scene relationships and objects by analyzing the rationales in each scene. The model is pre-trained on 100,000 cross-scene image pairs and validated on 10,000 image pairs. Both quantitative and qualitative experiments demonstrate the effectiveness of the proposed method. Our method has been released on GitHub: https://github.com/gavin-gqzhang/CS-VSL . Shichao Kan, Lu Shi 0004, Wanru Xu, Gaoyun An, Yi-Gang Cen |
Pattern Recognit. | 2 |
| 2025 | CellCircLoc: Deep Neural Network for Predicting and Explaining Cell Line-Specific CircRNA Subcellular LocalizationabstractThe subcellular localization of circular RNAs (circRNAs) is crucial for understanding their functional relevance and regulatory mechanisms. CircRNA subcellular localization exhibits variations across different cell lines, demonstrating the diversity and complexity of circRNA regulation within distinct cellular contexts. However, existing computational methods for predicting circRNA subcellular localization often ignore the importance of cell line specificity and instead train a general model on aggregated data from all cell lines. Considering the diversity and context-dependent behavior of circRNAs across different cell lines, it is imperative to develop cell line-specific models to accurately predict circRNA subcellular localization. In the study, we proposed CellCircLoc, a sequence-based deep learning model for circRNA subcellular localization prediction, which is trained for different cell lines. CellCircLoc utilizes a combination of convolutional neural networks, Transformer blocks, and bidirectional long short-term memory to capture both sequence local features and long-range dependencies within the sequences. In the Transformer blocks, CellCircLoc uses an attentive convolution mechanism to capture the importance of individual nucleotides. Extensive experiments demonstrate the effectiveness of CellCircLoc in accurately predicting circRNA subcellular localization across different cell lines, outperforming other computational models that do not consider cell line specificity. Moreover, the interpretability of CellCircLoc facilitates the discovery of important motifs associated with circRNA subcellular localization. Min Zeng 0004, Jingwei Lu, Chengqian Lu, Shichao Kan, Fei Guo 0001, Min Li 0007 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Multi-Modal Self-Perception Enhanced Large Language Model for 3D Region-of-Interest Captioning With Limited Dataabstract3D Region-of-Interest (RoI) Captioning involves translating a model's understanding of specific objects within a complex 3D scene into descriptive captions. Recent advancements in Large Language Models (LLMs) have shown great potential in this area. Existing methods capture the visual information from RoIs as input tokens for LLMs. However, this approach may not provide enough detailed information for LLMs to generate accurate region-specific captions. In this paper, we introduce Self-RoI, a Large Language Model with multi-modal self-perception capabilities for 3D RoI captioning. To ensure LLMs receive more precise and sufficient information, Self-RoI incorporates Implicit Textual Info. Perception to construct a multi-modal vision-language information. This module utilizes a simple mapping network to generate textual information about basic properties of RoI from vision-following response of LLMs. This textual information is then integrated with the RoI's visual representation to form a comprehensive multi-modal instruction for LLMs. Given the limited availability of 3D RoI-captioning data, we propose a two-stage training strategy to optimize Self-RoI efficiently. In the first stage, we align 3D RoI vision and caption representations. In the second stage, we focus on 3D RoI vision-caption interaction, using a disparate contrastive embedding module to improve the reliability of the implicit textual information and employing language modeling loss to ensure accurate caption generation. Our experiments demonstrate that Self-RoI significantly outperforms previous 3D RoI captioning models. Moreover, the Implicit Textual Info. Perception can be integrated into other multi-modal LLMs for performance enhancement. We will make our code available for further research. Lu Shi 0004, Shichao Kan, Yi Jin 0001, Linna Zhang, Yi-Gang Cen |
IEEE Trans. Multim. | 2 |
| 2025 | Low-Shot Unsupervised Visual Anomaly Detection via Sparse Feature RepresentationabstractVisual anomaly detection is an essential component in modern industrial manufacturing. Existing studies using notions of pairwise similarity distance between a test feature and nominal features have achieved great breakthroughs. However, the absolute similarity distance lacks certain generalizations, making it challenging to extend the comparison beyond the available samples. This limitation could potentially hamper anomaly detection performance in scenarios with limited samples. This article presents a novel sparse feature representation anomaly detection (SFRAD) framework, which formulates the anomaly detection as a sparse feature representation problem; and notably proposes an anomaly score by orthogonal matching pursuit (ASOMP) as a novel detection metric. Specifically, SFRAD calculates the Gaussian kernel distance between the test feature and its sparse representation in the nominal feature space for anomaly detection. Here, the orthogonal matching pursuit (OMP) algorithm is adopted to achieve the sparse feature representation. Moreover, to construct a low-redundancy memory bank storing the basis features for sparse representation, a novel basis feature sampling (BFS) algorithm is proposed by considering both the maximum coverage and the optimum feature representation simultaneously. As a result, SFRAD incorporates both the advantages of absolute similarity and linear representation; and this enhances the generalization in low-shot scenarios. Extensive experiments on the MVTec anomaly detection (MVTec AD), Kolektor surface-defect dataset (KolektorSDD), Kolektor surface-defect dataset 2 (KolektorSDD2), MVTec logical constraints anomaly detection (MVTec LOCO AD), Visual anomaly (VISA), Modified national institute of standards and technology (MNIST), and CIFAR-10 datasets demonstrate that our proposed SFRAD outperforms the previous methods and achieves state-of-the-art unsupervised anomaly detection performance. Notably, significantly improved outcomes and results have also been achieved on low-shot anomaly detection. Code is available at https://github.com/fanghuisky/SFRAD. Fanghui Zhang, Haiyue Zhu, Yi-Gang Cen, Shichao Kan, Linna Zhang, Prahlad Vadakkepat, Tong Heng Lee |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Global Contrastive Learning with High-Quality Data in Large Vision-Language Models for Pathological Question AnsweringabstractPathological question answering (PQA) is vital in computational pathology, as it involves interpreting pathological images and answering questions posed by humans. This interaction offers an effective means of engaging with users and enhancing the understanding of pathology-related information. Recent methods developed using large vision-language models (LVLMs), such as QUILT-LLAVA, have made significant progress in advancing PQA. However, existing models, such as those using the QUILT-1M dataset, neglect the quality of the training set during the fine-tuning stage, leading to sub-optimal performance. We recognize that high-quality training data can significantly enhance model performance. Therefore, we design a model-based data filtering strategy to remove images with obvious impurities from the instruction fine-tuning dataset. The filtered high-quality images are then used to fine-tune the model. Additionally, most existing vision-language alignment strategies focus primarily on aligning local features through next-word prediction, leading to a relatively homogeneous granularity in inter-modal alignment. To address this issue, we propose a global-wise alignment module, which introduces global-level contrastive learning during the pretraining stage to establish multi-granularity alignment between pathological images and language descriptions. Based on the above two processes, we design our method, named Global Contrastive Learning with High-Quality Data (GCL-HQD) for pathological question answering in LVLMs. Extensive experiments on two types of experimental settings demonstrate the effectiveness of the GCL-HQD method. Hulin Kuang, Suoni Liu, Shichao Kan |
BIBM | 5 |
| 2024 | Aligning Multimodal Biomedical Images and Language via One Large Vision-Language ModelabstractLarge Vision-Language Models (LVLMs) have garnered substantial attention in the biomedical image analysis domain due to their robust vision understanding capabilities. However, current methods rely heavily on dataset- and modality-specific fine-tuning. This involves tuning separate models for each dataset and biomedical modality. In this paper, we introduce a method for aligning multimodal biomedical images and language using a single LVLM, dubbed UniMed-LVLM. Specifically, we devise a General Projection Module (GPM) by integrating multiple image projection branches and implementing dynamic routing between the vision encoder and language decoder within the LLaVA-Med framework. Subsequently, we progressively align multiple biomedical modalities using a Parameter-Efficient Fine-Tuning (PEFT) technique known as Low-Rank Adaptation (LoRA). The model is initially trained on the LLaVA-Med dataset and then fine-tuned on four biomedical image analysis datasets: PathVqa, Slake, VqaRad, and Fitzpatrick17k, enabling the simultaneous analysis of radiology, pathology, and dermatology images. A single model is fine-tuned on three modalities across these datasets and evaluated on all test sets. Experimental results show that UniMed-LVLM improves the average evaluation score by 1.88% across the four datasets, validating its effectiveness in handling multimodal biomedical images. Min Zeng 0004, Jinfeng Ding, Yixiong Liang, Ruiqing Zheng, Min Li 0007, Shichao Kan |
BIBM | 8 |
| 2024 | FedGCA: Global Consistent Augmentation Based Single-Source Federated Domain GeneralizationabstractFederated Domain Generalization (FedDG) aims to train the global model for generalization ability to unseen domains with multi-domain training samples. However, clients in federated learning networks are often confined to a single, non-IID domain due to inherent sampling and temporal limitations. The lack of cross-domain interaction and the in-domain divergence impede the learning of domain-common features and limit the effectiveness of existing FedDG, referred to as the single-source FedDG (sFedDG) problem. To address this, we introduce the Federated Global Consistent Augmentation (FedGCA) method, which incorporates a style-complement module to augment data samples with diverse domain styles. To ensure the effective integration of augmented samples, FedGCA employs both global guided semantic consistency and class consistency, mitigating inconsistencies from local semantics within individual clients and classes across multiple clients. The conducted extensive experiments demonstrate the superiority of FedGCA. Yuan Liu 0038, Shichao Kan, Jianxin Wang 0001 |
ICME | 5 |
| 2024 | FedMMR: Multi-Modal Federated Learning via Missing Modality ReconstructionabstractFederated Learning (FL) presents a decentralized learning approach for privacy preservation. While many research focuses on uni-modal FL, a more intricate version, multimodal FL, uncovers fundamental attributes. Clients might only gather specific modalities, incurring missing modalities in multi-modal FL. The missing modality problem incurs modality heterogeneity among clients and loses inter-modal connection within one client. To address these issues, we introduce a novel multi-modal FL method called Federated Missing Modality Reconstruction (FedMMR), which tackles these dual challenges through two distinct facets. First, we devise a cross-modal reconstruction policy for synthesizing absent modalities from other modalities. This aligns the data feature spaces across clients, thereby alleviating bias in resultant local models. Subsequently, we steer local models to maintain an inter-modal awareness of both existing and reconstructed modalities, recognizing their potential as complementary components. Comprehensive results show that FedMMR outperforms existing FL baselines. Yuan Liu 0038, Shichao Kan, Yixiong Liang, Jianxin Wang 0001 |
ICME | 4 |
| 2024 | HRDecoder: High-Resolution Decoder Network for Fundus Image Lesion Segmentation
Ziyuan Ding, Yixiong Liang, Shichao Kan, Qing Liu 0003 |
MICCAI (9) | 3 |
| 2024 | CAKE: a flexible self-supervised framework for enhancing cell visualization, clustering and rare cell identificationabstractSingle cell sequencing technology has provided unprecedented opportunities for comprehensively deciphering cell heterogeneity. Nevertheless, the high dimensionality and intricate nature of cell heterogeneity have presented substantial challenges to computational methods. Numerous novel clustering methods have been proposed to address this issue. However, none of these methods achieve the consistently better performance under different biological scenarios. In this study, we developed CAKE, a novel and scalable self-supervised clustering method, which consists of a contrastive learning model with a mixture neighborhood augmentation for cell representation learning, and a self-Knowledge Distiller model for the refinement of clustering results. These designs provide more condensed and cluster-friendly cell representations and improve the clustering performance in term of accuracy and robustness. Furthermore, in addition to accurately identifying the major type cells, CAKE could also find more biologically meaningful cell subgroups and rare cell types. The comprehensive experiments on real single-cell RNA sequencing datasets demonstrated the superiority of CAKE in visualization and clustering over other comparison methods, and indicated its extensive application in the field of cell heterogeneity analysis. Contact: Ruiqing Zheng. ([email protected]). Jin Liu 0012, Weixing Zeng, Shichao Kan, Min Li 0007, Ruiqing Zheng |
Briefings Bioinform. | 3 |
| 2023 | Singularformer: Learning to Decompose Self-Attention to Linearize the Complexity of TransformerabstractTransformers achieve excellent performance in a variety of domains since they can capture long-distance dependencies through the self-attention mechanism. However, self-attention is computationally costly due to its quadratic complexity and high memory consumption. In this paper, we propose a novel Transformer variant (Singularformer) that uses neural networks to learn the singular value decomposition process of the attention matrix to design a linear-complexity and memory-efficient global self-attention mechanism. Specifically, we decompose the attention matrix into the product of three matrix factors based on singular value decomposition and design neural networks to learn these matrix factors, then the associative law of matrix multiplication is used to linearize the calculation of self-attention. The above procedure allows us to compute self-attention as two-dimensional reduction processes in the first and second token dimensional spaces, followed by a multi-head self-attention computational process on the first dimensional reduced token features. Experimental results on 8 real-world datasets demonstrate that Singularformer performs favorably against the other Transformer variants with lower time and space complexity. Our source code is publicly available at https://github.com/CSUBioGroup/Singularformer. Yifan Wu 0008, Shichao Kan, Min Zeng 0004, Min Li 0007 |
IJCAI | 2 |
| 2023 | Confidence-Aware Contrastive Learning for Semantic SegmentationabstractRecently supervised contrastive learning (SCL) has achieved remarkable progress in semantic segmentation. Nevertheless, prior works have often necessitated a substantial number of samples to attain satisfactory performance, leading to a significant increase in training overhead. In this work, we leverage the idea of reweighting each pair to reduce the demand for large numbers of training samples in contrastive learning and propose a novel loss, dubbed confidence-aware contrastive (CAC) loss, which adaptively reweights each pair according to the predicted confidence for semantic segmentation. To alleviate the misalignment between supervised learning and contrastive learning, we further introduce an extra weight branch with a stop-gradient operator to generate the pair weights. Moreover, we present a confidence-aware marginal anchor sampling method for the calculation of supervised contrastive loss which focuses on marginal rather than the hardest pairs. Coupling with our method consistently improves the performance of various models (e.g. HRNet, OCRNet, SegFormer) on Cityscapes, ADE20K, PASCAL-Context, and COCO-Stuff datasets. Compared to existing SCL-based methods, the proposed method achieves competitive or even better results without relying on a memory bank or a large number of samples. Our code is at https://github.com/CVIU-CSU/Confidence-Aware-Contrastive-Loss. Lele Lv, Qing Liu 0003, Shichao Kan, Yixiong Liang |
ACM Multimedia | 3 |
| 2023 | POAR: Towards Open Vocabulary Pedestrian Attribute RecognitionabstractPedestrian attribute recognition (PAR) aims to predict the attributes of a target pedestrian. Recent methods often address the PAR problem by training a multi-label classifier with predefined attribute classes, but they can hardly exhaust all possible pedestrian attributes in the real world. To tackle this problem, we propose a novel Pedestrian Open-Attribute Recognition (POAR) approach by formulating the problem as a task of image-text search. Our approach employs a Transformer-based Encoder with a Masking Strategy (TEMS) to focus on the attributes of specific pedestrian parts (e.g., head, upper body, lower body, feet, etc.), and introduces a set of attribute tokens to encode the corresponding attributes into visual embeddings. Each attribute category is described as a natural language sentence and encoded by the text encoder. Then, we compute the similarity between the visual and text embeddings to find the best attribute descriptions for the input images. To handle multiple attributes of a single pedestrian, we propose a Many-To-Many Contrastive (MTMC) loss with masked tokens. In addition, we propose a Grouped Knowledge Distillation (GKD) method to minimize the disparity between visual embeddings and unseen attribute text embeddings. We evaluate our proposed method on three PAR datasets with an open-attribute setting. The results demonstrate the effectiveness of our method as a strong baseline for the POAR task. Our code is available at https://github.com/IvyYZ/POAR. Yue Zhang 0065, Suchen Wang, Shichao Kan, Zhenyu Weng, Yi-Gang Cen, Yap-Peng Tan |
ACM Multimedia | 3 |
| 2023 | End-to-end feature diversity person search with rank constraint of cross-class matrix
Yue Zhang 0065, Shuqin Wang 0001, Shichao Kan, Yi-Gang Cen, Linna Zhang |
Neurocomputing | 3 |
| 2023 | Multi-layer capsule network with joint dynamic routing for fire recognition
Yuming Wu, Shichao Kan, Yongfang Xie |
Image Vis. Comput. | 3 |
| 2023 | Multiscale spatial temporal attention graph convolution network for skeleton-based anomaly behavior detection
Shichao Kan, Fanghui Zhang, Yi-Gang Cen, Linna Zhang, Damin Zhang |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Contrastive Bayesian Analysis for Deep Metric LearningabstractRecent methods for deep metric learning have been focusing on designing different contrastive loss functions between positive and negative pairs of samples so that the learned feature embedding is able to pull positive samples of the same class closer and push negative samples from different classes away from each other. In this work, we recognize that there is a significant semantic gap between features at the intermediate feature layer and class labels at the final output layer. To bridge this gap, we develop a contrastive Bayesian analysis to characterize and model the posterior probabilities of image labels conditioned by their features similarity in a contrastive learning setting. This contrastive Bayesian analysis leads to a new loss function for deep metric learning. To improve the generalization capability of the proposed method onto new classes, we further extend the contrastive Bayesian loss with a metric variance constraint. Our experimental results and ablation studies demonstrate that the proposed contrastive Bayesian metric learning method significantly improves the performance of deep metric learning in both supervised and pseudo-supervised scenarios, outperforming existing methods by a large margin. Shichao Kan, Zhiquan He, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | A graph model-based multiscale feature fitting method for unsupervised anomaly detection
Fanghui Zhang, Shichao Kan, Damin Zhang, Yi-Gang Cen, Linna Zhang, Vladimir Mladenovic |
Pattern Recognit. | 2 |
| 2022 | Coded Residual Transform for Generalizable Deep Metric LearningabstractA fundamental challenge in deep metric learning is the generalization capability of the feature embedding network model since the embedding network learned on training classes need to be evaluated on new test classes. To address this challenge, in this paper, we introduce a new method called coded residual transform (CRT) for deep metric learning to significantly improve its generalization capability. Specifically, we learn a set of diversified prototype features, project the feature map onto each prototype, and then encode its features using their projection residuals weighted by their correlation coefficients with each prototype. The proposed CRT method has the following two unique characteristics. First, it represents and encodes the feature map from a set of complimentary perspectives based on projections onto diversified prototypes. Second, unlike existing transformer-based feature representation approaches which encode the original values of features based on global correlation analysis, the proposed coded residual transform encodes the relative differences between the original features and their projected prototypes. Embedding space density and spectral decay analysis show that this multi perspective projection onto diversified prototypes and coded residual representation are able to achieve significantly improved generalization capability in metric learning. Finally, to further enhance the generalization performance, we propose to enforce the consistency on their feature similarity matrices between coded residual transforms with different sizes of projection prototypes and embedding dimensions. Our extensive experimental results and ablation studies demonstrate that the proposed CRT method outperform the state-of-the-art deep metric learning methods by large margins and improving upon the current best method by up to 4.28% on the CUB dataset. Shichao Kan, Yixiong Liang, Min Li 0007, Yi-Gang Cen, Jianxin Wang 0001, Zhihai He |
NeurIPS | 1 |
| 2022 | A GAN-based input-size flexibility model for single image dehazing
Shichao Kan, Yue Zhang 0065, Fanghui Zhang, Yi-Gang Cen |
Signal Process. Image Commun. | 1 |
| 2022 | Local Semantic Correlation Modeling Over Graph Neural Networks for Deep Feature Embedding and Image RetrievalabstractDeep feature embedding aims to learn discriminative features or feature embeddings for image samples which can minimize their intra-class distance while maximizing their inter-class distance. Recent state-of-the-art methods have been focusing on learning deep neural networks with carefully designed loss functions. In this work, we propose to explore a new approach to deep feature embedding. We learn a graph neural network to characterize and predict the local correlation structure of images in the feature space. Based on this correlation structure, neighboring images collaborate with each other to generate and refine their embedded features based on local linear combination. Graph edges learn a correlation prediction network to predict the correlation scores between neighboring images. Graph nodes learn a feature embedding network to generate the embedded feature for a given image based on a weighted summation of neighboring image features with the correlation scores as weights. Our extensive experimental results under the image retrieval settings demonstrate that our proposed method outperforms the state-of-the-art methods by a large margin, especially for top-1 recalls. Shichao Kan, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He |
IEEE Trans. Image Process. | 1 |
| 2021 | Spatial Assembly Networks for Image Representation LearningabstractIt has been long recognized that deep neural networks are sensitive to changes in spatial configurations or scene structures. Image augmentations, such as random translation, cropping, and resizing, can be used to improve the robustness of deep neural networks under spatial transforms. However, changes in object part configurations, spatial layout of object, and scene structures of the images may still result in major changes in the their feature representations generated by the network, creating significant challenges for various visual learning tasks, including representation or metric learning, image classification and retrieval. In this work, we introduce a new learnable module, called spatial assembly network (SAN), to address this important issue. This SAN module examines the input image and performs a learned re-organization and assembly of feature points from different spatial locations conditioned by feature maps from previous network layers so as to maximize the discriminative power of the final feature representation. This differentiable module can be flexibly incorporated into existing network architectures, improving their capabilities in handling spatial variations and structural changes of the image scene. We demonstrate that the proposed SAN module is able to significantly improve the performance of various metric / representation learning, image retrieval and classification tasks, in both supervised and unsupervised learning scenarios. Yang Li 0091, Shichao Kan, Jianhe Yuan, Wenming Cao 0001, Zhihai He |
CVPR | 2 |
| 2021 | Relative Order Analysis and Optimization for Unsupervised Deep Metric LearningabstractIn unsupervised learning of image features without labels, especially on datasets with fine-grained object classes, it is often very difficult to tell if a given image belongs to one specific object class or another, even for human eyes. However, we can reliably tell if image C is more similar to image A than image B. In this work, we propose to explore how this relative order can be used to learn discriminative features with an unsupervised metric learning method. Instead of resorting to clustering or self-supervision to create pseudo labels for an absolute decision, which often suffers from high label error rates, we construct reliable relative orders for groups of image samples and learn a deep neural network to predict these relative orders. During training, this relative order prediction network and the feature embedding network are tightly coupled, providing mutual constraints to each other to improve overall metric learning performance in a cooperative manner. During testing, the predicted relative orders are used as constraints to optimize the generated features and refine their feature distance-based image retrieval results using a constrained optimization procedure. Our experimental results demonstrate that the proposed relative orders for unsupervised learning (ROUL) method is able to significantly improve the performance ofunsupervised deep metric learning. Shichao Kan, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He |
CVPR | 1 |
| 2021 | Cross-domain Person Re-identification Based on the Sample Relation Guidance
Yue Zhang 0065, Fanghui Zhang, Shichao Kan, Linna Zhang, Jiaping Zong, Yi-Gang Cen |
ICIG (2) | 3 |
| 2021 | Block-based image matching for image retrieval
Ruizhen Zhao, Liequan Liang, Xinwei Zheng, Yi-Gang Cen, Shichao Kan |
J. Vis. Commun. Image Represent. | 6 |
| 2021 | Learned Model Composition With Critical Sample Look-Ahead for Semi-Supervised Learning on Small Sets of Labeled SamplesabstractIn this work, we propose to push the performance limit of semi-supervised learning on very small sets of labeled samples by developing a new method called learned model composition with critical sample look-ahead (LMCS). Training efficient deep neural networks on much smaller sets of labeled samples is a challenging problem. With a small labeled set, the initial network suffers from low accuracy. Based on this error-prone network, the subsequent semi-supervised learning process will be fragile and unstable. To address this issue, we propose to introduce a look-ahead master model to identify the correct direction of model evolution to effectively guide the semi-supervised learning process of the student model. Specifically, our proposed LMCS method explores two major ideas. First, it introduces a new learned model composition structure so that we can compose a more efficient master network from student models of past iterations through a network learning process. Second, we develop a new method, called confined maximum entropy search, to discover new critical samples near the model decision boundary and provide the master model with look-ahead access to these samples to enhance its guidance capability. Our extensive experimental results demonstrate that the proposed LMCS method outperforms the state-of-the-art semi-supervised learning methods, especially on small sets of labeled samples. For example, on the CIFAR-10 dataset, with a very small set of 80 labeled samples, our method outperforms Google's MixMatch method, reducing the error rate by more than 10%. Yang Li 0091, Shichao Kan, Wenming Cao 0001, Zhihai He |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Zero-Shot Learning to Index on Semantic Trees for Scalable Image RetrievalabstractIn this study, we develop a new approach, called zero-shot learning to index on semantic trees (LTI-ST), for efficient image indexing and scalable image retrieval. Our method learns to model the inherent correlation structure between visual representations using a binary semantic tree from training images which can be effectively transferred to new test images from unknown classes. Based on predicted correlation structure, we construct an efficient indexing scheme for the whole test image set. Unlike existing image index methods, our proposed LTI-ST method has the following two unique characteristics. First, it does not need to analyze the test images in the query database to construct the index structure. Instead, it is directly predicted by a network learnt from the training set. This zero-shot capability is critical for flexible, distributed, and scalable implementation and deployment of the image indexing and retrieval services at large scales. Second, unlike the existing distance-based index methods, our index structure is learnt using the LTI-ST deep neural network with binary encoding and decoding on a hierarchical semantic tree. Our extensive experimental results on benchmark datasets and ablation studies demonstrate that the proposed LTI-ST method outperforms existing index methods by a large margin while providing the above new capabilities which are highly desirable in practice. Shichao Kan, Yi-Gang Cen, Vladimir Mladenovic, Yang Li 0091, Zhihai He |
IEEE Trans. Image Process. | 1 |
| 2020 | Unsupervised Deep Metric Learning with Transformed Attention Consistency and Contrastive Clustering Loss
Yang Li 0091, Shichao Kan, Zhihai He |
ECCV (11) | 2 |
| 2020 | Metric learning-based kernel transformer with triplets and label constraints for feature fusion
Shichao Kan, Linna Zhang, Zhihai He, Yi-Gang Cen, Shiming Chen 0001, Jikun Zhou |
Pattern Recognit. | 1 |
| 2019 | A supervised learning to index model for approximate nearest neighbor image retrieval
Shichao Kan, Xinwei Zheng, Yi-Gang Cen, Zhenmin Zhu, Hengyou Wang |
Signal Process. Image Commun. | 1 |
| 2019 | Supervised Deep Feature Embedding With Handcrafted FeatureabstractImage representation methods based on deep convolutional neural networks (CNNs) have achieved the state-of-the-art performance in various computer vision tasks, such as image retrieval and person re-identification. We recognize that more discriminative feature embeddings can be learned with supervised deep metric learning and handcrafted features for image retrieval and similar applications. In this paper, we propose a new supervised deep feature embedding with a handcrafted feature model. To fuse handcrafted feature information into CNNs and realize feature embeddings, a general fusion unit is proposed (called Fusion-Net). We also define a network loss function with image label information to realize supervised deep metric learning. Our extensive experimental results on the Stanford online products' data set and the in-shop clothes retrieval data set demonstrate that our proposed methods outperform the existing state-of-the-art methods of image retrieval by a large margin. Moreover, we also explore the applications of the proposed methods in person re-identification and vehicle re-identification; the experimental results demonstrate both the effectiveness and efficiency of the proposed methods. Shichao Kan, Yi-Gang Cen, Zhihai He, Zhi Zhang 0005, Linna Zhang |
IEEE Trans. Image Process. | 1 |
| 2018 | Deep Proposal and Detection Networks for Road Damage Detection and ClassificationabstractRoad maintenance and management is an important task for the social administration. Inspecting the road conditions is the basis of the maintenance, and efficient and accurate inspection is able to perform road repair timely as well as reduce the maintenance cost. Traditional road damage detection have to use high-performance sensors, which are very costly. With the development of the deep learning methods, it is possible to detect the roads efficiently directly with images. In this paper, we propose a simple but efficient object detection model for road damage detection. We adopt the state-of-the-art deep objective detection model including Faster-RCNN and SSD for solving this problem. Experiments convey that the proposed model achieves good detection results and win the challange of IEEE BigData Road Damage Detection Challenge. The codes of the model can be available at https://github.com/kanshichao/DPDN_for_Road_Damage_Detection_and_Classification. Yanbo J. Wang, Shichao Kan, Chenyue Lu |
IEEE BigData | 3 |
| 2017 | SURF binarization and fast codebook construction for image retrieval
Shichao Kan, Yi-Gang Cen, Viacheslav V. Voronin, Vladimir Mladenovic, Ming Zeng 0012 |
J. Vis. Commun. Image Represent. | 1 |