VLDB 2026 Research / reviewers in the wild / expert
Chenlong Gao
dblp:202/6120
· DBLP profile ↗
16ranked-venue papers
0as first author
13since 2021 · last 2025
0000-0001-8827-7271ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mitigating Pervasive Modality Absence Through Multimodal Generalization and RefinementabstractThe performance of multimodal models often deteriorates when modality absence occurs. The absence disrupts the learned inter-modal correlations, resulting in biased multimodal representations. This challenge is especially pronounced when the absence is pervasive, affecting both the training and inference phases. Recent studies have attempted to reconstruct the missing information; however, most of them require complete supervision, which is seldom available in scenarios of pervasive absence. The quality of reconstruction remains a critical issue. Alternatively, others aim to learn robust representations from the available modalities but the substantial variations and biases are not fully addressed. This paper introduces the Multimodal Generalization and Refinement (MGR) framework to mitigate the issue of pervasive modality absence. MGR begins by acquiring generalized multimodal representations and iteratively refines them to recognize and calibrate the biased representations. Initially, multimodal samples with absence are embedded through foundation models, and MGR integrates independent unimodal features to further enhance generalization. Additionally, a novel mixed-context prompt is adopted to identify biases in both features and correlations. A redistribution operation can then refine these biases through graph pooling, culminating in robust and calibrated multimodal representations, which are suitable for downstream tasks. Comprehensive experiments on four benchmark datasets demonstrate that the proposed MGR framework outperforms state-of-the-art methods, effectively mitigating the impact of pervasive modality absence. Wuliang Huang, Yiqiang Chen 0001, Xinlong Jiang, Chenlong Gao, Qian Chen 0023, Yifan Wang 0027 |
AAAI | 4 |
| 2025 | VersaFusion: A Versatile Diffusion-Based Framework for Fine-Grained Image Editing and EnhancementabstractText-to-image (T2I) diffusion models have achieved remarkable progress in generating realistic images from textual descriptions. However, ensuring consistent high-quality image generation with complete backgrounds, object appearance, and optimal texture rendering remains challenging. This paper presents a novel fine-grained pixel-level image editing method based on pre-trained diffusion models. The proposed dual-branch architecture, consisting of Guidance and Generation branches, employs U-Net Denoisers and Self-Attention mechanisms. An improved DDIM-like inversion method obtains the latent representation, followed by multiple denoising steps. Cross-branch interactions, such as KV Replacement, Classifier Guidance, and Feature Correspondence, enable precise control while preserving image fidelity. The iterative refinement and reconstruction process facilitates finegrained editing control, supporting attribute modification, image outpainting, style transfer, and face synthesis with Clickand-Drag style editing using masks. Experimental results demonstrate the effectiveness of the proposed approach in enhancing the quality and controllability of T2I-generated images, surpassing existing methods while maintaining attractive computational complexity for practical real-world applications. Haocun Ye, Xinlong Jiang, Chenlong Gao, Bingyu Wang, Wuliang Huang |
AAAI | 3 |
| 2025 | FairFHTL: Achieving Task-Agnostic Fairness in Federated Hetero-Task LearningabstractFederated Hetero-Task Learning (FHTL) enables the simultaneous learning of multiple heterogeneous tasks on federated learning clients, offering enhanced flexibility. However, the inconsistency between optimization objectives and evaluation metrics for these heterogeneous tasks poses challenges in achieving performance fairness among clients. This study proposes a fairness-aware FHTL method, FairFHTL. It employs adversarial multi-task representation learning at the client level to learn the task-independent shared model. Consequently, it solves optimization objectives inspired by fair resource allocation on the server side to determine the update direction of the global shared model, ultimately achieving task-independent fair performance balance. Extensive experiments on three multi-task datasets demonstrate that FairFHTL significantly enhances performance across the majority of tasks compared to conventional federated learning and FHTL methods. Moreover, compared with other fairness-aware federated learning approaches, FairFHTL maintains a more uniform performance distribution across all tasks. Yiqiang Chen 0001, Xinlong Jiang, Wuliang Huang, Qian Chen 0023, Chenlong Gao, Zhirui Wang 0004, Bingjie Yan |
ICME | 6 |
| 2025 | Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and ApplicationabstractLarge Language Models (LLMs) have showcased exceptional capabilities in various domains, attracting significant interest from both academia and industry. Despite their impressive performance, the substantial size and computational demands of LLMs pose considerable challenges for practical deployment, particularly in environments with limited resources. The endeavor to compress language models while maintaining their accuracy has become a focal point of research. Among the various methods, knowledge distillation has emerged as an effective technique to enhance inference speed without greatly compromising performance. This article presents a thorough survey from three aspects: method, evaluation, and application, exploring knowledge distillation techniques tailored specifically for LLMs. Specifically, we divide the methods into white-box KD and black-box KD to better illustrate their differences. Furthermore, we also explored the evaluation tasks and distillation effects between different distillation methods and proposed directions for future research. Through in-depth understanding of the latest advancements and practical applications, this survey provides valuable resources for researchers, paving the way for sustained progress in this field. Chuanpeng Yang, Yao Zhu 0003, Wang Lu 0003, Yidong Wang 0003, Qian Chen 0023, Chenlong Gao, Bingjie Yan, Yiqiang Chen 0001 |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2024 | EyeGraphGPT: Knowledge Graph Enhanced Multimodal Large Language Model for Ophthalmic Report GenerationabstractAutomatic generation of ophthalmic reports holds significant potential to lessen clinicians’ workload, enhance work efficiency, and alleviate the imbalance between clinicians and patients. Recent advancements in multimodal large language models, represented by GPT-4, have demonstrated remarkable performance in the general domain. However, training such models necessitates a substantial amount of paired image-text data, yet paired ophthalmic data is limited, and ophthalmic reports are laden with specialized terminologies, making it challenging to transfer the training paradigm to the ophthalmic domain. In this paper, we propose EyeGraphGPT, a knowledge graph enhanced multimodal large language model for ophthalmic report generation. Specifically, we construct a knowledge graph by leveraging the knowledge from a medical database and expertise from ophthalmic experts to model relationships among ophthalmic diseases, enhancing the model’s focus on key disease information. We then perform relation-aware modal alignment to incorporate knowledge graph features into visual features, and further enhance modality collaboration through visual instruction fine-tuning to adapt the model to the ophthalmic domain. Our experiments on a real-world dataset demonstrates that EyeGraphGPT outperforms previous state-of-the-art models, highlighting its superiority in scenarios with limited medical data and extensive specialized terminologies. Xinlong Jiang, Chenlong Gao, Weiwei Dai, Bingyu Wang, Bingjie Yan, Wuliang Huang |
BIBM | 3 |
| 2024 | Buffalo: Biomedical Vision-Language Understanding with Cross-Modal Prototype and Federated Foundation Model CollaborationabstractFederated learning (FL) enables collaborative learning across multiple biomedical data silos with multimodal foundation models while preserving privacy. Due to the heterogeneity in data processing and collection methodologies across diverse medical institutions and the varying medical inspections patients undergo, modal heterogeneity exists in practical scenarios, where severe modal heterogeneity may even prevent model training. With privacy considerations, data transfer cannot be permitted, restricting knowledge exchange among different clients. To trickle these issues, we propose a cross-modal prototype imputation method for visual-language understanding (Buffalo) with only a slight increase in communication cost, which can improve the performance of fine-tuning general foundation models for downstream biomedical tasks. We conducted extensive experiments on medical report generation and biomedical visual question-answering tasks. The results demonstrate that Buffalo can fully utilize data from all clients to improve model generalization compared to other modal imputation methods in three modal heterogeneity scenarios, approaching or even surpassing the performance in the ideal scenario without missing modality. Bingjie Yan, Qian Chen 0023, Yiqiang Chen 0001, Xinlong Jiang, Wuliang Huang, Bingyu Wang, Zhirui Wang 0004, Chenlong Gao |
CIKM | 8 |
| 2024 | Model Trip: Enhancing Privacy and Fairness in Model Fusion Across Multi-Federations for Trustworthy Global HealthcareabstractFederated Learning has emerged as a revolutionary innovation in the evolving landscape of global healthcare, fostering collaboration among institutions and facilitating collaborative data analysis. As practical applications continue to proliferate, numerous federations have formed in different regions. The optimization and sustainable development of federation-pretrained models have emerged as new challenges. These challenges primarily encompass privacy, population shift and data dependency, which may lead to severe consequences such as the leakage of sensitive information within models and training samples, unfair model performance and resource burdens. To tackle these issues, we propose FairFusion, a cross-federation model fusion approach that enhances privacy and fairness. FairFusion operates across federations within a Model Trip paradigm, integrating knowledge from diverse federations to continually enhance model performance. Through federated model fusion, multi-objective quantification and optimization, FairFusion obtains trustworthy solutions that excel in utility, privacy and fairness. We conduct comprehensive experiments on three public real-world healthcare datasets. The results demonstrate that FairFusion achieves outstanding model fusion performance in terms of utility and fairness across various model structures and subgroups with sensitive attributes while guaranteeing model privacy. Qian Chen 0023, Yiqiang Chen 0001, Bingjie Yan, Xinlong Jiang, Xiaojin Zhang 0002, Yan Kang 0001, Wuliang Huang, Chenlong Gao, Lixin Fan, Qiang Yang 0001 |
ICDE | 9 |
| 2024 | MixPrompt: Enhancing Generalizability and Adversarial Robustness for Vision-Language Models via Prompt Fusion
Yunli Chen, Chenlong Gao |
ICIC (9) | 6 |
| 2024 | Correlation-Driven Multi-Modality Graph Decomposition for Cross-Subject Emotion RecognitionabstractMulti-modality physiological signal-based emotion recognition has attracted increasing attention as its capacity to capture human affective states comprehensively. Due to multi-modality heterogeneity and cross-subject divergence, practical applications struggle with generalizing models across individuals. Effectively addressing both issues requires mitigating the gap between multimodal signals while acquiring generalizable representations across subjects. However, existing approaches often handle these dual challenges separately, resulting in suboptimal generalization. This study introduces a novel framework, termed Correlation-Driven Multi-Modality Graph Decomposition (CMMGD). The proposed CMMGD initially captures adaptive cross-modal correlations. It connects each unimodal graph to a multimodal mixed graph. To simultaneously address the dual challenges, it incorporates a correlation-driven graph decomposition module that decomposes the mixed graph into concordant and discrepant subgraphs based on the correlations. The decomposed concordant subgraph encompasses consistently activated features across modalities and subjects during emotion elicitation, unveiling a generalizable subspace. Additionally, we design a Multi-Modality Graph Regularized Transformer (MGRT) backbone specifically tailored for multimodal physiological signals. The MGRT can alleviate the over-smoothing issue and mitigate over-reliance on any single modality. Extensive experiments demonstrate that CMMGD outperforms the state-of-the-art methods by 1.79% and 2.65% on DEAP and MAHNOB-HCI datasets, respectively, under the leave-one-subject-out cross-validation strategy. Wuliang Huang, Yiqiang Chen 0001, Xinlong Jiang, Chenlong Gao, Qian Chen 0023, Bingjie Yan, Yifan Wang 0027, Jianrong Yang |
ACM Multimedia | 4 |
| 2024 | FedBone: Towards Large-Scale Federated Multi-Task Learning
Xinlong Jiang, Chenlong Gao, Wuliang Huang |
J. Comput. Sci. Technol. | 5 |
| 2023 | FedTAM: Decentralized Federated Learning with a Feature Attention Based Multi-teacher Knowledge Distillation for HealthcareabstractFederated learning has emerged as a powerful technique for training robust models while preserving data privacy and security. However, real-world applications, especially in domains like healthcare, often face challenges due to non-independent and non-identically distributed (non-iid) data across different institutions. Additionally, the heterogeneity of data and the absence of a trusted central server further hinder collaborative efforts among medical institutions. Our paper introduces a novel federated learning approach called FedTAM, which incorporates cyclic model transfer and feature attention-based multi-teacher knowledge distillation. FedTAM is designed to tailor personalized models for individual clients within a decentralized federated learning setting, where data distribution is non-iid. Notably, this method enables student clients to selectively acquire the most pertinent and valuable knowledge from teacher clients through feature attention mechanism while filtering out irrelevant information. We conduct extensive experiments across five benchmark healthcare datasets and one public image classification dataset with feature shifts. Our results conclusively demonstrate that our method achieves remarkable accuracy improvements when compared to state-of-the-art approaches. This affirms the potential of FedTAM to significantly enhance federated learning performance, especially in challenging real-world contexts like healthcare. Tingting Mou, Xinlong Jiang, Bingjie Yan, Qian Chen 0023, Wuliang Huang, Chenlong Gao, Yiqiang Chen 0001 |
ICPADS | 8 |
| 2023 | AFL-CS: Asynchronous Federated Learning with Cosine Similarity-based Penalty Term and AggregationabstractHorizontal Federated Learning offers a means to develop machine learning models in the realm of medical application while preserving the confidentiality and security of patient data. However, due to the substantial heterogeneity of the devices in medical institution, traditional synchronous federated aggregation methods result in a noticeable decrease in training efficiency, thereby impacting the application and deployment of federated learning. Asynchronous Federated Learning (AFL) model aggregation methods can mitigate this problem but present new challenges in terms of convergence stability and speed. In this paper, we propose a cosine similarity-based layer-wise penalty term and asynchronous model aggregation method AFL-CS, which considers the global model convergence direction during local training. Compared with existing AFL aggregation methods, AFL-CS can achieve faster and more consistent convergence direction to superior performance especially in non-iid settings with high statistical heterogeneity, even reaching and exceeding synchronous FL. Bingjie Yan, Xinlong Jiang, Yiqiang Chen 0001, Chenlong Gao, Xuequn Liu |
ICPADS | 4 |
| 2021 | Deep unsupervised multi-modal fusion network for detecting driver distraction
Yiqiang Chen 0001, Chenlong Gao |
Neurocomputing | 3 |
| 2020 | WeDA: Designing and Evaluating A Scale-driven Wearable Diagnostic Assessment System for Children with ADHDabstractAttention Deficit Hyperactivity Disorder (ADHD) is one of the most common mental disorders affecting children. Because the etiology of ADHD is complex and its symptoms are not specific, there is a lack of feasible quantitative diagnostic methods. Pursuing objective and non-invasive detection methods and standards is of great practical significance to prevent the development of the disease. In this study, we aim to address one specific concern about the objectivity and quantification of ADHD diagnosis. Over a year, we iteratively designed and tested WeDA, a scale-driven wearable diagnostic assessment system. This system contains an Android computer machine with a large touchscreen, a suite of 3D printed interactive devices, and six wearable motion sensors. We implement ten diagnostic tasks drawing on the symptoms of ADHD based on DSM-5. The experimental results of classifying children with ADHD and typically developing children and subjective evaluations from doctors, parents, and children validate the effectiveness and acceptability of WeDA. Xinlong Jiang, Yiqiang Chen 0001, Wuliang Huang, Chenlong Gao, Yunbing Xing |
CHI | 5 |
| 2019 | A Novel Feature Incremental Learning Method for Sensor-Based Activity RecognitionabstractRecognizing activities of daily living is an important research topic for health monitoring and elderly care. However, most existing activity recognition models only work with static and pre-defined sensor configurations. Enabling an existing activity recognition model to adapt to the emergence of new sensors in a dynamic environment is a significant challenge. In this paper, we propose a novel feature incremental learning method, namely the Feature Incremental Random Forest (FIRF), to improve the performance of an existing model with a small amount of data on newly appeared features. It consists of two important components - 1) a mutual information based diversity generation strategy (MIDGS) and 2) a feature incremental tree growing mechanism (FITGM). MIDGS enhances the internal diversity of random forests, while FITGM improves the accuracy of individual decision trees. To evaluate the performance of FIRF, we conduct extensive experiments on three well-known public datasets for activity recognition. Experimental results demonstrate that FIRF is significantly more accurate and efficient compared with other state-of-the-art methods. It has the potential to allow the dynamic exploitation of new sensors in changing environments. Chunyu Hu 0001, Yiqiang Chen 0001, Xiaohui Peng 0002, Han Yu 0001, Chenlong Gao, Lisha Hu |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2017 | A multistage collaborative filtering method for fall detectionabstractFalls threaten the health and life of the elders heavily because they lead to injuries or even death. Therefore, a reliable monitoring and alarm mechanism is desperately in need to guarantee the quality of elders' life. In this paper, we propose a multistage machine learning method to perform fall detection and solve the false alarm and missing alarm problem in traditional fall detection methods. Our proposed method consists of three stages: 1) threshold filtering, 2) ELM classifier, and 3) orientation-based filtering. Our method utilizes a high-precision triaxial accelerometer to collect the relevant information. After filtered by our three-stage method, the signal can be determined whether it is a fall or not. Experimental results demonstrate that: different from the traditional state-of-art methods with a single machine learning classifier, our method can greatly reduce the missing alarm and false alarm rate on the premise of high accuracy for all detection. Yiqiang Chen 0001, Lisha Hu, Chenlong Gao, Chunyu Hu 0001, Jianfei Shen |
IJCNN | 4 |