Kelu Yao

dblp:205/8747 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-4891-3197ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 3Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic Consistency
abstract
Cross-modal Knowledge Distillation has demonstrated promising performance on paired modalities with strong semantic connections, referred to as Symmetric Cross-modal Knowledge Distillation (SCKD). However, implementing SCKD becomes exceedingly constrained in real-world scenarios due to the limited availability of paired modalities. To this end, we investigate a general and effective knowledge learning concept under weak semantic consistency, dubbed Asymmetric Cross-modal Knowledge Distillation (ACKD), aiming to bridge modalities with limited semantic overlap. Nevertheless, the shift from strong to weak semantic consistency improves flexibility but exacerbates challenges in knowledge transmission costs, which we rigorously verified based on optimal transport theory. To mitigate the issue, we further propose a framework, namely SemBridge, integrating a Student-Friendly Matching module and a Semantic-aware Knowledge Alignment module. The former leverages self-supervised learning to acquire semantic-based knowledge and provide personalized instruction for each student sample by dynamically selecting the relevant teacher samples. The latter seeks the optimal transport path by employing Lagrangian optimization. To facilitate the research, we curate a benchmark dataset derived from two modalities, namely Multi-Spectral (MS) and asymmetric RGB images, tailored for remote sensing scene classification. Comprehensive experiments exhibit that our framework achieves state-of-the-art performance compared with 7 existing approaches on 6 different model architectures across various datasets.
Riling Wei, Kelu Yao, Chuanguang Yang, Jin Wang 0039, Zhuoyan Gao, Chao Li 0028
AAAI2
2026 Space Computing Constellation: System Architecture, Implementations, and Challenges
abstract
Low Earth Orbit (LEO) satellite constellations have experienced rapid growth in recent years, driven by their potential to deliver global, high-bandwidth Internet services with low latency. Beyond connectivity, LEO constellations also offer promising opportunities to enable in-orbit processing of space-native data to support a wide range of emerging space applications. In this context, the concept of space computing has been proposed, a paradigm that seamlessly integrates networking and computing to provide computing-as-a-service anytime and anywhere in space. However, the inherent characteristics of satellite constellations, such as dynamic network topologies, constrained system resources, and the harsh space environment, pose significant challenges in achieving this vision. This paper outlines the system architecture and the key enabling technologies for space computing, including spaceborne computers, laser communications, spaceborne router, distributed operating systems, and onboard AI. We also present the implementation of an open space computing platform, the 3-Body computing constellation, along with the in-orbit experimental results that demonstrate the advantages of multi-satellite distributed computing. Furthermore, we outline future research directions essential for advancing toward a truly interconnected, autonomous, and intelligent space computing system.
Hua Wang 0011, Kelu Yao, Luqi Gong, Yichao Jin 0001, Yuan Liu 0030, Junxiao Xue, Zhiguo Wan, Chao Li 0028, Zhifeng Zhao
IEEE Internet Things J.2
2026 Visual-language active search for wide-area remote sensing imagery
Kelu Yao
Pattern Recognit.2
2025 Forensics-Bench: A Comprehensive Forgery Detection Benchmark Suite for Large Vision Language Models
abstract
Recently, the rapid development of AIGC has significantly boosted the diversities of fake media spread in the Internet, posing unprecedented threats to social security, politics, law, and etc. To detect the ever-increasingly diverse malicious fake media in the new era of AIGC, recent studies have proposed to exploit Large Vision Language Models (LVLMs) to design robust forgery detectors due to their impressive performance on a wide range of multimodal tasks. However, it still lacks a comprehensive benchmark designed to comprehensively assess LVLMs' discerning capabilities on forgery media. To fill this gap, we present Forensics-Bench, a new forgery detection evaluation benchmark suite to assess LVLMs across massive forgery detection tasks, requiring comprehensive recognition, location and reasoning capabilities on diverse forgeries. Forensics-Bench comprises 63, 292 meticulously curated multi-choice visual questions, covering 112 unique forgery detection types from 5 perspectives: forgery semantics, forgery modalities, forgery tasks, forgery types and forgery models. We conduct thorough evaluations on 22 open-sourced LVLMs and 3 proprietary models GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet, highlighting the significant challenges of comprehensive forgery detection posed by Forensics-Bench. We anticipate that Forensics-Bench will motivate the community to advance the frontier of LVLMs, striving for all-around forgery detectors in the era of AIGC. The deliverables will be updated here.
Jin Wang 0039, Chenghui Lv, Shichao Dong 0001, Kelu Yao, Chao Li 0028, Wenqi Shao
CVPR6
2025 ECG-guided individual identification via PPG
abstract
Photoplethsmography (PPG)-based individual identification aiming at recognizing humans via intrinsic cardiovascular activities has raised extensive attention due to its high security and resistance to mimicry. However, this kind of technology witnesses unpromising results due to the limitation of low information density. To this end, electrocardiogram (ECG) signals have been introduced as a novel modality to enhance the density of input information. Specifically, a novel cross-modal knowledge distillation framework is implemented to propagate discriminate knowledge from ECG modality to PPG modality without incurring additional computational demands at the inference phase. Furthermore, to ensure efficient knowledge propagation, Contrastive Language–Image Pre-training (CLIP)-based knowledge alignment and cross-knowledge assessment modules are proposed respectively. Comprehensive experiments are conducted and results show our framework outperforms the baseline model with the improvement of 2.8% and 3.0% in terms of overall accuracy on seen- and unseen individual recognitions.
Riling Wei, Kelu Yao, Chuanguang Yang, Chao Li 0028
ICASSP3
2025 Distillation-boosted heterogeneous architecture search for aphid counting
Shengqin Jiang, Qian Jie, Fengna Cheng, Kelu Yao
Expert Syst. Appl.5
2024 Diagnosing the Compositional Knowledge of Vision Language Models from a Game-Theoretic View
abstract
Compositional reasoning capabilities are usually considered as fundamental skills to characterize human perception. Recent studies show that current Vision Language Models (VLMs) surprisingly lack sufficient knowledge with respect to such capabilities. To this end, we propose to thoroughly diagnose the composition representations encoded by VLMs, systematically revealing the potential cause for this weakness. Specifically, we propose evaluation methods from a novel game-theoretic view to assess the vulnerability of VLMs on different aspects of compositional understanding, e.g., relations and attributes. Extensive experimental results demonstrate and validate several insights to understand the incapabilities of VLMs on compositional reasoning, which provide useful and reliable guidance for future studies. The deliverables will be updated here.
Jin Wang 0039, Shichao Dong 0001, Yapeng Zhu, Kelu Yao, Chao Li 0028
ICML4
2024 Embracing Adaptation: An Effective Dynamic Defense Strategy Against Adversarial Examples
abstract
Existing adversarial example defense methods are static, meaning they remain unchanged once training is completed, regardless of how attack methods change. Consequently, static defense methods are highly vulnerable to adaptive attacks. We argue that to counter more formidable attacks, models should continually adapt to various attack methods. We propose a novel dynamic defense approach. Initially, we use Gaussian Mixture Models (GMM) to obtain structural information of the data, which is combined with model prediction information to generate pseudo-labels for optimizing inputs. Subsequently, we employ information maximization and enhanced mean predictions as optimization objectives, utilizing a hierarchical optimization approach to refine the model. Meanwhile, we propose a sample-efficient optimization strategy that reduces the total number of samples in the test data stream for reverse updating and improves the efficiency. Notably, our method can be directly applied to pre-trained models without the need for accessing training data or retraining the model. Therefore, our approach is training-data-agnostic and model-agnostic, easily applicable to existing adversarially trained models, significantly enhancing the resilience of various models against white-box, black-box, and adaptive attacks across diverse datasets. We have conducted extensive experiments to validate the state-of-the-art of our proposed method. The pseudo-code can be found in the appendix.
Shenglin Yin, Kelu Yao, Jieyi Long
ACM Multimedia2
2023 AGAIN: Adversarial Training with Attribution Span Enlargement and Hybrid Feature Fusion
abstract
The deep neural networks (DNNs) trained by adversarial training (AT) usually suffered from significant robust generalization gap, i.e., DNNs achieve high training robustness but low test robustness. In this paper, we propose a generic method to boost the robust generalization of AT methods from the novel perspective of attribution span. To this end, compared with standard DNNs, we discover that the generalization gap of adversarially trained DNNs is caused by the smaller attribution span on the input image. In other words, adversarially trained DNNs tend to focus on specific visual concepts on training images, causing its limitation on test robustness. In this way, to enhance the robustness, we propose an effective method to enlarge the learned attribution span. Besides, we use hybrid feature statistics for feature fusion to enrich the diversity of features. Extensive experiments show that our method can effectively improves robustness of adversarially trained DNNs, outperforming previous SOTA methods. Furthermore, we provide a theoretical analysis of our method to prove its effectiveness.
Shenglin Yin, Kelu Yao, Sheng Shi, Yangzhou Du
CVPR2
2023 Towards Understanding the Generalization of Deepfake Detectors from a Game-Theoretical View
abstract
This paper aims to explain the generalization of deep-fake detectors from the novel perspective of multi-order interactions among visual concepts. Specifically, we propose three hypotheses: 1. Deepfake detectors encode multi-order interactions among visual concepts, in which the low-order interactions usually have substantially negative contributions to deepfake detection. 2. Deepfake detectors with better generalization abilities tend to encode low-order interactions with fewer negative contributions. 3. Generalized deepfake detectors usually weaken the negative contributions of low-order interactions by suppressing their strength. Accordingly, we design several mathematical metrics to evaluate the effect of low-order interaction for deepfake detectors. Extensive comparative experiments are conducted, which verify the soundness of our hypotheses. Based on the analyses, we further propose a generic method, which directly reduces the toxic effects of low-order interactions to improve the generalization of deepfake detectors to some extent.
Kelu Yao, Jin Wang 0039, Boyu Diao, Chao Li 0028
ICCV1
2022 Interpretable Generative Adversarial Networks
abstract
Learning a disentangled representation is still a challenge in the field of the interpretability of generative adversarial networks (GANs). This paper proposes a generic method to modify a traditional GAN into an interpretable GAN, which ensures that filters in an intermediate layer of the generator encode disentangled localized visual concepts. Each filter in the layer is supposed to consistently generate image regions corresponding to the same visual concept when generating different images. The interpretable GAN learns to automatically discover meaningful visual concepts without any annotations of visual concepts. The interpretable GAN enables people to modify a specific visual concept on generated images by manipulating feature maps of the corresponding filters in the layer. Our method can be broadly applied to different types of GANs. Experiments have demonstrated the effectiveness of our method.
Chao Li 0028, Kelu Yao, Jin Wang 0039, Boyu Diao, Yongjun Xu 0001, Quanshi Zhang
AAAI2
2019 Prediction and Study of the Applicability of Medical Gels to Patients
abstract
Gel is a post-operative cleaning material with antibacterial effect, which helps patients recover after surgery. It is more and more popular in surgery, but it is still controversial in use. This study collected the electronic medical records of patients in a hospital for nearly three years, using a combination of a variety of special selection methods to process data and using random forest, support vector machine, LightGBM and XGBoost and other machine learning methods to predict the suitability of patients. The results show that polysaccharide gel is not suitable for all people, whether to use it should consider different situations. This paper has studied the applicability of medical gels to patients, and established a predictability model to provide data support for the clinical application of this expensive medical material.
Bo Liu 0024, Mengmeng Huang, Kelu Yao, Xiaolu Fei, Qing Wang 0003
COMPSAC (2)3
2018 Gastric Pathology Image Recognition Based on Deep Residual Networks
abstract
Gastric cancer is a malignant neoplasm with a high mortality rate in the world. Nearly one million new cases occur each year. The most important measure to diagnose gastric cancer is the detection and treatment of diseases early. Gastric cancer detection is currently performed by pathologists reviewing large expanses of biological tissues, but this process is labor intensive and error-prone. In this paper, a framework for automatically detection of tumors in gastric pathology image (slide) has been proposed based on deep learning. A deep residual network with 50 layers is built by identity mapping on a dataset of pathology images. The proposed method makes the training of models easier and improves the generalization performance. Finally, the experimental results show that the F-score of our method achieves 96%. The research in auto-classification of gastric pathology images has great value for gastric cancer detection in clinical medicine.
Bo Liu 0024, Kelu Yao, Mengmeng Huang, Yong Li 0037
COMPSAC (2)2
2017 The Investigation on Effectiveness Evaluation Methods for One Medical Material Used for Surgery Patients Based on Electronic Medical Records Data
abstract
As the development of Electronic Medical Records in hospitals, more and more "real world data" became available for clinical researches, especially for the clinical effectiveness evaluation. Besides traditional statistical methods, more Machine Learning methods are also used to analyze the data. In this study, one high value medical consumable -gel, which is used in large quantities in cleaning surgical incision, is analyzed using both a traditional statistical method and a Machine Learning method to evaluate its clinical applicability, efficacy and safety. The Electronic Medical Records for three years are collected including patient gender, age, quantity of gel usages, surgical incision grades, preoperative diagnosis, and information about postoperative recoveries and so on. Through the two analysis methods, the difference among incision healing, antibiotic use, and the postoperative hospital days after using gel or common saline in the wound are analyzed. The results show that the two methods can reveal different aspects in clinical evaluation. The decision tree classification can provide valuable suggestions for the reasonable use of the gel and making reasonable medical policy with the applicable conditions of gels.
Kelu Yao, Xiaolu Fei, WangQing
COMPSAC (2)2