Yuxing Lu

dblp:299/1429 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-8207-4411ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 GlyphShield: Document Watermarking for the Physical World via Vector Typeface Synthesis
Yuxing Lu, Han Fang 0004, Sijing Xie, Luyu Yuan, Chengxin Zhao
AAAI2
2026 MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics
abstract
Yuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi, Rui Peng, Shuang Zeng, Xingyu Hu, Jinzhuo Wang, May Dongmei Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi, Rui Peng 0006, Shuang Zeng, Jinzhuo Wang, May D. Wang
ACL (1)1
2025 Knowledge Graph and Large Language Model for Metabolomics
abstract
The advancements in Knowledge Graphs (KGs) and Large Language Models (LLMs) are driving transformative changes across various research fields, including metabolomics. These tools present exceptional opportunities to elucidate complex metabolic pathways and identify biomarkers essential to biological systems. My research focuses on harnessing the potential of KGs and LLMs within metabolomics, specifically making interactions between them and with biological researches. KGs, with their structured representation of metabolic entities and relationships, provide a robust foundation for managing extensive multimodal metabolomic knowledge. Recently, I developed a metabolite-centric knowledge graph and explored innovative methodologies to leverage KGs and LLMs for enhancing predictive modeling in clinical settings. My future research aims to fully exploit the capabilities of KGs and LLMs in metabolomics, advancing our understanding and applications in this field.
Yuxing Lu
AAAI1
2025 END^2: Robust Dual-Decoder Watermarking Framework Against Non-Differentiable Distortions
abstract
DNN-based watermarking methods have rapidly advanced, with the ``Encoder-Noise Layer-Decoder'' (END) framework being the most widely used. To ensure end-to-end training, the noise layer in the framework must be differentiable. However, real-world distortions are often non-differentiable, leading to challenges in end-to-end training. Existing solutions only treat the distortion perturbation as additive noise, which does not fully integrate the effect of distortion in training. To better incorporate non-differentiable distortions into training, we propose a novel dual-decoder architecture (END^2). Unlike conventional END architecture, our method employs two structurally identical decoders: the Teacher Decoder, processing pure watermarked images, and the Student Decoder, handling distortion-perturbed images. The gradient is backpropagated only through the Teacher Decoder branch to optimize the encoder thus bypassing the problem of non-differentiability. To ensure resistance to arbitrary distortions, we enforce alignment of the two decoders' feature representations by maximizing the cosine similarity between their intermediate vectors on a hypersphere. Extensive experiments demonstrate that our scheme outperforms state-of-the-art algorithms under various non-differentiable distortions. Moreover, even without the differentiability constraint, our method surpasses baselines with a differentiable noise layer. Our approach is effective and easily implementable across all END architectures, enhancing practicality and generalizability.
Han Fang 0004, Yuxing Lu, Chengxin Zhao
AAAI3
2025 Towards Doctor-Like Reasoning: Medical RAG Fusing Knowledge with Patient Analogy through Textual Gradients
abstract
Existing medical RAG systems mainly leverage knowledge from medical knowledge bases, neglecting the crucial role of experiential knowledge derived from similar patient cases - a key component of human clinical reasoning. To bridge this gap, we propose DoctorRAG, a RAG framework that emulates doctor-like reasoning by integrating both explicit clinical knowledge and implicit case-based experience. DoctorRAG enhances retrieval precision by first allocating conceptual tags for queries and knowledge sources, together with a hybrid retrieval mechanism from both relevant knowledge and patient. In addition, a Med-TextGrad module using multi-agent textual gradients is integrated to ensure that the final output adheres to the retrieved knowledge and patient query. Comprehensive experiments on multilingual, multitask datasets demonstrate that DoctorRAG significantly outperforms strong baseline RAG models and gains improvements from iterative refinements. Our approach generates more accurate, relevant, and comprehensive responses, taking a step towards more doctor-like medical reasoning systems.
Yuxing Lu, Gecheng Fu, Xukai Zhao, Sin Yee Goi, Jinzhuo Wang
NeurIPS1
2025 KARMA: Leveraging Multi-Agent LLMs for Automated Knowledge Graph Enrichment
abstract
Maintaining comprehensive and up-to-date knowledge graphs (KGs) is critical for modern AI systems, but manual curation struggles to scale with the rapid growth of scientific literature. This paper presents KARMA, a novel framework employing multi-agent large language models (LLMs) to automate KG enrichment through structured analysis of unstructured text. Our approach employs nine collaborative agents, spanning entity discovery, relation extraction, schema alignment, and conflict resolution that iteratively parse documents, verify extracted knowledge, and integrate it into existing graph structures while adhering to domain-specific schema. Experiments on 1,200 PubMed articles from three different domains demonstrate the effectiveness of KARMA in knowledge graph enrichment, with the identification of up to 38,230 new entities while achieving 83.1\% LLM-verified correctness and reducing conflict edges by 18.6\% through multi-layer assessments.
Yuxing Lu, Xukai Zhao, Rui Peng 0006, Jinzhuo Wang
NeurIPS1
2025 KINDLE: Knowledge-Guided Distillation for Prior-Free Gene Regulatory Network Inference
abstract
Gene regulatory network (GRN) inference serves as a cornerstone for deciphering cellular decision-making processes. Early approaches rely exclusively on gene expression data, thus their predictive power remain fundamentally constrained by the vast combinatorial space of potential gene-gene interactions. Subsequent methods integrate prior knowledge to mitigate this challenge by restricting the solution space to biologically plausible interactions. However, we argue that the effectiveness of these approaches is contingent upon the precision of prior information and the reduction in the search space will circumscribe the models' potential for novel biological discoveries. To address these limitations, we introduce KINDLE, a three-stage framework that decouples GRN inference from prior knowledge dependencies. KINDLE trains a teacher model that integrates prior knowledge with temporal gene expression dynamics and subsequently distills this encoded knowledge to a student model, enabling accurate GRN inference solely from expression data without access to any prior. KINDLE achieves state-of-the-art performance across four benchmark datasets. Notably, it successfully identifies key transcription factors governing mouse embryonic development and precisely characterizes their functional roles. In mouse hematopoietic stem cell data, KINDLE accurately predicts fate transition outcomes following knockout of two critical regulators (Gata1 and Spi1). These biological validations demonstrate our framework's dual capability in maintaining topological inference precision while preserving discovery potential for novel biological mechanisms.
Rui Peng 0006, Qichen Sun, Yuxing Lu, Ziru Liu, Jinzhuo Wang
NeurIPS4
2025 Ultra-high Resolution Watermarking Framework Resistant to Extreme Cropping and Scaling
abstract
Recent developments in DNN-based image watermarking techniques have achieved impressive results in protecting digital content. However, most existing methods are constrained to low-resolution images as they need to encode the entire image, leading to prohibitive memory and computational costs when applied to high-resolution images. Moreover, they lack robustness to distortions prevalent in large-image transmission, such as extreme scaling and random cropping. To address these issues, we propose a novel watermarking method based on implicit neural representations (INRs). Leveraging the properties of INRs, our method employs resolution-independent coordinate sampling mechanism to generate watermarks pixel-wise, achieving ultra-high resolution watermark generation with fixed and limited memory and computational resources. This design ensures strong robustness in watermark extraction, even under extreme cropping and scaling distortions. Additionally, we introduce a hierarchical multi-scale coordinate embedding and a low-rank watermark injection strategy to ensure high-quality watermark generation and robust decoding. Experimental results demonstrate that our method significantly outperforms existing schemes in terms of both robustness and computational efficiency while preserving high image quality. Our approach achieves an accuracy greater than 98\% in watermark extraction with only 0.4\% of the image area in 2K images. These results highlight the effectiveness of our method, making it a promising solution for large-scale and high-resolution image watermarking applications.
Luyu Yuan, Han Fang 0004, Yuxing Lu, Sijing Xie, Chengxin Zhao
NeurIPS4
2025 Generalized and Invariant Single-Neuron In-Vivo Activity Representation Learning
abstract
In computational neuroscience, models representing single-neuron in-vivo activity have become essential for understanding the functional identities of individual neurons. These models, such as implicit representation methods based on Transformer architectures, contrastive learning frameworks, and variational autoencoders, aim to capture the invariant and intrinsic computational features of single neurons. The learned single-neuron computational role representations should remain invariant across changing environment and are affected by their molecular expression and location. Thus, the representations allow for in vivo prediction of the molecular cell types and anatomical locations of single neurons, facilitating advanced closed-loop experimental designs. However, current models face the problem of limited generalizability. This is due to batch effects caused by differences in experimental design, animal subjects, and recording platforms. These confounding factors often lead to overfitting, reducing the robustness and practical utility of the models across various experimental scenarios. Previous studies have not rigorously evaluated how well the models generalize to new animals or stimulus conditions, creating a significant gap in the field. To solve this issue, we present a comprehensive experimental protocol that explicitly evaluates model performance on unseen animals and stimulus types. Additionally, we propose a model-agnostic adversarial training strategy. In this strategy, a discriminator network is used to eliminate batch-related information from the learned representations. The adversarial framework forces the representation model to focus on the intrinsic properties of neurons, thereby enhancing generalizability. Our approach is compatible with all major single-neuron representation models and significantly improves model robustness. This work emphasizes the importance of generalization in single-neuron representation models and offers an effective solution, paving the way for the practical application of computational models in vivo. It also shows potential for building unified atlases based on single-neuron in vivo activity.
Yuxing Lu, Zhengrui Guo, Can Liao, Yifan Bu, Fangxu Zhou, Jinzhuo Wang
NeurIPS2
2024 Dual-Color Granularity Alignment for Text-Based Person Search
abstract
Text-based Person Search (TBPS) aims to retrieve the person images based on the given text descriptions. Due to the heterogeneity between modalities and the fine granularity of the person, it is challenging to address the task. Existing methods often overlook granularity consistency across different color channels, which means there’s much potential to enhance retrieval performance. In this paper, we propose a Dual-Color Granularity Alignment (DCGA) method for Text-Based Person Search. DCGA harnesses both color and grayscale information to address issues of color reliance and granularity consistency. Moreover, by employing an improved CR Loss with grayscale information used as an additional weak supervision, DCGA addresses intra-class variance and dataset scarcity. Extensive experiments have demonstrated that our proposed DCGA method achieves state-of-the-art results on all three public datasets.
Yuxing Lu, Ge Jiao
ICASSP2
2024 Concentrated Reasoning and Unified Reconstruction for Multi-Modal Media Manipulation
abstract
Detecting and Grounding Multi-Modal Media Manipulation (DGM4) is an emerging task that aims to identify and locate manipulated elements in both textual and visual media. Given the complexity of this task, the model requires more sophisticated reasoning capabilities to align multi-modal features and capture forgery traces. To this end, we propose a Concentrated reasoning and Unified reconstruction framework (CrUr) for DGM4. Instead of adhering to traditional hierarchical reasoning paradigms, we directly carry out all inference tasks using integrated multi-modal features. Specifically, we extract and align features at a finer granularity, capturing subtle differences that may indicate manipulation by leveraging advanced mask signal modeling. Moreover, to adapt to fine-grained reasoning tasks, we design a transformer-based Reconstruction Harmonizer to facilitate more complex interactions among the reconstructed features, ultimately obtaining integrated features. Experimental results on the DGM4datasets show that our method achieves state-of-the-art performances.
Yuxing Lu, Ge Jiao
ICASSP2
2024 Multiscale Scoring Model for Enhanced Urban Perception Evaluation
abstract
Effective urban management, renewal, and development rely on identifying low-quality areas within the city. However, previous studies have been limited by low-volume handcraft surveys and a dearth of data sources, making it difficult to understand human perception within the urban environment. In this paper, we propose a powerful yet simple scoring model to perform street view image recognition and evaluation which utilizes both global information and feature-level semantic information of street elements, resulting in a high-precision perception model on 6 indexes (Beautiful, Lively, Safe, Wealthy, Boring, and Depressing) from Place Pulse 2.0 dataset. The model is then independently applied to a large-scale and fine-grained evaluation task of 4,384 street view images in Shameen Region, Guangzhou, providing valuable perception details and decision-making support for urban planning for the local government. We believe our work will accelerate the digitization and intelligent transformation of municipal engineering.
Xukai Zhao, Yuxing Lu, Jinzhuo Wang
ICASSP2
2024 Enhancing Multimodal Knowledge Graph Representation Learning through Triple Contrastive Learning
Yuxing Lu, Jinzhuo Wang
IJCAI1
2024 An integrated deep learning approach for assessing the visual qualities of built environments utilizing street view images
Xukai Zhao, Yuxing Lu, Guangsi Lin
Eng. Appl. Artif. Intell.2
2023 MoTIF: a Method for Trustworthy Dynamic Multimodal Learning on Omics
abstract
Omics data are inherently multimodal. The existing multimodal learning methods mainly focus on exploiting complementary information across multiple modalities and integrating them via unified representations. However, few studies have focused on the interpretability of features and modalities and the reliability of results, which are crucial in specific domains such as precision medicine and the life sciences. We propose a Multi-omics Trustworthy Integration Framework (MoTIF) to improve the reliability of multimodal learning models by adding dynamic feature selection and modality selection modules and introducing uncertainty score metrics in the classification process to indicate the reliability of model results, which adhere to our Trustworthy Multimodal Integration (TMI) rule. We conduct exhaustive experiments on five multi-omics datasets derived from TCGA. Results demonstrate that MoTIF can improve the performance of multi-omics classification tasks and provide a more detailed explanation of the model’s internal mechanism and the trustworthiness of the classification results. Code for MoTIF is available at https://github.com/YuxingLu613/MoTIF.
Yuxing Lu, Rui Peng 0006, Jinzhuo Wang, Bingheng Jiang
BIBM1
2023 Multiomics dynamic learning enables personalized diagnosis and prognosis for pancancer and cancer subtypes
abstract
Artificial intelligence (AI) approaches in cancer analysis typically utilize a 'one-size-fits-all' methodology characterizing average patient responses. This manner neglects the diverse conditions in the pancancer and cancer subtypes of individual patients, resulting in suboptimal outcomes in diagnosis and treatment. To overcome this limitation, we shift from a blanket application of statistics to a focus on the explicit recognition of patient-specific abnormalities. Our objective is to use multiomics data to empower clinicians with personalized molecular descriptions that allow for customized diagnosis and interventions. Here, we propose a highly trustworthy multiomics learning (HTML) framework that employs multiomics self-adaptive dynamic learning to process each sample with data-dependent architectures and computational flows, ensuring personalized and trustworthy patient-centering of cancer diagnosis and prognosis. Extensive testing on a 33-type pancancer dataset and 12 cancer subtype datasets underscored the superior performance of HTML compared with static-architecture-based methods. Our findings also highlighting the potential of HTML in elucidating complex biological pathogenesis and paving the way for improved patient-specific care in cancer treatment.
Yuxing Lu, Rui Peng 0006, Lingkai Dong, Renjie Wu 0009, Jinzhuo Wang
Briefings Bioinform.1
2023 MedKPL: A heterogeneous knowledge enhanced prompt learning framework for transferable diagnosis
abstract
Artificial Intelligence (AI) based diagnosis systems have emerged as powerful tools to reform traditional medical care. Each clinician now wants to have his own intelligent diagnostic partner to expand the range of services he can provide. However, the implementation of intelligent decision support systems based on clinical note has been hindered by the lack of extensibility of end-to-end AI diagnosis algorithms. When reading a clinical note, expert clinicians make inferences with relevant medical knowledge, which serve as prompts for making accurate diagnoses. Therefore, external medical knowledge is commonly employed as an augmentation for medical text classification tasks. Existing methods, however, cannot integrate knowledge from various knowledge sources as prompts nor can fully utilize explicit and implicit knowledge. To address these issues, we propose a Medical Knowledge-enhanced Prompt Learning (MedKPL) diagnostic framework for transferable clinical note classification. Firstly, to overcome the heterogeneity of knowledge sources, such as knowledge graphs or medical QA databases, MedKPL uniform the knowledge relevant to the disease into text sequences of fixed format. Then, MedKPL integrates medical knowledge into the prompt designed for context representation. Therefore, MedKPL can integrate knowledge into the models to enhance diagnostic performance and effectively transfer to new diseases by using relevant disease knowledge. The results of our experiments on two medical datasets demonstrate that our method yields superior medical text classification results and performs better in cross-departmental transfer tasks under few-shot or even zero-shot settings. These findings demonstrate that our MedKPL framework has the potential to improve the interpretability and transferability of current diagnostic systems.
Yuxing Lu, Xiaohong Liu 0007, Zongxin Du, Yuanxu Gao
J. Biomed. Informatics1