Weiwei Dai

dblp:58/1122 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multidimensional Haptic Perception and Quantification Method for Ophthalmic Surgery Training
abstract
Virtual ophthalmic surgical training is a cost-effective and time-efficient paradigm. However, the absence of haptic feedback in virtual ophthalmic surgery systems limits the realism and effectiveness of training. To address this challenge, this study proposes a task-driven method tailored to quantify haptic information in ophthalmic cataract surgery. This work designs a virtual cataract surgery interaction scenario and simulates three key haptic interactions: damping when moving the surgical tool within the eyeball, stiffness during corneal incision and phacoemulsification, and pull during continuous curvilinear capsulorhexis. Using a task-driven and subjective-objective integrated data analysis approach, this work quantifies the just noticeable threshold (JNT) and just noticeable difference (JND). Experimental results demonstrate that the proposed method effectively quantifies the perceptual thresholds and perceptual difference for varied individuals. These findings provide empirical evidence for establishing a more realistic and effective haptic feedback in virtual ophthalmic surgical training.
Yang Gu 0001, He Yuan, Maoyan Li, Yiqiang Chen 0001, Weiwei Dai
Int. J. Hum. Comput. Interact.7
2025 Semantic-oriented Visual Prompt Learning for Class Incremental Learning
abstract
Class-incremental learning (CIL) enables models to continuously learn new classes while addressing catastrophic forgetting. With the introduction of pre-trained models, new tuning paradigms have emerged for CIL. This paper revisits parameter-efficient fine-tuning (PEFT) methods in the context of incremental learning. Prior studies reveal that PEFT methods’ extended parameters do not directly contribute to semantic perception, limiting performance with significant category and domain gaps. To address this, we propose semantic-oriented visual prompt learning (SVPL), which enhances semantic perception and improves task-specific knowledge extraction. SVPL assigns learnable prompts to each class, using a contrastive group alignment to align prompts to task-specific semantic spaces, thus preserving relationships between old and new knowledge. Additionally, hierarchical semantic delivery allows the semantic transformation of prompt groups from shallow to deep layers to facilitate efficient knowledge mining and enable effective learning of new knowledge. Extensive experimental results on five benchmarks demonstrate the superior performance of our methods.
Shuai Guo 0001, Yang Gu 0001, Yingwei Zhang 0002, Weining Weng, Weiwei Dai, Yiqiang Chen 0001
ICASSP7
2025 Distilling Closed-Source LLM's Knowledge for Locally Stable and Economic Biomedical Entity Linking
Yihao Ai, Zhiyuan Ning 0001, Weiwei Dai, Pengfei Wang 0008, Yi Du 0010, Wenjuan Cui, Kunpeng Liu 0001, Yuanchun Zhou
ICIC (25)3
2025 A Geometric Constraints based Bayesian Neural Network for Virtual Surgery Assessment
abstract
Virtual surgical assessment evaluates a surgeon’s technical skills by analyzing instrument motion data collected during virtual surgery. Although existing methods classify skills accurately from motion data, they struggle to generalize to unfamiliar surgical patterns because surgeons use varied techniques. To address these limitations, this paper proposes a Geometric Constraints based Bayesian Neural Network (GeomBNN) model for surgical assessment. The geometric constraints enforce orthogonality among nonlinear components of the motion data, enhancing the model’s capacity to capture complex surgical patterns. The Bayesian neural network’s probabilistic framework quantifies predictive uncertainty, enabling adaptation to complex procedures and improving assessment robustness. Experiments on four datasets show that GeomBNN achieves higher classification accuracy than existing methods. What is more, the model exhibits progressive feature attention, dynamically focusing on the most discriminative surgical skill features during training, further validating its effectiveness.
He Yuan, Weiwei Dai, Danmin Cao, Yang Gu 0001
IJCNN3
2025 Grade-Skewed Domain Adaptation via Asymmetric Bi-Classifier Discrepancy Minimization for Diabetic Retinopathy Grading
abstract
Diabetic retinopathy (DR) is a leading cause of preventable low vision worldwide. Deep learning has exhibited promising performance in the grading of DR. Certain deep learning strategies have facilitated convenient regular eye check-ups, which are crucial for managing DR and preventing severe visual impairment. However, the generalization performance on cross-center, cross-vendor, and cross-user test datasets is compromised due to domain shift. Furthermore, the presence of small lesions and the imbalanced grade distribution, resulting from the characteristics of DR grading (e.g., the progressive nature of DR disease and the design of grading standards), complicates image-level domain adaptation for DR grading. The general predictions of the models trained on grade-skewed source domains will be significantly biased toward the majority grades, which further increases the adaptation difficulty. We formulate this problem as a grade-skewed domain adaptation challenge. Under the grade-skewed domain adaptation problem, we propose a novel method for image-level supervised DR grading via Asymmetric Bi-Classifier Discrepancy Minimization (ABiD). First, we propose optimizing the feature extractor by minimizing the discrepancy between the predictions of the asymmetric bi-classifier based on two classification criteria to encourage the exploration of crucial features in adjacent grades and stretch the distribution of adjacent grades in the latent space. Moreover, the classifier difference is maximized by using the forward and inverse distribution compensation mechanism to locate easily confused instances, which avoids pseudo-label bias on the target domain. The experimental results on two public DR datasets and one private DR dataset demonstrate that our method outperforms state-of-the-art methods significantly.
Yang Gu 0001, Shuai Guo 0001, Shijie Wen, Nianfeng Shi, Weiwei Dai, Yiqiang Chen 0001
IEEE Trans. Medical Imaging7
2024 EyeGraphGPT: Knowledge Graph Enhanced Multimodal Large Language Model for Ophthalmic Report Generation
abstract
Automatic generation of ophthalmic reports holds significant potential to lessen clinicians’ workload, enhance work efficiency, and alleviate the imbalance between clinicians and patients. Recent advancements in multimodal large language models, represented by GPT-4, have demonstrated remarkable performance in the general domain. However, training such models necessitates a substantial amount of paired image-text data, yet paired ophthalmic data is limited, and ophthalmic reports are laden with specialized terminologies, making it challenging to transfer the training paradigm to the ophthalmic domain. In this paper, we propose EyeGraphGPT, a knowledge graph enhanced multimodal large language model for ophthalmic report generation. Specifically, we construct a knowledge graph by leveraging the knowledge from a medical database and expertise from ophthalmic experts to model relationships among ophthalmic diseases, enhancing the model’s focus on key disease information. We then perform relation-aware modal alignment to incorporate knowledge graph features into visual features, and further enhance modality collaboration through visual instruction fine-tuning to adapt the model to the ophthalmic domain. Our experiments on a real-world dataset demonstrates that EyeGraphGPT outperforms previous state-of-the-art models, highlighting its superiority in scenarios with limited medical data and extensive specialized terminologies.
Xinlong Jiang, Chenlong Gao, Weiwei Dai, Bingyu Wang, Bingjie Yan, Wuliang Huang
BIBM5
2024 Information Retrieval Optimization for Non-Exemplar Class Incremental Learning
abstract
Existing non-example class-incremental learning (NECIL) methods usually utilize a combination strategy of replay mechanism and knowledge distillation. However, this combination strategy only focuses on the preservation of old information quantitatively, ignoring the preservation quality. When the old knowledge has wrong redundant information, catastrophic forgetting is more likely to occur. Therefore, obtaining adequate information without impurities as much as possible and removing invalid or even harmful information has become an effective solution to improve the performance of NECIL. This process is consistent with the information bottleneck (IB) theory. Thus, we propose a new NECIL method based on the IB framework. By using the different information obtained from the new and old class samples and the implicit knowledge in the teacher model training process, the error of harmful redundant information learned is eliminated. Specifically, we propose two optimization strategies that align with the two optimization processes of the information bottleneck. Firstly, we employ a pseudo-prototype selection mechanism that selectively incorporates pseudo-samples into the learning process of new and old categories, thus enhancing the distinction between new and old categories and diminishing the mutual information between the input and intermediate features. Secondly, we introduce an attention-based feature distillation method that regulates the distillation strength between feature pairs based on their similarity, thereby augmenting the mutual information between intermediate features and output prediction. Extensive experiments on three benchmarks demonstrate that the proposed method exhibits significant incremental performance improvements over existing methods.
Shuai Guo 0001, Yang Gu 0001, Yingwei Zhang 0002, Weining Weng, Weiwei Dai, Yiqiang Chen 0001
CIKM7
2024 PrivFusion: Privacy-Preserving Model Fusion via Decentralized Federated Graph Matching
abstract
Model fusion is becoming a crucial component in the context of model-as-a-service scenarios, enabling the delivery of high-quality model services to local users. However, this approach introduces privacy risks and imposes certain limitations on its applications. Ensuring secure model exchange and knowledge fusion among users becomes a significant challenge in this setting. To tackle this issue, we propose PrivFusion, a novel architecture that preserves privacy while facilitating model fusion under the constraints of local differential privacy. PrivFusion leverages a graph-based structure, enabling the fusion of models from multiple parties without additional training. By employing randomized mechanisms, PrivFusion ensures privacy guarantees throughout the fusion process. To enhance model privacy, our approach incorporates a hybrid local differentially private mechanism and decentralized federated graph matching, effectively protecting both activation values and weights. Additionally, we introduce a perturbation filter adapter to alleviate the impact of randomized noise, thereby recovering the utility of the fused model. Through extensive experiments conducted on diverse image datasets and real-world healthcare applications, we provide empirical evidence showcasing the effectiveness of PrivFusion in maintaining model performance while preserving privacy. Our contributions offer valuable insights and practical solutions for secure and collaborative data analysis within the domain of privacy-preserving model fusion.
Qian Chen 0023, Yiqiang Chen 0001, Xinlong Jiang, Weiwei Dai, Wuliang Huang, Bingjie Yan, Wang Lu 0003
IEEE Trans. Knowl. Data Eng.5
2023 RDKG: A Reinforcement Learning Framework for Disease Diagnosis on Knowledge Graph
abstract
Automatic disease diagnosis from symptoms has attracted much attention in medical practices. It can assist doctors and medical practitioners in narrowing down disease candidates, reducing testing costs, improving diagnosis efficiency, and more importantly, saving human lives. Existing research has made significant progress in diagnosing disease but was limited by the gap between interpretability and accuracy. To fill this gap, in this paper, we propose a method called Reinforced Disease Diagnosis on Knowlege Graph (RDKG). Specifically, we first construct a knowledge graph containing all information from electronic medical records. To capture informative embeddings, we propose an enhanced knowledge graph embedding method that can embed information outside the knowledge graph into entity embedding. Then we transform the automatic disease diagnosis task into a Markov decision process on the knowledge graph. After that, we design a reinforcement learning method with a soft reward mechanism and a pruning strategy to solve the Markov decision process. We accomplish automated disease diagnosis by finding a path from symptoms to disease. The experimental results show that our model can effectively utilize heterogeneous information in the knowledge graph to complete the automatic disease diagnosis. Besides, our model demonstrates supreme performance in both accuracy and interpretability.
Shipeng Guo, Kunpeng Liu 0001, Pengfei Wang 0008, Weiwei Dai, Yi Du 0010, Yuanchun Zhou, Wenjuan Cui
ICDM4
2023 NEEDED: Introducing Hierarchical Transformer to Eye Diseases Diagnosis
abstract
With the development of natural language processing tech- niques(NLP), automatic diagnosis of eye diseases using ophthalmology electronic medical records (OEMR) has become possible. It aims to evaluate the condition of both eyes of a patient respectively, and we formulate it as a particular multi-label classification task in this paper. Although there are a few related studies in other diseases, automatic diagnosis of eye diseases exhibits unique characteristics. First, descriptions of both eyes are mixed up in OEMR documents, with both free text and templated asymptomatic descriptions, resulting in sparsity and clutter of information. Second, OEMR documents contain multiple parts of descriptions and have long document lengths. Third, it is critical to provide explainability to the disease diagnosis model. To overcome those challenges, we present an effective automatic eye disease diagnosis framework, NEEDED. In this framework, a preprocessing module is integrated to improve the density and quality of information. Then, we design a hierarchical transformer structure for learning the contextualized representations of each sentence in the OEMR document. For the diagnosis part, we propose an attention-based predictor that enables traceable diagnosis by obtaining disease-specific information. Experiments on the real dataset and comparison with several baseline models show the advantage and explainability of our framework.
Xu Ye, Meng Xiao 0001, Zhiyuan Ning 0001, Weiwei Dai, Wenjuan Cui, Yi Du 0010, Yuanchun Zhou
SDM4
2023 GLIM-Net: Chronic Glaucoma Forecast Transformer for Irregularly Sampled Sequential Fundus Images
abstract
Chronic Glaucoma is an eye disease with progressive optic nerve damage. It is the second leading cause of blindness after cataract and the first leading cause of irreversible blindness. Glaucoma forecast can predict future eye state of a patient by analyzing the historical fundus images, which is helpful for early detection and intervention of potential patients and avoiding the outcome of blindness. In this paper, we propose a GLaucoma forecast transformer based on Irregularly saMpled fundus images named GLIM-Net to predict the probability of developing glaucoma in the future. The main challenge is that the existing fundus images are often sampled at irregular times, making it difficult to accurately capture the subtle progression of glaucoma over time. We therefore introduce two novel modules, namely time positional encoding and time-sensitive MSA (multi-head self-attention) modules, to address this challenge. Unlike many existing works that focus on prediction for an unspecified future time, we also propose an extended model which is further capable of prediction conditioned on a specific future time. The experimental results on the benchmark dataset SIGF show that the accuracy of our method outperforms the state-of-the-art models. In addition, the ablation experiments also confirm the effectiveness of the two modules we propose, which can provide a good reference for the optimization of Transformer models.
Xiaoyan Hu 0010, Ling-Xiao Zhang, Lin Gao 0004, Weiwei Dai, Xiaoguang Han 0001, Yukun Lai, Yiqiang Chen 0001
IEEE Trans. Medical Imaging4