Yutao Dou

dblp:299/3317 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0001-9990-690XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SMILE: A Scale-aware Multiple Instance Learning Method for Multicenter STAS Lung Cancer Histopathology Diagnosis
abstract
Spread through air spaces (STAS) represents a newly identified aggressive pattern in lung cancer, which is known to be associated with adverse prognostic factors and complex pathological features. Pathologists currently rely on time-consuming manual assessments, which are highly subjective and prone to variation. This highlights the urgent need for automated and precise diagnostic solutions. 2,970 lung cancer tissue slides are comprised from multiple centers, re-diagnosed them, and constructed and publicly released three lung cancer STAS datasets: STAS-CSU (hospital), STAS-TCGA, and STAS-CPTAC. All STAS datasets provide corresponding pathological feature diagnoses and related clinical data. To address the bias, sparse and heterogeneous nature of STAS, we propose an scale-aware multiple instance learning(SMILE) method for STAS diagnosis of lung cancer. By introducing a scale-adaptive attention mechanism, the SMILE can adaptively adjust high-attention instances, reducing over-reliance on local regions and promoting consistent detection of STAS lesions. Extensive experiments show that SMILE achieved competitive diagnostic results on STAS-CSU, diagnosing 251 and 319 STAS samples in CPTAC and TCGA, respectively, surpassing clinical average AUC. The 11 open baseline results are the first to be established for STAS research, laying the foundation for the future expansion, interpretability, and clinical integration of computational pathology technologies. The datasets and code are available at https://github.com/panliangrui/IJCAI25.
Liangrui Pan, Xiaoyu Li 0008, Yutao Dou, Qiya Song, Jiadi Luo, Qingchun Liang, Shaoliang Peng
IJCAI3
2025 Pre-trained Latent Diffusion Model-based Image Enhancement with Cross-Fusion Transformer Network for Expression Recognition
abstract
Facial Expression Recognition (FER) has broad application potential in fields such as education, human-computer interaction, healthcare, and Online monitoring. However, the FER task faces significant challenges, including missing, incomplete, and tilted facial samples in the data, as well as inter-class similarity and intra-class disparity in facial expressions, which pose challenges to recognition. To address these issues, we propose a two-stage method called LDM-POSTER, which integrates Latent Diffusion Models (LDM) and Prompt Engineering for data augmentation. The enhanced images are then fed into a dual-stream Pyramid Cross-Fusion Transformer Network (POSTER) for recognition. Specifically, the LDM and prompt engineering stages effectively repair missing or incomplete facial regions and correct tilted poses, resulting in more complete, clear, and information-rich input data. In the recognition stage, POSTER utilizes a transformer-based cross-fusion structure to integrate facial landmark features with global image features, effectively distinguishing inter-class similarities. By focusing on salient facial regions and employing a pyramid structure to achieve scale invariance, it alleviates the impact of intra-class disparity on recognition accuracy.Within extensive experiments demonstrate that LDM-POSTER achieves significant performance improvements in FER scenarios with missing facial key parts and complex structures, providing a robust and efficient solution for FER.
Yutao Dou, Changfeng He, Xianliang Chen, Jiansong Zhou, Shaoliang Peng
IJCNN2
2025 DepMambaformer: Integrating Bidirectional State Space Duality Model with Multimodal Attention for Depression Detection
Changfeng He, Yutao Dou, Shaoliang Peng
ISBRA (1)2
2025 Spatio-temporal masked autoencoder-based phonetic segments classification from ultrasound
Xi Dan, Kele Xu, Yihang Zhou, Chuanguang Yang, Yutao Dou, Cheng Yang 0004
Speech Commun.6
2024 Autonomous Pharmaceutical Care with Large Language Models
abstract
In modern healthcare systems, physicians often diagnose and prescribe within their specific fields of expertise. However, this compartmentalized approach may not fully assess the overall health condition of patients, leading to potential risks associated with polypharmacy, such as drug-drug interactions. Consequently, pharmaceutical care becomes a critical systemic task, typically managed by clinical pharmacists. Despite their crucial role, clinical pharmacists face challenges in providing personalized medication guidance due to the vast and rapidly evolving body of medical knowledge. With the advancement of artificial intelligence and large language models (LLMs), these technologies have demonstrated significant potential in pharmaceutical care by effectively reducing medication errors and adverse drug events, and alleviating the workload of healthcare professionals. However, existing LLMs still struggle with the accurate interpretation and automated execution of tasks in complex clinical scenarios like pharmaceutical care. To address these challenges, we developed the Shennong-Agent system, a multi-agent framework for LLMs with integrated multimodal inputs. Shennong-Agent is enabled to analyze and segment tasks using chained reasoning and autonomously perform complex pharmacy care tasks using tools such as knowledge retrieval and web search. Rigorously evaluated by medical experts, this system not only surpasses existing LLMs in performance but also enhances its capabilities through reinforcement learning with human feedback.
Yutao Dou, Zike Deng, Tao Xing, Shaoliang Peng
BIBM1
2024 EMO-Mamba: Multimodal Selective Structured State Space Model for Depression Detection
abstract
Depression is a severe mental disorder that significantly affects health and quality of life, making early detection critical for effective treatment and management. In traditional clinical diagnostics, doctors and mental health professionals typically assess emotional states by observing patients’ facial expressions and listening to their speech. Changes in facial expressions can reflect emotional responses, while vocal features such as tone and pace may indicate underlying emotional distress or psychological conditions. Artificial intelligence technologies are adept at efficiently extracting and analyzing these multimedia signals. However, existing methods often overlook critical information when processing long-time series data and struggle to handle multiple modalities simultaneously. To address these issues, we propose the EMO-Mamba, which utilizes selective mechanisms and state space models (SSM) to filter important features and effectively maintain memory over long time-series. Additionally, we introduced a multimodal data fusion framework to integrate key features from various modals, thereby enhancing model performance. Evaluations on the multimodal public datasets D-Vlog showed that our method achieved accuracy of 75.54%, demonstrating its effectiveness across diverse data environments and optimizing detection outcomes.
Tao Xing, Yutao Dou, Jiansong Zhou, Xianliang Chen, Shaoliang Peng
BIBM2
2024 Temporal Inconsistency-Based Active Learning
abstract
Deep supervised learning has demonstrated strong capabilities; however, such progress relies on massive and expensive data annotation. Active Learning (AL) has been introduced to selectively annotate samples, thus reducing the human labeling effort. Previous AL research has focused on employing recently trained models to design sampling strategies, based on uncertainty or representativeness. Drawing inspiration from the issue of model forgetting, we propose a novel AL framework called Temporal Inconsistency-Based Active Learning (TIR-AL). In this framework, multiple snapshots of the models across consecutive cycles are jointly utilized to select samples with higher temporal inconsistency, by computing the proposed self-weighted nuclear norm metric. Furthermore, we introduce a consistency regularization term to mitigate the issue of forgetting. Together, these components make full use of the potential of data and facilitate effective interaction within the AL loop. To demonstrate the efficacy of TIR-AL, we conducted a set of experiments illustrating how our approach outperforms state-of-the-art methods without incurring any additional training costs.
Tianjiao Wan, Yutao Dou, Kele Xu, Zijian Gao, Bo Ding 0001, Huaimin Wang 0001
ICASSP2
2024 Report-Concept Textual-Prompt Learning for Enhancing X-ray Diagnosis
abstract
Despite significant advances in image-text medical visual language modeling, the high cost of fine-grained annotation of images to align radiology reports has led current approaches to focus primarily on semantic alignment between the image and the full report, neglecting the critical diagnostic information contained in the text. This is insufficient in medical scenarios demanding high explainability. To address this problem, in this paper, we introduce radiology reports as images in prompt learning. Specifically, we extract key clinical concepts, lesion locations, and positive labels from easily accessible radiology reports and combine them with an external medical knowledge base to form fine-grained self-supervised signals. Moreover, we propose a novel Report-Concept Textual-Prompt Learning ( RC-TPL ), which aligns radiology reports at multiple levels. In the inference phase, the report-level and concept-level prompts provide rich global and local semantic understanding for X-ray images. Extensive experiments on X-ray image datasets demonstrate the superior performance of our approach with respect to various baselines, especially in the presence of scarce imaging data. Our study not only significantly improves the accuracy of data-constrained medical X-ray diagnosis, but also demonstrates how the integration of domain-specific conceptual knowledge can enhance the explainability of medical image analysis.
Xiongjun Zhao, Guanting Li, Yutao Dou, Shaoliang Peng
ACM Multimedia5
2024 PEACE: A Dataset of Pharmaceutical Care for Cancer Pain Analgesia Evaluation and Medication Decision
abstract
Over half of cancer patients experience long-term pain management challenges. Recently, interest has grown in systems for cancer pain treatment effectiveness assessment (TEA) and medication recommendation (MR) to optimize pharmacological care. These systems aim to improve treatment effectiveness by recommending personalized medication plans based on comprehensive patient information. Despite progress, current systems lack multidisciplinary treatment (MDT) team assessments of treatment and the patient's perception of medication, crucial for effective cancer pain management. Moreover, managing cancer pain medication requires multiple adjustments to the treatment plan based on the patient's evolving condition, a detail often missing in existing datasets. To tackle these issues, we designed the PEACE dataset specifically for cancer pain medication research. It includes detailed pharmacological care records for over 38,000 patients, covering demographics, clinical examination, treatment outcomes, medication plans, and patient self-perceptions. Unlike existing datasets, PEACE records not only long-term and multiple follow-ups both inside and outside hospitals but also includes patients' self-assessments of medication effects and the impact on their lives. We conducted a proof-of-concept study with 13 machine learning algorithms on the PEACE dataset for the TEA (classification task) and MR (regression task). These experiments provide valuable insights into the potential of the PEACE dataset for advancing personalized cancer pain management. The dataset is accessible at: [https://github.com/YTYTYD/PEACE].
Yutao Dou
NeurIPS1
2024 Optimization of the parallel semi-Lagrangian scheme to overlap computation with communication based on grouping levels in YHGSM
Dazheng Liu, Wenjuan Liu, Liangrui Pan, Yutao Dou
CCF Trans. High Perform. Comput.4
2023 An Attention-based Label Mapping and Multi-factor Domain Adaptation Approach for ACS Prediction
abstract
Acute Coronary Syndrome (ACS), an emergent medical condition, is intricately linked to environmental factors like air pollution and meteorological conditions. Harnessing regional environmental data, such as weather metrics, can promptly forecast ACS incidence rates, enabling optimised medical resource allocation and increased patient recovery rates. However, the prediction task is rendered complex due to disparities in data collection capabilities across institutions, yielding datasets with analogous features but significant label variations, impeding the application of universal models. Challenges abound due to the heterogeneity of multi-factor data, temporal alignment disparities, and the intricacies of sparse data. To address these challenges, this paper introduces the Domain Adaptation with Multi-factor Associative Structures (DAMAS), a time-series domain adaptation approach based on multi-factor sparse associative frameworks. Augmented by an isomorphic attention-driven variable label mapping scheme and combined with multi-layer perceptrons, our approach skilfully negotiates label imbalances. This results in refined prediction precision connecting environmental factors to regional ACS incidences.
Yutao Dou, Xiongjun Zhao, Kun Xie 0001, Guo Chen 0001, Shaoliang Peng
BIBM1
2023 ParaMET: A Parallel Framework for Efficient Medical Data Extraction on Tianhe-NG Supercomputer
abstract
In the burgeoning realm of data-driven medical research, the escalating scale and intricacy of contemporary medical datasets frequently surpass the processing capabilities of traditional computational environments. Specifically, I/O bottlenecks have emerged as pivotal constraints in several research areas. In this paper, a data service framework is introduced that harnesses supercomputers to parallelly access multi-modal datasets and supports multi-node processing, called ParaMET. To enhance user accessibility, a web-based user interface has been integrated, allowing permitted researchers to effortlessly interact via their laptops, complemented by an API suite available through SDK for those adept with supercomputing for deeper data manipulation. Through extensive empirical validation, our framework manifests a remarkable performance elevation in multi-node supercomputing settings, achieving acceleration of up to approximately 1000x compared to existing methods.
Yutao Dou, Yangtao Zheng, Dazheng Liu, Keqin Li 0001, Sheng Xiao, Shaoliang Peng
BIBM1
2022 Transformer-Based Unsupervised Learning for Early Detection of Sepsis (Student Abstract)
abstract
A 6-hour early detection of sepsis leads to a significant increase in the chance of surviving it. Previous sepsis early detection studies have focused on improving the performance of supervised learning algorithms while ignoring the potential correlation in data mining, and there was no reliable method to deal with the problem of incomplete data. In this paper, we proposed the Denoising Transformer AutoEncoder (DTAE) for the first time combining transformer and unsupervised learning. DTAE can learn the correlation of the features required for early detection of sepsis without the label. This method can effectively solve the problems of data sparsity and noise and discover the potential correlation of features by adding DTAE enhancement module without modifying the existing algorithms. Finally, the experimental results show that the proposed method improves the existing algorithms and achieves the best results of early detection.
Yutao Dou, Wei Li 0058, Albert Y. Zomaya
AAAI1
2021 LUNAR : Drug Screening for Novel Coronavirus Based on Representation Learning Graph Convolutional Network
abstract
An outbreak of COVID-19 that began in late 2019 was caused by a novel coronavirus(SARS-CoV-2). It has become a global pandemic. As of June 9, 2020, it has infected nearly 7 million people and killed more than 400,000, but there is no specific drug. Therefore, there is an urgent need to find or develop more drugs to suppress the virus. Here, we propose a new nonlinear end-to-end model called LUNAR. It uses graph convolutional neural networks to automatically learn the neighborhood information of complex heterogeneous relational networks and combines the attention mechanism to reflect the importance of the sum of different types of neighborhood information to obtain the representation characteristics of each node. Finally, through the topology reconstruction process, the feature representations of drugs and targets are forcibly extracted to match the observed network as much as possible. Through this reconstruction process, we obtain the strength of the relationship between different nodes and predict drug candidates that may affect the treatment of COVID-19 based on the known targets of COVID-19. These selected candidate drugs can be used as a reference for experimental scientists and accelerate the speed of drug development. LUNAR can well integrate various topological structure information in heterogeneous networks, and skillfully combine attention mechanisms to reflect the importance of neighborhood information of different types of nodes, improving the interpretability of the model. The area under the curve(AUC) of the model is 0.949 and the accurate recall curve (AUPR) is 0.866 using 10-fold cross-validation. These two performance indexes show that the model has superior predictive performance. Besides, some of the drugs screened out by our model have appeared in some clinical studies to further illustrate the effectiveness of the model.
Deshan Zhou, Shaoliang Peng, Wu Zhong, Yutao Dou
IEEE ACM Trans. Comput. Biol. Bioinform.5