Feiyan Liu

dblp:134/7425 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DSFDU: Detection of unicode modifier letter obfuscated commands in Living-Off-the-Land attacks
Feiyan Liu, Guanglu Sun
Comput. Secur.2
2025 Fooling Machine's Eyes: Unicode Modifier Letter Evasion Attack
abstract
By analyzing command-line arguments during process execution, Endpoint Detection and Response (EDR) and Security Information and Event Management (SIEM) tools can detect potential malicious commands and identify Advanced Persistent Threats (APT). In practice, most implementations rely on heuristics and regular expression matching. However, attackers can bypass detection by obfuscating commands using Unicode modifier letters. This paper investigates the mechanism behind this evasion technique. Through reverse engineering, we identify the root cause and discover an internationalization API vulnerability in Windows. The impact assessment reveals that 424 system programs are potentially at risk. We further examine command arguments containing internationalized domain names with modifier letters. Code analysis traces how these domains resolve across operating systems, exposing additional attack surfaces. Based on these findings, we formally define modifier letter evasion attacks. We then propose a threat model, which allows attackers to construct evasion commands that remain executable while bypassing command-line detection. Furthermore, we design evasion cases and validate them through proof-of-concept experiments. Results show that high-risk cases can bypass certain leading EDR and SIEM tools. To mitigate this attack, we propose a detection approach, develop a corresponding tool, and suggest defense measures.
Chao Gao 0021, Guanglu Sun, Feiyan Liu
ACSAC4
2025 Seeing Through Ambiguity: Effective Video-guided Machine Translation via Chaotic Fusion and Causally Aligned Spatio-temporal Attention
abstract
Video-guided machine translation (VMT) involves taking text and video modalities as inputs, leveraging visual context to resolve the semantic ambiguities for improving the translation quality. This task remains challenging due to the difficulty of effective cross-modal integration and visual grounding. To address the issues, we propose a novel VMT model that combines temporal video and spatial keyframe streams by providing complementary visual cues. We develop a chaotic fusion mechanism to integrate modality-specific representations from various modalities that help capture semantic interactions between visual and textual cues. To improve visual grounding, a causally aligned spatio-temporal attention mechanism is also designed to enhance semantic alignment by refining decoder-side attention over the video and keyframe streams, respectively. We further propose PolyVTE, an evaluation dataset targeting polysemous ambiguities in VMT. Results on VATEX and PolyVTE datasets show that our model outperforms state-of-the-art models. The results also prove that using keyframe and video modalities significantly improves disambiguation capabilities. The PolyVTE dataset is available at https://github.com/zheng5d/PolyVTE.
Feiyan Liu, Xiaoli Wang 0002
ACM Multimedia2
2025 EDRMM: enhancing drug recommendation via multi-granularity and multi-attribute representation
abstract
BACKGROUND: Drug recommendation is a crucial application of artificial intelligence in medical practice. Although many models have been proposed to solve this task, two challenges remain unresolved: (i) most existing models use all historical visits as input, overlooking fine-grained correlations between historical and current information; (ii) Electronic Health Records (EHRs) are underutilized, with only partial information considered to describe patient conditions. To tackle the challenges, we propose a novel drug recommendation model, denoted by EDRMM, which incorporates multi-granularity and multi-attribute information into representation learning. We develop a longitudinal attribute-level history selection mechanism to effectively identify fine-grained historical information that is highly relevant to a patient's current clinical conditions. We analyze the impact of key Electronic Health Record (EHR) attributes, demonstrating that incorporating such attributes into patient representations can further boost performance. We also design an adaptive global Drug-Drug Interaction (DDI) risk regularization term for the DDI loss function to better balance accuracy and safety during training. RESULTS: Experimental results show that our model achieves state-of-the-art performance on a widely used MIMIC-III dataset. CONCLUSIONS: EDRMM overcomes two key drug recommendation limitations through three innovations: (1) Dynamic attribute-level history selection, which retrieves relevant features and filters out noise, (2) The integration of multi-attribute EHR with attribute-specific encoding strategies to generate comprehensive patient representations, and (3) Hybrid optimization balancing accuracy and safety via adaptive DDI regularization. The combination of these three innovations enables the proposed EDRMM to achieve the best recommendation performance on the MIMIC-III dataset.
Feiyan Liu, Yibo Xie, Dongxiang Zhang
BMC Bioinform.1
2025 Using domain-specific keyword features to enhance deep learning-based pressure vessel inspection problem identification
Feiyan Liu, Yibin Jin, Zechen Liu
Eng. Appl. Artif. Intell.3
2024 MHGRL: An Effective Representation Learning Model for Electronic Health Records
abstract
Electronic health records (EHRs) serve as a digital repository storing comprehensive medical information about patients. Representation learning for EHRs plays a crucial role in healthcare applications. In this paper, we propose a Multimodal Heterogeneous Graph-enhanced Representation Learning, denoted as MHGRL, aimed at learning effective EHR representations. To address the challenge posed by data insufficiency of EHRs, MHGRL utilizes a multimodal heterogeneous graph to model an EHR. Specifically, we construct a heterogeneous graph for each EHR and enrich it by incorporating multimodal information with medical ontology and textual notes. With the integration of pre-trained model, graph neural network, and attention mechanism, MHGRL effectively incorporates both node attributes and structural information across a multimodal heterogeneous graph. Moreover, we employ contrastive learning to ensure the consistency of representations for similar EHRs and improve the model robustness. The experimental results show that MHGRL outperforms all baselines on two real clinical datasets in downstream tasks, including EHR clustering and disease prediction. The code is available at https://github.com/emmali808/MHGRL.
Feiyan Liu, Liangzhi Li 0004, Xiaoli Wang 0002, Jinsong Su, Yiming Qian
LREC/COLING1
2024 BESTMVQA: A Benchmark Evaluation System for Medical Visual Question Answering
Xiaojie Hong, Zixin Song, Liangzhi Li 0004, Xiaoli Wang 0002, Feiyan Liu
ECML/PKDD (9)5
2024 OEHR: An Orthopedic Electronic Health Record Dataset
Yibo Xie, Kaifan Wang, Feiyan Liu, Xiaoli Wang 0002, Guofeng Huang
SIGIR4
2022 OVQA: A Clinically Generated Visual Question Answering Dataset
abstract
Medical visual question answering (Med-VQA) is a challenging problem that aims to take a medical image and a clinical question about the image as input and output a correct answer in natural language. Current medical systems often require large-scale and high-quality labeled data for training and evaluation. To address the challenge, we present a new dataset, denoted by OVQA, which is generated from electronic medical records. We develop a semi-automatic data generation tool for constructing the dataset. First, medical entities are automatically extracted from medical records and filled into predefined templates for generating question and answer pairs. These pairs are then combined with medical images extracted from corresponding medical records, to generate candidates for visual question answering (VQA). The candidates are finally verified with high-quality labels annotated by experienced physicians. To evaluate the quality of OVQA, we conduct comprehensive experiments on state-of-the-art methods for the Med-VQA task to our dataset. The results show that our OVQA can be used as a benchmarking dataset for evaluating existing Med-VQA systems. The dataset can be downloaded from http://47.94.174.82/.
Yefan Huang, Xiaoli Wang 0002, Feiyan Liu, Guofeng Huang
SIGIR3