VLDB 2026 Research / reviewers in the wild / expert
Yi Guan
dblp:69/5722
· DBLP profile ↗
68ranked-venue papers
1as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 20 since 2021Artificial intelligence and machine learning · 29 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Invariant Feature Learning for Counterfactual Watch-time Prediction in Video RecommendationabstractVideo recommendation systems heavily rely on user watch time feedback, making accurate watch time prediction a crucial task. However, this task inherently suffers from bias, as recommendation models tend to favor long-duration videos to maximize watch time. This issue, known as duration bias in the watch-time prediction context, can be explained from a causal perspective, where video duration acts as a confounder. Recent works address this bias using backdoor adjustment, isolating the direct effect of content on watch time from observational data. These methods typically discretize video duration into groups, estimate group-wise effects, and then aggregate them via a unified prediction model. However, this aggregation strategy is prone to model misspecification due to feature distribution shift across groups. In this paper, we reinterpret the problem through the lens of invariant learning and propose a novel framework: Duration-Invariant Feature Learning (DIFL). DIFL employs a kernel-based regularization that enforces representation invariance across duration groups, reducing sensitivity to group design and improving generalization. This enables more accurate modeling of the direct causal effect and making counterfactual inference. Extensive experiments on both public and real large-scale production datasets demonstrate the effectiveness of our approach, which achieves SOTA performance. Chenghou Jin, Yixin Ren, Hongxu Ma 0001, Yewei Xia, Yi Guan, Hao Zhang 0079, Jiandong Ding, Jihong Guan, Shuigeng Zhou |
AAAI | 5 |
| 2026 | AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Modelsabstractn the agricultural domain, the deployment of large language models (LLMs) is hindered by the lack of training data and evaluation benchmarks. To mitigate this issue, we propose AgriEval, the first comprehensive Chinese agricultural benchmark with three main characteristics: (1) Comprehensive Capability Evaluation. AgriEval covers six major agriculture categories and 29 subcategories within agriculture, addressing four core cognitive scenarios—memorization, understanding, inference, and generation. (2) High-Quality Data. The dataset is curated from university-level examinations and assignments, providing a natural and robust benchmark for assessing the capacity of LLMs to apply knowledge and make expert-like decisions. (3) Diverse Formats and Extensive Scale. AgriEval comprises 14,697 multiple-choice questions and 2,167 open-ended question-and-answer questions, establishing it as the most extensive agricultural benchmark available to date. We also present comprehensive experimental results over 51 open-source and commercial LLMs. The experimental results reveal that most existing LLMs struggle to achieve 60 percent accuracy, underscoring the developmental potential in agricultural LLMs. Additionally, we conduct extensive experiments to investigate factors influencing model performance and propose strategies for enhancement. Lian Yan, Haotian Wang 0007, Tianyang Sun, Liangliang Liu 0002, Yi Guan, Jingchi Jiang |
AAAI | 7 |
| 2026 | Stabilized neural ordinary differential equation for text classification in natural language processing
Linfang Dai, Shiqin Ou, Zhenyuan Guo, Shiping Wen 0001, Yi Guan |
Neurocomputing | 5 |
| 2025 | Agri-CM³: A Chinese Massive Multi-modal, Multi-level Benchmark for Agricultural Understanding and ReasoningabstractMulti-modal Large Language Models (MLLMs) integrating images, text, and speech can provide farmers with accurate diagnoses and treatment of pests and diseases, enhancing agricultural efficiency and sustainability. However, existing benchmarks lack comprehensive evaluations, particularly in multi-level reasoning, making it challenging to identify model limitations. To address this issue, we introduce Agri-CM^3, an expert-validated benchmark assessing MLLMs’ understanding and reasoning in agricultural management. It includes 3,939 images and 15,901 multi-level multiple-choice questions with detailed explanations. Evaluations of 45 MLLMs reveal significant gaps. Even GPT-4o achieves only 63.64% accuracy, falling short in fine-grained reasoning tasks. Analysis across three reasoning levels and seven compositional abilities highlights key challenges in accuracy and cognitive understanding. Our study provides insights for advancing MLLMs in agricultural management, driving their development and application. Code and data are available at https://github.com/HIT-Kwoo/Agri-CM3. Haotian Wang 0007, Yi Guan, Fanshu Meng, Chao Zhao 0002, Lian Yan, Yang Yang 0041, Jingchi Jiang |
ACL (1) | 2 |
| 2025 | T1D-MLLM: Multimodal Large Language Model and Cross-Scenario Dataset for Multi-Scenario Management of Type 1 DiabetesabstractThe management of Type 1 Diabetes (T1D) involves comprehensive scenarios, including blood glucose prediction, risk assessment, and insulin dosing control. However, the heterogeneity of information requirements, task objectives, and behavioral logic poses challenges in constructing unified T1D management systems with conventional deep learning models, mainly attributed to insufficient capability of feature alignment and lack of high-quality multi-scenario T1D data. In this paper, we propose T1D-MLLM, the first multimodal large language model designed for unified multi-scenario T1D management, as well as construct LCT1D, a large-scale and cross-scenario T1D dataset. Specifically, T1D-MLLM integrates time-series physiological data with natural language descriptions to capture longterm dependencies across multiple management scenarios while also enhancing fine-grained perception of time-series. Meanwhile, to overcome data scarcity, we proposed a multimodal data generation paradigm based on expert strategies. By constructing task templates and applying a rule-driven alignment mechanism, we generated 150,000 high-quality expert samples with individualized physiological parameters, which provide rich and diverse training samples, significantly improving the T1D-MLLM's capabilities in heterogeneous feature alignment and cross-scenario inference. These experiments demonstrate the effectiveness of the T1D-MLLM in multi-scenarios of various tasks as a unified system, with an excellent performance that surpasses both opensource and proprietary models. Liangliang Liu 0002, Yi Guan, Rujia Shen, Guowei Zheng, Chaoran Kong, Jingchi Jiang |
BIBM | 2 |
| 2025 | Blood Glucose Forecasting Via Fusing Intra- and Inter-Variable VariationsabstractBlood glucose (BG) forecasting aims to help people with Type-1 diabetes (T1D) avoid hyperglycemia or hypoglycemia, which plays a crucial role in medical monitoring. Despite the advancements in deep learning methods for BG forecasting, their ability to predict long-term time series remains limited, and they cannot fully meet the demand for BG forecasting. This limitation stems from the failure to account for both intra- and inter-variable variations simultaneously. To address this challenge, we introduce the$\text{Fi}^{2}$VBlock, which exploits the frequency perspective to fuse intra- and intervariable variations. After transforming to the frequency domain using the Frequency Transform Module, the Frequency Cross Attention between the real and imaginary parts is designed to obtain enhanced frequency representations and capture intravariable variations. In addition, inception blocks are employed to integrate information, thus capturing correlations across different variables. Our backbone network,$\text{Fi}^{2} \mathrm{V}$, employs a residual architecture by concatenating multiple$\text{Fi}^{2}$VBlocks, thereby avoiding degradation problems. Experimental evaluations reveal that$\text{Fi}^{2} \mathrm{V}$outperforms other baselines on the T1DMS and Dnurse datasets and demonstrates zero-shot generalization across patients. Rujia Shen, Yi Guan, Liangliang Liu 0002, Jingchi Jiang |
BIBM | 2 |
| 2025 | AD2QT: Online Task Allocation Based on Transformer and Deep Reinforcement Learning in Mobile Crowdsensing
Yuhong Tan, Tao Peng 0011, Zeyu Chi, Xingyi Wu, Yi Guan |
ICIC (9) | 5 |
| 2025 | Causal discovery based on hierarchical reinforcement learning
Jingchi Jiang, Rujia Shen, Chao Zhao 0002, Yi Guan, Xuehui Yu, Xuelian Fu |
Expert Syst. Appl. | 4 |
| 2025 | Modeling clinical thinking based on knowledge hypergraph attention network and prompt learning for disease prediction
Yang Yang 0137, Xin Li 0012, Haotian Wang 0007, Yi Guan, Jingchi Jiang |
Expert Syst. Appl. | 5 |
| 2025 | Learning to break: Knowledge-enhanced reasoning in multi-agent debate system
Haotian Wang 0007, Xiyuan Du, Weijiang Yu, Qianglong Chen, Kun Zhu 0025, Lian Yan, Yi Guan |
Neurocomputing | 8 |
| 2025 | Quality-Controllable automatic construction method of Chinese knowledge graph for medical decision-making applications
Yang Yang 0137, Yi Guan, Haotian Wang 0007, Jingchi Jiang, Huaizhang Shi, Xiguang Liu |
Inf. Process. Manag. | 4 |
| 2025 | Knowledge assimilation: Implementing knowledge-guided agricultural large language model
Jingchi Jiang, Lian Yan, Zhenbo Xia, Haotian Wang 0007, Yang Yang 0137, Yi Guan |
Knowl. Based Syst. | 7 |
| 2025 | Hierarchical Causal Discovery From Large-Scale Observed VariablesabstractIt is a long-standing question to discover causal relations from observed variables in many empirical sciences. However, current causal discovery methods are inefficient when dealing with large-scale observed variables due to challenges in conditional independence (CI) tests or complex computations of acyclicity, and may even fail altogether. To address the efficiency issue in causal discovery from large-scale observed variables, we propose a Hierarchical Causal Discovery (HCD) framework with a bilevel policy that handles this issue by boosting existing models. Specifically, the high-level policy first finds a causal cut set to partition observed variables into several causal clusters and releases the clusters to the low-level policy. The low-level policy applies any causal discovery method to process these causal clusters in parallel and obtain intra-cluster structures for subsequently inter-cluster structure merging in the high-level policy. To avoid missing inter-cluster edges, we theoretically demonstrate the feasibility of causal cluster cut and inter-cluster structure merging. We also prove the completeness and correctness of HCD for causal discovery. Experiments on both synthetic and real-world datasets demonstrate that HCD consistently and significantly enhances the efficiency and effectiveness of existing advanced methods. Rujia Shen, Muhan Li, Chao Zhao 0002, Boran Wang, Yi Guan, Jie Liu 0001, Jingchi Jiang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Dialogues Are Not Just Text: Modeling Cognition for Dialogue Coherence EvaluationabstractThe generation of logically coherent dialogues by humans relies on underlying cognitive abilities. Based on this, we redefine the dialogue coherence evaluation process, combining cognitive judgment with the basic text to achieve a more human-like evaluation. We propose a novel dialogue evaluation framework based on Dialogue Cognition Graph (DCGEval) to implement the fusion by in-depth interaction between cognition modeling and text modeling. The proposed Abstract Meaning Representation (AMR) based graph structure called DCG aims to uniformly model four dialogue cognitive abilities. Specifically, core-semantic cognition is modeled by converting the utterance into an AMR graph, which can extract essential semantic information without redundancy. The temporal and role cognition are modeled by establishing logical relationships among the different AMR graphs. Finally, the commonsense knowledge from ConceptNet is fused to express commonsense cognition. Experiments demonstrate the necessity of modeling human cognition for dialogue evaluation, and our DCGEval presents stronger correlations with human judgments compared to other state-of-the-art evaluation metrics. Jia Su 0004, Yang Yang 0137, Zipeng Gao, Xinyu Duan, Yi Guan |
AAAI | 6 |
| 2024 | ARRS: Adaptive Representation and Relevance Scoring Enhance Whole Slide Image Classification using Multi-Instance LearningabstractThe classification of whole slide images (WSIs) is crucial in computational pathology and has significant clinical implications. Due to the extremely high resolution of WSIs and the lack of detailed lesion annotations, Multiple instance learning (MIL) has recently shown great promise for WSI classification by modeling WSIs as "bags" and treating cropped patches as "instances". However, using pre-trained feature extractors often leads to biased instance representations as the data used to pre-train these models differ significantly from histopathology data. Furthermore, since focusing on only certain instances may lead to overlooking important details, it is crucial to comprehensively assess the relevance of all instances for positive instance selection. In this paper, we propose a weakly supervised method to enhance WSI classification using adaptive representation and an instance relevance scoring strategy. To address the issue of biased data representation, we introduce an adaptive representation designed to enhance features relevant to lesion regions. This involves an adaptive block that transforms input features to better represent these critical characteristics, while simultaneously applying an attention-based probability distribution to maintain consistency between the transformed features. Additionally, we propose an instance relevance scoring strategy that assigns importance scores to each instance based on its contribution to the classification. Two publicly available datasets, CAMELYON-16 and TCGA-NSCLC, are used to validate the proposed method. The experimental results show that our proposed method outperforms existing state-of-the-art approaches in WSI classification. Chaoran Kong, Jingchi Jiang, Yi Guan, Xiguang Liu, Haiyan You, Yunyun Cao, Yang Yang 0041 |
BIBM | 4 |
| 2024 | Forecasting Influenza Like Illness based on White-Box TransformersabstractInfluenza seriously endangers human health and even causes a large number of deaths every year. Transformers for Influenza-like illness (ILI) forecasting have recently been proven effective. However, these end-to-end deep models are mathematically almost black-box, hindering us from inferring the specific roles and functionalities of each layer within the models, which constitutes a common key challenge in deep neural networks. At the same time, explainability helps trust and use AI systems effectively. In this paper, we propose an efficient ILI forecasting framework incorporating a patching design and variable-channel pairs, which can accommodate any white-box transformer, thereby endowing ILI forecasting with both interpretability and analytical accuracy. Through extensive experimental validation, leveraging white-box transformers such as CRATE, our White-box Time Series Transformer (WhiteTST) framework achieves the state-of-the-art accuracy on ILI datasets. We visually present the self-attention maps within WhiteTST to indicate further explainability. Our results suggest a pathway for designing white-box foundational models for ILI forecasting that concurrently exhibit high accuracy and interpretability. The code is available online in https://github.com/HITshenrj/WhiteTST. Rujia Shen, Yaoxiong Lin, Boran Wang, Liangliang Liu 0002, Yi Guan, Jingchi Jiang |
BIBM | 5 |
| 2024 | An interactive food recommendation system using reinforcement learning
Liangliang Liu 0002, Yi Guan, Rujia Shen, Guowei Zheng, Xuelian Fu, Xuehui Yu, Jingchi Jiang |
Expert Syst. Appl. | 2 |
| 2024 | Knowledge-based dynamic prompt learning for multi-label disease diagnosis
Jing Xie 0012, Xin Li 0012, Yi Guan, Jingchi Jiang, Xitong Guo |
Knowl. Based Syst. | 4 |
| 2024 | Robust expansion of phylogeny for fast-growing genome sequence dataabstractMassive sequencing of SARS-CoV-2 genomes has urged novel methods that employ existing phylogenies to add new samples efficiently instead of de novo inference. 'TIPars' was developed for such challenge integrating parsimony analysis with pre-computed ancestral sequences. It took about 21 seconds to insert 100 SARS-CoV-2 genomes into a 100k-taxa reference tree using 1.4 gigabytes. Benchmarking on four datasets, TIPars achieved the highest accuracy for phylogenies of moderately similar sequences. For highly similar and divergent scenarios, fully parsimony-based and likelihood-based phylogenetic placement methods performed the best respectively while TIPars was the second best. TIPars accomplished efficient and accurate expansion of phylogenies of both similar and divergent sequences, which would have broad biological applications beyond SARS-CoV-2. TIPars is accessible from https://tipars.hku.hk/ and source codes are available at https://github.com/id-bioinfo/TIPars. Yongtao Ye, Marcus H. Shum, Joseph L. Tsui, Guangchuang Yu, Huachen Zhu, Joseph T. Wu, Yi Guan, Tommy Tsan-Yuk Lam |
PLoS Comput. Biol. | 8 |
| 2024 | EIRAD: An Evidence-Based Dialogue System With Highly Interpretable Reasoning Path for Automatic DiagnosisabstractDialogue System for Medical Diagnosis (DSMD) based on reinforcement learning (RL) can simulate patient-doctor interactions, playing a crucial role in clinical diagnosis. However, due to the complexity of disease etiology, DSMD faces the challenges of low efficiency in diagnostic evidence search. Moreover, solely RL-based DSMS, without the constraints of professional medical knowledge, often generates irrational, meaningless, or even erroneous symptom inquiries, leading to poor interpretability of diagnostic path and high misdiagnosis rates. To address these issues, we propose anEvidence-based dialogue system with highlyInterpretableReasoning path forAutomaticDiagnosis (EIRAD) grounded in medical knowledge graph (MKG). Specifically, our automated diagnostic model captures key symptoms for suspected diseases by explicitly leveraging the topology of MKG, enhancing the interpretability and accuracy of diagnosis. To expedite the retrieval of factual evidence, we develop two mechanisms: 1) Mapping mechanism between the entity set of MKG and DSMD's diagnostic evidence and diseases. According to the patient's symptoms, EIRAD prunes irrelevant disease and symptom nodes from the MKG, which can truncate the invalid action of RL-based DSMD. 2) Reward Mechanism of integrating the effectiveness of symptom inquiry and the accuracy of disease diagnosis. The comprehensive reward system is suitable for intelligent consultation, which can effectively drive DSMD to accelerate evidence collection. Experimental results demonstrate that our model significantly outperforms competitive benchmark methods in symptom inquiry efficiency and diagnostic accuracy. Lian Yan, Yi Guan, Haotian Wang 0007, Yang Yang 0137, Boran Wang, Jingchi Jiang |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Disease Diagnosis based on Multiple Semantic Relationship Prompt SubgraphabstractExisting disease diagnosis models are often trained on clinical data such as electronic medical records (EMRs), which are lack of knowledge guidance. In this paper, we propose a new disease diagnosis model based on Multiple Semantic Relationship Prompt Subgraph (MSRPS), which can efficiently incorporate medical knowledge and utilize pretrained language models (PLMs) to achieve reasonable and accurate diagnosis results. Firstly, the MSRPS is extracted from the medical knowledge graph based on the patient’s personalized condition. Then graph encoder is used to generate a continuous prompt learning template based on graph neural network (GNN). Finally, a disease diagnosis model with prompt learning template under pre-trained and prompt paradigm is constructed to predict diagnosis results. Compared with the fine-tuning approach, this method needs fewer trainable parameters and less training data but achieve better performance. Experiments are conduct on two multi-label disease diagnosis datasets: ChineseEMR-50 and MIMIC-III-50. Results demonstrate that our model can be used in various pre-trained models and achieve the state-of-the-art results. Jing Xie 0012, Xin Li 0012, Yi Guan, Xitong Guo |
BIBM | 4 |
| 2023 | Efficient Evidence-Based Dialogue System for Medical DiagnosisabstractWith the rise of intelligent medical assistance, the Dialogue System for Medical Diagnosis(DSMD) guided by reinforcement learning(RL) has gained much attention. However, currently available medical dialogue datasets suffer from insufficient diagnostic evidence caused by sparse symptoms, making it difficult to reproduce the evidence-based process of doctors in differential diagnosis and disease confirmation. Moreover, purely data-driven RL often involves extensive and blind trial-and-error, leading to inquiries about irrelevant symptoms to the patient’s chief complaints in limited dialogue turns, further exacerbating the issue of inadequate diagnostic evidence. To enhance the quantity and effectiveness of potential symptom collection in DSMD, we first construct a more comprehensive medical dialogue dataset CMD based on electronic medical records. The diversity of diseases and symptoms mentioned in the dialogue context of CMD surpasses that of existing public datasets. Furthermore, to enhance the efficiency of diagnostic evidence collection in DSMD, inspired by the logic of symptom inquiries in doctor-patient interactions, we combine experiential diagnostic knowledge with a specialized medical knowledge graph to constrain the inquiry of symptoms via RL, eliminating the introduction of symptoms unrelated to the patient. Experimental results demonstrate that our model significantly outperforms competitive benchmark methods in terms of diagnostic accuracy and the efficiency of symptom inquiries. Our codes and the CMD dataset are available at https://github.com/YanPioneer/EBAD. Lian Yan, Yi Guan, Haotian Wang 0007, Jingchi Jiang |
BIBM | 2 |
| 2023 | Clarifying Confusion in Acne Severity Grading via Visualized Aggregation and Separation of Deep RepresentationabstractSeverity grading plays a vitally important role in the diagnosis and treatment of acne. However, due to its special characteristics such as similar samples, unclear class boundaries and imbalanced categories, the diagnosis is extremely easy to be confused. In this paper, we propose a novel, simple and intuitive loss function, namely Aggregation Separation Loss (ASLoss), as an adjunct for classification loss to clarify the common easily-confused cases. The ASLoss mines the commonalities of the same severity and the gaps among different severities in deep feature spaces. To demonstrate the generality of the proposed ASLoss, we also validate ASLoss on another common easily-confused task of expression recognition. The experimental results show that representations extracted by ASLoss are sufficiently clear and distinguishable, the performance of various popular methods can be improved significantly by ASLoss, the optimal network reaches the state-of-the-art and diagnostic level of dermatologists and the ASLoss can be generalized to improve the performance on other easily-confused tasks. Zeming Zhang, Jingchi Jiang, Ruyue Dong, Chaoran Kong, Yi Guan, Xiguang Liu, Haiyan You |
BIBM | 5 |
| 2023 | Interpretable Diagnosis of Face Acne via Complementation Learning of Evidence Localization and Severity Level GradingabstractAcne seriously affects people’s daily lives. Several studies of automated acne diagnosis either lack reasonable interpretation to support the diagnosis or ignore evidence in the diagnosis process. In this paper, we propose an interpretable diagnosis framework for face acne. This framework uses complementation learning of evidence localization and severity level grading to boost both streams by sharing the features that support each other. Evidence localization learns to identify lesion areas for supporting the diagnosis stream as well as providing interpretation. Severity level grading learns to recognize the diagnosis result and also provides reference and rectification for evidence localization. Experimental results show that complementation learning improves both evidence localization and severity level grading, the lesion areas from evidence localization can support the diagnosis and provide interpretations, and the diagnosis framework reaches the state-of-the-art level and the diagnostic performance of dermatologists. Zeming Zhang, Jingchi Jiang, Chaoran Kong, Yi Guan, Xiguang Liu, Haiyan You |
BIBM | 5 |
| 2023 | DECAF: An interpretable deep cascading framework for ICU mortality prediction
Jingchi Jiang, Xuehui Yu, Boran Wang, Linjiang Ma, Yi Guan |
Artif. Intell. Medicine | 5 |
| 2023 | DED: Diagnostic Evidence Distillation for acne severity grading on face images
Jingchi Jiang, Dongxin Chen, Yi Guan, Xiguang Liu, Haiyan You |
Expert Syst. Appl. | 5 |
| 2023 | LHP: Logical hypergraph link prediction
Yang Yang 0137, Yi Guan, Haotian Wang 0007, Chaoran Kong, Jingchi Jiang |
Expert Syst. Appl. | 3 |
| 2023 | ARLPE: A meta reinforcement learning framework for glucose regulation in type 1 diabetics
Xuehui Yu, Yi Guan, Lian Yan, Shulang Li, Xuelian Fu, Jingchi Jiang |
Expert Syst. Appl. | 2 |
| 2022 | Contextual Policy Transfer in Meta-Reinforcement Learning via Active Learning
Jingchi Jiang, Lian Yan, Xuehui Yu, Yi Guan |
WISA | 4 |
| 2022 | Unified Fine-Grained Biomedical Entity Recognition as a Combination of Boundary Detection and Sequence GenerationabstractBiomedical Named Entity Recognition (BioNER) is a critical component of biomedical information extraction. NER is more challenging in the biomedical domain because of fine-grained entity types and more common nested and discontinuous entity forms. However, none of the BioNER datasets contains a large amount of all three entity forms, including flat, nested, and discontinuous. Not to mention that there is a unified BioNER model for dealing with the above three entity forms simultaneously. Methods in the public domain only focus on identifying text spans and ignore distinguishing fine-grained entity types. To address these issues, we propose a unified framework based on our own BioNER dataset CCNER, which innovatively models the BioNER task as a combination of boundary recognition and sequence generation. CCNER is a comprehensive and fine-grained BioNER dataset, where the proportion of discontinuous, nested, and flat entities in the dataset is 8.9%, 52.6%, and 38.5%, respectively. Meanwhile, it includes five fine-grained entity types. Our proposed framework includes two modules which are boundary detection and entity generation. In the boundary detection module, we propose a sample-based span representation method to determine fine-grained entity boundaries better. Finally, we conduct experiments on four datasets and achieve competitive results1.1Code is available at https://github.com/lx-hit/BioNER. Yang Yang 0137, Mingchen Ye, Yi Guan, Xuehui Yu, Jingchi Jiang |
BIBM | 4 |
| 2022 | Acne Severity Grading on Face Images via Extraction and Guidance of Prior KnowledgeabstractAcne Vulgaris seriously affects people’s daily life. In this paper, we propose a face acne grading framework which is a new paradigm to solve the image classification problem where the number and type of small objects are the evidence. This framework includes two components: prior knowledge extraction and prior knowledge guided network. The prior knowledge extraction uses an excellent segmentation method to predict the lesion areas as prior knowledge. The prior knowledge guided network fuses the prior knowledge and its corresponding image to grade the severity. The experiment results demonstrate that our framework achieves the state-of-the-art and diagnosis level of dermatologists. Jingchi Jiang, Dongxin Chen, Yi Guan, Xiguang Liu, Haiyan You, Xue Cheng |
BIBM | 5 |
| 2022 | CGPG-GAN: An Acne Lesion Inpainting Model for Boosting Downstream DiagnosisabstractThe collection and publication of medical images on the face are quite difficult because of the invasion of privacy. Meanwhile, it takes a major expenditure of time and effort to manually label large-scale face images covered with s o many fine skin lesions. In this work, a multi-class object large-scale image inpainting model Class-Guided PG-GAN (CGPG-GAN) is proposed and its application in boosting downstream model performances is explored. This model is applied on face acne lesion inpainting where the image size is very large and missing areas are different types of lesions. The experiment results show that our method is superior to some existing methods and can improve the performance of downstream diagnosis remarkably. Jingchi Jiang, Dongxin Chen, Yi Guan, Xiguang Liu, Haiyan You, Xue Cheng |
BIBM | 5 |
| 2022 | Multi-scale Label Attention Network based on Abductive Causal Graph for Disease DiagnosisabstractThe auxiliary disease diagnosis based on electronic medical records is of great significance, providing doctors with diagnostic advice and avoiding misdiagnosis. Existing work on disease diagnosis mainly utilizes deep learning models to extract sequence information in electronic medical records, ignoring the interpretability of results and the structural knowledge, especially causal knowledge. In our work, we propose a multiscale label attention network based on abductive causal graph (MSLAN-ACG) to improve model accuracy and interpretability of results. First, we construct multiple encoders in the multiscale label attention network, which can extract n-gram segment information of different lengths for each disease. Meanwhile, to enhance the interpretability of results, we visualize the weight score of different segments for disease results. Second, we propose a disease representation method by defining an abductive causal graph and then using graph convolutional network for knowledge fusion on this graph. The information propagation based on abductive causal graph is consistent with the actual abductive reasoning process from symptoms to diseases, making the model more reasonable. The effectiveness of our model is demonstrated by achieving state-of-the-art results on MIMICIII-50 and ChineseEMR datasets. Haotian Wang 0007, Yi Guan, Linjiang Ma, Xin Li 0012, Jing Xie 0012, Jingchi Jiang |
BIBM | 2 |
| 2022 | Causal Coupled Mechanisms: A Control Method with Cooperation and Competition for Complex SystemabstractComplex systems are ubiquitous in the real world and tend to have complicated and poorly understood dynamics. For their control issues, the challenge is to guarantee accuracy, robustness, and generalization in such bloated and troubled environments. Fortunately, a complex system can be divided into multiple modular structures that human cognition appears to exploit. Inspired by this cognition, a novel control method, Causal Coupled Mechanisms (CCMs), is proposed that explores the cooperation in division and competition in combination. Our method employs the theory of hierarchical reinforcement learning (HRL), in which 1) the high-level policy with competitive awareness divides the whole complex system into multiple functional mechanisms, and 2) the low-level policy finishes the control task of each mechanism. Specifically for cooperation, a cascade control module helps the series operation of CCMs, and a forward coupled reasoning module is used to recover the coupling information lost in the division process. On both synthetic systems and a real-world biological regulatory system, the CCM method achieves robust and state-of-the-art control results even with unpredictable random noise. Moreover, generalization results show that reusing prepared specialized CCMs helps to perform well in environments with different confounders and dynamics. Xuehui Yu, Yi Guan, Xinmiao Yu, Jingchi Jiang |
BIBM | 2 |
| 2022 | Gated Tree-based Graph Attention Network (GTGAT) for medical knowledge graph reasoning
Jingchi Jiang, Tao Wang 0073, Boran Wang, Linjiang Ma, Yi Guan |
Artif. Intell. Medicine | 5 |
| 2022 | ggmsa: a visual exploration tool for multiple sequence alignment and associated dataabstractThe identification of the conserved and variable regions in the multiple sequence alignment (MSA) is critical to accelerating the process of understanding the function of genes. MSA visualizations allow us to transform sequence features into understandable visual representations. As the sequence-structure-function relationship gains increasing attention in molecular biology studies, the simple display of nucleotide or protein sequence alignment is not satisfied. A more scalable visualization is required to broaden the scope of sequence investigation. Here we present ggmsa, an R package for mining comprehensive sequence features and integrating the associated data of MSA by a variety of display methods. To uncover sequence conservation patterns, variations and recombination at the site level, sequence bundles, sequence logos, stacked sequence alignment and comparative plots are implemented. ggmsa supports integrating the correlation of MSA sequences and their phenotypes, as well as other traits such as ancestral sequences, molecular structures, molecular functions and expression levels. We also design a new visualization method for genome alignments in multiple alignment format to explore the pattern of within and between species variation. Combining these visual representations with prime knowledge, ggmsa assists researchers in discovering MSA and making decisions. The ggmsa package is open-source software released under the Artistic-2.0 license, and it is freely available on Bioconductor (https://bioconductor.org/packages/ggmsa) and Github (https://github.com/YuLab-SMU/ggmsa). Tingze Feng, Shuangbin Xu, Fangluan Gao, Tommy T. Lam, Tianzhi Wu, Huina Huang, Li Zhan, Yi Guan, Zehan Dai, Guangchuang Yu |
Briefings Bioinform. | 11 |
| 2021 | AdvPicker: Effectively Leveraging Unlabeled Data via Adversarial Discriminator for Cross-Lingual NERabstractWeile Chen, Huiqiang Jiang, Qianhui Wu, Börje Karlsson, Yi Guan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Weile Chen, Huiqiang Jiang, Qianhui Wu, Börje Karlsson 0001, Yi Guan |
ACL/IJCNLP (1) | 5 |
| 2021 | An Acne Grading Framework on Face Images via Skin Attention and SFNetabstractSeverity level grading is a vitally important step to make correct diagnoses and personalized treatment schemes for acne, which is mainly carried out in two ways: criterion-based lesion counting and experience-based global estimation. In this paper, the global estimation of acne severity grading is studied by Convolutional Neural Networks (CNNs) and a unified acne grading framework that can diagnose referring to different grading criteria is proposed. Firstly, an adaptive image preprocessing method that can efficiently reduce the background noise and emphasize the skin information is proposed. Next, an innovative CNN structure SFNet, which fuses local skin features with global features to effectively enhance the perception of color gaps between skin and lesion, is presented. The proposed framework is verified on two datasets with different acne grading criteria. Experimental results show that the accuracy of the proposed framework reaches 84.52% exceeding the state-of-the-art method by 1.7% and reaches the diagnostic level of a professional dermatologist. Yi Guan, Haiyan You, Xue Cheng, Jingchi Jiang |
BIBM | 2 |
| 2021 | Alignment free sequence comparison methods and reservoir host predictionabstractMOTIVATION: The emergence and subsequent pandemic of the SARS-CoV-2 virus raised urgent questions about its origin and, particularly, its reservoir host. These types of questions are long-standing problems in the management of emerging infectious diseases and are linked to virus discovery programs and the prediction of viruses that are likely to become zoonotic. Conventional means to identify reservoir hosts have relied on surveillance, experimental studies and phylogenetics. More recently, machine learning approaches have been applied to generate tools to swiftly predict reservoir hosts from sequence data. RESULTS: Here, we extend a recent work that combined sequence alignment and a mixture of alignment-free approaches using a gradient boosting machines machine learning model, which integrates genomic traits and phylogenetic neighbourhood signatures to predict reservoir hosts. We add a more uniform approach by applying Machine Learning with Digital Signal Processing-based structural patterns. The extended model was applied to an existing virus/reservoir host dataset and to the SARS-CoV-2 and related viruses and generated an improvement in prediction accuracy. AVAILABILITY AND IMPLEMENTATION: The source code used in this work is freely available at https://github.com/bill1167/hostgbms. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bill Lee, David K. Smith 0001, Yi Guan |
Bioinform. | 3 |
| 2021 | Clinical decision-making framework against over-testing based on modeling implicit evaluation criteria
Yang Yang 0041, Hongxing Huo, Jingchi Jiang, Xuemei Sun, Yi Guan, Xitong Guo, Shengping Liu |
J. Biomed. Informatics | 5 |
| 2020 | Medical knowledge embedding based on recursive neural network for multi-disease diagnosis
Jingchi Jiang, Huanzheng Wang, Jing Xie 0012, Xitong Guo, Yi Guan, Qiubin Yu |
Artif. Intell. Medicine | 5 |
| 2020 | Learning an expandable EMR-based medical knowledge network to enhance clinical diagnosis
Jing Xie 0012, Jingchi Jiang, Yehan Wang, Yi Guan, Xitong Guo |
Artif. Intell. Medicine | 4 |
| 2020 | Sampled-Data Synchronization of Network Systems in Industrial ManufactureabstractThis paper proposes a novel control strategy for the synchronization of network systems. The designed distributed controllers adopt the communication channels to exchange information. The designed controller for each heterogeneous node includes two parts: 1) the reference generator (RG) to copy the dynamics of the leader and 2) adaptive regulator (AR) to achieve synchronization purpose. Under the action of sampled-data control law, outputs of all RGs converge to the output of the leader. The closed-loop system of the leader and all RGs are be equivalently written as the interaction of an operator and a linear time-invariant system. The small gain theorem is utilized to calculate the upper bound of the sampling intervals. Furthermore, the integral quadratic constraints can provide the passivity-type property of the operator and give the less conservative results. Meanwhile, the AR can ensure that nonidentical node tracks its exosystem. Thus, all nonidentical nodes and the leader achieve output synchronization. The proposed control strategy is similar to the separation principle, which includes two steps. Finally, a numerical example is given to demonstrated the effectiveness of the proposed control strategy. Yuanqing Wu 0003, Yanzhou Li, Shenghuang He, Yi Guan |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2019 | Classifying medical relations in clinical text via convolutional neural networks
Bin He 0005, Yi Guan |
Artif. Intell. Medicine | 2 |
| 2018 | Convolutional Gated Recurrent Units for Medical Relation Classification
Bin He 0005, Yi Guan |
BIBM | 2 |
| 2018 | EMR-based medical knowledge representation and inference via Markov random fields and distributed representation learning
Chao Zhao 0002, Jingchi Jiang, Yi Guan, Xitong Guo, Bin He 0005 |
Artif. Intell. Medicine | 3 |
| 2018 | A multitask bi-directional RNN model for named entity recognition on Chinese electronic medical recordsabstractBACKGROUND: Electronic Medical Record (EMR) comprises patients' medical information gathered by medical stuff for providing better health care. Named Entity Recognition (NER) is a sub-field of information extraction aimed at identifying specific entity terms such as disease, test, symptom, genes etc. NER can be a relief for healthcare providers and medical specialists to extract useful information automatically and avoid unnecessary and unrelated information in EMR. However, limited resources of available EMR pose a great challenge for mining entity terms. Therefore, a multitask bi-directional RNN model is proposed here as a potential solution of data augmentation to enhance NER performance with limited data. METHODS: A multitask bi-directional RNN model is proposed for extracting entity terms from Chinese EMR. The proposed model can be divided into a shared layer and a task specific layer. Firstly, vector representation of each word is obtained as a concatenation of word embedding and character embedding. Then Bi-directional RNN is used to extract context information from sentence. After that, all these layers are shared by two different task layers, namely the parts-of-speech tagging task layer and the named entity recognition task layer. These two tasks layers are trained alternatively so that the knowledge learned from named entity recognition task can be enhanced by the knowledge gained from parts-of-speech tagging task. RESULTS: The performance of our proposed model has been evaluated in terms of micro average F-score, macro average F-score and accuracy. It is observed that the proposed model outperforms the baseline model in all cases. For instance, experimental results conducted on the discharge summaries show that the micro average F-score and the macro average F-score are improved by 2.41% point and 4.16% point, respectively, and the overall accuracy is improved by 5.66% point. CONCLUSIONS: In this paper, a novel multitask bi-directional RNN model is proposed for improving the performance of named entity recognition in EMR. Evaluation results using real datasets demonstrate the effectiveness of the proposed model. Shanta Chowdhury, Xishuang Dong, Lijun Qian, Xiangfang Li, Yi Guan, Jinfeng Yang, Qiubin Yu |
BMC Bioinform. | 5 |
| 2017 | Transfer bi-directional LSTM RNN for named entity recognition in Chinese electronic medical recordsabstractIn this paper, a transfer bi-directional recurrent neural networks (RNN) is proposed for named entity recognition (NER) in Chinese electronic medical records (EMRs) that aims to extract medical knowledge such as phrases recording diseases and treatments automatically. We propose a two-step procedure where the first step is to train a shallow bi-directional RNN in the general domain, and the second step is to transfer knowledge from the general domain to train a deeper bi-directional RNN for recognizing medical concepts from Chinese EMRs. Specifically, this is achieved by initializing the shallow parts of the deeper network in the second step with parameter weights from the bi-directional RNN trained in the first step. Then the deeper networks are re-trained on the Chinese EMRs. Experimental results show that NER performances are improved by the transferred knowledge significantly. Xishuang Dong, Shanta Chowdhury, Lijun Qian, Yi Guan, Jinfeng Yang, Qiubin Yu |
Healthcom | 4 |
| 2017 | Building a comprehensive syntactic and semantic corpus of Chinese clinical texts
Bin He 0005, Bin Dong 0003, Yi Guan, Jinfeng Yang, Qiubin Yu, Jianyi Cheng, Chunyan Qu |
J. Biomed. Informatics | 3 |
| 2017 | Learning and inference in knowledge-based probabilistic model for medical diagnosis
Jingchi Jiang, Xueli Li, Chao Zhao 0002, Yi Guan, Qiubin Yu |
Knowl. Based Syst. | 4 |
| 2016 | ActivityHijacker: Hijacking the Android Activity Component for Sensitive DataabstractAs many malicious apps have been distributed in Android markets, it is urgent to fix the vulnerabilities exploited by these apps and develop effective mitigation methods. In this paper, we identify a common vulnerability in many Android apps. This vulnerability is rooted in an unprotected Android component, called "Activity". Utilizing this vulnerability, a malicious app can stealthily and cannily collect sensitive user data. To better understand this issue, we have built "ActivityHijacker", an app that can detect the right moment to hijack the Activity component and intercept a user's password while it is being inputted in real time. To assess the prevalence of this vulnerability, we further analyzed 8 Android OSs (version 2.x to 5.x) and 22 banking apps collected in January 2016 from various Android markets. Our results show that all these Android OSs and apps are susceptible to the vulnerability; 14 out of 22 banking apps can be hijacked without any prompts. We further present a mitigation mechanism that restricts the Activity component to authorized apps. Chenglong Li 0001, Yi Guan, Yibo Xue, Yingfei Dong |
ICCCN | 3 |
| 2016 | DroidChain: A novel Android malware detection method based on behavior chains
Chenglong Li 0001, Zhenlong Yuan, Yi Guan, Yibo Xue |
Pervasive Mob. Comput. | 4 |
| 2015 | CRFs based de-identification of medical recordsabstractDe-identification is a shared task of the 2014 i2b2/UTHealth challenge. The purpose of this task is to remove protected health information (PHI) from medical records. In this paper, we propose a novel de-identifier, WI-deId, based on conditional random fields (CRFs). A preprocessing module, which tokenizes the medical records using regular expressions and an off-the-shelf tokenizer, is introduced, and three groups of features are extracted to train the de-identifier model. The experiment shows that our system is effective in the de-identification of medical records, achieving a micro-F1 of 0.9232 at the i2b2 strict entity evaluation level. Bin He 0005, Yi Guan, Jianyi Cheng, Keting Cen, Wenlan Hua |
J. Biomed. Informatics | 2 |
| 2014 | Representing Words as LymphocytesabstractSimilarity between words is becoming a generic problem for many applications of computational linguistics, and computing word similarities is determined by word representations. Inspired by the analogies between words and lymphocytes, a lymphocyte-style word representation is proposed. The word representation is built on the basis of dependency syntax of sentences and represent word context as head properties and dependent properties of the word. Lymphocyte-style word representations are evaluated by computing the similarities between words, and experiments are conducted on the Penn Chinese Treebank 5.1. Experimental results indicate that the proposed word representations are effective. Jinfeng Yang, Yi Guan, Xishuang Dong, Bin He 0005 |
AAAI | 2 |
| 2014 | Developing a linguistically annotated corpus of Chinese electronic medical recordabstractElectronic Medical Record (EMR) is the material base of smart healthcare, its automatic analysis is dependent on nature language processing (NLP) technologies. Syntactic analysis, as the basic technology of NLP, can be used to convert the free text of EMR to structured text. However, research on syntactic analysis, even Chinese word segmentation and part-of-speech (POS) tagging on Chinese electronic Medical record (CEMR), is currently at a blank stage because of the lack of annotated corpus on CEMR. To resolve this problem, we propose the annotated scheme from Chinese word segmentation to syntactic analysis, and built the first syntactically annotated corpus of CEMR. Through analyzing the annotated CEMR, we find it has stronger grammatical regularity and particular statistical distribution. These finds are taken advantage to improve the Stanford parser and develop a state-of-the-art Chinese word segmentation and POS tagging system for CEMR. The evaluation results show a substantial benefit to statistical machine learning models from the annotated CEMR. Fangfang Zhao, Yi Guan |
BIBM | 3 |
| 2014 | Words Are Analogous To Lymphocytes: A Multi-Word-Agent Autonomous Learning Model
Jinfeng Yang, Xishuang Dong, Yi Guan |
ICSEng | 3 |
| 2014 | Transfer learning based clinical concept extraction on data from multiple sources
Xinbo Lv, Yi Guan, Benyang Deng |
J. Biomed. Informatics | 2 |
| 2013 | Reserved Self-training: A Semi-supervised Sentiment Classification Method for Chinese Microblogs
Xishuang Dong, Yi Guan, Jinfeng Yang |
IJCNLP | 3 |
| 2013 | Multi-robot task allocation using CNP combines with neural network
Quande Yuan, Yi Guan, Bingrong Hong, Xiangping Meng |
Neural Comput. Appl. | 2 |
| 2012 | Set-Similarity Joins Based Semi-supervised Sentiment Analysis
Xishuang Dong, Qibo Zou, Yi Guan |
ICONIP (1) | 3 |
| 2011 | Automatically Generating Questions from Queries for Community-based Question Answering
Haifeng Wang 0001, Ting Liu 0001, Yi Guan |
IJCNLP | 5 |
| 2008 | A New Measurement of Systematic SimilarityabstractThe relationship of similarity may be the most universal relationship that exists between every two objects in either the material world or the mental world. Although similarity modeling has been the focus of cognitive science for decades, many theoretical and realistic issues are still under controversy. In this paper, a new theoretical framework that conforms to the nature of similarity and incorporates the current similarity models into a universal model is presented. The new model, i.e., the systematic similarity model, which is inspired by the contrast model of similarity and structure mapping theory in cognitive psychology, is the universal similarity measurement that has many potential applications in text, image, or video retrieval. The text relevance ranking experiments undertaken in this research tentatively show the validity of the new model. Yi Guan, Xiaolong Wang 0001, Qiang Wang 0001 |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2007 | A Probabilistic Approach to Syntax-based Reordering for Statistical Machine Translation
Chi-Ho Li, Dongdong Zhang 0001, Mu Li 0001, Ming Zhou 0001, Yi Guan |
ACL | 6 |
| 2007 | Using Maximum Entropy Model to Extract Protein-Protein Interaction Information from Biomedical Literature
Chengjie Sun, Lei Lin 0001, Xiaolong Wang 0001, Yi Guan |
ICIC (1) | 4 |
| 2007 | Exploiting residue-level and profile-level interface propensities for usage in binding sites prediction of proteinsabstractBACKGROUND: Recognition of binding sites in proteins is a direct computational approach to the characterization of proteins in terms of biological and biochemical function. Residue preferences have been widely used in many studies but the results are often not satisfactory. Although different amino acid compositions among the interaction sites of different complexes have been observed, such differences have not been integrated into the prediction process. Furthermore, the evolution information has not been exploited to achieve a more powerful propensity. RESULT: In this study, the residue interface propensities of four kinds of complexes (homo-permanent complexes, homo-transient complexes, hetero-permanent complexes and hetero-transient complexes) are investigated. These propensities, combined with sequence profiles and accessible surface areas, are inputted to the support vector machine for the prediction of protein binding sites. Such propensities are further improved by taking evolutional information into consideration, which results in a class of novel propensities at the profile level, i.e. the binary profiles interface propensities. Experiment is performed on the 1139 non-redundant protein chains. Although different residue interface propensities among different complexes are observed, the improvement of the classifier with residue interface propensities can be negligible in comparison with that without propensities. The binary profile interface propensities can significantly improve the performance of binding sites prediction by about ten percent in term of both precision and recall. CONCLUSION: Although there are minor differences among the four kinds of complexes, the residue interface propensities cannot provide efficient discrimination for the complicated interfaces of proteins. The binary profile interface propensities can significantly improve the performance of binding sites prediction of protein, which indicates that the propensities at the profile level are more accurate than those at the residue level. Qiwen Dong, Xiaolong Wang 0001, Lei Lin 0001, Yi Guan |
BMC Bioinform. | 4 |
| 2007 | Recent advances on NLP research in Harbin Institute of Technology
Tiejun Zhao, Yi Guan, Ting Liu 0001, Qiang Wang 0001 |
Frontiers Comput. Sci. China | 2 |
| 2006 | Conditional Random Fields Based Label Sequence and Information Feedback
Wei Jiang 0036, Yi Guan, Xiaolong Wang 0001 |
ICIC (2) | 2 |
| 2004 | A Study of Semi-discrete Matrix Decomposition for LSI in Automated Text Categorization
Qiang Wang 0001, Xiaolong Wang 0001, Yi Guan |
IJCNLP | 3 |