Lili Jiang 0002

dblp:49/3022-2 · DBLP profile ↗
← Back
26ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-7788-3986ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 13 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 SynNER: Synergizing Large and Small Language Models for Few-Shot Nested NER
abstract
Large language models (LLMs) encounter challenges when addressing few-shot nested named entity recognition (NER) tasks. Traditional LLM-based approaches typically either prompt the model to generate entity words or types in sentences based on entity categories or word spans, or directly extract all entities of specific types present in the sentences. These methods often suffer from issues such as low query efficiency or suboptimal accuracy. This paper introduces an innovative framework, SynNER, which synergizes small and large language models to address these limitations. Initially, a small language model identifies low-confidence word spans, which are then refined and refined by a large language model. To simultaneously ensure recognition accuracy and improve the query efficiency of the LLM, we propose a Batch-Prompt strategy and an Entity Indexing method. These techniques enable the LLM to process multiple test instances simultaneously while maintaining precise correction results. Experimental results demonstrate that our method achieves significant performance gains on benchmark datasets, offering a cost-effective solution for few-shot nested NER tasks.
Hong Ming, Jiaoyun Yang, Lili Jiang 0002, Ning An 0001
IJCNN4
2025 SynNER: Synergizing Large and Small Language Models for Few-Shot Nested NER
abstract
Large language models (LLMs) encounter challenges when addressing few-shot nested named entity recognition (NER) tasks. Traditional LLM-based approaches typically either prompt the model to generate entity words or types in sentences based on entity categories or word spans, or directly extract all entities of specific types present in the sentences. These methods often suffer from issues such as low query efficiency or suboptimal accuracy. This paper introduces an innovative framework, SynNER, which synergizes small and large language models to address these limitations. Initially, a small language model identifies low-confidence word spans, which are then refined and refined by a large language model. To simultaneously ensure recognition accuracy and improve the query efficiency of the LLM, we propose a Batch-Prompt strategy and an Entity Indexing method. These techniques enable the LLM to process multiple test instances simultaneously while maintaining precise correction results. Experimental results demonstrate that our method achieves significant performance gains on benchmark datasets, offering a cost-effective solution for few-shot nested NER tasks.
Hong Ming, Jiaoyun Yang, Lili Jiang 0002, Ning An 0001
IJCNN4
2025 Incomplete Multi-View Drug Recommendation via Multi-Level Representation Learning and Curriculum Learning
abstract
The drug recommendation task aims to provide effective and safe prescription decision support for clinical treatment based on patients' past Electronic Health Records (EHR). However, the prevalent phenomenon of missing views in multi-source heterogeneous EHR data may cause performance degradation. This is due to the lack of sufficient information and increased learning difficulties, which limit the practical effectiveness of drug recommendation models in medical applications. In this paper, we emphasize the problems of incompleteness in practical drug recommendation and propose the Incomplete Multi-View Drug Recommendation model via Multi-Level Representation Learning and Curriculum Learning named IMDR. In particular, IMDR employs a Multi-Level Representation Learning architecture equipped with a Medical Code-Level Drug Knowledge Infusion Module and a Visit-Level Cross-View Information Module for patient representation learning to overcome the information loss caused by incomplete data. And then, a Gaussian-guided curriculum learning strategy is proposed to assist the learning process of IMDR with a novel difficulty measure to achieve effective progressive learning under missing medical views. Systematic evaluation on two large-scale real-world medical datasets, MIMIC-III and MIMIC-IV, demonstrates that IMDR reduces the Drug-Drug Interaction (DDI) rate by 2.97% compared to existing state-of-the-art drug recommendation baselines, while achieving significant improvements of 3.29% and 1.97% in Jaccard similarity scores and F1 score, respectively. Furthermore, compared to advanced incomplete multi-view learning (IML) models, IMDR's advantages in Jaccard similarity scores and F1 score further expand to 4.03% and 2.41%.
Ning Liu 0014, Yunsen Tang, Haitao Yuan 0002, Hongtao Lv, Lili Jiang 0002, Zhen Li 0049, Wei Zhang 0056, Jianyong Wang 0001
KDD (2)5
2025 Harnessing high-quality pseudo-labels for robust few-shot nested named entity recognition
Hong Ming, Jiaoyun Yang, Lili Jiang 0002, Ning An 0001
Eng. Appl. Artif. Intell.4
2025 Mitigating prototype shift: Few-shot nested named entity recognition with prototype-attention contrastive learning
Hong Ming, Jiaoyun Yang, Lili Jiang 0002, Ning An 0001
Expert Syst. Appl.4
2024 LPNER: Label Prompt for Few-shot Nested Named Entity Recognition
Jiaoyun Yang, Zhihan Zhu, Hong Ming, Lili Jiang 0002, Ning An 0001
ACML4
2024 Few-shot nested named entity recognition
Hong Ming, Jiaoyun Yang, Fang Gui, Lili Jiang 0002, Ning An 0001
Knowl. Based Syst.4
2021 ICDAR 2021 Competition on Multimodal Emotion Recognition on Comics Scenes
Vincent Nguyen 0001, Xuan-Son Vu, Christophe Rigaud, Lili Jiang 0002, Jean-Christophe Burie
ICDAR (4)4
2021 Context-based image explanations for deep neural networks
abstract
With the increased use of machine learning in decision-making scenarios, there has been a growing interest in explaining and understanding the outcomes of machine learning models. Despite this growing interest, existing works on interpretability and explanations have been mostly intended for expert users. Explanations for general users have been neglected in many usable and practical applications (e.g., image tagging, caption generation). It is important for non-technical users to understand features and how they affect an instance-specific prediction to satisfy the need for justification. In this paper, we propose a model-agnostic method for generating context-based explanations aiming for general users. We implement partial masking on segmented components to identify the contextual importance of each segment in scene classification tasks. We then generate explanations based on feature importance. We present visual and text-based explanations: (i) saliency map presents the pertinent components with a descriptive textual justification, (ii) visual map with a color bar graph showing the relative importance of each feature for a prediction. Evaluating the explanations using a user study (N = 50), we observed that our proposed explanation method visually outperformed existing gradient and occlusion based methods. Hence, our proposed explanation method could be deployed to explain models’ decisions to non-expert users in real-world applications.
Sule Anjomshoae, Daniel Omeiza, Lili Jiang 0002
Image Vis. Comput.3
2020 Multimodal Review Generation with Privacy and Fairness Awareness
abstract
Users express their opinions towards entities (e.g., restaurants) via online reviews which can be in diverse forms such as text, ratings, and images.Modeling reviews are advantageous for user behavior understanding which, in turn, supports various user-oriented tasks such as recommendation, sentiment analysis, and review generation.In this paper, we propose MG-PriFair, a multimodal neural-based framework, which generates personalized reviews with privacy and fairness awareness.Motivated by the fact that reviews might contain personal information and sentiment bias, we propose a novel differentially private (dp)-embedding model for training privacy guaranteed embeddings and an evaluation approach for sentiment fairness in the food-review domain.Experiments on our novel review dataset show that MG-PriFair is capable of generating plausibly long reviews while controlling the amount of exploited user data and using the least sentimentbiased word embeddings.To the best of our knowledge, we are the first to bring user privacy and sentiment fairness into the review generation task.The dataset and source codes are available at https
Xuan-Son Vu, Thanh-Son Nguyen 0001, Duc-Trong Le, Lili Jiang 0002
COLING4
2020 A Web-Based Platform for Mining and Ranking Association Rules
Addi Ait-Mlouk, Lili Jiang 0002
ECIR (2)2
2020 Privacy-Preserving Visual Content Tagging using Graph Transformer Networks
abstract
With the rapid growth of Internet media, content tagging has become an important topic with many multimedia understanding applications, including efficient organisation and search. Nevertheless, existing visual tagging approaches are susceptible to inherent privacy risks in which private information may be exposed unintentionally. The use of anonymisation and privacy-protection methods is desirable, but with the expense of task performance. Therefore, this paper proposes an end-to-end framework (SGTN) using Graph Transformer and Convolutional Networks to significantly improve classification and privacy preservation of visual data. Especially, we employ several mechanisms such as differential privacy based graph construction and noise-induced graph transformation to protect the privacy of knowledge graphs. Our approach unveils new state-of-the-art on MS-COCO dataset in various semi-supervised settings. In addition, we showcase a real experiment in the education domain to address the automation of sensitive document tagging. Experimental results show that our approach achieves an excellent balance of model accuracy and privacy preservation on both public and private datasets.
Xuan-Son Vu, Duc-Trong Le, Christoffer Edlund, Lili Jiang 0002, Hoang D. Nguyen
ACM Multimedia4
2019 dpUGC: Learn Differentially Private Representation for User Generated Contents (Best Paper Award, Third Place, Shared)
Xuan-Son Vu, Son N. Tran, Lili Jiang 0002
CICLing (1)3
2019 Graph-based Interactive Data Federation System for Heterogeneous Data Retrieval and Analytics
abstract
Given the increasing number of heterogeneous data stored in relational databases, file systems or cloud environment, it needs to be easily accessed and semantically connected for further data analytic. The potential of data federation is largely untapped, this paper presents an interactive data federation system (https://vimeo.com/319473546) by applying large-scale techniques including heterogeneous data federation, natural language processing, association rules and semantic web to perform data retrieval and analytics on social network data. The system first creates a Virtual Database (VDB) to virtually integrate data from multiple data sources. Next, a RDF generator is built to unify data, together with SPARQL queries, to support semantic data search over the processed text data by natural language processing (NLP). Association rule analysis is used to discover the patterns and recognize the most important co-occurrences of variables from multiple data sources. The system demonstrates how it facilitates interactive data analytic towards different application scenarios (e.g., sentiment analysis, privacy-concern analysis, community detection).
Xuan-Son Vu, Addi Ait-Mlouk, Erik Elmroth, Lili Jiang 0002
WWW4
2019 Microarray Missing Value Imputation: A Regularized Local Learning Method
abstract
Microarray experiments on gene expression inevitably generate missing values, which impedes further downstream biological analysis. Therefore, it is key to estimate the missing values accurately. Most of the existing imputation methods tend to suffer from the over-fitting problem. In this study, we propose two regularized local learning methods for microarray missing value imputation. Motivated by the grouping effect of $L_{2}$L2 regularization, after selecting the target gene, we train an $L_{2}$L2 Regularized Local Least Squares imputation model (RLLSimpute_L2) on the target gene and its neighbors to estimate the missing values of the target gene. Furthermore, RLLSimpute_L2 imputes the missing values in an ascending order based on the associated missing rate with each target gene. This contributes to fully utilizing the previously estimated values. Besides $L_{2}$L2, we further explore $L_{1}$L1 regularization and propose an $L_{1}$L1 Regularized Local Least Squares imputation model (RLLSimpute_L1). To evaluate their effectiveness, we conducted extensive experimental studies on six benchmark datasets covering both time series and non-time series cases. Nine state-of-the-art imputation methods are compared with RLLSimpute_L2 and RLLSimpute_L1 in terms of three performance metrics. The comparative experimental results indicate that RLLSimpute_L2 outperforms its competitors by achieving smaller imputation errors and better structure preservation of differentially expressed genes.
Aiguo Wang 0002, Ye Chen 0011, Ning An 0001, Jing Yang 0008, Lian Li 0001, Lili Jiang 0002
IEEE ACM Trans. Comput. Biol. Bioinform.6
2018 Self-adaptive Privacy Concern Detection for User-Generated Content
Xuan-Son Vu, Lili Jiang 0002
CICLing (1)2
2018 Lexical-semantic resources: yet powerful resources for automatic personality classification
abstract
In this paper, we aim to reveal the impact of lexical-semantic resources, used in particular for word sense disambiguation and sense-level semantic categorization, on automatic personality classification task.While stylistic features (e.g., part-of-speech counts) have been shown their power in this task, the impact of semantics beyond targeted word lists is relatively unexplored.We propose and extract three types of lexical-semantic features, which capture high-level concepts and emotions, overcoming the lexical gap of word n-grams.Our experimental results are comparable to state-of-the-art methods, while no personality-specific resources are required.
Xuan-Son Vu, Lucie Flek, Lili Jiang 0002, Iryna Gurevych
GWC3
2017 Personality-based Knowledge Extraction for Privacy-preserving Data Analysis
abstract
In this paper, we present a differential privacy preserving approach, which extracts personality-based knowledge to serve privacy guarantee data analysis on personal sensitive data. Based on the approach, we further implement an end-to-end privacy guarantee system, KaPPA, to provide researchers iterative data analysis on sensitive data. The key challenge for differential privacy is determining a reasonable amount of privacy budget to balance privacy preserving and data utility. Most of the previous work applies unified privacy budget to all individual data, which leads to insufficient privacy protection for some individuals while over-protecting others. In KaPPA, the proposed personality-based privacy preserving approach automatically calculates privacy budget for each individual. Our experimental evaluations show a significant trade-off of sufficient privacy protection and data utility.
Xuan-Son Vu, Lili Jiang 0002, Anders Brändström, Erik Elmroth
K-CAP2
2014 Toward detection of aliases without string similarity
Ning An 0001, Lili Jiang 0002, Jianyong Wang 0001, Ping Luo 0001, Min Wang 0001, Bing Nan Li
Inf. Sci.2
2014 Senti-LSSVM: Sentiment-Oriented Multi-Relation Extraction with Latent Structural SVM
abstract
Extracting instances of sentiment-oriented relations from user-generated web documents is important for online marketing analysis. Unlike previous work, we formulate this extraction task as a structured prediction problem and design the corresponding inference as an integer linear program. Our latent structural SVM based model can learn from training corpora that do not contain explicit annotations of sentiment-bearing expressions, and it can simultaneously recognize instances of both binary (polarity) and ternary (comparative) relations with regard to entity mentions of interest. The empirical evaluation shows that our approach significantly outperforms state-of-the-art systems across domains (cameras and movies) and across genres (reviews and forum posts). The gold standard corpus that we built will also be a valuable resource for the community.
Lizhen Qu, Yi Zhang 0003, Rui Wang 0005, Lili Jiang 0002, Rainer Gemulla, Gerhard Weikum
Trans. Assoc. Comput. Linguistics4
2013 GRIAS: An Entity-Relation Graph Based Framework for Discovering Entity Aliases
abstract
Recognizing the various aliases of an entity is a critical task for many applications, including Web search, recommendation system, and e-discovery. The goal of this paper is to accurately identify entity aliases, especially the long tail ones in the unstructured data. Our solution GRIAS (abbr. for a Graph-based framework for discovering entity Aliases) is motivated by the entity relationships collected from both the structured and unstructured data. These relationships help to build an entity-relation graph, and the graph-based similarity is calculated between an entity and its alias candidates which are first chosen by our proposed candidate selection method. Extensive experimental results on two real-world datasets demonstrate both the effectiveness and efficiency of the proposed framework.
Lili Jiang 0002, Ping Luo 0001, Jianyong Wang 0001, Yuhong Xiong, Bingduan Lin, Min Wang 0001, Ning An 0001
ICDM1
2013 YaLi: a crowdsourcing plug-in for NERD
abstract
We demonstrate the YaLi browser plug-in which discovers named entities in Web pages and provides background knowledge about them. The plug-in is implemented with two purposes. From a user perspective, it enriches the browsing experience with entities, helping users with their information needs. From the research perspective, we aim to improve the methods that are used for named entity recognition and disambiguation (NERD) by leveraging the plug-in as an implicit crowdsourcing platform. YaLi tracks the system's errors and the users' corrections, and also gathers implicit training data for improving NERD accuracy.
Yafang Wang, Lili Jiang 0002, Johannes Hoffart, Gerhard Weikum
SIGIR2
2012 Towards alias detection without string similarity: an active learning based approach
abstract
Entity aliases commonly exist and accurately detecting these aliases plays a vital role in various applications. In this paper, we use an active-learning-based method to detect aliases without string similarity. To minimize the cost on pairwise comparison, a subset-based method restricts the alias selection within a small-scale entity set. Within each generated entity set, an active learning based logistic regression classifier is employed to predict whether a candidate is the alias of a given entity. The experimental results on three datasets clearly demonstrate that our proposed approach can effectively detect this kind of entity aliases.
Lili Jiang 0002, Jianyong Wang 0001, Ping Luo 0001, Ning An 0001, Min Wang 0001
SIGIR1
2010 GRAPE: a system for disambiguating and tagging people names in web search
abstract
Name ambiguity is a big challenge in people information retrieval and has received considerable attention, especially with the increasing volume of Web data in recent years. In this demo, we present a system, GRAPE, which is capable of finding people related information over the Web. The salient features of our system are people name disambiguation and people tag presentation, which effectively distinguish different people entities sharing the same name and uniquely represent each namesake with a cluster of tags, such as occupation, birthdate, and organization.
Lili Jiang 0002, Wei Shen 0004, Jianyong Wang 0001, Ning An 0001
WWW1
2009 GRAPE: A Graph-Based Framework for Disambiguating People Appearances in Web Search
abstract
Finding information about people using search engines is one of the most common activities on the Web. However, search engines usually return a long list of Web pages, which may be relevant to many namesakes, especially given the explosive growth of Web data. To address the challenge caused by name ambiguity in Web people search, this paper proposes a novel graph-based framework, GRAPE (abbr. a Graph-based fRamework for disAmbiguating People appEarances in Web search). In GRAPE, people tag information (e.g., people name, organization, and email address) surrounding the queried people name is extracted from the search results, a graph-based unsupervised algorithm is then developed to cluster the extracted tags, where a new method, Cohesion, is introduced to measure the importance of a tag for clustering, and each final cluster of tags represents a unique people entity. Experimental results show that our proposed framework outperforms the state-of-the-art Web people name disambiguation approaches.
Lili Jiang 0002, Jianyong Wang 0001, Ning An 0001, Shengyuan Wang 0001, Jian Zhan, Lian Li 0001
ICDM1
2009 Two birds with one stone: a graph-based framework for disambiguating and tagging people names in web search
abstract
The ever growing volume of Web data makes it increasingly challenging to accurately find relevant information about a specific person on the Web. To address the challenge caused by name ambiguity in Web people search, this paper explores a novel graph-based framework to both disambiguate and tag people entities in Web search results. Experimental results demonstrate the effectiveness of the proposed framework in tag discovery and name disambiguation.
Lili Jiang 0002, Jianyong Wang 0001, Ning An 0001, Shengyuan Wang 0001, Jian Zhan, Lian Li 0001
WWW1