VLDB 2026 Research / reviewers in the wild / expert
Xuan Wang 0008
dblp:34/4799-8
· DBLP profile ↗
24ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-1381-8958ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning EvaluationabstractWhile small language models (SLMs) have shown promise on various reasoning tasks, their ability to judge the correctness of answers remains unclear compared to large language models (LLMs). Prior work on LLM-as-a-judge frameworks typically relies on comparing candidate answers against ground-truth labels or other candidate answers using predefined metrics like entailment. However, this approach is inherently indirect and difficult to fully automate, offering limited support for fine-grained and scalable evaluation of reasoning outputs. In this work, we propose JudgeBoard, a novel evaluation pipeline that directly queries models to assess the correctness of candidate answers without requiring extra answer comparisons. We focus on two core reasoning domains: mathematical reasoning and science/commonsense reasoning, and construct task-specific evaluation leaderboards using both accuracy-based ranking and an Elo-based rating system across five benchmark datasets, enabling consistent model comparison as judges rather than comparators. To improve judgment performance in lightweight models, we propose MAJ (Multi-Agent Judging), a novel multi-agent evaluation framework that leverages multiple interacting SLMs with distinct reasoning profiles to approximate LLM-level judgment accuracy through collaborative deliberation. Experimental results reveal a significant performance gap between SLMs and LLMs in isolated judging tasks. However, our MAJ framework substantially improves the reliability and consistency of SLMs. On the MATH dataset, MAJ using smaller-sized models as backbones performs comparatively well or even better than their larger-sized counterparts. Our findings highlight that multi-agent SLM systems can potentially match or exceed LLM performance in judgment tasks, with implications for scalable and efficient assessment. Zhenyu Bi, Gaurav Srivastava 0012, Swastik Roy, Morteza Ziyadi, Xuan Wang 0008 |
AAAI | 7 |
| 2026 | LLM4Cell: Taxonomy and Evaluation of LLM and Agentic Models for Single-Cell BiologyabstractSajib Acharjee Dip, Adrika Zafor, Bikash Kumar Paul, Uddip Acharjee Shuvo, Muhit Islam Emon, Xuan Wang, Liqing Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sajib Acharjee Dip, Adrika Zafor, Bikash Kumar Paul, Uddip Acharjee Shuvo, Muhit Islam Emon, Xuan Wang 0008, Liqing Zhang 0002 |
ACL (1) | 6 |
| 2026 | Semantic Prompting: Agentic Incremental Narrative Refinement through Spatial Semantic InteractionabstractInteractive spatial layouts empower users to synthesize information and organize findings for sensemaking. While Large Language Models (LLMs) can automate narrative generation from spatial layouts, current collage-based and re-generation methods struggle to support the incremental spatial refinements inherent to the sensemaking process. We identify three critical gaps in existing spatial-textual generation: interaction-revision misalignment, human-LLM intent misalignment, and lack of granular customization. To address these, we introduce Semantic Prompting, a framework for spatial refinement that perceives semantic interactions, reasons about refinement intent, and performs targeted positional revisions. We implemented S-prism to realize this framework. The empirical evaluation demonstrated that S-prism effectively enhanced the precision of interaction-revision refinement. A user study (N = 14) highlighted how participants leveraged S-prism for incremental formalization through interactive steering. Results showed that users valued its efficient, adaptable, and trustworthy support, which effectively strengthens human-LLM intent alignment. Xuxin Tang, Ibrahim Asadullah Tahmid, Eric Krokos, Kirsten Whitley, Xuan Wang 0008, Chris North 0001 |
AVI | 5 |
| 2025 | DEBATE, TRAIN, EVOLVE: Self-Evolution of Language Model ReasoningabstractLarge language models (LLMs) have improved significantly in their reasoning through extensive training on massive datasets.However, relying solely on additional data for improvement is becoming increasingly impractical, highlighting the need for models to autonomously enhance their reasoning without external supervision.In this paper, we propose DEBATE, TRAIN, EVOLVE (DTE), a novel ground truthfree training framework that uses multi-agent debate traces to evolve a single language model.We also introduce a new prompting strategy REFLECT-CRITIQUE-REFINE, to improve debate quality by explicitly instructing agents to critique and refine their reasoning.Extensive evaluations on seven reasoning benchmarks with six open-weight models show that our DTE framework achieve substantial improvements, with an average accuracy gain of 8.92% on the GSM-PLUS dataset.Furthermore, we observe strong cross-domain generalization, with an average accuracy gain of 5.8% on all other benchmarks, suggesting that our method captures general reasoning capabilities.Our framework code and trained models are publicly available at https://github.com/ctrl- gaurav/Debate-Train-Evolve. Gaurav Srivastava 0012, Zhenyu Bi, Xuan Wang 0008 |
EMNLP | 4 |
| 2025 | ThinkSLM: Towards Reasoning in Small Language ModelsabstractReasoning has long been viewed as an emergent property of large language models (LLMs).However, recent studies challenge this assumption, showing that small language models (SLMs) can also achieve competitive reasoning performance.This paper introduces THINKSLM, the first extensive benchmark to systematically evaluate and study the reasoning abilities of SLMs trained from scratch or derived from LLMs through quantization, pruning, and distillation.We first establish a reliable evaluation criterion comparing available methods and LLM judges against our human evaluations.Then we present a study evaluating 72 diverse SLMs from six major model families across 17 reasoning benchmarks.We repeat all our experiments three times to ensure a robust assessment.Our findings show that: 1) reasoning ability in SLMs is strongly influenced by training methods and data quality rather than solely model scale; 2) quantization preserves reasoning capability, while pruning significantly disrupts it; 3) larger models consistently exhibit higher robustness against adversarial perturbations and intermediate reasoning, but certain smaller models closely match or exceed the larger models' performance.Our findings challenge the assumption that scaling is the only way to achieve strong reasoning.Instead, we foresee a future where SLMs with strong reasoning capabilities can be developed through structured training or post-training compression. Gaurav Srivastava 0012, Shuxiang Cao, Xuan Wang 0008 |
EMNLP | 3 |
| 2025 | Prediction of gene regulatory connections with joint single-cell foundation models and graph-based learningabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) data offers unprecedented opportunities to infer gene regulatory networks (GRNs) at a fine-grained resolution, shedding light on cellular phenotypes at the molecular level. However, the high sparsity, noise, and dropout events inherent in scRNA-seq data pose significant challenges for accurate and reliable GRN inference. The rapid growth in experimentally validated transcription factor-DNA binding data has enabled supervised machine learning methods, which rely on known regulatory interactions to learn patterns, and achieve high accuracy in GRN inference by framing it as a gene regulatory link prediction task. This study addresses the gene regulatory link prediction problem by learning vectorized representations at the gene level to predict missing regulatory interactions. However, a higher performance of supervised learning methods requires a large amount of known TF-DNA binding data, which is often experimentally expensive and therefore limited in amount. Advances in large-scale pre-training and transfer learning provide a transformative opportunity to address this challenge. In this study, we leverage large-scale pre-trained models, trained on extensive scRNA-seq datasets and known as single-cell foundation models (scFMs). These models are combined with joint graph-based learning to establish a robust foundation for gene regulatory link prediction. RESULTS: We propose scRegNet, a novel and effective framework that leverages scFMs with joint graph-based learning for gene regulatory link prediction. scRegNet achieves state-of-the-art results in comparison with nine baseline methods on seven scRNA-seq benchmark datasets. Additionally, scRegNet is more robust than the baseline methods on noisy training data. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/sindhura-cs/scRegNet. Sindhura Kommu, Yizhi Wang 0009, Yue Joseph Wang, Xuan Wang 0008 |
Bioinform. | 4 |
| 2024 | TTM-RE: Memory-Augmented Document-Level Relation ExtractionabstractDocument-level relation extraction aims to categorize the association between any two entities within a document.We find that previous methods for document-level relation extraction are ineffective in exploiting the full potential of large amounts of training data with varied noise levels.For example, in the ReDo-cRED benchmark dataset, state-of-the-art methods trained on the large-scale, lower-quality, distantly supervised training data generally do not perform better than those trained solely on the smaller, high-quality, human-annotated training data.To unlock the full potential of large-scale noisy training data for documentlevel relation extraction, we propose TTM-RE, a novel approach that integrates a trainable memory module, known as the Token Turing Machine, with a noisy-robust loss function that accounts for the positive-unlabeled setting.Extensive experiments on ReDocRED, a benchmark dataset for document-level relation extraction, reveal that TTM-RE achieves state-ofthe-art performance (with an absolute F1 score improvement of over 3%).Ablation studies further illustrate the superiority of TTM-RE in other domains (the ChemDisGene dataset in the biomedical domain) and under highly unlabeled settings. Chufan Gao, Xuan Wang 0008, Jimeng Sun 0001 |
ACL (1) | 2 |
| 2024 | EEG2Text: Open Vocabulary EEG-to-Text Translation with Multi-View TransformerabstractDeciphering the intricacies of the human brain has captivated curiosity for centuries. Recent strides in Brain-Computer Interface (BCI) technology, particularly using motor imagery, have restored motor functions such as reaching, grasping, and walking in paralyzed individuals. However, unraveling natural language from brain signals remains a formidable challenge. Electroencephalography (EEG) is a non-invasive technique used to record electrical activity in the brain by placing electrodes on the scalp. Previous studies of EEG-to-text decoding have achieved high accuracy on small closed vocabularies, but still fall short of high accuracy when dealing with large open vocabularies. We propose a novel method, EEG2Text, to improve the accuracy of open vocabulary EEG-to-text decoding. Specifically, EEG2Text leverages EEG pre-training to enhance the learning of semantics from EEG signals and proposes a multi-view transformer to model the EEG signal processing by different t spatial regions of the brain. Experiments show that EEG2Text has superior performance, outperforming the state-of-the-art baseline methods by a large margin of up to 5% in absolute BLEU and ROUGE scores. EEG2Text shows great potential for a high-performance open-vocabulary brain-to-text system to facilitate communication. Our code is available at: https://github.com/ForeverNightmare/EEG2Text. Daniel Hajialigol, Benny Antony, Aiguo Han, Xuan Wang 0008 |
IEEE Big Data | 5 |
| 2024 | OntoType: Ontology-Guided and Pre-Trained Language Model Assisted Fine-Grained Entity TypingabstractFine-grained entity typing (FET), which assigns entities in text with context-sensitive, fine-grained semantic types, is a basic but important task for knowledge extraction from unstructured text. FET has been studied extensively in natural language processing and typically relies on human-annotated corpora for training, which is costly and difficult to scale. Recent studies explore the utilization of pre-trained language models (PLMs) as a knowledge base to generate rich and context-aware weak supervision for FET. However, a PLM still requires direction and guidance to serve as a knowledge base as they often generate a mixture of rough and fine-grained types, or tokens unsuitable for typing. In this study, we vision that an ontology provides a semantics-rich, hierarchical structure, which will help select the best results generated by multiple PLM models and head words. Specifically, we propose a novel annotation-free, ontology-guided FET method, OntoType, which follows a type ontological structure, from coarse to fine, ensembles multiple PLM prompting results to generate a set of type candidates, and refines its type resolution, under the local context with a natural language inference model. Our experiments on the Ontonotes, FIGER, and NYT datasets using their associated ontological structures demonstrate that our method outperforms the state-of-the-art zero-shot fine-grained entity typing methods as well as a typical LLM method, ChatGPT. Our error analysis shows that refinement of the existing ontology structures will further improve fine-grained entity typing. Tanay Komarlu, Minhao Jiang, Xuan Wang 0008, Jiawei Han 0001 |
KDD | 3 |
| 2022 | REACTCLASS: Cross-Modal Supervision for Subword-Guided Reactant Entity ClassificationabstractWe propose REACTCLASS that automatically maps the low-level concrete chemical entities into the high-level reactant groups without human effort for training data annotation. REACTCLASS is designed to take two special characteristics of the chemical molecules into consideration. The first characteristic is that each chemical molecule can be represented in two modalities: a chemical name in the text and a molecular structure in the graph. We propose to use cross-modal supervision to automatically create the training data for chemical name classification in the text via molecular structure matching in the graph. The second characteristic is that there is a knowledge-aware subword correlation between the surface names of the chemical entities to be classified and that of the reactant groups as class labels. We propose to train a classification model based on the subword cross-attention map between each chemical name and the corresponding reaction group. Experiments demonstrate that REACTCLASS is highly effective, achieving state-of-the-art performance in classifying the chemical names into human-defined reactant groups without requiring human effort for training data annotation. Xuan Wang 0008, Vivian Hu, Minhao Jiang, Yu Zhang 0044, Jinfeng Xiao, Danielle Cherrice Loving, Heng Ji 0001, Martin D. Burke, Jiawei Han 0001 |
BIBM | 1 |
| 2022 | New Frontiers of Scientific Text Mining: Tasks, Data, and ToolsabstractExploring the vast amount of rapidly growing scientific text data is highly beneficial for real-world scientific discovery. However, scientific text mining is particularly challenging due to the lack of specialized domain knowledge in natural language context, complex sentence structures in scientific writing, and multi-modal representations of scientific knowledge. This tutorial presents a comprehensive overview of recent research and development on scientific text mining, focusing on the biomedical and chemistry domains. First, we introduce the motivation and unique challenges of scientific text mining. Then we discuss a set of methods that perform effective scientific information extraction, such as named entity recognition, relation extraction, and event extraction. We also introduce real-world applications such as textual evidence retrieval, scientific topic contrasting for drug discovery, and molecule representation learning for reaction prediction. Finally, we conclude our tutorial by demonstrating, on real-world datasets (COVID-19 and organic chemistry literature), how the information can be extracted and retrieved, and how they can assist further scientific discovery. We also discuss the emerging research problems and future directions for scientific text mining. Xuan Wang 0008, Heng Ji 0001, Jiawei Han 0001 |
KDD | 1 |
| 2022 | Seed-Guided Topic Discovery with Out-of-Vocabulary SeedsabstractYu Zhang, Yu Meng, Xuan Wang, Sheng Wang, Jiawei Han. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yu Zhang 0044, Yu Meng 0001, Xuan Wang 0008, Sheng Wang 0012, Jiawei Han 0001 |
NAACL-HLT | 3 |
| 2021 | InterHG: an Interpretable and Accurate Model for Hypothesis GenerationabstractHypothesis generation, which tries to identify implicit associations between two concepts, has attracted much attention due to its ability of linking key concepts scattered in different articles and enriching plausible new hypotheses. Among existing approaches for hypothesis generation, matrix factorization based methods have achieved start-of-the-art performance. However, matrix factorization based methods suffer from the following limitations: 1) Bridge concepts are determined only as a post-hoc analysis of matrix factorization results; 2) The embeddings of concepts by matrix factorization cannot be explained, and thus it is hard to understand whether the concepts are linked in a semantically meaningful way. To overcome these limitations, we propose an interpretable and accurate hypothesis generation model (InterHG), which improves both accuracy and interpretability compared with existing methods. First, we propose to explicitly model the relationship between bridge concepts and given concept pairs, and conduct tensor factorization to identify link concepts. This reduces information loss and improves accuracy compared with post-hoc approaches. Second, we leverage the description of categories in the tensor factorization, which can output concept embedding as a weighted combination of known categories. With this meaningful embedding representation, medical researchers are able to check the correctness of the suggested link concepts for a given concept pair. We conduct experiments based on MeSH terms (a controlled vocabulary of biomedical concepts) extracted from MEDLINE corpus and category information obtained from UMLS (a comprehensive biomedical concept database). Results demonstrate that the proposed InterHG is highly accurate and produces meaningful embeddings for explanations. Haoyu Wang 0004, Xuan Wang 0008, Yaqing Wang 0001, Guangxu Xun, Kishlay Jha, Jing Gao 0004 |
BIBM | 2 |
| 2021 | Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-TrainingabstractWe study the problem of training named entity recognition (NER) models using only distantly-labeled data, which can be automatically obtained by matching entity mentions in the raw text with entity types in a knowledge base.The biggest challenge of distantlysupervised NER is that the distant supervision may induce incomplete and noisy labels, rendering the straightforward application of supervised learning ineffective.In this paper, we propose (1) a noise-robust learning scheme comprised of a new loss function and a noisy label removal step, for training NER models on distantly-labeled data, and (2) a self-training method that uses contextualized augmentations created by pre-trained language models to improve the generalization ability of the NER model.On three benchmark datasets, our method achieves superior performance, outperforming existing distantlysupervised NER models by significant margins 1 . Yu Meng 0001, Yunyi Zhang 0001, Jiaxin Huang 0001, Xuan Wang 0008, Yu Zhang 0044, Heng Ji 0001, Jiawei Han 0001 |
EMNLP (1) | 4 |
| 2021 | ChemNER: Fine-Grained Chemistry Named Entity Recognition with Ontology-Guided Distant SupervisionabstractScientific literature analysis needs fine-grained named entity recognition (NER) to provide a wide range of information for scientific discovery.For example, chemistry research needs to study dozens to hundreds of distinct, fine-grained entity types, making consistent and accurate annotation difficult even for crowds of domain experts.On the other hand, domain-specific ontologies and knowledge bases (KBs) can be easily accessed, constructed, or integrated, which makes distant supervision realistic for fine-grained chemistry NER.In distant supervision, training labels are generated by matching mentions in a document with the concepts in the knowledge bases (KBs).However, this kind of KB-matching suffers from two major challenges: incomplete annotation and noisy annotation.We propose CHEMNER, an ontologyguided, distantly-supervised method for finegrained chemistry NER to tackle these challenges.It leverages the chemistry type ontology structure to generate distant labels with novel methods of flexible KB-matching and ontology-guided multi-type disambiguation.It significantly improves the distant label generation for the subsequent sequence labeling model training.We also provide an expertlabeled, chemistry NER dataset with 62 finegrained chemistry types (e.g., chemical compounds and chemical reactions).Experimental results show that CHEMNER is highly effective, outperforming substantially the stateof-the-art NER methods (with .25 absolute F1 score improvement). Xuan Wang 0008, Vivian Hu, Xiangchen Song, Shweta Garg 0004, Jinfeng Xiao, Jiawei Han 0001 |
EMNLP (1) | 1 |
| 2020 | Fine-Grained Named Entity Recognition with Distant Supervision in COVID-19 LiteratureabstractBiomedical named entity recognition (BioNER) is a fundamental step for mining COVID-19 literature. Existing BioNER datasets cover a few common coarse-grained entity types (e.g., genes, chemicals, and diseases), which cannot be used to recognize highly domain-specific entity types (e.g., animal models of diseases) or emerging ones (e.g., coronaviruses) for COVID-19 studies. We present CORD-NER, a fine-grained named entity recognized dataset of COVID-19 literature (up until May 19, 2020). CORD-NER contains over 12 million sentences annotated via distant supervision. Also included in CORD-NER are 2,000 manually-curated sentences as a test set for performance evaluation. CORD-NER covers 75 fine-grained entity types. In addition to the common biomedical entity types, it covers new entity types specifically related to COVID-19 studies, such as coronaviruses, viral proteins, evolution, and immune responses. The dictionaries of these fine-grained entity types are collected from existing knowledge bases and human-input seed sets. We further present DISTNER, a distantly supervised NER model that relies on a massive unlabeled corpus and a collection of dictionaries to annotate the COVID-19 corpus. DISTNER provides a benchmark performance on the CORD-NER test set for future research. Xuan Wang 0008, Xiangchen Song, Bangzheng Li, Kang Zhou 0002, Qi Li 0012, Jiawei Han 0001 |
BIBM | 1 |
| 2020 | Pattern-enhanced Named Entity Recognition with Distant SupervisionabstractSupervised deep learning methods have achieved state-of-the-art performance on the task of named entity recognition (NER). However, such methods suffer from high cost and low efficiency in training data annotation, leading to highly specialized NER models that cannot be easily adapted to new domains. Recently, distant supervision has been applied to replace human annotation, thanks to the fast development of domain-specific knowledge bases. However, the generated noisy labels pose significant challenges in learning effective neural models with distant supervision. We propose PatNER, a distantly supervised NER model that effectively deals with noisy distant supervision from domain-specific dictionaries. PatNER does not require human-annotated training data but only relies on unlabeled data and incomplete domain-specific dictionaries for distant supervision. It incorporates the distant labeling uncertainty into the neural model training to enhance distant supervision. We go beyond the traditional sequence labeling framework and propose a more effective fuzzy neural model using the tie-or-break tagging scheme for the NER task. Extensive experiments on three benchmark datasets in two domains demonstrate the power of PatNER. Case studies on two additional real-world datasets demonstrate that PatNER improves the distant NER performance in both entity boundary detection and entity type recognition. The results show a great promise in supporting high quality named entity recognition with domain-specific dictionaries on a wide variety of entity types. Xuan Wang 0008, Yingjun Guan, Yu Zhang 0044, Qi Li 0012, Jiawei Han 0001 |
IEEE BigData | 1 |
| 2020 | Textual Evidence Mining via Spherical Heterogeneous Information Network EmbeddingabstractScientific literature, as one of the major knowledge resources, provides abundant textual evidence that has great potential to support high-quality scientific hypothesis validation. In this paper, we study the problem of textual evidence mining in scientific literature: given a scientific hypothesis as a query triplet, find the textual evidence sentences in scientific literature that support the input query. A critical challenge for textual evidence mining in scientific literature is to retrieve high-quality textual evidence without human supervision. Because it is non-trivial to obtain a large set of human-annotated articles containing evidence sentences in scientific literature. To tackle this challenge, we propose EvidenceMiner, a high-quality textual evidence retrieval method for scientific literature without human-annotated training examples. To achieve high-quality textual evidence retrieval, we leverage heterogeneous information from both existing knowledge bases and massive unstructured text. We propose to construct a large heterogeneous information network (HIN) to build connections between the user-input queries and the candidate evidence sentences. Based on the constructed HIN, we propose a novel HIN embedding method that directly embeds the nodes onto a spherical space to improve the retrieval performance. Quantitative experiments on a huge biomedical literature corpus (over 4 million sentences) demonstrate that EvidenceMiner significantly outperforms baseline methods for unsupervised textual evidence retrieval. Case studies also demonstrate that our HIN construction and embedding greatly benefit many downstream applications such as textual evidence interpretation and synonym meta-pattern discovery. Xuan Wang 0008, Yu Zhang 0044, Aabhas Chauhan, Qi Li 0012, Jiawei Han 0001 |
IEEE BigData | 1 |
| 2020 | Minimally Supervised Categorization of Text with MetadataabstractDocument categorization, which aims to assign a topic label to each document, plays a fundamental role in a wide variety of applications. Despite the success of existing studies in conventional supervised document classification, they are less concerned with two real problems: (1)the presence of metadata : in many domains, text is accompanied by various additional information such as authors and tags. Such metadata serve as compelling topic indicators and should be leveraged into the categorization framework; (2)label scarcity: labeled training samples are expensive to obtain in some cases, where categorization needs to be performed using only a small set of annotated data. In recognition of these two challenges, we propose MetaCat, a minimally supervised framework to categorize text with metadata. Specifically, we develop a generative process describing the relationships between words, documents, labels, and metadata. Guided by the generative model, we embed text and metadata into the same semantic space to encode heterogeneous signals. Then, based on the same generative process, we synthesize training samples to address the bottleneck of label scarcity. We conduct a thorough evaluation on a wide range of datasets. Experimental results prove the effectiveness of MetaCat over many competitive baselines. Yu Zhang 0044, Yu Meng 0001, Jiaxin Huang 0001, Frank F. Xu, Xuan Wang 0008, Jiawei Han 0001 |
SIGIR | 5 |
| 2019 | Distantly Supervised Biomedical Named Entity Recognition with Dictionary ExpansionabstractState-of-the-art biomedical named entity recognition (BioNER) systems apply supervised machine learning models (i.e., relying on human effort for training data annotation) which are not easy to be generalized to new entity types and datasets. We propose a distantly supervised approach, AutoBioNER, that automatically recognizes biomedical entities from massive corpora with user-input dictionaries. AutoBioNER does not need any human annotated data. It relies on incomplete entity dictionaries to provide seeds for each entity type and performs a novel entity set expansion step for corpus-level new entity recognition and dictionary completion. The expanded dictionaries are used as distant supervision to train a neural model for BioNER. Experimental results show that AutoBioNER achieves the best performance among the methods that only use dictionaries with no additional human effort on BioNER benchmark datasets. It is also demonstrated that the dictionary expansion step plays an important role in the great performances. Xuan Wang 0008, Yu Zhang 0044, Qi Li 0012, Xiang Ren 0001, Jingbo Shang, Jiawei Han 0001 |
BIBM | 1 |
| 2019 | HiGitClass: Keyword-Driven Hierarchical Classification of GitHub RepositoriesabstractGitHub has become an important platform for code sharing and scientific exchange. With the massive number of repositories available, there is a pressing need for topic-based search. Even though the topic label functionality has been introduced, the majority of GitHub repositories do not have any labels, impeding the utility of search and topic-based analysis. This work targets the automatic repository classification problem as keyword-driven hierarchical classification. Specifically, users only need to provide a label hierarchy with keywords to supply as supervision. This setting is flexible, adaptive to the users' needs, accounts for the different granularity of topic labels and requires minimal human effort. We identify three key challenges of this problem, namely (1) the presence of multi-modal signals; (2) supervision scarcity and bias; (3) supervision format mismatch. In recognition of these challenges, we propose the HiGitClass framework, comprising of three modules: heterogeneous information network embedding; keyword enrichment; topic modeling and pseudo document generation. Experimental results on two GitHub repository collections confirm that HiGitClass is superior to existing weakly-supervised and dataless hierarchical classification methods, especially in its ability to integrate both structured and unstructured data for repository classification. Code and datasets related to this paper are available at https://github.com/yuzhimanhua/HiGitClass. Yu Zhang 0044, Frank F. Xu, Yu Meng 0001, Xuan Wang 0008, Qi Li 0012, Jiawei Han 0001 |
ICDM | 5 |
| 2019 | Cross-type biomedical named entity recognition with deep multi-task learningabstractMOTIVATION: State-of-the-art biomedical named entity recognition (BioNER) systems often require handcrafted features specific to each entity type, such as genes, chemicals and diseases. Although recent studies explored using neural network models for BioNER to free experts from manual feature engineering, the performance remains limited by the available training data for each entity type. RESULTS: We propose a multi-task learning framework for BioNER to collectively use the training data of different types of entities and improve the performance on each of them. In experiments on 15 benchmark BioNER datasets, our multi-task model achieves substantially better performance compared with state-of-the-art BioNER systems and baseline neural sequence labeling models. Further analysis shows that the large performance gains come from sharing character- and word-level information among relevant biomedical entities across differently labeled corpora. AVAILABILITY AND IMPLEMENTATION: Our source code is available at https://github.com/yuzhimanhua/lm-lstm-crf. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xuan Wang 0008, Yu Zhang 0044, Xiang Ren 0001, Yuhao Zhang 0004, Marinka Zitnik, Jingbo Shang, Curt Langlotz, Jiawei Han 0001 |
Bioinform. | 1 |
| 2018 | PENNER: Pattern-enhanced Nested Named Entity Recognition in Biomedical Literature
Xuan Wang 0008, Yu Zhang 0044, Qi Li 0012, Cathy H. Wu, Jiawei Han 0001 |
BIBM | 1 |
| 2018 | Pattern Discovery for Wide-Window Open Information Extraction in Biomedical Literature
Qi Li 0012, Xuan Wang 0008, Yu Zhang 0044, Fei Ling, Cathy H. Wu, Jiawei Han 0001 |
BIBM | 2 |