EDBT 2026 Demo / reviewers in the wild / expert
Sheng Wang 0012
dblp:85/1868-12
· DBLP profile ↗
32ranked-venue papers
2as first author
25since 2021 · last 2025
0000-0002-0439-5199ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GrInAdapt: Source-Free Multi-target Domain Adaptation for Retinal Vessel Segmentation
Zixuan Liu 0001, Aaron Honjaya, Yuekai Xu, Hefu Pan, Xin Wang 0113, Linda G. Shapiro, Sheng Wang 0012, Ruikang K. Wang |
MICCAI (5) | 8 |
| 2024 | H2D: Hierarchical Heterogeneous Graph Learning Framework for Drug-Drug Interaction PredictionabstractAccurately predicting Drug-Drug Interactions (DDIs) is critical to designing effective drug combination therapies. Recently, Artificial Intelligence (AI)-powered DDI prediction approaches have emerged as a new paradigm. However, most existing methods oversimplify the complex hierarchical structure within molecules and overlook the multi-source heterogeneous information external to molecules, limiting their modeling and predictive capabilities. To address this, we propose a H ierarchical H eterogeneous graph learning framework for D DI prediction, namely H2D. H2D employs an internal-to-external, local-to-global hierarchical perspective, exploiting intra-molecular multi-granularity structures and inter-molecular biomedical interactions to mutually enhance across hierarchical levels. Extensive experimental results demonstrate H2D's effectiveness on three real-world DDI prediction tasks (binary-class, multi-class, and multi-label). In sum, H2D achieves state-of-the-art performance in DDI prediction by leveraging the multi-scale graph structures, opening up new avenues in AI-powered DDI prediction. Ran Zhang 0008, Xuezhi Wang 0004, Sheng Wang 0012, Kunpeng Liu 0001, Yuanchun Zhou, Pengfei Wang 0008 |
CIKM | 3 |
| 2024 | A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryabstractIn many scientific fields, large language models (LLMs) have revolutionized the way text and other modalities of data (e.g., molecules and proteins) are handled, achieving superior performance in various applications and augmenting the scientific discovery process.Nevertheless, previous surveys on scientific LLMs often concentrate on one or two fields or a single modality.In this paper, we aim to provide a more holistic view of the research landscape by unveiling cross-field and cross-modal connections between scientific LLMs regarding their architectures and pretraining techniques.To this end, we comprehensively survey over 260 scientific LLMs, discuss their commonalities and differences, as well as summarize pre-training datasets and evaluation tasks for each field and modality.Moreover, we investigate how LLMs have been deployed to benefit scientific discovery.Resources related to this survey are available at https://github.com/yuzhimanhua/ Awesome-Scientific-Language-Models. Yu Zhang 0044, Xiusi Chen, Bowen Jin, Sheng Wang 0012, Shuiwang Ji, Wei Wang 0010, Jiawei Han 0001 |
EMNLP | 4 |
| 2024 | MMICL: Empowering Vision-language Model with Multi-Modal In-Context LearningabstractSince the resurgence of deep learning, vision-language models (VLMs) enhanced by large language models (LLMs) have grown exponentially in popularity.
However, while LLMs can utilize extensive background knowledge and task information with in-context learning, most VLMs still struggle with understanding complex multi-modal prompts with multiple images, making VLMs less effective in downstream vision-language tasks.
In this paper, we address the limitation above by 1) introducing vision-language Model with **M**ulti-**M**odal **I**n-**C**ontext **L**earning(MMICL), a new approach to allow the VLM to deal with multi-modal inputs efficiently; 2) proposing a novel context scheme to augment the in-context learning ability of the VLM; 3) constructing the Multi-modal In-Context Learning (MIC) dataset, designed to enhance the VLM's ability to understand complex multi-modal prompts.
Our experiments confirm that MMICL achieves new state-of-the-art zero-shot performance on a wide range of general vision-language tasks, especially for complex benchmarks, including MME and MMBench. Our analysis demonstrates that MMICL effectively tackles the challenge of complex multi-modal prompt understanding and emerges the impressive ICL ability. Furthermore, we observe that MMICL successfully alleviates language bias in VLMs, a common issue for VLMs that often leads to hallucination when faced with extensive textual context.
Our code, dataset, dataset tool, and model are available at https://github.com/PKUnlp-icler/MIC. Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma 0001, Kaikai An, Liang Chen 0024, Zixuan Liu 0001, Sheng Wang 0012, Wenjuan Han, Baobao Chang |
ICLR | 8 |
| 2024 | Enhancing Hi-C contact matrices for loop detection with Capricorn: a multiview diffusion modelabstractMOTIVATION: High-resolution Hi-C contact matrices reveal the detailed three-dimensional architecture of the genome, but high-coverage experimental Hi-C data are expensive to generate. Simultaneously, chromatin structure analyses struggle with extremely sparse contact matrices. To address this problem, computational methods to enhance low-coverage contact matrices have been developed, but existing methods are largely based on resolution enhancement methods for natural images and hence often employ models that do not distinguish between biologically meaningful contacts, such as loops and other stochastic contacts. RESULTS: We present Capricorn, a machine learning model for Hi-C resolution enhancement that incorporates small-scale chromatin features as additional views of the input Hi-C contact matrix and leverages a diffusion probability model backbone to generate a high-coverage matrix. We show that Capricorn outperforms the state of the art in a cross-cell-line setting, improving on existing methods by 17% in mean squared error and 26% in F1 score for chromatin loop identification from the generated high-coverage data. We also demonstrate that Capricorn performs well in the cross-chromosome setting and cross-chromosome, cross-cell-line setting, improving the downstream loop F1 score by 14% relative to existing methods. We further show that our multiview idea can also be used to improve several existing methods, HiCARN and HiCNN, indicating the wide applicability of this approach. Finally, we use DNA sequence to validate discovered loops and find that the fraction of CTCF-supported loops from Capricorn is similar to those identified from the high-coverage data. Capricorn is a powerful Hi-C resolution enhancement method that enables scientists to find chromatin features that cannot be identified in the low-coverage contact matrix. AVAILABILITY AND IMPLEMENTATION: Implementation of Capricorn and source code for reproducing all figures in this paper are available at https://github.com/CHNFTQ/Capricorn. Tangqi Fang, Yifeng Liu 0004, Addie Woicik, Minsi Lu, Anupama Jha, Xiao Wang 0013, Gang Li 0034, Borislav H. Hristov, Zixuan Liu 0001, William Stafford Noble, Sheng Wang 0012 |
Bioinform. | 12 |
| 2023 | GraphPrompt: Graph-Based Prompt Templates for Biomedical Synonym PredictionabstractIn the expansion of biomedical dataset, the same category may be labeled with different terms, thus being tedious and onerous to curate these terms. Therefore, automatically mapping synonymous terms onto the ontologies is desirable, which we name as biomedical synonym prediction task. Unlike biomedical concept normalization (BCN), no clues from context can be used to enhance synonym prediction, making it essential to extract graph features from ontology. We introduce an expert-curated dataset OBO-syn encompassing 70 different types of concepts and 2 million curated concept-term pairs for evaluating synonym prediction methods. We find BCN methods perform weakly on this task for not making full use of graph information. Therefore, we propose GraphPrompt, a prompt-based learning approach that creates prompt templates according to the graphs. GraphPrompt obtained 37.2% and 28.5% improvement on zero-shot and few-shot settings respectively, indicating the effectiveness of these graph-based prompt templates. We envision that our method GraphPrompt and OBO-syn dataset can be broadly applied to graph-based NLP tasks, and serve as the basis for analyzing diverse and accumulating biomedical data. All the data and codes are avalible at: https://github.com/HanwenXuTHU/GraphPrompt Jiayou Zhang, Shizhuo Zhang, Megh Manoj Bhalerao, Yucong Liu, Sheng Wang 0012 |
AAAI | 8 |
| 2023 | SST: Semantic and Structural Transformers for Hierarchy-aware Language Models in E-commerceabstractHierarchies are common structures used to organize data, such as e-commerce hierarchies associated with product data. With these product hierarchies, we aim to learn hierarchy-aware product text embeddings to improve fine-tuning performance on a variety of downstream e-commerce tasks. Existing methods leverage hierarchies by either aligning the text embeddings to separate hierarchical embeddings or by aligning the hierarchical information implicitly within a unified text Transformer. Although these models optimize to predict hierarchy information, performing further fine-tuning on new tasks is non-trivial. To bridge this gap, we propose a pre-training architecture to implicitly encode the hierarchy within the product text and then directly leverage a sub-set of the pre-training model during fine-tuning. Pre-training is done through Semantic and Structural Transformers (SST) where the Semantic-Transformer first encodes the product text into a contextual embedding, which is then used by the Structural-Transformer to infer the product’s path in the hierarchy. Fine-tuning is done using only the initial Semantic-Transformer, now that hierarchy-aware text embeddings are learned. With this design, we eliminate the need of linking each fine-tuning dataset with corresponding hierarchies. This leads to fine-tuning performance improvements on critical e-commerce downstream tasks over the existing state-of-the-art hierarchy models, even when hierarchy data $is$ available during fine-tuning. Moreover, this improvement is consistent even after augmenting our baseline models to support fine-tuning. We conclude by discussing how such implicit structural encodings can be leveraged beyond the e-commerce domain. Karan Samel, Houyu Zhang, Jun Ma 0029, Haoming Jiang, Qing Ping, Sheng Wang 0012, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi |
IEEE Big Data | 6 |
| 2023 | ForeSeer: Product Aspect Forecasting Using Temporal Graph EmbeddingabstractDeveloping text mining approaches to mine aspects from customer reviews has been well-studied due to its importance in understanding customer needs and product attributes. In contrast, it remains unclear how to predict the future emerging aspects of a new product that currently has little review information. This task, which we named product aspect forecasting, is critical for recommending new products, but also challenging because of the missing reviews. Here, we propose ForeSeer, a novel textual mining and product embedding approach progressively trained on temporal product graphs for this novel product aspect forecasting task. ForeSeer transfers reviews from similar products on a large product graph and exploits these reviews to predict aspects that might emerge in future reviews. A key novelty of our method is to jointly provide review, product, and aspect embeddings that are both time-sensitive and less affected by extremely imbalanced aspect frequencies. We evaluated ForeSeer on a real-world product review system containing 11,536,382 reviews and 11,000 products over 3 years. We observe that ForeSeer substantially outperformed existing approaches with at least 49.1% AUPRC improvement under the real setting where aspect associations are not given. ForeSeer further improves future link prediction on the product graph and the review aspect association prediction. Collectively, Foreseer offers a novel framework for review forecasting by effectively integrating review text, product network, and temporal information, opening up new avenues for online shopping recommendation and e-commerce applications. Zixuan Liu 0001, Gaurush Hiranandani, Kun Qian 0018, Edward W. Huang, Yi Xu 0011, Belinda Zeng, Karthik Subbian, Sheng Wang 0012 |
CIKM | 8 |
| 2023 | Graph-Aware Language Model Pre-Training on a Large Graph Corpus Can Help Multiple Graph ApplicationsabstractModel pre-training on large text corpora has been demonstrated effective for various downstream applications in the NLP domain. In the graph mining domain, a similar analogy can be drawn for pre-training graph models on large graphs in the hope of benefiting downstream graph applications, which has also been explored by several recent studies. However, no existing study has ever investigated the pre-training of text plus graph models on large heterogeneous graphs with abundant textual information (a.k.a. large graph corpora) and then fine-tuning the model on different related downstream applications with different graph schemas. To address this problem, we propose a framework of graph-aware language model pre-training (GaLM) on a large graph corpus, which incorporates large language models and graph neural networks, and a variety of fine-tuning methods on downstream applications. We conduct extensive experiments on Amazon's real internal datasets and large public datasets. Comprehensive empirical results and in-depth analysis demonstrate the effectiveness of our proposed methods along with lessons learned. Da Zheng 0004, Jun Ma 0029, Houyu Zhang, Vassilis N. Ioannidis, Xiang Song 0003, Qing Ping, Sheng Wang 0012, Carl Yang 0001, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi |
KDD | 8 |
| 2023 | Supervised biological network alignment with graph neural networksabstractMOTIVATION: Despite the advances in sequencing technology, massive proteins with known sequences remain functionally unannotated. Biological network alignment (NA), which aims to find the node correspondence between species' protein-protein interaction (PPI) networks, has been a popular strategy to uncover missing annotations by transferring functional knowledge across species. Traditional NA methods assumed that topologically similar proteins in PPIs are functionally similar. However, it was recently reported that functionally unrelated proteins can be as topologically similar as functionally related pairs, and a new data-driven or supervised NA paradigm has been proposed, which uses protein function data to discern which topological features correspond to functional relatedness. RESULTS: Here, we propose GraNA, a deep learning framework for the supervised NA paradigm for the pairwise NA problem. Employing graph neural networks, GraNA utilizes within-network interactions and across-network anchor links for learning protein representations and predicting functional correspondence between across-species proteins. A major strength of GraNA is its flexibility to integrate multi-faceted non-functional relationship data, such as sequence similarity and ortholog relationships, as anchor links to guide the mapping of functionally related proteins across species. Evaluating GraNA on a benchmark dataset composed of several NA tasks between different pairs of species, we observed that GraNA accurately predicted the functional relatedness of proteins and robustly transferred functional annotations across species, outperforming a number of existing NA methods. When applied to a case study on a humanized yeast network, GraNA also successfully discovered functionally replaceable human-yeast protein pairs that were documented in previous studies. AVAILABILITY AND IMPLEMENTATION: The code of GraNA is available at https://github.com/luo-group/GraNA. Kerr Ding, Sheng Wang 0012, Yunan Luo |
Bioinform. | 2 |
| 2023 | Gemini: memory-efficient integration of hundreds of gene networks with high-order poolingabstractMOTIVATION: The exponential growth of genomic sequencing data has created ever-expanding repositories of gene networks. Unsupervised network integration methods are critical to learn informative representations for each gene, which are later used as features for downstream applications. However, these network integration methods must be scalable to account for the increasing number of networks and robust to an uneven distribution of network types within hundreds of gene networks. RESULTS: To address these needs, we present Gemini, a novel network integration method that uses memory-efficient high-order pooling to represent and weight each network according to its uniqueness. Gemini then mitigates the uneven network distribution through mixing up existing networks to create many new networks. We find that Gemini leads to more than a 10% improvement in F1 score, 15% improvement in micro-AUPRC, and 63% improvement in macro-AUPRC for human protein function prediction by integrating hundreds of networks from BioGRID, and that Gemini's performance significantly improves when more networks are added to the input network collection, while Mashup and BIONIC embeddings' performance deteriorates. Gemini thereby enables memory-efficient and informative network integration for large gene networks and can be used to massively integrate and analyze networks in other domains. AVAILABILITY AND IMPLEMENTATION: Gemini can be accessed at: https://github.com/MinxZ/Gemini. Addie Woicik, Mingxin Zhang 0008, Sara Mostafavi, Sheng Wang 0012 |
Bioinform. | 5 |
| 2023 | POPDx: an automated framework for patient phenotyping across 392 246 individuals in the UK Biobank studyabstractOBJECTIVE: For the UK Biobank, standardized phenotype codes are associated with patients who have been hospitalized but are missing for many patients who have been treated exclusively in an outpatient setting. We describe a method for phenotype recognition that imputes phenotype codes for all UK Biobank participants. MATERIALS AND METHODS: POPDx (Population-based Objective Phenotyping by Deep Extrapolation) is a bilinear machine learning framework for simultaneously estimating the probabilities of 1538 phenotype codes. We extracted phenotypic and health-related information of 392 246 individuals from the UK Biobank for POPDx development and evaluation. A total of 12 803 ICD-10 diagnosis codes of the patients were converted to 1538 phecodes as gold standard labels. The POPDx framework was evaluated and compared to other available methods on automated multiphenotype recognition. RESULTS: POPDx can predict phenotypes that are rare or even unobserved in training. We demonstrate substantial improvement of automated multiphenotype recognition across 22 disease categories, and its application in identifying key epidemiological features associated with each phenotype. CONCLUSIONS: POPDx helps provide well-defined cohorts for downstream studies. It is a general-purpose method that can be applied to other biobanks with diverse but incomplete data. Sheng Wang 0012, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 2 |
| 2022 | DisenCite: Graph-Based Disentangled Representation Learning for Context-Specific Citation GenerationabstractCiting and describing related literature are crucial to scientific writing. Many existing approaches show encouraging performance in citation recommendation, but are unable to accomplish the more challenging and onerous task of citation text generation. In this paper, we propose a novel disentangled representation based model DisenCite to automatically generate the citation text through integrating paper text and citation graph. A key novelty of our method compared with existing approaches is to generate context-specific citation text, empowering the generation of different types of citations for the same paper. In particular, we first build and make available a graph enhanced contextual citation dataset (GCite) with 25K edges in different types characterized by citation contained sections over 4.8K research papers. Based on this dataset, we encode each paper according to both textual contexts and structure information in the heterogeneous citation graph. The resulted paper representations are then disentangled by the mutual information regularization between this paper and its neighbors in graph. Extensive experiments demonstrate the superior performance of our method comparing to state-of-the-art approaches. We further conduct ablation and case studies to reassure that the improvement of our method comes from generating the context-specific citation through incorporating the citation graph. Yifan Wang 0014, Yiping Song, Chaoran Cheng, Wei Ju 0001, Ming Zhang 0004, Sheng Wang 0012 |
AAAI | 7 |
| 2022 | Textomics: A Dataset for Genomics Data Summary GenerationabstractSummarizing biomedical discovery from genomics data using natural languages is an essential step in biomedical research but is mostly done manually.Here, we introduce Textomics, a novel dataset of genomics data description, which contains 22,273 pairs of genomics data matrices and their summaries.Each summary is written by the researchers who generated the data and associated with a scientific paper.Based on this dataset, we study two novel tasks: generating textual summary from a genomics data matrix and vice versa.Inspired by the successful applications of k nearest neighbors in modeling genomics data, we propose a kNN-Vec2Text model to address these tasks and observe substantial improvement on our dataset.We further illustrate how Textomics can be used to advance other applications, including evaluating scientific paper embeddings and generating masked templates for scientific paper understanding.Textomics serves as the first benchmark for generating textual summaries for genomics data and we envision it will be broadly applied to other biomedical and natural language processing applications.1 Mu-Chun Wang, Zixuan Liu 0001, Sheng Wang 0012 |
ACL (1) | 3 |
| 2022 | Deep Graph Mutual Learning for Cross-domain Recommendation
Yifan Wang 0014, Weiping Song, Jiangke Fan, Sheng Wang 0012, Ming Zhang 0004 |
DASFAA (2) | 10 |
| 2022 | MetaFill: Text Infilling for Meta-Path Generation on Heterogeneous Information NetworksabstractHeterogeneous Information Network (HIN) is essential to study complicated networks containing multiple edge types and node types.Meta-path, a sequence of node types and edge types, is the core technique to embed HINs.Since manually curating meta-paths is timeconsuming, there is a pressing need to develop automated meta-path generation approaches.Existing meta-path generation approaches cannot fully exploit the rich textual information in HINs, such as node names and edge type names.To address this problem, we propose MetaFill, a text-infilling-based approach for meta-path generation.The key idea of MetaFill is to formulate meta-path identification problem as a word sequence infilling problem, which can be advanced by Pretrained Language Models (PLMs).We observed the superior performance of MetaFill against existing meta-path generation methods and graph embedding methods that do not leverage meta-paths in both link prediction and node classification on two real-world HIN datasets.We further demonstrated how MetaFill can accurately classify edges in the zero-shot setting, where existing approaches cannot generate any meta-paths.MetaFill exploits PLMs to generate meta-paths for graph embedding, opening up new avenues for language model applications in graph analysis. Zequn Liu, Kefei Duan, Ming Zhang 0004, Sheng Wang 0012 |
EMNLP | 6 |
| 2022 | HE-SNE: Heterogeneous Event Sequence-based Streaming Network Embedding for Dynamic BehaviorsabstractLarge amounts of user behavior data provide opportunities for user behavior modeling and have great potential in many downstream applications such as advertising and anomaly detection. Compared with traditional methods, embedding-based methods are used more often recently because of their efficiency and scalability. These methods build a “behavior-entity” bipartite graph and learn static embeddings for nodes in the graph. However, behavior patterns in the real world could not be static because entity properties such as user interests usually evolve along with time. In this paper, we formulate user behaviors as a temporal event sequence and propose a stream network embedding approach to capture the evolving nature of user behaviors. Representation of each event is built and used to update the embeddings of nodes. Two contextual behavior modeling tasks are studied for dynamic user behaviors, and experimental results with real-world data demonstrate the effectiveness of our proposed approach over several competitive baselines. Yifan Wang 0014, Jianhao Shen, Yiping Song, Sheng Wang 0012, Ming Zhang 0004 |
IJCNN | 4 |
| 2022 | Graph-in-Graph Network for Automatic Gene Ontology Description GenerationabstractGene Ontology (GO) is the primary gene function knowledge base that enables computational tasks in biomedicine. The basic element of GO is a term, which includes a set of genes with the same function. Existing research efforts of GO mainly focus on predicting gene term associations. Other tasks, such as generating descriptions of new terms, are rarely pursued. In this paper, we propose a novel task: GO term description generation. This task aims to automatically generate a sentence that describes the function of a GO term belonging to one of the three categories, i.e., molecular function, biological process, and cellular component. To address this task, we propose a Graph-in-Graph network that can efficiently leverage the structural information of GO. The proposed network introduces a two-layer graph: the first layer is a graph of GO terms where each node is also a graph (gene graph). Such a Graph-in-Graph network can derive the biological functions of GO terms and generate proper descriptions. To validate the effectiveness of the proposed network, we build three large-scale benchmark datasets. By incorporating the proposed Graph-in-Graph network, the performances of seven different sequence-to-sequence models can be substantially boosted across all evaluation metrics, with up to 34.7%, 14.5%, and 39.1% relative improvements in BLEU, ROUGE-L, and METEOR, respectively. Bang Yang, Chenyu You, Xian Wu 0001, Shen Ge, Adelaide Woicik, Sheng Wang 0012 |
KDD | 7 |
| 2022 | Brain-Aware Replacements for Supervised Contrastive Learning in Detection of Alzheimer's Disease
Mehmet Saygin Seyfioglu, Zixuan Liu 0001, Pranav Kamath, Sadjyot Gangolli, Sheng Wang 0012, Thomas J. Grabowski, Linda G. Shapiro |
MICCAI (1) | 5 |
| 2022 | Seed-Guided Topic Discovery with Out-of-Vocabulary SeedsabstractYu Zhang, Yu Meng, Xuan Wang, Sheng Wang, Jiawei Han. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yu Zhang 0044, Yu Meng 0001, Xuan Wang 0008, Sheng Wang 0012, Jiawei Han 0001 |
NAACL-HLT | 4 |
| 2022 | ProTranslator: Zero-Shot Protein Function Prediction Using Textual Description
Sheng Wang 0012 |
RECOMB | 2 |
| 2022 | scPretrain: multi-task self-supervised learning for cell-type classificationabstractMOTIVATION: Rapidly generated scRNA-seq datasets enable us to understand cellular differences and the function of each individual cell at single-cell resolution. Cell-type classification, which aims at characterizing and labeling groups of cells according to their gene expression, is one of the most important steps for single-cell analysis. To facilitate the manual curation process, supervised learning methods have been used to automatically classify cells. Most of the existing supervised learning approaches only utilize annotated cells in the training step while ignoring the more abundant unannotated cells. In this article, we proposed scPretrain, a multi-task self-supervised learning approach that jointly considers annotated and unannotated cells for cell-type classification. scPretrain consists of a pre-training step and a fine-tuning step. In the pre-training step, scPretrain uses a multi-task learning framework to train a feature extraction encoder based on each dataset's pseudo-labels, where only unannotated cells are used. In the fine-tuning step, scPretrain fine-tunes this feature extraction encoder using the limited annotated cells in a new dataset. RESULTS: We evaluated scPretrain on 60 diverse datasets from different technologies, species and organs, and obtained a significant improvement on both cell-type classification and cell clustering. Moreover, the representations obtained by scPretrain in the pre-training step also enhanced the performance of conventional classifiers, such as random forest, logistic regression and support-vector machines. scPretrain is able to effectively utilize the massive amount of unlabeled data and be applied to annotating increasingly generated scRNA-seq datasets. AVAILABILITY AND IMPLEMENTATION: The data and code underlying this article are available in scPretrain: Multi-task self-supervised learning for cell type classification, at https://github.com/ruiyi-zhang/scPretrain and https://zenodo.org/record/5802306. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yunan Luo, Jianzhu Ma, Ming Zhang 0004, Sheng Wang 0012 |
Bioinform. | 5 |
| 2021 | Graphine: A Dataset for Graph-aware Terminology Definition GenerationabstractPrecisely defining the terminology is the first step in scientific communication.Developing neural text generation models for definition generation can circumvent the laborintensity curation, further accelerating scientific discovery.Unfortunately, the lack of large-scale terminology definition dataset hinders the process toward definition generation.In this paper, we present a large-scale terminology definition dataset Graphine covering 2,010,648 terminology definition pairs, spanning 227 biomedical subdisciplines.Terminologies in each subdiscipline further form a directed acyclic graph, opening up new avenues for developing graph-aware text generation models.We then proposed a novel graphaware definition generation model Graphex that integrates transformer with graph neural network.Our model outperforms existing text generation models by exploiting the graph structure of terminologies.We further demonstrated how Graphine can be used to evaluate pretrained language models, compare graph representation learning methods and predict sentence granularity.We envision Graphine to be a unique resource for definition generation and many other NLP tasks in biomedicine. 1 Zequn Liu, Shukai Wang, Yiyang Gu, Ming Zhang 0004, Sheng Wang 0012 |
EMNLP (1) | 6 |
| 2021 | Auto-Encoding Knowledge Graph for Unsupervised Medical Report GenerationabstractMedical report generation, which aims to automatically generate a long and coherent report of a given medical image, has been receiving growing research interests. Existing approaches mainly adopt a supervised manner and heavily rely on coupled image-report pairs. However, in the medical domain, building a large-scale image-report paired dataset is both time-consuming and expensive. To relax the dependency on paired data, we propose an unsupervised model Knowledge Graph Auto-Encoder (KGAE) which accepts independent sets of images and reports in training. KGAE consists of a pre-constructed knowledge graph, a knowledge-driven encoder and a knowledge-driven decoder. The knowledge graph works as the shared latent space to bridge the visual and textual domains; The knowledge-driven encoder projects medical images and reports to the corresponding coordinates in this latent space and the knowledge-driven decoder generates a medical report given a coordinate in this space. Since the knowledge-driven encoder and decoder can be trained with independent sets of images and reports, KGAE is unsupervised. The experiments show that the unsupervised KGAE generates desirable medical reports without using any image-report training pairs. Moreover, KGAE can also work in both semi-supervised and supervised settings, and accept paired images and reports in training. By further fine-tuning with image-report pairs, KGAE consistently outperforms the current state-of-the-art models on two datasets. Chenyu You, Xian Wu 0001, Shen Ge, Sheng Wang 0012, Xu Sun 0001 |
NeurIPS | 5 |
| 2021 | Disease gene prediction with privileged information and heteroscedastic dropoutabstractMOTIVATION: Recently, machine learning models have achieved tremendous success in prioritizing candidate genes for genetic diseases. These models are able to accurately quantify the similarity among disease and genes based on the intuition that similar genes are more likely to be associated with similar diseases. However, the genetic features these methods rely on are often hard to collect due to high experimental cost and various other technical limitations. Existing solutions of this problem significantly increase the risk of overfitting and decrease the generalizability of the models. RESULTS: In this work, we propose a graph neural network (GNN) version of the Learning under Privileged Information paradigm to predict new disease gene associations. Unlike previous gene prioritization approaches, our model does not require the genetic features to be the same at training and test stages. If a genetic feature is hard to measure and therefore missing at the test stage, our model could still efficiently incorporate its information during the training process. To implement this, we develop a Heteroscedastic Gaussian Dropout algorithm, where the dropout probability of the GNN model is determined by another GNN model with a mirrored GNN architecture. To evaluate our method, we compared our method with four state-of-the-art methods on the Online Mendelian Inheritance in Man dataset to prioritize candidate disease genes. Extensive evaluations show that our model could improve the prediction accuracy when all the features are available compared to other methods. More importantly, our model could make very accurate predictions when >90% of the features are missing at the test stage. AVAILABILITY AND IMPLEMENTATION: Our method is realized with Python 3.7 and Pytorch 1.5.0 and method and data are freely available at: https://github.com/juanshu30/Disease-Gene-Prioritization-with-Privileged-Information-and-Heteroscedastic-Dropout. Juan Shu, Yu Li 0006, Sheng Wang 0012, Bowei Xi, Jianzhu Ma |
Bioinform. | 3 |
| 2020 | DisenHAN: Disentangled Heterogeneous Graph Attention Network for RecommendationabstractHeterogeneous information network has been widely used to alleviate sparsity and cold start problems in recommender systems since it can model rich context information in user-item interactions. Graph neural network is able to encode this rich context information through propagation on the graph. However, existing heterogeneous graph neural networks neglect entanglement of the latent factors stemming from different aspects. Moreover, meta paths in existing approaches are simplified as connecting paths or side information between node pairs, overlooking the rich semantic information in the paths. In this paper, we propose a novel disentangled heterogeneous graph attention network DisenHAN for top-N recommendation, which learns disentangled user/item representations from different aspects in a heterogeneous information network. In particular, we use meta relations to decompose high-order connectivity between node pairs and propose a disentangled embedding propagation layer which can iteratively identify the major aspect of meta relations. Our model aggregates corresponding aspect features from each meta relation for the target user/item. With different layers of embedding propagation, DisenHAN is able to explicitly capture the collaborative filtering effect semantically. Extensive experiments on three real-world datasets show that DisenHAN consistently outperforms state-of-the-art approaches. We further demonstrate the effectiveness and interpretability of the learned disentangled representations via insightful case studies and visualization. Yifan Wang 0014, Suyao Tang, Yuntong Lei, Weiping Song, Sheng Wang 0012, Ming Zhang 0004 |
CIKM | 5 |
| 2019 | GRep: Gene Set Representation via Gaussian Embedding
Sheng Wang 0012, Emily R. Flynn, Russ B. Altman |
RECOMB | 1 |
| 2017 | Framing Electronic Medical Records as Polylingual Documents in Query Expansion
Edward W. Huang, Sheng Wang 0012, Doris J. Lee, Runshun Zhang, Baoyan Liu, Xuezhong Zhou, ChengXiang Zhai |
AMIA | 2 |
| 2016 | A conditional probabilistic model for joint analysis of symptoms, diseases, and herbs in traditional Chinese medicine patient recordsabstractTraditional Chinese medicine (TCM) can provide important complementary medical care to modern medicine, and is widely practiced in China and many other countries. Unfortunately, due to its empirical nature and history of trial and error, effective diagnosis and prescription methods are not well-defined. This setback results in a significant challenge in retaining, sharing, and inheriting knowledge among physicians. In this paper, we propose a new asymmetric probabilistic model for the joint analysis of symptoms, diseases, and herbs in patient records to discover and extract latent TCM knowledge. We base our model on the comprehensive evaluation of modern medicine and TCM-specific symptoms in addition to herb prescriptions for particular diseases. Experimental results on a large dataset demonstrate the effectiveness of the proposed model for discovering useful knowledge and its potential clinical applications. Sheng Wang 0012, Edward W. Huang, Runshun Zhang, Baoyan Liu, Xuezhong Zhou, ChengXiang Zhai |
BIBM | 1 |
| 2016 | Early identification of adverse drug reactions from search log dataabstractThe timely and accurate identification of adverse drug reactions (ADRs) following drug approval is a persistent and serious public health challenge. Aggregated data drawn from anonymized logs of Web searchers has been shown to be a useful source of evidence for detecting ADRs. However, prior studies have been based on the analysis of established ADRs, the existence of which may already be known publically. Awareness of these ADRs can inject existing knowledge about the known ADRs into online content and online behavior, and thus raise questions about the ability of the behavioral log-based methods to detect new ADRs. In contrast to previous studies, we investigate the use of search logs for the early detection of known ADRs. We use a large set of recently labeled ADRs and negative controls to evaluate the ability of search logs to accurately detect ADRs in advance of their publication. We leverage the Internet Archive to estimate when evidence of an ADR first appeared in the public domain and adjust the index date in a backdated analysis. Our results demonstrate how search logs can be used to detect new ADRs, the central challenge in pharmacovigilance. Ryen W. White, Sheng Wang 0012, Apurv Pant, Rave Harpaz, Pushpraj Shukla, Walter Sun, William DuMouchel, Eric Horvitz |
J. Biomed. Informatics | 2 |
| 2014 | SUIT: A Supervised User-Item Based Topic Model for Sentiment AnalysisabstractProbabilistic topic models have been widely used for sentiment analysis. However, most of existing topic methods only model the sentiment text, but do not consider the user, who expresses the sentiment, and the item, which the sentiment is expressed on. Since different users may use different sentiment expressions for different items, we argue that it is better to incorporate the user and item information into the topic model for sentiment analysis. In this paper, we propose a new Supervised User-Item based Topic model, called SUIT model, for sentiment analysis. It can simultaneously utilize the textual topic and latent user-item factors. Our proposed method uses the tensor outer product of text topic proportion vector, user latent factor and item latent factor to model the sentiment label generalization. Extensive experiments are conducted on two datasets: review dataset and microblog dataset. The results demonstrate the advantages of our model. It shows significant improvement compared with supervised topic models and collaborative filtering methods. Fangtao Li, Sheng Wang 0012, Shenghua Liu, Ming Zhang 0004 |
AAAI | 2 |
| 2013 | Measuring Strength of Ties in Social Network
Dakui Sheng, Sheng Wang 0012, Ziqi Wang 0002, Ming Zhang 0004 |
APWeb | 3 |