Junmei Wang

dblp:99/948 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 7 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Utilizing large language models for integrating document-level contextual semantic into pseudo-relevance feedback
abstract
Pseudo-Relevance Feedback (PRF) is a key technique in information retrieval (IR). Traditional implementations rely on statistical information, such as term frequency, for precise matching and relevance assessment. However, these methods struggle to fully capture the deep semantic integrity of query terms, especially in handling polysemy, high semantic relevance, and long-document comprehension. To address these challenges, this paper innovatively proposes a large language model-assisted PRF probabilistic model. The model first employs a precise matching algorithm to evaluate and determine the term-level weights, and then uses a large language model to encode the contextual relationships within the query and feedback documents, thereby accurately acquiring the global semantic weights of terms relevant to the query at the document level. By adjusting a balancing factor to allocate weights between these two components, the model comprehensively selects expanded terms for constructing a new query representation and executing query expansion (QE). This model not only facilitates approximate matching through the integration of global semantic features of documents but also effectively combines with the precise matching information of traditional PRF models, enabling a comprehensive and accurate optimization of queries from a broader perspective. To validate effectiveness, extensive empirical analyses on five TREC datasets assess performance across key metrics such as MAP, P@10, NDCG, and MRR. Experimental results show significant improvements over baseline models. Comparative analyses and case studies confirm that the expanded terms maintain high semantic relevance and consistency with the original query while preserving diversity and effectively capturing global document semantics, establishing an efficient QE mechanism.
Min Pan, Wenrui Xiong, Junmei Wang, Feng Deng, Ellen Anne Huang, Jinguang Chen, Jimmy Huang 0001
Knowl. Based Syst.4
2025 Advancing promiscuous aggregating inhibitor analysis with intelligent machine learning classification
abstract
Small molecules have been playing a crucial role in drug discovery; however, some exhibit nonspecific inhibitory effects during hit screening due to the formation of colloidal aggregators. Such false positives often lead to significant research costs and time investment. Therefore, to identify potential aggregating compounds efficiently and accurately at an early stage of drug discovery, we employed several machine learning techniques to develop classification models for identifying promiscuous aggregating inhibitors. Using a training dataset of 10 000 aggregators and 10 000 nonaggregators, models were trained by combining four different molecular representations with various machine learning algorithms. We found that the best-performing model is the one that employs path-based FP2 fingerprints in conjunction with the cubic support vector machine algorithm, which achieved the highest accuracy and area under the receiver operating characteristic curve values for both the validation and test datasets while maintaining high sensitivity and specificity levels (>0.93). Additionally, we have proposed a new model interpretation method, global sensitivity analysis (GSA), to complement the well-recognized SHapley Additive exPlanations analysis. Several comparative studies have shown that GSA is a time-efficient and accurate approach for identifying crucial descriptors that contribute to model prediction, especially in the scenario where the dataset contains a substantial number of data entries with a limited set of descriptors. Our models as well as GSA findings can provide useful guidance on screening library design to minimize false positives.
Luxuan Wang, Beihong Ji, Jingchen Zhai, Junmei Wang
Briefings Bioinform.4
2025 NNKcat: deep neural network to predict catalytic constants (Kcat) by integrating protein sequence and substrate structure with enhanced data imbalance handling
abstract
Catalytic constant (Kcat) is to describe the efficiency of catalyzing reactions. The Kcat value of an enzyme-substrate pair indicates the rate an enzyme converts saturated substrates into product during the catalytic process. However, it is challenging to construct robust prediction models for this important property. Most of the existing models, including the one recently published by Nature Catalysis (Li et al.), are suffering from the overfitting issue. In this study, we proposed a novel protocol to construct Kcat prediction models, introducing an intermedia step to separately develop substrate and protein processors. The substrate processor leverages analyzing Simplified Molecular Input Line Entry System (SMILES) strings using a graph neural network model, attentive FP, while the protein processor abstracts protein sequence information utilizing long short-term memory architecture. This protocol not only mitigates the impact of data imbalance in the original dataset but also provides greater flexibility in customizing the general-purpose Kcat prediction model to enhance the prediction accuracy for specific enzyme classes. Our general-purpose Kcat prediction model demonstrates significantly enhanced stability and slightly better accuracy (R2 value of 0.54 versus 0.50) in comparison with Li et al.'s model using the same dataset. Additionally, our modeling protocol enables personalization of fine-tuning the general-purpose Kcat model for specific enzyme categories through focused learning. Using Cytochrome P450 (CYP450) enzymes as a case study, we achieved the best R2 value of 0.64 for the focused model. The high-quality performance and expandability of the model guarantee its broad applications in enzyme engineering and drug research & development.
Jingchen Zhai, Xiguang Qi, Lianjin Cai, Haocheng Tang, Junmei Wang
Briefings Bioinform.7
2025 A knowledge-based approach for pseudo-relevance feedback by exploiting semantic relevance
Junmei Wang, Jimmy Huang 0001, Luyun Wang, Jiajia Wang 0009
Knowl. Inf. Syst.1
2025 Utilizing Large Language Model for Conversational Information Seeking via Dual-Query Generation and Joint-Encoding
abstract
Conversational retrieval leverages multi-turn conversations to meet users’ information needs, and accurately understanding the new intent has become a significant challenge in this field. Recently, the language comprehension and reasoning capabilities of large language models (LLMs) offer a viable solution to these challenges. In this article, we propose a new Dual-Query Generation and Joint-Encoding method by utilizing LLM for Conversational Information Seeking, abbreviated as DQ-CIS. Specifically, we propose a dual-query generation approach that leverages both open source and closed source LLMs to generate two complementary queries: a full-rewrite query that preserves the context semantics of the conversation and a condensed-rewrite query that emphasizes the core intent of the current query. Additionally, to better express the semantic information of the query, we propose a dual-query joint-encoding method, which enhances the thematic expression of query vectors by treating the dual-query as semantic complementary. A query coverage fine-tuned semantic matching method is also introduced to improve result relevance and ranking by fine-tuning the original retrieval scores by ColBERT. We conducted a number of experiments on seven publicly available conversational retrieval datasets. The results show that compared with other models, DQ-CIS has strong competitiveness in both retrieval efficiency and retrieval results.
Junmei Wang, Fengjing Zhang, Xiadan Chen, Puyu He, Ellen Anne Huang, Jimmy Huang 0001
ACM Trans. Inf. Syst.1
2023 SPRF: A semantic Pseudo-relevance Feedback enhancement for information retrieval via ConceptNet
Min Pan, Quanli Pei, Teng Li 0012, Ellen Anne Huang, Junmei Wang, Jimmy Huang 0001
Knowl. Based Syst.6
2023 Model Checking of Possibilistic Linear-Time Properties Based on Generalized Possibilistic Decision Processes
abstract
Model checking possibilistic linear-time properties was investigated by Li in 2017. However, nondeterminism of the system is absent in previous studies. Therefore, in order to permit both possibilistic and nondeterministic choices, we use the generalized possibilistic decision process (GPDP) as a model of the system. First, the definition of GPDP describing the behavior of nondeterministic system is given in detail, the resolution of nondeterminism is performed by using the notion of schedulers, and the semantics of generalized possibilistic linear-temporal logic (GPoLTL) with schedulers are defined. Second, we study possibilistic model checking of some fuzzy linear-time properties under GPDP. Since there are many (infinite) schedulers satisfying a certain linear timing property in a given state of a GPDP, it is particularly critical to study the optimal strategy and its corresponding possible measure, which is called extremal possibility model checking. For some special fuzzy linear-time properties, such as constrained reachability, step-bounded constrained reachability, reachability, always reachability, repeated reachability, persistence reachability, we present complete solution to the optimal (including maximum and minimum cases) possibilistic model checking of the above reachability using the fixpoint techniques. We also introduce fuzzy$\omega$-regular properties in GPDP and show that their model checking can be simplified by repeated reachability. The algorithms for model checking are also provided. Additionally, an example is presented to illustrate the methods described in the article.
Yongming Li 0001, Wuniu Liu, Junmei Wang, Xianfeng Yu
IEEE Trans. Fuzzy Syst.3
2022 A probabilistic framework for integrating sentence-level semantics via BERT into pseudo-relevance feedback
Min Pan, Junmei Wang, Jimmy Huang 0001, Angela Jennifer Huang, Jinguang Chen
Inf. Process. Manag.2
2021 Machine learning on ligand-residue interaction profiles to significantly improve binding affinity prediction
abstract
Structure-based virtual screenings (SBVSs) play an important role in drug discovery projects. However, it is still a challenge to accurately predict the binding affinity of an arbitrary molecule binds to a drug target and prioritize top ligands from an SBVS. In this study, we developed a novel method, using ligand-residue interaction profiles (IPs) to construct machine learning (ML)-based prediction models, to significantly improve the screening performance in SBVSs. Such a kind of the prediction model is called an IP scoring function (IP-SF). We systematically investigated how to improve the performance of IP-SFs from many perspectives, including the sampling methods before interaction energy calculation and different ML algorithms. Using six drug targets with each having hundreds of known ligands, we conducted a critical evaluation on the developed IP-SFs. The IP-SFs employing a gradient boosting decision tree (GBDT) algorithm in conjunction with the MIN + GB simulation protocol achieved the best overall performance. Its scoring power, ranking power and screening power significantly outperformed the Glide SF. First, compared with Glide, the average values of mean absolute error and root mean square error of GBDT/MIN + GB decreased about 38 and 36%, respectively. Second, the mean values of squared correlation coefficient and predictive index increased about 225 and 73%, respectively. Third, more encouragingly, the average value of the areas under the curve of receiver operating characteristic for six targets by GBDT, 0.87, is significantly better than that by Glide, which is only 0.71. Thus, we expected IP-SFs to have broad and promising applications in SBVSs.
Beihong Ji, Xibing He, Jingchen Zhai, Viet Hoang Man, Junmei Wang
Briefings Bioinform.6
2021 Landscape of drug-resistance mutations in kinase regulatory hotspots
abstract
More than 48 kinase inhibitors (KIs) have been approved by Food and Drug Administration. However, drug-resistance (DR) eventually occurs, and secondary mutations have been found in the previously targeted primary-mutated cancer cells. Cancer and drug research communities recognize the importance of the kinase domain (KD) mutations for kinasopathies. So far, a systematic investigation of kinase mutations on DR hotspots has not been done yet. In this study, we systematically investigated four types of representative mutation hotspots (gatekeeper, G-loop, αC-helix and A-loop) associated with DR in 538 human protein kinases using large-scale cancer data sets (TCGA, ICGC, COSMIC and GDSC). Our results revealed 358 kinases harboring 3318 mutations that covered 702 drug resistance hotspot residues. Among them, 197 kinases had multiple genetic variants on each residue. We further computationally assessed and validated the epidermal growth factor receptor mutations on protein structure and drug-binding efficacy. This is the first study to provide a landscape view of DR-associated mutation hotspots in kinase's secondary structures, and its knowledge will help the development of effective next-generation KIs for better precision medicine.
Pora Kim, Junmei Wang, Zhongming Zhao
Briefings Bioinform.3
2021 In silico binding profile characterization of SARS-CoV-2 spike protein and its mutants bound to human ACE2 receptor
abstract
Severe acute respiratory syndrome coronavirus (SARS-CoV-2), a novel coronavirus, has brought an unprecedented pandemic to the world and affected over 64 million people. The virus infects human using its spike glycoprotein mediated by a crucial area, receptor-binding domain (RBD), to bind to the human ACE2 (hACE2) receptor. Mutations on RBD have been observed in different countries and classified into nine types: A435S, D364Y, G476S, N354D/D364Y, R408I, V341I, V367F, V483A and W436R. Employing molecular dynamics (MD) simulation, we investigated dynamics and structures of the complexes of the prototype and mutant types of SARS-CoV-2 spike RBDs and hACE2. We then probed binding free energies of the prototype and mutant types of RBD with hACE2 protein by using an end-point molecular mechanics Poisson Boltzmann surface area (MM-PBSA) method. According to the result of MM-PBSA binding free energy calculations, we found that V367F and N354D/D364Y mutant types showed enhanced binding affinities with hACE2 compared to the prototype. Our computational protocols were validated by the successful prediction of relative binding free energies between prototype and three mutants: N354D/D364Y, V367F and W436R. Thus, this study provides a reliable computational protocol to fast assess the existing and emerging RBD mutations. More importantly, the binding hotspots identified by using the molecular mechanics generalized Born surface area (MM-GBSA) free energy decomposition approach can guide the rational design of small molecule drugs or vaccines free of drug resistance, to interfere with or eradicate spike-hACE2 binding.
Xibing He, Jingchen Zhai, Beihong Ji, Viet Hoang Man, Junmei Wang
Briefings Bioinform.6
2020 A Pseudo-relevance feedback framework combining relevance matching and semantic matching for information retrieval
Junmei Wang, Min Pan, Tingting He 0003, Xinhui Tu
Inf. Process. Manag.1
2014 P-loop Conformation Governed Crizotinib Resistance in G2032R-Mutated ROS1 Tyrosine Kinase: Clues from Free Energy Landscape
abstract
Tyrosine kinases are regarded as excellent targets for chemical drug therapy of carcinomas. However, under strong purifying selection, drug resistance usually occurs in the cancer cells within a short term. Many cases of drug resistance have been found to be associated with secondary mutations in drug target, which lead to the attenuated drug-target interactions. For example, recently, an acquired secondary mutation, G2032R, has been detected in the drug target, ROS1 tyrosine kinase, from a crizotinib-resistant patient, who responded poorly to crizotinib within a very short therapeutic term. It was supposed that the mutation was located at the solvent front and might hinder the drug binding. However, a different fact could be uncovered by the simulations reported in this study. Here, free energy surfaces were characterized by the drug-target distance and the phosphate-binding loop (P-loop) conformational change of the crizotinib-ROS1 complex through advanced molecular dynamics techniques, and it was revealed that the more rigid P-loop region in the G2032R-mutated ROS1 was primarily responsible for the crizotinib resistance, which on one hand, impaired the binding of crizotinib directly, and on the other hand, shortened the residence time induced by the flattened free energy surface. Therefore, both of the binding affinity and the drug residence time should be emphasized in rational drug design to overcome the kinase resistance.
Huiyong Sun, Youyong Li, Junmei Wang, Tingjun Hou
PLoS Comput. Biol.4
2006 A Partition-Based Approach to Graph Mining
abstract
Existing graph mining algorithms typically assume that databases are relatively static and can fit into the main memory. Mining of subgraphs in a dynamic environment is currently beyond the scope of these algorithms. To bridge this gap, we first introduce a partition-based approach called PartMiner for mining graphs. The PartMiner algorithm finds the frequent subgraphs by dividing the database into smaller and more manageable units, mining frequent subgraphs on these smaller units and finally combining the results of these units to losslessly recover the complete set of subgraphs in the database. Next, we extend PartMiner to handle updates in the dynamic environment. Experimental results indicate that PartMiner is effective and scalable in finding frequent subgraphs, and outperforms existing algorithms in the presence of updates.
Junmei Wang, Wynne Hsu, Mong-Li Lee, Chang Sheng
ICDE1
2005 A framework for mining topological patterns in spatio-temporal databases
abstract
Mining topological patterns in spatial databases has received a lot of attention. However, existing work typically ignores the temporal aspect and suffers from certain efficiency problems. They are not scalable for mining topological patterns in spatio-temporal databases. In this paper, we study the problem for mining topological patterns by incorporating the temporal aspect in the mining process. We introduce a summary-structure that records the instances' count information of a feature in a region within a time window. Using this structure, we design an algorithm, TopologyMiner, to find interesting topological patterns without the need to generate candidates. Experimental results show that TopologyMiner is effective and scalable in finding topological patterns and outperforms Apriori-like algorithm by a few orders of magnitudes.
Junmei Wang, Wynne Hsu, Mong-Li Lee
CIKM1
2005 Mining Generalized Spatio-Temporal Patterns
Junmei Wang, Wynne Hsu, Mong-Li Lee
DASFAA1
2004 Discovering Geographical Features for Location-Based Services
Junmei Wang, Wynne Hsu, Mong-Li Lee
DASFAA1
2004 FlowMiner: Finding Flow Patterns in Spatio-Temporal Databases
abstract
The widespread use of spatio-temporal databases and applications has fuelled an urgent need to discover interesting time and space patterns in such databases. While much work has been done in discovering time/sequence patterns or spatial patterns, discovering of patterns involving both time and space dimensions is still in its infancy, We introduce the concept of flow patterns. Flow patterns are intended to describe the change of events over space and time. These flow patterns are useful to the understanding of many real-life applications. We present a disk-based algorithm, FlowMiner, which utilizes temporal relationships and spatial relationships amid events to generate flow patterns. Our performance study shows that FlowMiner is both scalable and efficient. Experiments on real-life datasets also reveal interesting flow patterns.
Junmei Wang, Wynne Hsu, Mong-Li Lee, Jason Tsong-Li Wang
ICTAI1