VLDB 2026 Research / reviewers in the wild / expert
Huimin Zhao 0003
dblp:91/6896-3
· DBLP profile ↗
29ranked-venue papers
10as first author
5since 2021 · last 2025
0000-0002-6471-9837ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 12 · 6 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Predicting reward-based crowdfunding success with multimodal data: A theory-guided framework
Liqian Bao, Zongxi Liu, Shuaiyong Xiao, Huimin Zhao 0003 |
Inf. Manag. | 5 |
| 2024 | A deep learning and clustering-based topic consistency modeling framework for matching health information supply and demandabstractAbstract Improving health literacy through health information dissemination is one of the most economical and effective mechanisms for improving population health. This process needs to fully accommodate the thematic suitability of health information supply and demand and reduce the impact of information overload and supply–demand mismatch on the enthusiasm of health information acquisition. We propose a health information topic modeling analysis framework that integrates deep learning methods and clustering techniques to model the supply‐side and demand‐side topics of health information and to quantify the thematic alignment of supply and demand. To validate the effectiveness of the framework, we have conducted an empirical analysis on a dataset with 90,418 pieces of textual data from two prominent social networking platforms. The results show that the supply of health information in general has not yet met the demand, the demand for health information has not yet been met to a considerable extent, especially for disease‐related topics, and there is clear inconsistency between the supply and demand sides for the same health topics. Public health policy‐making departments and content producers can adjust their information selection and dissemination strategies according to the distribution of identified health topics, thereby improving the effectiveness of public health information dissemination. Dongxiao Gu, Huimin Zhao 0003, Xuejie Yang, Min Li 0081, Changyong Liang |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2023 | Trust antecedents in online reviews across national cultures
Youngeui Kim, Mark Srite, Huimin Zhao 0003 |
Decis. Support Syst. | 3 |
| 2022 | My Real Avatar has a Doctor Appointment in the Wepital: A System for Persistent, Efficient, and Ubiquitous Medical Care
Huimin Zhao 0003, Patrick Sanvanson, Nitin Walia, Hemant K. Jain 0001, Reza Shaker |
Inf. Manag. | 2 |
| 2022 | Trust Decision-Making in Online Social Communities: A Network-Based ModelabstractThe unique characteristics of online social communities call for a reexamination and adaptation of established behavioral theories of trust decision-making. Guided by relevant social science and computational graph theories, we propose a conceptual model of trust decision-making in online social networks. This is the first study that integrates the existing graph-based view of trust decision-making in social networks into socio-psychological theories of trust to provide a richer understanding of trusting decisions in online social networks. We introduce new behavioral antecedents of trusting decisions, and redefine and integrate existing graph-based concepts to develop our proposed conceptual model. We introduce new behavioral antecedents of trusting decisions that have not been identified in previous research. We also identify novel operationalization methods to measure behavioral trust-inducing factors for online social networks. Our empirical findings indicate that both behavioral and network-specific trust decision-making factors should be considered in studying trusting decisions in online social networks. Gelareh Towhidi, Atish P. Sinha, Mark Srite, Huimin Zhao 0003 |
J. Comput. Inf. Syst. | 4 |
| 2020 | A case-based ensemble learning system for explainable breast cancer recurrence prediction
Dongxiao Gu, Kaixiang Su, Huimin Zhao 0003 |
Artif. Intell. Medicine | 3 |
| 2019 | Assessing product competitive advantages from the perspective of customers by mining user-generated content on social media
Cuiqing Jiang, Huimin Zhao 0003 |
Decis. Support Syst. | 3 |
| 2018 | Using contextual features and multi-view ensemble learning in product defect identification from online discussion forums
Cuiqing Jiang, Huimin Zhao 0003 |
Decis. Support Syst. | 3 |
| 2017 | A case-based reasoning system based on weighted heterogeneous value distance metric for breast cancer diagnosis
Dongxiao Gu, Changyong Liang, Huimin Zhao 0003 |
Artif. Intell. Medicine | 3 |
| 2017 | Adapting sentiment lexicons to domain-specific social media texts
Shuyuan Deng, Atish P. Sinha, Huimin Zhao 0003 |
Decis. Support Syst. | 3 |
| 2015 | Analyzing Positioning Strategies in Sponsored Search Auctions Under CTR-Based Quality ScoringabstractQuality score (QS) plays a critical role in sponsored search advertising (SSA) auctions, and in practice is closely correlated to the historical click-through rate (CTR) of an advertisement. The CTR-QS correlation may impose great influence on advertisers' positioning strategies of selecting the targeting slots in the sponsored list. In the literature, however, QS is implicitly assumed to be an independent variable and exogenously assigned by Web search engines, so that little theoretical or managerial insights can be offered to help understand the positioning dynamics in SSA auctions with CTR-QS correlation. We strive to bridge this research gap in this paper. Based on a discrete time-dependent optimal control model, which explicitly captures the relationship between the historical CTR and QS, we determine the optimal strategy for revenue-maximizing advertisers' QS-based positioning decisions through a policy-iteration-based numerical approximation method. We also investigate two practically-used heuristic strategies, namely the greedy and farsighted positioning strategies, aiming to examine and help understand advertisers' real-world positioning dynamics. Our analysis indicates that both the optimal and greedy positioning strategies lead advertisers to monotonically increase or decrease their targeting slots over time, which may cause a polarization trend emerging in SSA markets. Meanwhile, the farsighted positioning strategy can accelerate the polarization. Our simulations show that both the greedy and farsighted strategies have good revenue performance. Our findings indicate that advertisers should monotonically adjust their targeting positions to maximize their revenue in CTR-QS correlated SSA auctions. Our findings also highlight the need for Web search engine companies to set a lowered weight for historical CTRs or use position-normalized CTRs in their QS measurements, so as to suppress the polarization trend. Yong Yuan 0003, Daniel Dajun Zeng, Huimin Zhao 0003, Linjing Li |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2013 | Cross-Correlation Measure for Mining Spatio-Temporal PatternsabstractSpatio-temporal data mining is finding applications in many domains, such as public health, public safety, financial fraud detection, transportation, and product lifecycle management. Correlation analysis is an important spatio-temporal mining technique for unveiling spatial and temporal relationships among multiple event types. This paper presents a new measure for assessing and analyzing spatio-temporal cross-correlations. This measure extends Ripley’s a widely used measure of spatial correlation, with an additional temporal dimension. Empirical studies using real-world data show that the new measure can lead to a more discriminating and flexible spatio-temporal data analysis framework. In contrast with its predecessor, this measure also allows the discovery of leading (and potentially causal) event types whose occurrences precede those of other event types. Findings from analyses employing this measure may bear important managerial implications. James Ma, Daniel Dajun Zeng, Huimin Zhao 0003 |
J. Database Manag. | 3 |
| 2012 | Mining actionable behavioral rules
Peng Su 0001, Wenji Mao, Daniel Dajun Zeng, Huimin Zhao 0003 |
Decis. Support Syst. | 4 |
| 2011 | Mining actionable behavioral rules from group dataabstractMany security-related applications can benefit from constructing models to predict the behavior of an entity. However, such models do not provide the user with explicit knowledge that can be directly used to influence the behavior for his/her interest. This type of knowledge is called actionable knowledge. Actionability is a very important aspect of the interestingness of mined patterns. In this paper, we formally define a new problem of mining actionable behavioral rules from group data. We also propose an algorithm for solving the problem. Using terrorism group data, our experiment shows the validity of our approach as well as the practical value of our defined problem in security informatics. Peng Su 0001, Wenji Mao, Daniel Dajun Zeng, Huimin Zhao 0003 |
ISI | 4 |
| 2011 | An extended tuning method for cost-sensitive regression and forecasting
Huimin Zhao 0003, Atish P. Sinha, Gaurav Bansal |
Decis. Support Syst. | 1 |
| 2011 | Ontology for Developing Web Sites for Natural Disaster Management: Methodology and ImplementationabstractRecent natural disasters have highlighted the need for disaster preparedness, planning, and management. Hurricane Katrina demonstrated the usefulness of Web sites in dealing with natural disasters. However, little is known about the necessary contents and structures of Web-based information systems for natural disaster management. In this paper, we focus on developing an ontology structure of elements for Web-based disaster management systems. Web elements are identified, following a grounded-theory approach, from an inventory of 6032 Web pages drawn from 100 disaster management Web sites. Selected semi-structured data representation approaches are used to organize the resulting ontology structure, which consists of 2094 Web elements. The ontology structure is further coded into a Web-based system, allowing easy online access. Chen-Huei Chou, Huimin Zhao 0003 |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2010 | Collaborative filtering in social tagging systems based on joint item-tag recommendationsabstractTapping into the wisdom of the crowd, social tagging can be considered an alternative mechanism - as opposed to Web search - for organizing and discovering information on the Web. Effective tag-based recommendation of information items, such as Web resources, is a critical aspect of this social information discovery mechanism. A precise understanding of the information structure of social tagging systems lies at the core of an effective tag-based recommendation method. While most of the existing research either implicitly or explicitly assumes a simple tripartite graph structure for this purpose, we propose a comprehensive information structure to capture all types of co-occurrence information in the tagging data. Based on the proposed information structure, we further propose a unified user profiling scheme to make full use of all available information. Finally, supported by our proposed user profile, we propose a novel framework for collaborative filtering in social tagging systems. In our proposed framework, we first generate joint item-tag recommendations, with tags indicating topical interests of users in target items. These joint recommendations are then refined by the wisdom from the crowd and projected to the item space for final item recommendations. Evaluation using three real-world datasets shows that our proposed recommendation approach significantly outperformed state-of-the-art approaches. Jing Peng 0006, Daniel Dajun Zeng, Huimin Zhao 0003, Fei-Yue Wang 0001 |
CIKM | 3 |
| 2010 | Commercial Internet filters: Perils and opportunities
Chen-Huei Chou, Atish P. Sinha, Huimin Zhao 0003 |
Decis. Support Syst. | 3 |
| 2010 | Matching Attributes across Overlapping Heterogeneous Data Sources Using Mutual InformationabstractIdentifying matching attributes across heterogeneous data sources is a critical and time-consuming step in integrating the data sources. In this paper, the author proposes a method for matching the most frequently encountered types of attributes across overlapping heterogeneous data sources. The author uses mutual information as a unified measure of dependence on various types of attributes. An example is used to demonstrate the utility of the proposed method, which is useful in developing practical attribute matching tools. Huimin Zhao 0003 |
J. Database Manag. | 1 |
| 2009 | Effects of feature construction on classification performance: An empirical study in bank failure prediction
Huimin Zhao 0003, Atish P. Sinha |
Expert Syst. Appl. | 1 |
| 2008 | Entity matching across heterogeneous data sources: An approach based on constrained cascade generalization
Huimin Zhao 0003, Sudha Ram |
Data Knowl. Eng. | 1 |
| 2008 | Incorporating domain knowledge into data mining classifiers: An application in indirect lending
Atish P. Sinha, Huimin Zhao 0003 |
Decis. Support Syst. | 2 |
| 2007 | Combining schema and instance information for integrating heterogeneous data sources
Huimin Zhao 0003, Sudha Ram |
Data Knowl. Eng. | 1 |
| 2007 | A multi-objective genetic programming approach to developing Pareto optimal decision trees
Huimin Zhao 0003 |
Decis. Support Syst. | 1 |
| 2006 | Elitist and Ensemble Strategies for Cascade GeneralizationabstractSeveral methods have been proposed for cascading other classification algorithms with decision tree learners to alleviate the representational bias of decision trees and, potentially, to improve classification accuracy. Such cascade generalization of decision trees increases the flexibility of the decision boundaries between classes and promotes better fitting of the training data. However, more flexible models may not necessarily lead to more predictive power. Because of potential overfitting problems, the true classification accuracy on test data may not increase. Recently, a generic method for cascade generalization has been proposed. The method uses a parameter — the maximum cascading depth — to constrain the degree that other classification algorithms are cascaded with decision tree learners. A method for efficiently learning a collection (i.e., a forest) of generalized decision trees, each with other classification algorithms cascaded to a particular depth, also has been developed. In this article, we propose several new strategies, including elitist and ensemble (weighted or unweighted), for using the various decision trees in such a collection in the prediction phase. Our empirical evaluation using 32 data sets in the UCI machine learning repository shows that, on average, the elitist strategy outperforms the weighted full ensemble strategy, which, in turn, outperforms the unweighted full ensemble strategy. However, no strategy is universally superior across all applications. Since the same training process can be used to evaluate the various strategies, we recommend that several promising strategies be evaluated and compared before selecting the one to use for a given application. Huimin Zhao 0003, Atish P. Sinha, Sudha Ram |
J. Database Manag. | 1 |
| 2005 | Entity identification for heterogeneous database integration--a multiple classifier system approach and empirical evaluation
Huimin Zhao 0003, Sudha Ram |
Inf. Syst. | 1 |
| 2005 | An Efficient Algorithm for Generating Generalized Decision ForestsabstractA shortcoming of univariate decision tree learners is that they do not learn intermediate concepts and select only one of the input features in the branching decision at each intermediate tree node. It has been empirically demonstrated that cascading other classification methods, which learn intermediate concepts, with decision tree learners can alleviate such representational bias of decision trees and potentially improve classification performance. However, a more complex model that fits training data better may not necessarily perform better on unseen data, commonly referred to as the overfitting problem. To find the most appropriate degree of such cascade generalization, a decision forest (i.e., a set of decision trees with other classification models cascaded to different degrees) needs to be generated, from which the best decision tree can then be identified. In this paper, the authors propose an efficient algorithm for generating such decision forests. The algorithm uses an extended decision tree data structure and constructs any node that is common to multiple decision trees only once. The authors have empirically evaluated the algorithm using 32 data sets for classification problems from the University of California, Irvine (UCI) machine learning repository and report on results demonstrating the efficiency of the algorithm in this paper. Huimin Zhao 0003, Atish P. Sinha |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2004 | Constrained Cascade Generalization of Decision TreesabstractWhile decision tree techniques have been widely used in classification applications, a shortcoming of many decision tree inducers is that they do not learn intermediate concepts, i.e., at each node, only one of the original features is involved in the branching decision. Combining other classification methods, which learn intermediate concepts, with decision tree inducers can produce more flexible decision boundaries that separate different classes, potentially improving classification accuracy. We propose a generic algorithm for cascade generalization of decision tree inducers with the maximum cascading depth as a parameter to constrain the degree of cascading. Cascading methods proposed in the past, i.e., loose coupling and tight coupling, are strictly special cases of this new algorithm. We have empirically evaluated the proposed algorithm using logistic regression and C4.5 as base inducers on 32 UCI data sets and found that neither loose coupling nor tight coupling is always the best cascading strategy and that the maximum cascading depth in the proposed algorithm can be tuned for better classification accuracy. We have also empirically compared the proposed algorithm and ensemble methods such as bagging and boosting and found that the proposed algorithm performs marginally better than bagging and boosting on the average. Huimin Zhao 0003, Sudha Ram |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2003 | A Spanning Tree Based Approach to Identifying Web Services
Hemant K. Jain 0001, Huimin Zhao 0003, Nageswara Rao Chinta |
ICWS | 2 |