VLDB 2026 Research / reviewers in the wild / expert
Zhanhuai Li
dblp:05/299
· DBLP profile ↗
84ranked-venue papers
0as first author
17since 2021 · last 2024
0009-0003-6936-5745ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 39 · 6 since 2021Artificial intelligence and machine learning · 23 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 1 since 2021Systems, architecture and hardware · 8 · 1 since 2021Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Exploratory Query Optimization for Template-Based SQL WorkloadsabstractSQL query optimization aims to choose an optimal Query Execution Plan (QEP) for a query. The existing optimizers usually choose the plan with the minimal execution cost. However, in some real scenarios (e.g., cloud OLAP), it is a business imperative to minimize the cost of an entire workload. Unfortunately, the greedy optimizers are prone to generating execution plans that are only suboptimal from the workload perspective. In this paper, we propose a novel QEP optimization approach for template-based SQL that aims to minimize the execution cost of an entire workload instead of individual queries. While choosing an execution plan for a query, the proposed approach considers a plan's impact on future query execution as well as its own execution cost. We first define an exploratory metric, Plan Exploration Value (PEV), to measure a plan's potential benefit to future query execution. Then, we model exploratory plan selection as an optimization problem with cost constraints, and present an efficient algorithm based on dynamic programming for its solution. Finally, we evaluate the performance of the proposed approach by a comparative study on benchmark workloads. Our extensive experiments have shown that it generates more efficient workload execution plans than the existing alternatives. Jieming Feng, Zhanhuai Li, Qun Chen 0001 |
ICDE | 2 |
| 2024 | Adaptive deep learning for entity disambiguation via knowledge-based risk analysis
Youcef Nafa, Qun Chen 0001, Boyi Hou, Zhanhuai Li |
Expert Syst. Appl. | 4 |
| 2024 | Automating localized learning for cardinality estimation based on XGBoost
Jieming Feng, Zhanhuai Li, Qun Chen 0001, Hailong Liu 0004 |
Knowl. Inf. Syst. | 2 |
| 2024 | Enhancing Enterprise Credit Risk Assessment with Cascaded Multi-level Graph Representation Learning
Lingyun Song, Yacong Tan, Zhanhuai Li, Xuequn Shang 0001 |
Neural Networks | 4 |
| 2024 | A Multi-Group Multi-Stream attribute Attention network for fine-grained zero-shot learning
Lingyun Song, Xuequn Shang 0001, Ruizhi Zhou, Jun Liu 0002, Jie Ma 0001, Zhanhuai Li, Mingxuan Sun 0001 |
Neural Networks | 6 |
| 2024 | Thorough Data Pruning for Join Query in Database SystemabstractThe improvement of robustness and efficiency for multi-way equijoin query is challenging, no-matter for centralized database systems or distributed database systems. Due to lots of unnecessary data existing during query processing, these two metrics will be seriously reduced. If we can thoroughly prune unnecessary data in advance, the robustness and efficiency will be highly improved. However, the pruning power of current strategies, such as predicate push-down and algebraic equivalence, is limited. We present deepDP, a powerful, generalized, and efficient strategy for data pruning. deepDP builds multiple independent pruning spaces by generating longest transitive closures and applies appropriate data pruning strategy for each pruning space. For thoroughly pruning unnecessary data, deepDP employs$\alpha \cdot \beta$pruning strategy to clean each pruning space based on a newly designed statistic information-Hollow Range and re-shuffles the elements in all pruned spaces for maximizing robustness and efficiency benefits meanwhile minimizing the invasion. We implement deepDP in PostgreSQL but are not limited to it, and evaluate deepDP on TPC-H, JOB, and our synthesis benchmark–DHR. The experiment results show that compared to traditional data pruning strategy, deepDP can improve multi-way equijoin query on efficiency by 3.5x. Jin-Tao Gao, Zhanhuai Li, Sun Jian |
IEEE Trans. Sustain. Comput. | 2 |
| 2023 | Adaptive deep learning for entity resolution by risk analysis
Qun Chen 0001, Zhaoqiang Chen, Youcef Nafa, Tianyi Duan, Wei Pan 0007, Lijun Zhang 0003, Zhanhuai Li |
Knowl. Based Syst. | 7 |
| 2023 | Supervised Gradual Machine Learning for Aspect-Term Sentiment AnalysisabstractAbstract Recent work has shown that Aspect-Term Sentiment Analysis (ATSA) can be effectively performed by Gradual Machine Learning (GML). However, the performance of the current unsupervised solution is limited by inaccurate and insufficient knowledge conveyance. In this paper, we propose a supervised GML approach for ATSA, which can effectively exploit labeled training data to improve knowledge conveyance. It leverages binary polarity relations between instances, which can be either similar or opposite, to enable supervised knowledge conveyance. Besides the explicit polarity relations indicated by discourse structures, it also separately supervises a polarity classification DNN and a binary Siamese network to extract implicit polarity relations. The proposed approach fulfills knowledge conveyance by modeling detected relations as binary features in a factor graph. Our extensive experiments on real benchmark data show that it achieves the state-of-the-art performance across all the test workloads. Our work demonstrates clearly that, in collaboration with DNN for feature extraction, GML outperforms pure DNN solutions. Yanyan Wang 0005, Qun Chen 0001, Murtadha H. M. Ahmed, Zhaoqiang Chen, Wei Pan 0007, Zhanhuai Li |
Trans. Assoc. Comput. Linguistics | 7 |
| 2022 | A deep grouping fusion neural network for multimedia content understandingabstractAbstract How Deep Neural Networks (DNNs) best cope with the understanding of multimedia contents still remains an open problem, mainly due to two factors. First, conventional DNNs cannot effectively learn the representations of the images with sparse visual information. For example, the images describing knowledge concepts in textbooks. Second, existing DNNs cannot effectively capture the fine‐grained interactions between the images and text descriptions. To address these issues, we propose a deep Cross‐Media Grouping Fusion Network (CMGFN), which mainly has two distinctive properties: 1) CMGFN can effectively learn visual features from the images with sparse visual information. This is achieved by first progressively adjusting the attention of convolution filters to valuable visual regions, and then enhancing the use of key visual information in feature construction. 2) By a cross‐media grouping co‐attention mechanism, CMGFN can effectively use the interactions between visual features of different semantics and textual descriptions, to learn cross‐media features representing different fine‐grained semantics in different groups. Empirical studies demonstrate that CMGFN not only achieves state‐of‐the‐art performance on the multimedia documents containing sparse visual information, but also shows superior general applicability on other multimedia data, e.g., the multimedia fake news. Lingyun Song, Mengzhen Yu, Xuequn Shang 0001, Yu Lu 0003, Jun Liu 0002, Zhanhuai Li |
IET Image Process. | 7 |
| 2022 | Adaptive deep learning for network intrusion detection by risk analysis
Lijun Zhang 0003, Zhaoqiang Chen, Tianwei Liu, Qun Chen 0001, Zhanhuai Li |
Neurocomputing | 6 |
| 2022 | Active deep learning on entity resolution by risk sampling
Youcef Nafa, Qun Chen 0001, Zhaoqiang Chen, Tianyi Duan, Zhanhuai Li |
Knowl. Based Syst. | 7 |
| 2022 | Gradual Machine Learning for Entity ResolutionabstractUsually considered as a classification problem, entity resolution (ER) can be very challenging on real data due to the prevalence of dirty values. The state-of-the-art solutions for ER were built on a variety of learning models (most notably deep neural networks), which require lots of accurately labeled training data. Unfortunately, high-quality labeled data usually require expensive manual work, and are therefore not readily available in many real scenarios. In this paper, we propose a novel learning paradigm for ER, calledgradual machine learning, which aims to enable effective machine labeling without the requirement for manual labeling effort. It begins with some easy instances in a task, which can be automatically labeled by the machine with high accuracy, and then gradually labels more challenging instances by iterative factor graph inference. In gradual machine learning, the hard instances in a task are gradually labeled in small stages based on the estimated evidential certainty provided by the labeled easier instances. Our extensive experiments on real data have shown that the performance of the proposed approach is considerably better than its unsupervised alternatives, and highly competitive compared to the state-of-the-art supervised techniques. Using ER as a test case, we demonstrate that gradual machine learning is a promising paradigm potentially applicable to other challenging classification tasks requiring extensive labeling effort. Boyi Hou, Qun Chen 0001, Yanyan Wang 0005, Youcef Nafa, Zhanhuai Li |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2021 | A Semi-structured Data Classification Model with Integrating Tag Sequence and Ngram
Lijun Zhang 0003, Ning Li 0022, Wei Pan 0007, Zhanhuai Li |
DASFAA (2) | 4 |
| 2021 | An Overview on Supervised Semi-structured Data ClassificationabstractMany collaboratively building resources, such as Wikipedia, Weibo and Quora, exist in the form of semi-structured data. The semi-structured data has been widely used in areas such as data integration, data distribution, data storage, data management, information retrieval and knowledge management. For large volumes of semi-structured data on the Web, semi-structured data classification technique can group them into different categories by their structure and/or content information. Supervised semi-structured data classification plays an important role in many applications. This paper provides an overview of the literature in the area of supervised semi-structured data classification. A general framework for semi-structured data classification is presented, which is mainly composed of two steps: feature extraction and model building. Several different representation models of semi-structured data are discussed, mainly including rooted labeled tree model, feature vector space model and feature set model. A large selection of semi-structured data classification approaches are reviewed in detail from two aspects: based on structure only and based on both structure and content. Finally, several future research directions for semistructured data classification are presented. Lijun Zhang 0003, Ning Li 0022, Zhanhuai Li |
DSAA | 3 |
| 2021 | WOBTree: a write-optimized B+-tree for non-volatile memory
Zhanhuai Li, Xiao Zhang 0014, Xiaonan Zhao, Song Jiang 0001 |
Frontiers Comput. Sci. | 2 |
| 2021 | Aspect-level sentiment analysis based on gradual machine learning
Yanyan Wang 0005, Qun Chen 0001, Jiquan Shen, Boyi Hou, Murtadha H. M. Ahmed, Zhanhuai Li |
Knowl. Based Syst. | 6 |
| 2021 | Joint Inference for Aspect-Level Sentiment Analysis by Deep Neural Networks and Linguistic HintsabstractThe state-of-the-art techniques for aspect-level sentiment analysis focused on feature modeling using a variety of deep neural networks (DNN). Unfortunately, their performance may still fall short of expectation in real scenarios due to the semantic complexity of natural languages. Motivated by the observation that many linguistic hints (e.g., sentiment words and shift words) are reliable polarity indicators, we propose a joint framework, SenHint, which can seamlessly integrate the output of deep neural networks and the implications of linguistic hints in a unified model based on Markov logic network (MLN). SenHint leverages the linguistic hints for multiple purposes: (1) to identify the easy instances, whose polarities can be automatically determined by the machine with high accuracy; (2) to capture the influence of sentiment words on aspect polarities; (2) to capture the implicit relations between aspect polarities. We present the required techniques for extracting linguistic hints, encoding their implications as well as the output of DNN into the unified model, and joint inference. Finally, we have empirically evaluated the performance of SenHint on both English and Chinese benchmark datasets. Our extensive experiments have shown that compared to the state-of-the-art DNN techniques, SenHint can effectively improve polarity detection accuracy by considerable margins. Yanyan Wang 0005, Qun Chen 0001, Murtadha H. M. Ahmed, Zhanhuai Li, Wei Pan 0007, Hailong Liu 0004 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Towards Interpretable and Learnable Risk Analysis for Entity ResolutionabstractMachine-learning-based entity resolution has been widely studied. However, some entity pairs may be mislabeled by machine learning models and existing studies do not study the risk analysis problem -- predicting and interpreting which entity pairs are mislabeled. In this paper, we propose an interpretable and learnable framework for risk analysis, which aims to rank the labeled pairs based on their risks of being mislabeled. We first describe how to automatically generate interpretable risk features, and then present a learnable risk model and its training technique. Finally, we empirically evaluate the performance of the proposed approach on real data. Our extensive experiments have shown that the learning risk model can identify the mislabeled pairs with considerably higher accuracy than the existing alternatives. Zhaoqiang Chen, Qun Chen 0001, Boyi Hou, Zhanhuai Li, Guoliang Li 0001 |
SIGMOD Conference | 4 |
| 2020 | An adaptive strategy for statistics collecting in distributed database
Jin-Tao Gao, Wenjie Liu 0006, Zhanhuai Li |
Frontiers Comput. Sci. | 3 |
| 2020 | A new fragments allocating method for join query in distributed database
Jin-Tao Gao, Zhanhuai Li, Wenjie Liu 0006, Yantao Yue |
Frontiers Comput. Sci. | 2 |
| 2020 | A general fragments allocation method for join query in distributed database
Jin-Tao Gao, Wenjie Liu 0006, Zhanhuai Li |
Inf. Sci. | 3 |
| 2020 | Constructing domain-dependent sentiment dictionary for sentiment analysis
Murtadha H. M. Ahmed, Qun Chen 0001, Zhanhuai Li |
Neural Comput. Appl. | 3 |
| 2020 | r-HUMO: A Risk-Aware Human-Machine Cooperation Framework for Entity Resolution with Quality GuaranteesabstractEven though many approaches have been proposed for entity resolution (ER), it remains very challenging to enforce quality guarantees. To this end, we propose a risk-aware HUman-Machine cOoperation framework for ER, denoted by r-HUMO. Built on the existing HUMO framework, r-HUMO similarly enforces both precision and recall guarantees by partitioning an ER workload between the human and the machine. However, r-HUMO is the first solution that optimizes the process of human workload selection from a risk perspective. It iteratively selects human workload by real-time risk analysis based on the human-labeled results as well as the prespecified machine metric. In this paper, we first introduce the r-HUMO framework and then present the risk model to prioritize the instances for manual inspection. Finally, we empirically evaluate r-HUMO's performance on real data. Our extensive experiments show that r-HUMO is effective in enforcing quality guarantees, and compared with the state-of-the-art alternatives, it can achieve desired quality control with reduced human cost. Boyi Hou, Qun Chen 0001, Zhaoqiang Chen, Youcef Nafa, Zhanhuai Li |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | Hint-Embedding Attention-Based LSTM for Aspect Identification Sentiment Analysis
Murtadha H. M. Ahmed, Qun Chen 0001, Yanyan Wang 0005, Zhanhuai Li |
PRICAI (2) | 4 |
| 2019 | Gradual Machine Learning for Entity ResolutionabstractUsually considered as a classification problem, entity resolution can be very challenging on real data due to the prevalence of dirty values. The state-of-the-art solutions for ER were built on a variety of learning models (most notably deep neural networks), which require lots of accurately labeled training data. Unfortunately, high-quality labeled data usually require expensive manual work, and are therefore not readily available in many real scenarios. In this demo, we propose a novel learning paradigm for ER, called gradual machine learning, which aims to enable effective machine labeling without the requirement for manual labeling effort. It begins with some easy instances in a task, which can be automatically labeled by the machine with high accuracy, and then gradually labels more challenging instances based on iterative factor graph inference. In gradual machine learning, the hard instances in a task are gradually labeled in small stages based on the estimated evidential certainty provided by the labeled easier instances. Our extensive experiments on real data have shown that the proposed approach performs considerably better than its unsupervised alternatives, and its performance is also highly competitive compared to the state-of-the-art supervised techniques. Using ER as a test case, we demonstrate that gradual machine learning is a promising paradigm potentially applicable to other challenging classification tasks requiring extensive labeling effort. Video: https://youtu.be/99bA9aamsgk Boyi Hou, Qun Chen 0001, Jiquan Shen, Ping Zhong 0004, Yanyan Wang 0005, Zhaoqiang Chen, Zhanhuai Li |
WWW | 8 |
| 2019 | Statistical relational learning based automatic data cleaning
Weibang Li, Zhanhuai Li, Mengtian Cui |
Frontiers Comput. Sci. | 3 |
| 2019 | An efficient parallel algorithm of N-hop neighborhoods on graphs in distributed environment
Wenjie Liu 0006, Zhanhuai Li |
Frontiers Comput. Sci. | 2 |
| 2019 | Reducing partition skew on MapReduce: an incremental allocation approach
Zhuo Wang 0002, Qun Chen 0001, Bo Suo, Wei Pan 0007, Zhanhuai Li |
Frontiers Comput. Sci. | 5 |
| 2019 | Identification of cancer subtypes by integrating multiple types of transcriptomics data with deep learning in breast cancer
Xuequn Shang 0001, Zhanhuai Li |
Neurocomputing | 3 |
| 2018 | GraphU: A Unified Vertex-Centric Parallel Graph Processing PlatformabstractMany synchronous and asynchronous distributed platforms based on the Bulk Synchronous Parallel (BSP) model have been built for large-scale vertex -centric graph processing. Unfortunately, a program designed for a synchronous platform may not work properly on an asynchronous one. As a result, given the same problem, end users may be required to design different parallel algorithms for different platforms. Recently, we have proposed a unified programming model, DFA-G (Deterministic Finite Automaton for Graph processing), which expresses the computation at a vertex as a series of message-driven state transitions. It has the attractive property that any program modeled after it can run properly across synchronous and asynchronous platforms. In this demo, we first propose a framework of complexity analysis for DFA-G automaton and show that it can significantly facilitate complexity analysis on asynchronous programs. Due to the existing BSP platforms' deficiency in supporting efficient DFA-G execution, we then develop a new prototype platform, GraphU. GraphU was built on the popular open-source Giraph project. But it entirely removes synchronization barriers and decouples remote communication from vertex computation. Finally, we empirically evaluate the performance of various DFA-G programs on GraphU by a comparative study. Our experiments validate the efficacy of the proposed complexity analysis approach and the efficiency of GraphU. Qun Chen 0001, Zhuo Wang 0002, Murtadha H. M. Ahmed, Zhanhuai Li |
ICDCS | 5 |
| 2018 | Enabling Quality Control for Entity Resolution: A Human and Machine Cooperation FrameworkabstractEven though many machine algorithms have been proposed for entity resolution, it remains very challenging to find a solution with quality guarantees. In this paper, we propose a novel HUman and Machine cOoperation (HUMO) framework for entity resolution (ER), which divides an ER workload between the machine and the human. HUMO enables a mechanism for quality control that can flexibly enforce both precision and recall levels. We introduce the optimization problem of HUMO, minimizing human cost given a quality requirement, and then present three optimization approaches: a conservative baseline one purely based on the monotonicity assumption of precision, a more aggressive one based on sampling and a hybrid one that can take advantage of the strengths of both previous approaches. Finally, we demonstrate by extensive experiments on real and synthetic datasets that HUMO can achieve high-quality results with reasonable return on investment (ROI) in terms of human cost, and it performs considerably better than the state-of-the-art alternatives in quality control. Zhaoqiang Chen, Qun Chen 0001, Fengfeng Fan, Yanyan Wang 0005, Zhuo Wang 0002, Youcef Nafa, Zhanhuai Li, Hailong Liu 0004, Wei Pan 0007 |
ICDE | 7 |
| 2018 | OC-Cache: An Open-channel SSD Based Cache for Multi-Tenant SystemsabstractIn a multi-tenant cloud environment, tenants are usually hosted by virtual machines. Cloud providers deploy multiple virtual machines on a physical server to better utilize physical resources including CPU, memory, and storage devices. SSDs are often used as an I/O cache shared among the tenants for large storage systems using hard disk drives (HDDs) as their main storage devices, which can receive much of SSD's performance benefit and HDD's cost advantage. A key challenge in the use of the shared cache is to ensure strong performance isolation and maintain its high utilization at the same time. However, conventional SSD cache management approaches cannot effectively address this challenge. In this paper, we propose OC-Cache, an open-channel SSD cache framework which utilizes SSD'd internal parallelism to adaptively allocate cache to tenants for both good performance isolation and high SSD utilization. In particular, OC-Cache uses a tenant's miss ratio curve to determine the amount of cache space allocation and where the allocation is (in dedicated or shared SSD channels) and dynamically manages cache space according to the workload characteristics. Experiments show that OC-Cache significantly reduces interference among tenants, and maintains high utilization of the SSD cache. Zhanhuai Li, Xiao Zhang 0014, Xiaonan Zhao, Xingsheng Zhao, Song Jiang 0001 |
IPCCC | 2 |
| 2018 | FSObserver: A Performance Measurement and Monitoring Tool for Distributed Storage Systems
Xiao Zhang 0014, Lanxin Kong, Shunyi Zhu, Zhanhuai Li, Xiaonan Zhao |
NPC | 4 |
| 2018 | BCDForest: a boosting cascade deep forest model towards the classification of cancer subtypes based on gene expression dataabstractBACKGROUND: The classification of cancer subtypes is of great importance to cancer disease diagnosis and therapy. Many supervised learning approaches have been applied to cancer subtype classification in the past few years, especially of deep learning based approaches. Recently, the deep forest model has been proposed as an alternative of deep neural networks to learn hyper-representations by using cascade ensemble decision trees. It has been proved that the deep forest model has competitive or even better performance than deep neural networks in some extent. However, the standard deep forest model may face overfitting and ensemble diversity challenges when dealing with small sample size and high-dimensional biology data. RESULTS: In this paper, we propose a deep learning model, so-called BCDForest, to address cancer subtype classification on small-scale biology datasets, which can be viewed as a modification of the standard deep forest model. The BCDForest distinguishes from the standard deep forest model with the following two main contributions: First, a named multi-class-grained scanning method is proposed to train multiple binary classifiers to encourage diversity of ensemble. Meanwhile, the fitting quality of each classifier is considered in representation learning. Second, we propose a boosting strategy to emphasize more important features in cascade forests, thus to propagate the benefits of discriminative features among cascade layers to improve the classification performance. Systematic comparison experiments on both microarray and RNA-Seq gene expression datasets demonstrate that our method consistently outperforms the state-of-the-art methods in application of cancer subtype classification. CONCLUSIONS: The multi-class-grained scanning and boosting strategy in our model provide an effective solution to ease the overfitting challenge and improve the robustness of deep forest model working on small-scale data. Our model provides a useful approach to the classification of cancer subtypes by using deep learning on high-dimensional and small-scale biology data. Shuhui Liu, Zhanhuai Li, Xuequn Shang 0001 |
BMC Bioinform. | 3 |
| 2018 | The New Hardware Development Trend and the Challenges in Data Management and AnalysisabstractHardware techniques and environments underwent significant transformations in the field of information technology, represented by high-performance processors and hardware accelerators characterized by abundant heterogeneous parallelism, nonvolatile memory with hybrid storage hierarchies, and RDMA-enabled high-speed network. Recent hardware trends in these areas deeply affect data management and analysis applications. In this paper, we first introduce the development trend of the new hardware in computation, storage, and network dimensions. Then, the related research techniques which affect the upper data management system design are reviewed. Finally, challenges and opportunities are addressed for the key technologies of data management and analysis in new hardware environments. Wei Pan 0007, Zhanhuai Li, Chuliang Weng |
Data Sci. Eng. | 2 |
| 2018 | Automatic Web-based relational data imputation
Hailong Liu 0004, Zhanhuai Li, Qun Chen 0001, Zhaoqiang Chen |
Frontiers Comput. Sci. | 2 |
| 2018 | GL-RF: a reconciliation framework for label-free entity resolution
Yaoli Xu, Zhanhuai Li, Qun Chen 0001, Fengfeng Fan |
Frontiers Comput. Sci. | 2 |
| 2018 | Reasoning about attribute value equivalence in relational data
Fengfeng Fan, Zhanhuai Li, Qun Chen 0001, Lei Chen 0002 |
Inf. Syst. | 2 |
| 2018 | Relational data imputation with quality guarantee
Fengfeng Fan, Zhanhuai Li, Qun Chen 0001, Lei Chen 0002 |
Inf. Sci. | 2 |
| 2018 | An efficient theta-join query processing in distributed environment
Wenjie Liu 0006, Zhanhuai Li |
J. Parallel Distributed Comput. | 2 |
| 2017 | Towards the classification of cancer subtypes by using cascade deep forest model in gene expression dataabstractThe classification of cancer subtypes is of great importance in cancer disease diagnosis and therapy. Many supervised learning methods have been applied to classification of cancer subtypes in the past few years, especially of deep learning based methods. Recently, a deep forest model has been proposed as an alternative of deep neural networks to learn hyper-representations by using cascade ensemble decision trees, and it has been proved that deep forest model has competitive or even better performance than deep neural networks. However, the original deep forest may face under-fitting and ensemble diversity problems when dealing with small sample size, and high-dimension biology data. It is important to improve the deep forest model to work better on small-scale biology data. In this paper, we propose a deep learning model to follow the mission of cancer subtype classification on small-scale biology data sets, which can be viewed as modification of original deep forest model. Our model distinguishes from the original deep forest model with two main contributions: First, a named multi-class-scanning method is proposed to train multiple simple binary classifiers to encourage diversity of ensemble. Meanwhile, the fitting quality of each classifier is considered in representations learning. Second, we propose a boosting strategy to emphasize more important features in cascade forests of representations learning, thus to propagate the benefits of discriminative features among layers to improve the overall classification performance. Systematical experiments on both microarray and RNA-seq data sets demonstrate that our method consistently outperforms the most state-of-the-art classification methods in application of cancer subtype classifications. Shuhui Liu, Zhanhuai Li, Xuequn Shang 0001 |
BIBM | 3 |
| 2017 | POOLSIDE: An Online Probabilistic Knowledge Base for Shopping Decision SupportabstractWe present POOLSIDE, an online PrObabilistic knOwLedge base for ShoppIng DEcision support, that provides with the on-target recommendation service based on explicit user requirement. With a natural language interface, POOLSIDE can answer question in real-time. We present how to construct the knowledge base and how to enable real-time response in POOLSIDE. Finally, we demonstrate that Poolside can give high-quality product recommendations with high efficiency.(The demo video can be accessed via the link:https://www.youtube.com/watch?v=D8ALi11CUcc) Ping Zhong 0004, Zhanhuai Li, Qun Chen 0001, Yanyan Wang 0005, Lianping Wang, Murtadha H. M. Ahmed, Fengfeng Fan |
CIKM | 2 |
| 2017 | A Human-and-Machine Cooperative Framework for Entity Resolution with Quality GuaranteesabstractFor entity resolution, it remains very challenging to find the solution with quality guarantees as measured by both precision and recall. In this demo, we propose a HUman-and-Machine cOoperative framework, denoted by HUMO, for entity resolution. Compared with the existing approaches, HUMO enables a flexible mechanism for quality control that can enforce both precision and recall levels. We also introduce the problem of minimizing human cost given a quality requirement and present corresponding optimization techniques. Finally, we demo that HUMO achieves high-quality results with reasonable return on investment (ROI) in terms of human cost on real datasets. Zhaoqiang Chen, Qun Chen 0001, Zhanhuai Li |
ICDE | 3 |
| 2017 | Parallelizing maximal clique and k-plex enumeration over graph data
Zhuo Wang 0002, Qun Chen 0001, Boyi Hou, Bo Suo, Zhanhuai Li, Wei Pan 0007, Zachary G. Ives |
J. Parallel Distributed Comput. | 5 |
| 2016 | Discovering Approximate Functional Dependencies from Distributed Big Data
Weibang Li, Zhanhuai Li, Qun Chen 0001, Tao Jiang 0030, Zhilei Yin |
APWeb (2) | 2 |
| 2016 | Parallelizing Maximal Clique Enumeration Over Graph Data
Qun Chen 0001, Zhuo Wang 0002, Bo Suo, Zhanhuai Li, Zachary G. Ives |
DASFAA (2) | 5 |
| 2016 | Towards Scalable Subgraph Pattern Matching over Big Graphs on MapReduceabstractBig graph-structured data pervade our world, ranging from microworld such as gene regulatory networks to macroworld such as social networks. Subgraph matching is a fundamental operation for many graph applications, such as graph database and graph mining. However, existing sequential algorithms have limited applicability on large graphs because of the inherent NP-completeness of subgraph isomorphism and distributed graph storage. Therefore, there is a need to parallelize subgraph matching over big graph data in a distributed environment. With MapReduce as the backdrop, this paper proposes a new approach, named ParMa, for efficient subgraph matching on distributed platforms. It consists of alternate computation and communication phases. We first build a cost model and then propose approaches to optimize the execution process. Instead of existing parallel approaches which only considers intermediate result size, the proposed cost model takes the number of iteration invocations as the primary cost. Based on this, our optimizations mainly focus on the aspects that affects iteration number throughout the execution of matching. One is query decomposition. We propose an effective query decomposition approach to minimize the number of subqueries and their matches. The other is join processing. We introduce a suite of mechanisms, including join plan making, local join processing and join cost estimation, to join partial matches in an appropriate way to reduce its cost. Finally, our extensive experiments on both synthetic and real graphs demonstrated that ParMa outperforms the state-of-the-art solutions by considerable margins. Bo Suo, Zhanhuai Li, Qun Chen 0001, Wei Pan 0007 |
ICPADS | 2 |
| 2016 | Efficient Maximal Clique Enumeration Over Graph DataabstractIn a wide variety of emerging data-intensive applications, such as social network analysis, Web document clustering, entity resolution, and detection of consistently co-expressed genes in systems biology, the detection of dense subgraphs (cliques) is an essential component. Unfortunately, this problem is NP-Complete and thus computationally intensive at scale—hence there is a need for efficient processing, as well as the techniques for distributing the computation across multiple machines such that the computation, which is too time-consuming on a single machine, can be efficiently performed on a machine cluster given that it is large enough. In this paper, we propose a new algorithm (called GP) for maximal clique enumeration. It identifies cliques by the operation of binary graph partitioning, which iteratively divides a graph until each task is sufficiently small to be processed in parallel. Given a connected graph $$G=(V,E)$$ , the GP algorithm has a space complexity of O(|E|) and a time complexity of $$O(|E|\mu (G))$$ , where $$\mu (G)$$ represents the number of different cliques existing in G. We also present a hybrid algorithm, which can effectively leverage the advantages of both the GP algorithm and the classical Bron-and-Kerbosch (BK) algorithm. Then, we develop corresponding parallel solutions based on the GP and hybrid algorithms. Finally, we evaluate the performance of the proposed solutions on real and synthetic graph data. Our extensive experiments show that in both centralized and parallel setting, our proposed GP and hybrid approaches achieve considerably better performance than the state-of-the-art BK approach. Our parallel solutions are implemented and evaluated on MapReduce, a popular shared-nothing parallel framework, but can easily generalize to other shared-nothing or shared-memory parallel frameworks. Boyi Hou, Zhuo Wang 0002, Qun Chen 0001, Bo Suo, Zhanhuai Li, Zachary G. Ives |
Data Sci. Eng. | 6 |
| 2016 | Constrained query of order-preserving submatrix in gene expression data
Tao Jiang 0030, Zhanhuai Li, Xuequn Shang 0001, Weibang Li, Zhilei Yin |
Frontiers Comput. Sci. | 2 |
| 2016 | A probabilistic ranking framework for web-based relational data imputation
Zhaoqiang Chen, Qun Chen 0001, Zhanhuai Li, Lei Chen 0002 |
Inf. Sci. | 4 |
| 2015 | Towards Order-Preserving SubMatrix Search and Indexing
Tao Jiang 0030, Zhanhuai Li, Qun Chen 0001, Kai-Wen Li, Wei Pan 0007 |
DASFAA (2) | 2 |
| 2015 | OMEGA: An Order-Preserving SubMatrix Mining, Indexing and Search Tool
Tao Jiang 0030, Zhanhuai Li, Qun Chen 0001, Kai-Wen Li, Wei Pan 0007 |
ECML/PKDD (3) | 2 |
| 2015 | Discovering Functional Dependencies in Vertically Distributed Big Data
Weibang Li, Zhanhuai Li, Qun Chen 0001, Tao Jiang 0030, Hailong Liu 0004 |
WISE (2) | 2 |
| 2014 | Identification of protein complexes and functional modules in integrated PPI networksabstractMining the protein complexes and functional modules from protein-protein interaction (PPI) networks is vital to understand the mechanism of cellular components and protein functions. Most of the proposed methods had solely focused on static properties of the PPI networks since the available PPI data are static. However, cellular systems are highly dynamic. That is, the interactions of proteins are responsive to environmental cues to accomplish diverse cellular functions. It is important to consider the dynamic inherent within the PPI networks to identify protein complexes and functional modules. In addition, most computational methods did not distinguish between protein complexes and functional modules. It is important to distinguish between them since they are different protein organizations. In this paper, we propose a novel framework to analyze the PPI networks in dynamic conditions by integrating time-series gene expression profiles data and subcelluar localization data. The algorithm, CBMI, is developed to identify protein complexes in integrated PPI networks. By investigating multiple perspectives of proteins in the PPI networks, we identify the “dynamic” hubs in the PPI networks, and then present a new method to discover the functional modules in the PPI networks. The experimental results show that the integration of temporal gene expression data and subcelluar localization data with PPI data contributes to extracting the protein complexes more precisely. Comprehensive evaluations based on f-measure and functional annotations in MIPS database reveal that our algorithm, CBMI, outperforms other previous algorithms in identifying protein complexes, and the detected functional modules are statistically significant in terms of functional annotations. The proposed framework provides a new clue to distinguish between protein complexes and functional modules, and the developed algorithms can be an effective technique for the identification of them. Xuequn Shang 0001, Qingping Zhu, Mingkui Huang, Zhanhuai Li |
BIBM | 5 |
| 2013 | Parallel Partitioning and Mining Gene Expression Data with Butterfly Network
Tao Jiang 0030, Zhanhuai Li, Qun Chen 0001, Wei Pan 0007, Zhuo Wang 0002 |
DEXA (1) | 2 |
| 2012 | Mining Frequent Association Tag Sequences for Clustering XML Documents
Lijun Zhang 0003, Zhanhuai Li, Qun Chen 0001, Ning Li 0022, Ying Lou |
APWeb | 2 |
| 2012 | Semantic relevance ranking for XML keyword search
Ying Lou, Zhanhuai Li, Qun Chen 0001 |
Inf. Sci. | 2 |
| 2011 | Complex Event Processing over Unreliable RFID Data Streams
Yanming Nie, Zhanhuai Li, Qun Chen 0001 |
APWeb | 2 |
| 2011 | Event Detection over Live and Archived Streams
Shanglian Peng, Zhanhuai Li, Qun Chen 0001, Wei Pan 0007, Hailong Liu 0004, Yanming Nie |
WAIM | 2 |
| 2011 | MFCluster: Mining Maximal Fault-Tolerant Constant Row Biclusters in Microarray Dataset
Xuequn Shang 0001, Zhanhuai Li |
WAIM | 4 |
| 2010 | FDTM: Block Level Data Migration Policy in Tiered Storage System
Xiaonan Zhao, Zhanhuai Li, Leijie Zeng |
NPC | 2 |
| 2010 | Online Pattern Aggregation over RFID Data Streams
Hailong Liu 0004, Zhanhuai Li, Qun Chen 0001, Shanglian Peng |
WAIM | 2 |
| 2010 | Efficient Multiple Objects-Oriented Event Detection Over RFID Data Streams
Shanglian Peng, Zhanhuai Li, Qun Chen 0001, Hailong Liu 0004, Yanming Nie, Wei Pan 0007 |
WAIM | 2 |
| 2009 | Mining High-Correlation Association Rules for Inferring Gene Regulation Networks
Xuequn Shang 0001, Zhanhuai Li |
DaWaK | 3 |
| 2009 | Delay and Energy Efficiency Tradeoffs for Data Collections in Large Scale Wireless Sensor NetworksabstractIn this paper, we study efficient data collection and aggregation problem in wireless sensor networks. We first propose efficient distributed algorithms for data collection problem with approximately the minimum delay, or the minimum number of messages to be sent by all wireless nodes, or the minimum total energy consumption by all wireless nodes respectively. For example, given an algorithm A for data collection, let ¿T, ¿M, and ¿Ebe the approximation ratio of A in terms of time complexity, message complexity, and energy complexity respectively. We then show that, for data collection, there are networks of n nodes and maximum degree ¿, such that ¿M¿E= ¿(¿) for any algorithm. In addition, we analytically proved that all our proposed methods are either optimum or within constants factor of the optimum. We further present the message, energy, time complexity and studied the complexity tradeoffs for data aggregation problem. Xufei Mao, Ping Xu 0001, Guojun Dai, Zhanhuai Li |
MASS | 5 |
| 2009 | Probabilistic Modeling of Streaming RFID Data by Using Correlated Variable-duration HMMsabstractRadio frequency identification (RFID) has been widely deployed to track product flow in such fields as automated manufacture, retail and supply chain management. The special characteristics of streaming RFID data, combined with the specific scenarios of RFID applications, present numerous challenges in RFID stream processing, including noisy and incomplete data, temporal and spatial correlations and very huge volumes. In this paper, we present a probabilistic model, specifically correlated variable-duration hidden Markov models (CVD-HMMs), to capture uncertainty and correlations of locations of tagged objects. Based on this model, we can infer object locations from raw RFID streams. And our model can be self-tuned by learning its key parameters from sample RFID readings. Experimental results show that our proposed model and the preliminary inference techniques are effective. Yanming Nie, Zhanhuai Li, Shanglian Peng, Qun Chen 0001 |
SERA | 2 |
| 2009 | Optimization Techniques for RFID Complex Event Processing
Hailong Liu 0004, Qun Chen 0001, Zhanhuai Li |
J. Comput. Sci. Technol. | 3 |
| 2008 | Sequential Pattern Mining for Protein Function Prediction
Xuequn Shang 0001, Zhanhuai Li |
ADMA | 3 |
| 2008 | Optimizing Complex Event Processing over RFID Data StreamsabstractOne research question crucial to RFID technology's wider adoption is how to efficiently transform sequences of RFID readings into meaningful business events. Contrary to traditional events, RFID readings are usually of high volume and velocity, and have the attributes representing their reading objects, occurrence times and spots. Based on these characteristics and the non-deterministic finite automata (NFA) implementation framework, this paper studies the performance issues of RFID complex event processing and proposes corresponding optimization techniques. Our techniques include : (1) taking advantage of negation events or exclusiveness between events to prune intermediate results, thus reduce memory consumption; (2) with complex events' different selectivities, purposefully reordering the join operations between events to improve overall efficiency, thus achieve higher stream throughput; (3) utilizing the slot-based or B+-tree-based approach to optimize the processing performance with the time window constraint. We present these techniques' analytical results and validate their effectiveness through experiments. Qun Chen 0001, Zhanhuai Li, Hailong Liu 0004 |
ICDE | 2 |
| 2007 | RWAR: A Resilient Window-consistent Asynchronous Replication ProtocolabstractAsynchronous replication protocol is playing an increasingly important role in the design of a remote disaster-tolerance system. A resilient window-consistent asynchronous replication protocol (RWAR) is presented in this paper RWAR increases the synchronous feature of asynchronous replication protocol by setting replication space-windows. This can achieve widow-consistency and decrease the risk of the inconsistency between the primary and backup systems. Simultaneously, RWAR dynamically adjusts the size of every space-window by setting checkpoints behind space-windows and calculating the system bandwidth-utility. This can strengthen the resiliency and flexibility of every space-window and ensure the replication performance of the primary system. It's proved with experiments that RWAR affords trade-off between data consistency and replication performance. It is helpful to construct a practical replication-based disaster-tolerance system Zhanhuai Li, Wei Lin 0007 |
ARES | 2 |
| 2007 | An Efficient Encoding and Labeling Scheme for Dynamic XML Data
Zhanhuai Li, Rugui Yao |
DEXA | 2 |
| 2007 | A Fast Disaster Recovery Mechanism for Volume Replication Systems
Zhanhuai Li, Wei Lin 0007 |
HPCC | 2 |
| 2007 | The Design of Finite State Machine for Asynchronous Replication Protocol
Zhanhuai Li, Wei Lin 0007, Minglei Hei, Jianhua Hao |
ICIC (2) | 2 |
| 2007 | New Sampling-Based Summary Structures for Sliding Windows over Data Streams
Longbo Zhang, Zhanhuai Li, Guangyuan Zhao |
ICIC (3) | 2 |
| 2007 | Supporting Multi-attribute Queries in Peer-to-Peer Data Management SystemsabstractSupporting relational query processing or dealing with spatial objects in peer-to-peer(P2P) data management systems needs multi-attribute exact match query processing and multi-attribute range query processing. A scheme to support these queries in P2P data management systems is proposed. By using a multi-attribute order-preserving hash mapping based on a virtual partition tree and indexing the generated keys of the multi-attribute data using P-Grid, data are partitioned dynamically among the dynamic set of peers. After that, a multi-attribute exact match query algorithm and two multi-attribute range query algorithms based on this partitioning strategy are proposed. Finally, two load balancing mechanisms are designed to ensure load balancing when the scheme works in a situation where data distribution in the multi-attribute data space is extremely skewed. Initial analysis shows that this work is effective and efficient. Zhanhuai Li, Longbo Zhang |
PDCAT | 2 |
| 2006 | Improving the Performance of Data Stream Classifiers by Mining Recurring Contexts
Zhanhuai Li, Yang Zhang 0010, Longbo Zhang |
ADMA | 2 |
| 2006 | The Practical Method of Fractal Dimensionality Reduction Based on Z-Ordering Technique
Guanghui Yan, Zhanhuai Li |
ADMA | 2 |
| 2005 | Joining Associative Classifier for Medical ImagesabstractOne of the best prevention measures against breast cancer is the early tumor detection in digital mammography. Detecting tumor in mammography is a difficult task because of their size and the high content of similar patterns in the image. This brings the necessity of creating automatic tools to find whether a mammography present tumor or not. In this paper we join association rule classifier with rough set theory which we call the joining associative classifier (JAC) to mining digital mammography. The experimental results shows that this joining associative classifier performance at 77.48% of classifying accuracy which is higher than 69.11% using associative classifier only. At the same time, the number of rules decreased distinctively. Moreover, the experiments we conducted demonstrate the use and effectiveness of association rule mining in image categorization. Jiang Yun, Zhanhuai Li, Wang Yong, Longbo Zhang |
HIS | 2 |
| 2005 | DRC-BK: Mining Classification Rules by Using Boolean Kernels
Yang Zhang 0010, Zhanhuai Li, Kebin Cui |
ICCSA (1) | 2 |
| 2004 | A New Approach for Selecting Attributes Based on Rough Set Theory
Zhanhuai Li, Yang Zhang 0010 |
IDEAL | 2 |
| 2004 | Modeling of Moving Objects and Querying Videos by TrajectoriesabstractContent-based retrieval of videos in databases is a technique which has attracted considerable research interest during the last years (Guojun, 1999). Most papers in this area discuss how to retrieve some kinds of video scenes involving moving objects according to some example trajectories of these objects according to Nabil (1998) and Aghbari et al. (2000). In this paper, we present a model to describe the trajectories of moving objects and use it to retrieve video scenes. Yan Jianfeng, Zhanhuai Li |
MMM | 2 |
| 2004 | DRC-BK: Mining Classification Rules with Help of SVM
Yang Zhang 0010, Zhanhuai Li, Kebin Cui |
PAKDD | 2 |
| 2003 | Improving the Performance of Text Classifiers by Using Association Features
Yang Zhang 0010, Lijun Zhang 0003, Zhanhuai Li, Yan Jianfeng |
ISMIS | 3 |
| 2003 | The Concept of Attribute Dimension and Corresponding OperationsabstractMember attribute is used to describe the property of dimension members. It is not fully understood or well defined by OLAP research community. We focus on a special kind of member attributes that could also be used as dimensions called attribute dimensions. To facilitate this kind of multidimensional data modeling from real-world applications, the classic multidimensional data structure is extended and a group of algebraic operations are introduced to formulate corresponding multidimensional queries. In this extended model, the attribute dimension is regarded as a special 'view' and can be stated either statically or dynamically based on member attribute. With this approach, both ROLAP and MOLAP can benefit from storage saving and reduced processing time. Compared with current OLAP products and research papers, built-in integrity restraint on member attribute and multidimensional data set makes this extended model unique. Zhanhuai Li |
Web Intelligence | 2 |