Qun Chen 0001

dblp:16/1091-1 · DBLP profile ↗
← Back
45ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0001-9556-755XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 27 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5Systems, architecture and hardware · 3Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Adaptive deep learning for slide-level multi-label biomarker prediction in breast cancer whole slide images via misprediction risk analysis
Gul Sheeraz, Wangwang Li, Qun Chen 0001, Fengjin Zhou
Expert Syst. Appl.4
2024 Towards Exploratory Query Optimization for Template-Based SQL Workloads
abstract
SQL query optimization aims to choose an optimal Query Execution Plan (QEP) for a query. The existing optimizers usually choose the plan with the minimal execution cost. However, in some real scenarios (e.g., cloud OLAP), it is a business imperative to minimize the cost of an entire workload. Unfortunately, the greedy optimizers are prone to generating execution plans that are only suboptimal from the workload perspective. In this paper, we propose a novel QEP optimization approach for template-based SQL that aims to minimize the execution cost of an entire workload instead of individual queries. While choosing an execution plan for a query, the proposed approach considers a plan's impact on future query execution as well as its own execution cost. We first define an exploratory metric, Plan Exploration Value (PEV), to measure a plan's potential benefit to future query execution. Then, we model exploratory plan selection as an optimization problem with cost constraints, and present an efficient algorithm based on dynamic programming for its solution. Finally, we evaluate the performance of the proposed approach by a comparative study on benchmark workloads. Our extensive experiments have shown that it generates more efficient workload execution plans than the existing alternatives.
Jieming Feng, Zhanhuai Li, Qun Chen 0001
ICDE3
2024 Few-shot image classification based on gradual machine learning
Xianming Kuang, Feiyu Liu, Kehao Wang 0003, Lijun Zhang 0003, Qun Chen 0001
Expert Syst. Appl.6
2024 Adaptive deep learning for entity disambiguation via knowledge-based risk analysis
Youcef Nafa, Qun Chen 0001, Boyi Hou, Zhanhuai Li
Expert Syst. Appl.2
2024 Automating localized learning for cardinality estimation based on XGBoost
Jieming Feng, Zhanhuai Li, Qun Chen 0001, Hailong Liu 0004
Knowl. Inf. Syst.3
2023 Adaptive deep learning for entity resolution by risk analysis
Qun Chen 0001, Zhaoqiang Chen, Youcef Nafa, Tianyi Duan, Wei Pan 0007, Lijun Zhang 0003, Zhanhuai Li
Knowl. Based Syst.1
2023 Supervised Gradual Machine Learning for Aspect-Term Sentiment Analysis
abstract
Abstract Recent work has shown that Aspect-Term Sentiment Analysis (ATSA) can be effectively performed by Gradual Machine Learning (GML). However, the performance of the current unsupervised solution is limited by inaccurate and insufficient knowledge conveyance. In this paper, we propose a supervised GML approach for ATSA, which can effectively exploit labeled training data to improve knowledge conveyance. It leverages binary polarity relations between instances, which can be either similar or opposite, to enable supervised knowledge conveyance. Besides the explicit polarity relations indicated by discourse structures, it also separately supervises a polarity classification DNN and a binary Siamese network to extract implicit polarity relations. The proposed approach fulfills knowledge conveyance by modeling detected relations as binary features in a factor graph. Our extensive experiments on real benchmark data show that it achieves the state-of-the-art performance across all the test workloads. Our work demonstrates clearly that, in collaboration with DNN for feature extraction, GML outperforms pure DNN solutions.
Yanyan Wang 0005, Qun Chen 0001, Murtadha H. M. Ahmed, Zhaoqiang Chen, Wei Pan 0007, Zhanhuai Li
Trans. Assoc. Comput. Linguistics2
2022 Adaptive deep learning for network intrusion detection by risk analysis
Lijun Zhang 0003, Zhaoqiang Chen, Tianwei Liu, Qun Chen 0001, Zhanhuai Li
Neurocomputing5
2022 Active deep learning on entity resolution by risk sampling
Youcef Nafa, Qun Chen 0001, Zhaoqiang Chen, Tianyi Duan, Zhanhuai Li
Knowl. Based Syst.2
2022 Gradual Machine Learning for Entity Resolution
abstract
Usually considered as a classification problem, entity resolution (ER) can be very challenging on real data due to the prevalence of dirty values. The state-of-the-art solutions for ER were built on a variety of learning models (most notably deep neural networks), which require lots of accurately labeled training data. Unfortunately, high-quality labeled data usually require expensive manual work, and are therefore not readily available in many real scenarios. In this paper, we propose a novel learning paradigm for ER, calledgradual machine learning, which aims to enable effective machine labeling without the requirement for manual labeling effort. It begins with some easy instances in a task, which can be automatically labeled by the machine with high accuracy, and then gradually labels more challenging instances by iterative factor graph inference. In gradual machine learning, the hard instances in a task are gradually labeled in small stages based on the estimated evidential certainty provided by the labeled easier instances. Our extensive experiments on real data have shown that the performance of the proposed approach is considerably better than its unsupervised alternatives, and highly competitive compared to the state-of-the-art supervised techniques. Using ER as a test case, we demonstrate that gradual machine learning is a promising paradigm potentially applicable to other challenging classification tasks requiring extensive labeling effort.
Boyi Hou, Qun Chen 0001, Yanyan Wang 0005, Youcef Nafa, Zhanhuai Li
IEEE Trans. Knowl. Data Eng.2
2021 Aspect-level sentiment analysis based on gradual machine learning
Yanyan Wang 0005, Qun Chen 0001, Jiquan Shen, Boyi Hou, Murtadha H. M. Ahmed, Zhanhuai Li
Knowl. Based Syst.2
2021 Joint Inference for Aspect-Level Sentiment Analysis by Deep Neural Networks and Linguistic Hints
abstract
The state-of-the-art techniques for aspect-level sentiment analysis focused on feature modeling using a variety of deep neural networks (DNN). Unfortunately, their performance may still fall short of expectation in real scenarios due to the semantic complexity of natural languages. Motivated by the observation that many linguistic hints (e.g., sentiment words and shift words) are reliable polarity indicators, we propose a joint framework, SenHint, which can seamlessly integrate the output of deep neural networks and the implications of linguistic hints in a unified model based on Markov logic network (MLN). SenHint leverages the linguistic hints for multiple purposes: (1) to identify the easy instances, whose polarities can be automatically determined by the machine with high accuracy; (2) to capture the influence of sentiment words on aspect polarities; (2) to capture the implicit relations between aspect polarities. We present the required techniques for extracting linguistic hints, encoding their implications as well as the output of DNN into the unified model, and joint inference. Finally, we have empirically evaluated the performance of SenHint on both English and Chinese benchmark datasets. Our extensive experiments have shown that compared to the state-of-the-art DNN techniques, SenHint can effectively improve polarity detection accuracy by considerable margins.
Yanyan Wang 0005, Qun Chen 0001, Murtadha H. M. Ahmed, Zhanhuai Li, Wei Pan 0007, Hailong Liu 0004
IEEE Trans. Knowl. Data Eng.2
2020 Towards Interpretable and Learnable Risk Analysis for Entity Resolution
abstract
Machine-learning-based entity resolution has been widely studied. However, some entity pairs may be mislabeled by machine learning models and existing studies do not study the risk analysis problem -- predicting and interpreting which entity pairs are mislabeled. In this paper, we propose an interpretable and learnable framework for risk analysis, which aims to rank the labeled pairs based on their risks of being mislabeled. We first describe how to automatically generate interpretable risk features, and then present a learnable risk model and its training technique. Finally, we empirically evaluate the performance of the proposed approach on real data. Our extensive experiments have shown that the learning risk model can identify the mislabeled pairs with considerably higher accuracy than the existing alternatives.
Zhaoqiang Chen, Qun Chen 0001, Boyi Hou, Zhanhuai Li, Guoliang Li 0001
SIGMOD Conference2
2020 Constructing domain-dependent sentiment dictionary for sentiment analysis
Murtadha H. M. Ahmed, Qun Chen 0001, Zhanhuai Li
Neural Comput. Appl.2
2020 r-HUMO: A Risk-Aware Human-Machine Cooperation Framework for Entity Resolution with Quality Guarantees
abstract
Even though many approaches have been proposed for entity resolution (ER), it remains very challenging to enforce quality guarantees. To this end, we propose a risk-aware HUman-Machine cOoperation framework for ER, denoted by r-HUMO. Built on the existing HUMO framework, r-HUMO similarly enforces both precision and recall guarantees by partitioning an ER workload between the human and the machine. However, r-HUMO is the first solution that optimizes the process of human workload selection from a risk perspective. It iteratively selects human workload by real-time risk analysis based on the human-labeled results as well as the prespecified machine metric. In this paper, we first introduce the r-HUMO framework and then present the risk model to prioritize the instances for manual inspection. Finally, we empirically evaluate r-HUMO's performance on real data. Our extensive experiments show that r-HUMO is effective in enforcing quality guarantees, and compared with the state-of-the-art alternatives, it can achieve desired quality control with reduced human cost.
Boyi Hou, Qun Chen 0001, Zhaoqiang Chen, Youcef Nafa, Zhanhuai Li
IEEE Trans. Knowl. Data Eng.2
2019 Hint-Embedding Attention-Based LSTM for Aspect Identification Sentiment Analysis
Murtadha H. M. Ahmed, Qun Chen 0001, Yanyan Wang 0005, Zhanhuai Li
PRICAI (2)2
2019 Gradual Machine Learning for Entity Resolution
abstract
Usually considered as a classification problem, entity resolution can be very challenging on real data due to the prevalence of dirty values. The state-of-the-art solutions for ER were built on a variety of learning models (most notably deep neural networks), which require lots of accurately labeled training data. Unfortunately, high-quality labeled data usually require expensive manual work, and are therefore not readily available in many real scenarios. In this demo, we propose a novel learning paradigm for ER, called gradual machine learning, which aims to enable effective machine labeling without the requirement for manual labeling effort. It begins with some easy instances in a task, which can be automatically labeled by the machine with high accuracy, and then gradually labels more challenging instances based on iterative factor graph inference. In gradual machine learning, the hard instances in a task are gradually labeled in small stages based on the estimated evidential certainty provided by the labeled easier instances. Our extensive experiments on real data have shown that the proposed approach performs considerably better than its unsupervised alternatives, and its performance is also highly competitive compared to the state-of-the-art supervised techniques. Using ER as a test case, we demonstrate that gradual machine learning is a promising paradigm potentially applicable to other challenging classification tasks requiring extensive labeling effort. Video: https://youtu.be/99bA9aamsgk
Boyi Hou, Qun Chen 0001, Jiquan Shen, Ping Zhong 0004, Yanyan Wang 0005, Zhaoqiang Chen, Zhanhuai Li
WWW2
2019 Reducing partition skew on MapReduce: an incremental allocation approach
Zhuo Wang 0002, Qun Chen 0001, Bo Suo, Wei Pan 0007, Zhanhuai Li
Frontiers Comput. Sci.2
2018 GraphU: A Unified Vertex-Centric Parallel Graph Processing Platform
abstract
Many synchronous and asynchronous distributed platforms based on the Bulk Synchronous Parallel (BSP) model have been built for large-scale vertex -centric graph processing. Unfortunately, a program designed for a synchronous platform may not work properly on an asynchronous one. As a result, given the same problem, end users may be required to design different parallel algorithms for different platforms. Recently, we have proposed a unified programming model, DFA-G (Deterministic Finite Automaton for Graph processing), which expresses the computation at a vertex as a series of message-driven state transitions. It has the attractive property that any program modeled after it can run properly across synchronous and asynchronous platforms. In this demo, we first propose a framework of complexity analysis for DFA-G automaton and show that it can significantly facilitate complexity analysis on asynchronous programs. Due to the existing BSP platforms' deficiency in supporting efficient DFA-G execution, we then develop a new prototype platform, GraphU. GraphU was built on the popular open-source Giraph project. But it entirely removes synchronization barriers and decouples remote communication from vertex computation. Finally, we empirically evaluate the performance of various DFA-G programs on GraphU by a comparative study. Our experiments validate the efficacy of the proposed complexity analysis approach and the efficiency of GraphU.
Qun Chen 0001, Zhuo Wang 0002, Murtadha H. M. Ahmed, Zhanhuai Li
ICDCS2
2018 Enabling Quality Control for Entity Resolution: A Human and Machine Cooperation Framework
abstract
Even though many machine algorithms have been proposed for entity resolution, it remains very challenging to find a solution with quality guarantees. In this paper, we propose a novel HUman and Machine cOoperation (HUMO) framework for entity resolution (ER), which divides an ER workload between the machine and the human. HUMO enables a mechanism for quality control that can flexibly enforce both precision and recall levels. We introduce the optimization problem of HUMO, minimizing human cost given a quality requirement, and then present three optimization approaches: a conservative baseline one purely based on the monotonicity assumption of precision, a more aggressive one based on sampling and a hybrid one that can take advantage of the strengths of both previous approaches. Finally, we demonstrate by extensive experiments on real and synthetic datasets that HUMO can achieve high-quality results with reasonable return on investment (ROI) in terms of human cost, and it performs considerably better than the state-of-the-art alternatives in quality control.
Zhaoqiang Chen, Qun Chen 0001, Fengfeng Fan, Yanyan Wang 0005, Zhuo Wang 0002, Youcef Nafa, Zhanhuai Li, Hailong Liu 0004, Wei Pan 0007
ICDE2
2018 Automatic Web-based relational data imputation
Hailong Liu 0004, Zhanhuai Li, Qun Chen 0001, Zhaoqiang Chen
Frontiers Comput. Sci.3
2018 GL-RF: a reconciliation framework for label-free entity resolution
Yaoli Xu, Zhanhuai Li, Qun Chen 0001, Fengfeng Fan
Frontiers Comput. Sci.3
2018 Reasoning about attribute value equivalence in relational data
Fengfeng Fan, Zhanhuai Li, Qun Chen 0001, Lei Chen 0002
Inf. Syst.3
2018 Relational data imputation with quality guarantee
Fengfeng Fan, Zhanhuai Li, Qun Chen 0001, Lei Chen 0002
Inf. Sci.3
2017 POOLSIDE: An Online Probabilistic Knowledge Base for Shopping Decision Support
abstract
We present POOLSIDE, an online PrObabilistic knOwLedge base for ShoppIng DEcision support, that provides with the on-target recommendation service based on explicit user requirement. With a natural language interface, POOLSIDE can answer question in real-time. We present how to construct the knowledge base and how to enable real-time response in POOLSIDE. Finally, we demonstrate that Poolside can give high-quality product recommendations with high efficiency.(The demo video can be accessed via the link:https://www.youtube.com/watch?v=D8ALi11CUcc)
Ping Zhong 0004, Zhanhuai Li, Qun Chen 0001, Yanyan Wang 0005, Lianping Wang, Murtadha H. M. Ahmed, Fengfeng Fan
CIKM3
2017 A Human-and-Machine Cooperative Framework for Entity Resolution with Quality Guarantees
abstract
For entity resolution, it remains very challenging to find the solution with quality guarantees as measured by both precision and recall. In this demo, we propose a HUman-and-Machine cOoperative framework, denoted by HUMO, for entity resolution. Compared with the existing approaches, HUMO enables a flexible mechanism for quality control that can enforce both precision and recall levels. We also introduce the problem of minimizing human cost given a quality requirement and present corresponding optimization techniques. Finally, we demo that HUMO achieves high-quality results with reasonable return on investment (ROI) in terms of human cost on real datasets.
Zhaoqiang Chen, Qun Chen 0001, Zhanhuai Li
ICDE2
2017 Parallelizing maximal clique and k-plex enumeration over graph data
Zhuo Wang 0002, Qun Chen 0001, Boyi Hou, Bo Suo, Zhanhuai Li, Wei Pan 0007, Zachary G. Ives
J. Parallel Distributed Comput.2
2016 Discovering Approximate Functional Dependencies from Distributed Big Data
Weibang Li, Zhanhuai Li, Qun Chen 0001, Tao Jiang 0030, Zhilei Yin
APWeb (2)3
2016 Parallelizing Maximal Clique Enumeration Over Graph Data
Qun Chen 0001, Zhuo Wang 0002, Bo Suo, Zhanhuai Li, Zachary G. Ives
DASFAA (2)1
2016 Towards Scalable Subgraph Pattern Matching over Big Graphs on MapReduce
abstract
Big graph-structured data pervade our world, ranging from microworld such as gene regulatory networks to macroworld such as social networks. Subgraph matching is a fundamental operation for many graph applications, such as graph database and graph mining. However, existing sequential algorithms have limited applicability on large graphs because of the inherent NP-completeness of subgraph isomorphism and distributed graph storage. Therefore, there is a need to parallelize subgraph matching over big graph data in a distributed environment. With MapReduce as the backdrop, this paper proposes a new approach, named ParMa, for efficient subgraph matching on distributed platforms. It consists of alternate computation and communication phases. We first build a cost model and then propose approaches to optimize the execution process. Instead of existing parallel approaches which only considers intermediate result size, the proposed cost model takes the number of iteration invocations as the primary cost. Based on this, our optimizations mainly focus on the aspects that affects iteration number throughout the execution of matching. One is query decomposition. We propose an effective query decomposition approach to minimize the number of subqueries and their matches. The other is join processing. We introduce a suite of mechanisms, including join plan making, local join processing and join cost estimation, to join partial matches in an appropriate way to reduce its cost. Finally, our extensive experiments on both synthetic and real graphs demonstrated that ParMa outperforms the state-of-the-art solutions by considerable margins.
Bo Suo, Zhanhuai Li, Qun Chen 0001, Wei Pan 0007
ICPADS3
2016 Efficient Maximal Clique Enumeration Over Graph Data
abstract
In a wide variety of emerging data-intensive applications, such as social network analysis, Web document clustering, entity resolution, and detection of consistently co-expressed genes in systems biology, the detection of dense subgraphs (cliques) is an essential component. Unfortunately, this problem is NP-Complete and thus computationally intensive at scale—hence there is a need for efficient processing, as well as the techniques for distributing the computation across multiple machines such that the computation, which is too time-consuming on a single machine, can be efficiently performed on a machine cluster given that it is large enough. In this paper, we propose a new algorithm (called GP) for maximal clique enumeration. It identifies cliques by the operation of binary graph partitioning, which iteratively divides a graph until each task is sufficiently small to be processed in parallel. Given a connected graph $$G=(V,E)$$ , the GP algorithm has a space complexity of O(|E|) and a time complexity of $$O(|E|\mu (G))$$ , where $$\mu (G)$$ represents the number of different cliques existing in G. We also present a hybrid algorithm, which can effectively leverage the advantages of both the GP algorithm and the classical Bron-and-Kerbosch (BK) algorithm. Then, we develop corresponding parallel solutions based on the GP and hybrid algorithms. Finally, we evaluate the performance of the proposed solutions on real and synthetic graph data. Our extensive experiments show that in both centralized and parallel setting, our proposed GP and hybrid approaches achieve considerably better performance than the state-of-the-art BK approach. Our parallel solutions are implemented and evaluated on MapReduce, a popular shared-nothing parallel framework, but can easily generalize to other shared-nothing or shared-memory parallel frameworks.
Boyi Hou, Zhuo Wang 0002, Qun Chen 0001, Bo Suo, Zhanhuai Li, Zachary G. Ives
Data Sci. Eng.3
2016 A probabilistic ranking framework for web-based relational data imputation
Zhaoqiang Chen, Qun Chen 0001, Zhanhuai Li, Lei Chen 0002
Inf. Sci.2
2015 Towards Order-Preserving SubMatrix Search and Indexing
Tao Jiang 0030, Zhanhuai Li, Qun Chen 0001, Kai-Wen Li, Wei Pan 0007
DASFAA (2)3
2015 OMEGA: An Order-Preserving SubMatrix Mining, Indexing and Search Tool
Tao Jiang 0030, Zhanhuai Li, Qun Chen 0001, Kai-Wen Li, Wei Pan 0007
ECML/PKDD (3)3
2015 Discovering Functional Dependencies in Vertically Distributed Big Data
Weibang Li, Zhanhuai Li, Qun Chen 0001, Tao Jiang 0030, Hailong Liu 0004
WISE (2)3
2013 Parallel Partitioning and Mining Gene Expression Data with Butterfly Network
Tao Jiang 0030, Zhanhuai Li, Qun Chen 0001, Wei Pan 0007, Zhuo Wang 0002
DEXA (1)3
2012 Mining Frequent Association Tag Sequences for Clustering XML Documents
Lijun Zhang 0003, Zhanhuai Li, Qun Chen 0001, Ning Li 0022, Ying Lou
APWeb3
2012 Semantic relevance ranking for XML keyword search
Ying Lou, Zhanhuai Li, Qun Chen 0001
Inf. Sci.3
2011 Complex Event Processing over Unreliable RFID Data Streams
Yanming Nie, Zhanhuai Li, Qun Chen 0001
APWeb3
2011 Event Detection over Live and Archived Streams
Shanglian Peng, Zhanhuai Li, Qun Chen 0001, Wei Pan 0007, Hailong Liu 0004, Yanming Nie
WAIM4
2010 Online Pattern Aggregation over RFID Data Streams
Hailong Liu 0004, Zhanhuai Li, Qun Chen 0001, Shanglian Peng
WAIM3
2010 Efficient Multiple Objects-Oriented Event Detection Over RFID Data Streams
Shanglian Peng, Zhanhuai Li, Qun Chen 0001, Hailong Liu 0004, Yanming Nie, Wei Pan 0007
WAIM4
2009 Probabilistic Modeling of Streaming RFID Data by Using Correlated Variable-duration HMMs
abstract
Radio frequency identification (RFID) has been widely deployed to track product flow in such fields as automated manufacture, retail and supply chain management. The special characteristics of streaming RFID data, combined with the specific scenarios of RFID applications, present numerous challenges in RFID stream processing, including noisy and incomplete data, temporal and spatial correlations and very huge volumes. In this paper, we present a probabilistic model, specifically correlated variable-duration hidden Markov models (CVD-HMMs), to capture uncertainty and correlations of locations of tagged objects. Based on this model, we can infer object locations from raw RFID streams. And our model can be self-tuned by learning its key parameters from sample RFID readings. Experimental results show that our proposed model and the preliminary inference techniques are effective.
Yanming Nie, Zhanhuai Li, Shanglian Peng, Qun Chen 0001
SERA4
2009 Optimization Techniques for RFID Complex Event Processing
Hailong Liu 0004, Qun Chen 0001, Zhanhuai Li
J. Comput. Sci. Technol.2
2008 Optimizing Complex Event Processing over RFID Data Streams
abstract
One research question crucial to RFID technology's wider adoption is how to efficiently transform sequences of RFID readings into meaningful business events. Contrary to traditional events, RFID readings are usually of high volume and velocity, and have the attributes representing their reading objects, occurrence times and spots. Based on these characteristics and the non-deterministic finite automata (NFA) implementation framework, this paper studies the performance issues of RFID complex event processing and proposes corresponding optimization techniques. Our techniques include : (1) taking advantage of negation events or exclusiveness between events to prune intermediate results, thus reduce memory consumption; (2) with complex events' different selectivities, purposefully reordering the join operations between events to improve overall efficiency, thus achieve higher stream throughput; (3) utilizing the slot-based or B+-tree-based approach to optimize the processing performance with the time window constraint. We present these techniques' analytical results and validate their effectiveness through experiments.
Qun Chen 0001, Zhanhuai Li, Hailong Liu 0004
ICDE1