Haitian Chen

dblp:246/6539 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 Unsupervised Dense Retrieval with Conterfactual Contrastive Learning
abstract
Efficiently retrieving a concise set of candidates from a large doc- ument corpus remains a pivotal challenge in Information Retrieval (IR). Neural retrieval models, particularly dense retrieval models built with transformers and pretrained language models, have been popular due to their superior performance. However, criticisms have also been raised on their lack of explainability and vulnerability to adversarial attacks. In response to these challenges, we propose to improve the robustness of dense retrieval models by enhancing their sensitivity of fine-grained relevance signals. A model achieving sensitivity in this context should exhibit high variances when doc- uments' key passages determining their relevance to queries have been modified, while maintaining low variances for other changes in irrelevant passages. This sensitivity allows a dense retrieval model to produce robust results with respect to attacks that try to promote documents without actually increasing their relevance. It also makes it possible to analyze which part of a document is actually relevant to a query, and thus improve the explainability of the retrieval model. Motivated by causality and counterfactual analysis, we propose a se- ries of counterfactual regularization methods based on game theory and unsupervised learning with counterfactual passages. Specifically, we first introduce a cooperative game theory-based counterfactual passage extraction method, identifying the key passages that can influence relevance. Then we propose several subsequent unsuper- vised learning tasks, based on these counterfactual passages, serve to regularize the model's learning process to improve the robustness and sensitivity. Experiments show that, our method can extract key passages without reliance on the passage-level relevance annotations. Moreover, the regularized dense retrieval models exhibit heightened robustness against adversarial attacks, surpassing the state-of-the-art anti-attack methods.
Haitian Chen, Qingyao Ai, Yujia Zhou 0002, Xiao Wang 0043, Yiqun Liu 0001, Fen Lin 0002, Qin Liu 0022
WSDM1
2024 Towards Online and Safe Configuration Tuning with Semi-supervised Anomaly Detection
abstract
The performance of modern database management systems highly relies on hundreds of adjustable knobs. Traditionally, these knobs are manually adjusted by database administrators, a process that is both inefficient and ineffective for tuning large-scale databases in cloud environments. Recent research has explored the use of machine learning techniques to enable the automatic tuning of database configurations. Although most existing learning-based methods achieve satisfactory results on static workloads, they often experience performance degradation and low sampling efficiency in real-world environments. According to our study, this is primarily due to a lack of safety guarantees during the configuration sampling process. To address the aforementioned issues, we propose SafeTune, an online tuning system that adapts to dynamic workloads. Our core idea is to filter out a large number of configurations with potential risks during the configuration sampling process. We employ a two-stage filtering approach: The first stage utilizes a semi-supervised outlier ensemble with feature learning to achieve high-quality feature representation. The second stage employs a ranking-based classifier to refine the filtering process. In addition, to alleviate the cold-start problem, we leverage the historical tuning experience to provide high-quality initial samples during the initialization phase. We conducted comprehensive evaluations on static and dynamic workloads. In comparison to offline baseline methods, SafeTune reduces 95.6%-98.6% unsafe configuration suggestions. In contrast with state-of-the-art methods, SafeTune has improved cumulative performance by 10.5%-46.6% and tuning speed by 15.1%-35.4%.
Haitian Chen, Xu Chen 0023, Zibo Liang, Xiushi Feng, Jiandong Xie, Han Su 0001, Kai Zheng 0001
CIKM1
2024 DACE: A Database-Agnostic Cost Estimator
abstract
Cost estimation is of great importance in query optimization. However, traditional optimizers compute the cost based on heuristics, sacrificing accuracy for efficiency. In recent years, learning-based cost estimation models have achieved high accuracy. However, their poor robustness and inefficiency lead to their failure to meet the needs of practical scenarios. We propose a lightweight and Database-Agnostic Cost Estimation model (DACE) to address the above limitations. To further improve the effectiveness of DACE, we design a tree-structure-based loss adjustment strategy to learn sub-plan information and solve the information redundancy problem. As a pretrained estimator, DACE can efficiently make accurate predictions on unseen databases. For more complex scenarios, we fine-tune DACE with LoRA. The excellent efficiency allows DACE to adapt to challenging scenarios with minimal effort. As a pretrained encoder, DACE can improve the accuracy and robustness of other cost estimation models through knowledge integration and solve the notorious cold start problem. Extensive experiments have shown that DACE's accuracy, efficiency, and robustness are much better than existing methods.
Zibo Liang, Xu Chen 0023, Yuyang Xia, Runfan Ye, Haitian Chen, Jiandong Xie, Kai Zheng 0001
ICDE5
2023 Continual Trajectory Prediction with Uncertainty-Aware Generative Memory Replay
abstract
A reliable autonomous driving system should take safe and efficient actions in constantly changing traffic. This requires the trajectory prediction model to continuously learn from incoming data and adapt to new scenarios. In the context of rapidly growing data volume, existing trajectory prediction models must retrain on all datasets to avoid forgetting previously learned knowledge when facing additional data from new environments. In contrast, the paradigm of continual learning solely necessitates training on new data, saving a significant amount of training overhead. Therefore, it is crucial to equip the trajectory prediction model with the ability of continual learning. In this paper, inspired by rehearsal and pseudo-rehearsal methods in continual learning, we propose a continual trajectory prediction framework with uncertainty-aware generative memory replay, CTP-UGR. Our framework effectively avoids excessive memory space requirements while generating trajectory data that is authentic, representative and discriminative for continual learning. Extensive experiments on two real-world datasets demonstrate our proposed CTP-UGR significantly outperforms other baselines in terms of both accuracy and catastrophic forgetting. Besides, our framework can be combined with other state-of-the-art trajectory prediction models to achieve better performance.
Xiushi Feng, Shuncheng Liu 0001, Haitian Chen, Kai Zheng 0001
ICDM3
2023 Behavior Modeling for Point of Interest Search
abstract
With the increasing popularity of location-based services, the point-of-interest (POI) search has received considerable attention in recent years. Existing studies on POI search mostly focus on how to construct better retrieval models to retrieve the relevant POI based on query-POI matching. However, user behavior in POI search, i.e., how users examine the search engine result page (SERP), is mostly underexplored. A good understanding of user behavior is well-recognized as a key to develop effective user models and retrieval models to improve the search quality. Therefore, in this paper, we propose to investigate user behavior in POI search with a lab study in which users' eye movements and their implicit feedback on the SERP are collected. Based on the collected data, we analyze (1) query-level user behavior patterns in POI search, i.e., examination and interactions on SERP; (2) session-level user behavior patterns in POI search, i.e., query reformulation, termination of search, etc. Our work sheds light on user behavior in POI search and could potentially benefit future studies on related research topics.
Haitian Chen, Qingyao Ai, Zhijing Wu 0001, Yiqun Liu 0001, Min Zhang 0006, Shaoping Ma, Naiqiang Tan
SIGIR1
2023 Deep Learning-Based Bloom Filter for Efficient Multi-key Membership Testing
abstract
Abstract Multi-key membership testing plays a crucial role in computing systems and networking applications, encompassing web search, mail systems, distributed databases, firewalls, and network routing. Traditional approaches, such as the Bloom filter, encounter limitations within this specific context. Addressing these challenges, we propose the Multi-key Learned Bloom Filter (MLBF), a hybrid method that combines machine learning techniques with the Bloom filter. The MLBF introduces a value-interaction-based multi-key classifier and a multi-key Bloom filter. Furthermore, we introduce an Interval-based MLBF approach, which categorizes keys into specific intervals based on data distribution to minimize the False Positive Rate (FPR). Additionally, MLBF incorporates an out-of-distribution (OOD) detection component to identify data shifts. Through extensive experimental evaluations on three authentic datasets, we demonstrate the superiority of the proposed MLBF in terms of FPR and query efficiency.
Haitian Chen, Yunchuan Li, Yan Zhao 0008, Rui Zhou 0015, Kai Zheng 0001
Data Sci. Eng.1
2023 LEON: A New Framework for ML-Aided Query Optimization
abstract
Query optimization has long been a fundamental yet challenging topic in the database field. With the prosperity of machine learning (ML), some recent works have shown the advantages of reinforcement learning (RL) based learned query optimizer. However, they suffer from fundamental limitations due to the data-driven nature of ML. Motivated by the ML characteristics and database maturity, we propose LEON -a framework for ML-aidEd query OptimizatioN. LEON improves the expert query optimizer to self-adjust to the particular deployment by leveraging ML and the fundamental knowledge in the expert query optimizer. To train the ML model, a pairwise ranking objective is proposed, which is substantially different from the previous regression objective. To help the optimizer to escape the local minima and avoid failure, a ranking and uncertainty-based exploration strategy is proposed, which discovers the valuable plans to aid the optimizer. Furthermore, an ML model-guided pruning is proposed to increase the planning efficiency without hurting too much performance. Extensive experiments offer evidence that the proposed framework can outperform the state-of-the-art methods in terms of end-to-end latency performance, training efficiency, and stability.
Xu Chen 0023, Haitian Chen, Zibo Liang, Shuncheng Liu 0001, Kai Zeng 0002, Han Su 0001, Kai Zheng 0001
Proc. VLDB Endow.2
2020 Preference-based Evaluation Metrics for Web Image Search
abstract
Following the success of Cranfield-like evaluation approaches to evaluation in web search, web image search has also been evaluated with absolute judgments of (graded) relevance. However, recent research has found that collecting absolute relevance judgments may be difficult in image search scenarios due to the multi-dimensional nature of relevance for image results. Moreover, existing evaluation metrics based on absolute relevance judgments do not correlate well with search users' satisfaction perceptions in web image search.
Xiaohui Xie, Jiaxin Mao, Yiqun Liu 0001, Maarten de Rijke, Haitian Chen, Min Zhang 0006, Shaoping Ma
SIGIR5
2019 Optimization Location Selection Analysis of Energy Storage Unit in Energy Internet System Based on Tabu Search
abstract
Energy Internet has become the theme of the new round of industrial revolution. Energy storage, as a key technical support for the development of energy Internet, has always been of concern to numerous people, since the energy Internet consists of various energy networks that can provide energy support for different energy subnetworks. Therefore, the energy storage unit is in a crucial position in the entire energy network. This paper points out the importance of various energy storage technologies in the energy Internet. An energy storage unit location analysis method based on Tabu search algorithm is proposed to reduce the network energy loss, pressing mimizing network loss as constraint on the location of the energy storage unit as a search target. The Tabu search algorithm is programmed using Matlab and is used to search for the location of energy storage unit in the IEEE example. Besides, the optimal node solution is obtained, which verifies the feasibility of this algorithm to analyze the location selection of energy storage unit in the energy Internet. This paper has some reference value for the coordinated optimization of energy storage units in the energy Internet.
Haitian Chen, Yanwei Ji, Shunjiang Wang, Weichun Ge, Anlong Su
Int. J. Softw. Eng. Knowl. Eng.1