VLDB 2026 Research / reviewers in the wild / expert
Haixia Han
dblp:250/5128
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2024
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Small Language Model Can Self-CorrectabstractGenerative Language Models (LMs) such as ChatGPT have exhibited remarkable performance across various downstream tasks. Nevertheless, one of their most prominent drawbacks is generating inaccurate or false information with a confident tone. Previous studies have devised sophisticated pipelines and prompts to induce large LMs to exhibit the capability for self-correction. However, large LMs are explicitly prompted to verify and modify their answers separately rather than completing all steps spontaneously like humans. Moreover, these complex prompts are extremely challenging for small LMs to follow. In this paper, we introduce the Intrinsic Self-Correction (ISC) in generative language models, aiming to correct the initial output of LMs in a self-triggered manner, even for those small LMs with 6 billion parameters. Specifically, we devise a pipeline for constructing self-correction data and propose Partial Answer Masking (PAM), aiming to endow the model with the capability for intrinsic self-correction through fine-tuning. We conduct experiments using LMs with parameters sizes ranging from 6 billion to 13 billion in two tasks, including commonsense reasoning and factual knowledge reasoning. Our experiments demonstrate that the outputs generated using ISC outperform those generated without self-correction. We believe that the output quality of even small LMs can be further improved by empowering them with the ability to intrinsic self-correct. Haixia Han, Jiaqing Liang, Jie Shi 0010, Qianyu He, Yanghua Xiao |
AAAI | 1 |
| 2023 | Thompson Sampling with Time-Varying Reward for Contextual Bandits
Cairong Yan, Hualu Xu, Haixia Han, Yanting Zhang 0001, Zijian Wang 0010 |
DASFAA (2) | 3 |
| 2023 | MB-DP: A Multi-behavior Recommendation Model Integrating Dynamic PreferencesabstractMulti-behavior recommendation has gained significant attention in recent years for its ability to outperform singlebehavior models.Current research related to multi-behavior models leaves room for improvement in the following two areas.First, the noise carried by individual behaviors and the additional noise generated during behavior processing is often overlooked, and these can ultimately degrade recommendation performance.Second, the specific time period of behavioral interactions and the frequency of interactions within that time period are also not taken into account.To address the above limitations, we propose a multibehavior recommendation model integrating dynamic preferences (MB-DP) that captures dynamic interests while smoothing and denoising multi-behavior information.MB-DP extracts low and high-order semantics from various behaviors and unifies the measurements to generate interaction predictions.Additionally, it analyzes the interaction time and frequency of each behavior using gated recurrent units to capture the dynamic preferences of users and improve the prediction values.Extensive experimental results on two real-world datasets show that MB-DP significantly improves recommendation performance compared to the state-ofthe-art baselines. Cairong Yan, Xiaopeng Guan, Haixia Han, Zhaohui Zhang 0001 |
SEKE | 3 |
| 2023 | Enhancing Multi-Behavior Recommendations Through Capturing Dynamic PreferencesabstractMulti-behavior recommendation has gained significant attention in recent years for its ability to outperform single-behavior models. Current research related to multi-behavior models leaves room for improvement in the following two areas. First, the noise carried by individual behaviors and the additional noise generated during behavior processing is often overlooked, and these can ultimately degrade recommendation performance. Second, the specific time period of behavioral interactions and the frequency of interactions within that time period are also not taken into account. To address the above limitations, we propose a multi-behavior recommendation model integrating dynamic preferences (MB-DP) that captures dynamic interests while smoothing and denoising multi-behavior information. MB-DP extracts low and high-order semantics from various behaviors and unifies the measurements to generate interaction predictions. Additionally, it analyzes the interaction time and frequency of each behavior using gated recurrent units to capture the dynamic preferences of users and improve the prediction values. Extensive experimental results on two real-world datasets show that MB-DP significantly improves recommendation performance compared to the state-of-the-art baselines. Cairong Yan, Xiaopeng Guan, Haixia Han, Zhaohui Zhang 0001, Yanting Zhang 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2022 | Dynamic clustering based contextual combinatorial multi-armed bandit for online recommendationabstractRecommender systems still face a trade-off between exploring new items to maximize user satisfaction and exploiting those already interacted with to match user interests. This problem is widely recognized as the exploration/exploitation (EE) dilemma, and the multi-armed bandit (MAB) algorithm has proven to be an effective solution. As the scale of users and items in real-world application scenarios increases, their purchase interactions become sparser. Then three issues need to be investigated when building MAB-based recommender systems. First, large-scale users and sparse interactions increase the difficulty of user preference mining. Second, traditional bandits model items as arms and cannot deal with ever-growing items effectively. Third, widely used Bernoulli-based reward mechanisms only feedback 0 or 1, ignoring rich implicit feedback such as behaviors like click and add-to-cart. To address these problems, we propose an algorithm named Dynamic Clustering based Contextual Combinatorial Multi-Armed Bandits (DC3MAB), which consists of three configurable key components. Specifically, a dynamic user clustering strategy enables different users in the same cluster to cooperate in estimating the expected rewards of arms. A dynamic item partitioning approach based on collaborative filtering significantly reduces the scale of arms and produces a recommendation list instead of one item to provide diversity. In addition, a multi-class reward mechanism based on fine-grained implicit feedback helps better capture user preferences. Extensive empirical experiments on three real-world datasets demonstrate the superiority of our proposed DC3MAB over state-of-the-art bandits (On average, +75.8% in F1 and +54.3% in cumulative reward). The source code is available at https://github.com/HaixHan/DC3MAB. Cairong Yan, Haixia Han, Yanting Zhang 0001, Dandan Zhu 0001, Yongquan Wan |
Knowl. Based Syst. | 2 |
| 2021 | Two-Phase Multi-armed Bandit for Online RecommendationabstractPersonalized online recommendations strive to adapt their services to individual users by making use of both item and user information. Despite recent progress, the issue of balancing exploitation-exploration (EE) [1] remains challenging. In this paper, we model the personalized online recommendation of e-commence as a two-phase multi-armed bandit problem. This is the first time that “big arm” and “small arm” are introduced into multi-armed bandit (MAB), and a two-stage strategy is adopted to provide target users with the most suitable recommendation list. In the first phase, MAB is used to obtain an item subset that users may be interested in from a large number of items. We use item categories as arms instead of individual items in existing related models to control the arm scale and reduce computational complexity. In the second phase, we directly use the items generated in the first phase as arms of MAB and obtain rewards through fine-grained implicit feedback from users. Empirical studies on three real-world datasets show that our proposed method TPBandit performs better than state-of-the-art bandit-based recommendation methods in several evaluation metrics such as Precision, Recall, and Hit Ratio. Moreover, the two-phase method improves the recommendation performance by nearly 50% compared to the one-phase method in the best case. Cairong Yan, Haixia Han, Zijian Wang 0010, Yanting Zhang 0001 |
DSAA | 2 |