EDBT 2026 Demo / reviewers in the wild / expert
Hanfeng Liu
dblp:00/588
· DBLP profile ↗
7ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 36% Representation and self-supervised learning · 18% Information extraction and text analysis · 18% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › robustness › multimodal robustness
missing modality robustness |
0.9 | 1 | 2025 | RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition · IJCAI 2025 |
Natural language and speech › Information extraction and text analysis › emotion recognition
multimodal emotion recognition |
0.9 | 1 | 2025 | RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition · IJCAI 2025 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.9 | 1 | 2025 | RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition · IJCAI 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition · IJCAI 2025 |
Data mining › predictive modeling
classification |
0.5 | 1 | 2021 | Enhancing SVMs with Problem Context Aware Pipeline · KDD 2021 |
Data mining › predictive modeling › classification
support vector machine |
0.5 | 1 | 2021 | Enhancing SVMs with Problem Context Aware Pipeline · KDD 2021 |
Machine learning › Efficient and distributed learning › hardware acceleration
GPU training |
0.4 | 1 | 2020 | ThunderGBM: Fast GBDTs and Random Forests on GPUs · J. Mach. Learn. Res. 2020 |
Machine learning › Kernel, tree and ensemble methods › gradient boosting
gradient boosting decision tree |
0.4 | 1 | 2020 | ThunderGBM: Fast GBDTs and Random Forests on GPUs · J. Mach. Learn. Res. 2020 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest |
0.4 | 1 | 2020 | ThunderGBM: Fast GBDTs and Random Forests on GPUs · J. Mach. Learn. Res. 2020 |
Data mining › text mining
sentiment analysis |
0.1 | 1 | 2021 | Enhancing SVMs with Problem Context Aware Pipeline · KDD 2021 |
Methods — techniques the papers use, named apart from their topics
reconstruction · 0.9contrastive learning · 0.9adversarial training · 0.9subproblem construction · 0.5kernel hyperparameter tuning · 0.5feature customization · 0.5data balancing · 0.5random forest · 0.4gradient boosting · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating Multi-Output GBDTs with GPUsabstractGradient Boosted Decision Trees (GBDTs) have demonstrated good performance in many data science competitions. This paper studies multidimensional output GBDT (GBDT-MO) training which can handle complex dependencies between input features and multidimensional outputs. Due to the multiple output dimensionality, the training time and memory consumption of GBDT-MO are significantly higher than those of the single-dimensional output GBDT training, which makes training GBDT-MO challenging, especially for large-scale datasets. In this paper, we propose a novel GPU-accelerated GBDT-MO training system to speed up the training while maintaining competitive model quality. Our system leverages the GPU to efficiently construct decision trees, especially by dynamically choosing efficient histogram building methods at different stages of the training, enabling scalable and efficient GBDT-MO training. Furthermore, our system supports multi-GPU training on a single machine by partitioning features across GPUs and using communication-efficient synchronization strategies, allowing further scalability for high-dimensional datasets. We evaluate the performance of our proposed GBDT-MO system on several datasets. Compared with CPU-based implementations, our system achieves a speedup ranging from 30 × to 190 ×. Moreover, our system outperforms the state-of-the-art GPU-based baselines by a speedup ranging from 1.7 × to 170 ×, while maintaining strong predictive performance compared with the state-of-the-art methods. Hanfeng Liu, Xuemei Peng, Zeyi Wen |
ICPP | 1 |
| 2025 | RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion RecognitionabstractConventional Multi-modal multi-label emotion recognition (MMER) assumes complete access to visual, textual, and acoustic modalities. However, real-world multi-party settings often violate this assumption, as non-speakers frequently lack acoustic and textual inputs, leading to a significant degradation in model performance. Existing approaches also tend to unify heterogeneous modalities into a single representation, overlooking each modality’s unique characteristics. To address these challenges, we propose RAMer (Reconstruction-based Adversarial Model for Emotion Recognition), which refines multi-modal representations by not only exploring modality commonality and specificity but crucially by leveraging reconstructed features, enhanced by contrastive learning, to overcome data incompleteness and enrich feature quality. RAMer also introduces a personality auxiliary task to complement missing modalities using modality-level attention, improving emotion reasoning. To further strengthen the model's ability to capture label and modality interdependency, we propose a stack shuffle strategy to enrich correlations between labels and modality-specific features. Experiments on three benchmarks, i.e., MEmoR, CMU-MOSEI, and M³ED, demonstrate that RAMer achieves state-of-the-art performance in dyadic and multi-party MMER scenarios. Yizhang Zhu, Hanfeng Liu, Zeyi Wen, Nan Tang 0001, Yuyu Luo |
IJCAI | 3 |
| 2021 | FastPSO: Towards Efficient Swarm Intelligence Algorithm on GPUsabstractParticle Swarm Optimization (PSO) has been widely used in various optimization tasks (e.g., neural architecture search and autonomous vehicle navigation), because it can solve non-convex optimization problems with simplicity and efficacy. However, the PSO algorithm is often time-consuming to use, especially for high-dimensional problems, which hinders its applicability in time-critical applications. In this paper, we propose novel techniques to accelerate the PSO algorithm with GPUs. To mitigate the efficiency bottleneck, we formally model the PSO optimization as a process of element-wise operations on matrices. Based on the modeling, we develop an efficient GPU algorithm to perform the element-wise operations in massively parallel using the tensor cores and shared memory. Moreover, we propose a series of novel techniques to improve our proposed algorithm, including (i) GPU resource-aware thread creation to prevent creating too many threads when the number of particles/dimensions is large; (ii) designing parallel techniques to initialize swarm particles with fast random number generation; (iii) exploiting GPU memory caching to manage swarm information instead of allocating new memory and (iv) developing a schema to support customized swarm evaluation functions. We conduct extensive experiments on four optimization applications to study the efficiency of our algorithm called “FastPSO”. Experimental results show that FastPSO consistently outperforms the existing CPU-based PSO libraries by two orders of magnitude, and transcends the existing GPU-based implementation by 5 to 7 times, while achieving better or competitive optimization results. Hanfeng Liu, Zeyi Wen, Wei Cai 0002 |
ICPP | 1 |
| 2021 | Enhancing SVMs with Problem Context Aware PipelineabstractIn recent years, many data mining practitioners have treated deep neural networks (DNNs) as a standard recipe of creating the state-of-the-art solutions. As a result, models like Support Vector Machines (SVMs) have been overlooked. While the results from DNNs are encouraging, DNNs also come with their huge number of parameters in the model and overheads in long training/inference time. SVMs have excellent properties such as convexity, good generality and efficiency. In this paper, we propose techniques to enhance SVMs with an automatic pipeline which exploits the context of the learning problem. The pipeline consists of several components including data aware subproblem construction, feature customization, data balancing among subproblems with augmentation, and kernel hyper-parameter tuner. Comprehensive experiments show that our proposed solution is more efficient, while producing better results than the other SVM based approaches. Additionally, we conduct a case study of our proposed solution on a popular sentiment analysis problem---the aspect term sentiment analysis (ATSA) task. The study shows that our SVM based solution can achieve competitive predictive accuracy to DNN (and even majority of the BERT) based approaches. Furthermore, our solution is about 40 times faster in inference and has 100 times fewer parameters than the models using BERT. Our findings can encourage more research work on conventional machine learning techniques which may be a good alternative for smaller model size and faster training/inference. Zeyi Wen, Zhishang Zhou, Hanfeng Liu, Bingsheng He, Xia Li 0007, Jian Chen 0011 |
KDD | 3 |
| 2020 | ThunderGBM: Fast GBDTs and Random Forests on GPUsabstractGradient Boosting Decision Trees (GBDTs) and Random Forests (RFs) have been used in many real-world applications. They are often a standard recipe for building state-of-the-art solutions to machine learning and data mining problems. However, training and prediction are very expensive computationally for large and high dimensional problems. This article presents an efficient and open source software toolkit called ThunderGBM which exploits the high-performance Graphics Processing Units (GPUs) for GBDTs and RFs. ThunderGBM supports classification, regression and ranking, and can run on single or multiple GPUs of a machine. Our experimental results show that ThunderGBM outperforms the existing libraries while producing similar models, and can handle high dimensional problems where existing GPU-based libraries fail. Documentation, examples, and more details about ThunderGBM are available at https://github.com/xtra-computing/thundergbm. Zeyi Wen, Hanfeng Liu, Jiashuai Shi, Qinbin Li, Bingsheng He, Jian Chen 0011 |
J. Mach. Learn. Res. | 2 |
| 2017 | Chinese Question Classification Based on Semantic Joint Features
Xia Li 0007, Hanfeng Liu, Shengyi Jiang |
NLPCC | 2 |
| 2007 | 0-SM: A fast algorithm for mining Candidate Clusters in Pattern-based ClusteringabstractUnlike traditional clustering methods that focus on grouping objects with similar values on a set of dimensions, clustering by pattern similarity finds objects that exhibit a coherent pattern of rise and fall in subspaces. Pattern-based clustering extends the concept of traditional clustering and benefits a wide range of applications, including large scale scientific data analysis, target marketing, Web usage analysis, etc. However, state-of-the-art pattern-based clustering methods (e.g., the sigma-pCluster algorithm), mining candidate clusters mostly by comparing each pair of attributes and objects, which have reduced the efficiency and makes them inappropriate for many real-life applications. This paper present a fast algorithm for mining candidate clusters. We called it Zero-Sub-Matrix. It has a better efficiency than previous algorithms. Hanfeng Liu |
CIDM | 3 |