Hanfeng Liu

dblp:00/588 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 36% Representation and self-supervised learning · 18% Information extraction and text analysis · 18%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness › multimodal robustness
missing modality robustness
0.912025
RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition · IJCAI 2025
Natural language and speech › Information extraction and text analysis › emotion recognition
multimodal emotion recognition
0.912025
RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition · IJCAI 2025
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.912025
RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition · IJCAI 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition · IJCAI 2025
Data mining › predictive modeling
classification
0.512021
Enhancing SVMs with Problem Context Aware Pipeline · KDD 2021
Data mining › predictive modeling › classification
support vector machine
0.512021
Enhancing SVMs with Problem Context Aware Pipeline · KDD 2021
Machine learning › Efficient and distributed learning › hardware acceleration
GPU training
0.412020
ThunderGBM: Fast GBDTs and Random Forests on GPUs · J. Mach. Learn. Res. 2020
Machine learning › Kernel, tree and ensemble methods › gradient boosting
gradient boosting decision tree
0.412020
ThunderGBM: Fast GBDTs and Random Forests on GPUs · J. Mach. Learn. Res. 2020
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest
0.412020
ThunderGBM: Fast GBDTs and Random Forests on GPUs · J. Mach. Learn. Res. 2020
Data mining › text mining
sentiment analysis
0.112021
Enhancing SVMs with Problem Context Aware Pipeline · KDD 2021

Methods — techniques the papers use, named apart from their topics

reconstruction · 0.9contrastive learning · 0.9adversarial training · 0.9subproblem construction · 0.5kernel hyperparameter tuning · 0.5feature customization · 0.5data balancing · 0.5random forest · 0.4gradient boosting · 0.4
YearPublicationVenuePosition
2025 Accelerating Multi-Output GBDTs with GPUs
abstract
Gradient Boosted Decision Trees (GBDTs) have demonstrated good performance in many data science competitions. This paper studies multidimensional output GBDT (GBDT-MO) training which can handle complex dependencies between input features and multidimensional outputs. Due to the multiple output dimensionality, the training time and memory consumption of GBDT-MO are significantly higher than those of the single-dimensional output GBDT training, which makes training GBDT-MO challenging, especially for large-scale datasets. In this paper, we propose a novel GPU-accelerated GBDT-MO training system to speed up the training while maintaining competitive model quality. Our system leverages the GPU to efficiently construct decision trees, especially by dynamically choosing efficient histogram building methods at different stages of the training, enabling scalable and efficient GBDT-MO training. Furthermore, our system supports multi-GPU training on a single machine by partitioning features across GPUs and using communication-efficient synchronization strategies, allowing further scalability for high-dimensional datasets. We evaluate the performance of our proposed GBDT-MO system on several datasets. Compared with CPU-based implementations, our system achieves a speedup ranging from 30 × to 190 ×. Moreover, our system outperforms the state-of-the-art GPU-based baselines by a speedup ranging from 1.7 × to 170 ×, while maintaining strong predictive performance compared with the state-of-the-art methods.
Hanfeng Liu, Xuemei Peng, Zeyi Wen
ICPP1
2025 RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition
abstract
Conventional Multi-modal multi-label emotion recognition (MMER) assumes complete access to visual, textual, and acoustic modalities. However, real-world multi-party settings often violate this assumption, as non-speakers frequently lack acoustic and textual inputs, leading to a significant degradation in model performance. Existing approaches also tend to unify heterogeneous modalities into a single representation, overlooking each modality’s unique characteristics. To address these challenges, we propose RAMer (Reconstruction-based Adversarial Model for Emotion Recognition), which refines multi-modal representations by not only exploring modality commonality and specificity but crucially by leveraging reconstructed features, enhanced by contrastive learning, to overcome data incompleteness and enrich feature quality. RAMer also introduces a personality auxiliary task to complement missing modalities using modality-level attention, improving emotion reasoning. To further strengthen the model's ability to capture label and modality interdependency, we propose a stack shuffle strategy to enrich correlations between labels and modality-specific features. Experiments on three benchmarks, i.e., MEmoR, CMU-MOSEI, and M³ED, demonstrate that RAMer achieves state-of-the-art performance in dyadic and multi-party MMER scenarios.
Yizhang Zhu, Hanfeng Liu, Zeyi Wen, Nan Tang 0001, Yuyu Luo
IJCAI3
2021 FastPSO: Towards Efficient Swarm Intelligence Algorithm on GPUs
abstract
Particle Swarm Optimization (PSO) has been widely used in various optimization tasks (e.g., neural architecture search and autonomous vehicle navigation), because it can solve non-convex optimization problems with simplicity and efficacy. However, the PSO algorithm is often time-consuming to use, especially for high-dimensional problems, which hinders its applicability in time-critical applications. In this paper, we propose novel techniques to accelerate the PSO algorithm with GPUs. To mitigate the efficiency bottleneck, we formally model the PSO optimization as a process of element-wise operations on matrices. Based on the modeling, we develop an efficient GPU algorithm to perform the element-wise operations in massively parallel using the tensor cores and shared memory. Moreover, we propose a series of novel techniques to improve our proposed algorithm, including (i) GPU resource-aware thread creation to prevent creating too many threads when the number of particles/dimensions is large; (ii) designing parallel techniques to initialize swarm particles with fast random number generation; (iii) exploiting GPU memory caching to manage swarm information instead of allocating new memory and (iv) developing a schema to support customized swarm evaluation functions. We conduct extensive experiments on four optimization applications to study the efficiency of our algorithm called “FastPSO”. Experimental results show that FastPSO consistently outperforms the existing CPU-based PSO libraries by two orders of magnitude, and transcends the existing GPU-based implementation by 5 to 7 times, while achieving better or competitive optimization results.
Hanfeng Liu, Zeyi Wen, Wei Cai 0002
ICPP1
2021 Enhancing SVMs with Problem Context Aware Pipeline
abstract
In recent years, many data mining practitioners have treated deep neural networks (DNNs) as a standard recipe of creating the state-of-the-art solutions. As a result, models like Support Vector Machines (SVMs) have been overlooked. While the results from DNNs are encouraging, DNNs also come with their huge number of parameters in the model and overheads in long training/inference time. SVMs have excellent properties such as convexity, good generality and efficiency. In this paper, we propose techniques to enhance SVMs with an automatic pipeline which exploits the context of the learning problem. The pipeline consists of several components including data aware subproblem construction, feature customization, data balancing among subproblems with augmentation, and kernel hyper-parameter tuner. Comprehensive experiments show that our proposed solution is more efficient, while producing better results than the other SVM based approaches. Additionally, we conduct a case study of our proposed solution on a popular sentiment analysis problem---the aspect term sentiment analysis (ATSA) task. The study shows that our SVM based solution can achieve competitive predictive accuracy to DNN (and even majority of the BERT) based approaches. Furthermore, our solution is about 40 times faster in inference and has 100 times fewer parameters than the models using BERT. Our findings can encourage more research work on conventional machine learning techniques which may be a good alternative for smaller model size and faster training/inference.
Zeyi Wen, Zhishang Zhou, Hanfeng Liu, Bingsheng He, Xia Li 0007, Jian Chen 0011
KDD3
2020 ThunderGBM: Fast GBDTs and Random Forests on GPUs
abstract
Gradient Boosting Decision Trees (GBDTs) and Random Forests (RFs) have been used in many real-world applications. They are often a standard recipe for building state-of-the-art solutions to machine learning and data mining problems. However, training and prediction are very expensive computationally for large and high dimensional problems. This article presents an efficient and open source software toolkit called ThunderGBM which exploits the high-performance Graphics Processing Units (GPUs) for GBDTs and RFs. ThunderGBM supports classification, regression and ranking, and can run on single or multiple GPUs of a machine. Our experimental results show that ThunderGBM outperforms the existing libraries while producing similar models, and can handle high dimensional problems where existing GPU-based libraries fail. Documentation, examples, and more details about ThunderGBM are available at https://github.com/xtra-computing/thundergbm.
Zeyi Wen, Hanfeng Liu, Jiashuai Shi, Qinbin Li, Bingsheng He, Jian Chen 0011
J. Mach. Learn. Res.2
2017 Chinese Question Classification Based on Semantic Joint Features
Xia Li 0007, Hanfeng Liu, Shengyi Jiang
NLPCC2
2007 0-SM: A fast algorithm for mining Candidate Clusters in Pattern-based Clustering
abstract
Unlike traditional clustering methods that focus on grouping objects with similar values on a set of dimensions, clustering by pattern similarity finds objects that exhibit a coherent pattern of rise and fall in subspaces. Pattern-based clustering extends the concept of traditional clustering and benefits a wide range of applications, including large scale scientific data analysis, target marketing, Web usage analysis, etc. However, state-of-the-art pattern-based clustering methods (e.g., the sigma-pCluster algorithm), mining candidate clusters mostly by comparing each pair of attributes and objects, which have reduced the efficiency and makes them inappropriate for many real-life applications. This paper present a fast algorithm for mining candidate clusters. We called it Zero-Sub-Matrix. It has a better efficiency than previous algorithms.
Hanfeng Liu
CIDM3