VLDB 2026 Research / reviewers in the wild / expert
Hua Gao
dblp:64/1565
· DBLP profile ↗
26ranked-venue papers
9as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Text-driven 3D human motion generation for pose estimation using dual-transformer architecture
Rizwan Abbas, Hua Gao, Xi Li 0001 |
Comput. Aided Des. | 2 |
| 2025 | Robust outlier detection method based on local entropy and global density
Kaituo Zhang, Bingyang Zhang, Wei Huang 0015, Hua Gao, Rongchun Wan |
Expert Syst. Appl. | 4 |
| 2024 | HashNeck is a Boosting Tool for Deep Learning to HashingabstractThe goal of hashing for image and video retrieval is to encode multimedia data into compact binary codes, allowing for efficient approximate nearest neighbor search by ensuring that similar images or videos have closely related codes in Hamming space. To improve the effectiveness of hashing, we propose introducing a classification task to assist in training the hash network and enhance the discriminability of predicted hash codes. Unlike conventional multi-task learning approaches, we propose a HashNeck structure for the classification branch that utilizes the similarities between the expected and predicted hash codes to determine whether a neuron should participate in the classification task. By only guiding neurons with correctly predicted hash codes through the classification task, we effectively resolve the conflict between the hash and classification tasks. We evaluated the effectiveness of our proposed method on benchmark image and video datasets, including ImageNet100, MS COCO, NUS-WIDE, UCF-101, and HMDB51. The experimental results on image and video retrieval tasks demonstrate that our method outperforms state-of-the-art hashing methods in terms of retrieval performance. These compelling results demonstrate the superiority of our algorithm and its potential for improving the field of deep learning to hashing. Hua Gao, Chenchen Hu, Guang Han 0002, Jiafa Mao, Wei Huang 0015, Kaiyuan Wan |
ICMR | 1 |
| 2024 | Point-level feature learning based on vision transformer for occluded person re-identificationabstractPerson re-identification is challenging due to the presence of variations in pose and occlusion, which significantly impact the matching of visual features across different camera views and pose considerable difficulty for accurate person re-identification. This paper proposes a novel method for occluded person re-identification by introducing point-level feature learning based on vision transformers. Our approach utilizes a pose estimator to detect the keypoints of the human body and employs these points to locate intermediate features. These intermediate features of keypoints are input to a pose-based transformer branch to learn point-level features. Then, we design a part-based transformer branch to learn part-level features that capture visual features of different image parts, further enhancing the discriminative power of the learned features. Additionally, we employ a global branch to learn the global-level feature by treating the person's image as a single entity. Finally, we integrate point-level, part-level, and global-level features to represent a person's features. The experimental results on occluded and partial person re-identification datasets demonstrate the effectiveness of our proposed approach in improving re-identification. Our approach shows potential for improving person re-identification in scenarios with occlusion and pose variations. Hua Gao, Chenchen Hu, Guang Han 0002, Jiafa Mao, Wei Huang 0015, Qiu Guan |
Image Vis. Comput. | 1 |
| 2023 | Harmonic enhancement using learnable comb filter for light-weight full-band speech enhancement model
Xiaohuai Le, Yiqing Guo, Xianjun Xia, Hua Gao, Yijian Xiao, Piao Ding, Shenyi Song |
INTERSPEECH | 8 |
| 2023 | A Novel Method for Identifying Bipolar Disorder Based on Diagnostic Texts
Hua Gao, Kaikai Chi |
PRCV (4) | 1 |
| 2023 | Deep Depression Detection Based on Feature Fusion and Result Fusion
Hua Gao, Kaikai Chi |
PRCV (4) | 1 |
| 2023 | Q-learning-based sequential recovery of interdependent power-communication network after cascading failures
Hua Gao |
Neural Comput. Appl. | 4 |
| 2023 | Generalized complex kernel least-mean-square algorithm with adaptive kernel widths
Zezhen Huang, Hua Gao |
Neural Comput. Appl. | 3 |
| 2021 | Box Regression-Guided Anchor-free for Robust Visual TrackingabstractThe Siamese tracker-based approach has achieved significant success in recent years. However, these approaches do not consider the different requirements for input feature in classification and regression branches. The regression branch needs feature information slightly larger than the object region, while the classification branch needs to avoid classification failure caused by the introduction of background information. In this paper, we present a novel Box Regression-Guided Anchor-free for Robust Visual Tracking. Firstly, a scale-aware regression module is designed to satisfy the feature requirements of the regression branch, which can capture feature information of various scales. Secondly, regression-guided classification module is applied to aligning the feature between the regression result and correlation feature, thereby avoiding the introduction of background information to classification branch. In addition, the new correlation operation is introduced to gain more superb correlation feature. Comparsion experimental exhibits that the proposed tracker achieves promising results in five challenging benchmark tests, including GOT-10K, OTB-2015, VOT-2018, VOT-2019 and TrackingNet, and run at an average speed of 60 FPS in real-time. Sixian Chan 0001, Xiaolong Zhou 0001, Cong Bai, Hua Gao, Shengyong Chen |
SMC | 5 |
| 2020 | gmRAD: an integrated SNP calling pipeline for genetic mapping with RADseq across a hybrid populationabstractRestriction site-associated DNA sequencing (RADseq) is a powerful technology that has been extensively applied in population genetics, phylogenetics and genetic mapping. Although many software packages are available for ecological and evolutionary studies, a few effective tools are available for extracting genotype data with RADseq for genetic mapping, a prerequisite for quantitative trait locus mapping, comparative genomics and genome scaffold assembly. Here, we present an integrated pipeline called gmRAD for generating single nucleotide polymorphism (SNP) genotypes from RADseq data, de novo, across a genetic mapping population derived by crossing two parents. As an analytical strategy, the software takes five steps to implement the whole algorithms, including clustering the first (forward) reads of each parent, building two parental references, generating parental SNP catalogs, calling SNP genotypes across all individuals and filtering the genotype data for genetic linkage mapping. All the steps can be completed with a simple command line, but they can be also performed optionally if prerequisite files are available. To validate its application, we also performed a real data analysis with RADseq data from an F1 hybrid population derived by crossing Populus deltoides and Populus simonii. The software gmRAD is freely available at https://github.com/tongchf/gmRAD. Hainan Wu, Wenguo Yang, Hua Gao, Chunfa Tong |
Briefings Bioinform. | 5 |
| 2017 | Automated industry classification with deep learningabstractIn this paper, we present a novel technique for industry classification. Leveraging EverString's API to construct a database of companies labeled with the industries to which they belong, we train a deep neural network to predict the industries of novel companies. We examine the capacity of our model to predict six-digit NAICS codes, as well as the ability of our model architecture to adapt to other industry segmentation schemas. Additionally, we investigate the ability of our model to generalize despite the presence of noise in the labels in our training set. Finally, we explore the possibility of increasing predictive precision by thresholding based on the confidence scores that our model outputs along with its predictions. We find that our approach yields six-digit NAICS code predictions that surpass the precision of gold-standard databases. Sam Wood, Rohit Muthyala, Yixing Qin, Nilaj Rukadikar, Amit Rai, Hua Gao |
IEEE BigData | 7 |
| 2017 | Action Units and Their Cross-Correlations for Prediction of Cognitive Load during DrivingabstractDriving requires the constant coordination of many body systems and full attention of the person. Cognitive distraction (subsidiary mental load) of the driver is an important factor that decreases attention and responsiveness, which may result in human error and accidents. In this paper, we present a study of facial expressions of such mental diversion of attention. First, we introduce a multi-camera database of 46 people recorded while driving a simulator in two conditions, baseline and induced cognitive load using a secondary task. Then, we present an automatic system to differentiate between the two conditions, where we use features extracted from Facial Action Unit (AU) values and their cross-correlations in order to exploit recurring synchronization and causality patterns. Both the recording and detection system are suitable for integration in a vehicle and a real-world application, e.g., an early warning system. We show that when the system is trained individually on each subject we achieve a mean accuracy and F-score of$\sim 95$percent, and for the subject independent tests$\sim 68$percent accuracy and$\sim 66$percent F-score, with person-specific normalization to handle subject dependency. Based on the results, we discuss the universality of the facial expressions of such states and possible real-world uses of the system. Anil Yüce, Hua Gao, Gabriel Louis Cuendet, Jean-Philippe Thiran |
IEEE Trans. Affect. Comput. | 2 |
| 2017 | A Regression-Based User Calibration Framework for Real-Time Gaze EstimationabstractEye movements play a very significant role in human-computer interaction (HCI) as they are natural and fast, and contain important cues for human cognitive state and visual attention. Over the last two decades, many techniques have been proposed to accurately estimate the gaze. Among these, video-based remote eye trackers have attracted much interest, since they enable nonintrusive gaze estimation. To achieve high estimation accuracies for remote systems, user calibration is inevitable in order to compensate for the estimation bias caused by person-specific eye parameters. Although several explicit and implicit user calibration methods have been proposed to ease the calibration burden, the procedure is still cumbersome and needs further improvement. In this paper, we present a comprehensive analysis of regression-based user calibration techniques. We propose a novel weighted least squares regression-based user calibration method together with a real-time cross-ratio based gaze estimation framework. The proposed system enables to obtain high estimation accuracy with minimum user effort, which leads to user-friendly HCI applications. Experimental results conducted on both simulations and user experiments show that our framework achieves a significant performance improvement over the state-of-the-art user calibration methods when only a few points are available for the calibration. Nuri Murat Arar, Hua Gao, Jean-Philippe Thiran |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Profit Maximizing Route Recommendation for Vehicle Sharing Requests
Hua Gao, Heli Sun, Xiaolin Jia |
APWeb (2) | 3 |
| 2015 | Towards Convenient Calibration for Cross-Ratio Based Gaze EstimationabstractEye gaze movements are considered as a salient modality for human computer interaction applications. Recently, cross-ratio (CR) based eye tracking methods have attracted increasing interest because they provide remote gaze estimation using a single uncalibrated camera. However, due to the simplification assumptions in CR-based methods, their performance is lower than the model-based approaches [8]. Several efforts have been made to improve the accuracy by compensating for the assumptions with subject specific calibration. This paper presents a CR-based automatic gaze estimation system that accurately works under natural head movements. A subject-specific calibration method based on regularized least-squares regression (LSR) is introduced for achieving higher accuracy compared to other state-of-the-art calibration methods. Experimental results also show that the proposed calibration method generalizes better when fewer calibration points are used. This enables user friendly applications with minimum calibration effort without sacrificing too much accuracy. In addition, we adaptively fuse the estimation of the point of regard (PoR) from both eyes based on the visibility of eye features. The adaptive fusion scheme reduces accuracy error by around 20% and also increases the estimation coverage under natural head movements. Nuri Murat Arar, Hua Gao, Jean-Philippe Thiran |
WACV | 2 |
| 2014 | Detecting emotional stress from facial expressions for driving safetyabstractMonitoring the attentive and emotional status of the driver is critical for the safety and comfort of driving. In this work a real-time non-intrusive monitoring system is developed, which detects the emotional states of the driver by analyzing facial expressions. The system considers two negative basic emotions, anger and disgust, as stress related emotions. We detect an individual emotion in each video frame and the decision on the stress level is made on sequence level. Experimental results show that the developed system operates very well on simulated data even with generic models. An additional pose normalization step reduces the impact of pose mismatch due to camera setup and pose variation, and hence improves the detection accuracy further. Hua Gao, Anil Yüce, Jean-Philippe Thiran |
ICIP | 1 |
| 2014 | Extending explicit shape regression with mixed feature channels and pose priorsabstractFacial feature detection offers a wide range of applications, e.g. in facial image processing, human computer interaction, consumer electronics, and the entertainment industry. These applications impose two antagonistic key requirements: high processing speed and high detection accuracy. We address both by expanding upon the recently proposed explicit shape regression [1] to (a) allow usage and mixture of different feature channels, and (b) include head pose information to improve detection performance in non-cooperative environments. Using the publicly available “wild” datasets LFW [10] and AFLW [11], we show that using these extensions outperforms the baseline (up to 10% gain in accuracy at 8% IOD) as well as other state-of-the-art methods. Matthias Richter 0003, Hua Gao, Hazim Kemal Ekenel |
WACV | 2 |
| 2012 | Face Alignment Using a Ranking Model based on Regression TreesabstractIn this work, we exploit the regression trees-based ranking model, which has been successfully applied in the domain of web-search ranking, to build appearance models for face alignment. The model is an ensemble of regression trees which is learned with gradient boosting. The MCT (Modified Census Transform) as well as its unbinarized version PCT (Pseudo Census Transform) are used as features due to their robustness to illumination changes. To avoid the overfitting problem in gradient boosting, we use random trees to initialize the boosting. The Nelder Mead’s simplex method is applied for fitting the learned model. We compare the proposed regression trees-based pointwise ranking model to pairwise ranking model. Experiments show that the proposed model improves both robustness and accuracy for face alignment. Hua Gao, Hazim Kemal Ekenel, Rainer Stiefelhagen |
BMVC | 1 |
| 2012 | A ranking model for face alignment with Pseudo Census Transform
Hua Gao, Hazim Kemal Ekenel, Rainer Stiefelhagen |
ICPR | 1 |
| 2012 | Multi-view facial expression recognition using local appearance features
Nikolas Hesse, Tobias Gehrig, Hua Gao, Hazim Kemal Ekenel |
ICPR | 3 |
| 2011 | Boosting Pseudo Census Transform Features for Face AlignmentabstractFace alignment using deformable face model has attracted broad interest in recent years for its wide range of applications in facial analysis. Previous work has shown that discriminative deformable models have better generalization capacity compared to generative models [8, 9]. In this paper, we present a new discriminative face model based on boosting pseudo census transform features. This feature is considered to be less sensitive to illumination changes, which yields a more robust alignment algorithm. The alignment is based on maximizing the scores of boosted strong classifier, which indicate whether the current alignment is a correct or incorrect one. The proposed approach has been evaluated extensively on several databases. The experimental results show that our approach generalizes better on unseen data compared to the Haar feature-based approach. Moreover, its training procedure is much faster due to the low dimensionality of the configuration space of the proposed feature. Hua Gao, Hazim Kemal Ekenel, Mika Fischer, Rainer Stiefelhagen |
BMVC | 1 |
| 2010 | Multi-resolution Local Appearance-Based Face VerificationabstractFacial analysis based on local regions/blocks usually outperforms holistic approaches because it is less sensitive to local deformations and occlusions. Moreover, modeling local features enables us to avoid the problem of high dimensionality of feature space. In this paper, we model the local face blocks with Gabor features and project them into a discriminant identity space. The similarity score of a face pair is determined by fusion of the local classifiers. To acquire complementary information in different scales of face images, we integrate the local decisions from various image resolutions. The proposed multi-resolution block based face verification system is evaluated on the experiment 4 of Face Recognition Grand Challenge (FRGC) version 2.0. We obtained 92.5% verification [email protected]% FAR, which is the highest performance reported on this experiment so far in the literature. Hua Gao, Hazim Kemal Ekenel, Mika Fischer, Rainer Stiefelhagen |
ICPR | 1 |
| 2008 | Face recognition for smart interactionsabstractIn this paper, face recognition systems that have been developed for smart interactions at the interACT Research Center is presented. The face recognition efforts at the interACT Research Center consist of development of a fast and robust face recognition algorithm and fully automatic face recognition systems that can be deployed for real-life smart interaction applications. The face recognition algorithm is based on appearances of local facial regions that are represented with discrete cosine transform coefficients. Many fully automatic face recognition systems have been developed based on this algorithm. Among these systems two of the portable ones will be shown as interactive demos. Moreover, demo videos will be shown for the other systems. Hazim Kemal Ekenel, Mika Fischer, Hua Gao, Lorant Szasz-Toth, Rainer Stiefelhagen |
FG | 3 |
| 2007 | Face Recognition for Smart InteractionsabstractIn this paper an overview of face recognition research activities at the interACT Research Center is given. The face recognition efforts at the interACT Research Center consist of development of a fast and robust face recognition algorithm and fully automatic face recognition systems that can be deployed for real-life smart interaction applications. The face recognition algorithm is based on appearances of local facial regions that are represented with discrete cosine transform coefficients. Three fully automatic face recognition systems have been developed that are based on this algorithm. The first one is the "door monitoring system" that observes the entrance of a room and identifies the subjects while they are entering the room. The second one is the "portable face recognition system" that aims at environment-free face recognition and recognizes the user of a machine. The third system, "3D face recognition system", performs fully automatic face recognition on 3D range data. Hazim Kemal Ekenel, Johannes Stallkamp, Hua Gao, Mika Fischer, Rainer Stiefelhagen |
ICME | 3 |
| 2007 | 3-D Face Recognition Using Local Appearance-Based ModelsabstractIn this paper, we present a local appearance-based approach for 3-D face recognition. In the proposed algorithm, we first register the 3-D point clouds to provide a dense correspondence between faces. Afterwards, we analyze two mapping techniques—the closest-point mapping and the ray-casting mapping, to construct depth images from the corresponding well-registered point clouds. The depth images that are obtained are then divided into local regions where the discrete cosine transformation is performed to extract local information. The local features are combined at the feature level for classification. Experimental results on the FRGC version 2.0 face database show that the proposed algorithm performs superior to the well-known face recognition algorithms. Hazim Kemal Ekenel, Hua Gao, Rainer Stiefelhagen |
IEEE Trans. Inf. Forensics Secur. | 2 |