VLDB 2026 Research / reviewers in the wild / expert
Haoxiang Wang 0002
dblp:155/8045-2
· DBLP profile ↗
14ranked-venue papers
1as first author
8since 2021 · last 2023
0000-0003-4474-838XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Attention Based Relation Network for Facial Action Units RecognitionabstractFacial action unit (AU) recognition is essential to facial expression analysis. Since there are highly positive or negative correlations between AUs, some existing AU recognition works have focused on modeling AU relations. However, previous relationship-based approaches typically embed predefined rules into their models and ignore the impact of various AU relations in different crowds. In this paper, we propose a novel Attention Based Relation Network (ABRNet) for AU recognition, which can automatically capture AU relations without unnecessary or even disturbing predefined rules. ABRNet uses several relation learning layers to automatically capture different AU relations. The learned AU relation features are then fed into a self-attention fusion module, which aims to refine individual AU features with attention weights to enhance the feature robustness. Furthermore, we propose an AU relation dropout strategy and AU relation loss (AUR-Loss) to better model AU relations, which can further improve AU recognition. Extensive experiments show that our approach achieves state-of-the-art performance on the DISFA and DISFA+ datasets. Haoxiang Wang 0002, Jiawang Liu |
ICASSP | 2 |
| 2022 | DTE-Net: Dual Temporal Excitation Network for Video Violence RecognitionabstractVideo-based violence recognition has become a crucial topic owing to the development of surveillance cameras. However, with the extra temporal dimension and no precision range of violent video data, violence recognition is a challenging problem. In this study, we propose a dual temporal excitation network (DTE-Net) consisting of a shift temporal adaptive module (STAM) and a sparse object interaction transformer (SOI-Tr) module. The STAM extracts coarse-grained local and global temporal information by fusing shift module with temporal adaptive modeling module. The SOI-Tr module utilizes important object attention to excite fine-grained global temporal representation reasoning. In addition, we create a multi-class violence (MCV) dataset of video clips extracted from real-world scenes to address the limitation of poorly diversified categories in most existing violence datasets. Finally, we also conduct extensive experiments on five violence datasets, including the MCV, and the results show that our network outperforms state-of-the-art performance. Wenwei Yan, Haoxiang Wang 0002, Jun Xuan, Yuxuan Tang, Aihua Mao |
ICME | 2 |
| 2022 | Graph based emotion recognition with attention pooling for variable-length utterances
Jiawang Liu, Haoxiang Wang 0002 |
Neurocomputing | 2 |
| 2022 | Convolution by Multiplication: Accelerated Two- Stream Fourier Domain Convolutional Neural Network for Facial Expression RecognitionabstractFacial expression plays an important role in human communication as a type of nonverbal language and has been widely used in various areas such as psychology, human-computer interaction and robotics. Nowadays, convolutional neural network is a promising approach for facial expression recognition. However, convolutional layers can be time-consuming and computationally expensive because a large number of parameters participate in the calculations and need to be updated during training. To improve the performance of deep neural network in facial expression recognition and accelerate training and calculation, we propose a novel framework which adopts efficient element-wise multiplication to replace traditional convolution. To disentangle reliable feature representation for more effective recognition and further enhance the recognition performance while maintaining the efficiency, we propose a representation scheme which can retain informative feature components while removing unreliable ones in Fourier domain based on the proposed multiplication framework. Extensive comparison and ablation studies are conducted on several benchmark datasets, which shows the efficiency and effectiveness of the proposed model. Xingming Zhang 0001, Xiangyuan Lan, Haoxiang Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Cohesive Multi-Modality Feature Learning and Fusion for COVID-19 Patient Severity PredictionabstractThe outbreak of coronavirus disease (COVID-19) has been a nightmare to citizens, hospitals, healthcare practitioners, and the economy in 2020. The overwhelming number of confirmed cases and suspected cases put forward an unprecedented challenge to the hospital's capacity of management and medical resource distribution. To reduce the possibility of cross-infection and attend a patient according to his severity level, expertly diagnosis and sophisticated medical examinations are often required but hard to fulfil during a pandemic. To facilitate the assessment of a patient's severity, this paper proposes a multi-modality feature learning and fusion model for end-to-end covid patient severity prediction using the blood test supported electronic medical record (EMR) and chest computerized tomography (CT) scan images. To evaluate a patient's severity by the co-occurrence of salient clinical features, the High-order Factorization Network (HoFN) is proposed to learn the impact of a set of clinical features without tedious feature engineering. On the other hand, an attention-based deep convolutional neural network (CNN) using pre-trained parameters are used to process the lung CT images. Finally, to achieve cohesion of cross-modality representation, we design a loss function to shift deep features of both-modality into the same feature space which improves the model's performance and robustness when one modality is absent. Experimental results demonstrate that the proposed multi-modality feature learning and fusion model achieves high performance in an authentic scenario. Jinzhao Zhou, Xingming Zhang 0001, Ziwei Zhu 0005, Xiangyuan Lan, Lunkai Fu, Haoxiang Wang 0002, Hanchun Wen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Graph Isomorphism Network for Speech Emotion Recognition
Jiawang Liu, Haoxiang Wang 0002 |
Interspeech | 2 |
| 2021 | A Speech Emotion Recognition Framework for Better Discrimination of Confusions
Jiawang Liu, Haoxiang Wang 0002 |
Interspeech | 2 |
| 2021 | Facial Expression Recognition Using Frequency Neural NetworkabstractFacial expression recognition has become a newly-emerging topic in recent decades, which has important value in the field of human-computer interaction. In this paper, we present a deep learning based approach, named frequency neural network (FreNet), for facial expression recognition. Different from convolutional neural network in spatial domain, FreNet inherits the advantages of processing image in frequency domain, such as efficient computation and spatial redundancy elimination. First, we propose the learnable multiplication kernel and construct multiple multiplication layers to learn features in frequency domain. Second, a summarization layer is proposed following multiplication layers to further yield high-level features. Third, based on the property of discrete cosine transform (DCT), we utilize multiplication layers and summarization layer to construct the Basic-FreNet, which can yield high-level features on the widely used DCT feature. Finally, to further achieve better performance on Basic-FreNet, we propose the Block-FreNet in which the weight-shared multiplication kernel is designed for feature learning and the block sub-sampling is designed for dimension reduction. The experimental results show that the Block-FreNet not only achieves superior performance, but also greatly reduces the computational cost. To our best knowledge, the proposed approach is the first attempt to fill in the blank of frequency based deep learning model for facial expression recognition. Xingming Zhang 0001, Xiping Hu, Siqi Wang 0001, Haoxiang Wang 0002 |
IEEE Trans. Image Process. | 5 |
| 2020 | TEAN: Timeliness enhanced attention network for session-based recommendation
Dongpei Chen, Xingming Zhang 0001, Haoxiang Wang 0002 |
Neurocomputing | 3 |
| 2019 | Modeling Missing Data Based on Neural Fuzzy Inference for Implicit RecommendationabstractAs implicit feedback can be tracked automatically and is easy to collect, the implicit recommendation attracts more researcher attentions. However, the uncertainty of the implicit feedback meaning poses a great challenge to the implicit recommendation. Especially for the missing data, we are not sure whether the users dislike or just have not seen the items. It may lead to bias of predictions. In this paper, we propose Neural Fuzzy Inference based on User preference and Item popularity (UI-NFI) algorithm to model the missing data in implicit recommendation. First, we use fuzzy set theory to represent user preference and item popularity that get from the history interactions and side information. Furthermore, neural fuzzy inference is proposed to predict the exposure possibility of missing data. Based on the fuzzy inference model, UI-NFI and matrix factorization model perform joint learning to predict. Experimental results show that our model has better performance compared to the other implicit recommendation algorithms. Xingming Zhang 0001, Haoxiang Wang 0002 |
ICTAI | 3 |
| 2019 | A deep variational matrix factorization method for recommendation on large scale sparse dataset
Xingming Zhang 0001, Haoxiang Wang 0002, Dongpei Chen |
Neurocomputing | 3 |
| 2019 | Facial expression recognition via region-based convolutional fusion network
Yingsheng Ye, Xingming Zhang 0001, Yubei Lin, Haoxiang Wang 0002 |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | A Research on Fast Face Feature Points Detection on Smart Mobile DevicesabstractWe explore how to leverage the performance of face feature points detection on mobile terminals from 3 aspects. First, we optimize the models used in SDM algorithms via PCA and Spectrum Clustering. Second, we propose an evaluation criterion using Linear Discriminative Analysis to choose the best local feature descriptions which plays a critical role in feature points detection. Third, we take advantage of multicore architecture of mobile terminal and parallelize the optimized SDM algorithm to improve the efficiency further. The experiment observations show that our final accomplished GPC‐SDM (improved Supervised Descent Method using spectrum clustering, PCA, and GPU acceleration) suppresses the memory usage, which is beneficial and efficient to meet the real‐time requirements. Xiaohe Li, Xingming Zhang 0001, Haoxiang Wang 0002 |
Wirel. Commun. Mob. Comput. | 3 |
| 2008 | Service-oriented approach to collaborative visualizationabstractAbstract This paper presents a new service‐oriented approach to the design and implementation of visualization systems in a Grid computing environment. The approach evolves the traditional dataflow visualization system, based on processes communicating via shared memory or sockets, into an environment in which visualization Web services can be linked in a pipeline using the subscription and notification services available in Globus Toolkit 4. A specific aim of our design is to support collaborative visualization, allowing a geographically distributed research team to work collaboratively on visual analysis of data. A key feature of the system is the use of a formal description of the visualization pipeline, using the skML language first developed in the gViz e‐Science project. This description is shared by all collaborators in a session. In co‐operation with the e‐Viz project, we generate user interfaces for the visualization services automatically from the skML description. The new system is called notification‐service‐based collaborative visualization. A simple prototype has been built and is used to illustrate the concepts. Copyright © 2008 John Wiley & Sons, Ltd. Haoxiang Wang 0002, Ken Brodlie, James W. Handley, Jason D. Wood |
Concurr. Comput. Pract. Exp. | 1 |