VLDB 2026 Research / reviewers in the wild / expert
Haimin Zhang 0001
dblp:25/8840-1
· DBLP profile ↗
26ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0002-0021-3634ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MEGG: replay via maximally extreme GGscore in incremental learning for neural recommendation modelsabstractAbstract Neural collaborative filtering (NCF)-based recommendation models have been widely adopted in practical recommender systems due to their effectiveness. However, these models are typically developed under the static deep learning paradigm, where training is conducted on fixed datasets with the implicit assumption of a static data distribution. This approach is ill-suited for dynamic environments, such as those encountered in real-world platforms, where user preferences and collaborative filtering patterns evolve continuously. To address this limitation, incremental learning-a paradigm designed to integrate new knowledge while preserving previously learned information-emerges as a promising alternative. Despite its potential, the direct application of conventional incremental learning methods, which are prevalent in domains like computer vision and natural language processing, is hindered by unique challenges in recommender systems. These include the distinct task paradigm, data complexity, and sparsity issues. Moreover, existing incremental learning approaches tailored for neural recommendation models remain scarce and often suffer from limited generalizability. To bridge this gap, we propose an innovative experience replay-based incremental learning framework specifically designed for neural recommendation models, termed Replay Samples with Maximally Extreme GGscore (MEGG). At the core of MEGG is a novel metric, the GGscore, which quantifies the influence of individual samples on model training. By selectively replaying samples with the most extreme GGscores, our method effectively mitigates catastrophic forgetting, thereby maintaining high predictive performance over time. A key advantage of MEGG lies in its data-centric nature, which renders it agnostic to the underlying model architecture. This ensures broad applicability across various neural recommendation models and seamless integration with existing incremental learning frameworks to further enhance performance. Extensive experiments conducted on three neural recommendation models across four benchmark datasets demonstrate the superior effectiveness of MEGG compared to state-of-the-art methods. Furthermore, additional evaluations highlight its scalability, efficiency, and robustness. The implementation of MEGG will be made publicly available upon acceptance. Yunxiao Shi, Shuo Yang 0006, Haimin Zhang 0001, Li Wang 0064, Yongze Wang, Qiang Wu 0001, Min Xu 0001 |
Data Min. Knowl. Discov. | 3 |
| 2026 | Differential Encoding for Improved Representation Learning Over GraphsabstractCombining the message-passing paradigm with the global attention mechanism has emerged as an effective framework for learning over graphs. The message-passing paradigm and the global attention mechanism basically generate embeddings of nodes by taking the sum of information from a node's local neighbourhood and from the entire graph, respectively. However, this simple summation aggregation approach fails to distinguish between the information from a node itself or from the node's neighbours. Therefore, there exists information lost at each layer of embedding generation, and this information lost could be accumulated and become more serious in deeper model layers. In this paper, we present a differential encoding method to address the issue of information lost. Instead of simply taking the sum to aggregate local or global information, we explicitly encode the difference between the information from a node itself and that from the node's local neighbours (or from the rest of the entire graph nodes). The obtained differential encoding is then combined with the original aggregated representation to generate the updated node embedding. By combining differential encodings, the representational ability of generated node embeddings is improved, and therefore the model performance is improved. The differential encoding method is empirically evaluated on different graph tasks on seven benchmark datasets. The results show that it is a general method that improves the message-passing update and the global attention update, advancing the state-of-the-art performance for graph representation learning on these benchmark datasets. Haimin Zhang 0001, Jiahao Xia 0001, Min Xu 0001 |
IEEE Trans. Big Data | 1 |
| 2026 | Beyond KAN: Introducing KarSein for Adaptive High-Order Feature Interaction Modeling in CTR PredictionabstractModeling high-order feature interactions is crucial for Click-Through Rate (CTR) prediction, yet traditional approaches typically predefine a maximum interaction order and exhaustively enumerate feature combinations up to that order. This paradigm depends heavily on prior domain knowledge to delimit the interaction space and incurs substantial computational overhead. As a result, conventional CTR models face a persistent tension between enriching representations with complex high-order interactions and keeping computation tractable. To address this dual challenge, this study introduces the Kolmogorov–Arnold Represented Sparse Efficient Interaction Network (KarSein). Drawing inspiration from the learnable activation mechanism in the Kolmogorov–Arnold Network (KAN), KarSein leverages this mechanism to adaptively transform low-order basic features into high-order feature interactions, offering a novel approach to feature interaction modeling. KarSein extends the capabilities of KAN by introducing a more efficient architecture that significantly reduces computational costs while accommodating 2D embedding vectors as feature inputs. Furthermore, it overcomes the limitation of KAN’s its inability to spontaneously capture multiplicative relationships among features. Extensive experiments highlight the superiority of KarSein, demonstrating its ability to surpass not only the vanilla implementation of KAN in CTR prediction tasks but also other baseline methods. Remarkably, KarSein achieves exceptional predictive accuracy while maintaining a highly compact parameter size and minimal computational overhead. Moreover, KarSein retains the key advantages of KAN, such as strong interpretability and structural sparsity. As the first systematic adaptation of KAN to CTR prediction, KarSein offers a practical, parameter-efficient, and interpretable alternative for modeling complex feature interactions in large-scale recommendation systems. Yunxiao Shi, Wujiang Xu, Haimin Zhang 0001, Qiang Wu 0001, Min Xu 0001 |
ACM Trans. Inf. Syst. | 3 |
| 2025 | STPM: Spatial-Temporal Point Mamba for Activity Recognition Using mmWave Radar Point CloudsabstractHuman activity recognition using millimeter-wave radar point clouds has emerged as a promising visual privacy-preserving sensing paradigm, transitioning from multi-domain Doppler analysis to point cloud-based methods for richer spatial information. However, existing approaches face critical challenges in modeling temporal dependencies between consecutive frames and simultaneously capturing both local geometric structures and global spatial relationships. To address these challenges, we propose STPM (Spatial-Temporal Point Mamba), a novel framework that extends traditional State Space Models through three key innovations: (1) a bidirectional selective mechanism that captures comprehensive temporal dependencies while maintaining linear memory complexity, (2) a queue-based temporal processing strategy with theoretical guarantees for preventing error accumulation, and (3) a hierarchical grouping strategy that effectively models both local geometric details and global spatial contexts. Through extensive evaluations on RadHAR and MM-Fi datasets, STPM achieves state-of-the-art performance with 98.14% and 95.69% accuracy respectively, while reducing memory consumption by 35% compared to transformer-based alternatives. Extensive experiments demonstrate its effectiveness in distinguishing semantically similar but functionally distinct motions. Yingru Chen, Haimin Zhang 0001, Min Xu 0001 |
ICME | 3 |
| 2025 | Mitigating Knowledge Discrepancies among Multiple Datasets for Task-agnostic Unified Face AlignmentabstractAbstract Despite the similar structures of human faces, existing face alignment methods cannot learn unified knowledge from multiple datasets with different landmark annotations. The limited training samples in a single dataset commonly result in fragile robustness in this field. To mitigate knowledge discrepancies among different datasets and train a task-agnostic unified face alignment (TUFA) framework, this paper presents a strategy to unify knowledge from multiple datasets. Specifically, we calculate a mean face shape for each dataset. To explicitly align these mean shapes on an interpretable plane based on their semantics, each shape is then incorporated with a group of semantic alignment embeddings. The 2D coordinates of these aligned shapes can be viewed as the anchors of the plane. By encoding them into structure prompts and further regressing the corresponding facial landmarks using image features, a mapping from the plane to the target faces is finally established, which unifies the learning target of different datasets. Consequently, multiple datasets can be utilized to boost the generalization ability of the model. The successful mitigation of discrepancies also enhances the efficiency of knowledge transferring to a novel dataset, significantly boosts the performance of few-shot face alignment. Additionally, the interpretable plane endows TUFA with a task-agnostic characteristic, enabling it to locate landmarks unseen during training in a zero-shot manner. Extensive experiments are carried on seven benchmarks and the results demonstrate an impressive improvement in face alignment brought by knowledge discrepancies mitigation. The code is available at https://github.com/Jiahao-UTS/TUFA Jiahao Xia 0001, Min Xu 0001, Wenjian Huang 0001, Jianguo Zhang 0001, Haimin Zhang 0001, Chunxia Xiao |
Int. J. Comput. Vis. | 5 |
| 2025 | Passive Human Tracking With WiFi Point CloudsabstractIntegrated sensing and communication (ISAC) technology empowers WiFi to function as both sensors for wireless sensing and communication devices for data exchange. Currently, achieving accurate object tracking with commercial WiFi devices is still challenging due to the limited bandwidth, a small number of antennas, and the clock asynchronization in a bi-static setup. Many existing methods achieve tracking only via extracting a dominant Doppler frequency shift (DFS) from a moving person. However, since the human body is nonrigid, various body parts generate different DFSs, and different subcarriers can exhibit varying Doppler characteristics in a multipath environment. This work presents WiDFS2.0, an enhanced real-time tracking scheme that leverages the micro-Doppler effect to extract multiple signal features from various body parts of a moving person, represented as WiFi point clouds. Each point cloud consists of Doppler, Angle of Arrival, Range, and signal-to-noise ratio. We design a novel signal processing chain to extract the WiFi point clouds. Then, we refine these point clouds and implement an extended Kalman filter-based algorithm to track the person’s trajectory. Our experiments demonstrate that WiDFS2.0 can achieve real-time tracking with a median position error of 0.55 m, while determining the presence of a moving person with over 98% accuracy during tracking. Jian (Andrew) Zhang, Haimin Zhang 0001, Min Xu 0001, Y. Jay Guo |
IEEE Internet Things J. | 3 |
| 2024 | Enhancing Retrieval and Managing Retrieval: A Four-Module Synergy for Improved Quality and Efficiency in RAG SystemsabstractRetrieval-augmented generation (RAG) techniques leverage the in-context learning capabilities of large language models (LLMs) to produce more accurate and relevant responses. Originating from the simple ‘retrieve-then-read’ approach, the RAG framework has evolved into a highly flexible and modular paradigm. A critical component, the Query Rewriter module, enhances knowledge retrieval by generating a search-friendly query. This method aligns input questions more closely with the knowledge base. Our research identifies opportunities to enhance the Query Rewriter module to Query Rewriter+ by generating multiple queries to overcome the Information Plateaus associated with a single query and by rewriting questions to eliminate Ambiguity, thereby clarifying the underlying intent. We also find that current RAG systems exhibit issues with Irrelevant Knowledge; to overcome this, we propose the Knowledge Filter. These two modules are both based on the instruction-tuned Gemma-2B model, which together enhance response quality. The final identified issue is Redundant Retrieval; we introduce the Memory Knowledge Reservoir and the Retriever Trigger to solve this. The former supports the dynamic expansion of the RAG system’s knowledge base in a parameter-free manner, while the latter optimizes the cost for accessing external knowledge, thereby improving resource utilization and response efficiency. These four RAG modules synergistically improve the response quality and efficiency of the RAG system. The effectiveness of these modules has been validated through experiments and ablation studies across six common QA datasets. The source code can be accessed at https://github.com/Ancientshi/ERM4. Yunxiao Shi, Xing Zi, Zijing Shi, Haimin Zhang 0001, Qiang Wu 0001, Min Xu 0001 |
ECAI | 4 |
| 2024 | Center-bridged Interaction Fusion for hyperspectral and LiDAR classification
Lu Huo, Jiahao Xia 0001, Leijie Zhang, Haimin Zhang 0001, Min Xu 0001 |
Neurocomputing | 4 |
| 2024 | Unsupervised Part Discovery via Dual Representation AlignmentabstractObject parts serve as crucial intermediate representations in various downstream tasks, but part-level representation learning still has not received as much attention as other vision tasks. Previous research has established that Vision Transformer can learn instance-level attention without labels, extracting high-quality instance-level representations for boosting downstream tasks. In this paper, we achieve unsupervised part-specific attention learning using a novel paradigm and further employ the part representations to improve part discovery performance. Specifically, paired images are generated from the same image with different geometric transformations, and multiple part representations are extracted from these paired images using a novel module, named PartFormer. These part representations from the paired images are then exchanged to improve geometric transformation invariance. Subsequently, the part representations are aligned with the feature map extracted by a feature map encoder, achieving high similarity with the pixel representations of the corresponding part regions and low similarity in irrelevant regions. Finally, the geometric and semantic constraints are applied to the part representations through the intermediate results in alignment for part-specific attention learning, encouraging the PartFormer to focus locally and the part representations to explicitly include the information of the corresponding parts. Moreover, the aligned part representations can further serve as a series of reliable detectors in the testing phase, predicting pixel masks for part discovery. Extensive experiments are carried out on four widely used datasets, and our results demonstrate that the proposed method achieves competitive performance and robustness due to its part-specific attention. Jiahao Xia 0001, Wenjian Huang 0001, Min Xu 0001, Jianguo Zhang 0001, Haimin Zhang 0001, Ziyu Sheng, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Vital Sign Monitoring in Dynamic Environment via mmWave Radar and Camera FusionabstractContact-free vital sign monitoring, which uses wireless signals for recognizing human vital signs (i.e, breath and heartbeat), is an attractive solution to health and security. However, the subject’s body movement and the change in actual environments can result in inaccurate frequency estimation of heartbeat and respiratory. In this paper, we propose a robust mmWave radar and camera fusion system for monitoring vital signs, which can perform consistently well in dynamic scenarios, e.g., when some people move around the subject to be tracked, or a subject waves his/her arms and marches on the spot. Three major processing modules are developed in the system, to enable robust sensing. First, we utilize a camera to assist a mmWave radar to accurately localize the subjects of interest. Second, we exploit the calculated subject position to form transmitting and receiving beamformers, which can improve the reflected power from the targets and weaken the impact of dynamic interference. Third, we propose a weighted multi-channel Variational Mode Decomposition (WMC-VMD) algorithm to separate the weak vital sign signals from the dynamic ones due to subject’s body movement. Experimental results show that, the 90th percentile errors in respiration rate (RR) and heartbeat rate (HR) are less than 0.5 RPM (respirations per minute) and 6 BPM (beats per minute), respectively. Jian (Andrew) Zhang, Haimin Zhang 0001, Min Xu 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Towards High-Quality Photorealistic Image Style TransferabstractPreserving important textures of the content image and achieving prominent style transfer results remains a challenge in the field of image style transfer. This challenge arises from the entanglement between color and texture during the style transfer process. To address this challenge, we propose an end-to-end network that incorporates adaptive weighted least squares (AWLS) filter, iterative least squares (ILS) filter, and channel separation. Given a content image ($\mathcal {C}$) and a reference style image ($\mathcal {S}$), we begin by separating the RGB channels and utilizing ILS filter to decompose them into structure and texture layers. We then perform style transfer on the structural layers using WCT$^{2}$(incorporating wavelet pooling and unpooling techniques for whitening and coloring transforms) in the R, G, and B channels, respectively. We address the texture distortion caused by WCT$^{2}$with a texture enhancing (TE) module in the structural layer. Furthermore, we propose an estimating and compensating for the structure loss (ECSL) module. In the ECSL module, with the AWLS filter and the ILS filter, we estimate the texture loss caused by TE, convert the loss of the structural layer to the loss of the texture layer, and compensate for the loss in the texture layer. The final structural layer and the texture layer are merged into the channel style transfer results in the separated R, G, and B channels into the final style transfer result. Thereby, this enables a more complete texture preservation and a significant style transfer process. To evaluate our method, we utilize quantitative experiments using various metrics, including NIQE, AG, SSIM, PSNR, and a user study. The experimental results demonstrate the superiority of our approach over the previous state-of-the-art methods. Haimin Zhang 0001, Gang Fu 0003, Caoqing Jiang, Fei Luo 0004, Chunxia Xiao, Min Xu 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | SSFG: Stochastically Scaling Features and Gradients for Regularizing Graph Convolutional NetworksabstractGraph convolutional networks (GCNs) have been successfully applied in various graph-based tasks. In a typical graph convolutional layer, node features are updated by aggregating neighborhood information. Repeatedly applying graph convolutions can cause the oversmoothing issue, i.e., node features at deep layers converge to similar values. Previous studies have suggested that oversmoothing is one of the major issues that restrict the performance of GCNs. In this article, we propose a stochastic regularization method to tackle the oversmoothing problem. In the proposed method, we stochastically scale features and gradients (SSFG) by a factor sampled from a probability distribution in the training procedure. By explicitly applying a scaling factor to break feature convergence, the oversmoothing issue is alleviated. We show that applying stochastic scaling at the gradient level is complementary to that applied at the feature level to improve the overall performance. Our method does not increase the number of trainable parameters. When used together with ReLU, our SSFG can be seen as a stochastic ReLU activation function. We experimentally validate our SSFG regularization method on three commonly used types of graph networks. Extensive experimental results on seven benchmark datasets for four graph-based tasks demonstrate that our SSFG regularization is effective in improving the overall performance of the baseline graph networks. The code is available at https://github.com/vailatuts/SSFG-regularization. Haimin Zhang 0001, Min Xu 0001, Guoqiang Zhang 0003, Kenta Niwa |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Learning Graph Representations Through Learning and Propagating Edge FeaturesabstractGraph convolutional networks have achieved considerable success in various graph domain tasks. Recently, numerous types of graph convolutional networks have been developed. A typical rule for learning a node's feature in these graph convolutional networks is to aggregate node features from the node's local neighborhood. However, in these models, the interrelation information between adjacent nodes is not well-considered. This information could be helpful to learn improved node embeddings. In this article, we present a graph representation learning framework that generates node embeddings through learning and propagating edge features. Instead of aggregating node features from a local neighborhood, we learn a feature for each edge and update a node's representation by aggregating local edge features. The edge feature is learned from the concatenation of the edge's starting node feature, the input edge feature, and the edge's end node feature. Unlike node feature propagation-based graph networks, our model propagates different features from a node to its neighbors. In addition, we learn an attention vector for each edge in aggregation, enabling the model to focus on important information in each feature dimension. By learning and aggregating edge features, the interrelation between a node and its neighboring nodes is integrated in the aggregated feature, which helps learn improved node embeddings in graph representation learning. Our model is evaluated on graph classification, node classification, graph regression, and multitask binary graph classification on eight popular datasets. The experimental results demonstrate that our model achieves improved performance compared with a wide variety of baseline models. Haimin Zhang 0001, Jiahao Xia 0001, Guoqiang Zhang 0003, Min Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Robust Face Alignment via Inherent Relation Learning and Uncertainty EstimationabstractHuman tends to locate the facial landmarks with heavy occlusion by their relative position to the easily identified landmarks. The clue is defined as the landmark inherent relation while it is ignored by most existing methods. In this paper, we present Dynamic Sparse Local Patch Transformer (DSLPT), a novel face alignment framework for the inherent relation learning and uncertainty estimation. Unlike most existing methods that regress facial landmarks directly from global features, the DSLPT first generates a rough representation of each landmark from a local patch cropped from the feature map and then adaptively aggregates them by a case dependent inherent relation. Finally, the DSLPT predicts the coordinate and uncertainty of each landmark by regressing their probability distribution from the output features. Moreover, we introduce a coarse-to-fine framework to incorporate with DSLPT for an improved result. In the framework, the position and size of each patch are determined by the probability distribution of the corresponding landmark predicted in the previous stage. The dynamic patches will ensure a fine-grained landmark representation for inherent relation learning so that a rough prediction result can gradually converge to the target facial landmarks. We integrate the coarse-to-fine model into an end-to-end training pipeline and carry out experiments on the mainstream benchmarks. The results demonstrate that the DSLPT achieves state-of-the-art performance with much less computational complexity. The codes and models are available at https://github.com/Jiahao-UTS/DSLPT. Jiahao Xia 0001, Min Xu 0001, Haimin Zhang 0001, Jianguo Zhang 0001, Wenjian Huang 0001, Hu Cao, Shiping Wen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Multiscale Emotion Representation Learning for Affective Image RecognitionabstractRecognition of emotions conveyed in images has attracted increasing research attention. Recent studies show that leveraging local affective regions helps to improve the recognition performance. However, these studies do not consider features from the broad context of the local affective regions, which could provide useful information for learning improved emotion representations. In this paper, we present a region-based multiscale network that learns features for the local affective region as well as the broad context for affective image recognition. The proposed network consists of an affective region detection module and a multiscale feature learning module. The class activation mapping method is used to generate pseudo affective regions from a pretrained deep neural network to train the detection module. For the affective region outputted by the detection module, three-scale features are extracted and then encoded by a kernel-based graph attention network for final emotion classification. We show that integrating features from the broad context is effective in improving the recognition performance. We experimentally evaluate the proposed network for both multi-class emotion recognition and binary sentiment classification on different benchmark datasets. The experimental results demonstrate that the proposed network achieves improved or comparable performance as compared to previous state-of-the-art models. Haimin Zhang 0001, Min Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Recognition of Emotions in User-Generated Videos through Frame-Level Adaptation and Emotion Intensity LearningabstractRecognition of emotions in user-generated videos has attracted considerable research attention. Most existing approaches focus on learning frame-level features and fail to consider frame-level emotion intensities which are critical for video representation. In this research, we aim to extract frame-level features and emotion intensities through transferring emotional information from an image emotion dataset. To achieve this goal, we propose an end-to-end network for joint emotion recognition and intensity learning with unsupervised adversarial adaptation. The proposed network consists of a classification stream, an intensity learning stream and an adversarial adaptation module. The classification stream is used to generate pseudo intensity maps with the class activation mapping method to train the intensity learning subnetwork. The intensity learning stream is built upon an improved feature pyramid network in which features from different scales are cross-connected. The adversarial adaptation module is employed to reduce the domain difference between the source dataset and target video frames. By aligning cross domain features, we enable our network to learn on the source data while generalizing to video frames. Finally, we apply a weighted sum pooling method to frame-level features and emotion intensities to generate video-level features. We evaluate the proposed method on two benchmark datasets,i.e.,VideoEmotion-8 and Ekman-6. The experimental results show that the proposed method achieves improved performance compared to previous state-of-the-art methods. Haimin Zhang 0001, Min Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | An efficient multitask neural network for face alignment, head pose estimation and face trackingabstractWhile Convolutional Neural Networks (CNNs) have significantly boosted the performance of face related algorithms, maintaining accuracy and efficiency simultaneously in practical use remains challenging. The state-of-the-art methods employ deeper networks for better performance, which makes it less practical for mobile applications because of more parameters and higher computational complexity. Therefore, we propose an efficient multitask neural network, Alignment & Tracking & Pose Network (ATPN) for face alignment, face tracking and head pose estimation . Specifically, to achieve better performance with fewer layers for face alignment, we introduce a shortcut connection between shallow-layer and deep-layer features. We find the shallow-layer features are highly correspond to facial boundaries that can provide the structural information of face and it is crucial for face alignment. Moreover, we generate a cheap heatmap based on the face alignment result and fuse it with features to improve the performance of the other two tasks. Based on the heatmap, the network can utilize both geometric information of landmarks and appearance information for head pose estimation. The heatmap also provides attention clues for face tracking. The face tracking task also saves us the face detection procedure for each frame, which also significantly boost the real-time capability for video-based tasks. We experimentally validate ATPN on four benchmark datasets, WFLW, 300VW, WIDER Face and 300W-LP. The experimental results demonstrate that it achieves better performance with much less parameters and lower computational complexity compared to other light models. Jiahao Xia 0001, Haimin Zhang 0001, Shiping Wen 0001, Shuo Yang 0006, Min Xu 0001 |
Expert Syst. Appl. | 2 |
| 2021 | GEME: Dual-stream multi-task GEnder-based micro-expression recognition
Xuan Nie, Madhumita A. Takalkar, Mengyang Duan, Haimin Zhang 0001, Min Xu 0001 |
Neurocomputing | 4 |
| 2021 | Graph neural networks with multiple kernel ensemble attention
Haimin Zhang 0001, Min Xu 0001 |
Knowl. Based Syst. | 1 |
| 2021 | Weakly Supervised Emotion Intensity Prediction for Recomi/tmi40.htmlgnition of Emotions in ImagesabstractRecognition of emotions in images is attracting increasing research attention. Recent studies show that using local region information helps to improve the recognition performance. Intuitively, emotion intensity maps provide more detailed information than image regions. Inspired by this intuition, we propose an end-to-end deep neural network for image emotion recognition leveraging emotion intensity learning. The proposed network is composed of a first classification stream, an intensity prediction stream and a second classification stream. The intensity prediction stream is built on top of the feature pyramid network to extract multilevel features. The class activation mapping technique is used to generate pseudo intensity maps from the first classification stream to guide the proposed network for emotion intensity learning. The predicted intensity map is integrated into the second classification stream for final emotion recognition. The three streams are trained cooperatively to improve the performance. We evaluate the proposed network for both emotion recognition and sentiment classification on different benchmark datasets. The experimental results demonstrate that the proposed network achieves improved performance compared to previous state-of-the-art approaches. Haimin Zhang 0001, Min Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Improving the generalization performance of deep networks by dual pattern learning with adversarial adaptationabstractIn this paper, we present a dual pattern learning network architecture with adversarial adaptation (DPLAANet). Unlike conventional networks, the proposed network has two input branches and two loss functions. This architecture forces the network to learn robust features by analysing dual inputs. The dual input structure allows the network to have a considerably large number of image pairs, which can help address the overfitting issue due to limited training data . In addition, we propose to associate the two input branches with two random interest values during training. As a stochastic regularization technique, this method can improve the generalization performance . Moreover, we introduce to use the adversarial training approach to reduce the domain difference between fused image features and single image features . Extensive experiments on CIFAR-10, CIFAR-100, FI-8, the Google commands dataset, and MNIST demonstrate that our DPLAANets exhibit better performance than the baseline networks. The experimental results on subsets of CIFAR-10, CIFAR-100, and MNIST demonstrate that DPLAANets have a good generalization performance on small datasets. The proposed architecture can be easily extended to have more than two input branches. The experimental results on subsets of MNIST show that the architecture with three branches outperforms two branches when the training set is extremely small. Haimin Zhang 0001, Min Xu 0001 |
Knowl. Based Syst. | 1 |
| 2019 | Improving Micro-expression Recognition Accuracy Using Twofold Feature Extraction
Madhumita A. Takalkar, Haimin Zhang 0001, Min Xu 0001 |
MMM (1) | 2 |
| 2019 | Multi-level region-based Convolutional Neural Network for image emotion classification
Tianrong Rao, Haimin Zhang 0001, Min Xu 0001 |
Neurocomputing | 3 |
| 2018 | LSTM-based Flight Trajectory PredictionabstractSafety ranks the first in Air Traffic Management (ATM). Accurate trajectory prediction can help ATM to forecast potential dangers and effectively provide instructions for safely traveling. Most trajectory prediction algorithms work for land traffic, which rely on points of interest (POIs) and are only suitable for stationary road condition. Compared with land traffic prediction, flight trajectory prediction is very difficult because way-points are sparse and the flight envelopes are heavily affected by external factors. In this paper, we propose a flight trajectory prediction model based on a Long Short-Term Memory (LSTM) network. The four interacting layers of a repeating module in an LSTM enables it to connect the long-term dependencies to present predicting task. Applying sliding windows in LSTM maintains the continuity and avoids compromising the dynamic dependencies of adjacent states in the long-term sequences, which helps to improve accuracy of trajectory prediction. Taking time dimension into consideration, both 3-D (time stamp, latitude and longitude) and 4-D (time stamp, latitude, longitude and altitude) trajectories are predicted to prove the efficiency of our approach. The dataset we use was collected by ADS-B ground stations. We evaluate our model by widely used measurements, such as the mean absolute error (MAE), the mean relative error (MRE), the root mean square error (RMSE) and the dynamic warping time (DWT) methods. As Markov Model is the most popular in time series processing, comparisons among Markov Model (MM), weighted Markov Model (wMM) and our model are presented. Our model outperforms the existing models (MM and wMM) and provides a strong basis for abnormal detection and decision-making. Min Xu 0001, Quan Pan 0001, Bing Yan 0001, Haimin Zhang 0001 |
IJCNN | 5 |
| 2018 | Recognition of Emotions in User-Generated Videos With Kernelized FeaturesabstractRecognition of emotions in user-generated videos has attracted increasing research attention. Most existing approaches are based on spatial features extracted from video frames. However, due to the broad affective gap between spatial features of images and high-level emotions, the performance of existing approaches is restricted. To bridge the affective gap, we propose recognizing emotions in user-generated videos with kernelized features. We reformulate the equation of the discrete Fourier transform as a linear kernel function and construct a polynomial kernel function based on the linear kernel. The polynomial kernel is applied to spatial features of video frames to generate kernelized features. Compared with spatial features, kernelized features show superior discriminative capability. Moreover, we are the first to apply the sparse representation method to reduce the impact of noise contained in videos; this method helps contribute to performance improvement. Extensive experiments are conducted on two challenging benchmark datasets, that is, VideoEmotion-8 and Ekman-6. The experimental results demonstrate that the proposed method achieves state-of-the-art performance. Haimin Zhang 0001, Min Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2016 | Modeling temporal information using discrete fourier transform for recognizing emotions in user-generated videosabstractWith the widespread of user-generated Internet videos, emotion recognition in those videos attracts increasing research efforts. However, most existing works are based on framelevel visual features and/or audio features, which might fail to model the temporal information, e.g. characteristics accumulated along time. In order to capture video temporal information, in this paper, we propose to analyse features in frequency domain transformed by discrete Fourier transform (DFT features). Frame-level features are firstly extract by a pre-trained deep convolutional neural network (CNN). Then, time domain features are transferred and interpolated into DFT features. CNN and DFT features are further encoded and fused for emotion classification. By this way, static image features extracted from a pre-trained deep CNN and temporal information represented by DFT features are jointly considered for video emotion recognition. Experimental results demonstrate that combining DFT features can effectively capture temporal information and therefore improve emotion recognition performance. Our approach has achieved a state-of-the-art performance on the largest video emotion dataset (VideoEmotion-8 dataset), improving accuracy from 51.1% to 55.6%. Haimin Zhang 0001, Min Xu 0001 |
ICIP | 1 |