Wenjing Han

dblp:80/8392 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
12since 2021 · last 2027
0000-0003-3180-547XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2027 Anomaly detection in IMU based on quantile regression-temporal pattern attention-transformer and covariance matrix adaptation evolution strategy
Wenting Han, Wenjing Han, Bingqin Liu, Moyao Yu
Expert Syst. Appl.3
2025 A Differential Inclusion Approach for Learning Heterogeneous Sparsity in Neuroimaging Analysis
abstract
In voxel-based neuroimaging disease prediction, it was recently found that in addition to lesion features, there exists another type of feature called "Procedural Bias", which is introduced during preprocessing and can further improve the prediction power. However, traditional sparse learning methods fail to simultaneously capture both types of features due to their heterogeneity in sparsity types. Specifically, the lesion features are spatially coherent and suffer from volumetric degeneration, while the procedural bias refers to enlarged voxels that are dispersedly distributed. In this paper, we propose a new method based on differential inclusion, which generates a sparse regularized solution path on multiple parameters that are enforced with heterogeneous sparsity to capture lesion features and the procedural bias separately. Specifically, we employ Total Variation with a non-negative constraint for the parameter associated with degenerated and spatially coherent lesions; on the other hand, we impose $\ell_1$ sparsity with a non-positive constraint on the parameter related to enlarged and scatterly distributed procedural bias. We theoretically show that our method enjoys model selection consistency and $\ell_2$ consistency in estimation. The utility of our method is demonstrated by improved prediction power and interpretability in the early prediction of Alzheimer’s Disease.
Wenjing Han, Yueming Wu 0005, Xinwei Sun 0001, Lingjing Hu, Yizhou Wang 0001
AISTATS1
2025 Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation
abstract
Liwen Sun, James Jialun Zhao, Wenjing Han, Chenyan Xiong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Liwen Sun, James (Jialun) Zhao, Wenjing Han, Chenyan Xiong
NAACL (Long Papers)3
2024 Simultaneous Retrieval of Land Surface Temperature and Soil Moisture Using Multichannel Passive Microwave Data
abstract
Land surface temperature (LST) and soil moisture (SM) are two important parameters in land surface ecosystem at regional and global scale. The accurate acquisition of LST and SM can benefit various fields, including agriculture and climate which are closely related to human life. The independent retrievals of LST and SM from passive microwave observations are mutually restricted and highly dependent on auxiliary data. To solve this problem, a simulations retrieval method of LST and SM was proposed based on the characteristics of multi-frequency and dual-polarization. The simultaneous solution of LST and SM was realized by approximating and correcting the radiative transfer equation (RTE). The performance of the proposed method was evaluated using simulated data, resulting in a root mean square error (RMSE) of approximately 1.63 K and 0.063 m3/m3. This method was further used to retrieve LST and SM from AMSR-E observations. The retrieved LST was compared to the MODIS land surface temperature product under clear-sky, with RMSE of 5.68 K. The retrieved LST was validated using the ISD air temperature under cloudy-sky, with RMSE of 4.29 K. The accuracy of retrieved LST changes with the variation of vegetation. The retrieved SM was evaluated using the CCI soil moisture product and in-situ observations. The result shows that the accuracy ranges from 0.0157 to 0.1115 m3/m3with the change of vegetation. This study gives a feasible method to retrieve LST and SM simultaneously with reasonable accuracy.
Xiao-Jing Han, Na Yao, Pei Leng, Wenjing Han, Xueyuan Chen
IEEE Trans. Geosci. Remote. Sens.5
2023 Speaker-Aware Hierarchical Transformer For Personality Recognition In Multiparty Dialogues
abstract
Personality recognition is one of the core technologies in human-machine interaction, which has received increasing attention. Previous works mainly focus on essays or monologues, while personality traits reveal more in the interactions with others. Due to the lack of appropriate datasets, a few approaches aim to recognize personality traits in conversations, and most of them ignore interdependence between speakers and connection between conversations. In this paper, we create a multiparty dialogue-based personality dataset derived from CPED containing 1,195 data samples. We center on one speaker and extract related dialogues to compose each data sample annotated with speaker’s Big-Five traits, which is conducive to fully describe a center speaker using diverse cues of personality in different dialogues. Along the same lines, we propose a Speaker-aware Hierarchical Transformer named SH-Transformer to address above concerns, in which Personalized Embeddings (PE) adopt special tokens to distinguish center speakers in complete conversations and hierarchical Transformer capture diverse cues in utterances and conversations. Experimental results show that our method outperforms the non-interactive baseline by 1.38%, which confirms the necessity of considering both interactive information and diverse cues among dialogues. Our code will be released at github.com/Chloehxxx/SH-Transformer.
Wenjing Han, Yirong Chen, Xiaofen Xing, Guohua Zhou, Xiangmin Xu 0001
ICASSP1
2023 Human Action Recognition Method Based on Spatio-temporal Relationship
Haigang Yu, Wenjing Han
ICIG (2)4
2023 Lightweight human pose estimation algorithm based on polarized self-attention
Haigang Yu, Wenjing Han
Multim. Syst.5
2023 YOLO-ERF: lightweight object detector for UAV aerial images
Fengxi Sun, Wenjing Han
Multim. Syst.5
2023 A Novel Approach to All-Weather LST Estimation Using XGBoost Model and Multisource Data
abstract
Land surface temperature (LST) plays a crucial role in the physical and chemical processes of the land–atmosphere system. Remote sensing technology has greatly advanced the measurement of thermal infrared LST (TIR LST), which is the most widely utilized surface temperature product. However, cloud cover and mist often cause significant data loss in TIR LST. To address this issue and reconstruct the MYD11A1 LST under cloudy conditions, this study proposes an all-weather LST generation method based on the extreme gradient boosting (XGBoost) model. This method incorporates spatial-seamless passive microwave LST (PMW LST) to capture the nonlinear relationship between TIR LST and other variables. Compared to the MYD11A1 LST, the generated all-weather LST provides continuous spatial texture information without a significant boundary reconstruction effect, improving the accuracy of spatiotemporal variations in LST in China. In situ validation demonstrated the high accuracy of the generated all-weather LST, with mean$R^{2}$, bias, and unbiased root-mean-square error (ubRMSE) of 0.96 (0.91), 1.08 K (3.61 K), and 2.92 K (4.54 K) under clear (cloudy) daytime conditions, and 0.92 (0.95), −0.93 K (−2.96 K), and 3.09 K (3.04 K) under clear (cloudy) nighttime conditions. These results indicate the feasibility and reasonableness of the all-weather LST generation method developed in this study and affirm its ability to generate highly accurate all-weather LST.
Sibo Duan, Yihua Lian, Enyu Zhao, Hong Chen 0021, Wenjing Han
IEEE Trans. Geosci. Remote. Sens.5
2023 Generation of Spatial-Seamless AMSR2 Land Surface Temperature in China During 2012-2020 Using a Deep Neural Network
abstract
Land surface temperature (LST) reflects the cold and hot conditions of the land surface and is one of the most important geophysical parameters in the study and research of the land–atmosphere system. Passive microwave (PMW) is one of the primary techniques for obtaining spatially continuous LST at regional, continental, and global scales. However, there is an orbital gap in the LST retrieved from PMW (PMW LST) due to the scanning scheme of the PMW sensor, which limits the application of PMW LST, so it is necessary for the proposed some methods to fill the orbital gap of PMW LST. In this study, a new orbital gap-filling method based on a deep neural network (DNN) was developed to address the issue of PMW LST orbital gaps. This method first established the DNN model based on the nonlinear relationship between AMSR2 LST and 11 environmental variables and then used the DNN model to generate a new spatially continuous LST product, namely, DNN-LST, and, finally, used DNN-LST to fill the orbital gaps of AMSR2 LST to generate the daytime/nighttime spatially seamless gap-filled LST (GF-LST) product for China from 2012 to 2020. GF-LST can more correctly represent the spatiotemporal variation of surface temperature in China than AMSR2 LST because it has continuous spatial texture information and no obvious boundary reconstruction effect. After verifying the accuracy of GF-LST products through simulated gap region validation and in situ validation, it can be found that: 1) DNN-LST in simulated gap regions showed high accuracy during the daytime and nighttime on July 15, 2012–2020, and the mean values of bias and root mean square error (RMSE) compared with AMSR2 LST at day (night) were, respectively, −0.08 K (−0.22 K) and 1.89 K (2.23 K); 2) the accuracy of DNN-LST was the best in autumn (mean RMSE values of 1.43 K at day and 1.89 K at night) and the worst in winter (mean RMSE values of 2.35 K at day and 2.36 K at night), no matter during daytime or nighttime, in different seasons in 2015–2017; 3) the RMSE value of DNN-LST during nighttime was slightly higher than the RMSE value of DNN-LST during daytime; and 4) the accuracy of DNN-LST was equivalent to AMSR2 LST, that is, the unbiased RMSE (ubRMSE) of DNN-LST and AMSR2 LST was all about 4 K compared with in situ LST, but the ubRMSE of DNN-LST was slightly lower than AMSR2 LST. The above accuracy validation analysis shows that DNN-LST has good robustness and good spatial consistency with AMSR2 LST and can be well used to fill the orbital gap of AMSR2 LST to generate spatial seamless GF-LST product.
Yihua Lian, Sibo Duan, Wenjing Han, Meng Liu 0009
IEEE Trans. Geosci. Remote. Sens.4
2022 Parameter-Efficient Conformers via Sharing Sparsely-Gated Experts for End-to-End Speech Recognition
abstract
While transformers and their variant conformers show promising performance in speech recognition, the parameterized property leads to much memory cost during training and inference.Some works use cross-layer weight-sharing to reduce the parameters of the model.However, the inevitable loss of capacity harms the model performance.To address this issue, this paper proposes a parameter-efficient conformer via sharing sparsely-gated experts.Specifically, we use sparsely-gated mixture-of-experts (MoE) to extend the capacity of a conformer block without increasing computation.Then, the parameters of the grouped conformer blocks are shared so that the number of parameters is reduced.Next, to ensure the shared blocks with the flexibility of adapting representations at different levels, we design the MoE routers and normalization individually.Moreover, we use knowledge distillation to further improve the performance.Experimental results show that the proposed model achieves competitive performance with 1/3 of the parameters of the encoder, compared with the full-parameter model.
Wenjing Han, Kaituo Xu
INTERSPEECH3
2021 Lightweight Non-local High-Resolution Networks for Human Pose Estimation
Qixiang Sun, Xiaojie Yin, Yuzhe He, Wenjing Han
ICIG (2)7
2020 Ordinal Learning for Emotion Recognition in Customer Service Calls
abstract
Approaches toward ordinal speech emotion recognition (SER) tasks are commonly based on the categorical classification algorithms, where the rank-order emotions are arbitrarily treated as independent categories. To employ the ordinal information between emotional ranks, we propose to model the ordinal SER tasks under a COnsistent RAnk Logits (CORAL) based deep learning framework. Specifically, a multi-class ordinal SER task is transformed into a series of binary SER sub-tasks predicting whether an utterance's emotion is larger than a rank. All the sub-tasks are jointly solved by one single network with a mislabelling cost defined as the the sum of the individual cross-entropy loss for each sub-task. Having the VGGish as our basic network structure, via minimizing above CORAL based cost, a VGGish-CORAL network is implemented in this contribution. Experimental results on a real-world call center dataset and the widely used IEMOCAP corpus demonstrate the effectiveness of VGGish-CORAL compared to the categorical VGGish.
Wenjing Han, Björn W. Schuller, Huabin Ruan
ICASSP1
2018 Towards Temporal Modelling of Categorical Speech Emotion Recognition
abstract
To model the categorical speech emotion recognition task in a temporal manner, the first challenge arising is how to transfer the categorical label for each utterance into a label sequence.To settle this, we make a hypothesis that an utterance is consisting of emotional and non-emotional segments, and these non-emotional segments correspond to silent regions, short pauses, transitions between phonemes, unvoiced phonemes, etc.With this hypothesis, we propose to treat an utterance's label sequence as a chain of two states: the emotional state denoting the emotional frame and Null denoting the non-emotional frame.Then, we exploit a recurrent neural network based connectionist temporal classification model to automatically label and align an utterance's emotional segments with emotional labels, while non-emotional segments with Nulls.Experimental results on the IEMOCAP corpus validate our hypothesis and also demonstrate the effectiveness of our proposed method compared to the state-of-the-art algorithms.
Wenjing Han, Huabin Ruan, Haifeng Li 0001, Björn W. Schuller
INTERSPEECH1
2017 DCNN and DNN based multi-modal depression recognition
abstract
In this paper, we propose an audio visual multimodal depression recognition framework composed of deep convolutional neural network (DCNN) and deep neural network (DNN) models. For each modality, corresponding feature descriptors are input into a DCNN to learn high-level global features with compact dynamic information, which are then fed into a DNN to predict the PHQ-8 score. For multi-modal depression recognition, the predicted PHQ-8 scores from each modality are integrated in a DNN for the final prediction. In addition, we propose the Histogram of Displacement Range as a novel global visual descriptor to quantify the range and speed of the facial landmarks' displacements. Experiments have been carried out on the Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WOZ) dataset for the Depression Sub-challenge of the Audio-Visual Emotion Challenge (AVEC 2016), results show that the proposed multi-modal depression recognition framework obtains very promising results on both the development set and test set, which outperforms the state-of-the-art results.
Le Yang 0009, Dongmei Jiang, Wenjing Han, Hichem Sahli
ACII3
2013 An FPGA-Based Data Flow Engine for Gaussian Copula Model
abstract
The Gaussian Copula Model (GCM) plays an important role in the state-of-the-art financial analysis field for modeling the dependence of financial assets. However, the existing implementations of GCM are all computationallydemanding and time-consuming. In this paper, we propose a Dataflow Engine (DFE) design to accelerate the GCM computation. Specifically, a commonly used CPU-friendly GCM algorithm is converted into a fully-pipelined dataflow graph through four steps of optimization: recomposing the algorithm to be pipeline-friendly, removing unnecessary computation, sharing common computing results, and reducing the computing precision while maintaining the same level of accuracy for the computation results. The performance of the proposed DFE design is compared with three CPU-based implementations that are well-optimized. Experimental results show that our DFE solution not only generates fairly accurate result, but also achieves a maximum of 467x speedup over a single-thread CPU-based solution, 120x speedup over a multi-thread CPUbased solution, and 47x speedup over an MPI-based solution.
Huabin Ruan, Xiaomeng Huang, Haohuan Fu, Guangwen Yang 0002, Wayne Luk, Sébastien Racanière, Oliver Pell, Wenjing Han
FCCM8
2013 Active learning for dimensional speech emotion recognition
Wenjing Han, Haifeng Li 0001, Huabin Ruan, Lin Ma 0003, Jiayin Sun, Björn W. Schuller
INTERSPEECH1
2012 Automatic recognition of emotion evoked by general sound events
abstract
Without a doubt there is emotion in sound. So far, however, research efforts have focused on emotion in speech and music despite many applications in emotion-sensitive sound retrieval. This paper is an attempt at automatic emotion recognition of general sounds. We selected sound clips from different areas of the daily human environment and model them using the increasingly popular dimensional approach in the emotional arousal and valence space. To establish a reliable ground truth, we compare mean and median of four annotators with their evaluator weighted estimator. We discuss human labelers' consistency, feature relevance, and automatic regression. Results reach correlation coefficients of .61 (arousal) and .49 (valence).
Björn W. Schuller, Simone Hantke, Felix Weninger, Wenjing Han, Zixing Zhang 0001, Shri Narayanan
ICASSP4
2012 Preserving actual dynamic trend of emotion in dimensional speech emotion recognition
abstract
In this paper, we use the concept of dynamic trend of emotion to describe how a human's emotion changes over time, which is believed to be important for understanding one's stance toward current topic in interactions. However, the importance of this concept - to our best knowledge - has not been paid enough attention before in the field of speech emotion recognition (SER). Inspired by this, this paper aims to evoke researchers' attention on this concept and makes a primary effort on the research of predicting correct dynamic trend of emotion in the process of SER. Specifically, we propose a novel algorithm named Order Preserving Network (OPNet) to this end. First, as the key issue for OPNet construction, we propose employing a probabilistic method to define an emotion trend-sensitive loss function. Then, a nonlinear neural network is trained using the gradient descent as optimization algorithm to minimize the constructed loss function. We validated the prediction performance of OPNet on the VAM corpus, by mean linear error as well as a rank correlation coefficient γ as measures. Comparing to k-Nearest Neighbor and support vector regression, the proposed OPNet performs better on the preservation of actual dynamic trend of emotion.
Wenjing Han, Haifeng Li 0001, Florian Eyben, Lin Ma 0003, Jiayin Sun, Björn W. Schuller
ICMI1