EDBT 2026 Demo / reviewers in the wild / expert
Xiao Sun 0003
dblp:30/202-3
· DBLP profile ↗
83ranked-venue papers
22as first author
62since 2021 · last 2026
0000-0001-9750-7032ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 9 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 5 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 4 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorSystems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CausalSymptom: Learning Causal Disentangled Representation for Depression Severity Estimation on Transcribed Clinical Interviews
Mingzheng Li, Xiao Sun 0003, Xinke Wang, Feng-Qi Cui, Xun Yang 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | GRPose: Learning Graph Relations for Human Image Generation with Pose PriorsabstractRecent methods using diffusion models have made significant progress in human image generation with various control signals such as pose priors. However, existing efforts are still struggling to generate high-quality images with consistent pose alignment, resulting in unsatisfactory output. In this paper, we propose a framework that delves into the graph relations of pose priors to provide control information for human image generation. The main idea is to establish a graph topological structure between the pose priors and latent representation of diffusion models to capture the intrinsic associations between different pose parts. A Progressive Graph Integrator (PGI) is designed to learn the spatial relationships of the pose priors with the graph structure, adopting a hierarchical strategy within an Adapter to gradually propagate information across different pose parts. Besides, a pose perception loss is introduced based on a pretrained pose estimation network to minimize the pose differences. Extensive qualitative and quantitative experiments conducted on the Human-Art and LAION-Human datasets clearly demonstrate that our model can achieve significant performance improvement over the latest benchmark models. Xiangchen Yin, Donglin Di, Lei Fan 0007, Hao Li 0030, Wei Chen 0089, Gouxiao Fei, Yang Song 0001, Xiao Sun 0003, Xun Yang 0001 |
AAAI | 8 |
| 2025 | MultiAgentESC: A LLM-based Multi-Agent Collaboration Framework for Emotional Support ConversationabstractThe development of Emotional Support Conversation (ESC) systems is critical for delivering mental health support tailored to the needs of help-seekers.Recent advances in large language models (LLMs) have contributed to progress in this domain, while most existing studies focus on generating responses directly and overlook the integration of domain-specific reasoning and expert interaction.Therefore, in this paper, we propose a training-free Multi-Agent collaboration framework for ESC (Mul-tiAgentESC).The framework is designed to emulate the human-like process of providing emotional support through three stages: dialogue analysis, strategy deliberation, and response generation.At each stage, a multi-agent system is employed to iteratively enhance information understanding and reasoning, simulating real-world decision-making processes by incorporating diverse interactions among these expert agents.Additionally, we introduce a novel response-centered approach to handle the one-to-many problem on strategy selection, where multiple valid strategies are initially employed to generate diverse responses, followed by the selection of the optimal response through multi-agent collaboration.Experiments on the ESConv dataset reveal that our proposed framework excels at providing emotional support as well as diversifying support strategy selection 1 . Yangyang Xu 0002, Jinpeng Hu, Zhuoer Zhao, Zhangling Duan, Xiao Sun 0003, Xun Yang 0001 |
EMNLP | 5 |
| 2025 | DEFormer: DCT-driven Enhancement Transformer for Low-light Image and Dark VisionabstractLow-light image enhancement restores the colors and details of a single image and improves high-level visual tasks. However, restoring the lost details in the dark area is still a challenge relying only on the RGB domain. In this paper, we delve into frequency as a new clue into the model and propose a DCT-driven enhancement transformer (DEFormer) framework. First, we propose a learnable frequency branch (LFB) for frequency enhancement contains DCT processing and curvature-based frequency enhancement (CFE) to represent frequency features. Additionally, we propose a cross domain fusion (CDF) to reduce the differences between the RGB domain and the frequency domain. Our DEFormer has achieved superior results on the LOL and MIT-Adobe FiveK datasets, improving the dark detection performance. Xiangchen Yin, Zhenda Yu, Xin Gao 0028, Xiao Sun 0003 |
ICASSP | 4 |
| 2025 | Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) plays a critical role in enhancing user experience within human-computer interaction. However, existing methods are overwhelmed by temporal domain analysis, overlooking the valuable envelope structures of the frequency domain that are equally important for robust emotion recognition. To overcome this limitation, we propose TF-Mamba, a novel multi-domain framework that captures emotional expressions in both temporal and frequency dimensions. Concretely, we propose a temporal-frequency mamba block to extract temporal- and frequency-aware emotional features, achieving an optimal balance between computational efficiency and model expressiveness. Besides, we design a Complex Metric-Distance Triplet (CMDT) loss to enable the model to capture representative emotional clues for SER. Extensive experiments on the IEMOCAP and MELD datasets show that TF-Mamba surpasses existing methods in terms of model size and latency, providing a more practical solution for future SER applications. Fei Wang 0067, Kun Li 0008, Yanyan Wei, Shengeng Tang, Shu Zhao 0005, Xiao Sun 0003 |
ICASSP | 7 |
| 2025 | AMMSM: Adaptive Motion Magnification and Sparse Mamba for Micro-Expression RecognitionabstractMicro-expressions are typically regarded as unconscious manifestations of a person's genuine emotions. However, their short duration and subtle signals pose significant challenges for downstream recognition. We propose a multi-task learning framework named the Adaptive Motion Magnification and Sparse Mamba (AMMSM) to address this. This framework aims to enhance the accurate capture of micro-expressions through self-supervised subtle motion magnification, while the sparse spatial selection Mamba architecture combines sparse activation with the advanced Visual Mamba model to model key motion regions and their valuable representations more effectively. Additionally, we employ evolutionary search to optimize the magnification factor and the sparsity ratios of spatial selection, followed by fine-tuning to improve performance further. Extensive experiments on two standard datasets demonstrate that the proposed AMMSM achieves state-of-the-art (SOTA) accuracy and robustness. Xuxiong Liu, Tengteng Dong, Fei Wang 0067, Weijie Feng, Xiao Sun 0003 |
ICME | 5 |
| 2025 | Robust Facial Expression Recognition via Symmetry Exploitation and Adaptive Neighborhood Label CorrectionabstractFacial Expression Recognition (FER) datasets often contain ambiguous labels due to inter-class similarity, the inherent complexity of facial expressions, and subjectivity in human annotations. Although significant research has focused on reducing noise in facial expression data, facial expression recognition remains challenging due to the complexity of environmental factors and individual differences. In this paper, we propose a method for Facial Expression Recognition (FER) called Robust Facial Expression Recognition via Symmetry Exploitation and Adaptive Neighborhood Label Correction (SEANLC). Specifically, we exploit the natural symmetry of faces to enhance the model’s capacity to recognize key facial features. Furthermore, we generate an adaptive threshold for each class to divide the data into clean and noisy samples, and employ a neighborhood label correction strategy to the noisy samples to mitigate the impact of noisy labels. Extensive experiments demonstrate the effectiveness of our approach in reducing noise and ambiguity, significantly outperforming state-of-the-art methods in noisy label FER. Xiao Sun 0003, Xuxiong Liu |
IJCNN | 2 |
| 2025 | Wi-Pulmo: Commodity WiFi Can Capture Your Pulmonary Function Without Mouth ClingingabstractPulmonary function testing is a crucial examination for respiratory diseases. Current medical spirometers are bulky and inconvenient, while available portable spirometers are extremely expensive and often lack accuracy. Furthermore, both devices require direct contact, inevitably increasing the cross-infection risk. To tackle these challenges, we propose Wi-Pulmo, an end-to-end deep learning-based Wireless System that utilizes WiFi channel state information (CSI) to provide contact-free, convenient, cost-effective, and precise pulmonary function testing outside the clinical setting. Based on the analysis of thoracic and abdominal movement patterns, Wi-Pulmo first validates the feasibility of using WiFi to estimate pulmonary function. Then, Wi-Pulmo designs an efficient fine-grained sensing quality-based algorithm for complete exhalation segmentation. Additionally, a relevant interference-tolerant learning algorithm based on variational inference is proposed to accurately map the CSI of WiFi signals to pulmonary function. Extensive experiments achieved average monitoring error rates of 2.59% for normal subjects in daily scenarios and 5.87% for real patients in tertiary hospitals over a two-month period. These satisfactory results demonstrate the strong effectiveness and robustness of Wi-Pulmo. Furthermore, our findings in clinical reveal a close correlation between chronic diseases and pulmonary function. Peng Zhao 0024, Jinyang Huang, Xiang Zhang 0011, Zhi Liu 0002, Huan Yan 0004, Meng Wang 0037, Guohang Zhuang, Yutong Guo, Xiao Sun 0003, Meng Li 0006 |
IEEE Internet Things J. | 9 |
| 2025 | Empathy Level Alignment via Reinforcement Learning for Empathetic Response GenerationabstractEmpathetic response generation, aiming to understand the user’s situation and feelings and respond empathically, is crucial in building human-like dialogue systems. Traditional approaches typically employ maximum likelihood estimation as the optimization objective during training, yet fail to align the empathy levels between generated and target responses. To this end, we propose an empathetic response generation framework using reinforcement learning (EmpRL). The framework develops an effective empathy reward function and generates empathetic responses by maximizing the expected reward through reinforcement learning. EmpRL utilizes the pre-trained T5 model as the generator and further fine-tunes it to initialize the policy. To align the empathy levels between generated and target responses within a given context, an empathy reward function containing three empathy communication mechanisms—emotional reaction, interpretation, and exploration—is constructed using pre-designed and pre-trained empathy identifiers. During reinforcement learning training, the proximal policy optimization algorithm is used to fine-tune the policy, enabling the generation of empathetic responses. Both automatic and human evaluations demonstrate that the proposed EmpRL framework significantly improves the quality of generated responses, enhances the similarity in empathy levels between generated and target responses, and produces empathetic responses covering both affective and cognitive aspects. Hui Ma 0011, Bo Zhang 0121, Bo Xu 0009, Jian Wang 0021, Hongfei Lin, Xiao Sun 0003 |
IEEE Trans. Affect. Comput. | 6 |
| 2025 | Facial Expression Recognition With Vision Transformer Using Fused Shifted WindowsabstractFacial expressions contain massive affective information. Previous methods have focused on using diverse models based on CNN or Transformer to handle the facial expression recognition(FER) task. However, most of them treat the FER task as a general image classification task and neglect the impact of regions of interest (ROIs) for the performance of FER. To verify the influence of different ROIs, in this paper, we propose a vision Transformer based onFusedShiftedwindows (FSwin), called FSwin Transformer. The semantic information of the face, obtained by ROIs, guides the FSwin Transformer to focus more on the key regions. The fused shifted windows enable the model to perform global semantic interactions, allowing it to concentrate on key regions without losing the topological structural information of the entire face. Additionally, learnable parameters are introduced to learn the feature weights expressed by each ROI, helping the model dynamically adjust the attention distribution. We have conducted controlled experiments to quantitatively verify the impact of ROIs. And the experimental results show that with the increase of the number of ROIs, the accuracy of FER is significantly improved, demonstrating that the key ROIs play an important role in feature extraction. The results on Jaffe, CK+, FER2013, AffectNet have reached 99.8%, 99.0%, 74.8%, and 68.9%, respectively, which all set new state-of-the-art. Extensive cross-dataset experiments also show that the FSwin Transformer has good generalization ability, proving that our proposed model has a beneficial effect on FER tasks. Xiao Sun 0003, Rui Wang 0156, Shaokai Chen, Meng Wang 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | PsycoLLM: Enhancing LLM for Psychological Understanding and EvaluationabstractMental health has attracted substantial attention in recent years and large language model (LLM) can be an effective technology for alleviating this problem owing to its capability in text understanding and dialogue. However, existing research in this domain often suffers from limitations, such as training on datasets lacking crucial prior knowledge and evidence, and the absence of comprehensive evaluation methods. In this article, we propose a specialized psychological LLM, named PsycoLLM, trained on a proposed high-quality psychological dataset, including single-turn QA, multiturn dialogues, and knowledge-based QA. Specifically, we construct multi-turn dialogues through a three-step pipeline comprising multiturn QA generation, evidence judgment, and dialogue refinement. We augment this process with real-world psychological case backgrounds extracted from online platforms, enhancing the relevance and applicability of the generated data. Additionally, to compare the performance of PsycoLLM with other LLMs, we develop a comprehensive psychological benchmark based on authoritative psychological counseling examinations in China, which includes assessments of professional ethics, theoretical proficiency, and case analysis. The experimental results on the benchmark illustrate the effectiveness of PsycoLLM, which demonstrates superior performance compared with other LLMs. Jinpeng Hu, Tengteng Dong, Hui Ma 0011, Xiao Sun 0003, Dan Guo 0001, Xun Yang 0001, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | Temporal Gated Face Alignment Network for Camera-Based Physiological SensingabstractThe remote photoplethysmography (rPPG) technique estimates vital signs, such as heart rate (HR), by analyzing subtle skin color variation in facial videos induced by the pulse. However, it remains a critical challenge to robustly acquire cardiac pulse information in scenarios with head motion (e.g., rotation and swing), as it inevitably introduces interfering noise, such as facial geometric deformation and displacement. Most existing methods primarily focus on how to extract subtle pulse signals while neglecting the detrimental effects of noise, especially out-of-distribution motion patterns. In response to this, we propose a temporal gated face alignment network (TGFAN) to adaptively counteract motion noise in long-term video sequences. Specifically, the bidirectional temporal face alignment (BFA) block first captures interframe motion discrepancies to align the displaced face feature and then extracts motion-robust pulse features. Furthermore, we propose a learnable temporal gating mechanism that disentangles the features and dynamically guides motion-disturbed segments for feature alignment, thereby alleviating local head motion interference. Experimental evaluations on four benchmark datasets demonstrate our superior performance on both intradataset and cross-dataset tests. Qi Li 0045, Dan Guo 0001, Yuanen Zhou, Xiao Sun 0003, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Prompt Learning With Multiperspective Cues for Emotional Support Conversation SystemsabstractThe emotional support conversation (ESC) system is tailored to reduce the distress of individuals who are experiencing emotional challenges. To achieve this goal, the system must fully understand the emotional state of the user and employ appropriate strategies to provide effective emotional support. However, most existing methods predominantly focus on the user’s psychological state and the causes of user’s emotional problem, neglecting key cues, e.g., topics of conversation and listener’s psychological state. This leads to an incomplete understanding of the user’s situation and ineffective emotional support. This article presents a novel method that utilizesprompt learning withmultiperspectivecues (PMPC) to generate emotional support responses. Specifically, we extract multiperspective cues from the dialogue history and the causes of user’s emotional problem. To fully utilize the different cues within the ESC system, we categorize them into two groups:semantic enhancement cuesandsemantic constraint cues. Subsequently, we construct two prompts based on different categories of cues, which are designed to thoroughly understand user’s emotional predicament and generate an engaging support response, respectively. Experimental results on the ESConv dataset demonstrate that our proposed PMPC can surpass other approaches in both automatic and human evaluation metrics, providing compelling evidence of its efficacy and potential impact in real-world applications. Yangyang Xu 0002, Zhuoer Zhao, Xiao Sun 0003, Xun Yang 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | Facial Depression Estimation via Multi-Cue Contrastive LearningabstractVision-based depression estimation is an emerging yet impactful task, whose challenge lies in predicting the severity of depression from facial videos lasting at least several minutes. Existing methods primarily focus on fusing frame-level features to create comprehensive representations. However, they often overlook two crucial aspects: 1) inter- and intra-cue correlations, and 2) variations among samples. Hence, simply characterizing sample embeddings while ignoring to mine the relation among multiple cues leads to limitations. To address this problem, we propose a novel Multi-Cue Contrastive Learning (MCCL) framework to mine the relation among multiple cues for discriminative representation. Specifically, we first introduce a novel cross-characteristic attentive interaction module to model the relationship among multiple cues from four facial features (e.g., 3D landmarks, head poses, gazes, FAUs). Then, we propose a temporal segment attentive interaction module to capture the temporal relationships within each facial feature over time intervals. Moreover, we integrate contrastive learning to leverage the variations among samples by regarding the embeddings of inter-cue and intra-cue as positive pairs while considering embeddings from other samples as negative. In this way, the proposed MCCL framework leverages the relationships among the facial features and the variations among samples to enhance the process of multi-cue mining, thereby achieving more accurate facial depression estimation. Extensive experiments on public datasets, DAIC-WOZ, CMDC, and E-DAIC, demonstrate that our model not only outperforms the advanced depression methods but that the discriminative representations of facial behaviors provide potential insights about depression. Our code is available at:https://github.com/xkwangcn/MCCL.git Xinke Wang, Xiao Sun 0003, Mingzheng Li, Bin Hu 0001, Dan Guo 0001, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Fuzzy Membership-Driven Adaptive Sample Mining for Facial Expression RecognitionabstractFacial expression recognition (FER) datasets are plagued by fuzzy label uncertainty due to interclass similarities, compound emotions, and annotation subjectivity. Moreover, the complexity of recognition differs across various expression categories, making an undifferentiated method inequitable for each expression. In this work, we propose adaptive sample mining (ASM), a fuzzy membership-driven framework with two innovations: it constructs a category-specific dynamic threshold set as the average probabilities of samples that are correctly predicted, through which it explicitly separates clean and noisy subsets while leaving ambiguous samples as the intermediate fuzzy region between them. In addition, it introduces a rule based on multilocal region prediction consistency to reclassify samples from the clean subset into the ambiguous subset when local predictions conflict. Finally, the triregularization module implements tailored training approaches for these three subsets: the clean subset is trained using a supervised learning strategy to maximize the utility of high-quality annotations, the ambiguous subset adopts a mutual learning strategy to enhance discrimination capability, and the noisy subset employs an unsupervised learning technique so that it can minimize the influence from noisy labels. We performed comprehensive experiments on the RAF-DB, FERPlus, and AffectNet. The proposed ASM achieved 90.94% accuracy on RAF-DB and 86.81% on its synthetic noisy version. The ablation experiment and analysis demonstrate that our method can efficiently identify ambiguous and noisy labels. By weighting the classification loss based on fuzzy membership, it helps to make efficient use of the labels. This results in an improvement in accuracy. Xuxiong Liu, Liuwei An, Xiao Sun 0003 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2025 | UniEmoX: Cross-Modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion PerceptionabstractVisual emotion analysis holds significant research value in both computer vision and psychology. However, existing methods for visual emotion analysis suffer from limited generalizability due to the ambiguity of emotion perception and the diversity of data scenarios. To tackle this issue, we introduce UniEmoX, a cross-modal semantic-guided large-scale pretraining framework. Inspired by psychological research emphasizing the inseparability of the emotional exploration process from the interaction between individuals and their environment, UniEmoX integrates scene-centric and person-centric low-level image spatial structural information, aiming to derive more nuanced and discriminative emotional representations. By exploiting the similarity between paired and unpaired image-text samples, UniEmoX distills rich semantic knowledge from the CLIP model to enhance emotional embedding representations more effectively. To the best of our knowledge, this is the first large-scale pretraining framework that integrates psychological theories with contemporary contrastive learning and masked image modeling techniques for emotion analysis across diverse scenarios. Additionally, we develop a visual emotional dataset titled Emo8. Emo8 samples cover a range of domains, including cartoon, natural, realistic, science fiction and advertising cover styles, covering nearly all common emotional scenes. Comprehensive experiments conducted on seven benchmark datasets across two downstream tasks validate the effectiveness of UniEmoX. The source code is available at https://github.com/chincharles/u-emo. Xiao Sun 0003, Zhi Liu 0002 |
IEEE Trans. Image Process. | 2 |
| 2025 | Structure-Guided Diffusion Transformer for Low-Light Image EnhancementabstractWhile the diffusion transformer (DiT) has become a focal point of interest in recent years, its application in low-light image enhancement remains a blank area for exploration. Current methods recover the details from low-light images while inevitably amplifying the noise in images, resulting in poor visual quality. In this paper, we firstly introduce DiT into the low-light enhancement task and design a novel Structure-guided Diffusion Transformer based Low-light image enhancement (SDTL) framework. We compress the feature through wavelet transform to improve the inference efficiency of the model and capture the multi-directional frequency band. Then we propose a Structure Enhancement Module (SEM) that uses structural prior to enhance the texture and leverages an adaptive fusion strategy to achieve more accurate enhancement effect. In Addition, we propose a Structure-guided Attention Block (SAB) to pay more attention to texture-riched tokens and avoid interference from noisy areas in noise prediction. Extensive qualitative and quantitative experiments demonstrate that our method achieves SOTA performance on several popular datasets, validating the effectiveness of SDTL in improving image quality and the potential of DiT in low-light enhancement tasks. Xiangchen Yin, Zhenda Yu, Longtao Jiang, Xin Gao 0028, Xiao Sun 0003, Zhi Liu 0002, Xun Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | MCAN: An Efficient Multi-Task Network for Facial Expression AnalysisabstractWith the development of artificial intelligence, artificial intelligence technology is widely used in robots. For example, emotional computing robots need to be able to complete the function of facial expression analysis. The multi-task deep learning model MCAN proposed in this paper is designed to complete facial expression analysis for robots which can predict discrete expressions and dimensional measures. The class center loss function proposed in this paper can increase the inter-class distance while reducing the intra-class distance. In addition, multi-head attention network was improved to focus on different areas of the input image. Finally, feature pyramid network allows information exchange and fusion between feature maps at different levels. Experiments have proven that MCAN achieves excellent results on both the AffectNet dataset and the AFEW-VA dataset, and also speeds up inference. Rui Wang 0156, Qingjian Ni, Xiao Sun 0003 |
CSCWD | 5 |
| 2024 | WR-Former: Vision Transformer with Weight Reallocation Module for Robust Facial Expression RecognitionabstractFacial expression recognition(FER) is an exceedingly challenging task in the field of computer vision, primarily due to variant head poses, occlusions and illumination conditions. Previous methods focus on extracting more facial features to improve the performance on the FER benchmarks. However, most of them ignore the weights of the key regions. To address the above problem, this paper proposes a vision Transformer with weight reallocation module termed as WR-Former. The weight reallocation module utilizes a hybrid multi-head self-attention and can dynamically iterate during the training process. It guides the model to reallocate tokens based on the importance of different facial regions. Besides, a simple yet effective data augmentation method is proposed to expand training samples and improve the robustness. Experimental results and visualizations demonstrate that our WR-Former outperforms the previous state-of-the-art methods and has the ability to better extract the features conveyed by the key facial regions. Rui Wang 0156, Xiao Sun 0003 |
CSCWD | 4 |
| 2024 | DM-NAI: Dynamic Information Diffusion Model Incorporating Non-Adjacent Node InteractionabstractDescribing the dynamics of information diffusion within social networks poses a formidable challenge. Despite multiple endeavors aimed at addressing this issue, only a limited number of studies have effectively replicated and forecasted the evolving course of information diffusion. In this paper, we propose a novel model, DM-NAI, which not only considers the information transfer between adjacent users but also takes into account the information transfer between non-adjacent users to comprehensively depict the information diffusion process. Extensive experiments are conducted on six datasets to predict the information diffusion range and the diffusion trend of the social network. The experimental results demonstrate an average prediction accuracy range of 94.62% to 96.71%, respectively, significantly outperforming state-of-the-art solutions. This finding illustrates that considering information transmission between non-adjacent users helps DM-NAI achieve more accurate information diffusion predictions. Jinyang Huang, Xiang Zhang 0011, Peng Zhao 0024, Guohang Zhuang, Huan Yan 0004, Xiao Sun 0003, Meng Wang 0037 |
ICC | 8 |
| 2024 | UAPE: Information Propagation Model Based on User Attitude and Public Opinion EnvironmentabstractModeling the information propagation process in social networks is a challenging problem. Despite numerous attempts to address this issue, existing studies often assume that user attitudes have only one opportunity to alter during the information propagation process. Additionally, these studies tend to consider the transformation of user attitudes as solely influenced by a single user, overlooking the dynamic and evolving nature of user attitudes and the impact of the public opinion environment. In this paper, we propose a novel model, UAPE, which considers the influence of the aforementioned factors on the information propagation process. Specifically, UAPE regards the user's attitude towards the topic as dynamically changing, with the change jointly affected by multiple users simultaneously. Furthermore, the joint influence of multiple users can be considered as the impact of the public opinion environment. Extensive experimental results demonstrate that the model achieves an accuracy range of 91.62% to 94.01 %, surpassing the performance of existing research. Jinyang Huang, Xiang Zhang 0011, Peng Zhao 0024, Guohang Zhuang, Huan Yan 0004, Xiao Sun 0003, Meng Wang 0037 |
ICC | 8 |
| 2024 | LDIP: Real-time on-road object detection with depth estimation from a single imageabstractDetecting on-road objects with absolute depth information is one of the most crucial tasks in autonomous driving to ensure safety. Traditional 2D object detection aims to classify and locate objects in image space, but it cannot acquire in-depth information. While 3D object detection and pixel-level depth detection tasks can provide accurate depth information for objects, they are challenging to deploy in real-world scenarios due to their significant inference overhead. This paper proposes a novel deep learning-based model named the Location and Depth Information Perceptron (LDIP), designed to provide positional, categorical, and absolute depth information for given objects in the images.We first conducted model training and validation on the vehicle-side autonomous driving dataset—KITTI. The experimental results show that we achieved a 68.6% mAP in object recognition tasks and an RMSE of 0.101 and AbsRel of 2.327 in depth estimation tasks, all of which represent state-of-the-art performance in comparable tasks. Subsequently, we fine-tuned the trained model on DAIR, where the validated mAP, AbsRel, and RMSE reached 65.4%, 0.092, and 2.461 respectively. This demonstrates the robustness and generalization of our model across different types of road datasets.Moreover, in comparison to other models, our model is more compact while maintaining accuracy, achieving an inference speed of 70 frames per second on an NVIDIA 4060 GPU, thus making it deployable in practical scenarios. Relevant code is available at https://github.com/xcp-ustc/LDIP. Chengpeng Xu, Xiao Sun 0003, Yangyang Xu 0002, Ruolin Wang |
IROS | 2 |
| 2024 | FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial LandmarksabstractDepression is a prevalent mental health disorder that significantly impacts individuals' lives and well-being. Early detection and intervention are crucial for effective treatment and management of depression. Recently, there are many end-to-end deep learning methods leveraging the facial expression features for automatic depression detection. However, most current methods overlook the temporal dynamics of facial expressions. Although very recent 3DCNN methods remedy this gap, they introduce more computational cost due to the selection of CNN-based backbones and redundant facial features. To address the above limitations, by considering the timing correlation of facial expressions, we propose a novel framework called FacialPulse, which recognizes depression with high accuracy and speed. By harnessing the bidirectional nature and proficiently addressing long-term dependencies, the Facial Motion Modeling Module (FMMM) is designed in FacialPulse to fully capture temporal features. Since the proposed FMMM has parallel processing capabilities and has the gate mechanism to mitigate gradient vanishing, this module can also significantly boost the training speed. Besides, to effectively use facial landmarks to replace original images to decrease information redundancy, a Facial Landmark Calibration Module (FLCM) is designed to eliminate facial landmark errors to further improve recognition accuracy. Extensive experiments on the AVEC2014 dataset and MMDA dataset (a depression dataset) demonstrate the superiority of FacialPulse on recognition accuracy and speed, with the average MAE (Mean Absolute Error) decreased by 21% compared to baselines, and the recognition speed increased by 100% compared to state-of-the-art methods. Codes are released at https://github.com/volatileee/FacialPulse. Jinyang Huang, Jie Zhang 0042, Xin Liu 0104, Xiang Zhang 0011, Zhi Liu 0002, Peng Zhao 0024, Sigui Chen, Xiao Sun 0003 |
ACM Multimedia | 9 |
| 2024 | EmoAda: A Multimodal Emotion Interaction and Psychological Adaptation System
Tengteng Dong, Xinke Wang, Yishun Jiang, Xiao Sun 0003 |
MMM (4) | 6 |
| 2024 | Joint Multi-cue Learning for Emotion Recognition in Human-Computer Interaction
Xiao Sun 0003 |
PRCV (11) | 2 |
| 2024 | SCBG: Semantic-Constrained Bidirectional Generation for Emotional Support ConversationabstractThe Emotional Support Conversation (ESC) task aims to deliver consolation, encouragement, and advice to individuals undergoing emotional distress, thereby assisting them in overcoming difficulties. In the context of emotional support dialogue systems, it is of utmost importance to generate user-relevant and diverse responses. However, previous methods failed to take into account these crucial aspects, resulting in a tendency to produce universal and safe responses (e.g., “I do not know” and “I am sorry to hear that”). To tackle this challenge, a semantic-constrained bidirectional generation (SCBG) framework is utilized for generating more diverse and user-relevant responses. Specifically, we commence by selecting keywords that encapsulate the ongoing dialogue topics based on the context. Subsequently, a bidirectional generator generates responses incorporating these keywords. Two distinct methodologies, namely, statistics-based and prompt-based methods, are employed for keyword extraction. Experimental results on the ESConv dataset demonstrate that the proposed SCBG framework improves response diversity and user relevance while ensuring response quality. Yangyang Xu 0002, Zhuoer Zhao, Xiao Sun 0003 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2024 | CoMix: Confronting with Noisy Label Learning with Co-training Strategies on Textual MislabelingabstractThe existence of noisy labels is inevitable in real-world large-scale corpora. As deep neural networks are notably vulnerable to overfitting on noisy samples, this highlights the importance of the ability of language models to resist noise for efficient training. However, little attention has been paid to alleviating the influence of label noise in natural language processing. To address this problem, we present CoMix, a robust Noise-against training strategy taking advantage of Co-training that deals with textual annotation errors in text classification tasks. In our proposed framework, the original training set is first split into labeled and unlabeled subsets according to a sample partition criteria and then applies label refurbishment on the unlabeled subsets. We implement textual interpolation in hidden space between samples on the updated subsets. Meanwhile, we employ peer diverged networks simultaneously leveraging co-training strategies to avoid the accumulation of confirm bias. Experimental results on three popular text classification benchmarks demonstrate the effectiveness of CoMix in bolstering the network’s resistance to label mislabeling under various noise types and ratios, which also outperforms the state-of-the-art methods. Shu Zhao 0005, Zhuoer Zhao, Yangyang Xu 0002, Xiao Sun 0003 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2024 | Hypergraph Neural Network for Emotion Recognition in ConversationsabstractModeling conversational context is an essential step for emotion recognition in conversations. Existing works still suffer from insufficient utilization of local context information and remote context information. This article designs a hypergraph neural network, namely HNN-ERC, to better utilize local and remote contextual information. HNN-ERC combines the recurrent neural network with the conventional hypergraph neural network to strengthen connections between utterances and make each utterance receive information from other utterances better. The proposed model has empirically achieved state-of-the-art results on three benchmark datasets, demonstrating the effectiveness and superiority of the new model. Cheng Zheng 0006, Xiao Sun 0003 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2024 | Detecting Depression With Heterogeneous Graph Neural Network in Clinical Interview TranscriptabstractDepression has an intense impact on individuals, yet many cases go undiagnosed. Thus, it is imperative to design an effective model for the automated diagnosis of depression. However, existing methods do not adequately capture contextual information in a clinical interview. Inspired by the depression diagnosis process, we propose a new perspective on detecting depression as a dialog information extraction task. Specifically, this article constructs a heterogeneous graph that models the participant’s depression state and uses the graph attention network to aggregate the pieces of depressive clues. In addition, we use the focal loss as a loss function for dealing with class imbalance by reshaping the standard cross-entropy loss. Experimental results demonstrate that our proposed model depression state extraction with heterogeneous graph attention neural network (DSE-HGAT) surpasses the baseline models on the Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WOZ) dataset. Meanwhile, agreement analysis between our proposed model and the gold standard shows that it is moderate ($k$= 0.528,$p 0.05$). Overall, our model is very effective in identifying depression in the clinical interview transcript, which has the potential to assist doctors with medical conditions. Mingzheng Li, Xiao Sun 0003, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Dynamic Emotional Transition Sampling and Emotional Guidance of Individuals Based on ConversationabstractEmotion is the vital path to achieving strong artificial intelligence. Therefore, it is significant to study the emotional guiding and controlling theory to enhance system intelligence. Conversational datasets in social media contain useful information, including individuals exhibiting continuously changing emotion states in response to external stimuli, and are the foundation for research on artificial emotions. We define dialog-based emotional guidance in reinforcement learning to research the emotional guidance of the discrete emotion model. This article proposes three strategies to obtain the optimal policy: 1) given the current emotional transition matrix, use the emotional Markov decision process (E-MDP) algorithm to calculate the optimal stimulus policy for each target emotion; 2) given the emotional transition sequences, use the emotional Monte Carlo algorithm to calculate the optimal stimulus policy; and 3) given the emotional transition sequences, use the$Q$-learning algorithm to calculate the optimal stimulus policy. Besides, we improve the Markov chain Monte Carlo algorithm to sample emotional transition sequences and design a metric to evaluate the effectiveness of policies. Experimental results on three datasets show that these methods can more effectively guide emotion than traditional methods. Particularly, E-MDP achieves the best results while others can be more widely used in real-world scenarios. Xiao Sun 0003, Fuji Ren, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | SHADE: Speaker-History-Aware Dialog Generation Through Contrastive and Prompt LearningabstractTraditional persona-based dialog generation models typically rely on textual descriptions of persona information, but this approach is limited by this and its inability to capture all aspects of a persona’s speaking style. To solve this problem, we propose a novel dialog generation model called Speaker-History-Aware Dialogue GEneration (SHADE) through Contrastive and Prompt Learning, which utilizes contrastive learning to model speaking style from historical conversation, resulting in more personalized and distinctive response. Based on the idea of, Prompt learning, on the one hand, we embed small-scale trainable parameters according to Continuous Prompts to stimulate the knowledge of the pre-trained language model, so as to improve the feature extraction ability of the model. On the other hand, we add speaker tag to the decoder input according to Discrete Prompts to improve the consistency of speaking style. We evaluate our model on two popular text-only datasets, DailyDialog and PersonaChat. The results show that SHADE outperforms other multiturn hierarchical dialog generation models in terms of response consistency, logic, and diversity. SHADE is optimal for metrics PPL, NIST, ROUGEL, METEOR, BLUE2, etc. Our proposed SHADE offers a promising direction for persona-based dialog generation, addressing the limitation of existing approaches and paving the way for more personalized and engaging conversational agents in various applications. Futian Wang, Xiao Sun 0003 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | AST-GCN: Augmented Spatial Temporal Graph Convolutional Neural Network for Gait Emotion RecognitionabstractSkeleton-based methods have recently achieved good performance in deep learning-based gait emotion recognition (DL-GER). However, the current methods have two drawbacks that limit the ability to learn discriminative emotional features from gait. First, these methods do not exclude the effect of the subject’s walking orientation on emotion classification. Second, they do not sufficiently learn the implicit connections between the joints during human walking. In this paper, an augmented spatial-temporal graph convolutional neural network (AST-GCN) is introduced to solve these two problems. The interframe shift encoding (ISE) module acquires interframe shifts of joints to make the network sensitive to changes in emotion-related joint movements regardless of the subject’s walking orientation. A multichannel implicit connection inference method learns more implicit connection relations related to emotions. Notably, we unify current skeleton-based methods into a common framework that validates the most powerful feature representation capability of our AST-GCN from a theoretical perspective. In addition, we extend the skeleton-based gait dataset using posture estimation software. Experiments demonstrate that our AST-GCN outperforms state-of-the-art methods on three datasets on two tasks. Xiao Sun 0003, Zhengzheng Tu, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Channel-Wise Interactive Learning for Remote Heart Rate Estimation From Facial VideoabstractRemote photoplethysmography measurement (also called rPPG prediction) is a vision-based technique that allows for the non-contact monitoring of human physiological activity using facial video. However, precisely detecting subtle color changes on facial skin, especially in less-constrained real-life scenarios, remains a formidable challenge for rPPG prediction. In this work, we address a rPPG-based heart rate estimation task by proposing an end-to-end Channel-wise Interaction Network (CIN-rPPG), in which the core idea contains two specialized units: channel-temporal interactive learning (CIT) and channel-spatial interactive learning (CIS). The CITunit gets the periodicity of the rPPG signal by using temporal-wise shifting and channel-wise scaling to measure the interaction between channels and temporal dimensions. The CISunit does both spatial-wise scaling and channel-wise scaling at the same time to perform channel-spatial interaction. This is intended to reveal how rPPG-related visual responses are detected on the human face. We exploit the rPPG recovery through the alternation of CITand CISimplementations. The CIN-rPPG is completely conducted by convolutional operations on the sequential 2D feature maps of facial video in an end-to-end manner. Extensive experiments on three heart rate estimation datasets (UBFC-rPPG, PURE, and MMSE-HR) demonstrate that CIN-rPPG achieves state-of-the-art performance on both intra-dataset and cross-dataset testing. Qi Li 0045, Dan Guo 0001, Xilan Tian, Xiao Sun 0003, Haifeng Zhao 0001, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | KeystrokeSniffer: An Off-the-Shelf Smartphone Can Eavesdrop on Your Privacy From AnywhereabstractWith mobile phones becoming increasingly prevalent and embedding high-quality microphones, attackers have the ability to employ these microphones to eavesdrop user’s keyboard input. However, existing work usually assumes that keystroke eavesdropping is performed against known environments and victims, which inevitably makes attack systems lack generalization. To reveal the real threat of the acoustic signal-based attack strategy, this paper proposes a keystroke eavesdropping algorithm called KeystrokeSniffer, which is robust to unknown input environments and unknown victims. In particular, to mimic the real input environment of victims, an environment estimation algorithm is first designed by extracting the timbre-related characteristics to predict the keyboard type and identifying large-size key data from collected unlabeled samples to estimate the 3D microphone coordinates. Then, by imitating unknown environments and victim data, this algorithm achieves effective keystroke eavesdropping with a small training set. By further considering the commonalities of different keystroke habits, a robust feature extraction method that reflects the keystroke location is adopted to reduce the impact of individual input habits. Extensive experimental results using various commodity smartphones indicate that the scheme is capable of predicting keyboard input accurately under different unknown scenarios. Specifically, even when both the victims and keyboards are unknown, KeystrokeSniffer can still achieve high Top-5 accuracy, reaching 79.5% in predicting keystrokes and 96.7% in predicting meaningful words, which demonstrates KeystrokeSniffer has excellent generalization capabilities. By setting different parameter values of various impact factors, e.g., noise and hand length factors, the strong robustness of the system is demonstrated, which proves that KeystrokeSniffer can violate privacy in real situations. Jinyang Huang, Jia-Xuan Bai, Xiang Zhang 0011, Zhi Liu 0002, Yuanhao Feng, Jianchun Liu, Xiao Sun 0003, Mianxiong Dong, Meng Li 0006 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2023 | Conditional Convolution Residual Network for Efficient Super-Resolution
Yunsheng Guo, Jinyang Huang, Xiang Zhang 0011, Xiao Sun 0003, Yu Gu 0003 |
ICANN (10) | 4 |
| 2023 | Dynamic Memory-Based Continual Learning with Generating and Screening
Siying Tao, Jinyang Huang, Xiang Zhang 0011, Xiao Sun 0003, Yu Gu 0003 |
ICANN (3) | 4 |
| 2023 | GLM: A Model Based on Global-Local Joint Learning for Emotion Recognition from Gaits Using Dual-Stream Network
Xiao Sun 0003 |
ICIG (1) | 2 |
| 2023 | CoLRP: A Contrastive Learning Abstractive Text Summarization Method with ROUGE PenaltyabstractContrastive learning can reduce the impact of ex-posure bias associated with training using maximum likelihood estimation, which aims to pull together positive samples to increase the likelihood of high-quality summaries and push away irrelevant negative samples to reduce the likelihood of low-quality summaries. In contrastive learning-based text summarization methods, a standard method for selecting positive and negative samples is randomly selected within a batch. This method can lead to sampling bias to the extent that the consistency of the representation space is compromised. Therefore, we propose a new method to penalize false negatives based on ROUGE metric scores as weights to sample from the dynamic output of the model training process. The method calculates ROUGE metric scores for penalizing false negatives in real-time and can distinguish between positive and negative samples to ensure spatial consistency and alleviate exposure bias. Experimental results on XSum, CNN/DM, and Multi-News datasets show that our approach effectively improves the performance of the latest text summarization pre-training models. Caidong Tan, Xiao Sun 0003 |
IJCNN | 2 |
| 2023 | Attention and Relative Distance Alignment for Low-Resolution Facial Expression Recognition
Liuwei An, Xiao Sun 0003, Meng Wang 0001 |
PRCV (5) | 2 |
| 2023 | ASM: Adaptive Sample Mining for In-The-Wild Facial Expression Recognition
Xiao Sun 0003, Liuwei An, Meng Wang 0001 |
PRCV (5) | 2 |
| 2023 | Dynamic Facial Expression Recognition Based on Vision Transformer with Deformable ModuleabstractFacial expressions convey a great deal of information during human emotional interaction. However, due to the potential for various types of in-the-wild interferences, such as occlusions and variant head poses, dynamic facial expression recognition (DFER) has been a desperately complicated task. Previous methods focus on applying more robust models to extract the spatial-temporal features but ignore the impact of the key features of the regions of interest (ROIs). This inhibits further improvement of recognition accuracy. In this paper, we propose a 3D vision Transformer with a deformable module termed 3D-DSwin Transformer to guide our model to capture more discriminative features. The deformable module can gradually shift the deformable points to guide our model to pay more attention to the ROIs. A simple yet effective video augmentation method is proposed to expand the number of training samples and avoid overfitting. Visualizations and extensive experimental results demonstrate that our proposed 3D-DSwin Transformer has the ability to obtain the key feature maps, and outperforms the previous state-of-the-art methods on both the FERV39k and DFEW benchmarks. Rui Wang 0156, Xiao Sun 0003 |
SMC | 2 |
| 2023 | Robust facial expression recognition with global-local joint representation learning
Chunxiao Fan 0002, Jia Li 0013, Xiao Sun 0003 |
Multim. Syst. | 5 |
| 2023 | Micro-expression recognition with attention mechanism and region enhancement
Shixin Zheng, Xiao Sun 0003, Dan Guo 0001, Junjie Lang |
Multim. Syst. | 3 |
| 2023 | User-based Hierarchical Network of Sina Weibo Emotion AnalysisabstractEmotion analysis on Sina Weibo has a great impetus for government agencies to survey public opinion and enterprises to track market demand. Most of the existing emotion analysis work on Sina Weibo focuses on mining the information contained in a single Weibo, ignoring the problem of inaccurate information extraction caused by the lack of contextual information in Weibo texts. Inspired by humans judging user emotional states from Weibo texts, this article creates a Weibo text five-category emotion classification dataset based on active users and proposes a user-based hierarchical network for Weibo emotion analysis. First, use the multi-head attention mechanism and convolutional neural network set in the information extraction module to analyze a single Weibo text to fully extract the emotional information contained in the text; at the same time, through the moving window set in the relevant information capture module, obtain other Weibo texts posted by the same user within a period, and capture the effective correlation information between Weibo texts; then, the dual text representation obtained above is concatenated, and through the information interaction layer, the relevant information is retrieved again, and the text representation is updated; finally, the classifier output the five-category emotion labels corresponding to each Weibo text. We demonstrate the model’s effectiveness through experiments and analysis in the results. Qian Chen 0033, Xiao Sun 0003, Meng Wang 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | Speaker-Aware Interactive Graph Attention Network for Emotion Recognition in ConversationabstractRecently, E motion R ecognition in C onversation (ERC) has attracted much attention and has become a hot topic in the field of natural language processing. Conversation is conducted in chronological order; current utterance is more likely influenced by nearby utterances. At the same time, speaker dependency also plays a core role in the conversation dynamic. The combined effect of the sequence-aware information and the speaker-aware information makes the emotion’s dynamic change. However, past works used simple information fusion methods to model the two kinds of information but ignored their interactive influence. Thus, we propose a novel method entitled SIGAT ( S peaker-aware I nteractive G raph A ttention Ne t work) to solve the problem. The core module is a mutual interactive module in which a dual-connection (self-connection and interact-connection) graph attention network is constructed. The advantage of SIGAT is modeling the speaker-aware and sequence-aware information in a unified graph and updating them simultaneously. In this way, we model the interactive influence of them and obtain the final representations, which have richer contextual clues. Experimental results on the four public datasets demonstrate that SIGAT outperforms the state-of-the-art models. Yunwei Shi, Weifeng Liu 0014, Zhenhua Huang 0002, Xiao Sun 0003 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2023 | Exploiting Japanese-Chinese Cognates with Shared Private Representations for NMTabstractNeural machine translation has achieved remarkable progress over the past several years; however, little attention has been paid to machine translation (MT) between Japanese and Chinese, which share a large proportion of cognate words that can be utilized as additional linguistic knowledge to enhance translation performance. In this article, we seek to strengthen the semantic correlation between Japanese and Chinese by leveraging cognate words that share common Chinese characters. Specifically, we experiment with three strategies: (1) a shared vocabulary with cognate lexicon induction, which models the commonality between source and target cognates; (2) a shared private representation with a dynamic gating mechanism, which models the language-specific features on the source side; and (3) an embedding shortcut, which enables the decoder to access the shared private representation with shortest distance and aids the training process. The experiments and analysis presented in this article demonstrate that our proposed approaches can significantly improve the performance of both Japanese-to-Chinese and Chinese-to-Japanese translations and verify the effectiveness of exploiting Japanese–Chinese cognates for MT. Fuji Ren, Xiao Sun 0003, Degen Huang, Piao Shi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | Multilingual BERT-based Word Alignment By Incorporating Common Chinese CharactersabstractWord alignment is an important task of detecting translation equivalents between a sentence pair. Although word alignment is no longer necessarily needed for neural machine translation, it’s still useful in a wealth of applications, e.g., bilingual lexicon induction, constraint decoding, and so on. However, the most well-known word aligners are still Giza++ and fastAlign, both of which are implementations of traditional IBM models. To keep pace with the advance in NMT, there has been a surge of interest in replacing the IBM models with neural models. We follow this trend but aim to boost performance of word alignment between Japanese and Chinese, which share a large portion of Chinese characters. Our key idea is to leverage these common Chinese characters in both languages as an indicator for inferring alignment; i.e., the source and target words with the common Chinese characters should be most likely aligned. Following this idea, we propose three methods that leverage common Chinese characters to boost the mBERT-based word alignment, including reward factor, representation alignment, and contrastive training. Furthermore, we annotate and release a golden dataset for Japanese-Chinese word alignment. Experiments on the dataset show that our methods outperform several strong baselines in terms of AER score and verify the effectiveness of exploiting common Chinese characters. Xiao Sun 0003, Fuji Ren, Degen Huang, Piao Shi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | Channel Attention TextCNN with Feature Word Extraction for Chinese Sentiment AnalysisabstractChinese short text sentiment analysis can help understand society’s views on various hot topics. Many existing sentiment analysis methods are based on sentiment dictionaries. Still, sentiment dictionaries are easily affected by subjective factors. They require a lot of time to build as well as maintenance to prevent obsolescence. For the aim of extracting rich information within texts more effectively, we propose a Channel Attention TextCNN with Feature Word Extraction model (CAT-FWE). The feature word extraction module helps us choose words that affect the sentiment of reviews. Then, these words are integrated with multi-level semantic information to enhance the information of sentences. In addition, the channel attention textCNN module that is a promotion of traditional TextCNN tends to pay more attention to those meaningful features. It eliminates the impacts of features that do not make any sense effectively. We apply our CAT-FWE model to both fine-grained classification and binary classification tasks for Chinese short texts. Experiment results show that it can improve the performance of emotion recognition. Jiangwei Liu, Zian Yan, Sibao Chen 0001, Xiao Sun 0003, Bin Luo 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | DIEU: A Dynamic Interaction Emotion Unit for Emotion Recognition in ConversationabstractEmotion recognition in conversation (ERC) is challenging because the conversation takes place in real time and the speakers interact with each other. However, existing methods ignore the dynamic characteristics of interaction between speakers, and the problem of long-range context propagation still exists. In this article, we propose a dynamic interaction emotion unit to solve the preceding problems on the transcription of the conversation. First, we propose a main influence interval search algorithm to provide a dynamic interaction interval for each utterance. Then, we utilize the speaker-aware influence module and the two-stream context module to capture the dynamic interaction and the contextual information from this interval. Furthermore, to obtain the speaker state representation rich in emotional information, we propose a novel dynamic routing algorithm to fuse the preceding information. These well-integrated state representations also enable our model to capture contextual information at a longer distance. Experiments on multiple datasets demonstrate the effectiveness of the proposed method. Shu Zhao 0005, Weifeng Liu 0014, Jie Chen 0025, Xiao Sun 0003 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | MHG-ERC: Multi-hypergraph Feature Aggregation Network for Emotion Recognition in ConversationsabstractThe modeling of conversational context is an essential step in Emotion Recognition in Conversations (ERC). To maintain high performance and a low GPU memory consumption, this article proposes a new idea of using multiple hypergraphs to model the conversational context and designs a multi-hypergraph feature aggregation network for ERC. We use context window, speaker information, position information between utterances, and specific step size to construct different hyperedges. Then, various hypergraphs generated by different hyperedges are used to aggregate local and remote context information in turn. Experiments on two dialogue emotion datasets, IEMOCAP and MELD, demonstrate the effectiveness and superiority of this new model. In addition, our model requires only relatively low GPU memory consumption. Cheng Zheng 0006, Xiao Sun 0003 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | Information-Enhanced Hierarchical Self-Attention Network for Multiturn Dialog GenerationabstractTransformer structure has shown promising results in multiturn dialog generation. The self-attention mechanism can learn global dependencies but ignores local information, limiting the model’s ability to model context information. In this article, we propose an information-enhanced hierarchical self-attention network (IEHSA). In the word-level encoder, words in successive windows are automatically encoded as local information, and words with dependent words are automatically encoded as syntactic information, both of which are used to enhance word information and then feed into the self-attention mechanism. In the utterance-level encoder, adjacent utterance representations are automatically encoded as dialog structure information, and the self-attention mechanism is used to update the utterance representations. The context and masked response representation are then updated using the self-attention mechanism in the decoder. Finally, the correlation between context and reply is calculated and used in further decoding. We compared IEHSA with the current popular hierarchical model on several datasets, and the experiments show that the proposed method has substantial improvements in both metric-based and human evaluations. Xiao Sun 0003, Qian Chen 0033, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2022 | Mining Syntactic Relationships via Recursion and Wandering on A Dependency Tree for Aspect-Based Sentiment AnalysisabstractAspect-Based Sentiment Analysis aims to determine the emotional orientation of a particular aspect of an online comment. Considering that syntax and even sentences themselves are a specific graph structure, most researchers' recent work mainly focuses on the dependency tree. They construct sentences into tree structures according to dependency parser. However, because a sentence can contain many aspects, it is difficult to correctly associate each aspect with the corresponding important part of the sentence by directly using the original dependency tree. To strengthen this association, a reconstructed dependency tree with aspect as the root is constructed by fine-tuning the original dependency tree for each aspect. We propose a neural network model named ARWAT, which performs on the reconstructed tree to learn informative contextual words and grammatical information effectively. Extensive experiment results demonstrate the superior performance of our proposed model against multiple baselines on five benchmark datasets. Xiao Sun 0003 |
IJCNN | 2 |
| 2022 | VFL - A deep learning-based framework for classifying walking gaits into emotions
Xiao Sun 0003, Chunxiao Fan 0002 |
Neurocomputing | 1 |
| 2022 | Multi-stage and multi-branch network with similar expressions label distribution learning for facial expression recognition
Junjie Lang, Xiao Sun 0003, Jia Li 0013, Meng Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | MAN: Mining Ambiguity and Noise for Facial Expression Recognition in the Wild
Xiao Sun 0003, Jia Li 0013, Meng Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Emotional Conversation Generation Orientated Syntactically Constrained Bidirectional-Asynchronous FrameworkabstractThe field of open-domain conversation generation using deep neural networks has attracted increasing attention from researchers for several years. However, traditional neural language models tend to generate safe, generic reply with poor logic and no emotion. In this paper, an emotional conversation generation orientated syntactically constrained bidirectional-asynchronous framework called E-SCBA is proposed to generate meaningful (logical and emotional) reply. In E-SCBA, pre-generated emotion keyword and topic keyword are asynchronously introduced into the reply during the generation, and the process of decoding is much different from the most existing methods that generates reply from the first word to the end. A newly designed bidirectional-asynchronous decoder with the multi-stage strategy is proposed to support this idea, which ensures the fluency and grammaticality of reply by making full use of syntactic constraint. Through the experiments, the results show that our framework not only improves the diversity of replies, but gains a boost on both logic and emotion compared with baselines as well. Xiao Sun 0003, Jianhua Tao 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Emotional Conversation Generation With Bilingual Interactive DecodingabstractThe perception and expression of emotion, as well as the content of a text, are key factors in the success of conversational agents. However, previous models for conversation generation handled single-language pairs during training and testing, neglecting the complementary information from different languages. In this article, we propose a bilingual-aided interactive approach that can simultaneously and interactively generate bilingual emotional replies to monolingual posts. Specifically, the generation of one emotional reply relies on the output of the encoder, the generated tokens, and the interactive information from the other language decoder. The interactive approach includes 1) internal interaction to capture the change in implicit contextual information and 2) external interaction to balance grammaticality and the expression of emotion. Qualitative and quantitative experiments with NLPCC2017 show that our model performs better in terms of the content and emotion of replies than several state-of-the-art approaches. Xiao Sun 0003, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2022 | Personality Assessment Based on Multimodal Attention Network Learning With Category-Based Mean Square ErrorabstractPersonality analysis is widely used in occupational aptitude tests and entrance psychological tests. However, answering hundreds of questions at once seems to be a burden. Inspired by personality psychology, we propose a multimodal attention network with Category-based mean square error (CBMSE) for personality assessment. With this method, we can obtain information about one's behaviour from his or her daily videos, including his or her gaze distribution, speech features, and facial expression changes, to accurately determine personality traits. In particular, we propose a new approach to implementing an attention mechanism based on the facial Region of No Interest (RoNI), which can achieve higher accuracy and reduce the number of network parameters. Simultaneously, we use CBMSE, a loss function with a higher penalty for the fuzzy boundary in personality assessment, to help the network distinguish boundary data. After effective data fusion, this method achieves an average prediction accuracy of 92.07%, which is higher than any other state-of-the-art model on the dataset of the ChaLearn Looking at People challenge in association with ECCV 2016. Xiao Sun 0003, Jie Huang 0038, Shixin Zheng, Xuanheng Rao, Meng Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Embedding Extra Knowledge and A Dependency Tree Based on A Graph Attention Network for Aspect-based Sentiment AnalysisabstractAspect-based sentiment analysis analyses the fine-grained sentiment polarity of a particular attribute in a sentence. In addition to the fact that most papers have applied attention mechanisms and neural networks to fine-grained sentiment analysis, some studies have reported outstanding performance from graph-structured neural networks in aspect-based sentiment analysis. In this paper, we propose the use of a graph attention network (GAT) to embed external knowledge and grammatical relations (KD-GAT). First, we employ a GAT to extract the nodes and edges related to sentences in the knowledge graph. Second, we consider the influence of conjunctions and parts of speech based on the aspect-oriented syntactic dependency tree. Finally, we emphasize the effect of aspect position in the mechanism of multiple attention. Experiments on the SemEval 2014 and Twitter datasets show that our model can improve the ability to analyze sentiment and that the graph structure method can better integrate grammatical relations to understand a sentence. Xiao Sun 0003, Meng Wang 0001 |
IJCNN | 2 |
| 2021 | Multi-attention based Deep Neural Network with hybrid features for Dynamic Sequential Facial Expression Recognition
Xiao Sun 0003, Pingping Xia, Fuji Ren |
Neurocomputing | 1 |
| 2021 | Design and Analysis of a Human-Machine Interaction System for Researching Human's Dynamic EmotionabstractDynamic emotion is typically used to facilitate human–machine interactions. Conversational data from social media contain a considerable amount of useful information, and such data are the foundation for researching dynamic and artificial emotion. At present, most human–machine interaction systems focus on the complexity and accuracy of the dialog but neglect the emotional characteristics of the speaker. When generating a dialog considering the emotional personality of the interlocutor, controlling, and guiding the dialog to a specified direction are essential. This article presents a system for studying dynamic emotions in human-computer interaction from the perspective of emotional transfer and guidance. Based on the emotional state of the interlocutor and the distribution of emotional transfer, the process of emotional transfer is simulated and sampled, and the sequence of emotional guidance is generated. In this system, two algorithms are proposed. A generative Markov chain Monte Carlo (GEN-MCMC) algorithm is proposed to generate a variety of emotional transfer sequences that fit the talke’s personality dynamically based on the real-world dialog. Further, a guiding MCMC (GUI-MCMC) algorithm-based GEN-MCMC is proposed to generate the emotional guiding sequences. The generated emotional sequences by GEN-MCMC were evaluated in two aspects: 1) consistency and 2) diversity. The experimental results show that the GEN-MCMC algorithm performs better than the general sequence generation algorithm in terms of consistency and diversity in generating emotional states. The GUI-MCMC was able to generate a proper stimulus sequence when given the first and target emotions. An emotional stimulus sequence can simulate the emotional transfer of the interlocutor in the process of dialogue, and give the observer appropriate reference to guide and control the emotions of dialogue. The experimental results show that the proposed system can effectively model the dynamic emotion in emotional transfer and guidance, which can be further used to build chat robots, intelligent assistants, and human–machine interaction systems. The models can also be used for emotional induction and enhance the feel-good or feel-terrible factor in human–machine communication applications, such as medical treatment of mental diseases, interrogation, and psychological attack and defense. Xiao Sun 0003, Zhengmeng Pei, Chen Zhang 0013, Guoqiang Li 0001, Jianhua Tao 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2021 | Emotional dialog generation via multiple classifiers based on a generative adversarial networkabstractHuman-machine dialog generation is an essential topic of research in the field of natural language processing. Generating high-quality, diverse, fluent, and emotional conversation is a challenging task. Based on continuing advancements in artificial intelligence and deep learning, new methods have come to the forefront in recent times. In particular, the end-to-end neural network model provides an extensible conversation generation framework that has the potential to enable machines to understand semantics and automatically generate responses. However, neural network models come with their own set of questions and challenges. The basic conversational model framework tends to produce universal, meaningless, and relatively "safe" answers. Based on generative adversarial networks (GANs), a new emotional dialog generation framework called EMC-GAN is proposed in this study to address the task of emotional dialog generation. The proposed model comprises a generative and three discriminative models. The generator is based on the basic sequence-to-sequence (Seq2Seq) dialog generation model, and the aggregate discriminative model for the overall framework consists of a basic discriminative model, an emotion discriminative model, and a fluency discriminative model. The basic discriminative model distinguishes generated fake sentences from real sentences in the training corpus. The emotion discriminative model evaluates whether the emotion conveyed via the generated dialog agrees with a pre-specified emotion, and directs the generative model to generate dialogs that correspond to the category of the pre-specified emotion. Finally, the fluency discriminative model assigns a score to the fluency of the generated dialog and guides the generator to produce more fluent sentences. Based on the experimental results, this study confirms the superiority of the proposed model over similar existing models with respect to emotional accuracy, fluency, and consistency. The proposed EMC-GAN model is capable of generating consistent, smooth, and fluent dialog that conveys pre-specified emotions, and exhibits better performance with respect to emotional accuracy, consistency, and fluency compared to its competitors. Wei Chen 0089, Xinmiao Chen, Xiao Sun 0003 |
Virtual Real. Intell. Hardw. | 3 |
| 2020 | A ROI-guided deep architecture for robust facial expressions recognition
Xiao Sun 0003, Pingping Xia, Ling Shao 0001 |
Inf. Sci. | 1 |
| 2020 | A novel approach to generate a large scale of supervised data for short text sentiment analysis
Xiao Sun 0003, Jiajin He |
Multim. Tools Appl. | 1 |
| 2020 | Emotional Conversation Generation Based on a Bayesian Deep Neural NetworkabstractThe field of conversation generation using neural networks has attracted increasing attention from researchers for several years. However, traditional neural language models tend to generate a generic reply with poor semantic logic and no emotion. This article proposes an emotional conversation generation model based on a Bayesian deep neural network that can generate replies with rich emotions, clear themes, and diverse sentences. The topic and emotional keywords of the replies are pregenerated by introducing commonsense knowledge in the model. The reply is divided into multiple clauses, and then a multidimensional generator based on the transformer mechanism proposed in this article is used to iteratively generate clauses from two dimensions: sentence granularity and sentence structure. Subjective and objective experiments prove that compared with existing models, the proposed model effectively improves the semantic logic and emotional accuracy of replies. This model also significantly enhances the diversity of replies, largely overcoming the shortcomings of traditional models that generate safe replies. Xiao Sun 0003, Jia Li 0013, Xing Wei 0002, Changliang Li, Jianhua Tao 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2019 | Hybrid spatiotemporal models for sentiment classification via galvanic skin response
Xiao Sun 0003, Changliang Li, Fuji Ren |
Neurocomputing | 1 |
| 2019 | Improved SMOTE Algorithm to Deal with Imbalanced Activity Classes in Smart Homes
Shikai Guo, Rong Chen 0003, Xiao Sun 0003, Xiangxin Wang |
Neural Process. Lett. | 4 |
| 2019 | Scene Categorization Using Deeply Learned Gaze Shifting KernelabstractAccurately recognizing sophisticated sceneries from a rich variety of semantic categories is an indispensable component in many intelligent systems, e.g., scene parsing, video surveillance, and autonomous driving. Recently, there have emerged a large quantity of deep architectures for scene categorization, wherein promising performance has been achieved. However, these models cannot explicitly encode human visual perception toward different sceneries, i.e., the sequence of humans sequentially allocates their gazes. To solve this problem, we propose deep gaze shifting kernel to distinguish sceneries from different categories. Specifically, we first project regions from each scenery into the so-called perceptual space, which is established by combining color, texture, and semantic features. Then, a novel non-negative matrix factorization algorithm is developed which decomposes the regions' feature matrix into the product of the basis matrix and the sparse codes. The sparse codes indicate the saliency level of different regions. In this way, the gaze shifting path from each scenery is derived and an aggregation-based convolutional neural network is designed accordingly to learn its deep representation. Finally, the deep representations of gaze shifting paths from all the scene images are incorporated into an image kernel, which is further fed into a kernel SVM for scene categorization. Comprehensive experiments on six scenery data sets have demonstrated the superiority of our method over a series of shallow/deep recognition models. Besides, eye tracking experiments have shown that our predicted gaze shifting paths are 94.6% consistent with the real human gaze allocations. Xiao Sun 0003, Zepeng Wang 0003, Jie Chang 0001, Yiyang Yao, Ping Li 0006, Roger Zimmermann |
IEEE Trans. Cybern. | 1 |
| 2018 | A Syntactically Constrained Bidirectional-Asynchronous Approach for Emotional Conversation GenerationabstractTraditional neural language models tend to generate generic replies with poor logic and no emotion.In this paper, a syntactically constrained bidirectional-asynchronous approach for emotional conversation generation (E-SCBA) is proposed to address this issue.In our model, pre-generated emotion keywords and topic keywords are asynchronously introduced into the process of decoding.It is much different from most existing methods which generate replies from the first word to the last.Through experiments, the results indicate that our approach not only improves the diversity of replies, but gains a boost on both logic and emotion compared with baselines. Xiao Sun 0003 |
EMNLP | 2 |
| 2018 | Perceptual multi-channel visual feature fusion for scene categorization
Xiao Sun 0003, Zhenguang Liu, Yuxing Hu, Roger Zimmermann |
Inf. Sci. | 1 |
| 2018 | Camera-Assisted Video Saliency Prediction and Its ApplicationsabstractVideo saliency prediction is an indispensable yet challenging technique which can facilitate various applications, such as video surveillance, autonomous driving, and realistic rendering. Based on the popularity of embedded cameras, we in this paper predict region-level saliency from videos by leveraging human gaze locations recorded using a camera, (e.g., those equipped on an iMAC and laptop PC). Our proposed camera-assisted mechanism improves saliency prediction by discovering human attended regions inside a video clip. It is orthogonal to the current saliency models, i.e., any existing video/image saliency model can be boosted by our mechanism. First of all, the spatial-and temporal-level visual features are exploited collaboratively for calculating an initial saliency map. We notice that the current saliency models are not sufficiently adaptable to the variations in lighting, different view angles, and complicated backgrounds. Therefore, assisted by a camera tracking human gaze movements, a non-negative matrix factorization algorithm is designed to accurately localize the semantically/visually salient video regions perceived by humans. Finally, the learned human gaze locations as well as the initial saliency map are integrated to optimize video saliency calculation. Empirical results thoroughly demonstrated that: 1) our approach achieves the state-of-the-art video saliency prediction accuracy by outperforming 11 mainstream algorithms considerably and 2) our method can conveniently and successfully enhance video retargeting, quality estimation, and summarization. Xiao Sun 0003, Yuxing Hu, Ping Li 0006, Zhao Xie, Zhenguang Liu |
IEEE Trans. Cybern. | 1 |
| 2017 | Improved facial expression recognition method based on ROI deep convolutional neutral networkabstractThis paper, we proposed an improved facial expression recognition (FER) method based on region of interesting (ROI) to guide the convolutional neutral networks (CNN) focus on the areas associated with the expression. This method can not only augment the training data, the relationship between the different ROI areas is helpful to intensify the reliability of the predicted targets. In test stage, we investigated two recognition methods: identify the test image directly; implemented decision fusion strategy on ROI areas. The model we used is fine-tuned from pre-trained deep CNN instead of training from scratch. In addition, we presented an innovative region-based image augmentation method named artificial face to increase the limited database. This method using expression retargeting as an expression-preserving data augmentation which is specific for FER. The performance of the proposed method has been validated on the public CK+ databases. Xiao Sun 0003, Man Lv, Changqin Quan, Fuji Ren |
ACII | 1 |
| 2016 | Chinese micro-blog sentiment analysis based on semantic features and PAD modelabstractWith the increasing impact of social networks, microblog becomes important carrier of information and social interaction for human beings, which contains emotional states that have important research significance. We try to analysis the microblog text with the methods of emotional vocabulary, combining domain knowledge of psychology and affective computing, continuous dimension of emotion psychology PAD model which is adopted as basis of sentiment analysis. Emotional state inherent in the text is analyzed to obtain a more accurate result and achieve purposes of emotional analysis. At the same time, to achieve emotional microblog text computability from the aspect of personal characteristics. Experimental results show that the method can improve the microblog text sentiment analysis accuracy and precision. The method is able to get a good application in the different themes and different emotional features. Xiao Sun 0003, Kunxia Wang, Fuji Ren |
ICIS | 2 |
| 2016 | Detect the emotions of the public based on cascade neural network modelabstractAlong with the development of social network, more and more people know the world by reading news. The problem about what kind of emotion is inspired when people read news is very worthy of discussion. This paper will mix Deep Belief Networks (DBN) model and Support Vector Machine (SVM) to a hybrid neural network model by using the Contrast Divergence (CD) algorithm to estimate the weights when training a generating model, ensure that each layer of the Restricted Boltzmann Machine (RBM) mapping the features of the inputs to the best. At the same time, we cascade the last layer of DBN and a SVM classifier to adjust judging performance. And a set of tags will be attached to the top (Associative Memory), through a process of parameter tuning, learn the identifying weights to obtain a network for the task of text classification. The experimental results show that the hybrid neural network model works better than the traditional text categorization method based on simple characteristics (such as CHI), and it is more suitable for extracting text semantic characteristics. Xiao Sun 0003, Xiaoqi Peng, Fuji Ren |
ICIS | 1 |
| 2016 | Emotional element detection and tendency judgment based on mixed model with deep featuresabstractWith the rapid development of B2C e-commerce and the popularity of online shopping, the Web storages huge number of product reviews comment by customers. Product reviews contain subjective feelings of customers who have used some products, more and more customers browse a large number of online reviews in order to know other customers word-of-mouth of product and service to make an informed choice. Manufacturers also need accurate user feedback from product reviews to improve their goods. However, a large number of reviews made it difficult for manufacturers or potential customers to track the comments and suggestions that customers made. This paper presents a method for extracting emotional elements containing emotional objects and emotional words and their tendencies from product reviews based on mixed model. First we constructed conditional random fields (CRFs) to extract emotional elements, lead-in semantic and word meaning as features to improve the robustness of feature template and used rules for hierarchical filtering errors. Then we constructed support vector machine (SVM) to classify the emotional tendency of the fine-grained elements to achieve key information from product reviews. Deep semantic information imported based on neural network (NN) to improve the traditional bag of word model. Experimental results show that the proposed model with deep features efficiently improved the F-Measure. Xiao Sun 0003, Chongyuan Sun, Fuji Ren, Kunxia Wang |
ICIS | 1 |
| 2016 | A database for emotional interactions of the elderlyabstractEmotional interaction plays an important role in human-computer interaction domains. One of the major limitations in the study of emotion interaction is the lack of databases. This paper describes a database for emotion interactions of the elderly. The database was collected with audio and video from sixteen actors (8 female and 8 male) in daily conversations of TV series, which covers seven type emotions, namely, anger, boredom, happiness, sadness, surprise, neutrality and disgust. The database consists of 810 speech-video excerpts from 118 conversations. Emotion annotations were evaluated by 18 participants (10 men and 8 women). In addition, this databases record the transitions of actors' emotion in daily conversations. Kunxia Wang, ZongBao Zhu, Xiao Sun 0003 |
ICIS | 4 |
| 2016 | Sentiment analysis for Chinese microblog based on deep neural networks with convolutional extension features
Xiao Sun 0003, Fuji Ren |
Neurocomputing | 1 |
| 2016 | Detecting influenza states based on hybrid model with personal emotional factors from social networks
Xiao Sun 0003, Jia-qi Ye, Fuji Ren |
Neurocomputing | 1 |
| 2015 | Chinese microblog sentiment classification based on convolution neural network with content extension methodabstractRelated research for sentiment analysis on Chinese microblog is aiming at analyzing the emotion of posters. This paper presents a content extension method that combines post with its' comments into a microblog conversation for sentiment analysis. A new convolutional auto encoder which can extract contextual sentiment information from microblog conversation of the post is proposed. Furthermore, a DBN model, which is composed by several layers of RBM(Restricted Boltzmann Machine) stacked together, is implemented to extract some higher level feature for short text of a post. These RBM layers can encoder observed short text to learn hidden structures or semantics information for better feature representation. A ClassRBM (Classification RBM) layer, which is stacked on top of RBM layers, is adapted to achieve the final sentiment classification. The experiment results demonstrate that, with proper structure and parameter, the performance of the proposed deep learning method on sentiment classification is better than state-of-the-art surface learning models such as SVM or NB, which also proves that DBN is suitable for short-length document classification with the proposed feature dimensionality extension method. Xiao Sun 0003, Fuji Ren |
ACII | 1 |
| 2014 | Word Frequency Statistics Model for Chinese Base Noun Phrase Identification
Lu Kong, Fuji Ren, Xiao Sun 0003, Changqin Quan |
ICIC (2) | 3 |
| 2014 | Multi-strategy Based Sina Microblog Data Acquisition for Opinion Mining
Xiao Sun 0003, Jia-qi Ye, Fuji Ren |
ICIC (2) | 1 |
| 2011 | Chinese New Word Identification: A Latent Discriminative Model with Global Features
Xiao Sun 0003, Degen Huang, Fuji Ren |
J. Comput. Sci. Technol. | 1 |
| 2008 | HMM and CRF Based Hybrid Model for Chinese Lexical Analysis
Degen Huang, Xiao Sun 0003, Shidou Jiao, Lishuang Li, Zhuoye Ding, Ru Wan |
IJCNLP | 2 |