EDBT 2026 Demo / reviewers in the wild / expert
Minglei Li 0001
dblp:136/7341-1
· DBLP profile ↗
30ranked-venue papers
5as first author
20since 2021 · last 2025
0000-0002-1427-3507ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 14 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReMask-Animate: Refined Character Image Animation Using Mask-Guided AdaptersabstractPose-controlled human video generation is of significant interest and finds extensive applications in areas such as automated advertising and content creation on social media platforms. While existing methods employing pose sequences and reference images for human image animation have exhibited notable performance, they tend to encounter issues such as specific region blurring, background sharpening, and decreased identity consistency. In this paper, we introduce ReMask-Animate, which utilizes masks as additional priors to guide the model's local visual attention to specific areas, thereby alleviating feature confusion between different regions of the image. Three distinct mask-guided adapters are designed for cross-condition regional fusion of hand and face pose features, mitigating feature confusion between the foreground and background, and enhancing the visual consistency of character identity. Moreover, these lightweight adapters introduce minimal computational overhead and can be seamlessly integrated into specific layers of the backbone architecture. Extensive experiments show that our method outperforms state-of-the-art methods on five metrics in public datasets. Additionally, qualitative evaluations highlight a significant improvement in the quality of generated videos, demonstrating our approach's superiority. Xunzhi Xiang, Haiwei Xue, Zonghong Dai, Minglei Li 0001, Ye Yue, Fei Ma 0006, Weijiang Yu, Heng Chang, F. Richard Yu |
AAAI | 5 |
| 2025 | Identity-Preserving Audio-Driven Holistic Human Motion Video GenerationabstractGenerating realistic human motion videos is a pivotal challenge in advancing human-computer interaction. While existing approaches often focus on generating either head or gesture movements from audio, they lack unified control over full-body motion, frequently producing low-resolution and blurred outputs. Additionally, these methods struggle to maintain character identity throughout the generated content. In this paper, we introduce a novel framework that generates photorealistic, personalized human motion videos from audio by decoupling identity features. We integrate both visual features and voice timbre to enhance the preservation of character identity. Our approach follows a four-stage paradigm: (1) frame generation, (2) identity feature customization, (3) audio-motion modeling, and (4) motion-video rendering. Through the collaborative modeling of audio-motion and motion-video stages, our approach effectively maintains the consistency of character identity and background throughout the video, enhancing the realism and coherence of the generated video. Experimental results demonstrate that our framework delivers high-resolution videos with superior fidelity, establishing a new and effective baseline for holistic human motion video generation. Haiwei Xue, Zhensong Zhang, Minglei Li 0001, Zonghong Dai, Zhiyong Wu 0001 |
ICASSP | 3 |
| 2025 | VideoHumanMIB: Unlocking Appearance Decoupling for Video Human Motion In-betweeningabstractWe propose VideoHumanMIB, a novel framework for Video Human Motion In-betweening that enables seamless transitions between different motion video clips, facilitating the generation of longer and more natural digital human videos. While existing video frame interpolation methods work well for similar motions in adjacent frames, they often struggle with complex human movements, resulting in artifacts and unrealistic transitions. To address these challenges, we introduce a two-stage approach: First, we design an Appearance Reconstruction AutoEncoder to decouple appearance and motion information, extracting robust appearance-invariant features. Second, we develop an enhanced diffusion pretrained network that leverages both motion optical flow and human pose as guidance conditions, enabling the model to learn comprehensive latent distributions of possible motions. Rather than operating directly in pixel space, our model works in a learned latent space, allowing it to better capture the underlying motion dynamics. The framework is optimized with a dual-frame constraint loss and a motion flow loss to ensure temporal consistency and natural movement transitions. Extensive experiments demonstrate that our approach generates highly realistic transition sequences that significantly outperform existing methods, particularly in challenging scenarios with large motion variations. The proposed VideoHumanMIB establishes a new baseline for human motion synthesis and enables more natural and controllable digital human animation. Haiwei Xue, Zhensong Zhang, Minglei Li 0001, Zonghong Dai, F. Richard Yu, Fei Ma 0006, Zhiyong Wu 0001 |
IJCAI | 3 |
| 2025 | Human Motion Video Generation: A SurveyabstractHuman motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys in this field focus on individual methods, lacking a comprehensive overview of the entire generative process. This paper addresses this gap by providing an in-depth survey of human motion video generation, encompassing over ten sub-tasks, and detailing the five key phases of the generation process: input, motion planning, motion video generation, refinement, and output. Notably, this is the first survey that discusses the potential of large language models in enhancing human motion video generation. Our survey reviews the latest developments and technological trends in human motion video generation across three primary modalities: vision, text, and audio. By covering over two hundred papers, we offer a thorough overview of the field and highlight milestone works that have driven significant technological breakthroughs. Our goal for this survey is to unveil the prospects of human motion video generation and serve as a valuable resource for advancing the comprehensive applications of digital humans. Haiwei Xue, Xiangyang Luo 0002, Zhanghao Hu, Xin Zhang 0169, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li 0001, Jian Yang 0003, Fei Ma 0006, Zhiyong Wu 0001, Changpeng Yang, Zonghong Dai, F. Richard Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion ModelabstractCo-speech gestures, if presented in the lively form of videos, can achieve superior visual effects in human-machine interaction. While previous works mostly gener-ate structural human skeletons, resulting in the omission of appearance information, we focus on the direct gener-ation of audio-driven co-speech gesture videos in this work. There are two main challenges: 1) A suitable motion feature is needed to describe complex human movements with crucial appearance information. 2) Gestures and speech exhibit inherent dependencies and should be temporally aligned even of arbitrary length. To solve these problems, we present a novel motion-decoupled framework to gener-ate co-speech gesture videos. Specifically, we first intro-duce a well-designed nonlinear TPS transformation to ob-tain latent motion features preserving essential appearance information. Then a transformer-based diffusion model is proposed to learn the temporal correlation between gestures and speech, and performs generation in the latent motion space, followed by an optimal motion selection mod-ule to produce long-term coherent and consistent gesture videos. For better visual perception, we further design a refinement network focusing on missing details of cer-tain areas. Extensive experimental results show that our proposed framework significantly outperforms existing approaches in both motion and video-related evaluations. Our code, demos, and more resources are available at https://github.com/thuhcsi/S2G-MDDiffusion. Qiaochu Huang, Zhensong Zhang, Zhiyong Wu 0001, Minglei Li 0001, Songcen Xu |
CVPR | 7 |
| 2024 | Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion ModelsabstractAudio-driven co-speech human gesture generation has made remarkable advancements recently. However, most previous works only focus on single person audio-driven gesture generation. We aim at solving the problem of conversational co-speech gesture generation that considers multiple participants in a conversation, which is a novel and challenging task due to the difficulty of simultaneously incorporating semantic information and other relevant features from both the primary speaker and the interlocutor. To this end, we propose CoDiffuseGesture, a diffusion model-based approach for speech-driven interaction gesture generation via modeling bilateral conversational intention, emotion, and semantic context. Our method synthesizes appropriate interactive, speech-matched, high-quality gestures for conversational motions through the intention perception module and emotion reasoning module at the sentence level by a pretrained language model. Experimental results demonstrate the promising performance of the proposed method. Haiwei Xue, Zhensong Zhang, Zhiyong Wu 0001, Minglei Li 0001, Zonghong Dai, Helen M. Meng |
ICASSP | 5 |
| 2024 | Speak From Heart: An Emotion-Guided LLM-Based Multimodal Method for Emotional Dialogue GenerationabstractRecent advancements in Large Language Models~(LLMs) have greatly enhanced the generation capabilities of dialogue systems. However, progress on emotional expression during dialogues might be still limited, especially when capturing and processing the multimodal cues for emotional expression. Therefore, it is urgent to fully adapt the multimodal understanding ability and transferability of LLMs to enhance the emotional-oriented multimodal processing capabilities. To that end, in this paper, we propose a novel Emotion-Guided Multimodal Dialogue model based on LLM, termed ELMD. Specifically, to enhance the emotional expression ability of LLMs, our ELMD customizes an emotional retrieval module, which mainly provides appropriate response demonstration for LLM in understanding emotional context. Subsequently, a two-stage training strategy is proposed, founded on previous demonstration support, to support uncovering nuanced emotions behind multimodal information and constructing natural responses. Comprehensive experiments demonstrate the effectiveness and superiority of ELMD. Chenxiao Liu, Zheyong Xie, Sirui Zhao, Tong Xu 0001, Minglei Li 0001, Enhong Chen |
ICMR | 6 |
| 2024 | Bridging Gaps in Content and Knowledge for Multimodal Entity LinkingabstractMultimodal Entity Linking (MEL) aims to address the ambiguity in multimodal mentions and associate them with Multimodal Knowledge Graphs (MMKGs). Existing works primarily focus on designing multimodal interaction and fusion mechanisms to enhance the performance of MEL. However, these methods still overlook two crucial gaps within the MEL task. One is the content discrepancy between mentions and entities, manifested as uneven information density. The other is the knowledge gap, indicating insufficient knowledge extraction and reasoning during the linking process. To bridge these gaps, we propose a novel framework FissFuse, as well as a plug-and-play knowledge-aware re-ranking method KAR. Specifically, FissFuse collaborates with the Fission and Fusion branches, establishing dynamic features for each mention-entity pair and adaptively learning multimodal interactions to alleviate content discrepancy. Meanwhile, KAR is endowed with carefully crafted instruction for intricate knowledge reasoning, serving as re-ranking agents empowered by Large Language Models (LLMs). Extensive experiments on two well-constructed MEL datasets demonstrate outstanding performance of FissFuse compared with various baselines. Comprehensive evaluations and ablation experiments validate the effectiveness and generality of KAR. Pengfei Luo, Tong Xu 0001, Che Liu 0001, Suojuan Zhang, Linli Xu 0002, Minglei Li 0001, Enhong Chen |
ACM Multimedia | 6 |
| 2024 | SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local EditingabstractAudio-driven talking face generation aims to synthesize video with lip movements synchronized to input audio. However, current generative techniques face challenges in preserving intricate regional textures (skin, teeth). To address the aforementioned challenges, we propose a novel framework called SegTalker to decouple lip movements and image textures by introducing segmentation as intermediate representation. Specifically, given the mask of image employed by a parsing network, we first leverage the speech to drive the mask and generate talking segmentation. Then we disentangle semantic regions of image into style codes using a mask-guided encoder. Ultimately, we inject the previously generated talking segmentation and style codes into a mask-guided StyleGAN to synthesize video frame. In this way, most of textures are fully preserved. Moreover, our approach can inherently achieve background separation and facilitate mask-guided facial local editing. In particular, by editing the mask and swapping the region textures from a given reference image (e.g. hair, lip, eyebrows), our approach enables facial editing seamlessly when generating talking face video. Experiments demonstrate that our proposed approach can effectively preserve texture details and generate temporally consistent video while remaining competitive in lip synchronization. Quantitative and qualitative results on the HDTF and MEAD datasets illustrate the superior performance of our method over existing methods. Lingyu Xiong, Xize Cheng, Jintao Tan, Xianjia Wu, Xiandong Li, Fei Ma 0006, Minglei Li 0001, Huang Xu 0003 |
ACM Multimedia | 8 |
| 2023 | Co-speech Gesture Synthesis by Reinforcement Learning with Contrastive Pretrained RewardsabstractThere is a growing demand of automatically synthesizing co-speech gestures for virtual characters. However, it remains a challenge due to the complex relationship between input speeches and target gestures. Most existing works focus on predicting the next gesture that fits the data best, however, such methods are myopic and lack the ability to plan for future gestures. In this paper, we propose a novel reinforcement learning (RL) framework called RACER to generate sequences of gestures that maximize the overall satisfactory. RACER employs a vector quantized variational autoencoder to learn compact representations of gestures and a GPT-based policy architecture to generate coherent sequence of gestures autoregressively. In particular, we propose a contrastive pre-training approach to calculate the rewards, which integrates contextual information into action evaluation and successfully captures the complex relationships between multi-modal speech-gesture data. Experimental results show that our method significantly outperforms existing baselines in terms of both objective metrics and subjective human judgements. Demos can be found at https://github.com/RLracer/RACER.git. Mengchen Zhao, Yaqing Hou, Minglei Li 0001, Huang Xu 0003, Songcen Xu, Jianye Hao |
CVPR | 4 |
| 2023 | QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture GenerationabstractSpeech-driven gesture generation is highly challenging due to the random jitters of human motion. In addition, there is an inherent asynchronous relationship between human speech and gestures. To tackle these challenges, we introduce a novel quantization-based and phase-guided motion matching framework. Specifically, we first present a gesture VQ-VAE module to learn a codebook to summarize meaningful gesture units. With each code representing a unique gesture, random jittering problems are alleviated effectively. We then use Levenshtein distance to align diverse gestures with different speech. Levenshtein distance based on audio quantization as a similarity metric of corresponding speech of gestures helps match more appropriate gestures with speech, and solves the alignment problem of speech and gestures well. Moreover, we introduce phase to guide the optimal gesture matching based on the semantics of context or rhythm of audio. Phase guides when text-based or speech-based gestures should be performed to make the generated gestures more natural. Extensive experiments show that our method outperforms recent approaches on speech-driven gesture generation. Our code, database, pre-trained models and demos are available at https://github.com/YoungSeng/QPGesture. Zhiyong Wu 0001, Minglei Li 0001, Zhensong Zhang, Weihong Bao, Haolin Zhuang |
CVPR | 3 |
| 2023 | DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion ModelsabstractThe art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to the diversity of gestures and the difficulty of matching the rhythm and semantics of the gesture to the corresponding speech. To address these problems, we present DiffuseStyleGesture, a diffusion model based speech-driven gesture generation approach. It generates high-quality, speech-matched, stylized, and diverse co-speech gestures based on given speeches of arbitrary length. Specifically, we introduce cross-local attention and self-attention to the gesture diffusion pipeline to generate better speech matched and realistic gestures. We then train our model with classifier-free guidance to control the gesture style by interpolation or extrapolation. Additionally, we improve the diversity of generated gestures with different initial gestures and noise. Extensive experiments show that our method outperforms recent approaches on speech-driven gesture generation. Our code, pre-trained models, and demos are available at https://github.com/YoungSeng/DiffuseStyleGesture. Zhiyong Wu 0001, Minglei Li 0001, Zhensong Zhang, Weihong Bao, Long Xiao |
IJCAI | 3 |
| 2023 | UnifiedGesture: A Unified Gesture Synthesis Model for Multiple SkeletonsabstractThe automatic co-speech gesture generation draws much attention in computer animation. Previous works designed network structures on individual datasets, which resulted in a lack of data volume and generalizability across different motion capture standards. In addition, it is a challenging task due to the weak correlation between speech and gestures. To address these problems, we present UnifiedGesture, a novel diffusion model-based speech-driven gesture synthesis approach, trained on multiple gesture datasets with different skeletons. Specifically, we first present a retargeting network to learn latent homeomorphic graphs for different motion capture standards, unifying the representations of various gestures while extending the dataset. We then capture the correlation between speech and gestures based on a diffusion model architecture using cross-local attention and self-attention to generate better speech-matched and realistic gestures. To further align speech and gesture and increase diversity, we incorporate reinforcement learning on the discrete gesture units with a learned reward function. Extensive experiments show that UnifiedGesture outperforms recent approaches on speech-driven gesture generation in terms of CCA, FGD, and human-likeness. Zilin Wang 0002, Zhiyong Wu 0001, Minglei Li 0001, Zhensong Zhang, Qiaochu Huang, Songcen Xu, Changpeng Yang, Zonghong Dai |
ACM Multimedia | 4 |
| 2023 | Leveraging the Latent Diffusion Models for Offline Facial Multiple Appropriate Reactions GenerationabstractOffline Multiple Appropriate Facial Reaction Generation (OMAFRG) aims to predict the reaction of different listeners given a speaker, which is useful in the senario of human-computer interaction and social media analysis. In recent years, the Offline Facial Reactions Generation (OFRG) task has been explored in different ways. However, most studies only focus on the deterministic reaction of the listeners. The research of the non-deterministic (i.e. OMAFRG) always lacks of sufficient attention and the results are far from satisfactory. Compared with the deterministic OFRG tasks, the OMAFRG task is closer to the true circumstance but corresponds to higher difficulty for its requirement of modeling stochasticity and context. In this paper, we propose a new model named FRDiff to tackle this issue. Our model is developed based on the diffusion model architecture with some modification to enhance its ability of aggregating the context features. And the inherent property of stochasticity in diffusion model enables our model to generate multiple reactions. We conduct experiments on the datasets provided by the ACM Multimedia REACT2023 and obtain the second place on the board, which demonstrates the effectiveness of our method. Jun Yu 0001, Ji Zhao 0020, Guochen Xie, Fengxin Chen, Minglei Li 0001, Zonghong Dai |
ACM Multimedia | 7 |
| 2022 | A Multitask Learning Framework for Speaker Change Detection with Content Information from Unsupervised Speech DecompositionabstractSpeaker Change Detection (SCD) is a task of determining the time boundaries between speech segments of different speakers. SCD system can be applied to many tasks, such as speaker diarization, speaker tracking, and transcribing audio with multiple speakers. Recent advancements in deep learning lead to approaches that can directly detect the speaker change points from audio data at the frame-level based on neural network models. These approaches may be further improved by utilizing speaker information in the training data, and utilizing content information extracted in an unsupervised manner. This work proposes a novel framework for the SCD task, which utilizes a multitask learning architecture to leverage speaker information during the training stage, and adds the content information extracted from an unsupervised speech decomposition model to help detect the speaker change points. Experiment results show that the architecture of multitask learning with speaker information can improve the performance of SCD, and adding content information extracted from unsupervised speech decomposition model can further improve the performance. To the best of our knowledge, this work outperforms the state-of-the-art SCD results [1] on the AMI dataset. Danyang Zhao, Long Dang, Minglei Li 0001, Xixin Wu, Xunying Liu, Helen M. Meng |
ICASSP | 4 |
| 2022 | The ReprGesture entry to the GENEA Challenge 2022abstractThis paper describes the ReprGesture entry to the Generation and Evaluation of Non-verbal Behaviour for Embodied Agents (GENEA) challenge 2022. The GENEA challenge provides the processed datasets and performs crowdsourced evaluations to compare the performance of different gesture generation systems. In this paper, we explore an automatic gesture generation system based on multimodal representation learning. We use WavLM features for audio, FastText features for text and position and rotation matrix features for gesture. Each modality is projected to two distinct subspaces: modality-invariant and modality-specific. To learn inter-modality-invariant commonalities and capture the characters of modality-specific representations, gradient reversal layer based adversarial classifier and modality reconstruction decoders are used during training. The gesture decoder generates proper gestures using all representations and features related to the rhythm in the audio. Our code, pre-trained models and demo are available at https://github.com/YoungSeng/ReprGesture. Zhiyong Wu 0001, Minglei Li 0001, Mengchen Zhao, Jiuxin Lin, Liyang Chen, Weihong Bao |
ICMI | 3 |
| 2022 | A hierarchical interactive multi-channel graph neural network for technological knowledge flow forecasting
Huijie Liu 0001, Han Wu 0002, Le Zhang 0010, Runlong Yu, Ye Liu 0011, Chunli Liu 0001, Minglei Li 0001, Qi Liu 0003, Enhong Chen |
Knowl. Inf. Syst. | 7 |
| 2021 | Talking Face Generation Based on Information Bottleneck and Complementary RepresentationsabstractAudio-driven talking face generation is an active research direction in the field of virtual reality. The main challenge is that the generated lip shape of the speaker is out of sync with the input audio. To address this challenge, we propose a novel solution to synthesize lip-synchronized, high-quality, realistic video given input audio. We first decompose the target person's video frames into 3D face model parameters, and the information bottleneck is inserted into the audio-to-expression network to learn the mapping between audio features and expression parameters. Then, we replace the expression parameters in the target video frame with the extracted expression parameters from audio and re-render the face. Finally, we add high-level audio embedding extracted from the raw audio and lip landmarks embedding in the neural rendering network. The 3D face shapes, 2D landmarks, and audio embedding provide complementary information for the neural rendering network which guarantees the generation of lip-synchronized high-quality video portraits from the synthesized rendered faces. Experimental results show that compared with other talking face generation methods, our method is the best concerning lip synchronization with high video definition. Yiling Wu, Minglei Li 0001 |
CIKM | 3 |
| 2021 | Combining cross-modal knowledge transfer and semi-supervised learning for speech emotion recognition
Min Chen 0003, Jincai Chen, Yuan-Fang Li, Yiling Wu, Minglei Li 0001, Chuanbo Zhu 0002 |
Knowl. Based Syst. | 6 |
| 2021 | Improving Attention Model Based on Cognition Grounded Data for Sentiment AnalysisabstractAttention models are proposed in sentiment analysis and other classification tasks because some words are more important than others to train the attention models. However, most existing methods either use local context based information, affective lexicons, or user preference information. In this work, we propose a novel attention model trained by cognition grounded eye-tracking data. First,a reading prediction model is built using eye-tracking data as dependent data and other features in the context as independent data. The predicted reading time is then used to build a cognition grounded attention layer for neural sentiment analysis. Our model can capture attentions in context both in terms of words at sentence level as well as sentences at document level. Other attention mechanisms can also be incorporated together to capture other aspects of attentions, such as local attention, and affective lexicons. Results of our work include two parts. The first part compares our proposed cognition ground attention model with other state-of-the-art sentiment analysis models. The second part compares our model with an attention model based on other lexicon based sentiment resources. Evaluations show that sentiment analysis using cognition grounded attention model outperforms the state-of-the-art sentiment analysis methods significantly. Comparisons to affective lexicons also indicate that using cognition grounded eye-tracking data has advantages over other sentiment resources by considering both word information and context information. This work brings insight to how cognition grounded data can be integrated into natural language processing (NLP) tasks. Rong Xiang, Qin Lu 0001, Chu-Ren Huang, Minglei Li 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2018 | Phrase embedding learning based on external and internal context with compositionality constraint
Minglei Li 0001, Qin Lu 0001, Dan Xiong |
Knowl. Based Syst. | 1 |
| 2017 | A Cognition Based Attention Model for Sentiment AnalysisabstractAttention models are proposed in sentiment analysis because some words are more important than others.However, most existing methods either use local context based text information or user preference information.In this work, we propose a novel attention model trained by cognition grounded eye-tracking data.A reading prediction model is first built using eye-tracking data as dependent data and other features in the context as independent data.The predicted reading time is then used to build a cognition based attention (CBA) layer for neural sentiment analysis.As a comprehensive model, We can capture attentions of words in sentences as well as sentences in documents.Different attention mechanisms can also be incorporated to capture other aspects of attentions.Evaluations show the CBA based method outperforms the state-of-the-art local context based attention methods significantly.This brings insight to how cognition grounded data can be brought into NLP tasks. Qin Lu 0001, Rong Xiang, Minglei Li 0001, Chu-Ren Huang |
EMNLP | 4 |
| 2017 | Representation Learning of Multiword Expressions with Compositionality Constraint
Minglei Li 0001, Qin Lu 0001 |
KSEM | 1 |
| 2017 | Inferring Affective Meanings of Words from Word EmbeddingabstractAffective lexicon is one of the most important resource in affective computing for text. Manually constructed affective lexicons have limited scale and thus only have limited use in practical systems. In this work, we propose a regression-based method to automatically infer multi-dimensional affective representation of words via their word embedding based on a set of seed words. This method can make use of the rich semantic meanings obtained from word embedding to extract meanings in some specific semantic space. This is based on the assumption that different features in word embedding contribute differently to a particular affective dimension and a particular feature in word embedding contributes differently to different affective dimensions. Evaluation on various affective lexicons shows that our method outperforms the state-of-the-art methods on all the lexicons under different evaluation metrics with large margins. We also explore different regression models and conclude that the Ridge regression model, the Bayesian Ridge regression model and Support Vector Regression with linear kernel are the most suitable models. Comparing to other state-of-the-art methods, our method also has computation advantage. Experiments on a sentiment analysis task show that the lexicons extended by our method achieve better results than publicly available sentiment lexicons on eight sentiment corpora. The extended lexicons are publicly available for access. Minglei Li 0001, Qin Lu 0001, Lin Gui 0003 |
IEEE Trans. Affect. Comput. | 1 |
| 2016 | Domain-specific user preference prediction based on multiple user activitiesabstractInferring latent user preferences using both structured and unstructured data is an important social computing task. In this paper, we propose a user preference representation based on user activities embedded in unstructured data to better encode the homophily theory. The representation of an individual user is learned using a embedding based method to integrate latent user preferences in social media. The method has the ability to integrate a variety of user activities based cues from user comments, user social network (i.e; follower/followee connections) and user interested topics which are indicated by the topics a user has participated in. Experiments are conducted to evaluate the prediction of each user's favorite team as a part of user preferences in a dataset collected from the Hu-pu basketball discussion forum.1Results clearly indicate that our proposed user representation outperforms other user representation baselines. Integrating user social network and user interested topics with user comments can improve the overall performance of user preference prediction. Qin Lu 0001, Minglei Li 0001, Chu-Ren Huang |
IEEE BigData | 4 |
| 2016 | Emotion Corpus Construction Based on Selection from Hashtags
Minglei Li 0001, Qin Lu 0001, Wenjie Li 0002 |
LREC | 1 |
| 2016 | Syllable based DNN-HMM Cantonese Speech to Text System
Timothy Wong, Claire Li, Sam Lam, Billy Chiu, Qin Lu 0001, Minglei Li 0001, Dan Xiong, Roy Shing Yu, Vincent T. Y. Ng |
LREC | 6 |
| 2016 | Event Based Emotion Classification for News Articles
Minglei Li 0001, Qin Lu 0001 |
PACLIC | 1 |
| 2015 | Web Person Disambiguation Using Hierarchical Co-reference Model
Jian Xu 0014, Qin Lu 0001, Minglei Li 0001, Wenjie Li 0002 |
CICLing (1) | 3 |
| 2015 | A Novel Class Noise Estimation Method and Application in ClassificationabstractNoise in class labels of any training set can lead to poor classification results no matter what machine learning method is used. In this paper, we first present the problem of binary classification in the presence of random noise on the class labels, which we call class noise. To model class noise, a class noise rate is normally defined as a small independent probability of the class labels being inverted on the whole set of training data. In this paper, we propose a method to estimate class noise rate at the level of individual samples in real data. Based on the estimation result, we propose two approaches to handle class noise. The first technique is based on modifying a given surrogate loss function. The second technique eliminates class noise by sampling. Furthermore, we prove that the optimal hypothesis on the noisy distribution can approximate the optimal hypothesis on the clean distribution using both approaches. Our methods achieve over 87% accuracy on a synthetic non-separable dataset even when 40% of the labels are inverted. Comparisons to other algorithms show that our methods outperform state-of-the-art approaches on several benchmark datasets in different domains with different noise rates. Lin Gui 0003, Qin Lu 0001, Ruifeng Xu 0001, Minglei Li 0001, Qikang Wei |
CIKM | 4 |