Jia Jia 0001

dblp:71/2992-1 · DBLP profile ↗
← Back
138ranked-venue papers
10as first author
40since 2021 · last 2026
0009-0005-8449-278XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 113 · 6 first-author · 33 since 2021Artificial intelligence and machine learning · 49 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 4 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Emotion-Conditioned Motion Sub-spaces with Flow Matching for Real-Time Audio-Driven Talking Heads
abstract
Recent advances in audio-driven talking-head synthesis have brought lip-sync precision close to human perception, yet emotional fidelity and real-time inference remain open challenges. Existing pipelines typically disentangle lip articulation, facial expression, and head pose in latent space; this rigid factorization ignores the intrinsic coupling between articulation and affect — e.g., downward lip corners when sad—thus limiting expressiveness. We cast speech-conditioned facial motion as a sample from an emotion-conditioned distribution in a motion latent space. Concretely, we (i) learn a motion dictionary of orthogonal bases with an autoencoder via self-supervision, (ii) construct emotion-conditioned sub-spaces within the latent space, and (iii) design a layer-progressive cross-attention fusion module that modulates a flow-matching sampler with both audio and emotion signals. Only ten reverse ODE steps are required to generate a motion-latent trajectory, enabling real-time end-to-end latency. Extensive experiments on MEAD and RAVDESS show that our method outperforms recent GAN- and diffusion-based baselines in emotion accuracy while running at around 75 FPS on a single desktop GPU. The proposed framework delivers the first emotionally expressive Audio2Face system that simultaneously achieves lip-sync accuracy, affective realism, and real-time performance.
Haoyu Wang 0009, Xiaozhe Xin, Xiaoyu Qin 0001, Meiguang Jin, Junfeng Ma, Jia Jia 0001
AAAI7
2026 Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
abstract
Long video understanding (LVU) is challenging due to rich and complicated multimodal clues in long temporal range. Current methods adopt reasoning to improve the model's ability to analyze complex video clues in long videos via text-form reasoning. However, the existing literature suffers from the fact that the text-only reasoning under fixed video context may exacerbate hallucinations since detailed crucial clues are often ignored under limited video context length due to the temporal redundancy of long videos. To address this gap, we propose Video-TwG, a curriculum reinforced framework that employs a novel Think-with-Grounding paradigm, enabling video LLMs to actively decide when to perform on-demand grounding during interleaved text–video reasoning, selectively zooming into question-relevant clips only when necessary. Video-TwG can be trained end-to-end in a straightforward manner, without relying on complex auxiliary modules or heavily annotated reasoning traces. In detail, we design a Two-stage Reinforced Curriculum Strategy, where the model first learns think-with-grounding behavior on a small short-video GQA dataset with grounding labels, and then scales to diverse general QA data with videos of diverse domains to encourage generalization. Further, to handle complex think-with-grounding reasoning for various kinds of data, we propose the TwG-GRPO algorithm, which features the fine-grained grounding reward, self-confirmed pseudo reward, and accuracy-gated mechanism. Finally, we propose to construct a new TwG-51K dataset that facilitates training. Experiments on Video-MME, LongVideoBench, and MLVU show that Video-TwG consistently outperforms strong LVU baselines. Further ablation validates the necessity of our Two-stage Reinforced Curriculum Strategy and shows our TwG-GRPO better leverages diverse unlabeled data to improve grounding quality and reduce redundant groundings without sacrificing QA performance. https://github.com/hlchen23/Video-TwG
Houlun Chen, Xin Wang 0019, Guangyao Li 0001, Yuwei Zhou, Jia Jia 0001, Wenwu Zhu 0001
SIGIR6
2025 Minimal Impact ControlNet: Advancing Multi-ControlNet Integration
abstract
With the advancement of diffusion models, there is a growing demand for high-quality, controllable image generation, particularly through methods that utilize one or multiple control signals based on ControlNet. However, in current ControlNet training, each control is designed to influence all areas of an image, which can lead to conflicts when different control signals are expected to manage different parts of the image in practical applications. This issue is especially pronounced with edge-type control conditions, where regions lacking boundary information often represent low-frequency signals, referred to as silent control signals. When combining multiple ControlNets, these silent control signals can suppress the generation of textures in related areas, resulting in suboptimal outcomes. To address this problem, we propose Minimal Impact ControlNet. Our approach mitigates conflicts through three key strategies: constructing a balanced dataset, combining and injecting feature signals in a balanced manner, and addressing the asymmetry in the score function’s Jacobian matrix induced by ControlNet. These improvements enhance the compatibility of control signals, allowing for freer and more harmonious generation in areas with silent control signals.
Shikun Sun, Zixuan Wang 0026, Xubin Li, Tiezheng Ge, Zijie Ye, Xiaoyu Qin 0001, Junliang Xing, Bo Zheng 0007, Jia Jia 0001
ICLR10
2025 Localizing Step-by-Step: Multimodal Long Video Temporal Grounding with LLM
abstract
Video Temporal Grounding (VTG) localizes moments in untrimmed videos using natural language queries. Most VTG datasets focus on short videos, and existing approaches excel in short-term cross-modal matching but struggle with long VTG, where long-range temporal reasoning is required for complex events. Existing approaches typically output timestamp predictions without intermediate steps, limiting effective reasoning, whereas humans solve this step by step. To address this, we propose a long VTG framework, StepVTG, with multimodal visual and speech inputs, leveraging Large Language Models (LLMs) for step-by-step reasoning. Specifically, we transform task descriptions, speech, and visual inputs into text prompts. To enhance temporal reasoning, we introduce the Boundary-Perceptive Prompting strategy, which includes: i) a multiscale denoising Chain-of-Thought (CoT) combining global and local semantics with noise filtering, ii) validity principles to ensure LLMs generate reasonable, parsable predictions, and iii) one-shot In-Context Learning (ICL) to improve reasoning via imitation. For evaluation, we establish MM-LVTG, a new long VTG benchmark with multimodal inputs, and demonstrate through extensive experiments that StepVTG achieves state-of-the-art performance. It offers explainable reasoning steps for predictions and reveals potential in facilitating video understanding with off-the-shelf LLMs.
Houlun Chen, Xin Wang 0019, Hong Chen 0011, Zihan Song 0003, Jia Jia 0001, Wenwu Zhu 0001
ICME6
2025 Specify Privacy Yourself: Assessing Inference-Time Personalized Privacy Preservation Ability of Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities but raise significant privacy concerns due to their abilities to infer sensitive personal information from images with high precision. While current LVLMs are relatively well aligned to protect universal privacy, e.g., credit card data, we argue that privacy is inherently personalized and context-dependent. This work pivots towards a novel task: can LVLMs achieve Inference-Time Personalized Privacy Protection (ITP3), allowing users to dynamically specify privacy boundaries through language specifications? To this end, we present SPY-Bench, the first systematic assessment of ITP3 ability, which comprises (1) 32,700 unique samples with image-question pairs and personalized privacy instructions across 67 categories and 24 real-world scenarios, and (2) novel metrics grounded in user specifications and context awareness. Benchmarking the ITP3 ability of 21 SOTA LVLMs, we reveal that: (i) most models, even the top-performing o4-mini, perform poorly, with only ~24% compliance accuracy; (ii) they show quite limited contextual privacy understanding capability. Therefore, we implemented initial ITP3 alignment methods, including a novel Noise Contrastive Alignment variant which achieves 96.88% accuracy while maintaining reasonable general performance. These results mark an initial step towards the ethical deployment of more controllable LVLMs. Code and data are at https://github.com/achernarwang/specify-privacy-yourself.
Xingqi Wang 0003, Xiaoyuan Yi, Xing Xie 0001, Jia Jia 0001
ACM Multimedia4
2025 DEPO: Enhancing E-commerce Image Background Generation with Short Trajectory Direct Expected Preference Optimization
Shikun Sun, Zixuan Wang 0026, Xiaoyu Qin 0001, Tiezheng Ge, Bo Zheng 0007, Jia Jia 0001
ACM Multimedia8
2025 PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
Sihan Zhao, Zixuan Wang 0026, Tianyu Luan, Jia Jia 0001, Wentao Zhu 0004, Jiebo Luo 0001, Junsong Yuan 0001, Nan Xi
ACM Multimedia4
2025 HarmoniVox: Painting Voices to Match the Avatar's Soul
abstract
Imagine James Bond speaking like Mr. Bean---such a mismatch would create a jarring dissonance and break the viewer's immersion. Current research on virtual avatar animation has focused on modeling 3D geometry, appearance, motion generation, however, neglecting the harmony between speech prosody and the avatar's visual presentation and contextual environment. In this paper, we seek to bridge this gap by firstly identifying and defining the key elements necessary for achieving audiovisual harmony, such as appearance, expression, body posture, backgrounds and colors. Subsequently, we propose a method that jointly models semantic consistency in avatar animation, named HarmoniVox, specifically on crafting prosodic speech consistent with the avatar's essence from given visual image. To achieve this, we implement a technical framework with a mutual modal contrastive learning strategy, enhancing multimodal alignment in a coarse-to-fine fashion. To support this method, we establish a experimental dataset HarAvaSpeech comprising 28,929 image-audio pairs, designed to encompass expressive speech prosody and rich avatar visual presentations across a wide range of contexts. Leveraging this dataset, our experiments demonstrate that the proposed method outperforms the baselines in manipulating the nuanced tone and harmonious rhythm of speech with the avatar visual presentations, and reveal generalizability on out-of-domain cases. Demo would be provided in https://harmonivox.github.io/harmonivox/.
Songtao Zhou, Xiaoyu Qin 0001, Yixuan Zhou 0002, Qixin Wang 0002, Zeyu Jin, Zixuan Wang 0026, Zhiyong Wu 0001, Jia Jia 0001
ACM Multimedia8
2024 DanceCamera3D: 3D Camera Movement Synthesis with Music and Dance
abstract
Choreographers determine what the dances look like, while cameramen determine the final presentation of dances. Recently, various methods and datasets have show-cased the feasibility of dance synthesis. However, camera movement synthesis with music and dance remains an un-solved challenging problem due to the scarcity of paired data. Thus, we present DCM, a new multi-modal 3D dataset, which for the first time combines camera movement with dance motion and music audio. This dataset encom-passes 108 dance sequences (3.2 hours) of paired dance-camera-music data from the anime community, covering 4 music genres. With this dataset, we uncover that dance camera movement is multifaceted and human-centric, and possesses multiple influencing factors, making dance camera synthesis a more challenging task compared to camera or dance synthesis alone. To overcome these difficulties, we propose DanceCamera3D, a transformer-based diffusion model that incorporates a novel body attention loss and a condition separation strategy. For evaluation, we devise new metrics measuring camera movement quality, diversity, and dancer fidelity. Utilizing these metrics, we conduct extensive experiments on our DCM dataset, providing both quantitative and qualitative evidence showcasing the effectiveness of our DanceCamera3D model. Code and video demos are available at https://github.com/Carmenw1203/DanceCamera3D-Official.
Zixuan Wang 0026, Jia Jia 0001, Shikun Sun, Haozhe Wu, Jiaqing Zhou, Jiebo Luo 0001
CVPR2
2024 Inner Classifier-Free Guidance and Its Taylor Expansion for Diffusion Models
abstract
Classifier-free guidance (CFG) is a pivotal technique for balancing the diversity and fidelity of samples in conditional diffusion models. This approach involves utilizing a single model to jointly optimize the conditional score predictor and unconditional score predictor, eliminating the need for additional classifiers. It delivers impressive results and can be employed for continuous and discrete condition representations. However, when the condition is continuous, it prompts the question of whether the trade-off can be further enhanced. Our proposed inner classifier-free guidance (ICFG) provides an alternative perspective on the CFG method when the condition has a specific structure, demonstrating that CFG represents a first-order case of ICFG. Additionally, we offer a second-order implementation, highlighting that even without altering the training policy, our second-order approach can introduce new valuable information and achieve an improved balance between fidelity and diversity for Stable Diffusion.
Shikun Sun, Longhui Wei, Zhicai Wang, Zixuan Wang 0026, Junliang Xing, Jia Jia 0001, Qi Tian 0001
ICLR6
2024 LoRA-MER: Low-Rank Adaptation of Pre-Trained Speech Models for Multimodal Emotion Recognition Using Mutual Information
Yunrui Cai, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng
INTERSPEECH3
2024 VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
abstract
Recent AIGC systems possess the capability to generate digital multimedia content based on human language instructions, such as text, image and video. However, when it comes to speech, existing methods related to human instruction-to-speech generation exhibit two limitations. Firstly, they require the division of inputs into content prompt (transcript) and description prompt (style and speaker), instead of directly supporting human instruction. This division is less natural in form and does not align with other AIGC models. Secondly, the practice of utilizing an independent description prompt to model speech style, without considering the transcript content, restricts the ability to control speech at a fine-grained level. To address these limitations, we propose VoxInstruct, a novel unified multilingual codec language modeling framework that extends traditional text-to-speech tasks into a general human instruction-to-speech task. Our approach enhances the expressiveness of human instruction-guided speech generation and aligns the speech generation paradigm with other modalities. To enable the model to automatically extract the content of synthesized speech from raw text instructions, we introduce speech semantic tokens as an intermediate representation for instruction-to-content guidance. We also incorporate multiple Classifier-Free Guidance (CFG) strategies into our codec language model, which strengthens the generated speech following human instructions. Furthermore, our model architecture and training strategies allow for the simultaneous support of combining speech prompt and descriptive human instruction for expressive speech synthesis, which is a first-of-its-kind attempt. Codes, models and demos are at: https://github.com/thuhcsi/VoxInstruct.
Yixuan Zhou 0002, Xiaoyu Qin 0001, Zeyu Jin, Shuoyi Zhou, Shun Lei, Songtao Zhou, Zhiyong Wu 0001, Jia Jia 0001
ACM Multimedia8
2024 PlacidDreamer: Advancing Harmony in Text-to-3D Generation
abstract
Recently, text-to-3D generation has attracted significant attention, resulting in notable performance enhancements. Previous methods utilize end-to-end 3D generation models to initialize 3D Gaussians, multi-view diffusion models to enforce multi-view consistency, and text-to-image diffusion models to refine details with score distillation algorithms. However, these methods exhibit two limitations. Firstly, they encounter conflicts in generation directions since different models aim to produce diverse 3D assets. Secondly, the issue of over-saturation in score distillation has not been thoroughly investigated and solved. To address these limitations, we propose PlacidDreamer, a text-to-3D framework that harmonizes initialization, multi-view generation, and text-conditioned generation with a single multi-view diffusion model, while simultaneously employing a novel score distillation algorithm to achieve balanced saturation. To unify the generation direction, we introduce the Latent-Plane module, a training-friendly plug-in extension that enables multi-view diffusion models to provide fast geometry reconstruction for initialization and enhanced multi-view images to personalize the text-to-image diffusion model. To address the over-saturation problem, we propose to view score distillation as a multi-objective optimization problem and introduce the Balanced Score Distillation algorithm, which offers a Pareto Optimal solution that achieves both rich details and balanced saturation. Extensive experiments validate the outstanding capabilities of our PlacidDreamer. The code is available at https://github.com/HansenHuang0823/PlacidDreamer.
Shuo Huang 0005, Shikun Sun, Zixuan Wang 0026, Xiaoyu Qin 0001, Yanmin Xiong, Yuan Zhang 0020, Pengfei Wan 0001, Di Zhang 0026, Jia Jia 0001
ACM Multimedia9
2024 SpeechCraft: A Fine-Grained Expressive Speech Dataset with Natural Language Description
abstract
Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elaborate comprehension of speech style is urgently needed to facilitate insightful interplay between speech audio and natural language. However, constructing such datasets presents a major trade-off between large-scale data collection and high-quality annotation. To tackle this challenge, we propose an automatic speech annotation system for expressiveness interpretation that annotates in-the-wild speech clips with expressive and vivid human language descriptions. Initially, speech audios are processed by a series of expert classifiers and captioning models to capture diverse speech characteristics, followed by a fine-tuned LLaMA for customized annotation generation. Unlike previous tag/templet-based annotation frameworks with limited information and diversity, our system provides in-depth understandings of speech style through tailored natural language descriptions, thereby enabling accurate and voluminous data generation for large model training. With this system, we create SpeechCraft, a fine-grained bilingual expressive speech dataset. It is distinguished by highly descriptive natural language style prompts, containing approximately 2,000 hours of audio data and encompassing over two million speech clips. Extensive experiments demonstrate that the proposed dataset significantly boosts speech-language task performance in stylist speech synthesis and speech style understanding.
Zeyu Jin, Jia Jia 0001, Qixin Wang 0002, Kehan Li 0007, Shuoyi Zhou, Songtao Zhou, Xiaoyu Qin 0001, Zhiyong Wu 0001
ACM Multimedia2
2024 DanceCamAnimator: Keyframe-Based Controllable 3D Dance Camera Synthesis
abstract
Synthesizing camera movements from music and dance is highly challenging due to the contradicting requirements and complexities of dance cinematography. Unlike human movements, which are always continuous, dance camera movements involve both continuous sequences of variable lengths and sudden drastic changes to simulate the switching of multiple cameras. However, in previous works, every camera frame is equally treated and this causes jittering and unavoidable smoothing in post-processing. To solve these problems, we propose to integrate animator dance cinematography knowledge by formulating this task as a three-stage process: keyframe detection, keyframe synthesis, and tween function prediction. Following this formulation, we design a novel end-to-end dance camera synthesis framework DanceCamAnimator, which imitates human animation procedures and shows powerful keyframe-based controllability with variable lengths. Extensive experiments on the DCM dataset demonstrate that our method surpasses previous baselines quantitatively and qualitatively. Code will be available at https://github.com/Carmenw1203/DanceCamAnimator-Official.
Zixuan Wang 0026, Xiaoyu Qin 0001, Shikun Sun, Songtao Zhou, Jia Jia 0001, Jiebo Luo 0001
ACM Multimedia6
2024 Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization
abstract
Recent advancements in diffusion models trained on large-scale data have enabled the generation of indistinguishable human-level images, yet they often produce harmful content misaligned with human values, e.g., social bias, and offensive content. Despite extensive research on Large Language Models (LLMs), the challenge of Text-to-Image (T2I) model alignment remains largely unexplored. Addressing this problem, we propose LiVO (Lightweight Value Optimization), a novel lightweight method for aligning T2I models with human values. LiVO only optimizes a plug-and-play value encoder to integrate a specified value principle with the input prompt, allowing the control of generated images over both semantics and values. Specifically, we design a diffusion model-tailored preference optimization loss, which theoretically approximates the Bradley-Terry model used in LLM alignment but provides a more flexible trade-off between image quality and value conformity. To optimize the value encoder, we also develop a framework to automatically construct a text-image preference dataset of 86k (prompt, aligned image, violating image, value principle) samples. Without updating most model parameters and through adaptive value selection from the input prompt, LiVO significantly reduces harmful outputs and achieves faster convergence, surpassing several strong baselines and taking an initial step towards ethically aligned T2I models. Warning: This paper involves descriptions and images depicting discriminatory, pornographic, bloody, and horrific scenes.
Xingqi Wang 0003, Xiaoyuan Yi, Xing Xie 0001, Jia Jia 0001
ACM Multimedia4
2024 VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
abstract
Existing Video Corpus Moment Retrieval (VCMR) is limited to coarse-grained understanding that hinders precise video moment localization when given fine-grained queries. In this paper, we propose a more challenging fine-grained VCMR benchmark requiring methods to localize the best-matched moment from the corpus with other partially matched candidates. To improve the dataset construction efficiency and guarantee high-quality data annotations, we propose VERIFIED, an automatic \underline{V}id\underline{E}o-text annotation pipeline to generate captions with \underline{R}el\underline{I}able \underline{FI}n\underline{E}-grained statics and \underline{D}ynamics. Specifically, we resort to large language models (LLM) and large multimodal models (LMM) with our proposed Statics and Dynamics Enhanced Captioning modules to generate diverse fine-grained captions for each video. To filter out the inaccurate annotations caused by the LLM hallucination, we propose a Fine-Granularity Aware Noise Evaluator where we fine-tune a video foundation model with disturbed hard-negatives augmented contrastive and matching losses. With VERIFIED, we construct a more challenging fine-grained VCMR benchmark containing Charades-FIG, DiDeMo-FIG, and ActivityNet-FIG which demonstrate a high level of annotation quality. We evaluate several state-of-the-art VCMR models on the proposed dataset, revealing that there is still significant scope for fine-grained video understanding in VCMR.
Houlun Chen, Xin Wang 0019, Hong Chen 0011, Zeyang Zhang 0001, Bin Huang 0004, Jia Jia 0001, Wenwu Zhu 0001
NeurIPS7
2023 What Does Your Face Sound Like? 3D Face Shape towards Voice
abstract
Face-based speech synthesis provides a practical solution to generate voices from human faces. However, directly using 2D face images leads to the problems of uninterpretability and entanglement. In this paper, to address the issues, we introduce 3D face shape which (1) has an anatomical relationship between voice characteristics, partaking in the "bone conduction" of human timbre production, and (2) is naturally independent of irrelevant factors by excluding the blending process. We devise a three-stage framework to generate speech from 3D face shapes. Fully considering timbre production in anatomical and acquired terms, our framework incorporates three additional relevant attributes including face texture, facial features, and demographics. Experiments and subjective tests demonstrate our method can generate utterances matching faces well, with good audio quality and voice diversity. We also explore and visualize how the voice changes with the face. Case studies show that our method upgrades the face-voice inference to personalized custom-made voice creating, revealing a promising prospect in virtual human and dubbing applications.
Zhiyong Wu 0001, Ying Shan, Jia Jia 0001
AAAI4
2023 Shuffled Autoregression for Motion Interpolation
abstract
This work aims to provide a deep-learning solution for the motion interpolation task. Previous studies solve it with geometric weight functions. Some other works propose neural networks for different problem settings with consecutive pose sequences as input. However, motion interpolation is a more complex problem that takes isolated poses (e.g., only one start pose and one end pose) as input. When applied to motion interpolation, these deep learning methods have limited performance since they do not leverage the flexible dependencies between interpolation frames as the original geometric formulas do. To realize this interpolation characteristic, we propose a novel framework, referred to as Shuffled AutoRegression, which expands the autoregression to generate in arbitrary (shuffled) order and models any inter-frame dependencies as a directed acyclic graph. We further propose an approach to constructing a particular kind of dependency graph, with three stages assembled into an end-to-end spatial-temporal motion Transformer. Experimental results on one of the current largest datasets show that our model generates vivid and coherent motions from only one start frame to one end frame and outperforms competing methods by a large margin. The proposed model is also extensible to multiple keyframes’ motion interpolation tasks and other areas’ interpolation.
Shuo Huang 0005, Jia Jia 0001, Zongxin Yang, Wei Wang 0010, Haozhe Wu, Yi Yang 0001, Junliang Xing
ICASSP2
2023 MSNet: A Deep Architecture Using Multi-Sentiment Semantics for Sentiment-Aware Image Style Transfer
abstract
Sentiment plays an essential role in people’s perception of images. To incorporate the sentiment information into the image style transfer task for better sentiment-aware performance, we introduce a new task named sentiment-aware image style transfer. To solve this problem, we first introduce a novel Multi-Sentiment Semantics Space (MSS-Space) to capture the non-deterministic and complicated nature of sentiment semantics. With the MSS-Space, we establish tight associations between the visual attributes of images and the multi-sentiment semantics by minimizing their distance in MSS-Space and then propose the Multi-Sentiment Style Transfer Net (MSNet). Experiments demonstrate that, compared with three competing models, our proposed MSNet generates more explicit images and better preserves the integrity of salient objects, local details, and multi-sentiment. In particular, our model outperforms the state-of-the-art by +28.72% in terms of the top-3 accuracy on average.
Shikun Sun, Jia Jia 0001, Haozhe Wu, Zijie Ye, Junliang Xing
ICASSP2
2023 Salient Co-Speech Gesture Synthesizing with Discrete Motion Representation
abstract
Synthesizing co-speech gestures is challenging because the mapping from speech to gesticulation is inherently non-deterministic. When giving talks, people conduct not only gentle and rhythmic motions but also abrupt and salient gesticulations. Most previous research efforts, however, ignore this nature of co-speech gestures and synthesize deterministic results, producing over-smoothed movements with limited expressiveness. To address this issue, we propose a new co-speech gesture generation approach that produces high-quality salient gesticulations. Specifically, we build a discrete motion representation (DMR) space to bridge the speech-gesture mapping and the gesture generation stages. The incorporation of DMR enables random sampling in motion space and avoids the over-smooth problem in speech-gesture mapping. Based on DMR, we devise a novel multi-modal co-speech gesture synthesis model with temporal attention (MCGT). MCGT explicitly models DMR’s categorical distribution conditioned on the speech context, which captures complex context patterns and produces more salient gesticulations in sync with the context. In addition, we construct a new benchmark for evaluating salient motion quality in co-speech gestures, containing a large-scale co-speech gesture dataset with salient gesticulations. We also introduce a new metric, referred to as salient motion similarity, to evaluate the salient motion quality. Experiments demonstrate superior results from our approach over several competing baselines.
Zijie Ye, Jia Jia 0001, Haozhe Wu, Shuo Huang 0005, Shikun Sun, Junliang Xing
ICASSP2
2023 SDDM: Score-Decomposed Diffusion Models on Manifolds for Unpaired Image-to-Image Translation
abstract
Recent score-based diffusion models (SBDMs) show promising results in unpaired image-to-image translation (I2I). However, existing methods, either energy-based or statistically-based, provide no explicit form of the interfered intermediate generative distributions. This work presents a new score-decomposed diffusion model (SDDM) on manifolds to explicitly optimize the tangled distributions during image generation. SDDM derives manifolds to make the distributions of adjacent time steps separable and decompose the score function or energy guidance into an image "denoising" part and a content "refinement" part. To refine the image in the same noise level, we equalize the refinement parts of the score function and energy guidance, which permits multi-objective optimization on the manifold. We also leverage the block adaptive instance normalization module to construct manifolds with lower dimensions but still concentrated with the perturbed reference image. SDDM outperforms existing SBDM-based methods with much fewer diffusion steps on several I2I benchmarks.
Shikun Sun, Longhui Wei, Junliang Xing, Jia Jia 0001, Qi Tian 0001
ICML4
2023 Prosody Modeling with 3D Visual Information for Expressive Video Dubbing
Shansong Liu, Xu Li 0015, Haozhe Wu, Zhiyong Wu 0001, Ying Shan, Jia Jia 0001
INTERSPEECH7
2023 Curriculum-Listener: Consistency- and Complementarity-Aware Audio-Enhanced Temporal Sentence Grounding
abstract
Temporal Sentence Grounding aims to retrieve a video moment given a natural language query. Most existing literature merely focuses on visual information in videos without considering the naturally accompanied audio which may contain rich semantics. The few works considering audio simply regard it as an additional modality, overlooking that: i) it's non-trivial to explore consistency and complementarity between audio and visual; ii) such exploration requires handling different levels of information densities and noises in the two modalities. To tackle these challenges, we propose Adaptive Dual-branch Promoted Network (ADPN) to exploit such consistency and complementarity: i) we introduce a dual-branch pipeline capable of jointly training visual-only and audio-visual branches to simultaneously eliminate inter-modal interference; ii) we design Text-Guided Clues Miner (TGCM) to discover crucial locating clues via considering both consistency and complementarity during audio-visual interaction guided by text semantics; iii) we propose a novel curriculum-based denoising optimization strategy, where we adaptively evaluate sample difficulty as a measure of noise intensity in a self-aware fashion. Extensive experiments show the state-of-the-art performance of our method.
Houlun Chen, Xin Wang 0019, Xiaohan Lan, Hong Chen 0011, Xuguang Duan, Jia Jia 0001, Wenwu Zhu 0001
ACM Multimedia6
2023 AvatarFusion: Zero-shot Generation of Clothing-Decoupled 3D Avatars Using 2D Diffusion
abstract
Large-scale pre-trained vision-language models allow for the zero-shot text-based generation of 3D avatars. The previous state-of-the-art method utilized CLIP to supervise neural implicit models that reconstructed a human body mesh. However, this approach has two limitations. Firstly, the lack of avatar-specific models can cause facial distortion and unrealistic clothing in the generated avatars. Secondly, CLIP only provides optimization direction for the overall appearance, resulting in less impressive results. To address these limitations, we propose AvatarFusion, the first framework to use a latent diffusion model to provide pixel-level guidance for generating human-realistic avatars while simultaneously segmenting clothing from the avatar's body. AvatarFusion includes the first clothing-decoupled neural implicit avatar model that employs a novel Dual Volume Rendering strategy to render the decoupled skin and clothing sub-models in one space. We also introduce a novel optimization method, called Pixel-Semantics Difference-Sampling (PS-DS), which semantically separates the generation of body and clothes, and generates a variety of clothing styles. Moreover, we establish the first benchmark for zero-shot text-to-avatar generation. Our experimental results demonstrate that our framework outperforms previous approaches, with significant improvements observed in all metrics. Additionally, since our model is clothing-decoupled, we can exchange the clothes of avatars. Code are available on our project page https://hansenhuang0823.github.io/AvatarFusion.
Shuo Huang 0005, Zongxin Yang, Liangting Li, Yi Yang 0001, Jia Jia 0001
ACM Multimedia5
2023 HoloSinger: Semantics and Music Driven Motion Generation with Octahedral Holographic Projection
abstract
Lyrics and music are both significant for a singer to perform a song. Therefore, it is important in singer's motion generation to model both semantic and acoustic correlation with motions at the same time. In this paper, we propose HoloSinger, a novel comprehensive system that synthesizes singing motions according to the given song. Additionally, we present singing avatar with octahedral holographic projection. For singing motion generation, we introduce a Transformer-VAE generative model to decompose lyrics and music, then fuse their impacts to synthesize singer's motions. Extensive experiments and user studies show that our method automatically generates realistic motions that adhere to musical choreography and reflect the lyric semantics appropriately. Furthermore, we design a desktop-level holographic projection device with an octahedral structure. It achieves high-definition holographic projection effects with smaller volume, larger imaging area ratio, and the ability of real-time AI interaction.
Zeyu Jin, Zixuan Wang 0026, Qixin Wang 0002, Jia Jia 0001, Ye Bai 0001, Yi Zhao 0006, Hao Li 0078
ACM Multimedia4
2023 Versatile Face Animator: Driving Arbitrary 3D Facial Avatar in RGBD Space
abstract
Creating realistic 3D facial animation is crucial for various applications in the movie production and gaming industry, especially with the burgeoning demand in the metaverse. However, prevalent methods such as blendshape-based approaches and facial rigging techniques are time-consuming, labor-intensive, and lack standardized configurations, making facial animation production challenging and costly. In this paper, we propose a novel self-supervised framework, Versatile Face Animator, which combines facial motion capture with motion retargeting in an end-to-end manner, eliminating the need for blendshapes or rigs. Our method has the following two main characteristics: 1) we propose an RGBD animation module to learn facial motion from raw RGBD videos by hierarchical motion dictionaries and animate RGBD images rendered from 3D facial mesh coarse-to-fine, enabling facial animation on arbitrary 3D characters regardless of their topology, textures, blendshapes, and rigs; and 2) we introduce a mesh retarget module to utilize RGBD animation to create 3D facial animation by manipulating facial mesh with controller transformations, which are estimated from dense optical flow fields and blended together with geodesic-distance-based weights. Comprehensive experiments demonstrate the effectiveness of our proposed framework in generating impressive 3D facial animation results, highlighting its potential as a promising solution for the cost-effective and efficient production of facial animation in the metaverse.
Haoyu Wang 0009, Haozhe Wu, Junliang Xing, Jia Jia 0001
ACM Multimedia4
2023 Speech-Driven 3D Face Animation with Composite and Regional Facial Movements
abstract
Speech-driven 3D face animation poses significant challenges due to the intricacy and variability inherent in human facial movements. This paper emphasizes the importance of considering both the composite and regional natures of facial movements in speech-driven 3D face animation. The composite nature pertains to how speech-independent factors globally modulate speech-driven facial movements along the temporal dimension. Meanwhile, the regional nature alludes to the notion that facial movements are not globally correlated but are actuated by local musculature along the spatial dimension. It is thus indispensable to incorporate both natures for engendering vivid animation. To address the composite nature, we introduce an adaptive modulation module that employs arbitrary facial movements to dynamically adjust speech-driven facial movements across frames on a global scale. To accommodate the regional nature, our approach ensures that each constituent of the facial features for every frame focuses on the local spatial movements of 3D faces. Moreover, we present a non-autoregressive backbone for translating audio to 3D facial movements, which maintains high-frequency nuances of facial movements and facilitates efficient inference. Comprehensive experiments and user studies demonstrate that our method surpasses contemporary state-of-the-art approaches both qualitatively and quantitatively.
Haozhe Wu, Songtao Zhou, Jia Jia 0001, Junliang Xing
ACM Multimedia3
2023 Semantics2Hands: Transferring Hand Motion Semantics between Avatars
abstract
Human hands, the primary means of non-verbal communication, convey intricate semantics in various scenarios. Due to the high sensitivity of individuals to hand motions, even minor errors in hand motions can significantly impact the user experience. Real applications often involve multiple avatars with varying hand shapes, highlighting the importance of maintaining the intricate semantics of hand motions across the avatars. Therefore, this paper aims to transfer the hand motion semantics between diverse avatars based on their respective hand models. To address this problem, we introduce a novel anatomy-based semantic matrix (ASM) that encodes the semantics of hand motions. The ASM quantifies the positions of the palm and other joints relative to the local frame of the corresponding joint, enabling precise retargeting of hand motions. Subsequently, we obtain a mapping function from the source ASM to the target hand joint rotations by employing an anatomy-based semantics reconstruction network (ASRN). We train the ASRN using a semi-supervised learning strategy on the Mixamo and InterHand2.6M datasets. We evaluate our method in intra-domain and cross-domain hand motion retargeting tasks. The qualitative and quantitative results demonstrate the significant superiority of our ASRN over the state-of-the-arts. Code available at https://github.com/abcyzj/Semantics2Hand
Zijie Ye, Jia Jia 0001, Junliang Xing
ACM Multimedia2
2022 Learning from Designers: Fashion Compatibility Analysis Via Dataset Distillation
abstract
Learning fashion compatibility is of great significance to both academic research and industry, which serves as a key technique for many real applications like online shopping recommendation and clothing generation. In previous studies, user-generated data (e.g. outfits from social media platform) are usually used for learning item embeddings and further modeling the compatibility. However, due to the noisy and messy nature of such data, one can hardly learn a representation that can clearly characterize the fashion-related attributes (e.g. color, material). In this paper, we propose an Attention-based Dataset Distillation Graph Neural Network (ADD-GNN) to leverage the designer-generated data as a guidance on modeling the outfit compatibility. Specifically, we jointly optimize two components which distill knowledge from fashion designers for feature representation learning and model the overall compatibility through attention-based graph neural network. Experimental results on real world fashion datasets clearly demonstrate the superiority of our proposed ADD-GNN against several competitive baselines in outfit compatibility tasks, which proves the effectiveness of distilling knowledge from designers.
Yulan Chen, Zhiyong Wu 0001, Zheyan Shen, Jia Jia 0001
ICIP4
2022 Speaker Characteristics Guided Speech Synthesis
abstract
Talking head techniques are widely researched. Most of the previous works focus on the association among tones, prosody, and visual cues, such as head motion, lip movement, and gestures. However, it is widely believed the timbre, matching the voice with the speaker's identity, shall be considered, since people obtain speaker-specific information from both the auditory and visual modalities. This paper aims to generate proper voice characteristics in line with the speaker characteristics we select. We first select six speaker characteristics related to the voice qualities: gender, age, race, body mass index, face shape, and personality. We then train a Conditional Variational AutoEncoder with attention (attentionCVAE) model to infer speaker embeddings from speaker characteristics and employ a multi-speaker text-to-speech system to generate utterances of nonexistent speakers we set. Subjective tests indicate the proposed method successfully reconstructs real-world speaker embedding and generates realistic embedding from speaker characteristics. The further analysis uncovers how and to what extent the speaker characteristics influence the voice qualities of speakers.
Zhiyong Wu 0001, Jia Jia 0001
IJCNN3
2022 Towards Cross-speaker Reading Style Transfer on Audiobook Dataset
abstract
Cross-speaker style transfer aims to extract the speech style of the given reference speech, which can be reproduced in the timbre of arbitrary target speakers.Existing methods on this topic have explored utilizing utterance-level style labels to perform style transfer via either global or local scale style representations.However, audiobook datasets are typically characterized by both the local prosody and global genre, and are rarely accompanied by utterance-level style labels.Thus, properly transferring the reading style across different speakers remains a challenging task.This paper aims to introduce a chunk-wise multi-scale cross-speaker style model to capture both the global genre and the local prosody in audiobook speeches.Moreover, by disentangling speaker timbre and style with the proposed switchable adversarial classifiers, the extracted reading style is made adaptable to the timbre of different speakers.Experiment results confirm that the model manages to transfer a given reading style to new target speakers.With the support of local prosody and global genre type predictor, the potentiality of the proposed method in multi-speaker audiobook generation is further revealed.
Xiang Li 0105, Changhe Song, Xianhao Wei, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng
INTERSPEECH5
2022 Inferring Speaking Styles from Multi-modal Conversational Context by Multi-scale Relational Graph Convolutional Networks
abstract
To support applications of speech-driven interactive systems in various conversational scenarios, text-to-speech (TTS) synthesis needs to understand the conversational context and determine appropriate speaking styles in its synthesized speeches. These speaking styles are influenced by the dependencies between the multi-modal information in the context at both global scale (i.e. utterance level) and local scale (i.e. word level). However, the dependency modeling and speaking style inference at the local scale are largely missing in state-of-the-art TTS systems, resulting in the synthesis of incorrect or improper speaking styles. In this paper, to learn the dependencies in conversations at both global and local scales and to improve the synthesis of speaking styles, we propose a context modeling method which models the dependencies among the multi-modal information in context with multi-scale relational graph convolutional network (MSRGCN). The learnt multi-modal context information at multiple scales is then utilized to infer the global and local speaking styles of the current utterance for speech synthesis. Experiments demonstrate the effectiveness of the proposed approach, and ablation studies reflect the contributions from modeling multi-modal information and multi-scale dependencies.
Jingbei Li, Xixin Wu, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng, Qiao Tian 0001, Yuping Wang 0005, Yuxuan Wang 0002
ACM Multimedia5
2022 GroupDancer: Music to Multi-People Dance Synthesis with Style Collaboration
abstract
Different people dance in different styles. So when multiple people dance together, the phenomenon of style collaboration occurs: people need to seek common points while reserving differences in various dancing periods. Thus, we introduce a novel Music-driven Group Dance Synthesis task. Compared with single-people dance synthesis explored by most previous works, modeling the style collaboration phenomenon and choreographing for multiple people are more complicated and challenging. Moreover, the lack of sufficient records for conducting multi-people choreography in prior datasets further aggravates this problem. To address these issues, we construct a rich-annotated 3D Multi-Dancer Choreography dataset (MDC) and newly devise a metric SCEU for style collaboration evaluation. To our best knowledge, MDC is the first 3D dance dataset that collects both individual and collaborated music-dance pairs. Based on MDC, we present a novel framework, GroupDancer, consisting of three stages: Dancer Collaboration, Motion Choreography and Motion Transition. The Dancer Collaboration stage determines when and which dancers should collaborate their dancing styles from music. Afterward, the Motion Choreography stage produces a motion sequence for each dancer. Finally, the Motion Transition stage fills the gaps between the motions to achieve fluent and natural group dance. To make GroupDancer trainable from end to end and able to synthesize group dance with style collaboration, we propose mixed training and selective updating strategies. Comprehensive evaluations on the MDC dataset demonstrate that the proposed GroupDancer model can synthesize quite satisfactory group dance synthesis results with style collaboration.
Zixuan Wang 0026, Jia Jia 0001, Haozhe Wu, Junliang Xing, Jinghe Cai, Guowen Chen
ACM Multimedia2
2022 AI Carpet: Automatic Generation of Aesthetic Carpet Pattern
abstract
Stylized pattern generation is challenging and has received increasing attention in recent studies. However, it requires further exploration in pattern generation that matches the given scenes. This paper proposes a demonstration that automatically generates carpet patterns with input home scenes and other user preferences. Besides meeting the individual needs of users and providing highly editable output, the critical challenge of the system is to make the output pattern coordinate with the input home scene, which distinguishes our approach from others. The carpets generated by the system also permit easy modification or extension to various interior design styles.
Xingqi Wang 0003, Zeyu Jin, Shikun Sun, Jia Jia 0001
ACM Multimedia6
2021 Inferring Emotion from Large-scale Internet Voice Data: A Semi-supervised Curriculum Augmentation based Deep Learning Approach
abstract
Effective emotion inference from user queries helps to give a more personified response for Voice Dialogue Applications(VDAs). The tremendous amounts of VDA users bring in diverse emotion expressions. How to achieve a high emotion inferring performance from large-scale Internet Voice Data in VDAs? Traditionally, researches on speech emotion recognition are based on acted voice datasets, which have limited speakers but strong and clear emotion expressions. Inspired by this, in this paper, we propose a novel approach to leverage acted voice data with strong emotion expressions to enhance large-scale unlabeled internet voice data with diverse emotion expressions for emotion inferring. Specifically, we propose a novel semi-supervised multi-modal curriculum augmentation deep learning framework. First, to learn more general emotion cues, we adopt a curriculum learning based epoch-wise training strategy, which trains our model guided by strong and balanced emotion samples from acted voice data and sub-sequently leverages weak and unbalanced emotion samples from internet voice data.Second, to employ more diverse emotion expressions, we design a Multi-path Mix-match Multimodal Deep Neural Network(MMMD), which effectively learns feature representations for multiple modalities and trains labeled and unlabeled data in hybrid semi-supervised methods for superior generalization and robustness. Experiments on an internet voice dataset with 500,000 utterances show our method outperforms (+10.09% in terms of F1) several alternative baselines, while an acted corpus with 2,397 utterances contributes 4.35%. To further compare our method with state-of-the-art techniques in traditionally acted voice datasets, we also conduct experiments on public dataset IEMOCAP. The results reveal the effectiveness of the proposed approach.
Suping Zhou, Jia Jia 0001, Zhiyong Wu 0001, Wei Chen 0071, Shuo Huang 0005, Jialie Shen 0001
AAAI2
2021 PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback
abstract
Second language (L2) English learners often find it difficult to improve their pronunciations due to the lack of expressive and personalized corrective feedback. In this paper, we present Pronunciation Teacher (PTeacher), a Computer-Aided Pronunciation Training (CAPT) system that provides personalized exaggerated audio-visual corrective feedback for mispronunciations. Though the effectiveness of exaggerated feedback has been demonstrated, it is still unclear how to define the appropriate degrees of exaggeration when interacting with individual learners. To fill in this gap, we interview 100 L2 English learners and 22 professional native teachers to understand their needs and experiences. Three critical metrics are proposed for both learners and teachers to identify the best exaggeration levels in both audio and visual modalities. Additionally, we incorporate the personalized dynamic feedback mechanism given the English proficiency of learners. Based on the obtained insights, a comprehensive interactive pronunciation training course is designed to help L2 learners rectify mispronunciations in a more perceptible, understandable, and discriminative manner. Extensive user studies demonstrate that our system significantly promotes the learners’ learning efficiency.
Yaohua Bu, Hang Zhou 0009, Jia Jia 0001, Shengqi Chen 0001, Dachuan Shi, Haozhe Wu, Kun Li 0003, Zhiyong Wu 0001, Yuanchun Shi, Xiaobo Lu, Ziwei Liu 0002
CHI5
2021 Towards Multi-Scale Style Control for Expressive Speech Synthesis
abstract
This paper introduces a multi-scale speech style modeling method for end-to-end expressive speech synthesis.The proposed method employs a multi-scale reference encoder to extract both the global-scale utterance-level and the local-scale quasi-phoneme-level style features of the target speech, which are then fed into the speech synthesis model as an extension to the input phoneme sequence.During training time, the multiscale style model could be jointly trained with the speech synthesis model in an end-to-end fashion.By applying the proposed method to style transfer task, experimental results indicate that the controllability of the multi-scale speech style model and the expressiveness of the synthesized speech are greatly improved.Moreover, by assigning different reference speeches to extraction of style on each scale, the flexibility of the proposed method is further revealed.
Xiang Li 0105, Changhe Song, Jingbei Li, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng
Interspeech5
2021 Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face Synthesis
abstract
People talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth wide open, while the "solemn" style is more standardized and seldomly exhibits exaggerated motions. Due to such huge differences between different styles, it is necessary to incorporate the talking style into audio-driven talking face synthesis framework. In this paper, we propose to inject style into the talking face synthesis framework through imitating arbitrary talking style of the particular reference video. Specifically, we systematically investigate talking styles with our collected Ted-HD dataset and construct style codes as several statistics of 3D morphable model (3DMM) parameters. Afterwards, we devise a latent-style-fusion (LSF) model to synthesize stylized talking faces by imitating talking styles from the style codes. We emphasize the following novel characteristics of our framework: (1) It doesn't require any annotation of the style, the talking style is learned in an unsupervised manner from talking videos in the wild. (2) It can imitate arbitrary styles from arbitrary videos, and the style codes can also be interpolated to generate new styles. Extensive experiments demonstrate that the proposed framework has the ability to synthesize more natural and expressive talking styles compared with baseline methods.
Haozhe Wu, Jia Jia 0001, Haoyu Wang 0009, Yishun Dou, Qingshan Deng
ACM Multimedia2
2021 Controllable Emphatic Speech Synthesis based on Forward Attention for Expressive Speech Synthesis
abstract
In speech interaction scenarios, speech emphasis is essential for expressing the underlying intention and attitude. Recently, end-to-end emphatic speech synthesis greatly improves the naturalness of synthetic speech, but also brings new problems: 1) lack of interpretability for how emphatic codes affect the model; 2) no separate control of emphasis on duration and on intonation and energy. We propose a novel way to build an interpretable and controllable emphatic speech synthesis framework based on forward attention. Firstly, we explicitly model the local variation of speaking rate for emphasized words and neutral words with modified forward attention to manifest emphasized words in terms of duration. The 2-layers LSTM in decoder is further divided into attention-RNN and decoder-RNN to disentangle the influence of emphasis on duration and on intonation and energy. The emphasis information is injected into decoder-RNN for highlighting emphasized words in the aspects of intonation and energy. Experimental results have shown that our model can not only provide separate control of emphasis on duration and on intonation and energy, but also generate more robust and prominent emphatic speech with high quality and naturalness.
Liangqi Liu, Jiankun Hu, Zhiyong Wu 0001, Songfan Yang, Jia Jia 0001, Helen M. Meng
SLT6
2020 PEIA: Personality and Emotion Integrated Attentive Model for Music Recommendation on Social Media Platforms
abstract
With the rapid expansion of digital music formats, it's indispensable to recommend users with their favorite music. For music recommendation, users' personality and emotion greatly affect their music preference, respectively in a long-term and short-term manner, while rich social media data provides effective feedback on these information. In this paper, aiming at music recommendation on social media platforms, we propose a Personality and Emotion Integrated Attentive model (PEIA), which fully utilizes social media data to comprehensively model users' long-term taste (personality) and short-term preference (emotion). Specifically, it takes full advantage of personality-oriented user features, emotion-oriented user features and music features of multi-faceted attributes. Hierarchical attention is employed to distinguish the important factors when incorporating the latent representations of users' personality and emotion. Extensive experiments on a large real-world dataset of 171,254 users demonstrate the effectiveness of our PEIA model which achieves an NDCG of 0.5369, outperforming the state-of-the-art methods. We also perform detailed parameter analysis and feature contribution analysis, which further verify our scheme and demonstrate the significance of co-modeling of user personality and emotion in music recommendation.
Tiancheng Shen, Jia Jia 0001, Yan Li 0068, Yihui Ma, Yaohua Bu, Hanjie Wang, Tat-Seng Chua, Wendy Hall 0001
AAAI2
2020 Mining Unfollow Behavior in Large-Scale Online Social Networks via Spatial-Temporal Interaction
abstract
Online Social Networks (OSNs) evolve through two pervasive behaviors: follow and unfollow, which respectively signify relationship creation and relationship dissolution. Researches on social network evolution mainly focus on the follow behavior, while the unfollow behavior has largely been ignored. Mining unfollow behavior is challenging because user's decision on unfollow is not only affected by the simple combination of user's attributes like informativeness and reciprocity, but also affected by the complex interaction among them. Meanwhile, prior datasets seldom contain sufficient records for inferring such complex interaction. To address these issues, we first construct a large-scale real-world Weibo1 dataset, which records detailed post content and relationship dynamics of 1.8 million Chinese users. Next, we define user's attributes as two categories: spatial attributes (e.g., social role of user) and temporal attributes (e.g., post content of user). Leveraging the constructed dataset, we systematically study how the interaction effects between user's spatial and temporal attributes contribute to the unfollow behavior. Afterwards, we propose a novel unified model with heterogeneous information (UMHI) for unfollow prediction. Specifically, our UMHI model: 1) captures user's spatial attributes through social network structure; 2) infers user's temporal attributes through user-posted content and unfollow history; and 3) models the interaction between spatial and temporal attributes by the nonlinear MLP layers. Comprehensive evaluations on the constructed dataset demonstrate that the proposed UMHI model outperforms baseline methods by 16.44 on average in terms of precision. In addition, factor analyses verify that both spatial attributes and temporal attributes are essential for mining unfollow behavior.
Haozhe Wu, Jia Jia 0001, Yaohua Bu, Xiangnan He 0001, Tat-Seng Chua
AAAI3
2020 Cross-VAE: Towards Disentangling Expression from Identity For Human Faces
abstract
Facial expression and identity are two independent yet intertwined components for representing a face. For facial expression recognition, identity can contaminate the training procedure by providing tangled but irrelevant information. In this paper, we propose to learn clearly disentangled and discriminative features that are invariant of identities for expression recognition. However, such disentanglement normally requires annotations of both expression and identity on one large dataset, which is often unavailable. Our solution is to extend conditional VAE to a crossed version named Cross-VAE, which is able to use partially labeled data to disentangle expression from identity. We emphasis the following novel characteristics of our Cross-VAE: (1) It is based on an independent assumption that the two latent representations' distributions are orthogonal. This ensures both encoded representations to be disentangled and expressive. (2) It utilizes a symmetric training procedure where the output of each encoder is fed as the condition of the other. Thus two partially labeled sets can be jointly used. Extensive experiments show that our proposed method is capable of encoding expressive and disentangled features for facial expression. Compared with the baseline methods, our model shows an improvement of 3.56% on average in terms of accuracy.
Haozhe Wu, Jia Jia 0001, Lingxi Xie, Guo-Jun Qi, Yuanchun Shi, Qi Tian 0001
ICASSP2
2020 Enhancing Music Recommendation with Social Media Content: an Attentive Multimodal Autoencoder Approach
abstract
Music recommendation methods predict users' music preference primarily based on historical ratings. Meanwhile, manifold personal factors of users are also important for the problem, and research efforts have been made to improve the recommendation performance with auxiliary user information. As an important indicator of users' personal traits and states, the numerous social media content (e.g., texts, images and short videos), however, is still hardly exploited. In this work, we systematically study the utilization of multimodal social media content for music recommendation. We define groups of both targeted handcrafted features and generic deep features for each modality, and further propose an Attentive Multimodal Autoencoder approach (AMAE) to learn cross-modal latent representations from the extracted features. Attention mechanism is also employed to integrate users' global and contextual music preference with alterable weights. Experiments demonstrate remarkable improvement of recommendation performance (+2.40% in Hit Ratio and +3.30% in NDCG), manifesting the effectiveness of our AMAE approach, as well as the significance of incorporating social media content data in music recommendation.
Tiancheng Shen, Jia Jia 0001, Yan Li 0068, Hanjie Wang
IJCNN2
2020 Re-Weighted Interval Loss for Handling Data Imbalance Problem of End-to-End Keyword Spotting
Zhiyong Wu 0001, Daode Yuan, Jian Luan 0001, Jia Jia 0001, Helen M. Meng, Binheng Song
INTERSPEECH5
2020 Visual-speech Synthesis of Exaggerated Corrective Feedback
abstract
To provide more discriminative feedback for the second language (L2) learners to better identify their mispronunciation, we propose a method for exaggerated visual-speech feedback in computer-assisted pronunciation training (CAPT). The speech exaggeration is realized by an emphatic speech generation neural network based on Tacotron, while the visual exaggeration is accomplished by ADC Viseme Blending, namely increasing Amplitude of movement, extending the phone's Duration and enhancing the color Contrast. User studies show that exaggerated feedback outperforms non-exaggerated version on helping learners with pronunciation identification and pronunciation improvement.
Yaohua Bu, Shengqi Chen 0001, Jia Jia 0001, Kun Li 0003, Xiaobo Lu
ACM Multimedia5
2020 Aesthetic-Aware Image Style Transfer
abstract
Style transfer aims to synthesize an image which inherits the content of one image while preserving a similar style of the other one. The "style'' of an image usually refers to its unique feeling conveyed from visual features, which is highly related to the aesthetic effect of the image. Aesthetic effect can be mainly decomposed as two factors: colour and texture. Previous methods like Neural Style Transfer and Colour Transfer have shown strong abilities in transferring colour and texture features. However, such approaches neglect to further disentangle colour and texture, which makes some of unique aesthetic effects designed by human artists hard to express. In this paper, we propose a novel problem called Aesthetic-Aware Image Style Transfer task, which aims to transfer colour and texture separately and independently to manipulate the aesthetic effect of an image. We propose a novel Aesthetic-Aware Model-Optimisation-Based Style Transfer (AAMOBST) model to solve this problem. Specifically, AAMOBST is a multi-reference, two-path model. It uses different reference images to decide desired colour and texture features. It can segregate colour and texture into two distinct paths and transfer them independently. Qualitative and quantitative experiments show that our model can decide colour and texture features separately and is able to keep one of them fixed while changing the other one, which is not applicable for previous methods. Furthermore, on tasks that are applicable for previous methods (such as style transfer, colour-preserved transfer and colour-only transfer), our model shows comparable abilities with other baseline methods.
Jia Jia 0001, Bei Liu 0001, Yaohua Bu, Jianlong Fu
ACM Multimedia2
2020 ChoreoNet: Towards Music to Dance Synthesis with Choreographic Action Unit
abstract
Dance and music are two highly correlated artistic forms. Synthesizing dance motions has attracted much attention recently. Most previous works conduct music-to-dance synthesis via directly music to human skeleton keypoints mapping. Meanwhile, human choreographers design dance motions from music in a two-stage manner: they firstly devise multiple choreographic dance units (CAUs), each with a series of dance motions, and then arrange the CAU sequence according to the rhythm, melody and emotion of the music. Inspired by these, we systematically study such two-stage choreography approach and construct a dataset to incorporate such choreography knowledge. Based on the constructed dataset, we design a two-stage music-to-dance synthesis framework ChoreoNet to imitate human choreography procedure. Our framework firstly devises a CAU prediction model to learn the mapping relationship between music and CAU sequences. Afterwards, we devise a spatial-temporal inpainting model to convert the CAU sequence into continuous dance motions. Experimental results demonstrate that the proposed ChoreoNet outperforms baseline methods (0.622 in terms of CAU BLEU score and 1.59 in terms of user study score).
Zijie Ye, Haozhe Wu, Jia Jia 0001, Yaohua Bu, Wei Chen 0071
ACM Multimedia3
2020 Inferring Emphasis for Real Voice Data: An Attentive Multimodal Neural Network Approach
Suping Zhou, Jia Jia 0001, Wei Chen 0071, Jialie Shen 0001
MMM (2)2
2019 Learning Discriminative Features from Spectrograms Using Center Loss for Speech Emotion Recognition
abstract
Identifying the emotional state from speech is essential for the natural interaction of the machine with the speaker. However, extracting effective features for emotion recognition is difficult, as emotions are ambiguous. We propose a novel approach to learn discriminative features from variable length spectrograms for emotion recognition by cooperating soft-max cross-entropy loss and center loss together. The soft-max cross-entropy loss enables features from different emotion categories separable, and center loss efficiently pulls the features belonging to the same emotion category to their center. By combining the two losses together, the discriminative power will be highly enhanced, which leads to network learning more effective features for emotion recognition. As demonstrated by the experimental results, after introducing center loss, both the unweighted accuracy and weighted accuracy are improved by over 3% on Mel-spectrogram input, and more than 4% on Short Time Fourier Transform spectrogram input.
Dongyang Dai, Zhiyong Wu 0001, Runnan Li, Xixin Wu, Jia Jia 0001, Helen M. Meng
ICASSP5
2019 Dilated Residual Network with Multi-head Self-attention for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) plays an important role in intelligent speech interaction. One vital challenge in SER is to extract emotion-relevant features from speech signals. In state-of-the-art SER techniques, deep learning methods, e.g, Convolutional Neural Networks (CNNs), are widely employed for feature learning and have achieved significant performance. However, in the CNN-oriented methods, two performance limitations have raised: 1) the loss of temporal structure of speech in the progressive resolution reduction; 2) the ignoring of relative dependencies between elements in suprasegmental feature sequence. In this paper, we proposed the combining use of Dilated Residual Network (DRN) and Multi-head Self-attention to alleviate the above limitations. By employing DRN, the network can retain high resolution of temporal structure in feature learning, with similar size of receptive field to CNN based approach. By employing Multi-head Self-attention, the network can model the inner dependencies between elements with different positions in the learned suprasegmental feature sequence, which enhances the importing of emotion-salient information. Experiments on emotional benchmarking dataset IEMOCAP have demonstrated the effectiveness of the proposed framework, with 11.7% to 18.6% relative improvement to state-of-the-art approaches.
Runnan Li, Zhiyong Wu 0001, Jia Jia 0001, Sheng Zhao 0002, Helen M. Meng
ICASSP3
2019 A Compact Framework for Voice Conversion Using Wavenet Conditioned on Phonetic Posteriorgrams
abstract
Voice conversion can benefit from WaveNet vocoder with improvement in converted speech's naturalness and quality. However, nowadays approaches segregate the training of conversion module and WaveNet vocoder towards different optimization objectives, which might lead to the difficulty in model tuning and coordination. In this paper, we propose a compact framework to unify the conversion and the vocoder parts. Multi-head self-attention structure and bidirectional long short-term memory (BLSTM) recurrent neural network (RNN) are employed to encode speaker independent phonetic posteriorgrams (PPGs) into an intermediate representation which is used as the condition input of WaveNet to generate target speaker's waveform. In this way, we unify the conversion and vocoder parts into a compact system in which all parameters can be tuned simultaneously for global optimization. We compared the proposed method with the baseline system that consists of separately trained conversion module and WaveNet vocoder. Subjective evaluations show that the proposed method can achieve better results in both naturalness and speaker similarity.
Zhiyong Wu 0001, Runnan Li, Shiyin Kang, Jia Jia 0001, Helen M. Meng
ICASSP5
2019 Modality Attention for End-to-end Audio-visual Speech Recognition
abstract
Audio-visual speech recognition (AVSR) system is thought to be one of the most promising solutions for robust speech recognition, especially in noisy environment. In this paper, we propose a novel multimodal attention based method for audio-visual speech recognition which could automatically learn the fused representation from both modalities based on their importance. Our method is realized using state-of-the-art sequence-to-sequence (Seq2seq) architectures. Experimental results show that relative improvements from 2% up to 36% over the auditory modality alone are obtained depending on the different signal-to-noise-ratio (SNR). Compared to the traditional feature concatenation methods, our proposed approach can achieve better recognition performance under both clean and noisy conditions. We believe modality attention based end-to-end method can be easily generalized to other multimodal tasks with correlated information.
Wei Chen 0071, Jia Jia 0001
ICASSP5
2019 Modeling Emotion Influence Using Attention-based Graph Convolutional Recurrent Network
abstract
User emotion modeling is a vital problem of social media analysis. In previous studies, content and topology information of social networks have been considered in emotion modeling tasks, but the inflence of current emotion states of other users was not considered. We define emotion influence as the emotional impact from user’s friends in social networks, which is determined by both network structure and node attributes (the features of friends). In this paper, we try to model the emotion influence to help analyze user’s emotion. The key challenges to this problem are: 1) how to combine content features and network structures together to model emotion influence; 2) how to selectively focus on the major social network information related to emotion influence. To tackle these challenges, we propose an attention-based graph convolutional recurrent network to bring in emotion influence and content data. Firstly, we use an attention-based graph convolutional network to selectively aggregate the features of the user’s friends with specific attention. Then an LSTM model is used to learn user’s own content features and emotion influence. The model we proposed is more capable of quantifying the emotion influence in social networks as well as combining them together to analyze the user emotion status. We conduct emotion classification experiments to evaluate the effectiveness of our model on a real world dataset called Sina Weibo1. Results show that our model outperforms several state-of-the-art methods.
Yulan Chen, Jia Jia 0001, Zhiyong Wu 0001
ICMI2
2019 Design and Implementation of a Disambiguity Framework for Smart Voice Controlled Devices
abstract
With about 100 million people using it recently, SVCD(Smart Voice Controlled Device) are becoming demotic. Whether at home or in an office, usually, multiple appliances are under the control of a single SVCD and several people may manipulate an SVCD simultaneously. However, present SVCD fails to handle them appropriately. In this paper, we propose a novel framework for SVCD to eliminate orders’ ambiguity for single user or multi-user. We also design an algorithm combining Word2Vec and emotion detection for the device to wipe off ambiguity. Finally, we apply our framework into a virtual smart home scene and the performance of it indicates that our strategy resolves the problems commendably.
Kehua Lei, Jia Jia 0001, Cunjun Zhang
IJCAI3
2019 Towards Discriminative Representation Learning for Speech Emotion Recognition
abstract
In intelligent speech interaction, automatic speech emotion recognition (SER) plays an important role in understanding user intention. While sentimental speech has different speaker characteristics but similar acoustic attributes, one vital challenge in SER is how to learn robust and discriminative representations for emotion inferring. In this paper, inspired by human emotion perception, we propose a novel representation learning component (RLC) for SER system, which is constructed with Multi-head Self-attention and Global Context-aware Attention Long Short-Term Memory Recurrent Neutral Network (GCA-LSTM). With the ability of Multi-head Self-attention mechanism in modeling the element-wise correlative dependencies, RLC can exploit the common patterns of sentimental speech features to enhance emotion-salient information importing in representation learning. By employing GCA-LSTM, RLC can selectively focus on emotion-salient factors with the consideration of entire utterance context, and gradually produce discriminative representation for emotion inferring. Experiments on public emotional benchmark database IEMOCAP and a tremendous realistic interaction database demonstrate the outperformance of the proposed SER framework, with 6.6% to 26.7% relative improvement on unweighted accuracy compared to state-of-the-art techniques.
Runnan Li, Zhiyong Wu 0001, Jia Jia 0001, Yaohua Bu, Sheng Zhao 0002, Helen M. Meng
IJCAI3
2019 Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-Trained BERT
Dongyang Dai, Zhiyong Wu 0001, Shiyin Kang, Xixin Wu, Jia Jia 0001, Dan Su 0002, Dong Yu 0001, Helen M. Meng
INTERSPEECH5
2019 An Online Attention-Based Model for Speech Recognition
abstract
Attention-based end-to-end models such as Listen, Attend and Spell (LAS), simplify the whole pipeline of traditional automatic speech recognition (ASR) systems and become popular in the field of speech recognition.In previous work, researchers have shown that such architectures can acquire comparable results to state-of-the-art ASR systems, especially when using a bidirectional encoder and global soft attention (GSA) mechanism.However, bidirectional encoder and GSA are two obstacles for real-time speech recognition.In this work, we aim to stream LAS baseline by removing the above two obstacles.On the encoder side, we use a latency-controlled (LC) bidirectional structure to reduce the delay of forward computation.Meanwhile, an adaptive monotonic chunk-wise attention (AMoChA) mechanism is proposed to replace GSA for the calculation of attention weight distribution.Furthermore, we propose two methods to alleviate the huge performance degradation when combining LC and AMoChA.Finally, we successfully acquire an online LAS model, LC-AMoChA, which has only 3.5% relative performance reduction to LAS baseline on our internal Mandarin corpus.
Ruchao Fan, Wei Chen 0004, Jia Jia 0001, Gang Liu 0008
INTERSPEECH4
2019 One-Shot Voice Conversion with Global Speaker Embeddings
Zhiyong Wu 0001, Dongyang Dai, Runnan Li, Shiyin Kang, Jia Jia 0001, Helen M. Meng
INTERSPEECH6
2019 Understanding the Teaching Styles by an Attention based Multi-task Cross-media Dimensional Modeling
abstract
Teaching style plays an influential role in helping students to achieve academic success. In this paper, we explore a new problem of effectively understanding teachers' teaching styles. Specifically, we study 1) how to quantitatively characterize various teachers' teaching styles for various teachers and 2) how to model the subtle relationship between cross-media teaching related data (speech, facial expressions and body motions, content et al.) and teaching styles. Using the adjectives selected from more than 10,000 feedback questionnaires provided by an educational enterprise, a novel concept called Teaching Style Semantic Space (TSSS) is developed based on the pleasure-arousal dimensional theory to describe teaching styles quantitatively and comprehensively. Then a multi-task deep learning based model, Attention-based Multi-path Multi-task Deep Neural Network (AMMDNN), is proposed to accurately and robustly capture the internal correlations between cross-media features and TSSS. Based on the benchmark dataset, we further develop a comprehensive data set including 4,541 full-annotated cross-modality teaching classes. Our experimental results demonstrate that the proposed AMMDNN outperforms (+0.0842% in terms of the concordance correlation coefficient (CCC) on average) baseline methods. To further demonstrate the advantages of the proposed TSSS and our model, several interesting case studies are carried out, such as teaching styles comparison among different teachers and courses, and leveraging the proposed method for teaching quality analysis.
Suping Zhou, Jia Jia 0001, Yufeng Yin 0002, Xiang Li 0105, Zeyang Ye, Kehua Lei, Jialie Shen 0001
ACM Multimedia2
2019 Inferring Emotions From Large-Scale Internet Voice Data
abstract
As voice dialog applications (VDAs, e.g., Siri,11http://www.apple.com/ios/siri/. Cortana,22http://www.microsoft.com/en-us/mobile/campaign-cortana/. Google Now33http://www.google.com/landing/now/.) are increasing in popularity, inferring emotions from the large-scale internet voice data generated from VDAs can help give a more reasonable and humane response. However, the tremendous amounts of users in large-scale internet voice data lead to a great diversity of users accents and expression patterns. Therefore, the traditional speech emotion recognition methods, which mainly target acted corpora, cannot effectively handle the massive and diverse amount of internet voice data. To address this issue, we carry out a series of observations, find suitable emotion categories for large-scale internet voice data, and verify the indicators of the social attributes (query time, query topic, and users location) and emotion inferring. Based on our observations, two different strategies are employed to solve the problem. First, a deep sparse neural network model that uses acoustic information, textual information, and three indicators (a temporal indicator, descriptive indicator, and geo-social indicator) as the input is proposed. Then, to capture the contextual information, we propose a hybrid emotion inference model that includes long short-term memory to capture the acoustic features and a latent dirichlet allocation to extract text features. Experiments on 93 000 utterances collected from the Sogou Voice Assistant44http://yy.sogou.com. (Chinese Siri) validate the effectiveness of the proposed methodologies. Furthermore, we compare the two methodologies and give their advantages and disadvantages.
Jia Jia 0001, Suping Zhou, Yufeng Yin 0002, Boya Wu, Wei Chen 0071
IEEE Trans. Multim.1
2018 Lookine: Let the Blind Hear a Smile
abstract
It is believed that nonverbal visual information including facial expressions, facial micro-actions and head movements plays a significant role in fundamental social communication. Unfortunately it is regretful that the blind can not achieve such necessary information. Therefore, we propose a social assistant system, Lookine, to help them to go beyond this limitation. For Lookine, we apply the novel techniques including facial expression recognition, facial action recognition and head pose estimation, and obey barrier-free principles in our design. In experiments, the algorithm evaluation and user study prove that our system has promising accuracy, good real-time performance, and great user experience.
Yaohua Bu, Jia Jia 0001, Yuhan Tang, Xuan Zang
AAAI2
2018 Inferring Emotion from Conversational Voice Data: A Semi-Supervised Multi-Path Generative Neural Network Approach
abstract
To give a more humanized response in Voice Dialogue Applications (VDAs), inferring emotion states from users’ queries may play an important role. However, in VDAs, we have tremendous amount of VDA users and massive scale of unlabeled data with high dimension features from multimodal information, which challenge the traditional speech emotion recognition methods. In this paper, to better infer emotion from conversational voice data, we proposed a semi-supervised multi-path generative neural network. Specifically, first, we build a novel supervised multi-path deep neural network framework. To avoid high dimensional input, raw features are trained by groups in local classifiers. Then high-level features of each local classifiers are concatenated as input of a global classifier. These two kinds classifiers are trained simultaneously through a single objective function to achieve a more effective and discriminative emotion inferring. To further solve the labeled-data-scarcity problem, we extend the multi-path deep neural network to a generative model based on semi-supervised variational autoencoder (semi-VAE), which is able to train the labeled and unlabeled data simultaneously. Experiment based on a 24,000 real-world dataset collected from Sogou Voice Assistant (SVAD13) and a benchmark dataset IEMOCAP show that our method significantly outperforms the existing state-of-the-art results.
Suping Zhou, Jia Jia 0001, Yufei Dong, Yufeng Yin 0002, Kehua Lei
AAAI2
2018 Emphatic Speech Generation with Conditioned Input Layer and Bidirectional LSTMS for Expressive Speech Synthesis
abstract
By highlighting the focus of an utterance to draw attention, emphasis in speech interaction plays an important role for speaker intention expressing and understanding. Therefore, emphatic speech synthesis draws increasing interest in the text-to-speech (TTS) area. For emphatic speech synthesis, three problems still exist: 1) sparseness of emphatic speech data; 2) flexibility of trained model; 3) modelling shortage for secondary emphasis. Recently, recurrent neural networks (RNNs) and their bidirectional long short term memory (BLSTM) variants based statistical parametric speech synthesis (SPSS) systems have shown their adaptability and controllability in acoustic modelling thus can solve aforementioned problems. In this paper, we propose a novel conditional input layer for conventional BLSTM-RNN based approach combining using emphasis-specific vectors and linguistic features as input to produce emphatic speech trajectories. Experimental results from objective and subjective evaluations demonstrate the proposed approach can produce emphatic speech trajectories with high quality and naturalness only requiring an additional small-scale emphatic speech corpus.
Runnan Li, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng, Lianhong Cai
ICASSP4
2018 Understanding The Aesthetic Styles of Social Images
abstract
Aesthetic perception is nearly the most direct impact people could receive from images. Recent research on image understanding is mainly focused on image analysis, recognition and classification, regardless of the aesthetic meanings embedded in images. In this paper, we systematically study the problem of understanding the aesthetic styles of social images. First, we build a two-dimensional Image Aesthetic Space (IAS) to describe image aesthetic styles quantitatively and universally. Then, we propose a Bimodal Deep Autoen-coder with Cross Edges (BDA-CE) to deeply fuse the social image related features (i.e. images' visual features, tags' textual features). Connecting BDA-CE with a regression model, we are able to map the features to the IAS. The experimental results on the benchmark dataset we build with 120 thousand Flickr images show that our model outperforms (+5.5% in terms of MSE) alternative baselines. Furthermore, we conduct an interesting case study to demonstrate the advantages of our methods.
Yihui Ma, Jia Jia 0001, Yufan Hou, Yaohua Bu
ICASSP2
2018 Inferring Emotions from Image Social Networks Using Group-Based Factor Graph Model
abstract
Inferring emotions from image social networks is a hot research topic nowadays. For image social networks (Flickr, Instagram), there is an interesting phenomenon that people would like to establish or attend virtual groups and share images with different topics and emotions in different groups. Previous researches on inferring emotions usually focus on image content and user personalization, thus leading an interesting but challenging problem: whether virtual groups can influence members(users)` emotions. In this paper, we systematically study this problem from two aspects: 1) whether group homophily in users' emotions exists in image social networks; 2) how to model this subtle and complex group homophily in image social networks. Inspired by the study results of two aspects, we introduce group information to infer emotions in image social networks, and propose a novel Group-Based Factor Graph Model (G-FGM), incorporating image content, user personalization and group information to understand the emotions behind social images better. The experimental results on a dataset containing 218, 816 emotion-labeled images from Flickr show that our model outperforms (8.6-19.4% improvement in terms of F1-Measure) several baseline methods.
Wenjing Cai, Jia Jia 0001
ICME2
2018 Mental Health Computing via Harvesting Social Media Data
abstract
Mental health has become a general concern of people nowadays. It is of vital importance to detect and manage mental health issues before they turn into severe problems. Traditional psychological interventions are reliable, but expensive and hysteretic. With the rapid development of social media, people are increasingly sharing their daily lives and interacting with friends online. Via harvesting social media data, we comprehensively study the detection of mental wellness, with two typical mental problems, stress and depression, as specific examples. Initializing with binary user-level detection, we expand our research towards multiple contexts, by considering the trigger and level of mental health problems, and involving different social media platforms of different cultures. We construct several benchmark real-world datasets for analysis and propose a series of multi-modal detection models, whose effectiveness are verified by extensive experiments. We also make in-depth analysis to reveal the underlying online behaviors regarding these mental health issues.
Jia Jia 0001
IJCAI1
2018 Cross-Domain Depression Detection via Harvesting Social Media
abstract
Depression detection is a significant issue for human well-being. In previous studies, online detection has proven effective in Twitter, enabling proactive care for depressed users. Owing to cultural differences, replicating the method to other social media platforms, such as Chinese Weibo, however, might lead to poor performance because of insufficient available labeled (self-reported depression) data for model training. In this paper, we study an interesting but challenging problem of enhancing detection in a certain target domain (e.g. Weibo) with ample Twitter data as the source domain. We first systematically analyze the depression-related feature patterns across domains and summarize two major detection challenges, namely isomerism and divergency. We further propose a cross-domain Deep Neural Network model with Feature Adaptive Transformation & Combination strategy (DNN-FATC) that transfers the relevant information across heterogeneous domains. Experiments demonstrate improved performance compared to existing heterogeneous transfer methods or training directly in the target domain (over 3.4% improvement in F1), indicating the potential of our model to enable depression detection via social media for more countries with different cultural settings.
Tiancheng Shen, Jia Jia 0001, Guangyao Shen, Fuli Feng, Xiangnan He 0001, Huan-Bo Luan, Jie Tang 0001, Thanassis Tiropanis, Tat-Seng Chua, Wendy Hall 0001
IJCAI2
2018 Emotion Recognition from Variable-Length Speech Segments Using Deep Learning on Spectrograms
Xi Ma, Zhiyong Wu 0001, Jia Jia 0001, Mingxing Xu, Helen M. Meng, Lianhong Cai
INTERSPEECH3
2018 IcooBook: When the Picture Book for Children Encounters Aesthetics of Interaction
abstract
In this work, we propose a novel PCA (Perception & Cognition & Affection) model from the prospective of aesthetics in interaction. Based on PCA, we establish a new electronic interactive picture book for children, named IcooBook. At the first level of perception, the proposed IcooBook provides interfaces of multi-sensory interaction; at the second level of cognition, IcooBook builds immersive interactive scenes; at the third level of affection, IcooBook creates high-level interaction modes based on automatic emotion recognition. The research on user study had proved the effectiveness of IcooBook in helping children being focusing on reading, getting better understanding about the context, and further encouraging children to appreciate the beauty of deep affective interaction.
Yaohua Bu, Jia Jia 0001, Xiang Li 0105, Suping Zhou, Xiaobo Lu
ACM Multimedia2
2018 Inferring User Emotive State Changes in Realistic Human-Computer Conversational Dialogs
abstract
Human-computer conversational interactions are increasingly pervasive in real-world applications, such as chatbots and virtual assistants. The user experience can be enhanced through affective design of such conversational dialogs, especially in enabling the computer to understand the emotive state in the user's input, and to generate an appropriate system response within the dialog turn. Such a system response may further influence the user's emotive state in the subsequent dialog turn. In this paper, we focus on the change in the user's emotive states in adjacent dialog turns, to which we refer as user emotive state change. We propose a multi-modal, multi-task deep learning framework to infer the user's emotive states and emotive state changes simultaneously. Multi-task learning convolution fusion auto-encoder is applied to fuse the acoustic and textual features to generate a robust representation of the user's input. Long-short term memory recurrent auto-encoder is employed to extract features of system responses at the sentence-level to better capture factors affecting user emotive states. Multi-task learned structured output layer is adopted to model the dependency of user emotive state change, conditioned upon the user input's emotive states and system response in current dialog turn. Experimental results demonstrate the effectiveness of the proposed method.
Runnan Li, Zhiyong Wu 0001, Jia Jia 0001, Jingbei Li, Wei Chen 0071, Helen M. Meng
ACM Multimedia3
2018 MAHCI 2018: The 1st Workshop on Multimedia for Accessible Human Computer Interface
abstract
In the developing of advanced Human-Computer Interaction, multimedia technology plays a fundamental role to increase usability, and accessibility of computer interfaces. The first workshop on Multimedia for Accessible Human Computer Interface (MAHCI) provides a forum to both multimedia and HCI researchers to discuss the accessible human computer interface design, development, and evaluations with the state-of-the-art multimedia technology. It also enables multimedia community to expand its interaction with the HCI industry and broaden the scope of deploying multimedia technology in practical applications. The workshop features 5 papers which cover a number of novel applications and new methodologies in a half day program.
Xueliang Liu, Benoit Huet, Jia Jia 0001
ACM Multimedia4
2018 Dance with Melody: An LSTM-autoencoder Approach to Music-oriented Dance Synthesis
abstract
Dance is greatly influenced by music. Studies on how to synthesize music-oriented dance choreography can promote research in many fields, such as dance teaching and human behavior research. Although considerable effort has been directed toward investigating the relationship between music and dance, the synthesis of appropriate dance choreography based on music remains an open problem. There are two main challenges: 1) how to choose appropriate dance figures, i.e., groups of steps that are named and specified in technical dance manuals, in accordance with music and 2) how to artistically enhance choreography in accordance with music. To solve these problems, in this paper, we propose a music-oriented dance choreography synthesis method using a long short-term memory (LSTM)-autoencoder model to extract a mapping between acoustic and motion features. Moreover, we improve our model with temporal indexes and a masking method to achieve better performance. Because of the lack of data available for model training, we constructed a music-dance dataset containing choreographies for four types of dance, totaling 907,200 frames of 3D dance motions and accompanying music, and extracted multidimensional features for model training. We employed this dataset to train and optimize the proposed models and conducted several qualitative and quantitative experiments to select the best-fitted model. Finally, our model proved to be effective and efficient in synthesizing valid choreographies that are also capable of musical expression.
Taoran Tang, Jia Jia 0001, Hanyang Mao
ACM Multimedia2
2018 AniDance: Real-Time Dance Motion Synthesize to the Song
abstract
In this paper, we present a demo named AniDance that can synthesize dance motions with melody in real-time. When users sing a song or play one in their phone to AniDance, their melody will drive the 3D-space character to dance to create a lively dance animation. In practice, we conduct a music oriented 3D-space dance motion dataset by capturing real dance performances, using LSTM-autoencoder to identify the relation between music and dance. Based on these technologies, users can create valid choreographies that capable of musical expression, witch can promote their learning ability and interest in dance and music.
Taoran Tang, Hanyang Mao, Jia Jia 0001
ACM Multimedia3
2018 AI Painting: An Aesthetic Painting Generation System
abstract
There are many great works done in image generation. However, it is still an open problem how to generate a painting, which is meeting the aesthetic rules in specific style. Therefore, in this paper, we propose a demonstration to generate a specific painting based on users' input. In the system called AI Painting, we generate an original image from content text, transfer the image into a specific aesthetic effect, simulate the image into specific artistic genre, and illustrate the painting process.
Cunjun Zhang, Kehua Lei, Jia Jia 0001, Yihui Ma
ACM Multimedia3
2017 AniDraw: When Music and Dance Meet Harmoniously
Yaohua Bu, Taoran Tang, Jia Jia 0001, Songyao Wu, Yuming You
AAAI3
2017 A Virtual Personal Fashion Consultant: Learning from the Personal Preference of Fashion
abstract
Besides fashion, personalization is another important factor of wearing. How to balance fashion trend and personal preference to better appreciate wearing is a non-trivial task. In previous work we develop a demo, Magic Mirror, to recommend clothing collocation based on the fashion trend. However, the diversity of people’s aesthetics is huge. In order to meet different demand, Magic Mirror is upgraded in this paper, and it can give out recommendations by considering both the fashion trend and personal preference, and work as a private clothing consultant. For more suitable recommendation, the virtual consultant will learn users’ tastes and preferences from their behaviors by using Genetic algorithm. Users can get collocations or matched top/bottom recommendation after choosing occasion and style. They can also get a report about their fashion state and aesthetic standpoint on recent wearing.
Jingtian Fu, Yejun Liu, Jia Jia 0001, Yihui Ma, Fanhang Meng
AAAI3
2017 SenseRun: Real-Time Running Routes Recommendation towards Providing Pleasant Running Experiences
abstract
In this demo, we develop a mobile running application, SenseRun, to involve landscape experiences for routes recommendation. We firstly define landscape experiences, perceived enjoyment from landscape as motivators for running, by public natural area and traffic density. Based on landscape experiences, we categorize locations into 3 types (natural, leisure, traffic space) and set them with different basic weight. Real-time context factors (weather, season and hour of the day) are involved to adjust the weight. We propose a multi-attributes method to recommend routes with weight based on MVT (The Marginal Value Theorem) k-shortest-paths algorithm. We also use a landscape-awareness sounds algorithm as supplementary of landscape experiences. Experimental results improve that SenseRun can enhance running experiences and is helpful to promote regular physical activities.
Jiayu Long, Jia Jia 0001
AAAI2
2017 Towards Better Understanding the Clothing Fashion Styles: A Multimodal Deep Learning Approach
abstract
In this paper, we aim to better understand the clothing fashion styles. There remain two challenges for us: 1) how to quantitatively describe the fashion styles of various clothing, 2) how to model the subtle relationship between visual features and fashion styles, especially considering the clothing collocations. Using the words that people usually use to describe clothing fashion styles on shopping websites, we build a Fashion Semantic Space (FSS) based on Kobayashi's aesthetics theory to describe clothing fashion styles quantitatively and universally. Then we propose a novel fashion-oriented multimodal deep learning based model, Bimodal Correlative Deep Autoencoder (BCDA), to capture the internal correlation in clothing collocations. Employing the benchmark dataset we build with 32133 full-body fashion show images, we use BCDA to map the visual features to the FSS. The experiment results indicate that our model outperforms (+13% in terms of MSE) several alternative baselines, confirming that our model can better understand the clothing fashion styles. To further demonstrate the advantages of our model, we conduct some interesting case studies, including fashion trends analyses of brands, clothing collocation recommendation, etc.
Yihui Ma, Jia Jia 0001, Suping Zhou, Jingtian Fu, Yejun Liu, Zijian Tong
AAAI2
2017 Multi-Task Deep Learning for User Intention Understanding in Speech Interaction Systems
abstract
Speech interaction systems have been gaining popularity in recent years. The main purpose of these systems is to generate more satisfactory responses according to users' speech utterances, in which the most critical problem is to analyze user intention. Researches show that user intention conveyed through speech is not only expressed by content, but also closely related with users' speaking manners (e.g. with or without acoustic emphasis). How to incorporate these heterogeneous attributes to infer user intention remains an open problem. In this paper, we define Intention Prominence (IP) as the semantic combination of focus by text and emphasis by speech, and propose a multi-task deep learning framework to predict IP. Specifically, we first use long short-term memory (LSTM) which is capable of modeling long short-term contextual dependencies to detect focus and emphasis, and incorporate the tasks for focus and emphasis detection with multi-task learning (MTL) to reinforce the performance of each other. We then employ Bayesian network (BN) to incorporate multimodal features (focus, emphasis, and location reflecting users' dialect conventions) to predict IP based on feature correlations. Experiments on a data set of 135,566 utterances collected from real-world Sogou Voice Assistant illustrate that our method can outperform the comparison methods over 6.9-24.5% in terms of F1-measure. Moreover, a real practice in the Sogou Voice Assistant indicates that our method can improve the performance on user intention understanding by 7%.
Yishuang Ning, Jia Jia 0001, Zhiyong Wu 0001, Runnan Li, Yongsheng An, Helen M. Meng
AAAI2
2017 Learning cross-lingual knowledge with multilingual BLSTM for emphasis detection with limited training data
abstract
Bidirectional long short-term memory (BLSTM) recurrent neural network (RNN) has achieved state-of-the-art performance in many sequence processing problems given its capability in capturing contextual information. However, for languages with limited amount of training data, it is still difficult to obtain a high quality BLSTM model for emphasis detection, the aim of which is to recognize the emphasized speech segments from natural speech. To address this problem, in this paper, we propose a multilingual BLSTM (MTL-BLSTM) model where the hidden layers are shared across different languages while the softmax output layer is language-dependent. The MTL-BLSTM can learn cross-lingual knowledge and transfer this knowledge to both languages to improve the emphasis detection performance. Experimental results demonstrate our method can outperform the comparison methods over 2-15.6% and 2.9-15.4% on the English corpus and Mandarin corpus in terms of relative F1-measure, respectively.
Yishuang Ning, Zhiyong Wu 0001, Runnan Li, Jia Jia 0001, Mingxing Xu, Helen M. Meng, Lianhong Cai
ICASSP4
2017 A systematic approach to compute perceptual distribution of monosyllables
abstract
In speech understanding, perceptual computing is widely used to quantify the perceptual properties. Previous researches on perceptual computing mainly focused on the level of phonemes (i.e. consonants and vowels). However, perceptual measurement in the level of syllables is also needed in scenarios such as speech recognition. To tackle this problem, we propose a systematic approach to calculate the perceptual distribution of monosyllables. It is composed of three parts. First, we generate a feature vector from each monosyllable based on acoustic property. Second, we construct the perceptual space based on the perception distance of every two feature vectors. Third, we measure the perceptual distribution for these monosyllables based on the perceptual space and a constraint matrix. Experiments show that 1) cluster results are in accordance with articulation position category in acoustics, 2) recognition rate of audiometry is within the standard range of performance-intensity function, 3) distribution of paracusia is consistent with the computation results of perceptual distribution.
Yu-Hao Wu, Jia Jia 0001, Lianhong Cai
ICASSP2
2017 Inferring emotions from heterogeneous social media data: A Cross-media Auto-Encoder solution
abstract
Social media is rocking the world in recent year, which makes modeling social media contents important. However, the heterogeneity of social media data is the main constraint. This paper focuses on inferring emotions from large-scale social media data. Tweets on social media platform, always containing heterogeneous information from different combinations of modalities, are utilized to construct a cross-media dataset. How to integrate cross-media information and solve the problem of modality deficiency are main challenges. To address those challenges, this paper proposes a Cross-media Auto-Encoder(CAE) to infer emotions on cross-media data, and CAE is designed to reconstruct missing modalities and integrate heterogeneous representations. In our experiments, We employ a dataset of 226,113 tweets to infer emotions of tweets, and our method outperforms several machine learning methods (+11.11% in terms of F1-measure). Feature contribution analysis also verifies the importance of adopting cross-media features.
Shumei Zhang, Jia Jia 0001, Yishuang Ning
ICASSP2
2017 Depression Detection via Harvesting Social Media: A Multimodal Dictionary Learning Solution
abstract
Depression is a major contributor to the overall global burden of diseases. Traditionally, doctors diagnose depressed people face to face via referring to clinical depression criteria. However, more than 70% of the patients would not consult doctors at early stages of depression, which leads to further deterioration of their conditions. Meanwhile, people are increasingly relying on social media to disclose emotions and sharing their daily lives, thus social media have successfully been leveraged for helping detect physical and mental diseases. Inspired by these, our work aims to make timely depression detection via harvesting social media data. We construct well-labeled depression and non-depression dataset on Twitter, and extract six depression-related feature groups covering not only the clinical depression criteria, but also online behaviors on social media. With these feature groups, we propose a multimodal depressive dictionary learning model to detect the depressed users on Twitter. A series of experiments are conducted to validate this model, which outperforms (+3% to +10%) several baselines. Finally, we analyze a large-scale dataset on Twitter to reveal the underlying online behaviors between depressed and non-depressed users.
Guangyao Shen, Jia Jia 0001, Liqiang Nie, Fuli Feng, Cunjun Zhang, Tianrui Hu, Tat-Seng Chua, Wenwu Zhu 0001
IJCAI2
2017 Speech Emotion Recognition with Emotion-Pair Based Framework Considering Emotion Distribution Information in Dimensional Emotion Space
Xi Ma, Zhiyong Wu 0001, Jia Jia 0001, Mingxing Xu, Helen M. Meng, Lianhong Cai
INTERSPEECH3
2017 PIC2DISH: A Customized Cooking Assistant System
abstract
The art of cooking is always fascinating. Nevertheless, reproducing a delicious dish that one has never encountered before is not easy. Even if the name of dish is known and the corresponding recipe could be retrieved, the right ingredients for cooking the dish may not be available due to factors such as geography region or season. Furthermore, knowing how to cut, cook and control timing may be challenging for one whose has no cooking experience. In this paper, an all-around cooking assistant mobile app, named Pic2Dish, is developed to help users who would like to cook a dish but neither know the name of dish nor has cooking skill. Basically, by inputting a picture of the dish and the list of ingredients at hand, Pic2Dish automatically recognizes the dish name and recommends a customized recipe together with video clips to guide user on how to cook the dish. Importantly, the recommended recipe is modified from a retrieved recipe that best matches the given dish, with missing ingredients being replaced with the available ingredients that match dish context and taste. The whole process involves the recognition of dishes with convolutional neural network, classification of key and non-key ingredients, and context analysis of ingredient relationship and their cooking/cutting methods. The user studies, which recruit real users to cook dishes by using Pic2Dish, shows the usefulness of the app.
Yongsheng An, Jingjing Chen 0001, Chong-Wah Ngo, Jia Jia 0001, Huan-Bo Luan, Tat-Seng Chua
ACM Multimedia5
2017 Multi-scale Context Based Attention for Dynamic Music Emotion Prediction
abstract
Dynamic music emotion prediction is to recognize the continuous emotion information in music, which is necessary for music retrieval and recommendation. In this paper, we adopt the dimensional valence-arousal (V-A) emotion model to represent the dynamic emotion in music. In our opinion, music and V-A emotion label do not have the one-to-one correspondence in the time domain, while the expression of music emotion at one moment is the accumulation of previous music content for a period of time, so we propose Long Short-Term Memory (LSTM) based sequence-to-one mapping for dynamic music emotion prediction. Based on this sequence-to-one music emotion mapping, it is proved that different time scales' preceding content has an influence on the LSTM model's performance, so we further propose the Multi-scale Context based Attention (MCA) for dynamic music emotion prediction. We evaluate our proposed method on the database of Emotion in Music task at MediaEval 2015, and the results show that our proposed method outperforms most of the models using the same features and achieves a competitive performance with the state-of-the-art methods.
Xinxing Li, Mingxing Xu, Jia Jia 0001, Lianhong Cai
ACM Multimedia4
2017 Analyzing and Identifying Teens' Stressful Periods and Stressor Events From a Microblog
abstract
Increased health problems among adolescents caused by psychological stress have aroused worldwide attention. Long-standing stress without targeted assistance and guidance negatively impacts the healthy growth of adolescents, threatening the future development of our society. So far, research focused on detecting adolescent psychological stress revealed from each individual post on microblogs. However, beyond stressful moments, identifying teens' stressful periods and stressor events that trigger each stressful period is more desirable to understand the stress from appearance to essence. In this paper, we define the problem of identifying teens' stressful periods and stressor events from the open social media microblog. Starting from a case study of adolescents' posting behaviors during stressful school events, we build a Poisson-based probability model for the correlation between stressor events and stressful posting behaviors through a series of posts on Tencent Weibo (referred to as the microblog throughout the paper). With the model, we discover teens' maximal stressful periods and further extract details of possible stressor events that cause the stressful periods. We generalize and present the extracted stressor events in a hierarchy based on common stress dimensions and event types. Taking 122 scheduled stressful study-related events in a high school as the ground truth, we test the approach on 124 students' posts from January 1, 2012 to February 1, 2015 and obtain some promising experimental results: (stressful periods: recall 0.761, precision 0.737, and F1-measure 0.734) and (top-3 stressor events: recall 0.763, precision 0.756, and F1-measure 0.759). The most prominent stressor events extracted are in the self-cognition domain, followed by the school life domain. This conforms to the adolescent psychological investigation result that problems in school life usually accompanied with teens' inner cognition problems. Compared with the state-of-the-art top-1 personal life event detection approach, our stressor event detection method is 13.72% higher in precision, 19.18% higher in recall, and 16.50% higher in F1-measure, demonstrating the effectiveness of our proposed framework.
Qi Li 0006, Yuanyuan Xue, Liang Zhao 0023, Jia Jia 0001
IEEE J. Biomed. Health Informatics4
2017 Detecting Stress Based on Social Interactions in Social Networks
abstract
Psychological stress is threatening people's health. It is non-trivial to detect stress timely for proactive care. With the popularity of social media, people are used to sharing their daily activities and interacting with friends on social media platforms, making it feasible to leverage online social network data for stress detection. In this paper, we find that users stress state is closely related to that of his/her friends in social media, and we employ a large-scale dataset from real-world social platforms to systematically study the correlation of users' stress states and social interactions. We first define a set of stress-related textual, visual, and social attributes from various aspects, and then propose a novel hybrid model - a factor graph model combined with Convolutional Neural Network to leverage tweet content and social interaction information for stress detection. Experimental results show that the proposed model can improve the detection performance by 6-9 percent in F1-score. By further analyzing the social interaction data, we also discover several intriguing phenomena, i.e., the number of social structures of sparse connections (i.e., with no delta connections) of stressed users is around 14 percent higher than that of non-stressed users, indicating that the social structure of stressed users' friends tend to be less connected and less complicated than that of non-stressed users.
Huijie Lin, Jia Jia 0001, Jiezhong Qiu, Yongfeng Zhang 0003, Guangyao Shen, Lexing Xie, Jie Tang 0001, Tat-Seng Chua
IEEE Trans. Knowl. Data Eng.2
2017 Mobile Contextual Recommender System for Online Social Media
abstract
Exponential growth of media consumption in online social networks demands effective recommendation to improve the quality of experience especially for on-the-go mobile users. By means of large-scale trace-driven measurements over mobile Twitter traces from users, we reveal the significance of affective features in shaping users' social media behaviors. Existing recommender systems however, rarely support such psychological effect in real-life. To capture such effect, in this paper we propose Kaleido, a real mobile system that achieves an online social media recommendation solution by taking affective context into account. Specifically, we design a machine learning mechanism to infer the affective pulse of online social media. Furthermore, a cluster-based latent bias model (LBM) is provided for jointly training the affective pulse as well as user's behavior, location, and social contexts. Our comprehensive trace-driven experiments on Android prototype expose a superior prediction accuracy of 87 percent, which has 25 percent accuracy superior to existing mobile recommender systems. Moreover, by enabling users to offload their machine learning procedures to the deployed edge-cloud testbed, our system achieves speed-up of a factor of 1,000 against the local data training execution on smartphones.
Chao Wu 0002, Yaoxue Zhang, Jia Jia 0001, Wenwu Zhu 0001
IEEE Trans. Mob. Comput.3
2017 Inferring Emotional Tags From Social Images With User Demographics
abstract
Social images, which are images uploaded and shared on social networks, are used to express users’ emotions. Inferring emotional tags from social images is of great importance; it can benefit many applications, such as image retrieval and recommendation. Whereas previous related research has primarily focused on exploring image visual features, we aim to address this problem by studying whether user demographics make a difference regarding users’ emotional tags of social images. We first consider how to model the emotions of social images. Then, we investigate how user demographics, such as gender, marital status, and occupation, are related to the emotional tags of social images. A partially labeled factor graph model named the demographics factor graph model ( D-FGM ) is proposed to leverage the uncovered patterns. Experiments on a data set collected from the world's largest image sharing website Flickr 1 1 [Online]. Available: http://www.flickr.com/ confirm the accuracy of the proposed model. We also find some interesting phenomena. For example, men and women have different patterns to tag “anger” for social images.
Boya Wu, Jia Jia 0001, Yang Yang 0009, Peijun Zhao, Jie Tang 0001, Qi Tian 0001
IEEE Trans. Multim.2
2017 Trip Outfits Advisor: Location-Oriented Clothing Recommendation
abstract
When packing for a journey, have you ever asked “what clothes should I take with me?” Wearing appropriate and aesthetically pleasing clothing when traveling is a concern for many of us. Our data observation of photos from several popular travel websites reveals that people's choice of clothing items and their color combinations have strong correlations with the weather, the season, and the main type of attraction at the destination. This leads to an interesting and novel problem: can the correlation between clothing and locations be automatically learned from social photos and leveraged for location-oriented clothing recommendations? In this paper, we systematically study this problem and propose a hybrid multilabel convolutional neural network combined with the support vector machine (mCNN-SVM) approach to capture the intrinsic and complex correlations between clothing attributes and location attributes. Specifically, we adapt the CNN architecture to multilabel learning and fine-tune it using each fine-grained clothing item. Then, the recognized items are fed to the SVM to learn the correlations. Experiments on three fashion datasets and a benchmark journey outfit dataset show that our proposed approach outperforms several baselines by over 10.52-16.38% in terms of the mAP for clothing item recognition and outperforms several alternative methods by over 9.59-29.41% in terms of the mAP when ranking clothing by appropriateness for travel destinations. Finally, an interesting case study demonstrates the effectiveness of our method by answering what items to wear, how to match them, and how to dress in an aesthetically pleasing manner for a journey.
Xishan Zhang, Jia Jia 0001, Ke Gao 0012, Yongdong Zhang 0001, Dongming Zhang 0004, Jintao Li 0001, Qi Tian 0001
IEEE Trans. Multim.2
2016 Learning to Appreciate the Aesthetic Effects of Clothing
abstract
How do people describe clothing? The words like “formal”or "casual" are usually used. However, recent works often focus on recognizing or extracting visual features (e.g., sleeve length, color distribution and clothing pattern) from clothing images accurately. How can we bridge the gap between the visual features and the aesthetic words? In this paper, we formulate this task to a novel three-level framework: visual features(VF) - image-scale space (ISS) - aesthetic words space(AWS). Leveraging the art-field image-scale space served as an intermediate layer, we first propose a Stacked Denoising Autoencoder Guided by CorrelativeLabels (SDAE-GCL) to map the visual features to the image-scale space; and then according to the semantic distances computed byWordNet::Similarity, we map the most often used aesthetic words in online clothing shops to the image-scale space too. Employing upper body menswear images downloaded from several global online clothing shops as experimental data, the results indicate that the proposed three-level framework can help to capture the subtle relationship between visual features and aesthetic words better compared to several baselines. To demonstrate that our three-level framework and its implementation methods are universally applicable, we finally present some interesting analyses on the fashion trend of menswear in the last 10 years.
Jia Jia 0001, Guangyao Shen, Tao He 0016, Zhiyuan Liu 0001, Huan-Bo Luan
AAAI1
2016 Moodee: An Intelligent Mobile Companion for Sensing Your Stress from Your Social Media Postings
abstract
In this demo, we build a practical mobile application, Moodee,to help detect and release users’ psychological stress byleveraging users’ social media data in online social networks,and provide an interactive user interface to present users’and friends’ psychological stress states in an visualized andintuitional way.Given users’ online social media data as input, Moodee intelligentlyand automatically detects users’ stress states. Moreover,Moodee would recommend users with different linksto help release their stress. The main technology of this demois a novel hybrid model - a factor graph model combinedwith Deep Neural Network, which can leverage social mediacontent and social interaction information for stress detection.We think that Moodee can be helpful to people’s mentalhealth, which is a vital problem in modern world.
Huijie Lin, Jia Jia 0001, Enze Zhou, Jingtian Fu, Yejun Liu, Huan-Bo Luan
AAAI2
2016 Representation Learning of Knowledge Graphs with Entity Descriptions
abstract
Representation learning (RL) of knowledge graphs aims to project both entities and relations into a continuous low-dimensional space. Most methods concentrate on learning representations with knowledge triples indicating relations between entities. In fact, in most knowledge graphs there are usually concise descriptions for entities, which cannot be well utilized by existing methods. In this paper, we propose a novel RL method for knowledge graphs taking advantages of entity descriptions. More specifically, we explore two encoders, including continuous bag-of-words and deep convolutional neural models to encode semantics of entity descriptions. We further learn knowledge representations with both triples and descriptions. We evaluate our method on two tasks, including knowledge graph completion and entity classification. Experimental results on real-world datasets show that, our method outperforms other baselines on the two tasks, especially under the zero-shot setting, which indicates that our method is capable of building representations for novel entities according to their descriptions. The source code of this paper can be obtained from https://github.com/xrb92/DKRL.
Ruobing Xie, Zhiyuan Liu 0001, Jia Jia 0001, Huan-Bo Luan, Maosong Sun 0001
AAAI3
2016 Social Role-Aware Emotion Contagion in Image Social Networks
abstract
Psychological theories suggest that emotion represents the state of mind and instinctive responses of one’s cognitive system (Cannon 1927). Emotions are a complex state of feeling that results in physical and psychological changes that influence our behavior. In this paper, we study an interesting problem of emotion contagion in social networks. In particular, by employing an image social network (Flickr) as the basis of our study, we try to unveil how users’ emotional statuses influence each other and how users’ positions in the social network affect their influential strength on emotion. We develop a probabilistic framework to formalize the problem into a role-aware contagion model. The model is able to predict users’ emotional statuses based on their historical emotional statuses and social structures. Experiments on a large Flickr dataset show that the proposed model significantly outperforms (+31% in terms of F1-score) several alternative methods in predicting users’ emotional status. We also discover several intriguing phenomena. For example, the probability that a user feels happy is roughly linear to the number of friends who are also happy; but taking a closer look, the happiness probability is superlinear to the number of happy friends who act as opinion leaders (Page et al. 1999) in the network and sublinear in the number of happy friends who span structural holes (Burt 2001). This offers a new opportunity to understand the underlying mechanism of emotional contagion in online social networks.
Yang Yang 0009, Jia Jia 0001, Boya Wu, Jie Tang 0001
AAAI2
2016 Low level descriptors based DBLSTM bottleneck feature for speech driven talking avatar
abstract
Speech is bimodal in nature. There are close correlations between the acoustic speech signals and the visual gestures such as lip movements, facial expressions and head motions. For speech driven talking avatar, how to derive more representative acoustic features from which to predict more accurate and realistic visual gestures still remains the research problem. Inspired by the promising performance of low level descriptors (LLD) in speech emotion recognition, in this work, we investigate the usage of LLD feature for the task of speech driven talking avatar. Furthermore, visual gestures also demonstrate correlations with not only context information of past or future acoustic features (e.g. anticipatory co-articulation phenomena) but also textual information (e.g. textual hints for lip movement). To incorporate such information, we also propose to use deep bidirectional long short-term memory (DBLSTM) as the bottleneck feature extractor, which can combine LLD feature with contextual information. Experimental results indicate that the proposed LLD based DBLSTM bottleneck feature outperforms the conventional spectrum related features for the task of speech driven talking avatar, and more sophisticated contextual information can further improve the performance.
Xinyu Lan, Xu Li 0015, Yishuang Ning, Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Lianhong Cai
ICASSP6
2016 Inferring users' emotions for human-mobile voice dialogue applications
abstract
In this paper, we tackle the problem of inferring users' emotions in real-world Voice Dialogue Applications (VDAs, Siri1, Cortana2, etc.). We first conduct an investigation, indicating that besides the text information of users' queries, the acoustic information and query attributes are very important in inferring emotions in VDAs. To integrate the information above, we propose a Hybrid Emotion Inference Model (HEIM), which involves a Latent Dirichlet Allocation (LDA) to extract text features and a Long Short-Term Memory (LSTM) to model the acoustic features. To further improve accuracy, a Recurrent Autoencoder Guided by Query Attributes (RAGQA) which incorporates other emotion-related query attributes is proposed in HEIM to pre-train LSTM. The accuracy of HEIM on a data set collected from Sogou Voice Assistant3(Chinese Siri) containing 93,000 utterances achieves 75.2%, which outperforms state-of-the-art methods for 33.5–38.5%. Specifically, we discover that on average, the acoustic information enhances the performance for 46.6%, while query attributes further enhance the performance for 6.5%.
Boya Wu, Jia Jia 0001, Tao He 0016, Xiaoyuan Yi, Yishuang Ning
ICME2
2016 What Does Social Media Say about Your Stress?
Huijie Lin, Jia Jia 0001, Liqiang Nie, Guangyao Shen, Tat-Seng Chua
IJCAI2
2016 Phoneme Embedding and its Application to Speech Driven Talking Avatar Synthesis
Xu Li 0015, Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Xiaoyan Lou, Lianhong Cai
INTERSPEECH4
2016 Expressive Speech Driven Talking Avatar Synthesis with DBLSTM Using Limited Amount of Emotional Bimodal Data
Xu Li 0015, Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Xiaoyan Lou, Lianhong Cai
INTERSPEECH4
2016 Magic Mirror: A Virtual Fashion Consultant
abstract
What should I wear? We present Magic Mirror, a virtual fashion consultant, which can parse, appreciate and recommend the wearing. Magic Mirror is designed with a large display and Kinect to simulate the real mirror and interact with users in augmented reality. Internally, Magic Mirror is a practical appreciation system for automatic aesthetics-oriented clothing analysis. Specifically, we focus on the clothing collocation rather than the single one, the style (aesthetic words) rather than the visual features. We bridge the gap between the visual features and aesthetic words of clothing collocation to enable the computer to learn appreciating the clothing collocation. Finally, both object and subject evaluations verify the effectiveness of the proposed algorithm and Magic Mirror system.
Yejun Liu, Jia Jia 0001, Jingtian Fu, Yihui Ma, Zijian Tong
ACM Multimedia2
2016 Affective Contextual Mobile Recommender System
abstract
Exponential growth of media consumption in online social networks demands effective recommendation to improve the quality of experience especially for on-the-go mobile users. By means of large-scale trace-driven measurements over mobile Twitter traces from users, we reveal the significance of affective features in shaping users' social media behaviors. Existing recommender systems however, rarely support this psychological effect in real-life. To capture this effect, in this paper we propose Kaleido, a real mobile system to achieve an affect-aware learning-based social media recommendation.Specifically, we design a machine learning mechanism to infer the affective feature within media contents. Furthermore, a cluster-based latent bias model is provided for jointly training the affect, behavior and social contexts. Our comprehensive experiments on Android prototype expose a superior prediction accuracy of 82%, with more than 20% accuracy improvement over existing mobile recommender systems. Moreover, by enabling users to offload their machine learning procedures to the deployed edge-cloud testbed, our system achieves speed-up of a factor of 1,000 against the local data training execution on smartphones.
Chao Wu 0002, Jia Jia 0001, Wenwu Zhu 0001, Xu Chen 0004, Yaoxue Zhang
ACM Multimedia2
2016 Analysis of Teens' Chronic Stress on Micro-blog
Yuanyuan Xue, Qi Li 0006, Liang Zhao 0023, Jia Jia 0001, Feng Yu 0027, David A. Clifton
WISE (2)4
2016 Learning robust uniform features for cross-media social data by using cross autoencoders
Quan Guo, Jia Jia 0001, Guangyao Shen, Lei Zhang 0005, Lianhong Cai, Zhang Yi 0001
Knowl. Based Syst.2
2015 Understanding speaking styles of internet speech data with LSTM and low-resource training
abstract
Speech are widely used to express one's emotion, intention, desire, etc. in social network communication, deriving abundant of internet speech data with different speaking styles. Such data provides a good resource for social multimedia research. However, regarding different styles are mixed together in the internet speech data, how to classify such data remains a challenging problem. In previous work, utterance-level statistics of acoustic features are utilized as features in classifying speaking styles, ignoring the local context information. Long short-term memory (LSTM) recurrent neural network (RNN) has achieved exciting success in lots of research areas, such as speech recognition. It is able to retrieve context information for long time duration, which is important in characterizing speaking styles. To train LSTM, huge number of labeled training data is required. While for the scenario of internet speech data classification, it is quite difficult to get such large scale labeled data. On the other hand, we can get some publicly available data for other tasks (such as speech emotion recognition), which offers us a new possibility to exploit LSTM in the low-resource task. We adopt retraining strategy to train LSTM to recognize speaking styles in speech data by training the network on emotion and speaking style datasets sequentially without reset the weights of the network. Experimental results demonstrate that retraining improves the training speed and the accuracy of network in speaking style classification.
Xixin Wu, Zhiyong Wu 0001, Yishuang Ning, Jia Jia 0001, Lianhong Cai, Helen M. Meng
ACII4
2015 HMM-based emphatic speech synthesis for corrective feedback in computer-aided pronunciation training
abstract
This paper investigates the incorporation of hidden Markov model (HMM) based emphatic speech synthesis for audio exaggeration into an audio-visual speech synthesis framework for the corrective feedback in computer-aided pronunciation training (CAPT). To improve the voice quality of the synthetic emphatic speech, this paper proposes a new method for HMM training. In this method, the contextual questions for decision tree building are extended by considering the emphasis-related information. HMMs are then trained using a small scale emphatic corpus together with a large scale neutral corpus. The emphatic corpus is used to ensure the quality of the emphatic speech segments whereas the neutral corpus is to further improve the quality of both the non-emphatic speech segments and the emphatic ones. Finally, emphatic speech synthesis is achieved by extending the Flite+hts_engine. Experimental results show that our method can synthesize emphatic speech with high quality and make the feedback more discriminatively perceptible.
Yishuang Ning, Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng, Lianhong Cai
ICASSP3
2015 Understanding the emotions behind social images: Inferring with user demographics
abstract
Understanding the essential emotions behind social images is of vital importance: it can benefit many applications such as image retrieval and personalized recommendation. While previous related research mostly focuses on the image visual features, in this paper, we aim to tackle this problem by “linking inferring with users' demographics”. Specifically, we propose a partially-labeled factor graph model named D-FGM, to predict the emotions embedded in social images not only by the image visual features, but also by the information of users' demographics. We investigate whether users' demographics like gender, marital status and occupation are related to emotions of social images, and then leverage the uncovered patterns into modeling as different factors. Experiments on a data set from the world's largest image sharing website Flickr1 confirm the accuracy of the proposed model. The effectiveness of the users' demographics factors is also verified by the factor contribution analysis, which reveals some interesting behavioral phenomena as well.
Boya Wu, Jia Jia 0001, Yang Yang 0009, Peijun Zhao, Jie Tang 0001
ICME2
2015 MPHA: A Personal Hearing Doctor Based on Mobile Devices
abstract
As more and more people inquire to know their hearing level condition, audiometry is becoming increasingly important. However, traditional audiometric method requires the involvement of audiometers, which are very expensive and time consuming. In this paper, we present mobile personal hearing assessment (MPHA), a novel interactive mode for testing hearing level based on mobile devices. MPHA, 1) provides a general method to calibrate sound intensity for mobile devices to guarantee the reliability and validity of the audiometry system; 2) designs an audiometric correction algorithm for the real noisy audiometric environment. The experimental results show that MPHA is reliable and valid compared with conventional audiometric assessment.
Yu-Hao Wu, Jia Jia 0001, Wai-Kim Leung, Yejun Liu, Lianhong Cai
ICMI2
2015 Release Adolescent Stress by Virtual Chatting
Qi Li 0006, Yuanyuan Xue, Taoran Cheng, Shuangqing Xu, Jia Jia 0001
ICWE6
2015 Using tilt for automatic emphasis detection with Bayesian networks
Yishuang Ning, Zhiyong Wu 0001, Xiaoyan Lou, Helen M. Meng, Jia Jia 0001, Lianhong Cai
INTERSPEECH5
2015 Generating emphatic speech with hidden Markov model for expressive speech synthesis
Zhiyong Wu 0001, Yishuang Ning, Xiao Zang, Jia Jia 0001, Helen M. Meng, Lianhong Cai
Multim. Tools Appl.4
2015 Expressive talking avatar synthesis and animation
Lei Xie 0001, Jia Jia 0001, Helen M. Meng, Zhigang Deng 0001
Multim. Tools Appl.2
2015 Modeling Emotion Influence in Image Social Networks
abstract
We study emotion influence in large image social networks. We focus on users' emotions reflected by images that they have uploaded and social influence that plays a role in changing users' emotions. We first verify the existence of emotion influence in the image networks, and then propose a probabilistic factor graph based emotion influence model to answer the questions of “who influences whom”. Employing a real network from Flickr as the basis in our empirical study, we evaluate the effectiveness of different factors in the proposed model with in-depth data analysis. The learned influence is fundamental for social network analysis and can be applied to many applications. We consider using the influence to help predict users' emotions and our experiments can significantly improve the prediction accuracy (3.0-26.2 percent) over several alternative methods such as Naive Bayesian, SVM (Support Vector Machine) or traditional Graph Model. We further examine the behavior of the emotion influence model, and find that more social interactions correlate with higher emotion influence between two users, and the influence of negative emotions is stronger than positive ones.
Xiaohui Wang 0004, Jia Jia 0001, Jie Tang 0001, Boya Wu, Lianhong Cai, Lexing Xie
IEEE Trans. Affect. Comput.2
2014 How Do Your Friends on Social Media Disclose Your Emotions?
abstract
Extracting emotions from images has attracted much interest, in particular with the rapid development of social networks. The emotional impact is very important for understanding the intrinsic meanings of images. Despite many studies having been done, most existing methods focus on image content, but ignore the emotion of the user who published the image. One interesting question is: How does social effect correlate with the emotion expressed in an image? Specifically, can we leverage friends interactions (e.g., discussions) related to an image to help extract the emotions? In this paper, we formally formalize the problem and propose a novel emotion learning method by jointly modeling images posted by social users and comments added by their friends. One advantage of the model is that it can distinguish those comments that are closely related to the emotion expression for an image from the other irrelevant ones. Experiments on an open Flickr dataset show that the proposed model can significantly improve (+37.4% by F1) the accuracy for inferring user emotions. More interestingly, we found that half of the improvements are due to interactions between 1.0% of the closest friends.
Yang Yang 0009, Jia Jia 0001, Shumei Zhang, Boya Wu, Qicong Chen, Juan-Zi Li, Chunxiao Xing, Jie Tang 0001
AAAI2
2014 Helping Teenagers Relieve Psychological Pressures: A Micro-blog Based System
abstract
The rapid development of economy and society brings un-precedentedly intensive competition and adolescent psycho-logical pressures to current teenagers. If these psychological pressures could not be resolved properly, they will turn to mental problems, which will finally lead to serious conse-quences, such as suicide or aggressive behaviors. Tradition-al face-to-face psychological diagnosis and treatment cannot meet the demand of relieving teenagers ’ stress completely due to its lack of timeliness and diversity. With micro-blog becoming a popular media channel for teenagers ’ informa-tion acquisition, interaction, self-expression and emotion re-lease, we present a system called tHelper for sensing and easing teenagers ’ psychological pressures in study, commu-nication, affection, or self-recognition through micro-blog.
Qi Li 0006, Yuanyuan Xue, Jia Jia 0001
EDBT3
2014 Psychological stress detection from cross-media microblog data using Deep Sparse Neural Network
abstract
Long-term stress may lead to many severe physical and mental problems. Traditional psychological stress detection usually relies on the active individual participation, which makes the detection labor-consuming, time-costing and hysteretic. With the rapid development of social networks, people become more and more willing to share moods via microblog platforms. In this paper, we propose an automatic stress detection method from cross-media microblog data. We construct a three-level framework to formulate the problem. We first obtain a set of low-level features from the tweets. Then we define and extract middle-level representations based on psychological and art theories: linguistic attributes from tweets' texts, visual attributes from tweets' images, and social attributes from tweets' comments, retweets and favorites. Finally, a Deep Sparse Neural Network is designed to learn the stress categories incorporating the cross-media attributes. Experiment results show that the proposed method is effective and efficient on detecting psychological stress from microblog data.
Huijie Lin, Jia Jia 0001, Quan Guo, Yuanyuan Xue, Lianhong Cai
ICME2
2014 Acoustics, content and geo-information based sentiment prediction from large-scale networked voice data
abstract
Sentiment analysis from large-scale networked data attracts increasing attention in recent years. Most previous works on sentiment prediction mainly focus on text or image data. However, voice is the most natural and direct way to express people's sentiments in real-time. With the rapid development of smart phone voice dialogue applications (e.g., Siri and Sogou Voice Assistant), the large-scale networked voice data can help us better quantitatively understand the sentimental world we live in. In this paper, we study the problem of sentiment prediction from large-scale networked voice data. In particular, we first investigate the data observations and underlying sentiment patterns in human-mobile voice communication. Then we propose a deep sparse neural network (DSNN) model to incorporate acoustic features, content information and geo-information to automatically predict sentiments. The effectiveness of the proposed model is verified by the experiments on a real dataset from Sogou Voice Assistant application.
Zhu Ren, Jia Jia 0001, Quan Guo, Kuo Zhang 0001, Lianhong Cai
ICME2
2014 Using conditional random fields to predict focus word pair in spontaneous spoken English
abstract
This paper addresses the problem of automatically labeling focus word pairs in spontaneous spoken English, where a focus word pair refers to salient part of text or speech and the word motivating it.The prediction of focus word pairs is important for speech applications such as expressive text-tospeech (TTS) synthesis and speech recognition.It can also help in better textual and intention understanding for spoken dialog systems.Traditional approaches such as support vector machines (SVMs) prediction neglect the dependency between words and meet the obstacle of the imbalanced distribution of positive and negative samples of dataset.This paper introduces conditional random fields (CRFs) to the task of automatically predicting focus word pair from lexical, syntactic and semantic features.Furthermore, several new features related to syntactic and semantic information are proposed to achieve better performance.Experiments on the publicly available Switchboard corpus demonstrate that CRF model outperforms the baseline and SVM model for focus word pair prediction, and newly proposed features can further improve performance for CRF based predictor.Specifically, compared to the low recall rate of 11.31% achieved by the SVM model, the proposed CRF based predictor can yield a high recall rate of 70.88% with little impact on precision.
Xiao Zang, Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Lianhong Cai
INTERSPEECH4
2014 User-level psychological stress detection from social media using deep neural network
abstract
It is of significant importance to detect and manage stress before it turns into severe problems. However, existing stress detection methods usually rely on psychological scales or physiological devices, making the detection complicated and costly. In this paper, we explore to automatically detect individuals' psychological stress via social media. Employing real online micro-blog data, we first investigate the correlations between users' stress and their tweeting content, social engagement and behavior patterns. Then we define two types of stress-related attributes: 1) low-level content attributes from a single tweet, including text, images and social interactions; 2) user-scope statistical attributes through their weekly micro-blog postings, leveraging information of tweeting time, tweeting types and linguistic styles. To combine content attributes with statistical attributes, we further design a convolutional neural network (CNN) with cross autoencoders to generate user-scope content attributes from low-level content attributes. Finally, we propose a deep neural network (DNN) model to incorporate the two types of user-scope attributes to detect users' psychological stress. We test the trained model on four different datasets from major micro-blog platforms including Sina Weibo, Tencent Weibo and Twitter. Experimental results show that the proposed model is effective and efficient on detecting psychological stress from micro-blog data. We believe our model would be useful in developing stress detection tools for mental health agencies and individuals.
Huijie Lin, Jia Jia 0001, Quan Guo, Yuanyuan Xue, Qi Li 0006, Lianhong Cai
ACM Multimedia2
2014 Learning to Infer Public Emotions from Large-Scale Networked Voice Data
Zhu Ren, Jia Jia 0001, Lianhong Cai, Kuo Zhang 0001, Jie Tang 0001
MMM (1)2
2014 A computational cognition model of perception, memory, and judgment
Xiaolan Fu, Lianhong Cai, Ye Liu 0010, Jia Jia 0001, Zhang Yi 0001, Guozhen Zhao, Yong-Jin Liu 0001, Changxu Wu
Sci. China Inf. Sci.4
2014 Grading the Severity of Mispronunciations in CAPT Based on Statistical Analysis and Computational Speech Perception
Jia Jia 0001, Wai-Kim Leung, Yu-Hao Wu, Xiu-Long Zhang, Hao Wang 0077, Lianhong Cai, Helen M. Meng
J. Comput. Sci. Technol.1
2014 Head and facial gestures synthesis using PAD model for an expressive talking avatar
Jia Jia 0001, Zhiyong Wu 0001, Helen M. Meng, Lianhong Cai
Multim. Tools Appl.1
2014 Synthesizing English emphatic speech for multimodal corrective feedback in computer-aided pronunciation training
Zhiyong Wu 0001, Jia Jia 0001, Helen M. Meng, Lianhong Cai
Multim. Tools Appl.3
2013 Interpretable aesthetic features for affective image classification
abstract
Images can not only display contents themselves, but also convey emotions, e.g., excitement, sadness. Affective image classification is useful and hot in many fields such as computer vision and multimedia. Current researches usually consider the relationship model between images and emotions as a black box. They extract the traditional discursive visual features such as SIFT and wavelet textures, and use them directly upon various classification algorithms. However, these visual features are not interpretable, and people cannot know why such a set of features induce a particular emotion. And due to the highly subjective nature of images, the classification accuracies on these visual features are not satisfactory for a long time. We propose the interpretable aesthetic features to describe images inspired by art theories, which are intuitive, discriminative and easily understandable. Affective image classification based on these features can achieve higher accuracy, compared with the state-of-the-art. Specifically, the features can also intuitively explain why an image tends to convey a certain emotion. We also develop an emotion guided image gallery to demonstrate the proposed feature collection.
Xiaohui Wang 0004, Jia Jia 0001, Jiaming Yin, Lianhong Cai
ICIP2
2013 WeCard: a multimodal solution for making personalized electronic greeting cards
abstract
In this demo, we build a practical system, WeCard, to generate personalized multimodal electronic greeting cards based on parametric emotional talking avatar synthesis technologies. Given user-input greeting text and facial image, WeCard intelligently and automatically generate the personalized speech with expressive lip-motion synchronized facial animation. Besides the parametric talking avatar synthesis, WeCard incorporates two key technologies: 1) automatical face mesh generation algorithm based on MPEG-4 FAPs (Facial Animation Parameters) extracted by the face alignment algorithm; 2) emotional audio-visual speech synchronization algorithm based on DBN. More specifically, WeCard merges the users? preferred electronic card scene with emotional talking avatar animation, turning the final content into flash or video file that can be easily shared with friends. By this way, WeCard can help you make your multimodal greetings to be more attractive, beautiful, and sincere.
Huijie Lin, Jia Jia 0001, Hanyu Liao, Lianhong Cai
ACM Multimedia2
2013 Affective image adjustment with a single word
Xiaohui Wang 0004, Jia Jia 0001, Lianhong Cai
Vis. Comput.2
2012 Image Colorization with an Affective Word
Xiaohui Wang 0004, Jia Jia 0001, Hanyu Liao, Lianhong Cai
CVM2
2012 Hierarchical English Emphatic Speech Synthesis Based on HMM with Limited Training Data
abstract
Emphasis is an important form of expressiveness in speech. Hidden Markov model (HMM) based speech synthesis has shown great flexibility in generating expressive speech. This paper proposes a hierarchical model based on HMMs aiming at synthesizing emphatic speech of both high emphasis quality and high naturalness with the limited amount of data. Decision trees (DTs) are constructed with non-emphasis-related questions using both neutral and emphasis corpora. The data in each leaf node of the DTs are classified into 6 emphasis categories according to the emphasis-related questions. The data in the same emphasis category are grouped into one sub-node and are used to train one HMM. As there might be no data of some specific emphasis categories in the leaf nodes of the DTs, a method based on cost calculation is proposed to select a suitable HMM in the same leaf node for predicting parameters. Further a compensation model is proposed to adjust the predicted parameters. Experiments show that the proposed hierarchical model can synthesize emphatic speech with high quality for both naturalness and emphasis, using limited amount of training data. Index Terms: emphatic speech synthesis, hidden Markov model (HMM), hierarchy, compensation model 1.
Zhiyong Wu 0001, Helen M. Meng, Jia Jia 0001, Lianhong Cai
INTERSPEECH4
2012 Can we understand van gogh's mood?: learning to infer affects from images in social networks
abstract
Can we understand van Gogh's mood from his artworks? For many years, people have tried to capture van Gogh's affects from his artworks so as to understand the essential meaning behind the images and catch on why van Gogh created these works. In this paper, we study the problem of inferring affects from images in social networks. In particular, we aim to answer: What are the fundamental features that reflect the affects of the authors in images? How the social network information can be leveraged to help detect these affects? We propose a semi-supervised framework to formulate the problem into a factor graph model. Experiments on 20,000 random-download Flickr images show that our method can achieve a precision of 49% with a recall of 24% on inferring authors'affects into 16 categories. Finally, we demonstrate the effectiveness of the proposed method on automatically understanding van Gogh's Mood from his artworks, and inferring the trend of public affects around special event.
Jia Jia 0001, Sen Wu 0001, Xiaohui Wang 0004, Peiyun Hu, Lianhong Cai, Jie Tang 0001
ACM Multimedia1
2012 Understanding the emotional impact of images
abstract
No abstract available.
Xiaohui Wang 0004, Jia Jia 0001, Peiyun Hu, Sen Wu 0001, Jie Tang 0001, Lianhong Cai
ACM Multimedia2
2012 Comparison of adaptation methods for GMM-SVM based speech emotion recognition
abstract
The required length of the utterance is one of the key factors affecting the performance of automatic emotion recognition. To gain the accuracy rate of emotion distinction, adaptation algorithms that can be manipulated on short utterances are highly essential. Regarding this, this paper compares two classical model adaptation methods, maximum a posteriori (MAP) and maximum likelihood linear regression (MLLR), in GMM-SVM based emotion recognition, and tries to find which method can perform better on different length of the enrollment of the utterances. Experiment results show that MLLR adaptation performs better for very short enrollment utterances (with the length shorter than 2s) while MAP adaptation is more effective for longer utterances.
Jianbo Jiang, Zhiyong Wu 0001, Mingxing Xu, Jia Jia 0001, Lianhong Cai
SLT4
2012 Affective Image Colorization
Xiaohui Wang 0004, Jia Jia 0001, Hanyu Liao, Lianhong Cai
J. Comput. Sci. Technol.2
2011 Emotional Audio-Visual Speech Synthesis Based on PAD
abstract
Audio-visual speech synthesis is the core function for realizing face-to-face human-computer communication. While considerable efforts have been made to enable talking with computer like people, how to integrate the emotional expressions into the audio-visual speech synthesis remains largely a problem. In this paper, we adopt the notion of Pleasure-Displeasure, Arousal-Nonarousal, and Dominance-Submissiveness (PAD) 3-D-emotional space, in which emotions can be described and quantified from three different dimensions. Based on this new definition, we propose a unified model for emotional speech conversion using Boosting-Gaussian mixture model (GMM), as well as a facial expression synthesis model. We further present an emotional audio-visual speech synthesis approach. Specifically, we take the text and the target PAD values as input, and employ the text-to-speech (TTS) engine to first generate the neutral speeches. Then the Boosting-GMM is used to convert the neutral speeches to emotional speeches, and the facial expression is synthesized simultaneously. Finally, the acoustic features of the emotional speech are used to modulate the facial expression in the audio-visual speech. We designed three objective and five subjective experiments to evaluate the performance of each model and the overall approach. Our experimental results on audio-visual emotional speech datasets show that the proposed approach can effectively and efficiently synthesize natural and expressive emotional audio-visual speeches. Analysis on the results also unveil that the mutually reinforcing relationship indeed exists between audio and video information.
Jia Jia 0001, Lianhong Cai
IEEE Trans. Speech Audio Process.1
2010 Facial expression synthesis based on motion patterns learned from face database
abstract
Facial expression is the core function in face-to-face human-computer communication. In order to improve the accuracy and variety of the synthesized facial expressions, we propose a facial expression synthesis approach based on motion patterns learned from face database. We first define a set of partial facial regions: eyebrow, eye-lid, eyeball, upper lip, bottom lip and corner lip. For each region, we use a hierarchical clustering algorithm to learn the motion patterns from the face database. Then the patterns are used for parameterized synthesis of facial expression on the target face models. The experimental results show the effectiveness of the proposed approach.
Jia Jia 0001, Lianhong Cai
ICIP1
2007 Fake Finger Detection Based on Time-Series Fingerprint Image Analysis
Jia Jia 0001, Lianhong Cai
ICIC (1)1
2007 Fingerprint matching based on weighting method and the SVM
Jia Jia 0001, Lianhong Cai, Pinyan Lu, Xuhui Liu
Neurocomputing1