Yifei Yin

dblp:72/7108 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 34% 3D vision · 33% Representation and self-supervised learning · 17%
Human-computer interaction and pervasive computing
2 papers
Interaction techniques and input · 44% Collaborative and social computing · 33% Haptics and multimodal interaction · 23%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked prediction
1.012026
Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs · AAAI 2026
Computer vision › Vision and language
multimodal reasoning
1.012026
Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs · AAAI 2026
Computer vision › Vision and language
vision-language model
1.012026
Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs · AAAI 2026
Interaction techniques and input › mobile interaction › wearable device interaction
smartwatch interaction
0.812024
EmoWear: Exploring Emotional Teasers for Voice Message Interaction on Smartwatches · CHI 2024
Computer vision › 3D vision › 3d scene understanding
3d instance segmentation
0.712023
Hi4D: 4D Instance Segmentation of Close Human Interaction · CVPR 2023
Computer vision › Segmentation and scene understanding › instance segmentation
human instance segmentation
0.712023
Hi4D: 4D Instance Segmentation of Close Human Interaction · CVPR 2023
Computer vision › 3D vision
human mesh recovery
0.712023
Hi4D: 4D Instance Segmentation of Close Human Interaction · CVPR 2023
Computer vision › 3D vision
multi-person interaction
0.712023
Hi4D: 4D Instance Segmentation of Close Human Interaction · CVPR 2023
Audio and music processing
emotion recognition
0.612022
Multi-Classifier Interactive Learning for Ambiguous Speech Emotion Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2022
Audio and music processing › emotion recognition
speech emotion recognition
0.612022
Multi-Classifier Interactive Learning for Ambiguous Speech Emotion Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.312026
Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs · AAAI 2026
Haptics and multimodal interaction › haptic feedback
vibrotactile feedback
0.212022
VibEmoji: Exploring User-authoring Multi-modal Emoticons in Social Communication · CHI 2022

Methods — techniques the papers use, named apart from their topics

reinforcement fine-tuning · 1.0prior sampling · 1.0semantic processing · 0.8acoustic processing · 0.8neural implicit avatars · 0.7alternating optimization · 0.7recommendation · 0.6multi-classifier interactive learning · 0.6label distribution learning · 0.6field study · 0.6
YearPublicationVenuePosition
2026 Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs
abstract
Recent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet, a significant gap persists in their adaptation to real-world multimodal scenarios, most notably, vision-language tasks, due to a heavy focus on single-modal language settings. While efforts to transplant reinforcement learning techniques from NLP to Visual Language Models (VLMs) have emerged, these approaches often remain confined to perception-centric tasks or reduce images to textual summaries, failing to fully exploit visual context and commonsense knowledge, ultimately constraining the generalization of reasoning capabilities across diverse multimodal environments. To address this limitation, we introduce a novel fine-tuning task, Masked Prediction via Context and Commonsense (MPCC), which forces models to integrate visual context and commonsense reasoning by reconstructing semantically meaningful content from occluded images, thereby laying the foundation for generalized reasoning. To systematically evaluate the model’s performance in generalized reasoning, we developed a specialized evaluation benchmark, MPCC-Eval, and employed various fine-tuning strategies to guide reasoning. Among these, we introduced an innovative training method, Reinforcement Fine-Tuning with Prior Sampling, which not only enhances model performance but also improves its generalized reasoning capabilities in out-of-distribution (OOD) and cross-task scenarios.
Jiaao Yu 0001, Shenwei Li, Mingjie Han, Yifei Yin, Wenzheng Song, Chenghao Jia, Man Lan
AAAI4
2024 EmoWear: Exploring Emotional Teasers for Voice Message Interaction on Smartwatches
abstract
Voice messages, by nature, prevent users from gauging the emotional tone without fully diving into the audio content. This hinders the shared emotional experience at the pre-retrieval stage. Research scarcely explored “Emotional Teasers”—pre-retrieval cues offering a glimpse into an awaiting message’s emotional tone without disclosing its content. We introduce EmoWear, a smartwatch voice messaging system enabling users to apply 30 animation teasers on message bubbles to reflect emotions. EmoWear eases senders’ choice by prioritizing emotions based on semantic and acoustic processing. EmoWear was evaluated in comparison with a mirroring system using color-coded message bubbles as emotional cues (N=24). Results showed EmoWear significantly enhanced emotional communication experience in both receiving and sending messages. The animated teasers were considered intuitive and valued for diverse expressions. Desirable interaction qualities and practical implications are distilled for future design. We thereby contribute both a novel system and empirical knowledge concerning emotional teasers for voice messaging.
Pengcheng An, Jiawen Stefanie Zhu, Yifei Yin, Qingyuan Ma, Che Yan, Linghao Du, Jian Zhao 0010
CHI4
2023 Hi4D: 4D Instance Segmentation of Close Human Interaction
abstract
We propose Hi4D, a method and dataset for the automatic analysis of physically close human-human interaction under prolonged contact. Robustly disentangling several in-contact subjects is a challenging task due to occlusions and complex shapes. Hence, existing multi-view systems typically fuse 3D surfaces of close subjects into a single, connected mesh. To address this issue we leverage i) individually fitted neural implicit avatars; ii) an alternating optimization scheme that refines pose and surface through periods of close proximity; and iii) thus segment the fused raw scans into individual instances. From these instances we compile Hi4D dataset of 4D textured scans of 20 subject pairs, 100 sequences, and a total of more than 11 K frames. Hi4D contains rich interaction-centric annotations in 2D and 3D alongside accurately registered parametric body models. We define varied human pose and shape estimation tasks on this dataset and provide results from state-of-the-art methods on these benchmarks. Hi4D dataset can be found at https://ait.ethz.ch/Hi4D.
Yifei Yin, Manuel Kaufmann, Juan Jose Zarate, Jie Song 0006, Otmar Hilliges
CVPR1
2022 VibEmoji: Exploring User-authoring Multi-modal Emoticons in Social Communication
abstract
Emoticons are indispensable in online communications. With users’ growing needs for more customized and expressive emoticons, recent messaging applications begin to support (limited) multi-modal emoticons:, enhancing emoticons with animations or vibrotactile feedback. However, little empirical knowledge has been accumulated concerning how people create, share and experience multi-modal emoticons in everyday communication, and how to better support them through design. To tackle this, we developed VibEmoji, a user-authoring multi-modal emoticon interface for mobile messaging. Extending existing designs, VibEmoji grants users greater flexibility to combine various emoticons, vibrations, and animations on-the-fly, and offers non-aggressive recommendations based on these components’ emotional relevance. Using VibEmoji as a probe, we conducted a four-week field study with 20 participants, to gain new understandings from in-the-wild usage and experience, and extract implications for design. We thereby contribute to both a novel system and various insights for supporting users’ creation and communication of multi-modal emoticons.
Pengcheng An, Ziqi Zhou 0003, Qing Liu 0026, Yifei Yin, Linghao Du, Da-Yuan Huang, Jian Zhao 0010
CHI4
2022 A joint framework for mining discriminative and frequent visual representation
Ying Zhou 0021, Xuefeng Liang, Zhihui Liang, Yu Gu 0015, Yifei Yin
Neurocomputing7
2022 Multi-Classifier Interactive Learning for Ambiguous Speech Emotion Recognition
abstract
In recent years, speech emotion recognition technology is of great significance in widespread applications such as call centers, social robots and health care. Thus, the speech emotion recognition has been attracted much attention in both industry and academic. Since emotions existing in an entire utterance may have varied probabilities, speech emotion is likely to be ambiguous, which poses great challenges to recognition tasks. However, previous studies commonly assigned a single-label or multi-label to each utterance in certain. Therefore, their algorithms result in low accuracies because of the inappropriate representation. Inspired by the optimally interacting theory, we address the ambiguous speech emotions by proposing a novel multi-classifier interactive learning (MCIL) method. In MCIL, multiple different classifiers first mimic several individuals, who have inconsistent cognitions of ambiguous emotions, and construct new ambiguous labels (the emotion probability distribution). Then, they are retrained with the new labels to interact with their cognitions. This procedure enables each classifier to learn better representations of ambiguous data from others, and further improves the recognition ability. The experiments on three benchmark corpora (MAS, IEMOCAP, and FAU-AIBO) demonstrate that MCIL does not only improve each classifier’s performance, but also raises their recognition consistency from moderate to substantial.
Ying Zhou 0021, Xuefeng Liang, Yu Gu 0015, Yifei Yin, Longshan Yao
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Progressive Co-Teaching for Ambiguous Speech Emotion Recognition
abstract
Speech emotion recognition is a challenging task due to the ambiguity of emotion, which makes it difficult to learn the features of emotion data using machine learning algorithms. However, previous studies conventionally ignore the ambiguity of emotion and treat the emotion data as the same difficulty level, which results in low recognition accuracy. Motivated by human and animal learning studies, we propose a novel method named Progressive Co-teaching (PCT) to learn speech emotion features from simple to difficult. PCT method automatically identifies the difficulty level of data by itself using loss values, and then each network exchanges easy instances with small loss to peer network for early training. The rest instances with large loss are added gradually for later training. The experiment results demonstrate that our method achieves an improvement of 3.8% and 1.27% on MAS and IEMOCAP database than the state-of-the-arts, respectively.
Yifei Yin, Yu Gu 0015, Longshan Yao, Ying Zhou 0021, Xuefeng Liang
ICASSP1
2007 antiCODE: a natural sense-antisense transcripts database
abstract
BACKGROUND: Natural antisense transcripts (NATs) are endogenous RNA molecules that exhibit partial or complete complementarity to other RNAs, and that may contribute to the regulation of molecular functions at various levels. In recent years, large-scale NAT screens in several model organisms have produced much data, but there is no database to assemble all these data. AntiCODE intends to function as an integrated NAT database for this purpose. RESULTS: This release of antiCODE contains more than 30,000 non-redundant natural sense-antisense transcript pairs from 12 eukaryotic model organisms. In order to provide an integrated NAT research platform, efficient browser, search and Blast functions have been included to enable users to easily access information through parameters such as species, accession number, overlapping patterns, coding potential etc. In addition to the collected information, antiCODE also introduces a simple classification system to facilitate the study of natural antisense transcripts. CONCLUSION: Though a few similar databases also dealing with NATs have appeared lately, antiCODE is the most comprehensive among these, comprising almost all currently detected NAT pairs.
Yifei Yin, Yi Zhao 0013, Changning Liu, Shuguang Chen, Runsheng Chen
BMC Bioinform.1