Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Boya Wu

dblp:149/1290 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
2since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Video understanding and tracking · 38% Vision and language · 29% Speech recognition and synthesis · 16%
Databases, data mining, and information retrieval
3 papers
Web and social media mining · 80% Recommender systems · 20%

Topics — the 7 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
long video understanding
0.912025
MLVU: Benchmarking Multi-task Long Video Understanding · CVPR 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
0.912025
MLVU: Benchmarking Multi-task Long Video Understanding · CVPR 2025
Web and social media mining
social influence
0.422016
Social Role-Aware Emotion Contagion in Image Social Networks · AAAI 2016
How Do Your Friends on Social Media Disclose Your Emotions? · AAAI 2014
Natural language and speech › Speech recognition and synthesis › paralinguistic analysis
speech emotion recognition
0.412019
Inferring Emotions From Large-Scale Internet Voice Data · IEEE Trans. Multim. 2019
Web and social media mining › social media analysis
social image analysis
0.312017
Inferring Emotional Tags From Social Images With User Demographics · IEEE Trans. Multim. 2017
Computer vision › Video understanding and tracking
video question answering
0.312025
MLVU: Benchmarking Multi-task Long Video Understanding · CVPR 2025
Computer vision › Image recognition and object detection
image emotion analysis
0.122016
Social Role-Aware Emotion Contagion in Image Social Networks · AAAI 2016
How Do Your Friends on Social Media Disclose Your Emotions? · AAAI 2014

Methods — techniques the papers use, named apart from their topics

f1-score evaluation · 0.9multi-task evaluation · 0.9empirical study · 0.9probabilistic graphical model · 0.5long short-term memory · 0.4latent dirichlet allocation · 0.4deep sparse neural network · 0.4joint modeling · 0.4factor graph model · 0.3demographic modeling · 0.3
YearPublicationVenuePosition
2025 MLVU: Benchmarking Multi-task Long Video Understanding
abstract
The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in video types and evaluation tasks, and the inappropriateness for evaluating LVU performances. To address the above problems, we propose a new benchmark called MLVU (Multitask Long Video Understanding Benchmark) for the comprehensive and in-depth evaluation of LVU. MLVU presents the following critical values: 1) The substantial and flexible extension of video lengths, which enables the benchmark to evaluate LVU performance across a wide range of durations. 2) The inclusion of various video genres, such as movies, surveillance, egocentric videos, and cartoons, reflects the models’ LVU performances in different scenarios. 3) The development of diversified evaluation tasks, which enables a comprehensive examination of MLLMs’ key abilities in long-video understanding. The empirical study with 23 latest MLLMs reveals significant room for improvement in today’s technique, as all existing methods struggle with most of the evaluation tasks and exhibit severe performance degradation when handling longer videos. Additionally, it suggests that factors such as context length, image-understanding ability, and the choice of LLM backbone can play critical roles in future advancements. We anticipate that MLVU will advance the research of LVU by providing a comprehensive and in-depth analysis of MLLMs. The code and dataset can be accessed from https://github.com/JUNJIE99/MLVU.
Junjie Zhou 0001, Bo Zhao 0015, Boya Wu, Zhengyang Liang, Shitao Xiao, Minghao Qin, Yongping Xiong, Tiejun Huang 0001, Zheng Liu 0011
CVPR4
2025 Contrastive learning and prior knowledge-induced feature extraction network for prediction of high-risk recurrence areas in Gliomas
Boya Wu, Jianyun Cao, Yanchun Lv, Guohua Zhao, Ying Zhang 0095, Junguo Bu, Meiyan Huang
Medical Image Anal.1
2019 Inferring Emotions From Large-Scale Internet Voice Data
abstract
As voice dialog applications (VDAs, e.g., Siri,11http://www.apple.com/ios/siri/. Cortana,22http://www.microsoft.com/en-us/mobile/campaign-cortana/. Google Now33http://www.google.com/landing/now/.) are increasing in popularity, inferring emotions from the large-scale internet voice data generated from VDAs can help give a more reasonable and humane response. However, the tremendous amounts of users in large-scale internet voice data lead to a great diversity of users accents and expression patterns. Therefore, the traditional speech emotion recognition methods, which mainly target acted corpora, cannot effectively handle the massive and diverse amount of internet voice data. To address this issue, we carry out a series of observations, find suitable emotion categories for large-scale internet voice data, and verify the indicators of the social attributes (query time, query topic, and users location) and emotion inferring. Based on our observations, two different strategies are employed to solve the problem. First, a deep sparse neural network model that uses acoustic information, textual information, and three indicators (a temporal indicator, descriptive indicator, and geo-social indicator) as the input is proposed. Then, to capture the contextual information, we propose a hybrid emotion inference model that includes long short-term memory to capture the acoustic features and a latent dirichlet allocation to extract text features. Experiments on 93 000 utterances collected from the Sogou Voice Assistant44http://yy.sogou.com. (Chinese Siri) validate the effectiveness of the proposed methodologies. Furthermore, we compare the two methodologies and give their advantages and disadvantages.
Jia Jia 0001, Suping Zhou, Yufeng Yin 0002, Boya Wu, Wei Chen 0071
IEEE Trans. Multim.4
2017 Inferring Emotional Tags From Social Images With User Demographics
abstract
Social images, which are images uploaded and shared on social networks, are used to express users’ emotions. Inferring emotional tags from social images is of great importance; it can benefit many applications, such as image retrieval and recommendation. Whereas previous related research has primarily focused on exploring image visual features, we aim to address this problem by studying whether user demographics make a difference regarding users’ emotional tags of social images. We first consider how to model the emotions of social images. Then, we investigate how user demographics, such as gender, marital status, and occupation, are related to the emotional tags of social images. A partially labeled factor graph model named the demographics factor graph model ( D-FGM ) is proposed to leverage the uncovered patterns. Experiments on a data set collected from the world's largest image sharing website Flickr 1 1 [Online]. Available: http://www.flickr.com/ confirm the accuracy of the proposed model. We also find some interesting phenomena. For example, men and women have different patterns to tag “anger” for social images.
Boya Wu, Jia Jia 0001, Yang Yang 0009, Peijun Zhao, Jie Tang 0001, Qi Tian 0001
IEEE Trans. Multim.1
2016 Social Role-Aware Emotion Contagion in Image Social Networks
abstract
Psychological theories suggest that emotion represents the state of mind and instinctive responses of one’s cognitive system (Cannon 1927). Emotions are a complex state of feeling that results in physical and psychological changes that influence our behavior. In this paper, we study an interesting problem of emotion contagion in social networks. In particular, by employing an image social network (Flickr) as the basis of our study, we try to unveil how users’ emotional statuses influence each other and how users’ positions in the social network affect their influential strength on emotion. We develop a probabilistic framework to formalize the problem into a role-aware contagion model. The model is able to predict users’ emotional statuses based on their historical emotional statuses and social structures. Experiments on a large Flickr dataset show that the proposed model significantly outperforms (+31% in terms of F1-score) several alternative methods in predicting users’ emotional status. We also discover several intriguing phenomena. For example, the probability that a user feels happy is roughly linear to the number of friends who are also happy; but taking a closer look, the happiness probability is superlinear to the number of happy friends who act as opinion leaders (Page et al. 1999) in the network and sublinear in the number of happy friends who span structural holes (Burt 2001). This offers a new opportunity to understand the underlying mechanism of emotional contagion in online social networks.
Yang Yang 0009, Jia Jia 0001, Boya Wu, Jie Tang 0001
AAAI3
2016 Inferring users' emotions for human-mobile voice dialogue applications
abstract
In this paper, we tackle the problem of inferring users' emotions in real-world Voice Dialogue Applications (VDAs, Siri1, Cortana2, etc.). We first conduct an investigation, indicating that besides the text information of users' queries, the acoustic information and query attributes are very important in inferring emotions in VDAs. To integrate the information above, we propose a Hybrid Emotion Inference Model (HEIM), which involves a Latent Dirichlet Allocation (LDA) to extract text features and a Long Short-Term Memory (LSTM) to model the acoustic features. To further improve accuracy, a Recurrent Autoencoder Guided by Query Attributes (RAGQA) which incorporates other emotion-related query attributes is proposed in HEIM to pre-train LSTM. The accuracy of HEIM on a data set collected from Sogou Voice Assistant3(Chinese Siri) containing 93,000 utterances achieves 75.2%, which outperforms state-of-the-art methods for 33.5–38.5%. Specifically, we discover that on average, the acoustic information enhances the performance for 46.6%, while query attributes further enhance the performance for 6.5%.
Boya Wu, Jia Jia 0001, Tao He 0016, Xiaoyuan Yi, Yishuang Ning
ICME1
2015 Understanding the emotions behind social images: Inferring with user demographics
abstract
Understanding the essential emotions behind social images is of vital importance: it can benefit many applications such as image retrieval and personalized recommendation. While previous related research mostly focuses on the image visual features, in this paper, we aim to tackle this problem by “linking inferring with users' demographics”. Specifically, we propose a partially-labeled factor graph model named D-FGM, to predict the emotions embedded in social images not only by the image visual features, but also by the information of users' demographics. We investigate whether users' demographics like gender, marital status and occupation are related to emotions of social images, and then leverage the uncovered patterns into modeling as different factors. Experiments on a data set from the world's largest image sharing website Flickr1 confirm the accuracy of the proposed model. The effectiveness of the users' demographics factors is also verified by the factor contribution analysis, which reveals some interesting behavioral phenomena as well.
Boya Wu, Jia Jia 0001, Yang Yang 0009, Peijun Zhao, Jie Tang 0001
ICME1
2015 Modeling Emotion Influence in Image Social Networks
abstract
We study emotion influence in large image social networks. We focus on users' emotions reflected by images that they have uploaded and social influence that plays a role in changing users' emotions. We first verify the existence of emotion influence in the image networks, and then propose a probabilistic factor graph based emotion influence model to answer the questions of “who influences whom”. Employing a real network from Flickr as the basis in our empirical study, we evaluate the effectiveness of different factors in the proposed model with in-depth data analysis. The learned influence is fundamental for social network analysis and can be applied to many applications. We consider using the influence to help predict users' emotions and our experiments can significantly improve the prediction accuracy (3.0-26.2 percent) over several alternative methods such as Naive Bayesian, SVM (Support Vector Machine) or traditional Graph Model. We further examine the behavior of the emotion influence model, and find that more social interactions correlate with higher emotion influence between two users, and the influence of negative emotions is stronger than positive ones.
Xiaohui Wang 0004, Jia Jia 0001, Jie Tang 0001, Boya Wu, Lianhong Cai, Lexing Xie
IEEE Trans. Affect. Comput.4
2014 How Do Your Friends on Social Media Disclose Your Emotions?
abstract
Extracting emotions from images has attracted much interest, in particular with the rapid development of social networks. The emotional impact is very important for understanding the intrinsic meanings of images. Despite many studies having been done, most existing methods focus on image content, but ignore the emotion of the user who published the image. One interesting question is: How does social effect correlate with the emotion expressed in an image? Specifically, can we leverage friends interactions (e.g., discussions) related to an image to help extract the emotions? In this paper, we formally formalize the problem and propose a novel emotion learning method by jointly modeling images posted by social users and comments added by their friends. One advantage of the model is that it can distinguish those comments that are closely related to the emotion expression for an image from the other irrelevant ones. Experiments on an open Flickr dataset show that the proposed model can significantly improve (+37.4% by F1) the accuracy for inferring user emotions. More interestingly, we found that half of the improvements are due to interactions between 1.0% of the closest friends.
Yang Yang 0009, Jia Jia 0001, Shumei Zhang, Boya Wu, Qicong Chen, Juan-Zi Li, Chunxiao Xing, Jie Tang 0001
AAAI4