Meng Luo 0010

dblp:16/3121-10 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-2274-5719ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Information extraction and text analysis · 27% Vision and language · 23% Language models and text generation · 11%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 91% Collaborative and social computing · 9%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 27 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › sentiment analysis
aspect-based sentiment analysis
1.622025
The ACM Multimedia 2025 Grand Challenge of Multimodal Conversational Aspect-based Sentiment Analysis · ACM Multimedia 2025
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis · ACM Multimedia 2024
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
multimodal aspect-based sentiment analysis
1.622025
The ACM Multimedia 2025 Grand Challenge of Multimodal Conversational Aspect-based Sentiment Analysis · ACM Multimedia 2025
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis · ACM Multimedia 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
1.432025
On Path to Multimodal Generalist: General-Level and General-Bench · ICML 2025
CogMAEC'25: The 1st Workshop on Cognition-oriented Multimodal Affective and Empathetic Computing · ACM Multimedia 2025
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis · ACM Multimedia 2024
Natural language and speech › Information extraction and text analysis › sentiment analysis
sentiment flipping analysis
1.022025
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis · ACM Multimedia 2024
The ACM Multimedia 2025 Grand Challenge of Multimodal Conversational Aspect-based Sentiment Analysis · ACM Multimedia 2025
Computer vision › Vision and language › video grounding
spatio-temporal grounding
1.012026
Dr.V : A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-Grained Spatial-Temporal Grounding · Int. J. Comput. Vis. 2026
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.912025
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025
Natural language and speech › Question answering and dialogue systems › dialogue generation
empathetic response generation
0.912025
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark · WWW 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve Framework · ACL (1) 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
logic-based reasoning
0.912025
Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve Framework · ACL (1) 2025
Natural language and speech › Question answering and dialogue systems
multimodal dialogue system
0.912025
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark · WWW 2025
Computer vision › Vision and language › multimodal dialogue
multimodal dialogue understanding
0.912025
The ACM Multimedia 2025 Grand Challenge of Multimodal Conversational Aspect-based Sentiment Analysis · ACM Multimedia 2025
Computer vision › Video understanding and tracking › spatio-temporal understanding
spatio-temporal video understanding
0.912025
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025
Computer vision › Vision and language
video-language model
0.912025
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025
Human-AI interaction
affective computing
0.912025
CogMAEC'25: The 1st Workshop on Cognition-oriented Multimodal Affective and Empathetic Computing · ACM Multimedia 2025
Human-AI interaction › affective computing
multimodal affective computing
0.912025
CogMAEC'25: The 1st Workshop on Cognition-oriented Multimodal Affective and Empathetic Computing · ACM Multimedia 2025
Program synthesis and code generation
code generation with language models
0.912025
EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning · ICML 2025
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.812024
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis · ACM Multimedia 2024
Machine learning › Learning paradigms
class imbalance
0.712023
Towards Class-Balanced Privacy Preserving Heterogeneous Model Aggregation · IEEE Trans. Dependable Secur. Comput. 2023
Machine learning › Trustworthy machine learning
fairness
0.712023
Towards Class-Balanced Privacy Preserving Heterogeneous Model Aggregation · IEEE Trans. Dependable Secur. Comput. 2023
Machine learning › Efficient and distributed learning
federated and distributed training
0.712023
Towards Class-Balanced Privacy Preserving Heterogeneous Model Aggregation · IEEE Trans. Dependable Secur. Comput. 2023
Machine learning › Efficient and distributed learning › federated learning › model aggregation
heterogeneous model aggregation
0.712023
Towards Class-Balanced Privacy Preserving Heterogeneous Model Aggregation · IEEE Trans. Dependable Secur. Comput. 2023
Machine learning › Trustworthy machine learning
interpretability
0.312026
Dr.V : A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-Grained Spatial-Temporal Grounding · Int. J. Comput. Vis. 2026
Computer vision › Vision and language
affective reasoning
0.312025
CogMAEC'25: The 1st Workshop on Cognition-oriented Multimodal Affective and Empathetic Computing · ACM Multimedia 2025
Natural language and speech › Language models and text generation › alignment
preference alignment
0.312025
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025
Collaborative and social computing › computer-mediated communication › virtual communication
avatar-mediated communication
0.312025
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark · WWW 2025
Program synthesis and code generation › code generation with language models
fine-tuning for code generation
0.312025
EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning · ICML 2025
Privacy and data protection
privacy-preserving machine learning
0.212023
Towards Class-Balanced Privacy Preserving Heterogeneous Model Aggregation · IEEE Trans. Dependable Secur. Comput. 2023

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 2.5chain-of-empathetic reasoning · 1.7hierarchical perception-temporal-cognition framework · 1.0fine-grained grounding · 1.0panoptic sentiment sextuple extraction · 0.9large language model fine-tuning · 0.9hierarchical spatial-temporal modeling · 0.9execution-based code selection · 0.9direct preference optimization · 0.9benchmark construction · 0.9paraphrase-based verification · 0.8chain-of-sentiment reasoning · 0.8response-based aggregation · 0.7knowledge distillation · 0.7
YearPublicationVenuePosition
2026 Dr.V : A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-Grained Spatial-Temporal Grounding
Meng Luo 0010, Shengqiong Wu, Liqiang Jing, Tianjie Ju, Jinxiang Lai, Tianlong Wu, Xinya Du, Siyuan Yan, Jiebo Luo 0001, William Yang Wang, Hao Fei 0001, Mong-Li Lee, Wynne Hsu
Int. J. Comput. Vis.1
2025 Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve Framework
abstract
Jundong Xu, Hao Fei, Meng Luo, Qian Liu, Liangming Pan, William Yang Wang, Preslav Nakov, Mong-Li Lee, Wynne Hsu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jundong Xu, Hao Fei 0001, Meng Luo 0010, Qian Liu 0012, Liangming Pan, William Yang Wang, Preslav Nakov, Mong-Li Lee, Wynne Hsu
ACL (1)3
2025 On Path to Multimodal Generalist: General-Level and General-Bench
abstract
The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of language-based LLMs. Unlike their specialist predecessors, existing MLLMs are evolving towards a Multimodal Generalist paradigm. Initially limited to understanding multiple modalities, these models have advanced to not only comprehend but also generate across modalities. Their capabilities have expanded from coarse-grained to fine-grained multimodal understanding and from supporting singular modalities to accommodating a wide array of or even arbitrary modalities. To assess the capabilities of various MLLMs, a diverse array of benchmark test sets has been proposed. This leads to a critical question: Can we simply assume that higher performance across tasks indicates a stronger MLLM capability, bringing us closer to human-level AI? We argue that the answer is not as straightforward as it seems. In this project, we introduce an evaluation framework to delineate the capabilities and behaviors of current multimodal generalists. This framework, named General-Level, establishes 5-scale levels of MLLM performance and generality, offering a methodology to compare MLLMs and gauge the progress of existing systems towards more robust multimodal generalists and, ultimately, towards AGI (Artificial General Intelligence). Central to our framework is the use of Synergy as the evaluative criterion, categorizing capabilities based on whether MLLMs preserve synergy across comprehension and generation, as well as across multimodal interactions. To evaluate the comprehensive abilities of various generalists, we present a massive multimodal benchmark, General-Bench, which encompasses a broader spectrum of skills, modalities, formats, and capabilities, including over 700 tasks and 325,800 instances. The evaluation results that involve over 100 existing state-of-the-art MLLMs uncover the capability rankings of generalists, highlighting the challenges in reaching genuine AI. We expect this project to pave the way for future research on next-generation multimodal foundation models, providing a robust infrastructure to accelerate the realization of AGI. Project Page: https://generalist.top/, Leaderboard: https://generalist.top/leaderboard/, Benchmark: https://huggingface.co/General-Level/.
Hao Fei 0001, Yuan Zhou 0016, Juncheng Li 0006, Xiangtai Li, Qingshan Xu 0001, Bobo Li 0001, Shengqiong Wu, Yaoting Wang, Junbao Zhou, Jiahao Meng, Liangtao Shi, Minghe Gao, Daoan Zhang, Zhiqi Ge, Siliang Tang, Kaihang Pan, Yaobo Ye, Haobo Yuan, Tao Zhang 0042, Weiming Wu, Tianjie Ju, Zixiang Meng, Shilin Xu 0001, Liyu Jia, Meng Luo 0010, Jiebo Luo 0001, Tat-Seng Chua, Shuicheng Yan, Hanwang Zhang
ICML28
2025 EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning
abstract
As large language models (LLMs) play an increasingly important role in code generation, enhancing both correctness and efficiency has become crucial. Current methods primarily focus on correctness, often overlooking efficiency. To address this gap, we introduce SWIFTCODE to improve both aspects by fine-tuning LLMs on a high-quality dataset comprising correct and efficient code samples. Our methodology involves leveraging multiple LLMs to generate diverse candidate code solutions for various tasks across different programming languages. We then evaluate these solutions by directly measuring their execution time and memory usage through local execution. The code solution with the lowest execution time and memory consumption is selected as the final output for each task. Experimental results demonstrate significant improvements when fine-tuning with SWIFTCODE. For instance, Qwen2.5-Coder-7B-Instruct’s pass@1 score increases from 44.8% to 57.7%, while the average execution time for correct tasks decreases by 48.4%. SWIFTCODE offers a scalable and effective solution for advancing AI-driven code generation, benefiting both software development and computational problem-solving.
Dong Huang 0005, Guangtao Zeng, Jianbo Dai, Meng Luo 0010, Han Weng, Yuhao Qing, Heming Cui, Zhijiang Guo, Jie Zhang 0050
ICML4
2025 VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
Haojian Huang, Shengqiong Wu, Meng Luo 0010, Jinlan Fu, Xinya Du, Hanwang Zhang, Hao Fei 0001
ICML4
2025 CogMAEC'25: The 1st Workshop on Cognition-oriented Multimodal Affective and Empathetic Computing
abstract
The 1st Workshop on Cognition-oriented Multimodal Affective and Empathetic Computing (CogMAEC) was held at ACM Multimedia 2025. It focused on moving emotional AI beyond basic recognition toward deeper cognitive understanding. While traditional multimodal affective computing has emphasized simple emotion detection, the rise of multimodal large language models (MLLMs) has spurred interest in modeling how emotions emerge and evolve in context. The workshop gathered researchers on emotion reasoning, multimodal understanding, and human-computer empathy, exploring how machines can not only recognize emotions but also explain their causes and simulate human-like affective reasoning. The program featured invited talks, oral presentations, and posters spanning perception, interaction, causal modeling, and cognitive grounding. CogMAEC provided a platform to connect researchers across disciplines and foster future work on cognitively aware affective computing. Materials are available at https://CogMAEC.github.io/MM2025.
Hao Fei 0001, Bobo Li 0001, Meng Luo 0010, Qian Liu 0012, Lizi Liao, Fei Li 0021, Min Zhang 0005, Björn W. Schuller, Mong-Li Lee, Erik Cambria
ACM Multimedia3
2025 The ACM Multimedia 2025 Grand Challenge of Multimodal Conversational Aspect-based Sentiment Analysis
abstract
Understanding fine-grained sentiment dynamics in human conversations is a central goal for next-generation artificial intelligence, especially in scenarios where interactions are rich in both modalities and context. To advance research in this area, we organize the Multimodal Conversational Aspect-based Sentiment Analysis (MCABSA) challenge to the community of aspect-based sentiment analysis. The MCABSA challenge introduces two novel subtasks: 1) Panoptic Sentiment Sextuple Extraction, panoramically recognizing holder, target, aspect, opinion, sentiment, and rationale from multi-turn, multi-party multimodal dialogue; and 2) Sentiment Flipping Analysis, detecting the dynamic sentiment transformation throughout the conversation along with the causal reasons. To support these tasks, we present the PanoSent dataset, a high-quality, large-scale benchmark featuring multi-turn, multi-party dialogues annotated with both explicit and implicit sentiment elements across text, image, audio, and video modalities. PanoSent offers extensive real-world scenario coverage, providing a comprehensive testbed for multimodal conversational sentiment analysis. The challenge has attracted widespread participation from both academia and industry, with over 30 teams registered and more than 100 successful submissions. In this paper, we introduce the task, dataset, and evaluation settings, summarize the systems of the top teams, and discuss the findings of the participants. Further details of the challenge can be found at https://panosent.github.io/MM25-challenge.
Meng Luo 0010, Hao Fei 0001, Bobo Li 0001, Shengqiong Wu, Qian Liu 0012, Soujanya Poria, Erik Cambria, Mong-Li Lee, Wynne Hsu
ACM Multimedia1
2025 Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
abstract
Empathetic Response Generation (ERG) is one of the key tasks of the affective computing area, which aims to produce emotionally nuanced and compassionate responses to user's queries. However, existing ERG research is predominantly confined to the singleton text modality, limiting its effectiveness since human emotions are inherently conveyed through multiple modalities. To combat this, we introduce an avatar-based Multimodal ERG (MERG) task, entailing rich text, speech, and facial vision information. We first present a large-scale high-quality benchmark dataset, AvaMERG, which extends traditional text ERG by incorporating authentic human speech audio and dynamic talking-face avatar videos, encompassing a diverse range of avatar profiles and broadly covering various topics of real-world scenarios. Further, we deliberately tailor a system, named Empatheia, for MERG. Built upon a Multimodal Large Language Model (MLLM) with multimodal encoder, speech and avatar generators, Empatheia performs end-to-end MERG, with Chain-of-Empathetic reasoning mechanism integrated for enhanced empathy understanding and reasoning.Finally, we devise a list of empathetic-enhanced tuning strategies, strengthening the capabilities of emotional accuracy and content, avatar-profile consistency across modalities. Experimental results on AvaMERG data demonstrate that Empatheia consistently shows superior performance than baseline methods on both textual ERG and MERG. All data and code are open at https://AvaMERG.github.io/.
Han Zhang 0035, Zixiang Meng, Meng Luo 0010, Hong Han 0001, Lizi Liao, Erik Cambria, Hao Fei 0001
WWW3
2024 PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
abstract
While existing Aspect-based Sentiment Analysis (ABSA) has received extensive effort and advancement, there are still gaps in defining a more holistic research target seamlessly integrating multimodality, conversation context, fine-granularity, and also covering the changing sentiment dynamics as well as cognitive causal rationales. This paper bridges the gaps by introducing a multimodal conversational ABSA, where two novel subtasks are proposed: 1) Panoptic Sentiment Sextuple Extraction, panoramically recognizing holder, target, aspect, opinion, sentiment, rationale from multi-turn multi-party multimodal dialogue. 2) Sentiment Flipping Analysis, detecting the dynamic sentiment transformation throughout the conversation with the causal reasons. To benchmark the tasks, we construct PanoSent, a dataset annotated both manually and automatically, featuring high quality, large scale, multimodality, multilingualism, multi-scenarios, and covering both implicit&explicit sentiment elements. To effectively address the tasks, we devise a novel Chain-of-Sentiment reasoning framework, together with a novel multimodal large language model (namely Sentica) and a paraphrase-based verification mechanism. Extensive evaluations demonstrate the superiority of our methods over strong baselines, validating the efficacy of all our proposed methods. The work is expected to open up a new era for the ABSA community, and thus all our codes and data are open at https://PanoSent.github.io/.
Meng Luo 0010, Hao Fei 0001, Bobo Li 0001, Shengqiong Wu, Qian Liu 0012, Soujanya Poria, Erik Cambria, Mong-Li Lee, Wynne Hsu
ACM Multimedia1
2023 Towards Class-Balanced Privacy Preserving Heterogeneous Model Aggregation
abstract
Heterogeneous model aggregation (HMA) is an effective paradigm that integrates on-device trained models heterogeneous in architecture and target task into a comprehensive model. Recent works adopt knowledge distillation to amalgamate the knowledge of learned features and predictions from heterogeneous on-device models to realize HMA. However, most of them ignore that the disclosure of learned features exposes on-device models to privacy attacks. Moreover, the aggregated model may suffer from the imbalanced supervision caused by the uneven distribution of amalgamated knowledge about each class and show class bias. In this article, to address these issues, we propose a response-based class-balanced heterogeneous model aggregation mechanism, called CBHMA. It can effectively achieve HMA in a privacy-preserving manner and alleviate class bias in the aggregated model. Specifically, CBHMA aggregates on-device models by using only their response information to reduce their privacy leakage risk. To mitigate the impact of imbalanced supervision, CBHMA quantitatively measures the imbalanced supervision level for each class. Based on that, CBHMA customizes fine-grained misclassification costs for each class and utilizes such costs to adjust the importance of each class (more importance to classes with weaker supervision) in the response-based HMA algorithm. Extensive experiments on two real-world datasets demonstrate the effectiveness of CBHMA.
Xiaoyi Pang, Zhibo Wang 0001, Zeqing He, Peng Sun 0003, Meng Luo 0010, Ju Ren 0001, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.5