Yaoting Wang

dblp:258/1290 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Survey of Inductive Reasoning for Large Language Models
abstract
Kedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang, Siyu Yan, Xuecheng Wu, Yinqi Zhang, Qin Chen, Jie Zhou, Liang He, Biqing Qi, Linyang Li, Qipeng Guo, Xiaoming Shi, Wei Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Kedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang, Yinqi Zhang, Qin Chen 0001, Jie Zhou 0015, Liang He 0001, Biqing Qi, Linyang Li, Qipeng Guo, Wayne Zhang 0001
ACL (1)4
2025 AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
abstract
With the rapid advancement of Multi-modal Large Language Models (MLLMs), several diagnostic benchmarks have recently been developed to assess these models' multi-modal reasoning proficiency. However, these benchmarks are restricted to assessing primarily the visual aspect and do not examine the holistic audio-visual (AV) understanding. Moreover, currently, there are no benchmarks that investigate the capabilities of AVLLMs to calibrate their responses when presented with perturbed inputs. To this end, we introduce Audio-Visual Trustworthiness assessment Benchmark (AVTrustBench), comprising 600K samples spanning over 9 meticulously crafted tasks, evaluating the capabilities of AVLLMs across three distinct dimensions: Adversarial attack, Compositional reasoning, and Modality-specific dependency. Using our benchmark we extensively evaluate 13 state-of-the-art AVLLMs. The findings reveal that the majority of existing models fall significantly short of achieving human-like comprehension, offering valuable insights for future research directions. To alleviate the limitations in the existing approaches, we further propose a robust, model-agnostic calibrated audio-visual preference optimization based training strategy CAVPref, obtaining a gain up to 30.19% across all 9 tasks. We will publicly release our code and benchmark to facilitate future research in this direction.
Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta, Yaoting Wang, Mohamed Elhoseiny 0001, Ruohan Gao, Dinesh Manocha
ICCV4
2025 On Path to Multimodal Generalist: General-Level and General-Bench
abstract
The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of language-based LLMs. Unlike their specialist predecessors, existing MLLMs are evolving towards a Multimodal Generalist paradigm. Initially limited to understanding multiple modalities, these models have advanced to not only comprehend but also generate across modalities. Their capabilities have expanded from coarse-grained to fine-grained multimodal understanding and from supporting singular modalities to accommodating a wide array of or even arbitrary modalities. To assess the capabilities of various MLLMs, a diverse array of benchmark test sets has been proposed. This leads to a critical question: Can we simply assume that higher performance across tasks indicates a stronger MLLM capability, bringing us closer to human-level AI? We argue that the answer is not as straightforward as it seems. In this project, we introduce an evaluation framework to delineate the capabilities and behaviors of current multimodal generalists. This framework, named General-Level, establishes 5-scale levels of MLLM performance and generality, offering a methodology to compare MLLMs and gauge the progress of existing systems towards more robust multimodal generalists and, ultimately, towards AGI (Artificial General Intelligence). Central to our framework is the use of Synergy as the evaluative criterion, categorizing capabilities based on whether MLLMs preserve synergy across comprehension and generation, as well as across multimodal interactions. To evaluate the comprehensive abilities of various generalists, we present a massive multimodal benchmark, General-Bench, which encompasses a broader spectrum of skills, modalities, formats, and capabilities, including over 700 tasks and 325,800 instances. The evaluation results that involve over 100 existing state-of-the-art MLLMs uncover the capability rankings of generalists, highlighting the challenges in reaching genuine AI. We expect this project to pave the way for future research on next-generation multimodal foundation models, providing a robust infrastructure to accelerate the realization of AGI. Project Page: https://generalist.top/, Leaderboard: https://generalist.top/leaderboard/, Benchmark: https://huggingface.co/General-Level/.
Hao Fei 0001, Yuan Zhou 0016, Juncheng Li 0006, Xiangtai Li, Qingshan Xu 0001, Bobo Li 0001, Shengqiong Wu, Yaoting Wang, Junbao Zhou, Jiahao Meng, Liangtao Shi, Minghe Gao, Daoan Zhang, Zhiqi Ge, Siliang Tang, Kaihang Pan, Yaobo Ye, Haobo Yuan, Tao Zhang 0042, Weiming Wu, Tianjie Ju, Zixiang Meng, Shilin Xu 0001, Liyu Jia, Meng Luo 0010, Jiebo Luo 0001, Tat-Seng Chua, Shuicheng Yan, Hanwang Zhang
ICML8
2024 Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer
abstract
Never having seen an object and heard its sound simultaneously, can the model still accurately localize its visual position from the input audio? In this work, we concentrate on the Audio-Visual Localization and Segmentation tasks but under the demanding zero-shot and few-shot scenarios. To achieve this goal, different from existing approaches that mostly employ the encoder-fusion-decoder paradigm to decode localization information from the fused audio-visual feature, we introduce the encoder-prompt-decoder paradigm, aiming to better fit the data scarcity and varying data distribution dilemmas with the help of abundant knowledge from pre-trained models. Specifically, we first propose to construct a Semantic-aware Audio Prompt (SAP) to help the visual foundation model focus on sounding objects, meanwhile, the semantic gap between the visual and audio modalities is also encouraged to shrink. Then, we develop a Correlation Adapter (ColA) to keep minimal training efforts as well as maintain adequate knowledge of the visual foundation model. By equipping with these means, extensive experiments demonstrate that this new paradigm outperforms other fusion-based methods in both the unseen class and cross-dataset settings. We hope that our work can further promote the generalization study of Audio-Visual Localization and Segmentation in practical application scenarios. Project page: https://github.com/GeWu-Lab/Generalizable-Audio-Visual-Segmentation
Yaoting Wang, Weisong Liu, Guangyao Li 0001, Jian Ding 0001, Di Hu 0001
AAAI1
2024 Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation
Juncheng Ma, Peiwen Sun, Yaoting Wang, Di Hu 0001
ECCV (73)3
2024 Can Textual Semantics Mitigate Sounding Object Segmentation Preference?
Yaoting Wang, Peiwen Sun, Yuanchao Li, Honggang Zhang 0002, Di Hu 0001
ECCV (74)1
2024 Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
Yaoting Wang, Peiwen Sun, Dongzhan Zhou, Guangyao Li 0001, Honggang Zhang 0002, Di Hu 0001
ECCV (74)1
2024 A deep reinforcement learning based distributed multi-UAV dynamic area coverage algorithm for complex environment
Jian Xiao 0006, Guohui Yuan, Yuxi Xue, Jinhui He, Yaoting Wang, Yuanjiang Zou
Neurocomputing5
2024 Multi-agent cooperative area coverage: A two-stage planning approach based on reinforcement learning
Guohui Yuan, Jian Xiao 0006, Jinhui He, Honyu Jia, Yaoting Wang
Inf. Sci.5
2024 Synchronous composition and semantic line detection based on cross-attention
Qinggang Hou, Yongzhen Ke, Kai Wang 0065, Fan Qin 0001, Yaoting Wang
Multim. Syst.5
2024 Aesthetic feature design and aesthetic quality assessment for group photograph
abstract
Image aesthetics quality assessment has received extensive research in recent years, but there are still few studies on the aesthetic quality evaluation of group photograph of humans. In this work, we designed a set of high-level aesthetic features based on the experience and principles of group photography, including opened-eye, gaze, smile, facial occluded, facial orientation, facial blur, character center. Then we combined them and 83 generic aesthetic features to build two aesthetic assessment models. A large dataset of group photographs - GPD- annotated with the aesthetic score was constructed. The experimental result on the GPD shows that our features perform well for categorizing professional photos and snapshots and predicting the distinction of multiple group photographs of diverse human states under the same scene.The classification accuracy reached 70.97%, the discrimination metric we proposed reached 1.368, which was higher than the negative discrimination value of other methods.
Yaoting Wang, Yongzhen Ke, Kai Wang 0065, Cuijiao Zhang, Fan Qin 0001
Multim. Tools Appl.1
2024 A Model Learning Based Multiagent Flocking Collaborative Control Method for Stochastic Communication Environment
abstract
Improving the performance of flocking control policies in practical scenarios is of great value in promoting the practical application of multiagent flocking collaborative control algorithms. In this article, concerning the practicality of flocking algorithms in stochastic communication environments, we propose a model learning based multiagent flocking control algorithm. First, an agent motion model construction method based on sequential attention mechanisms is proposed to provide a more realistic agent motion model for environmental interaction. Considering the cooperation and equivalence of agents in the flocking task, a multiagent cooperative soft actor–critic (MACSAC) algorithm is proposed to optimize the control policy model. Then, a digital learning system for multiagent flocking collaborative control is constructed by combining the learned motion model with the MACSAC algorithm. Finally, we design a behavior reasoning (BR) model based on the prior control policy, and introduce the model into the MACSAC algorithm to infer the motion state of noncommunicating adjacent agents, which solves the problem of poor control policy caused by the information loss of observation state in stochastic communication environments. The experimental results indicate that the constructed digital learning system can effectively simulate the policy learning of multiagent flocking in actual environmental scenarios, and demonstrate that the designed BR model can effectively improve the performance of the MACSAC-based multiagent flocking collaborative control algorithm in stochastic communication environments.
Jian Xiao 0006, Chongjun Huang, Guohui Yuan, Yaoting Wang, Honyu Jia
IEEE Trans. Ind. Informatics4
2023 Spatial-invariant convolutional neural network for photographic composition prediction and automatic correction
Yaoting Wang, Yongzhen Ke, Kai Wang 0065
J. Vis. Commun. Image Represent.1
2022 A composition-oriented aesthetic view recommendation network supervised by the simplified golden ratio theory
Yaoting Wang, Yongzhen Ke, Kai Wang 0065, Fan Qin 0001
Expert Syst. Appl.1