Qilang Ye

dblp:372/0810 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0004-6907-702XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement
abstract
Understanding 3D scene-level affordances from natural language instructions is essential for enabling embodied agents to interact meaningfully in complex environments. However, this task remains challenging due to the need for semantic reasoning and spatial grounding. Existing methods mainly focus on object-level affordances or merely lift 2D predictions to 3D, neglecting rich geometric structure information in point clouds and incurring high computational costs. To address these limitations, we introduce Task-Aware 3D Scene-level Affordance segmentation (TASA), a novel geometry-optimized framework that jointly leverages 2D semantic cues and 3D geometric reasoning in a coarse-to-fine manner. To improve the affordance detection efficiency, TASA features a task-aware 2D affordance detection module to identify manipulable points from language and visual inputs, guiding the selection of task-relevant views. To fully exploit 3D geometric information, a 3D affordance refinement module is proposed to integrate 2D semantic priors with local 3D geometry, resulting in accurate and spatially coherent 3D affordance masks. Experiments on SceneFun3D demonstrate that TASA significantly outperforms the baselines in both accuracy and efficiency in scene-level affordance segmentation.
Qilang Ye, Yu Zhou 0015
AAAI3
2026 SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
abstract
Large Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action classification and description. However, when treating LLM as a recognizer, two questions arise: 1) How can LLMs understand the skeleton? 2) How can LLMs distinguish among actions? To address these problems, we introduce a novel paradigm named learning Skeleton representation with visual-motion knowledge for Action Recognition (SUGAR). In our pipeline, we first utilize off-the-shelf large-scale video models as a knowledge base to generate visual, motion information related to actions. Then, we propose to supervise skeleton learning through this prior knowledge to yield discrete representations. Finally, we use the LLM with untouched pre-training weights to understand these representations and generate the desired action targets and descriptions. Notably, we present a Temporal Query Projection (TQP) module to continuously model the skeleton signals with long sequences. Experiments on several skeleton-based action classification benchmarks demonstrate the efficacy of our SUGAR. Moreover, experiments on zero-shot scenarios show that SUGAR is more versatile than linear-based methods.
Qilang Ye, Yu Zhou 0015, Jie Zhang 0081, Xuanming Guo, Mingkui Tan, Weicheng Xie 0001, Yue Sun 0001, Tao Tan 0002, Xiaochen Yuan, Ghada Khoriba, Zitong Yu
AAAI1
2026 When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?
abstract
Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, which simulates an “Audio-Visual Confusion” scene by modifying the corresponding sound of an object in the video, e.g., mute the sounding object and ask MLLMs “Is there a/an {muted-object} sound”. Experimental results reveal that MLLMs, such as Qwen2.5-Omni and Gemini 2.5, struggle to discriminate non-existent audio due to visually dominated reasoning. Motivated by this observation, we introduce RL-CoMM, a Reinforcement Learning-based Collaborative Multi-MLLM that is built upon the Qwen2.5-Omni foundation. RL-CoMM includes two stages: 1) To alleviate visually dominated ambiguities, we introduce an external model, a Large Audio Language Model (LALM), as the reference model to generate audio-only reasoning. Then, we design a Step-wise Reasoning Reward function that enables MLLMs to self-improve audio-visual reasoning with the audio-only reference. 2) To ensure an accurate answer prediction, we introduce Answer-centered Confidence Optimization to reduce the uncertainty of potential heterogeneous reasoning differences. Extensive experiments on audio-visual question answering and audio-visual hallucination show that RL-CoMM improves accuracy by 10~30% over the baseline model with limited training data.
Qilang Ye, Jie Zhang 0081, Zitong Yu, Yu Zhou 0015
AAAI1
2026 Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification
abstract
Recent Audio-Visual Question Answering (AVQA) methods have advanced significantly.However, most AVQA methods lack effective mechanisms for handling missing modalities, suffering from severe performance degradation in real-world scenarios with data interruptions.Furthermore, prevailing methods for handling missing modalities predominantly rely on generative imputation to synthesize missing features.While partially effective, these methods tend to capture inter-modal commonalities but struggle to acquire unique, modalityspecific knowledge within the missing data, leading to hallucinations and compromised reasoning accuracy.To tackle these challenges, we propose R 2 ScP, a novel framework that shifts the paradigm of missing modality handling from traditional generative imputation to retrieval-based recovery.Specifically, we leverage cross-modal retrieval via unified semantic embeddings to acquire missing domain-specific knowledge.To maximize semantic restoration, we introduce a context-aware adaptive purification mechanism that eliminates latent semantic noise within the retrieved data.Additionally, we employ a two-stage training strategy to explicitly model the semantic relationships between knowledge from different sources.Extensive experiments demonstrate that R 2 ScP significantly improves AVQA and enhances robustness in modal-incomplete scenarios.1
Jiayu Zhang 0002, Shuo Ye, Qilang Ye, Zihan Song 0009, Jiajian Huang, Zitong Yu
ACL (1)3
2026 EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
abstract
Abstract Facial expression recognition (FER) has emerged as an important research topic in recent years. However, current FER paradigms face challenges in generalization, lack semantic information aligned with natural language, and struggle to process both images and videos within a unified framework. Multimodal Large Language Models (MLLMs) have recently achieved success, offering advantages in addressing these issues and potentially overcoming the limitations of current FER paradigms. Nonetheless, directly applying pre-trained MLLMs to FER remains challenging due to insufficient instruction datasets and the inability of vision encoders to extract fine-grained facial information. Our zero-shot evaluations of existing open-source MLLMs on FER reveal a significant performance gap compared to GPT-4V and state-of-the-art supervised methods. In this paper, we aim to enhance MLLMs’ capabilities in understanding facial expressions. We first introduce a facial expression recognition instruction dataset ( FERID ), which has 376k category instructions and 339k conversational instructions. We then propose a novel MLLM, named EMO-LLaMA , which incorporates facial priors from a pretrained facial analysis network to enhance its understanding of human facial information. Specifically, we design a Face Info Mining module to extract both global and local facial information. Furthermore, we utilize a handcrafted prompt to introduce age-gender-race attributes, considering the emotional differences across diverse human groups. Extensive experiments show that EMO-LLaMA achieves results comparable to or competitive with SOTA on both static and dynamic FER datasets. The instruction dataset and code will be available at https://github.com/xxtars/EMO-LLaMA .
Bohao Xing, Zitong Yu, Xin Liu 0012, Kaishen Yuan, Qilang Ye, Weicheng Xie 0001, Huanjing Yue, Heikki Kälviäinen
Int. J. Comput. Vis.5
2025 CAT+: Investigating and Enhancing Audio-Visual Understanding in Large Language Models
abstract
Multimodal Large Language Models (MLLMs) have gained significant attention due to their rich internal implicit knowledge for cross-modal learning. Although advances in bringing audio-visuals into LLMs have resulted in boosts for a variety of Audio-Visual Question Answering (AVQA) tasks, they still face two crucial challenges: 1) audio-visual ambiguity, and 2) audio-visual hallucination. Existing MLLMs can respond to audio-visual content, yet sometimes fail to describe specific objects due to the ambiguity or hallucination of responses. To overcome the two aforementioned issues, we introduce the CAT+, which enhances MLLM to ensure more robust multimodal understanding. We first propose the Sequential Question-guided Module (SQM), which combines tiny transformer layers and cascades Q-Formers to realize a solid audio-visual grounding. After feature alignment and high-quality instruction tuning, we introduce Ambiguity Scoring Direct Preference Optimization (AS-DPO) to correct the problem of CAT+ bias toward ambiguous descriptions. To explore the hallucinatory deficits of MLLMs in dynamic audio-visual scenes, we build a new Audio-visual Hallucination Benchmark, named AVHbench. This benchmark detects the extent of MLLM's hallucinations across three different protocols in the perceptual object, counting, and holistic description tasks. Extensive experiments across video-based understanding, open-ended, and close-ended AVQA demonstrate the superior performance of our method. The AVHbench is released at https://github.com/rikeilong/Bay-CAT.
Qilang Ye, Zitong Yu, Rui Shao 0001, Yawen Cui, Xiangui Kang, Xin Liu 0012, Philip Torr 0001, Xiaochun Cao
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
Qilang Ye, Zitong Yu, Rui Shao 0001, Xinyu Xie, Philip Torr 0001, Xiaochun Cao
ECCV (10)1
2024 3sG: Three-stage guidance for indoor human action recognition
abstract
Abstract Inference using skeleton to steer RGB videos is applicable to fine‐grained activities in indoor human action recognition (IHAR). However, existing methods that explore only spatial alignment are prone to bias, resulting in limited performance. The authors propose a Three‐stage Guidance (3sG) framework, leveraging skeleton knowledge to promote RGB in three stages. First, a soft shading image is proposed for alleviating background noise in videos, allowing the network to directly focus more on the motion region. Second, the authors propose to extract RGB frames of interest to reduce the computational effort. Furthermore, to explore more fully the complementary information between skeletons and RGB, the skeleton is coupled to the frame representation in a different spatial–temporal sharing pattern. Third, the global skeleton and skeleton‐guided RGB features are fed into the shared classifiers, which approximate the logit distributions of the two to enhance the performance in RGB unimodal. Finally, a fusion strategy that utilizes two learnable parameters to adaptively integrate the skeleton with the RGB is proposed. 3sG outperforms the state‐of‐the‐art results on the Toyota Smarthome dataset while it is more efficient than similar methods on the NTU RGB+D dataset.
Hai Nan, Qilang Ye, Zitong Yu
IET Image Process.2
2024 Pose-Promote: Progressive Visual Perception for Activities of Daily Living
abstract
Poses are effective in interpreting fine-grained human activities, especially when encountering complex visual information. Unimodal methods for action recognition unsatisfactorily to daily activities due to the lack of a more comprehensive perspective. Multimodal methods to combine pose and visual are still not exhaustive enough in mining complementary information. Therefore, we propose a Pose-promote (Ppromo) framework that utilizes a priori knowledge of pose joints to perceive visual information progressively. We first introduce a temporal promote module to activate each video segment using temporally synchronized joint weights. Then a spatial promote module is proposed to capture the key regions in visuals using the learned pose attentions. To further refine the bimodal associations, the global inter-promote module is proposed to align global pose-visual semantics at the feature granularity. Finally, a learnable late fusion strategy between visual and pose is applied for accurate inference. Ppromo achieves state-of-the-art performance on three publicly available datasets.
Qilang Ye, Zitong Yu
IEEE Signal Process. Lett.1