Tingyu Weng

dblp:258/7035 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-4760-5552ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Vision and language · 29% Video understanding and tracking · 24% 3D vision · 17%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
2.632025
UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface · NeurIPS 2025
Aligned Better, Listen Better for Audio-Visual Large Language Models · ICLR 2025
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding · ICCV 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
audio-visual large language model
0.912025
Aligned Better, Listen Better for Audio-Visual Large Language Models · ICLR 2025
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding
0.912025
Aligned Better, Listen Better for Audio-Visual Large Language Models · ICLR 2025
Computer vision › Segmentation and scene understanding
instance segmentation
0.912025
UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface · NeurIPS 2025
Computer vision › Video understanding and tracking
multimodal video understanding
0.912025
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding · ICCV 2025
Computer vision › Segmentation and scene understanding
semantic segmentation
0.912025
UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface · NeurIPS 2025
Computer vision › Video understanding and tracking
video representation learning
0.912025
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding · ICCV 2025
Computer vision › 3D vision
3d object recognition
0.812024
PartCom: Part Composition Learning for 3D Open-Set Recognition · Int. J. Comput. Vis. 2024
Machine learning › Trustworthy machine learning › open-world recognition
open-set recognition
0.812024
PartCom: Part Composition Learning for 3D Open-Set Recognition · Int. J. Comput. Vis. 2024
Computer vision › 3D vision › 3d shape analysis
3d shape understanding
0.712023
Decompose Novel into Known: Part Concept Learning For 3D Novel Class Discovery · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning
part-based representation learning
0.712023
Decompose Novel into Known: Part Concept Learning For 3D Novel Class Discovery · NeurIPS 2023
Computer vision › 3D vision › point cloud segmentation
point cloud semantic segmentation
0.712023
Context-Aware 3D Point Cloud Semantic Segmentation With Plane Guidance · IEEE Trans. Multim. 2023
Computer vision › Image recognition and object detection
object detection
0.312025
UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface · NeurIPS 2025
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatiotemporal representation
0.312025
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding · ICCV 2025
Machine learning › Transfer learning and domain adaptation › generalization to unseen classes
cross-category generalization
0.212023
Decompose Novel into Known: Part Concept Learning For 3D Novel Class Discovery · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

language interface · 0.9keyframe selection · 0.9instruction tuning · 0.9embedding retrieval · 0.9audio-visual multi-scale adapter · 0.9audio-visual interleaved merging · 0.94d rotary position embedding · 0.9plane relation network · 0.7part concept bank · 0.7contrastive learning · 0.7
YearPublicationVenuePosition
2025 DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
abstract
In recent years, the introduction of Multi-modal Large Language Models (MLLMs) into video understanding tasks has become increasingly prevalent. However, how to effectively integrate temporal information remains a critical research focus. Traditional approaches treat spatial and temporal information separately. Due to issues like motion blur, it is challenging to accurately represent the spatial information of rapidly moving objects. This can lead to temporally important regions being underemphasized during spatial feature extraction, which in turn hinders accurate spatio-temporal interaction and video understanding. To address this limitation, we propose an innovative video representation method called Dynamic-Image (DynImg). Specifically, we introduce a set of non-key frames as temporal prompts to highlight the spatial areas containing fast-moving objects. During the process of visual feature extraction, these prompts guide the model to pay additional attention to the fine-grained spatial features corresponding to these regions. Moreover, to maintain the correct sequence for DynImg, we employ a corresponding 4D video Rotary Position Embedding. This retains both the temporal and spatial adjacency of DynImg, helping MLLM understand the spatio-temporal order within this combined format. Experimental evaluations reveal that DynImg surpasses the state-of-the-art methods by approximately 2% across multiple video understanding benchmarks, proving the effectiveness of our temporal prompts in enhancing video comprehension.
Xiaoyi Bao, Chenwei Xie, Hao Tang 0005, Tingyu Weng
ICCV4
2025 Aligned Better, Listen Better for Audio-Visual Large Language Models
abstract
Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric settings. However, existing Video-LLMs and Audio-Visual Large Language Models (AV-LLMs) exhibit deficiencies in exploiting audio information, leading to weak understanding and hallucinations. To solve the issues, we delve into the model architecture and dataset. (1) From the architectural perspective, we propose a fine-grained AV-LLM, namely Dolphin. The concurrent alignment of audio and visual modalities in both temporal and spatial dimensions ensures a comprehensive and accurate understanding of videos. Specifically, we devise an audio-visual multi-scale adapter for multi-scale information aggregation, which achieves spatial alignment. For temporal alignment, we propose audio-visual interleaved merging. (2) From the dataset perspective, we curate an audio-visual caption \& instruction-tuning dataset, called AVU. It comprises 5.2 million diverse, open-ended data tuples (video, audio, question, answer) and introduces a novel data partitioning strategy. Extensive experiments show our model not only achieves remarkable performance in audio-visual understanding, but also mitigates potential hallucinations.
Shuailei Ma, Shijie Ma, Xiaoyi Bao, Chen-Wei Xie, Kecheng Zheng, Tingyu Weng, Siyang Sun
ICLR7
2025 UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface
abstract
Generalist models have achieved remarkable success in both language and vision-language tasks, showcasing the potential of unified modeling. However, effectively integrating fine-grained perception tasks like detection and segmentation into these models remains a significant challenge. This is primarily because these tasks often rely heavily on task-specific designs and architectures that can complicate the modeling process. To address this challenge, we present UFO, a framework that unifies fine-grained visual perception tasks through an open-ended language interface. By transforming all perception targets into the language space, UFO unifies object-level detection, pixel-level segmentation, and image-level vision-language tasks into a single model. Additionally, we introduce a novel embedding retrieval approach that relies solely on the language interface to support segmentation tasks. Our framework bridges the gap between fine-grained perception and vision-language tasks, significantly simplifying architectural design and training strategies while achieving comparable or superior performance to methods with intricate task-specific designs. After multi-task training on five standard visual perception datasets, UFO outperforms the previous state-of-the-art generalist models by 12.3 mAP on COCO instance segmentation and 3.3 mIoU on ADE20K semantic segmentation. Furthermore, our method seamlessly integrates with existing MLLMs, effectively combining fine-grained perception capabilities with their advanced language abilities, thereby achieving superior performance on the challenging reasoning segmentation. Code and models are available at https://github.com/nnnth/UFO.
Hao Tang 0005, Chen-Wei Xie, Xiaoyi Bao, Tingyu Weng, Pandeng Li, Liwei Wang 0001
NeurIPS5
2024 PartCom: Part Composition Learning for 3D Open-Set Recognition
Tingyu Weng, Jun Xiao 0005, Hao Pan 0001, Haiyong Jiang
Int. J. Comput. Vis.1
2023 Decompose Novel into Known: Part Concept Learning For 3D Novel Class Discovery
abstract
In this work, we address 3D novel class discovery (NCD) that discovers novel classes from an unlabeled dataset by leveraging the knowledge of disjoint known classes. The key challenge of 3D NCD is that learned features by known class recognition are heavily biased and hinder generalization to novel classes. Since geometric parts are more generalizable across different classes, we propose to decompose novel into known parts, coined DNIK, to mitigate the above problems. DNIK learns a part concept bank encoding rich part geometric patterns from known classes so that novel 3D shapes can be represented as part concept compositions to facilitate cross-category generalization. Moreover, we formulate three constraints on part concepts to ensure diverse part concepts without collapsing. A part relation encoding module (PRE) is also developed to leverage part-wise spatial relations for better recognition. We construct three 3D NCD tasks for evaluation and extensive experiments show that our method achieves significantly superior results than SOTA baselines (+11.7%, +14.1%, and +16.3% improvements on average for three tasks, respectively). Code and data will be released.
Tingyu Weng, Jun Xiao 0005, Haiyong Jiang
NeurIPS1
2023 Context-Aware 3D Point Cloud Semantic Segmentation With Plane Guidance
abstract
Point cloud segmentation is fundamental in under- standing 3D environments. However, most existing methods usually perform poorly on identifying boundaries of touching objects and large surfaces of objects. Planes in a scene usually act as supporting surfaces to separate touching objects and provide geometry priors to group points on a large surface as shown in Fig. 1. Besides, planes can roughly represent the structure of a scene, and are more efficient to encode holistic scene contexts than large scale point clouds. In light of the above advantages, we advise a plane-assisted module, coined3D-PAM, to enhance semantic segmentation of touching objects and large surface objects.3D-PAMconsists of a plane separation network (PS-Net) and a plane relation network (PR-Net).PS-Netfocuses on learning features that can robustly separate touching objects, e.g., a chair on a floor, as well as capture plane-based geometry priors to group points on a large plane, e.g., points of a desk.PR-Netencodes mutual plane relations as a proxy of a scene structure to capture holistic contexts.3D-PAMis designed as a plug-and-play module so that it can be easily plugged into any off-the-shelf semantic segmentation network. Extensive experiments demonstrate that the method achieves large segmentation improvements on several backbones, and accomplishes superior results on most categories when using a RandLA-Net backbone ($11/13$categories on S3DIS dataset and$15/20$categories on ScanNetv2 dataset). The project is available at GitHubhttps://github.com/windmillknight/Context-Aware-3D-Point-Cloud-Semantic-Segmentation-With-Plane-Guidance
Tingyu Weng, Jun Xiao 0005, Feilong Yan, Haiyong Jiang
IEEE Trans. Multim.1