Yang Liu 0293

dblp:51/3710-293 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0001-9982-9887ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Text-Conditional Visual-Language Alignment for Video Captioning
abstract
Video captioning remains a challenging task due to the diverse video content and the complex relationships between visual and textual elements. Recent efforts predominantly focus on multimodal architecture designs trained with paired video-caption data. Nonetheless, the learning paradigm suffers from the “one-to-many” corresponding problem, since one source video is mapped to multiple caption annotations. The difficulty of video captioning is further exacerbated by the poor-written captions, which mislead the captioner with irrelevant information. Essentially, the problem stems from the inadequate alignment between video and caption. In this work, we propose a Text-Conditional Alignment Transformer, which fully exploits the rich information provided by diverse labeled captions, and avoids the impacts of label ambiguity and noise. To alleviate the challenge of the “one-to-many” correspondence, we introduce Text-conditioned Video Encoding, which diversifies the video representation by emphasizing the spatial-temporal visual areas relevant to the given descriptions while filtering out redundant visual information. The refined video representation is well-aligned to match the corresponding text description, and naturally converts the “one-to-many” mapping to “one-to-one” mapping. To deal with the noisy annotations, we propose Quality-aware Caption Decoding. We first dynamically measure the qualities of different captions corresponding to the same video in a reference-free manner. Then the estimated qualities are further utilized as auxiliary signals, guiding the model to perform quality-aligned learning from noisy captions. We conduct extensive experiments on MSR-VTT, MSVD, VATEX and ActivityNet-Entities datasets, and demonstrate their consistent performance improvements compared to state-of-the-arts.
Wenhui Jiang 0001, Wenbin Guan, Zhizhen Li, Yuming Fang 0001, Yuxin Peng 0001, Yang Liu 0293
IEEE Trans. Circuits Syst. Video Technol.8
2025 Opinion-unaware blind stereoscopic image quality assessment: A comprehensive study
Jiebin Yan, Yuming Fang 0001, Xuelin Liu, Wenhui Jiang 0001, Yang Liu 0293
Pattern Recognit.5
2025 Separate, Locate, and Align: Determine Context Relation of Scene Text From Multiple Perspectives in TextVQA
abstract
Text-based Visual Question Answering (TextVQA) focuses on answering questions about the scene text in images. Most works in this field uses transformer based models to modeling the interaction of question and scene texts which means the scene texts will be treated as a natural language sentence and concatenated in reading order as a part of input. However, they ignore the fact that different from words in natural language sentence which have inherent context relation, the context relation of scene texts in images need to be determined. To tackle this problem, we propose a novel method named Separate, Locate and Align (SLA) that discriminate the context relation of scene texts from semantic, visual and spatial aspects. Specifically, based on scene texts with similar visual information (e.g. background color, font color, font style, etc.) having semantic contextual relations, we propose a Text Semantic Separate (TSS) module to discriminate the semantic relation between different scene texts according to their visual contextual information. Then, we introduce a Spatial Circle Position (SCP) module that helps the model discriminate the spatial relation between different scene texts. Last, we design a Visual Alignment (VA) module to help the model distinguish the visual relationships between different scene texts according to the color distribution differences. Extensive experiments show that our method outperforms existing alternatives on TextVQA and ST-VQA datasets without pre-training tasks.
Chengyang Fang, Wenhui Jiang 0001, Yuming Fang 0001, Yuxin Peng 0001, Yang Liu 0293
IEEE Trans. Circuits Syst. Video Technol.5
2025 Learning Comprehensive Visual Grounding for Video Captioning
abstract
The grounding accuracy of existing video captioners is still behind the expectation. The majority of existing methods perform grounded video captioning on sparse entity annotations. However, grounded captioning models rely on deliberate grounding annotations as supervision, which are relatively hard to obtain. Moreover, the captioning accuracy often suffers from degenerated object appearances on the annotated area such as motion blur and video defocus, and these models seldom consider the complex interactions among entities. In this paper, we propose a comprehensive visual grounding network to improve video captioning, by using inexpensive pseudo annotation while avoiding the need to collect large amounts of manual annotations. Specifically, the network consists of spatial-temporal entity grounding and action grounding. The proposed entity grounding encourages the attention mechanism to focus on informative spatial areas across video frames. The action grounding dynamically associates the verbs to related subjects and the corresponding context, which keeps fine-grained spatial and temporal details for action prediction. Both entity grounding and action grounding are formulated as a unified task guided by a soft grounding supervision. More importantly, the grounding objective is supervised by pseudo annotations automatically produced by a grounding annotation generation module, thus our model can be easily applied to the challenging dataset without any grounding annotation provided. We conduct extensive experiments on three benchmark datasets and demonstrate significant performance improvements of +2.4 CIDEr on MSR-VTT, +4.7 CIDEr on MSVD, and +5.1 CIDEr on ActivityNet-Entities compared to state-of-the-arts.
Wenhui Jiang 0001, Linxin Liu, Yuming Fang 0001, Yibo Cheng, Yuxin Peng 0001, Yang Liu 0293
IEEE Trans. Circuits Syst. Video Technol.6
2024 Comprehensive Visual Grounding for Video Description
abstract
The grounding accuracy of existing video captioners is still behind the expectation. The majority of existing methods perform grounded video captioning on sparse entity annotations, whereas the captioning accuracy often suffers from degenerated object appearances on the annotated area such as motion blur and video defocus. Moreover, these methods seldom consider the complex interactions among entities. In this paper, we propose a comprehensive visual grounding network to improve video captioning, by explicitly linking the entities and actions to the visual clues across the video frames. Specifically, the network consists of spatial-temporal entity grounding and action grounding. The proposed entity grounding encourages the attention mechanism to focus on informative spatial areas across video frames, albeit the entity is annotated in only one frame of a video. The action grounding dynamically associates the verbs to related subjects and the corresponding context, which keeps fine-grained spatial and temporal details for action prediction. Both entity grounding and action grounding are formulated as a unified task guided by a soft grounding supervision, which brings architecture simplification and improves training efficiency as well. We conduct extensive experiments on two challenging datasets, and demonstrate significant performance improvements of +2.3 CIDEr on ActivityNet-Entities and +2.2 CIDEr on MSR-VTT compared to state-of-the-arts.
Wenhui Jiang 0001, Yibo Cheng, Linxin Liu, Yuming Fang 0001, Yuxin Peng 0001, Yang Liu 0293
AAAI6
2024 Perceptual Quality Assessment of Omnidirectional Images: A Benchmark and Computational Model
abstract
Compared with traditional 2D images, omnidirectional images (also referred to as 360 ∘ images) have more complicated perceptual characteristics due to the particularities of imaging and display. How humans perceive omnidirectional images in an immersive environment and form the immersive quality of experience are important problems. Thus, it is crucial to measure the quality of omnidirectional images under different viewing conditions, which suffer from realistic distortions. In this article, we build a large-scale subjective assessment database for omnidirectional images and carry out a comprehensive psychophysical experiment to study the relationships between different factors (viewing conditions and viewing behaviors) and the perceptual quality of omnidirectional images. In addition, we collect both subjective ratings and head movement data. A thorough analysis of the collected subjective data is also provided, where we make several interesting findings. Moreover, with the proposed database, we propose a novel transformer-based omnidirectional image quality assessment model. To be consistent with the human viewing process, viewing conditions and behaviors are naturally incorporated into the proposed model. Specifically, the proposed model mainly consists of three parts: viewport sequence generation, multi-scale feature extraction, and perceptual quality prediction. Extensive experimental results conducted on the proposed database demonstrate the effectiveness of the proposed method over existing image quality assessment methods.
Xuelin Liu, Jiebin Yan, Yuming Fang 0001, Yang Liu 0293
ACM Trans. Multim. Comput. Commun. Appl.6
2023 UDNet: Uncertainty-aware deep network for salient object detection
Yuming Fang 0001, Jiebin Yan, Wenhui Jiang 0001, Yang Liu 0293
Pattern Recognit.5
2022 Perceptual Quality Assessment of Omnidirectional Images
abstract
Omnidirectional images, also called 360◦images, have attracted extensive attention in recent years, due to the rapid development of virtual reality (VR) technologies. During omnidirectional image processing including capture, transmission, consumption, and so on, measuring the perceptual quality of omnidirectional images is highly desired, since it plays a great role in guaranteeing the immersive quality of experience (IQoE). In this paper, we conduct a comprehensive study on the perceptual quality of omnidirectional images from both subjective and objective perspectives. Specifically, we construct the largest so far subjective omnidirectional image quality database, where we consider several key influential elements, i.e., realistic non-uniform distortion, viewing condition, and viewing behavior, from the user view. In addition to subjective quality scores, we also record head and eye movement data. Besides, we make the first attempt by using the proposed database to train a convolutional neural network (CNN) for blind omnidirectional image quality assessment. To be consistent with the human viewing behavior in the VR device, we extract viewports from each omnidirectional image and incorporate the user viewing conditions naturally in the proposed model. The proposed model is composed of two parts, including a multi-scale CNN-based feature extraction module and a perceptual quality prediction module. The feature extraction module is used to incorporate the multi-scale features, and the perceptual quality prediction module is designed to regress them to perceived quality scores. The experimental results on our database verify that the proposed model achieves the competing performance compared with the state-of-the-art methods.
Yuming Fang 0001, Jiebin Yan, Xuelin Liu, Yang Liu 0293
AAAI5
2022 Revisiting image captioning via maximum discrepancy competition
Boyang Wan, Wenhui Jiang 0001, Yuming Fang 0001, Minwei Zhu, Yang Liu 0293
Pattern Recognit.6
2022 Visual Cluster Grounding for Image Captioning
abstract
Attention mechanisms have been extensively adopted in vision and language tasks such as image captioning. It encourages a captioning model to dynamically ground appropriate image regions when generating words or phrases, and it is critical to alleviate the problems of object hallucinations and language bias. However, current studies show that the grounding accuracy of existing captioners is still far from satisfactory. Recently, much effort is devoted to improving the grounding accuracy by linking the words to the full content of objects in images. However, due to the noisy grounding annotations and large variations of object appearance, such strict word-object alignment regularization may not be optimal for improving captioning performance. In this paper, to improve the performance of both grounding and captioning, we propose a novel grounding model which implicitly links the words to the evidence in the image. The proposed model encourages the captioner to dynamically focus on informative regions of the objects, which could be either discriminative parts or full object content. With slacked constraints, the proposed captioning model can capture correct linguistic characteristics and visual relevance, and then generate more grounded image captions. In addition, we propose a novel quantitative metric for evaluating the correctness of the soft attention mechanism by considering the overall contribution of all object proposals when generating certain words. The proposed grounding model can be seamlessly plugged into most attention-based architectures without introducing inference complexity. We conduct extensive experiments on Flickr30k (Young et al., 2014) and MS COCO datasets (Lin et al., 2014), demonstrating that the proposed method consistently improves image captioning in both grounding and captioning. Besides, the proposed attention evaluation metric shows better consistency with the captioning performance.
Wenhui Jiang 0001, Minwei Zhu, Yuming Fang 0001, Guangming Shi, Yang Liu 0293
IEEE Trans. Image Process.6
2022 Subjective and Objective Quality of Experience of Free Viewpoint Videos
abstract
Free viewpoint videos (FVVs) provide immersive experiences for end-users, and they have been applied in many applications, such as movies, sports, and TV shows. However, the development of quantifying the quality of experience (QoE) of FVVs is still relatively slow due to the high costs of data collection and limited public databases. In this paper, we conduct a comprehensive study on FVV QoE. First, we construct the largest, to the best of our knowledge, FVV QoE database called Youku-FVV from two complex real scenarios, i. e., entertainment and sports. Specifically, Youku-FVV originates from the videos captured by dozens of real cameras arranged annularly. We use these videos to generate virtual viewpoints, which make up FVVs together with real views. In constructing the FVV QoE database, we consider both internal and external influencing factors of QoE, which correspond to FVV generation and playback, respectively. Besides, we make an initial attempt to train an efficient no reference FVV QoE prediction model using this database, where several sparse frame sampling strategies are validated. And we demonstrate the feasibility of striving for the balance between effectiveness and efficiency of FVV QoE prediction. The proposed FVV QoE database and source codes are publicly available at https://github.com/QTJiebin/FVV_QoE.
Jiebin Yan, Jing Li 0026, Yuming Fang 0001, Zhaohui Che, Xue Xia 0005, Yang Liu 0293
IEEE Trans. Image Process.6