Lida Yin

dblp:376/0532 · DBLP profile ↗
← Back
2ranked-venue papers in the field
0as first author
2since 2021 · last 2024
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2024 A Non-reference Just Recognized Distortion Prediction Framework for Object Detection Task
abstract
This work proposed a non-reference Just Recognized Distortion (JRD) prediction framework for object detection task based on Generative Adversarial Network (GAN) image generation. Inspired by the concept of Just Noticeable Difference (JND), the JRD is used to describe the threshold of image distortion acceptable for machine vision tasks. Centered around JRD, the proposed framework primarily consisted of a JRD image generation network and a residual-guided JRD regression network. Among them, the JRD image generation network was trained in conjunction with a multi-scale discriminator in the form of GAN. Additionally, we had compiled a dataset comprising over 130,000 images for the YOLOv7 object detection task, and it was used to validate the effectiveness of our proposed framework. Experiments indicated that the framework can approximate the JRD for object detection task with notable accuracy. Consequently, adopting our proposed framework for image compression could reduce the bitrate by 50% while maintaining high accuracy in task.
Yichen Liu 0006, Haibing Yin, Hongkui Wang, Xia Wang 0006, Lida Yin
DCC5
2024 BHSE-VQA: A Bidirectional Hierarchical Semantic Extraction Structure for Video Quality Assessment
abstract
The diversity of video content and unpredictability of distortions in user-generated content (UGC) videos pose a challenge for video quality assessment (VQA). Most existing methods are difficult to model complete visual perception loop to accurately capture video content and predict perceived quality. Thus, as shown in Figure 1 , this paper proposes a bidirectional hierarchical semantic extraction structure for VQA (BHSE-VQA), which simulates visual feedforward and feedback perception processes. Firstly, the feedforward and feedback multi-level network (FFMNet) based on the reverse hierarchy theory is designed to extract and adjust hierarchical semantic features on the bidirectional pathway. Then, considering the different effects of feature depth on visual perception results, this paper suggests a multi-level weight redistribution (MWR) strategy that makes use of the attention characteristics of the feedback outputs to realign the weights at each stage of the feedforward outputs. Finally, through a temporal attention fusion network (TAFNet), this paper further extracts the detailed features arising from feedforward and feedback perceptual differences and obtains quality scores. The experimental results show that the proposed model exceeds 0.845 on both SRCC and PLCC on YouTube-UGC database.
Longbin Mo, Haibing Yin, Hongkui Wang, Xia Wang 0006, Lida Yin, Tiansong Li
DCC5