Yiqi Gao

dblp:126/4401 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-3494-4539ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 More video-relevant paragraph captioning via Perturbed Attention Self-Distillation
Yiqi Gao, Wei Suo, Mengyang Sun, Le Liu 0008, Peng Wang 0015
Pattern Recognit.1
2024 Human-Centric Behavior Description in Videos: New Benchmark and Model
abstract
In the domain of video surveillance, describing the behavior of each individual within the video is becoming increasingly essential, especially in complex scenarios with multiple individuals present. This is because describing each individual's behavior provides more detailed situational analysis, enabling accurate assessment and response to potential risks, ensuring the safety and harmony of public places. Currently, video-level captioning datasets cannot provide fine-grained descriptions for each individual's specific behavior. However, mere descriptions at the video-level fail to provide an in-depth interpretation of individual behaviors, making it challenging to accurately determine the specific identity of each individual. To address this challenge, we construct a human-centric video surveillance captioning dataset, which provides detailed descriptions of the dynamic behaviors of 7,820 individuals. Specifically, we have labeled several aspects of each person, such as location, clothing, and interactions with other elements in the scene, and these people are distributed across 1,012 videos. Based on this dataset, we can link individuals to their respective behaviors, allowing for further analysis of each person's behavior in surveillance videos. Besides the dataset, we propose a novel video captioning approach that can describe individual behavior in detail on a person-level basis, achieving state-of-the-art results.
Lingru Zhou, Yiqi Gao, Manqing Zhang, Peng Wu 0015, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Multim.2
2023 S3C: Semi-Supervised VQA Natural Language Explanation via Self-Critical Learning
abstract
VQA Natural Language Explanation (VQA-NLE) task aims to explain the decision-making process of VQA models in natural language. Unlike traditional attention or gradient analysis, free-text rationales can be easier to understand and gain users' trust. Existing methods mostly use post-hoc or selfrationalization models to obtain a plausible explanation. However, these frameworks are bottle-necked by the following challenges: 1) the reasoning process cannot be faithfully responded to and suffer from the problem of logical inconsistency. 2) Human-annotated explanations are expensive and time-consuming to collect. In this paper, we propose a new Semi-Supervised VQA-NLE via Self-Critical Learning (S3C), which evaluates the candidate explanations by answering rewards to improve the logical consistency between answers and rationales. With a semi-supervised learning framework, the S3C can benefit from a tremendous amount of samples without human-annotated explanations. A large number of automatic measures and human evaluations all show the effectiveness of our method. Meanwhile, the framework achieves a new state-of-the-art performance on the two VQA-NLE datasets.
Wei Suo, Mengyang Sun, Weisong Liu, Yiqi Gao, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001
CVPR4
2022 A Simple and Robust Correlation Filtering Method for Text-Based Person Search
Wei Suo, Mengyang Sun, Kai Niu 0002, Yiqi Gao, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001
ECCV (35)4
2022 CapOnImage: Context-driven Dense-Captioning on Image
abstract
Existing image captioning systems are dedicated to generating narrative captions for images, which are spatially detached from the image in presentation.However, texts can also be used as decorations on the image to highlight the key points and increase the attractiveness of images.In this work, we introduce a new task called captioning on image (CapOn-Image) 1 , which aims to generate dense captions at different locations of the image based on contextual information.For this new task, we introduce a large-scale benchmark called CapOn-Image2M, which contains 2.1 million product images, each with an average of 4.8 spatially localized captions.To fully exploit the surrounding visual context to generate the most suitable caption for each location, we propose a multimodal pre-training model with multi-level pretraining tasks that progressively learn the correspondence between texts and image locations from easy to hard.To avoid generating redundant captions for nearby locations, we further enhance the location embedding with neighbor locations .Compared with other image captioning model variants, our model achieves the best results in both captioning accuracy and diversity aspects.
Yiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge, Yuning Jiang 0001, Peng Wang 0015
EMNLP1
2022 Dual-Level Decoupled Transformer for Video Captioning
abstract
Video captioning aims to understand the spatio-temporal semantic concept of the video and generate descriptive sentences. The de-facto approach to this task dictates a text generator to learn from offline-extracted motion or appearance features from pre-trained vision models. However, these methods may suffer from the so-called "couple" drawbacks on both video spatio-temporal representation and sentence generation. For the former, "couple" means learning spatio-temporal representation in a single model(3DCNN), resulting the problems named disconnection in task/pre-train domain and hard for end-to-end training. As for the latter, "couple" means treating the generation of visual semantic and syntax-related words equally. To this end, we present D2 - a dual-level decoupled transformer pipeline to solve the above drawbacks: (i) for video spatio-temporal representation, we decouple the process of it into "first-spatial-then-temporal" paradigm, releasing the potential of using dedicated model(e.g. image-text pre-training) to connect the pre-training and downstream tasks, and makes the entire model end-to-end trainable. (ii) for sentence generation, we propose Syntax-Aware Decoder to dynamically measure the contribution of visual semantic and syntax-related words. Extensive experiments on three widely-used benchmarks (MSVD, MSR-VTT and VATEX) have shown great potential of the proposed D2 and surpassed the previous methods by a large margin in the task of video captioning.
Yiqi Gao, Xinglin Hou, Wei Suo, Mengyang Sun, Tiezheng Ge, Yuning Jiang 0001, Peng Wang 0015
ICMR1
2022 Improving Image Captioning via Enhancing Dual-Side Context Awareness
abstract
Recent work on visual question answering demonstrate that grid features can work as well as region feature on vision language tasks. In the meantime, transformer-based model and its variants have shown remarkable performance on image captioning. However, the object-contextual information missing caused by the single granularity nature of grid feature on the encoder side, as well as the future contextual information missing due to the left2right decoding paradigm of transformer decoder, remains unexplored. In this work, we tackle these two problems by enhancing contextual information at dual-side:(i) at encoder side, we propose Context-Aware Self-Attention module, in which the key/value is expanded with adjacent rectangle region where each region contains two or more aggregated grid features; this enables grid feature with varying granularity, storing adequate contextual information for object with different scale. (ii) at decoder side, we incorporate a dual-way decoding strategy, in which left2right and right2left decoding are conducted simultaneously and interactively. It utilizes both past and future contextual information when generates current word. Combining these two modules with a vanilla transformer, our Context-Aware Transformer(CATNet) achieves a new state-of-the-art on MSCOCO benchmark.
Yiqi Gao, Ning Wang 0020, Wei Suo, Mengyang Sun, Peng Wang 0015
ICMR1
2014 Semiautonomous Vehicular Control Using Driver Modeling
abstract
Threat assessment during semiautonomous driving is used to determine when correcting a driver's input is required. Since current semiautonomous systems perform threat assessment by predicting a vehicle's future state while treating the driver's input as a disturbance, autonomous controller intervention is limited to a restricted regime. Improving vehicle safety demands threat assessment that occurs over longer prediction horizons wherein a driver cannot be treated as a malicious agent. In this paper, we describe a real-time semiautonomous system that utilizes empirical observations of a driver's pose to inform an autonomous controller that corrects a driver's input when possible in a safe manner. We measure the performance of our system using several metrics that evaluate the informativeness of the prediction and the utility of the intervention procedure. A multisubject driving experiment illustrates the usefulness, with respect to these metrics, of incorporating the driver's pose while designing a semiautonomous system.
Victor Shia, Yiqi Gao, Ramanarayan Vasudevan, Katherine Rose Driggs-Campbell, Theresa Lin, Francesco Borrelli, Ruzena Bajcsy
IEEE Trans. Intell. Transp. Syst.2
2013 Efficient seam carving for object removal
abstract
This paper introduces a new object removal approach for images based on discontinuous seam carving. Existing seam carving based object removal methods generally are time-consuming and oftentimes cause image distortion or cutting off many more seams than is necessary. In order to solve these limitations, our proposed method only considers the energy of all the pixels outside the target region. Firstly, based on the discontinuous seam carving, we calculate the energy map of both directions, up-down and bottom-up. Then we carve the seam respectively from the upper bound of the target region to up, from the lower bound of the target region to down and the middle part within the region. Experimental results prove that our proposed method outperform others in terms of efficiency and image quality significantly.
Bo Yan 0001, Yiqi Gao, Kairan Sun
ICIP2
2013 Robust Predictive Control for semi-autonomous vehicles with an uncertain driver model
abstract
A robust control design is proposed for the lane-keeping and obstacle avoidance of semiautonomous ground vehicles. A robust Model Predictive Controller (MPC) is used in order to enforce safety constraints with minimal control intervention. An uncertain driver model is used to obtain sets of predicted vehicle trajectories in closed-loop with the predicted driver's behavior. The robust MPC computes the smallest corrective steering action needed to keep the driver safe for all predicted trajectories in the set. Simulations of a driver approaching multiple obstacles, with uncertainty obtained from measured data, show the effect of the proposed framework.
Andrew Gray, Yiqi Gao, J. Karl Hedrick, Francesco Borrelli
Intelligent Vehicles Symposium2
2013 A Unified Approach to Threat Assessment and Control for Automotive Active Safety
abstract
This paper presents the design of a novel active safety system preventing unintended roadway departures. The proposed framework unifies threat assessment, stability, and control of passenger vehicles into a single combined optimization problem. A nonlinear model predictive control (MPC) problem is formulated, where nonlinear vehicle dynamics, in closed-loop with a driver model, is used to optimize the steering and braking actions needed to keep the driver safe. A model of the driver's nominal behavior is estimated based on his observed behavior. The driver commands the vehicle, whereas the safety system corrects the driver's steering and braking actions in case there is a risk that the vehicle will unintentionally depart from the road. The resulting predictive controller is always active, and mode switching is not necessary. We show simulation results detailing the behavior of the proposed controller and experimental results obtained by implementing the proposed framework on embedded hardware in a passenger vehicle. The results demonstrate the capability of the proposed controller to detect and avoid roadway departures while avoiding unnecessary interventions.
Andrew Gray, Mohammad Ali 0002, Yiqi Gao, J. Karl Hedrick, Francesco Borrelli
IEEE Trans. Intell. Transp. Syst.3
2012 Lowcomplexity content-aware image retargeting
abstract
Image retargeting plays a more and more significant role recently, thanks to the escalating diversity of display devices. In this paper, we present a novel low complexity content-aware image retargeting method, which can provide both high efficiency and quality. Specifically, the importance map is used as the basis of our method. The important regions tend to preserve its original size. Our method performs inter-row coherence filtering on importance maps in order to maintain the spatial coherence, and then directly utilizes the filtered importance maps to generate the scaling map. Experimental results show the proposed algorithm improves the efficiency significantly compared with other existing methods. At the same time, the resized image quality of our algorithm is as good as, if not better than, that of the other methods. As a result, our method possesses huge practical significance.
Kairan Sun, Bo Yan 0001, Yiqi Gao
ICIP3