VLDB 2026 Research / reviewers in the wild / expert
Yiqi Gao
dblp:126/4401
· DBLP profile ↗
12ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-3494-4539ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | More video-relevant paragraph captioning via Perturbed Attention Self-Distillation
Yiqi Gao, Wei Suo, Mengyang Sun, Le Liu 0008, Peng Wang 0015 |
Pattern Recognit. | 1 |
| 2024 | Human-Centric Behavior Description in Videos: New Benchmark and ModelabstractIn the domain of video surveillance, describing the behavior of each individual within the video is becoming increasingly essential, especially in complex scenarios with multiple individuals present. This is because describing each individual's behavior provides more detailed situational analysis, enabling accurate assessment and response to potential risks, ensuring the safety and harmony of public places. Currently, video-level captioning datasets cannot provide fine-grained descriptions for each individual's specific behavior. However, mere descriptions at the video-level fail to provide an in-depth interpretation of individual behaviors, making it challenging to accurately determine the specific identity of each individual. To address this challenge, we construct a human-centric video surveillance captioning dataset, which provides detailed descriptions of the dynamic behaviors of 7,820 individuals. Specifically, we have labeled several aspects of each person, such as location, clothing, and interactions with other elements in the scene, and these people are distributed across 1,012 videos. Based on this dataset, we can link individuals to their respective behaviors, allowing for further analysis of each person's behavior in surveillance videos. Besides the dataset, we propose a novel video captioning approach that can describe individual behavior in detail on a person-level basis, achieving state-of-the-art results. Lingru Zhou, Yiqi Gao, Manqing Zhang, Peng Wu 0015, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | S3C: Semi-Supervised VQA Natural Language Explanation via Self-Critical LearningabstractVQA Natural Language Explanation (VQA-NLE) task aims to explain the decision-making process of VQA models in natural language. Unlike traditional attention or gradient analysis, free-text rationales can be easier to understand and gain users' trust. Existing methods mostly use post-hoc or selfrationalization models to obtain a plausible explanation. However, these frameworks are bottle-necked by the following challenges: 1) the reasoning process cannot be faithfully responded to and suffer from the problem of logical inconsistency. 2) Human-annotated explanations are expensive and time-consuming to collect. In this paper, we propose a new Semi-Supervised VQA-NLE via Self-Critical Learning (S3C), which evaluates the candidate explanations by answering rewards to improve the logical consistency between answers and rationales. With a semi-supervised learning framework, the S3C can benefit from a tremendous amount of samples without human-annotated explanations. A large number of automatic measures and human evaluations all show the effectiveness of our method. Meanwhile, the framework achieves a new state-of-the-art performance on the two VQA-NLE datasets. Wei Suo, Mengyang Sun, Weisong Liu, Yiqi Gao, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
CVPR | 4 |
| 2022 | A Simple and Robust Correlation Filtering Method for Text-Based Person Search
Wei Suo, Mengyang Sun, Kai Niu 0002, Yiqi Gao, Peng Wang 0015, Yanning Zhang 0001, Qi Wu 0001 |
ECCV (35) | 4 |
| 2022 | CapOnImage: Context-driven Dense-Captioning on ImageabstractExisting image captioning systems are dedicated to generating narrative captions for images, which are spatially detached from the image in presentation.However, texts can also be used as decorations on the image to highlight the key points and increase the attractiveness of images.In this work, we introduce a new task called captioning on image (CapOn-Image) 1 , which aims to generate dense captions at different locations of the image based on contextual information.For this new task, we introduce a large-scale benchmark called CapOn-Image2M, which contains 2.1 million product images, each with an average of 4.8 spatially localized captions.To fully exploit the surrounding visual context to generate the most suitable caption for each location, we propose a multimodal pre-training model with multi-level pretraining tasks that progressively learn the correspondence between texts and image locations from easy to hard.To avoid generating redundant captions for nearby locations, we further enhance the location embedding with neighbor locations .Compared with other image captioning model variants, our model achieves the best results in both captioning accuracy and diversity aspects. Yiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge, Yuning Jiang 0001, Peng Wang 0015 |
EMNLP | 1 |
| 2022 | Dual-Level Decoupled Transformer for Video CaptioningabstractVideo captioning aims to understand the spatio-temporal semantic concept of the video and generate descriptive sentences. The de-facto approach to this task dictates a text generator to learn from offline-extracted motion or appearance features from pre-trained vision models. However, these methods may suffer from the so-called "couple" drawbacks on both video spatio-temporal representation and sentence generation. For the former, "couple" means learning spatio-temporal representation in a single model(3DCNN), resulting the problems named disconnection in task/pre-train domain and hard for end-to-end training. As for the latter, "couple" means treating the generation of visual semantic and syntax-related words equally. To this end, we present D2 - a dual-level decoupled transformer pipeline to solve the above drawbacks: (i) for video spatio-temporal representation, we decouple the process of it into "first-spatial-then-temporal" paradigm, releasing the potential of using dedicated model(e.g. image-text pre-training) to connect the pre-training and downstream tasks, and makes the entire model end-to-end trainable. (ii) for sentence generation, we propose Syntax-Aware Decoder to dynamically measure the contribution of visual semantic and syntax-related words. Extensive experiments on three widely-used benchmarks (MSVD, MSR-VTT and VATEX) have shown great potential of the proposed D2 and surpassed the previous methods by a large margin in the task of video captioning. Yiqi Gao, Xinglin Hou, Wei Suo, Mengyang Sun, Tiezheng Ge, Yuning Jiang 0001, Peng Wang 0015 |
ICMR | 1 |
| 2022 | Improving Image Captioning via Enhancing Dual-Side Context AwarenessabstractRecent work on visual question answering demonstrate that grid features can work as well as region feature on vision language tasks. In the meantime, transformer-based model and its variants have shown remarkable performance on image captioning. However, the object-contextual information missing caused by the single granularity nature of grid feature on the encoder side, as well as the future contextual information missing due to the left2right decoding paradigm of transformer decoder, remains unexplored. In this work, we tackle these two problems by enhancing contextual information at dual-side:(i) at encoder side, we propose Context-Aware Self-Attention module, in which the key/value is expanded with adjacent rectangle region where each region contains two or more aggregated grid features; this enables grid feature with varying granularity, storing adequate contextual information for object with different scale. (ii) at decoder side, we incorporate a dual-way decoding strategy, in which left2right and right2left decoding are conducted simultaneously and interactively. It utilizes both past and future contextual information when generates current word. Combining these two modules with a vanilla transformer, our Context-Aware Transformer(CATNet) achieves a new state-of-the-art on MSCOCO benchmark. Yiqi Gao, Ning Wang 0020, Wei Suo, Mengyang Sun, Peng Wang 0015 |
ICMR | 1 |
| 2014 | Semiautonomous Vehicular Control Using Driver ModelingabstractThreat assessment during semiautonomous driving is used to determine when correcting a driver's input is required. Since current semiautonomous systems perform threat assessment by predicting a vehicle's future state while treating the driver's input as a disturbance, autonomous controller intervention is limited to a restricted regime. Improving vehicle safety demands threat assessment that occurs over longer prediction horizons wherein a driver cannot be treated as a malicious agent. In this paper, we describe a real-time semiautonomous system that utilizes empirical observations of a driver's pose to inform an autonomous controller that corrects a driver's input when possible in a safe manner. We measure the performance of our system using several metrics that evaluate the informativeness of the prediction and the utility of the intervention procedure. A multisubject driving experiment illustrates the usefulness, with respect to these metrics, of incorporating the driver's pose while designing a semiautonomous system. Victor Shia, Yiqi Gao, Ramanarayan Vasudevan, Katherine Rose Driggs-Campbell, Theresa Lin, Francesco Borrelli, Ruzena Bajcsy |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2013 | Efficient seam carving for object removalabstractThis paper introduces a new object removal approach for images based on discontinuous seam carving. Existing seam carving based object removal methods generally are time-consuming and oftentimes cause image distortion or cutting off many more seams than is necessary. In order to solve these limitations, our proposed method only considers the energy of all the pixels outside the target region. Firstly, based on the discontinuous seam carving, we calculate the energy map of both directions, up-down and bottom-up. Then we carve the seam respectively from the upper bound of the target region to up, from the lower bound of the target region to down and the middle part within the region. Experimental results prove that our proposed method outperform others in terms of efficiency and image quality significantly. Bo Yan 0001, Yiqi Gao, Kairan Sun |
ICIP | 2 |
| 2013 | Robust Predictive Control for semi-autonomous vehicles with an uncertain driver modelabstractA robust control design is proposed for the lane-keeping and obstacle avoidance of semiautonomous ground vehicles. A robust Model Predictive Controller (MPC) is used in order to enforce safety constraints with minimal control intervention. An uncertain driver model is used to obtain sets of predicted vehicle trajectories in closed-loop with the predicted driver's behavior. The robust MPC computes the smallest corrective steering action needed to keep the driver safe for all predicted trajectories in the set. Simulations of a driver approaching multiple obstacles, with uncertainty obtained from measured data, show the effect of the proposed framework. Andrew Gray, Yiqi Gao, J. Karl Hedrick, Francesco Borrelli |
Intelligent Vehicles Symposium | 2 |
| 2013 | A Unified Approach to Threat Assessment and Control for Automotive Active SafetyabstractThis paper presents the design of a novel active safety system preventing unintended roadway departures. The proposed framework unifies threat assessment, stability, and control of passenger vehicles into a single combined optimization problem. A nonlinear model predictive control (MPC) problem is formulated, where nonlinear vehicle dynamics, in closed-loop with a driver model, is used to optimize the steering and braking actions needed to keep the driver safe. A model of the driver's nominal behavior is estimated based on his observed behavior. The driver commands the vehicle, whereas the safety system corrects the driver's steering and braking actions in case there is a risk that the vehicle will unintentionally depart from the road. The resulting predictive controller is always active, and mode switching is not necessary. We show simulation results detailing the behavior of the proposed controller and experimental results obtained by implementing the proposed framework on embedded hardware in a passenger vehicle. The results demonstrate the capability of the proposed controller to detect and avoid roadway departures while avoiding unnecessary interventions. Andrew Gray, Mohammad Ali 0002, Yiqi Gao, J. Karl Hedrick, Francesco Borrelli |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2012 | Lowcomplexity content-aware image retargetingabstractImage retargeting plays a more and more significant role recently, thanks to the escalating diversity of display devices. In this paper, we present a novel low complexity content-aware image retargeting method, which can provide both high efficiency and quality. Specifically, the importance map is used as the basis of our method. The important regions tend to preserve its original size. Our method performs inter-row coherence filtering on importance maps in order to maintain the spatial coherence, and then directly utilizes the filtered importance maps to generate the scaling map. Experimental results show the proposed algorithm improves the efficiency significantly compared with other existing methods. At the same time, the resized image quality of our algorithm is as good as, if not better than, that of the other methods. As a result, our method possesses huge practical significance. Kairan Sun, Bo Yan 0001, Yiqi Gao |
ICIP | 3 |