Jinqiang Cui

dblp:117/8652 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
15since 2021 · last 2025
0000-0002-7833-1876ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 9 since 2021Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory
abstract
Weichen Zhang, Chen Gao, Shiquan Yu, Ruiying Peng, Baining Zhao, Qian Zhang, Jinqiang Cui, Xinlei Chen, Yong Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chen Gao 0001, Shiquan Yu, Ruiying Peng, Baining Zhao, Jinqiang Cui, Xinlei Chen, Yong Li 0008
ACL (1)7
2025 UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
abstract
Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Weichen Zhang, Chen Gao, Yue Wang, Jinqiang Cui, Xinlei Chen, Yong Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Chen Gao 0001, Yue Wang 0007, Jinqiang Cui, Xinlei Chen, Yong Li 0008
ACL (1)9
2025 TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
abstract
Visual Place Recognition (VPR) is a crucial capability for long-term autonomous robots, enabling them to identify previously visited locations using visual information. However, existing methods remain limited in indoor settings due to the highly repetitive structures inherent in such environments. We observe that scene texts frequently appear in indoor spaces and can help distinguish visually similar but different places. This inspires us to propose TextInPlace, a simple yet effective VPR framework that integrates Scene Text Spotting (STS) to mitigate visual perceptual ambiguity in repetitive indoor environments. Specifically, TextInPlace adopts a dual-branch architecture within a local parameter sharing network. The VPR branch employs attention-based aggregation to extract global descriptors for coarse-grained retrieval, while the STS branch utilizes a bridging text spotter to detect and recognize scene texts. Finally, the discriminative texts are filtered to compute text similarity and re-rank the top-K retrieved images. To bridge the gap between current text-based repetitive indoor scene datasets and the typical scenarios encountered in robot navigation, we establish an indoor VPR benchmark dataset, called Maze-with-Text. Extensive experiments on both custom and public datasets demonstrate that TextInPlace achieves superior performance over existing methods that rely solely on appearance information. The dataset, code, and trained models are publicly available at https://github.com/HqiTao/TextInPlace.
Huaqi Tao, Bingxi Liu 0001, Calvin Chen, Tingjun Huang, Jinqiang Cui, Hong Zhang 0013
IROS6
2025 Open3D-VQA: A Benchmark for Embodied Spatial Concept Reasoning with Multimodal Large Language Model in Open Space
abstract
Spatial reasoning is a fundamental capability of multimodal large language models (MLLMs), yet their performance in open aerial environments remains underexplored. In this work, we present Open3D-VQA, a novel benchmark for evaluating MLLMs' ability to reason about complex spatial relationships from an aerial perspective. The benchmark comprises 73k QA pairs across seven general spatial reasoning tasks, offered in multiple-choice, true/false, and short-answer formats, and supports both visual and point cloud modalities. The questions are automatically generated from spatial relations extracted from both real-world and simulated aerial scenes. Evaluation on 13 popular MLLMs reveals that: 1) Models are generally better at answering questions about relative spatial relations than absolute distances, 2) 3D LLMs fail to demonstrate significant advantages over 2D LLMs, and 3) Fine-tuning solely on the simulated dataset can significantly improve the model's spatial reasoning performance in real-world scenarios. The benchmark, generation pipeline, and evaluation toolkit are released on this page.
Zile Zhou, Xuchen Liu 0001, Jianjie Fang, Chen Gao 0001, Jinqiang Cui, Yong Li 0008, Xinlei Chen, Xiao-Ping Zhang 0002
ACM Multimedia7
2025 Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
abstract
Humans can perceive and reason about spatial relationships from sequential visual observations, such as egocentric video streams. However, how pretrained models acquire such abilities, especially high-level reasoning, remains unclear. This paper introduces Embodied-R, a collaborative framework combining large-scale Vision-Language Models (VLMs) for perception and small-scale Language Models (LMs) for reasoning. Using Reinforcement Learning (RL) with a novel reward system considering think-answer logical consistency, the model achieves slow-thinking capabilities with limited computational resources. After training on only 5k embodied video samples, Embodied-R with a 3B LM matches state-of-the-art multimodal reasoning models (OpenAI-o1, Gemini-2.5-pro) on both in-distribution and out-of-distribution embodied spatial reasoning tasks. Embodied-R also exhibits emergent thinking patterns such as systematic analysis and contextual integration. We further explore research questions including response length, training on VLM, strategies for reward design, and differences in model generalization after SFT (Supervised Fine-Tuning) and RL training. The project page is available at: https://embodiedcity.github.io/Embodied-R/.
Baining Zhao, Ziyou Wang, Jianjie Fang, Chen Gao 0001, Fanhang Man, Jinqiang Cui, Xin Wang 0019, Xinlei Chen, Yong Li 0008, Wenwu Zhu 0001
ACM Multimedia6
2025 Location routing problem with interdependent mobile depot operations for post-disaster relief
Lei Jiao 0005, Zhihong Peng, Shuxin Ding, Jinqiang Cui
Expert Syst. Appl.5
2024 EnYOLO: A Real-Time Framework for Domain-Adaptive Underwater Object Detection with Image Enhancement
abstract
In recent years, significant progress has been made in the field of underwater image enhancement (UIE). However, its practical utility for high-level vision tasks, such as underwater object detection (UOD) in Autonomous Underwater Vehicles (AUVs), remains relatively unexplored. It may be attributed to several factors: (1) Existing methods typically employ UIE as a pre-processing step, which inevitably introduces considerable computational overhead and latency. (2) The process of enhancing images prior to training object detectors may not necessarily yield performance improvements. (3) The complex underwater environments can induce significant domain shifts across different scenarios, seriously deteriorating the UOD performance. To address these challenges, we introduce EnYOLO, an integrated real-time framework designed for simultaneous UIE and UOD with domain-adaptation capability. Specifically, both the UIE and UOD task heads share the same network backbone and utilize a lightweight design. Furthermore, to ensure balanced training for both tasks, we present a multi-stage training strategy aimed at consistently enhancing their performance. Additionally, we propose a novel domain-adaptation strategy to align feature embeddings originating from diverse underwater environments. Comprehensive experiments demonstrate that our framework not only achieves state-of-the-art (SOTA) performance in both UIE and UOD tasks, but also shows superior adaptability when applied to different underwater scenarios. Our efficiency analysis further highlights the substantial potential of our framework for onboard deployment.
Junjie Wen 0001, Jinqiang Cui, Benyun Zhao, Bingxin Han, Xuchen Liu 0001, Zhi Gao 0005, Ben M. Chen
ICRA2
2023 A Convolutional-Transformer Network for Crack Segmentation with Boundary Awareness
abstract
Cracks play a crucial role in assessing the safety and durability of manufactured buildings. However, the long and sharp topological features and complex background of cracks make the task of crack segmentation extremely challenging. In this paper, we propose a novel convolutional-transformer network based on encoder-decoder architecture to solve this challenge. Particularly, we designed a Dilated Residual Block (DRB) and a Boundary Awareness Module (BAM). The DRB pays attention to the local detail of cracks and adjusts the feature dimension for other blocks as needed. And the BAM learns the boundary features from the dilated crack label. Furthermore, the DRB is combined with a lightweight transformer that captures global information to serve as an effective encoder. Experimental results show that the proposed network performs better than state-of-the-art algorithms on two typical datasets. Datasets, code, and trained models are available for research at https://github.com/HqiTao/CT-crackseg.
Huaqi Tao, Bingxi Liu 0001, Jinqiang Cui, Hong Zhang 0013
ICIP3
2023 TJ-FlyingFish: Design and Implementation of an Aerial-Aquatic Quadrotor with Tiltable Propulsion Units
abstract
Aerial-aquatic vehicles are capable to move in the two most dominant fluids, making them more promising for a wide range of applications. We propose a prototype with special designs for propulsion and thruster configuration to cope with the vast differences in the fluid properties of water and air. For propulsion, the operating range is switched for the different mediums by the dual-speed propulsion unit, providing sufficient thrust and also ensuring output efficiency. For thruster configuration, thrust vectoring is realized by the rotation of the propulsion unit around the mount arm, thus enhancing the underwater maneuverability. This paper presents a quadrotor prototype of this concept and the design details and realization in practice.
Xuchen Liu 0001, Minghao Dou, Dongyue Huang, Songqun Gao, Ruixin Yan, Biao Wang 0004, Jinqiang Cui, Qinyuan Ren, LiHua Dou, Zhi Gao 0005, Jie Chen 0003, Ben M. Chen
ICRA7
2023 SyreaNet: A Physically Guided Underwater Image Enhancement Framework Integrating Synthetic and Real Images
abstract
Underwater image enhancement (UIE) is vital for high-level vision-related underwater tasks. Although learning-based UIE methods have made remarkable achievements in recent years, it's still challenging for them to consistently deal with various underwater conditions, which could be caused by: 1) the use of the simplified atmospheric image formation model in UIE may result in severe errors; 2) the network trained solely with synthetic images might have difficulty in generalizing well to real underwater images. In this work, we, for the first time, propose a framework SyreaNet for UIE that integrates both synthetic and real data under the guidance of the revised underwater image formation model and novel domain adaptation (DA) strategies. First, an underwater image synthesis module based on the revised model is proposed. Then, a physically guided disentangled network is designed to predict the clear images by combining both synthetic and real underwater images. The intra- and inter-domain gaps are abridged by fully exchanging the domain knowledge. Extensive experiments demonstrate the superiority of our framework over other state-of-the-art (SOTA) learning-based UIE methods qualitatively and quantitatively. The code and dataset are publicly available at https://github.com/RockWenJJ/SyreaNet.git.
Junjie Wen 0001, Jinqiang Cui, Zhenjun Zhao, Ruixin Yan, Zhi Gao 0005, LiHua Dou, Ben M. Chen
ICRA2
2022 Model-Based Reinforcement Learning with Self-attention Mechanism for Autonomous Driving in Dense Traffic
Junjie Wen 0001, Zuoquan Zhao, Jinqiang Cui, Ben M. Chen
ICONIP (2)3
2022 Design and Analysis of Truss Aerial Transportation System (TATS): The Lightweight Bar Spherical Joint Mechanism
abstract
In aerial cooperative transportation missions, it has been recognized that for small-sized but heavy payloads, the cable-suspended framework is a preferred manner. However, to maintain proper safe flight distances, cables always stay inclined, which implies that horizontal force components have to be generated by UAVs, and only partial thrust forces are used for gravity compensation. To overcome this drawback, in this paper, a new cooperative transportation system named Truss Aerial Transportation System (TATS) is proposed, where those horizontal forces can be internally compensated by the bar spherical joint structure. In the TATS, rigid bars can powerfully sustain the desired distances among UAVs for safe flight, resulting in a more compact and effective transportation system. Thanks to the structural advantage of the truss, the rigid bars can be made lightweight so as to minimize their induced gravity burden. The construction method of the proposed TATS is presented. The improvement in energy efficiency is analyzed and compared with the cable-suspended framework. Furthermore, the robustness property of a TATS configuration is evaluated by computing the margin capacity. Finally, a load test experiment is conducted on our made prototype, the results of which show the effectiveness and feasibility of the proposed TATS.
Qingkai Yang, Delong Wu, Shaozhun Wei, Jinqiang Cui, Hao Fang 0001
IROS6
2022 Few-Shot Scene Classification Using Auxiliary Objectives and Transductive Inference
abstract
Few-shot learning features the capability of generalizing from very few examples. To realize few-shot scene classification of optical remote sensing images, we propose a two-stage framework that first learns a general-purpose representation and then propagates knowledge in a transductive paradigm. Concretely, the first stage jointly learns a semantic class prediction task as well as two auxiliary objectives in a multi-task model. Therein, rotation prediction estimates the 2D transformation of an input, and contrastive prediction aims to pull together the positive pairs while pushing apart the negative pairs. The second stage aims to find an expected prototype having the minimal distance to all samples within the same class. Particularly, label propagation is applied to make joint prediction for both labeled and unlabeled data. Then the labeled set is expanded by those pseudo-labeled samples, thereby forming a rectified prototype to perform nearest-neighbor classification better. Extensive experiments on standard benchmarks including NWPU-RESISC45, AID, and WHU-RS-19 demonstrate that our method works effectively and achieves the best performance that significantly outperforms many state-of-the-art approaches.
Zhi Gao 0005, Can Li 0016, Jinqiang Cui
IEEE Geosci. Remote. Sens. Lett.6
2021 Single Image Deraining Integrating Physics Model and Density-Oriented Conditional GAN Refinement
abstract
Although advanced single image deraining methods have been proposed, their generalization ability to real-world images is usually limited, especially when dealing with rain patterns of different densities, shapes, and directions. In order to improve the robustness and generalization of these deraining methods, we propose a novel density-aware single image deraining method with gated multi-scale feature fusion, which consists of two stages. In the first stage, a sophisticated physics model is leveraged for initial deraining and a network branch is utilized for rain density estimation to guide the subsequent refinement. The second stage of model-independent refinement is realized using conditional Generative Adversarial Network (cGAN), attempting to eliminate artifacts and improve the restoration quality. Extensive experiments have been conducted on the representative synthetic rain datasets and real rain scenes, demonstrating the superiority of our method in terms of effectiveness and generalization ability, which outperforms the state-of-the-arts.
Min Cao 0001, Zhi Gao 0005, Bharath Ramesh 0001, Tiancan Mei, Jinqiang Cui
IEEE Signal Process. Lett.5
2021 A Two-Stage Density-Aware Single Image Deraining Method
abstract
Although advanced single image deraining methods have been proposed, one main challenge remains: the available methods usually perform well on specific rain patterns but can hardly deal with scenarios with dramatically different rain densities, especially when the impacts of rain streaks and the veiling effect caused by rain accumulation are heavily coupled. To tackle this challenge, we propose a two-stage density-aware single image deraining method with gated multi-scale feature fusion. In the first stage, a realistic physics model closer to real rain scenes is leveraged for initial deraining, and a network branch is also trained for rain density estimation to guide the subsequent refinement. The second stage of model-independent refinement is realized using conditional Generative Adversarial Network (cGAN), aiming to eliminate artifacts and improve the restoration quality. In particular, dilated convolutions are applied to extract rain features at multiple scales and gated feature fusion is exploited to better aggregate multi-level contextual information in both stages. Extensive experiments have been conducted on representative synthetic rain datasets and real rain scenes. Quantitative and qualitative results demonstrate the superiority of our method in terms of effectiveness and generalization ability, which outperforms the state-of-the-art.
Min Cao 0001, Zhi Gao 0005, Bharath Ramesh 0001, Tiancan Mei, Jinqiang Cui
IEEE Trans. Image Process.5
2016 3D motion planning for UAVs in GPS-denied unknown forest environment
abstract
In this paper, a decomposition hierarchic on-line motion planning approach consisting of path planning and trajectory generation is proposed for VTOL UAVs to fly in a GPS-denied unknown obstacle-rich environment such as forest and urban canyon. A closed-loop 3D path planning based on A* search algorithm is used to generate collision-free path and a 3D on-line trajectory generation based on maneuver automaton methods is used to generate a collision-free reference trajectory. The simulation and experiment on a VTOL UAV demonstrate the effectiveness of the proposed motion planning approach.
Fang Liao, Shupeng Lai, Yuchao Hu, Jinqiang Cui, Rodney Teo, Feng Lin 0003
Intelligent Vehicles Symposium4
2012 Construction and Modeling of a Variable Collective Pitch Coaxial UAV
Jinqiang Cui, Fei Wang 0018, Zhengyin Qian, Ben M. Chen, Tong Heng Lee
ICINCO (2)1