EDBT 2026 Demo / reviewers in the wild / expert
Qingyun Li
dblp:65/10015
· DBLP profile ↗
21ranked-venue papers
6as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using AgentsabstractBowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu, Qiushi Sun, Zehao Li, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Yian Wang, Qingyun Li, Yu Qiao, Zun Wang, Zichen Ding. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kaiming Jin, Zhaoyang Liu 0001, Qiushi Sun, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Yian Wang 0003, Qingyun Li, Yu Qiao 0001, Zun Wang 0001, Zichen Ding 0002 |
ACL (1) | 12 |
| 2025 | EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question AnsweringabstractWe introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text aware questions that reflect real user needs in outdoor driving and indoor house-keeping activities. The questions are designed to elicit identification and reasoning on scene text in an egocentric and dynamic environment. With EgoTextVQA, we comprehensively evaluate 10 prominent multimodal large language models. Currently, all models struggle, and the best results (Gemini 1.5 Pro) are around 33% accuracy, highlighting the severe deficiency of these techniques in egocentric QA assistance. Our further investigations suggest that precise temporal grounding and multi-frame reasoning, along with high resolution and auxiliary scene-text inputs, are key for better performance. With thorough analyses and heuristic suggestions, we hope EgoTextVQA can serve as a solid testbed for research in egocentric scene-text QA assistance. Our dataset is released at: https://github.com/zhousheng97/EgoTextVQA. Junbin Xiao, Qingyun Li, Yicong Li 0004, Xun Yang 0001, Dan Guo 0001, Meng Wang 0001, Tat-Seng Chua, Angela Yao |
CVPR | 3 |
| 2025 | OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with TextabstractImage-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains the capabilities of large language models during multimodal fine-tuning. However, the limited scale and diversity of current image-text interleaved data restrict the development of multimodal large language models. In this paper, we introduce OmniCorpus, a 10 billion-scale image-text interleaved dataset. Using an efficient data engine, we filter and extract large-scale high-quality documents, which contain 8.6 billion images and 1,696 billion text tokens. Compared to counterparts (e.g., MMC4, OBELICS), our dataset 1) has 15 times larger scales while maintaining good data quality; 2) features more diverse sources, including both English and non-English websites as well as video-centric websites; 3) is more flexible, easily degradable from an image-text interleaved format to pure text corpus and image-text pairs. Through comprehensive analysis and experiments, we validate the quality, usability, and effectiveness of the proposed dataset. We hope this could provide a solid data foundation for future multimodal model research. Qingyun Li, Zhe Chen 0017, Weiyun Wang, Wenhai Wang, Shenglong Ye, Zhenjiang Jin, Guanzhou Chen 0004, Yinan He, Zhangwei Gao, Erfei Cui, Jiashuo Yu, Hao Tian 0006, Bin Wang 0065, Xingjian Wei, Wei Li 0320, Wenjian Zhang, Bo Zhang 0069, Pinlong Cai |
ICLR | 1 |
| 2025 | Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuningabstractAdapting pre-trained foundation models for various downstream tasks has been prevalent in artificial intelligence. Due to the vast number of tasks and high costs, adjusting all parameters becomes unfeasible. To mitigate this, several fine-tuning techniques have been developed to update the pre-trained model weights in a more resource-efficient manner, such as through low-rank adjustments. Yet, almost all of these methods focus on linear weights, neglecting the intricacies of parameter spaces in higher dimensions like 4D.
Alternatively, some methods can be adapted for high-dimensional parameter space by compressing changes in the original space into two dimensions and then employing low-rank matrix adaptations. However, these approaches destructs the structural integrity of the involved high-dimensional spaces. To tackle the diversity of dimensional spaces across different foundation models and provide a more precise representation of the changes within these spaces, this paper introduces a generalized parameter-efficient fine-tuning framework, designed for various dimensional parameter space. Specifically, our method asserts that changes in each dimensional parameter space are based on a low-rank core space which maintains the consistent topological structure with the original space. It then models the changes through this core space alongside corresponding weights to reconstruct alterations in the original space. It effectively preserves the structural integrity of the change of original N-dimensional parameter space, meanwhile models it via low-rank tensor adaptation. Extensive experiments on computer vision, natural language processing and multi-modal tasks validate the effectiveness of our method. Chongjie Si, Xue Yang 0005, Zhengqin Xu, Qingyun Li, Jifeng Dai, Yu Qiao 0001, Xiaokang Yang 0001, Wei Shen 0002 |
ICLR | 5 |
| 2025 | Unleashing the Power of LLMs for Medical Video Answer Localization
Junbin Xiao, Qingyun Li, Yusen Yang, Liang Qiu 0002, Angela Yao |
MICCAI (7) | 2 |
| 2025 | InstructSAM: A Training-free Framework for Instruction-Oriented Remote Sensing Object RecognitionabstractLanguage-guided object recognition in remote sensing imagery is crucial for large-scale mapping and automated data annotation. However, existing open-vocabulary and visual grounding methods rely on explicit category cues, limiting their ability to handle complex or implicit queries that require advanced reasoning.
To address this issue, we introduce a new suite of tasks, including Instruction-Oriented Object Counting, Detection, and Segmentation (InstructCDS), covering open-vocabulary, open-ended, and open-subclass scenarios. We further present EarthInstruct, the first InstructCDS benchmark for earth observation. It is constructed from two diverse remote sensing datasets with varying spatial resolutions and annotation rules across 20 categories, necessitating models to interpret dataset-specific instructions.
Given the scarcity of semantically rich labeled data in remote sensing, we propose InstructSAM, a training-free framework for instruction-driven object recognition. InstructSAM leverages large vision-language models to interpret user instructions and estimate object counts, employs SAM2 for mask proposal, and formulates mask-label assignment as a binary integer programming problem. By integrating semantic similarity with counting constraints, InstructSAM efficiently assigns categories to predicted masks without relying on confidence thresholds. Experiments demonstrate that InstructSAM matches or surpasses specialized baselines across multiple tasks while maintaining near-constant inference time regardless of object count, reducing output tokens by 89\% and overall runtime by over 32\% compared to direct generation approaches. We believe the contributions of the proposed tasks, benchmark, and effective approach will advance future research in developing versatile object recognition systems. The code is available at https://VoyagerXvoyagerx.github.io/InstructSAM. Yijie Zheng, Weijie Wu, Qingyun Li, Aiai Ren |
NeurIPS | 3 |
| 2025 | Precision strike: Precise backdoor attack with dynamic trigger
Qingyun Li, Wei Chen 0006, Xiaotang Xu, Lifa Wu |
Comput. Secur. | 1 |
| 2025 | PointOBB-v3: Expanding Performance Boundaries of Single Point-Supervised Oriented Object Detection
Peiyuan Zhang, Xue Yang 0005, Yi Yu 0010, Qingyun Li, Yue Zhou 0005, Xiaosong Jia, Jingdong Chen, Xiang Li 0041, Junchi Yan, Yansheng Li 0001 |
Int. J. Comput. Vis. | 5 |
| 2024 | PointOBB: Learning Oriented Object Detection via Single Point SupervisionabstractSingle point-supervised object detection is gaining attention due to its cost-effectiveness. However, existing approaches focus on generating horizontal bounding boxes (HBBs) while ignoring oriented bounding boxes (OBBs) commonly used for objects in aerial images. This paper proposes PointOBB, the first single Point-based OBB generation method, for oriented object detection. PointOBB operates through the collaborative utilization of three distinctive views: an original view, a resized view, and a ro-tatedlflipped (rot/flp) view. Upon the original view, we leverage the resized and rot/flp views to build a scale augmentation module and an angle acquisition module, respectively. In the former module, a Scale-Sensitive Consistency (SSC) loss is designed to enhance the deep network's ability to perceive the object scale. For accurate object angle predictions, the latter module incorporates self-supervised learning to predict angles, which is associated with a scale-guided Dense-to-Sparse (DS) matching strategy for aggre-gating dense angles corresponding to sparse objects. The resized and rot/flp views are switched using a progressive multi- view switching strategy during training to achieve coupled optimization of scale and angle. Experimental re-sults on the DIOR-R and DOTA-v1.0 datasets demonstrate that PointOBB achieves promising performance, and significantly outperforms potential point-supervised baselines. Xue Yang 0005, Yi Yu 0010, Qingyun Li, Junchi Yan, Yansheng Li 0001 |
CVPR | 4 |
| 2024 | Point2RBox: Combine Knowledge from Synthetic Visual Patterns for End-to-End Oriented Object Detection with Single Point SupervisionabstractWith the rapidly increasing demand for oriented object detection (OOD), recent research involving weakly-supervised detectors for learning rotated box (RBox) from the horizontal box (HBox) has attracted more and more attention. In this paper, we explore a more challenging yet label-efficient setting, namely single point-supervised OOD, and present our approach called Point2RBox. Specifically, we propose to leverage two principles: 1) Synthetic pattern knowledge combination: By sampling around each labeled point on the image, we spread the object feature to synthetic visual patterns with known boxes to provide the knowledge for box regression. 2) Transform self-supervision: With a transformed input image (e.g. scaled/rotated), the output RBoxes are trained to follow the same transformation so that the network can perceive the relative size/rotation between objects. The detector is further enhanced by a few devised techniques to cope with peripheral issues, e.g. The anchor/layer assignment as the size of the object is not available in our point supervision setting. To our best knowledge, Point2RBox is the first end-to-end solution for point-supervised OOD. In particular, our method uses a lightweight paradigm, yet it achieves a competitive performance among point-supervised alternatives, 41.05%/27.62%/80.01% on DOTA/DIOR/HRSC datasets. Yi Yu 0010, Xue Yang 0005, Qingyun Li, Feipeng Da, Jifeng Dai, Yu Qiao 0001, Junchi Yan |
CVPR | 3 |
| 2024 | The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
Weiyun Wang, Yiming Ren 0001, Haowen Luo, Tiantong Li, Chenxiang Yan, Zhe Chen 0017, Wenhai Wang, Qingyun Li, Lewei Lu, Xizhou Zhu, Yu Qiao 0001, Jifeng Dai |
ECCV (33) | 8 |
| 2024 | The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open WorldabstractWe present the All-Seeing (AS) project: a large-scale dataset and model for recognizing and understanding everything in the open world.
Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a new dataset (AS-1B) with over 1.2 billion regions annotated with semantic tags, question-answering pairs, and detailed captions. It covers a wide range of 3.5 million common and rare concepts in the real world and has 132.2 billion tokens that describe the concepts and their attributes. Leveraging this new dataset, we develop the All-Seeing model (ASM), a unified framework for panoptic visual recognition and understanding. The model is trained with open-ended language prompts and locations, which allows it to generalize to various vision and language tasks with remarkable zero-shot performance, including both region- and image-level retrieval, region recognition, captioning, and question-answering. We hope that this project can serve as a foundation for vision-language artificial general intelligence research. Code is available at https://github.com/OpenGVLab/all-seeing. Weiyun Wang, Min Shi 0004, Qingyun Li, Wenhai Wang, Zhenhang Huang, Linjie Xing, Zhe Chen 0017, Hao Li 0069, Xizhou Zhu, Zhiguo Cao 0001, Tong Lu 0002, Jifeng Dai, Yu Qiao 0001 |
ICLR | 3 |
| 2024 | Co-Training Transformer for Remote Sensing Image Classification, Segmentation, and DetectionabstractSeveral fundamental remote sensing (RS) image processing tasks, including classification, segmentation, and detection, have been set to serve for manifold applications. In the RS community, the individual tasks have been studied separately for many years. However, the specialized models were only capable of a single task. They lacked the adaptability for generalizing to the other tasks. Moreover, Transformer exhibits a powerful generalization capacity because it has the property of dynamic feature weighting. Hence, there is a large potential of a uniform Transformer to learn multiple tasks simultaneously, i.e., multi-task learning (MTL). An MTL Transformer can combine knowledge from different tasks by sharing a uniform network. In this study, a general-purpose Transformer, which simultaneously processes the three tasks, is investigated for RS MTL. To build a Transformer capable of the three tasks, an MTL framework named RSCoTr is proposed. The framework uses a shared encoder to extract multi-scale features efficiently and three task-specific decoders to obtain different results. Moreover, a flexible training procedure named co-training is proposed. The MTL model is trained with multiple general data sets annotated for individual tasks. The co-training is as easy as training a specialized model for a single task. It can be developed into different learning strategies to meet various requirements. The proposed RSCoTr is trained jointly with various strategies on three challenging data sets of the three tasks. And the results demonstrate that the proposed MTL method achieves state-of-the-art performance in comparison with other competitive approaches. Code will be available at https://github.com/Li-Qingyun/RSCoTr. Qingyun Li, Yushi Chen 0002, Xin He 0004, Lingbo Huang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | ARS-DETR: Aspect Ratio-Sensitive Detection Transformer for Aerial Oriented Object DetectionabstractExisting oriented object detection in aerial images has progressed a lot in recent years and achieved a favorable success. However, high-precision oriented object detection in aerial images remains a challenging task. Some recent works have adopted the classification-based method to predict the angle in order to address boundary problem in angle. However, we have found that these works often neglect the sensitivity of objects with different aspect ratios to angle. At the same time, it is worth exploring a suitable way to improve the emerging transformer-based approaches in order to adapt them to oriented object detection. In this paper, we propose an Aspect Ratio Sensitive DEtection TRansformer, termed ARS-DETR, for oriented object detection in aerial images. Specifically, a new angle classification method, called Aspect Ratio aware Circle Smooth Label (AR-CSL), is proposed to smooth the angle label in a more reasonable way and discard the hyperparameter that introduced by previous work (e.g. CSL). Then, a rotated deformable attention module is designed to rotate the sampling points with the corresponding angles and eliminate the misalignment between region features and sampling points. Moreover, a dynamic weight coefficient according to the aspect ratio is adopted to calculate the angle loss. Comprehensive experiments on several challenging datasets demonstrate that our method achieves a competitive performance in the high-precision oriented object detection task. Yushi Chen 0002, Xue Yang 0005, Qingyun Li, Junchi Yan |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | H2RBox-v2: Incorporating Symmetry for Boosting Horizontal Box Supervised Oriented Object DetectionabstractWith the rapidly increasing demand for oriented object detection, e.g. in autonomous driving and remote sensing, the recently proposed paradigm involving weakly-supervised detector H2RBox for learning rotated box (RBox) from the more readily-available horizontal box (HBox) has shown promise. This paper presents H2RBox-v2, to further bridge the gap between HBox-supervised and RBox-supervised oriented object detection. Specifically, we propose to leverage the reflection symmetry via flip and rotate consistencies, using a weakly-supervised network branch similar to H2RBox, together with a novel self-supervised branch that learns orientations from the symmetry inherent in visual objects. The detector is further stabilized and enhanced by practical techniques to cope with peripheral issues e.g. angular periodicity. To our best knowledge, H2RBox-v2 is the first symmetry-aware self-supervised paradigm for oriented object detection. In particular, our method shows less susceptibility to low-quality annotation and insufficient training data compared to H2RBox. Specifically, H2RBox-v2 achieves very close performance to a rotation annotation trained counterpart -- Rotated FCOS: 1) DOTA-v1.0/1.5/2.0: 72.31%/64.76%/50.33% vs. 72.44%/64.53%/51.77%; 2) HRSC: 89.66% vs. 88.99%; 3) FAIR1M: 42.27% vs. 41.25%. Yi Yu 0010, Xue Yang 0005, Qingyun Li, Yue Zhou 0005, Feipeng Da, Junchi Yan |
NeurIPS | 3 |
| 2022 | Two-Branch Pure Transformer for Hyperspectral Image ClassificationabstractOwing to its capability of building long-range dependencies and global context connections, Transformer has been used for hyperspectral image (HSI) classification. However, most of the existing Transformer-based HSI spatial–spectral classification methods consist of a convolutional neural network (CNN) and Transformer, which are used to extract the local and global information, respectively. In this study, to fully explore the potential of Transformer, a pure Transformer is investigated for HSI classification. First, a spatial Transformer (Spa-TR) is designed for HSI spatial classification, which learns the spatial features locally and globally by adopting the window partition and shifted window schemes. Especially, the self-attention computations are limited within the local windows and cross-windows. Second, to fully use the abundant spectral information in HSIs, a two-branch pure Transformer (i.e., Spa-Spe-TR) is proposed, which includes a spectral Transformer (Spe-TR) and a Spa-TR. The spectral sequence features learned by Spe-TR and the spatial features generated by Spa-TR are effectively fused with a branch fusion strategy, which explicitly and automatically measures the importance between the joint spatial–spectral features and improves the discriminability of the joint features. Experimental results on the two widely used HSI datasets (i.e., Pavia and Indian Pines) demonstrate the efficacy of the proposed methods in comparison with other state-of-the-art approaches. Xin He 0004, Yushi Chen 0002, Qingyun Li |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Complementary Learning-Based Scene Classification of Remote Sensing Images With Noisy LabelsabstractRecently, many deep convolutional neural network (DCNN)-based methods have been proposed for remote sensing (RS) image scene classification (SC). In general, DCNNs obtain good generalization capabilities under the condition of correct labels. Unfortunately, the given samples are sometimes mislabeled. In this letter, the classification of RS images with noisy labels is investigated. First, complementary learning (CL), which learns from complementary labels rather than the original labels, is introduced for RS image classification with noisy labels. CL can decrease the probability of learning from incorrect information, and therefore, it is robust to noisy labels. Then, soft CL, which randomly disturbs the complementary labels of the training samples, is proposed to prevent the overfitting issue in training a DCNN. Moreover, an RS image scene classification framework combining ordinary learning (OL) and CL (RS-COCL) is proposed, which uses CL to obtain a good model and OL to fine-tune the deep model. Additionally, noisy labels filtering is used in RS-COCL (RS-COCL-NLF) to detected and corrected noisy samples. At last, soft CL is used in RS-COCL-NLF to obtain better classification performance. The proposed methods are tested on two widely used datasets (i.e., Northwestern Polytechnical University (NWPU)-RESISC45 and PatternNet) and the obtained results show that the proposed methods provide competitive classification accuracy compared to the state-of-the-art methods. Qingyun Li, Yushi Chen 0002, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | DHQN: a Stable Approach to Remove Target Network from Deep Q-learning NetworkabstractAs the first successful attempt to combine deep neural network and reinforcement learning, Deep Q-learning Network (DQN) draws a lot of attention from reinforcement learning researchers. One of the most important components of DQN is target network, which is used to stabilize learning process. When confront complex network structure, the existence of target network means extra memory resource to preserve the neural network weights and high computing cost to calculate target. Thus, we propose a Deep Hybrid Q-learning Network (DHQN) algorithm, which introduces an alternative approach, Random Hybrid Optimization (RHO), that can simplify DQN and attain a more stable and faster learning without a target network. We illustrate that RHO can decelerate divergence in the classical off-policy counterexample θ → 2θ problem. We also testify the effectiveness of DHQN in several control and Atari domains, which shows DHQN outperforms DQN without a target network and original DQN. Guang Yang 0006, Di'an Fei, Tian Huang, Qingyun Li, Xingguo Chen |
ICTAI | 5 |
| 2021 | A Random Forest Classification Algorithm Based Personal Thermal Sensation Model for Personalized Conditioning System in Office BuildingsabstractAbstract The personal thermal sensation model is used as the main component for personalized conditioning system, which is an effective method to fulfill thermal comfort requirements of the occupants, considering the energy consumption. The Random Forest classification algorithm based thermal sensation model is developed in this study, which combines indoor air quality parameters, personal information, physiological factors and occupancy preferences on selection of 7-level of sensation: cold, cool, slightly cool, neutral, slightly warm, warm and hot. Our model shows better functionality, as well as performance and factor selection. As a result, our method has achieved 70.2% accuracy, comparing with the 57.4% accuracy of support vector machine, and 67.7% accuracy of neutral network in an ASHRAE RP-884 database. Therefore, our newly developed model can be used in personalized thermal adjustment systems with intelligent control functions. Qingyun Li |
Comput. J. | 1 |
| 2016 | Achievable degrees of freedom of MIMO two-way X relay channel with delayed CSIT
Qingyun Li, Gang Wu 0001, Hongxiang Li 0001, Shaoqian Li |
Sci. China Inf. Sci. | 1 |
| 2014 | Adaptive and fast target detection in high-resolution SAR imageabstractIn this paper, a new adaptive and fast Constant false alarm rate (CFAR) target detection algorithm based on two level CFAR (TL-CFAR) detectors in high-resolution synthetic aperture radar (SAR) images is proposed. In the first level, the initial mask of targets is obtained by Cell Averaging CFAR (CA-CFAR) detector. In the second level, the precise parameters estimation of CFAR in the local window is implemented by removing those pixels that may belong to the neighboring targets which is identified from the first detector. The problem of high computational complexity of two levels CFAR detector is mitigated by introducing the integral image. Real SAR image data is used to verify the effectiveness of the proposed algorithm, and the results indicate that this algorithm can detect targets fast and precisely. Yihua Tan, Airong Sun, Qingyun Li |
IGARSS | 4 |