VLDB 2026 Research / reviewers in the wild / expert
Dongqi Tang
dblp:246/5788
· DBLP profile ↗
13ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoIT: Automated Image Tagging with Random Perturbation
Xuelin Zhu, Jianshu Li, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao |
Int. J. Comput. Vis. | 4 |
| 2026 | Multi-Label Image Classification via Contrastive Co-Occurrence LearningabstractMulti-label image classification is an essential task in computer vision that aims to identify multiple objects in images. Recently, there has been growing research interest in modeling the relationships between labels to enhance label representation learning. An intuitive approach is to train a network to estimate label co-occurrence probabilities in a supervised manner, which are then leveraged to guide the interactions between label representations. However, the extreme sparsity of label co-occurrence signals poses substantial challenges. To address this issue, we commence by examining the potential interaction behaviors between label representations under the guidance of ground-truth label co-occurrence signals. Inspired by our findings, a novel contrastive learning mechanism is crafted to mimic and enhance such behaviors, facilitating effective label representation interactions without relying on explicit supervision from label co-occurrence signals. Based on this, we develop a pioneering contrastive co-occurrence learning framework, which operates on the instance-level label co-occurrence graph for multi-label image classification. This framework involves sequential processes of label representation learning followed by co-occurrence perception learning. Cross-entropy loss for label classification learning and contrastive loss for co-occurrence perception learning are used to jointly optimize the entire framework end-to-end. In this way, label representations can interact effectively, fully perceiving their co-occurrence relationships at the instance level, thereby significantly improving the performance in label recognition. Extensive experiments on public benchmarks demonstrate the superiority of the proposed framework in multi-label image classification. Codes are available on https://github.com/jasonseu/CoCo. Xuelin Zhu, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao |
IEEE Trans. Image Process. | 3 |
| 2025 | Scalable Autoregressive Monocular Depth EstimationabstractThis paper proposes a new autoregressive model as an effective and scalable monocular depth estimator. Our idea is simple: We tackle the monocular depth estimation (MDE) task with an autoregressive prediction paradigm, based on two core designs. First, our depth autoregressive model (DAR) treats the depth map of different resolutions as a set of tokens, and conducts the low-to-high resolution autoregressive objective with a patch-wise causal mask. Second, our DAR recursively discretizes the entire depth range into more compact intervals, and attains the coarse-to-fine granularity autoregressive objective in an ordinal-regression manner. By coupling these two autoregressive objectives, our DAR establishes new state-of-the-art (SOTA) on KITTI and NYU Depth v2 by clear margins. Further, our scalable approach allows us to scale the model up to 2.0B and achieve the best RMSE of 1.799 on the KITTI dataset (5% improvement) compared to 1.896 by the current SOTA (Depth Anything). DAR further showcases zero-shot generalization ability on unseen datasets. These results suggest that DAR yields superior performance with an autoregressive prediction paradigm, providing a promising approach to equip modern autoregressive large models (e.g., GPT-4o) with depth estimation capabilities. Project page: https://depth-ar.github.io/. Dongqi Tang, Weiqiang Wang 0002, Danny Ziyi Chen, Jintai Chen, Jian Wu 0001 |
CVPR | 3 |
| 2025 | OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLM
Shuo Tong, Dongqi Tang, Weiqiang Wang 0002, Danny Ziyi Chen, Jintai Chen, Jian Wu 0001 |
ICCV | 4 |
| 2025 | TokenPacker: Efficient Visual Projector for Multimodal LLM
Wentong Li 0001, Yuqian Yuan, Jian Liu 0012, Dongqi Tang, Song Wang 0019, Jie Qin 0004, Jianke Zhu, Lei Zhang 0006 |
Int. J. Comput. Vis. | 4 |
| 2025 | Query-Based Knowledge Sharing for Open-Vocabulary Multi-Label ClassificationabstractIdentifying labels that are unseen during training, known as multi-label zero-shot learning, is a non-trivial task in computer vision. Recent studies have increasingly focused on utilizing vision-language pre-training (VLP) models to recognize unseen labels in an open-vocabulary manner. However, these approaches like knowledge distillation have offered only modest performance gains. The challenge of fully harnessing the potential of VLP models for effective multi-label zero-shot learning remains open. In this work, an advanced query-based knowledge sharing framework is proposed to explore the multi-modal knowledge from VLP models for open-vocabulary multi-label classification. Specifically, we introduce a set of label-agnostic query tokens that are designed to capture essential and informative visual knowledge from input images. These tokens are subsequently shared across all labels, allowing them to select pertinent one as visual clues for accurate recognition. Then, by integrating the pre-trained knowledge of VLP models, these query tokens, trained on seen labels, can be efficiently generalized to the recognition of unseen labels. Additionally, we reformulate ranking learning into a form of classification to enable the magnitude of feature vectors for prediction, which significantly benefits label recognition. Experiment results show that our framework outperforms state-of-the-art methods in multi-label zero-shot learning task by a significant margin, reaching 4.2% and 2.4% in mAP on the NUS-WIDE and Open Images datasets, respectively. Code and models are available at https://github.com/jasonseu/QKS . Xuelin Zhu, Dongqi Tang, Jiawei Ge 0002, Bo Liu 0004, Jiuxin Cao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Osprey: Pixel Understanding with Visual Instruction TuningabstractMultimodal large language models (MLLMs) have recently achieved impressive general-purpose vision-language capabilities through visual instruction tuning. However, current MLLMs primarily focus on image-level or box-level understanding, falling short in achieving fine-grained vision-language alignment at pixel level. Besides, the lack of mask-based instruction data limits their ad-vancements. In this paper, we propose Osprey, a mask-text instruction tuning approach, to extend MLLMs by incor-porating fine-grained mask regions into language instruction, aiming at achieving pixel-wise visual understanding. To achieve this goal, we first meticulously curate a mask-based region-text dataset with 724K samples, and then design a vision-language model by injecting pixel-level representation into LLM. Specifically, Osprey adopts a convolutional CLIP backbone as the vision encoder and employs a mask-aware visual extractor to extract precise visual mask features from high resolution input. Experimen-tal results demonstrate Osprey's superiority in various region understanding tasks, showcasing its new capability for pixel-level instruction tuning. In particular, Osprey can be integrated with Segment Anything Model (SAM) seamlessly to obtain multi-granularity semantics. The source code, dataset and demo can be found at https://github.com/CircleRadon/Osprey. Yuqian Yuan, Wentong Li 0001, Jian Liu 0012, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang 0006, Jianke Zhu |
CVPR | 4 |
| 2024 | Semantic-Guided Representation Enhancement for Multi-Label Image ClassificationabstractMulti-label image classification is an essential yet challenging task that requires to recognize multiple objects of images. To this end, recent studies have sought to acquire visual representations for each label by attention models, and then train binary classifiers for prediction. However, these methods have two major drawbacks: 1) They rely heavily on the precise alignments between two modalities, which is still challenging for current attention models; 2) They ignore patch-level representations rich in local object features, which are also of great importance for label recognition. In this paper, we propose a semantic-guided representation enhancement framework, which augments patch-level representations with object-level representations for robust label recognition. Concretely, the proposed framework consists of two significant components: 1) an inter-modal attention module that accounts for coarsely locating object regions and producing object-level representations for each label; 2) an intra-modal attention module that aggregates object representations to enhance patch representations based on their correlations. In this way, both local clues and global glances of objects are fully exploited simultaneously, rather than relying solely on object-level representations obtained by the inter-modal attention, thus improving the performance of label recognition. Experimental results show that our framework outperforms the state-of-the-art methods by 0.5%, 0.6%, 0.7% and 0.8% in mAP on Pascal VOC 2007, Microsoft COCO, NUS-WIDE and Visual Genome datasets, respectively. Codes and models are available on https://github.com/jasonseu/SGRE. Xuelin Zhu, Jianshu Li, Jiuxin Cao, Dongqi Tang, Bo Liu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Label-efficient Segmentation via Affinity PropagationabstractWeakly-supervised segmentation with label-efficient sparse annotations has attracted increasing research attention to reduce the cost of laborious pixel-wise labeling process, while the pairwise affinity modeling techniques play an essential role in this task. Most of the existing approaches focus on using the local appearance kernel to model the neighboring pairwise potentials. However, such a local operation fails to capture the long-range dependencies and ignores the topology of objects. In this work, we formulate the affinity modeling as an affinity propagation process, and propose a local and a global pairwise affinity terms to generate accurate soft pseudo labels. An efficient algorithm is also developed to reduce significantly the computational cost. The proposed approach can be conveniently plugged into existing segmentation networks. Experiments on three typical label-efficient segmentation tasks, i.e. box-supervised instance segmentation, point/scribble-supervised semantic segmentation and CLIP-guided semantic segmentation, demonstrate the superior performance of the proposed approach. Wentong Li 0001, Yuqian Yuan, Song Wang 0019, Wenyu Liu 0005, Dongqi Tang, Jian Liu 0012, Jianke Zhu, Lei Zhang 0006 |
NeurIPS | 5 |
| 2020 | SEE-LPR: A Semantic Segmentation Based End-to-End System for Unconstrained License Plate Detection and Recognition
Dongqi Tang, Ruo-Ze Liu |
MMM (1) | 1 |
| 2020 | An improved clear cell renal cell carcinoma stage prediction model based on gene setsabstractBACKGROUND: Clear cell renal cell carcinoma (ccRCC) is the most common subtype of renal cell carcinoma and accounts for cancer-related deaths. Survival rates are very low when the tumor is discovered in the late-stage. Thus, developing an efficient strategy to stratify patients by the stage of the cancer and inner mechanisms that drive the development and progression of cancers is critical in early prevention and treatment. RESULTS: In this study, we developed new strategies to extract important gene features and trained machine learning-based classifiers to predict stages of ccRCC samples. The novelty of our approach is that (i) We improved the feature preprocessing procedure by binning and coding, and increased the stability of data and robustness of the classification model. (ii) We proposed a joint gene selection algorithm by combining the Fast-Correlation-Based Filter (FCBF) search with the information value, the linear correlation coefficient, and variance inflation factor, and removed irrelevant/redundant features. Then the logistic regression-based feature selection method was used to determine influencing factors. (iii) Classification models were developed using machine learning algorithms. This method is evaluated on RNA expression value of clear cell renal cell carcinoma derived from The Cancer Genome Atlas (TCGA). The results showed that the result on the testing set (accuracy of 81.15% and AUC 0.86) outperformed state-of-the-art models (accuracy of 72.64% and AUC 0.81) and a gene set FJL-set was developed, which contained 23 genes, far less than 64. Furthermore, a gene function analysis was used to explore molecular mechanisms that might affect cancer development. CONCLUSIONS: The results suggested that our model can extract more prognostic information, and is worthy of further investigation and validation in order to understand the progression mechanism. Fangjun Li, Mu Yang, Mingqiang Zhang, Dongfeng Yuan, Dongqi Tang |
BMC Bioinform. | 7 |
| 2019 | GARN: A Novel Generative Adversarial Recognition Network for End-to-End Scene Character RecognitionabstractDeep neural networks have shown their powerful ability in scene character recognition tasks; however, in real life applications, it is often hard to find a large amount of high-quality scene character images for training these networks. In this paper, we proposed a novel end-to-end network named Generative Adversarial Recognition Networks (GARN) for accurate natural scene character recognition in an end-to-end way. The proposed GARN consists of a generation part and a classification part. For the generation part, the purpose is to produce diverse realistic samples to help the classifier overcome the overfitting problem. While in the classification part, a multinomial classifier is trained along with the generator in the form of a game to achieve better character recognition performance. That is, the proposed GARN has the ability to augment scene character data by its generation part and recognize scene characters by its classification part. It is trained in an adversarial way to improve recognition performance. The experimental results on benchmark datasets and the comparisons with the state-of-the-art methods show the effectiveness of the proposed GARN in scene character recognition. Dongqi Tang |
ICDAR | 2 |
| 2019 | Multimodal Image Captioning Through Combining Reinforced Cross Entropy Loss and Stochastic DeprecationabstractRecently, Cross Entropy Loss (CEL) has been proved to be useful in encoder-decoder based multimodal image captioning; however, it still faces the difficulty of inconsistency between optimizing function and evaluation metrics. In this paper, we propose a new approach for multimodal image captioning. It consists of 1) Reinforced Cross Entropy Loss (RCEL) to maximize the probability of ground truth captions and optimize evaluation metrics directly, and 2) Stochastic Deprecation (SD) to automatically select high-quality ground truth sentences without losing the diversity of corpus. The proposed RCEL and SD are generic and can improve the existing natural language generation models while combining them (RCEL-SD) can achieve the best result. Experimental results on the benchmark MSCOCO dataset show that the proposed RCEL-SD respectively outperforms CEL in terms of all the 7 evaluation metrics on three recent image captioning models. Dongqi Tang |
ICME | 3 |