VLDB 2026 Research / reviewers in the wild / expert
Jiaxin Qi
dblp:132/6164
· DBLP profile ↗
11ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0000-7127-7400ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gene Incremental Learning for Single-Cell TranscriptomicsabstractClasses, as fundamental elements of Computer Vision, have been extensively studied within incremental learning frameworks. In contrast, tokens, which play essential roles in many research fields, exhibit similar characteristics of growth, yet investigations into their incremental learning remain significantly scarce. This research gap primarily stems from the holistic nature of tokens in language, which imposes significant challenges on the design of incremental learning frameworks for them. To overcome this obstacle, in this work, we turn to a type of token, gene, for a large-scale biological dataset—single-cell transcriptomics—to formulate a pipeline for gene incremental learning and establish corresponding evaluations. We found that the forgetting problem also exists in gene incremental learning, thus we adapted existing class incremental learning methods to mitigate the forgetting of genes. Through extensive experiments, we demonstrated the soundness of our framework design and evaluations, as well as the effectiveness of the method adaptations. Finally, we provide a complete benchmark for gene incremental learning in single-cell transcriptomics. Jiaxin Qi, Jianqiang Huang 0001, Gaogang Xie |
AAAI | 1 |
| 2025 | A Simple and Comprehensive Benchmark for Single-Cell TranscriptomicsabstractSingle-cell transcriptomics describes complex molecular features at the individual cell level, serving various roles in biological research, such as enhancing gene expression and predicting drug responses. Due to transcriptomic data structurally resembling sequential data, many researchers have trained numerous transformers on extensive transcriptomic datasets. However, they have consistently neglected to explore the intrinsic properties of the data and the appropriateness of their chosen model architecture. In this paper, we carefully investigate the nature of transcriptomics, identifying three overlooked problems: 1) long-tailed data problem, 2) model selection problem, and 3) evaluation problem. Consequently, by applying the weighted sampling strategy, we address the long-tailed data problem and achieve consistent improvement across all settings. By adapting different model structures to transcriptomic data, we discover that transformers are not the only option. By developing three downstream tasks and fair evaluation metrics, we establish a simple and comprehensive benchmark to validate the effectiveness of models for transcriptomics. Through extensive experiments, we clarify the misunderstandings in the traditional methods and provide competitive baselines, thereby paving the way for future research in this field. Jiaxin Qi, Kailei Guo, Jianqiang Huang 0001, Gaogang Xie |
AAAI | 1 |
| 2025 | Graph Neural Networks as a Substitute for Transformers in Single-Cell TranscriptomicsabstractGraph Neural Networks (GNNs) and Transformers share significant similarities in their encoding strategies for interacting with features from nodes of interest, where Transformers use query-key scores and GNNs use edges. Compared to GNNs, which are unable to encode relative positions, Transformers leverage dynamic attention capabilities to better represent relative relationships, thereby becoming the standard backbones in large-scale sequential pre-training. However, the subtle difference prompts us to consider: if positions are no longer crucial, could we substitute Transformers with Graph Neural Networks in some fields such as Single-Cell Transcriptomics? In this paper, we first explore the similarities and differences between GNNs and Transformers, specifically in terms of relative positions. Additionally, we design a synthetic example to illustrate their equivalence where there are no relative positions between tokens in the sample. Finally, we conduct extensive experiments on a large-scale position-agnostic dataset-single-cell transcrip-tomics-finding that GNNs achieve competitive performance compared to Transformers while consuming fewer computation resources. These findings provide novel insights for researchers in the field of single-cell transcriptomics, challenging the prevailing notion that the Transformer is always the optimum choice. Jiaxin Qi, Jinli Ou, Jianqiang Huang 0001 |
BIBM | 1 |
| 2025 | A Comprehensive Benchmark for Electrocardiogram Time-SeriesabstractElectrocardiogram (ECG), a key bioelectrical time-series signal, is crucial for assessing cardiac health and diagnosing various diseases. Given its time-series format, ECG data is often incorporated into pre-training datasets for large-scale time-series model training. However, existing studies often overlook its unique characteristics and specialized downstream applications, which differ significantly from other time-series data, leading to an incomplete understanding of its properties. In this paper, we present an in-depth investigation of ECG signals and establish a comprehensive benchmark, which includes (1) categorizing its downstream applications into four distinct evaluation tasks, (2) identifying limitations in traditional evaluation metrics for ECG analysis, and introducing a novel metric; (3) benchmarking state-of-the-art time-series models and proposing a new architecture. Extensive experiments demonstrate that our proposed benchmark is comprehensive and robust. The results validate the effectiveness of the proposed metric and model architecture, which establish a solid foundation for advancing research in ECG signal analysis. Zhijiang Tang, Jiaxin Qi, Yuhua Zheng, Jianqiang Huang 0001 |
ACM Multimedia | 2 |
| 2024 | Fine-Tuning for Few-Shot Image Classification by Multimodal Prototype RegularizationabstractLarge pre-trained vision-language models, such as CLIP [1], have demonstrated remarkable performance in few shot image classification. To facilitate the rapid adaptation of CLIP in downstream tasks with limited visual samples, two primary frameworks have been proposed. The first framework centers on the image encoder and introduces a trainable visual classifier after the backbone to generate logits for each object class. Nevertheless, this framework heavily depends on limited visual features extracted by the pre-trained visual encoder, which can result in over-fitting issues. The second framework aims to optimize the text encoder by using trainable soft language prompts and computing logits for each class based on the similarity between image features and optimized prompt features. However, this framework encounters the issue of imperfect alignment between the representations extracted by the image and text encoders, making it difficult to fine-tune the language prompts using visual samples. This paper proposes a Multi- Modal Prototype Regularization (MMPR) method for CLIP based few-shot fine-tuning for image classification. MMPR can address the challenges of effectively utilizing both image and text features. MMPR fine-tunes a classifier and regularizes its weights using both image-based (ImgPR) and text-based (TexPR) prototypes. ImgPR represents the mean of image representations within the same class, derived from the image encoder, to distill specific visual distribution knowledge for classifier adaptation. TexPR represents the hand-crafted prompt associated with the class, derived from the text encoder, to incorporate general encyclopedic knowledge and mitigate visual over-fitting. MMPR significantly leverages both image and text information without increasing computational complexity during the inference stage compared to existing methods. Experimental results on various challenging public benchmarks demonstrate the superiority of the proposed MMPR method over state-of-the-art methods. Qianhao Wu, Jiaxin Qi, Hanwang Zhang, Jinhui Tang 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Deconfounded Visual GroundingabstractWe focus on the confounding bias between language and location in the visual grounding pipeline, where we find that the bias is the major visual reasoning bottleneck. For example, the grounding process is usually a trivial languagelocation association without visual reasoning, e.g., grounding any language query containing sheep to the nearly central regions, due to that most queries about sheep have ground-truth locations at the image center. First, we frame the visual grounding pipeline into a causal graph, which shows the causalities among image, query, target location and underlying confounder. Through the causal graph, we know how to break the grounding bottleneck: deconfounded visual grounding. Second, to tackle the challenge that the confounder is unobserved in general, we propose a confounder-agnostic approach called: Referring Expression Deconfounder (RED), to remove the confounding bias. Third, we implement RED as a simple language attention, which can be applied in any grounding method. On popular benchmarks, RED improves various state-of-the-art grounding methods by a significant margin. Code is available at: https://github.com/JianqiangH/Deconfounded_VG. Jianqiang Huang 0001, Jiaxin Qi, Qianru Sun, Hanwang Zhang |
AAAI | 3 |
| 2022 | Class Is Invariant to Context and Vice Versa: On Learning Invariance for Out-Of-Distribution Generalization
Jiaxin Qi, Kaihua Tang, Qianru Sun, Xian-Sheng Hua 0001, Hanwang Zhang |
ECCV (25) | 1 |
| 2022 | Invariant Feature Learning for Generalized Long-Tailed Classification
Kaihua Tang, Mingyuan Tao, Jiaxin Qi, Zhenguang Liu, Hanwang Zhang |
ECCV (24) | 3 |
| 2022 | A topography-aware approach to the automatic generation of urban road networksabstractExisting deep-learning tools for road network generation have limited applications in flat urban areas due to their overreliance on the geometric and spatial configurations of street networks and inadequate considerations of topographic information. This paper proposes a new method of street network generation based on a generative adversarial network by designing a pre-positioned geo-extractor module and a geo-merging bypath. The two improvements employ the complementary use of geometric configurations and topographic features to automate street network generation in both flat and hilly urban areas. Our experiments demonstrate that the improved model yields a more realistic prediction of street configurations than conventional image inpainting techniques. The model’s effectiveness is further enhanced when generating streets in hilly areas. Furthermore, the geo-extractor module provides insights from the computer vision perspective in recognizing when topographic information should be considered and which topographic information should receive more attention. Jiaxin Qi, Lubin Fan, Jianqiang Huang 0001, Ying Jin 0011, Tianren Yang |
Int. J. Geogr. Inf. Sci. | 2 |
| 2020 | Two Causal Principles for Improving Visual DialogabstractThis paper unravels the design tricks adopted by us, the champion team MReaL-BDAI, for Visual Dialog Challenge 2019: two causal principles for improving Visual Dialog (VisDial). By "improving", we mean that they can promote almost every existing VisDial model to the state-of-the-art performance on the leader-board. Such a major improvement is only due to our careful inspection on the causality behind the model and data, finding that the community has overlooked two causalities in VisDial. Intuitively, Principle 1 suggests: we should remove the direct input of the dialog history to the answer model, otherwise a harmful shortcut bias will be introduced; Principle 2 says: there is an unobserved confounder for history, question, and answer, leading to spurious correlations from training data. In particular, to remove the confounder suggested in Principle 2, we propose several causal intervention algorithms, which make the training fundamentally different from the traditional likelihood estimation. Note that the two principles are model-agnostic, so they are applicable in any VisDial model. Jiaxin Qi, Yulei Niu, Jianqiang Huang 0001, Hanwang Zhang |
CVPR | 1 |
| 2020 | "Reading" cities with computer vision: a new multi-spatial scale urban fabric dataset and a novel convolutional neural network solution for urban fabric classification tasksabstractThis paper builds on the proven track record of CNN-based pattern recognition and feature extraction methods, and reports a novel model that classifies urban fabric samples of metropolitan areas in terms of (1) which city they belong to, (2) what types of urban fabric they belong to, and (3) which historic period they originate from. Currently, such tasks require intensive manual work by senior professionals, and even then, inconsistencies and errors occur. Our work is based on a novel urban fabric dataset of four metropolitan areas with distinct typologies (linear development, open block, gated compound, medieval region, irregular grid and orthogonal gird), which consist of high resolution 3-dimensional built form data and hierarchical street networks. The classification model presented in this paper is the first that is capable of predicting the city origin, urban fabric pattern type and construction period. The novelty is also characterised by jointly considering urban fabric features across multiple spatial scales. The experiments demonstrate that this multi-scale approach can capture a wide range of urban fabric features across cities, urban fabric pattern types and development periods. We further find that the effectiveness can be enhanced by appending an auxiliary network for identifying the most appropriate combinations of the multiple spatial scales in line with the classification task. The dataset and model can massively scale up the productivity of researchers and professionals working on cities. Jiaxin Qi, Tianren Yang, Li Wan 0006, Ying Jin 0011 |
SIGSPATIAL/GIS | 2 |