VLDB 2026 Research / reviewers in the wild / expert
Xinqi Jiang
dblp:223/8091
· DBLP profile ↗
12ranked-venue papers
2as first author
11since 2021 · last 2026
0009-0002-7503-0231ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt GenerationabstractLarge language models (LLMs) are increasingly applied in specialized domains such as finance and healthcare, where they introduce unique safety risks.Domain-specific datasets of harmful prompts remain scarce and still largely rely on manual construction; public datasets mainly focus on explicit harmful prompts, which modern LLM defenses can often detect and refuse.In contrast, implicit harmful prompts-expressed through indirect domain knowledge-are harder to detect and better reflect real-world threats.We identify two challenges: transforming domain knowledge into actionable constraints and increasing the implicitness of generated harmful prompts.To address them, we propose an end-to-end framework that first performs knowledge-graphguided harmful prompt generation to systematically produce domain-relevant prompts, and then applies two-strategy obfuscation rewriting to convert explicit harmful prompts into implicit variants via direct and context-enhanced rewriting.This framework yields high-quality datasets combining strong domain relevance with implicitness, enabling more realistic redteaming and advancing LLM safety research.We release our code and datasets on GitHub. Huawei Zheng, Xinqi Jiang, Sen Yang 0008, Shouling Ji, Yingcai Wu, Dazhen Deng |
ACL (1) | 2 |
| 2026 | Learning a Semantic Similarity Orthogonal Space for Model-Level AI -Generated Image Source AttributionabstractABSTRACT The rapid evolution of diffusion models has established artificial intelligence generated image (AIGI) as a dominant paradigm in digital media, simultaneously escalating the risks of deepfakes and copyright infringement. While existing forensic methods focus primarily on distinguishing real from synthetic content, they fail to address the critical attribution challenge: identifying the specific model architecture responsible for a generated image. To bridge this gap, this paper proposes a novel AIGI source tracing framework capable of pinpointing source models in black‐box scenarios. Our approach is grounded in the hypothesis that images generated by the same model under identical semantic conditions exhibit consistent stylistic signatures and detail‐level artefacts. The framework operates through a ‘reconstruct‐and‐compare’ paradigm. First, we employ a CLIP‐based optimization algorithm to reconstruct the semantic prompt of the query image, enabling the generation of a reference sample from the candidate model. Second, to accurately measure the provenance similarity between the query and the reference, we introduce a Semantic Comparison Network based on Orthogonal Extension. This network utilizes a feature purification mechanism to decouple shared semantic content from model‐specific visual attributes. Furthermore, it employs a multi‐stage orthogonal training strategy to extract mutually independent feature subspaces, ensuring a comprehensive capture of diverse stylistic dimensions. Experimental results demonstrate that this framework effectively overcomes the limitations of current classifier‐based attribution, offering a robust solution for digital forensics and accountable AI governance. Yuchu Lin, Xinqi Jiang, Wei Wang 0077, Jinyu Tian 0001 |
Expert Syst. J. Knowl. Eng. | 2 |
| 2025 | ALCReg: Active Label Correction for Partial Point Cloud RegistrationabstractDeep point cloud registration methods encounter challenges due to partial overlaps and are heavily reliant on labeled data. In this paper, we propose ALCReg, an active label correction method for partial point cloud registration learning. ALCReg utilises a multimodal approach to generate pseudo labels, mitigating the cold-start issue in active learning. To ensure the diversity and representativeness of selected samples, we propose an inlier ratio based query strategy for manual correction. Furthermore, an innovative self-correction mechanism based on consistency is introduced, allowing the model to refine pseudo labels autonomously and further improve model performance. Experimental results on the 3DMatch and 3DLoMatch datasets demonstrate that ALCReg achieves comparable performance with the fully-supervised registration methods, even with only 5% of labeled samples, making it the first active learning method tailored for partial point cloud registration. Code is available at https://github.com/Jiang0903/ALCReg. Zongyi Xu, Xinqi Jiang, Shanshan Zhao 0001, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001 |
ICME | 2 |
| 2025 | CapsDT: Diffusion-Transformer for Capsule Robot ManipulationabstractVision-Language-Action (VLA) models have emerged as a prominent research area, showcasing significant potential across a variety of applications. However, their performance in endoscopy robotics, particularly endoscopy capsule robots that perform actions within the digestive system, remains unexplored. The integration of VLA models into endoscopy robots allows more intuitive and efficient interactions between human operators and medical devices, improving both diagnostic accuracy and treatment outcomes. In this work, we design CapsDT, a Diffusion Transformer model for capsule robot manipulation in the stomach. By processing interleaved visual inputs, and textual instructions, CapsDT can infer corresponding robotic control signals to facilitate endoscopy tasks. In addition, we developed a capsule endoscopy robot system, a capsule robot controlled by a robotic arm-held magnet, addressing different levels of four endoscopy tasks and creating corresponding capsule robot datasets within the stomach simulator. Comprehensive evaluations on various robotic tasks indicate that CapsDT can serve as a robust vision-language generalist, achieving state-of-the-art performance in various levels of endoscopy tasks while achieving a 26.25% success rate in real-world simulation manipulation. Xiting He, Mingwu Su, Xinqi Jiang, Long Bai 0008, Hongliang Ren 0001 |
IROS | 3 |
| 2025 | S2Reg: Structure-semantics collaborative point cloud registration
Zongyi Xu, Xinqi Jiang, Shiyang Cheng 0001, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001 |
Pattern Recognit. | 3 |
| 2025 | Weakly Supervised LiDAR Semantic Segmentation via Scatter Image AnnotationabstractWeakly supervised LiDAR semantic segmentation has made significant strides with limited labeled data. However, most existing methods focus on the network training under weak supervision, while efficient annotation strategies remain largely unexplored. To tackle this gap, we implement LiDAR semantic segmentation using scatter image annotation, effectively integrating an efficient annotation strategy with network training. Specifically, we propose employing scatter images to annotate LiDAR point clouds, combining a pre-trained optical flow estimation network with a foundational image segmentation model to rapidly propagate manual annotations into dense labels for both images and point clouds. Moreover, we propose ScatterNet, a network that includes three pivotal strategies to reduce the performance gap caused by such annotations. First, it utilizes dense semantic labels as supervision for the image branch, alleviating the modality imbalance between point clouds and images. Second, an intermediate fusion branch is proposed to obtain multimodal texture and structural features. Finally, a perception consistency loss is introduced to determine which information needs to be fused and which needs to be discarded during the fusion process. Extensive experiments on the nuScenes and SemanticKITTI datasets demonstrate that our method requires less than 0.02% of the labeled points to achieve over 95% of the performance of fully-supervised methods. Notably, our labeled points are only 5% of those used in the most advanced weakly supervised methods. Zongyi Xu, Xiaoshui Huang, Shanshan Zhao 0001, Xinqi Jiang, Xinbo Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | IGReg: Image-Geometry-Assisted Point Cloud Registration via Selective Correlation FusionabstractPoint cloud registration suffers from repeated patterns and low geometric structures in indoor scenes. The recent transformer utilises attention mechanism to capture the global correlations in feature space and improves the registration performance. However, for indoor scenarios, global correlation loses its advantages as it cannot distinguish real useful features and noise. To address this problem, we propose an image-geometry-assisted point cloud registration method by integrating image information into point features and selectively fusing the geometric consistency with respect to reliable salient areas. Firstly, an Intra-Image-Geometry fusion module is proposed to integrate the texture and structure information into the point feature space by the cross-attention mechanism. Initial corresponding superpoints are acquired as salient anchors in the source and target. Then, a selective correlation fusion module is designed to embed the correlations between the salient anchors and points. During training, the saliency location and selective correlation fusion modules exchange information iteratively to identify the most reliable salient anchors and achieve effective feature fusion. The obtained distinctive point cloud features allow for accurate correspondence matching, leading to the success of indoor point cloud registration. Extensive experiments are conducted on 3DMatch and 3DLoMatch datasets to demonstrate the outstanding performance of the proposed approach compared to the state-of-the-art, particularly in those geometrically challenging cases such as repetitive patterns and low-geometry regions. Zongyi Xu, Xinqi Jiang, Changjun Gu, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Semantic Segmentation Based on Vision Transformer via Interactive AttentionabstractSemantic segmentation is a fundamental task in the computer vision community that aims to achieve pixel-wise classification of images. Convolutional Neural Networks (CNNs) have been the backbone of typical semantic segmentation methods. However, the recent success of the Transformer architecture in natural language processing has led to its application in the field of image semantic segmentation. These methods mainly focus on learning more effective information through the encoder, while paying less attention to the decoder. In this paper, we propose a novel attention-based decoder module called the Attention In Attention (AIA) module. This module employs interactive attention to extract spatial and channel information and dynamically determine feature importance. Additionally, we propose the Feature Position Offset Estimation Module (FPOEM) to mitigate feature misalignment when features of different scales are fused. Experiments on two datasets, Cityscapes and ADE20K, show that the method proposed in this paper achieves state-of-the-art performance. Tao Qiu, Xinqi Jiang, Taiping Zhang |
SMC | 4 |
| 2023 | Object Detection via Multi-Scale Token Based on Vision TransformerabstractVisual transformers have achieved impressive performance on object detection. Traditional transformers only focus on multi-scale features between tokens and tokens. However, these methods do not pay attention to the fine-grained features inside a single token, which can lead to the loss of semantic information in the object detection task. To address this issue, we propose a novel network for the above problem, which consists of three components, (1) Internal Multiscale Token Module (IMTM) focuses on the receptive field size of each token and transforms the token dimension size to effectively extract more multiscale features within the self-attention layer, thereby improving the performance and generalization ability of the model. (2) Differential Filter Module (DFM) uses a convolutional network to focus on high-frequency information in the image, helping the Transformer to learn edge features and establish local context, while improving the model performance through residual connections. (3) Feature Fusion Module (FFM) enhances the local and global information extracted by the network by fusing information from different dimensions. Extensive experiments on PASCAL VOC shows that our proposed method can achieve a state-of-the-art performance on object detection. Tao Qiu, Xinqi Jiang, Zhaowei Shang, Taiping Zhang |
SMC | 3 |
| 2022 | A Novel GAN based on Progressive Growing Transformer with Capsule EmbeddingabstractGenerative Adversarial Networks (GANs) have achieved great improvement after using Convolutional Neural Networks (CNNs) instead of Multi-Layer Perceptrons (MLPs) to build network architecture. Recently, since Transformer architecture has performed well in compute vision, building a Transformer-based image generation network helps solve some of problems caused by CNNs e.g. CNNs-based GANs are difficult to train. On the other hand, the learning of positional encoding in the Transformer structure is often ignored in Transformer-based GANs. Capsule networks are usually considered to be able to learn position information in image features. Therefore, this paper constructs a Progressive Growing Transformer network with Capsule Embedding GAN (PGTCEGAN). The results from the proposed approach are promising with 3.59 FID and 3.92 FID on CelebA and LSUN-Church datasets respectively in image generation task. Xinqi Jiang, Taiping Zhang, Tao Qiu |
SMC | 1 |
| 2022 | Double Feature Pyramid Networks for Classification and Localization on Object DetectionabstractThe decoupled head for classification and localization have been proven powerful in the most of one-stage and two-stage detectors. However, most object detection algorithm share a Feature Pyramid Networks. We perform a thorough analysis about the effectiveness of Feature Pyramid Networks for these two tasks. The decoupled feature pyramid network performs better than the shared network. Going a step further, we found that the two tasks have different preferences for feature pyramid networks. For higher accuracy, we propose a Scene Parsing Pyramid Network for Classification and a Feature Pyramid Transformer Network for Localization. Scene Parsing Pyramid Network exploit the capability of global context information by different region based context aggregation through pyramid pooling module and pyramid attention feature extraction module. Feature Pyramid Transformer Network can capture the suitable contexts of objects residing in different scales. We evaluate our double Feature Pyramid Networks feature pyramid network in the object detection task by integrating it into the FCOS algorithm. The modified algorithm outperforms previous state-of-the-art feature pyramid based methods with a clear margin on both MS-COCO 2017 validation and test datasets. Taiping Zhang, Tao Qiu, Xinqi Jiang |
SMC | 5 |
| 2018 | User Rate and Energy Efficiency of HetNets Based on Poisson Cluster ProcessabstractHeterogeneous cellular networks (HetNets) consist of different tiers of base stations (BSs) to meet the ever-increasing mobile traffic demand. Random deployment of various BSs has mostly been assumed to satisfy a Poisson point process (PPP). However, low power small cells are usually clustered around the popular areas, and PPP does not reflect such a feature. To this end, in this paper, we consider base station (BS) cooperation and analyze user rate and energy efficiency of the HetNets based on a specifical Poisson cluster process (PCP). A calculable formula for spectral efficiency is derived. Simulation results are consistent with the numerical analysis, which confirms the accuracy of theoretical formulas. Xinqi Jiang, Fu-Chun Zheng |
VTC Spring | 1 |