Xing Jia

dblp:283/3213 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Low-Overhead dynamic codebook design for RIS-Aided multiuser MISO systems
Xing Jia, Jiancheng An 0001, Lu Gan 0003, Hongbin Li 0001
Signal Process.1
2025 MEVLERG: Medical Vision-Language Encoder for Report Generation
Songwen Pei, Kai Cong, Xing Jia
IEEE Big Data4
2024 Linking Adaptive Structure Induction and Neuron Filtering: A Spectral Perspective for Aspect-based Sentiment Analysis
abstract
Recently, it has been discovered that incorporating structure information (e.g., dependency trees) can improve the performance of aspect-based sentiment analysis (ABSA). The structure information is often obtained from off-the-shelf parsers, which are sub-optimal and unwieldy. Therefore, adaptively inducing task-specific structures is helpful in resolving this issue. In this work, we concentrate on adaptive graph structure induction for ABSA and explore the impact of neuron-level manipulation from a spectral perspective on structure induction. Specifically, we consider word representations from PLMs (pre-trained language models) as node features and employ a graph learning module to adaptively generate adjacency matrices, followed by graph neural networks (GNNs) to capture both node features and structural information. Meanwhile, we propose the Neuron Filtering (NeuLT), a method to conduct neuron-level manipulations on word representations in the frequency domain. We conduct extensive experiments on three public datasets to observe the impact of NeuLT on structure induction and ABSA. The results and further analysis demonstrate that performing neuron-level manipulation through NeuLT can shorten Aspects-sentiment Distance of induced structures and be beneficial to improve the performance of ABSA. The effects of our method can achieve or come close to SOTA (state-of-the-art) performance.
Maoyi Wang, Yun Xiong, Xing Jia, Zhonglei Guo
LREC/COLING5
2024 Beyond the Known: Novel Class Discovery for Open-World Graph Learning
Yun Xiong, Juncheng Fang, Xixi Wu, Dongxiao He, Xing Jia, Bingchen Zhao, Philip S. Yu
DASFAA (6)6
2024 Dynamic Codebook for Reconfigurable Intelligent Surface-Aided Multiuser MISO Communications
abstract
Reconfigurable intelligent surface (RIS) have emerged as a transformative technology capable of reshaping wireless channels to significantly enhance the efficiency of wireless communication networks in a cost-effective manner. However, the prevailing RIS reflection coefficient optimization scheme presents a significant challenge due to its dependence on channel state information (CSI), which results in excessive pilot overhead and error propagation. To address this issue, this paper proposes a probability update (PU) based dynamic codebook for RIS-aided multiuser multiple-input single-output (MU-MISO) communication systems. Specifically, we implement a learning-from-training strategy that dynamically updates the codebook independently of CSI. This process involves assigning a probability vector to the RIS reflecting elements to generate the codebook, with subsequent iterative updates to the probability vector based on codebook training outcomes. Moreover, numerical results illustrate that the proposed scheme can effectively cater to diverse system by flexibly balancing the training overhead and system performance. Finally, despite channel estimation errors, the proposed scheme outperforms passive beamforming and the existing codebook schemes, while significantly reducing training overhead and implementation complexity.
Xing Jia, Jiancheng An 0001, Xiaoqian Lu, Zhengwu Xu, Lu Gan 0003, Chau Yuen
GLOBECOM1
2024 Stacked Intelligent Metasurface Enabled Near-Field Multiuser Beamfocusing in the Wave Domain
abstract
Intelligent surfaces represent a breakthrough technology capable of customizing the wireless channel cost-effectively. However, the existing works generally focus on planar wavefront, neglecting near-field spherical wavefront characteristics caused by large array aperture and high operation frequencies in the terahertz (THz). Additionally, the single-layer reconfigurable intelligent surface (RIS) lacks the signal processing ability to mitigate the computational complexity at the base station (BS). To address this issue, we introduce a novel stacked intelligent metasurfaces (SIM) comprised of an array of programmable metasurface layers. The SIM aims to substitute conventional digital baseband architecture to execute computing tasks with ultra-low processing delay, albeit with a reduced number of radio-frequency (RF) chains and low-resolution digital-to-analog converters. In this paper, we present a SIM-aided multiuser multiple-input single-output (MU-MISO) near-field system, where the SIM is integrated into the BS to perform beamfocusing in the wave domain and customize an end-to-end channel with minimized inter-user interference. Finally, the numerical results demonstrate that near-field communication achieves superior spatial gain over the far-field, and the SIM effectively suppresses inter-user interference as the wireless signals propagate through it.
Xing Jia, Jiancheng An 0001, Hao Liu 0069, Lu Gan 0003, Marco Di Renzo, Mérouane Debbah, Chau Yuen
VTC Spring1
2024 Two-Stage Registration for Optical and SAR Images With Combined Features and Graph Neural Networks
abstract
The registration of optical and synthetic aperture radar (SAR) images is crucial for multisource remote sensing image analysis and application. Due to different imaging mechanisms, the repeatable key points between optical and SAR images are scarce. Recent works have mainly focused on improving feature detection and description, but the matching strategies still rely on neighboring search, which only considers the visual appearance of key points while neglecting the spatial relationships. To resolve the issue, this letter proposes a two-stage registration algorithm based on combined features and graph neural networks (GNNs). First, to obtain more repeatable points, a convolutional neural network (CNN) is designed by combining detector-free and detector-based strategies, of which the former generates feature points with a grid-like distribution and the latter detects key points with a distinctive appearance. Second, GNNs are utilized for feature matching. The positional information of feature points is embedded into descriptors, and the information from other feature points is aggregated through an attention-based context aggregation mechanism to enrich feature descriptions. Third, a two-stage registration framework is adopted to raise the registration accuracy. Finally, the experimental results show that the proposed method performs excellently for the registration of optical and SAR images, maintaining a high matching success rate (SR) and accuracy under various conditions, including large-scale rotations, scale changes, and even perspective transformations.
Weishuang Wu, Chengchen Ning, Jiao Guo, Tinghao Zhang, Xing Jia
IEEE Geosci. Remote. Sens. Lett.5
2024 DynamiSE: dynamic signed network embedding for link prediction
Haiting Sun, Yun Xiong, Yao Zhang 0009, Yali Xiang, Xing Jia, Haofen Wang
Mach. Learn.6
2023 Plug-and-Play Feature Generation for Few-Shot Medical Image Classification
abstract
Few-shot learning (FSL) presents immense potential in enhancing model generalization and practicality for medical image classification with limited training data; however, it still faces the challenge of severe overfitting in classifier training due to distribution bias caused by the scarce training samples. To address the issue, we propose MedMFG, a flexible and lightweight plug-and-play method designed to generate sufficient class-distinctive features from limited samples. Specifically, MedMFG first re-represents the limited prototypes to assign higher weights for more important information features. Then, the prototypes are variationally generated into abundant effective features. Finally, the generated features and prototypes are together to train a more generalized classifier. Experiments demonstrate that MedMFG outperforms the previous state-of-the-art methods on cross-domain benchmarks involving the transition from natural images to medical images, as well as medical images with different lesions. Notably, our method achieves over 10% performance improvement compared to several baselines. Fusion experiments further validate the adaptability of MedMFG, as it seamlessly integrates into various backbones and baselines, consistently yielding improvements of over 2.9% across all results.
Huifang Du, Xing Jia, Shuyong Gao, Yan Teng 0002, Haofen Wang
BIBM3
2023 Bridging the Gap: Cross-modal Knowledge Driven Network for Radiology Report Generation
abstract
Radiology report generation aims to generate medical reports based on given medical images, which can alleviate the workload of radiologists and has attracted significant research interest in recent years. However, existing studies have struggled to bridge the gap between the two different modalities (i.e. image and text) and generate clinically accurate reports. This is primarily due to the challenges in modelling the crossmodal mappings and the inefficiency of transferring knowledge across modalities. To address these challenges, in this paper, we propose to leverage a pre-constructed knowledge graph as a shared matrix that bridges the gap between visual and textual information, facilitating cross-modal knowledge transfer. This shared knowledge matrix effectively captures cross-modal mappings and aligns information between images and texts, thereby bridging the gap between modalities. Specifically, we propose a new module for knowledge distillation and preservation that integrates relevant knowledge representations into both visual and textual inputs, facilitating intuitive cross-modal knowledge interaction and enhancing the clinical accuracy of the generated reports. Experimental results on two benchmark datasets show the effectiveness of our method, outperforming state-of-the-arts in report generation.
Beichen Kang, Yao Zhang 0009, Yun Xiong, Xing Jia, Jianbo Jiao
BIBM4
2023 DynamiSE: Dynamic Signed Network Embedding for Link Prediction
abstract
In real-world scenarios, dynamic signed networks are ubiquitous where edges have positive and negative sign semantics and evolve over time. Encoding the dynamics and sign semantics of the network simultaneously is challenging. Moreover, over-smoothing is inevitably introduced by the learning of network dynamics. Targeting this gap, we propose Dynamic Signed Network Embedding (DynamiSE), which effectively integrates the balance theory and ordinary differential equation (ODE) into node representation learning to construct a deeper dynamic signed graph neural network and capture the complex sign semantics formed by the two types of edges.
Haiting Sun, Yun Xiong, Yao Zhang 0009, Yali Xiang, Xing Jia, Haofen Wang
DSAA6
2023 Multi-Modal Fusion with Semantic Supervision for Radiology Report Generation
abstract
Radiology report generation, one way of analyzing radiology images, is to generate a textual report automatically for the given image, and it is of great significance to assist diagnosis and alleviate the workload of radiologists. Some report generation methods have been therefore proposed. However, these methods suffer from the problem of low-quality generation, because of the visual and textual bias and training with text similarity oriented objective. To solve this problem, we propose a novel radiology report generation model with multi-modal fusion and semantic supervision, namely MS-Gen. MS-Gen consists of two main components, i.e., the semantic-visual fusion module and the semantic weighted contrastive loss. Specifically, the main idea of the semantic-visual fusion module is to make use of the domain-specific prior knowledge contained in a large pre-trained visual-language model and also the complementary nature between the image and text modalities. Moreover, a novel optimization term, i.e., the semantic weighted contrastive loss, is proposed to guide the optimization process with semantic similarity objective, and further enforce the generated reports with higher clinical accuracy. Extensive experiments conducted on two real datasets of IU X-Ray and MIMIC-CXR demonstrate the effectiveness of MS-Gen.
Xing Jia, Yun Xiong, Yao Zhang 0009
ECAI1
2022 Few-Shot Radiology Report Generation via Knowledge Transfer and Multi-modal Alignment
abstract
Automatic radiology report generation aims at generating informative text from the given medical image, which could assist diagnosis and lighten the workload of radiologists. While some models have been proposed to study on this task, few of them paid attention to the radiology report generation for rare diseases, except for RareGen which solved this problem by enhancing the semantic representations of rare diseases. However, there still exist several problems to be addressed. The first lies in that modeling the correlations among diseases by current studies can result in the problem of frequency bias, which can affect the detection of rare diseases. The second lies in that how to get better representations of disease regions, so as to benefit their corresponding report generation in the decoding stage. To tackle these challenges, we propose a new few-shot radiology report generation model, namely FS-Gen. FS-Gen is assembled with one module for more effective detection of rare diseases in the encoding stage, and the other module for the better representation generation of disease regions in the decoding stage. Specifically, in the encoding stage, a cascade visual enhancement module is proposed to strengthen the correlations among diseases, without incurring the problem of frequency bias. On the other hand, in the decoding stage, a co-referential aligned topic generation module is introduced to simultaneously capture the location and semantic information of disease regions, by aligning the multimodal representations. Extensive experiments are conducted on real-world medical image datasets to demonstrate the effectiveness of our model.
Xing Jia, Yun Xiong, Jiawei Zhang 0001, Yao Zhang 0009, Yangyong Zhu, Philip S. Yu
BIBM1
2021 Radiology Report Generation for Rare Diseases via Few-shot Transformer
abstract
Reliable automatic radiology report generation is highly desired to reduce the labor-intensive and error-prone workload for healthcare workers. While some multi-modal learning models have been proposed to study on this task, few of them paid attention to the radiology report generation for rare diseases, except for RareGen which solved this problem by enhancing the semantic representations of rare diseases. However, there still exist several open problems to be addressed. The first lies in the low proportion of disease regions in an image, making the visual information redundant or irrelevant to rare diseases to be encoded. The second lies in that correlations modeled in the encoding stage may not be effectively decoded in the decoding stage due to the multi-modal representation. To address these two issues, we propose a few-shot Transformer radiology report generation model, namely TransGen, for rare diseases. It integrates the advantages of Transformer with two key modules assembled. Specifically, in the encoding stage, a Semantic-aware Visual Learning (SVL) module is introduced to capture the regions of rare diseases. Following that, in the decoding stage, a Memory Augmented Semantic Enhancement (MASE) module is proposed to enhance intermediate representations. It could make full use of the semantic information contained in the historical-generated sentences to benefit report generation involving rare diseases. Extensive experiments have been conducted on two public datasets of IU X-Ray and MIMIC-CXR to demonstrate the effectiveness of our proposed model.
Xing Jia, Yun Xiong, Jiawei Zhang 0001, Yao Zhang 0009, Suzanne V. Blackley, Yangyong Zhu, Chunlei Tang
BIBM1
2020 Few-shot Radiology Report Generation for Rare Diseases
abstract
Automatic radiology report generation that interprets medical images and writes their diagnostic reports is in high demand, as the manual written-report can be laborintensive and error-prone. By this context so far, some radiology report generation models have been proposed already which can hardly detect rare diseases accurately due to insufficient training data of such diseases. Radiology report generation task is therefore severely challenged while involving the rare disease. To tackle this problem, we propose a few-shot Radiology report Generation model, namely RareGen, assembled with two components for better semantic representations learning which can benefit rare disease detection and their diagnosis report generation. Specifically, a few-shot learning generative network is introduced for generating artificial medical instances for rare diseases. Moreover, a disease graph convolution is proposed to model and strengthen the intrinsic correlations among diseases, which allows knowledge transfer from regular diseases to those rare diseases. To the best of our knowledge, this is the first study that focuses on rare disease diagnosis report generation from radiology data. Extensive experiments are conducted to demonstrate the effectiveness of our model.
Xing Jia, Yun Xiong, Jiawei Zhang 0001, Yao Zhang 0009, Yangyong Zhu
BIBM1