Qian Ma 0003

dblp:49/3103-3 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0001-6473-9523ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multimodal Knowledge Graph Completion via Relation-Aware Negative Sampling with Diffusion-based Interpolation
Qian Ma 0003, Linfei Dai, Zhongming Yao, Yu Gu 0002, Tianyi Li 0005, Christian S. Jensen, Ge Yu 0001
Proc. VLDB Endow.1
2026 Modeling Relational Logic Circuits for and-Inverter Graph Convolutional Network
abstract
The automation of logic circuit design enhances chip performance, energy efficiency, and reliability, and is widely applied in the field of Electronic Design Automation (EDA). And-Inverter Graphs (AIGs) efficiently represent, optimize, and verify the functional characteristics of digital circuits, enhancing the efficiency of EDA development. Due to the complex structure and large scale of nodes in real-world AIGs, accurate modeling is challenging, leading to existing work lacking the ability to jointly model functional and structural characteristics, as well as insufficient dynamic information propagation capability. To address the aforementioned challenges, we propose AIGer, with the aim to enhance the expression of AIGs and thereby improve the efficiency of EDA development. Specifically, AIGer consists of two components: 1) Node logic feature initialization embedding component and 2) AIGs feature learning network component. The node logic feature initialization embedding component projects logic nodes, such as AND and NOT, into independent semantic spaces, to enable effective node embedding for subsequent processing. Building upon this, the AIGs feature learning network component employs a heterogeneous graph convolutional network, designing dynamic relationship weight matrices and differentiated information aggregation approaches to better represent the original structure and information of AIGs. The combination of these two components enhances AIGer’s ability to jointly model functional and structural characteristics and improves its message passing capability, thereby strengthening its expressive power for AIGs and enhancing the development efficiency of logic circuits. Experimental results indicate that AIGer outperforms the current best models in the Signal Probability Prediction (SPP) task, improving MAE and MSE by 18.95% and 44.44%, respectively. In the Truth Table Distance Prediction (TTDP) task, AIGer achieves improvements of 33.57% and 14.79% in MAE and MSE, respectively, compared to the best-performing models1.
Weihao Sun, Shikai Guo, Qian Ma 0003, Hui Li 0014, Yongpeng Weng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2026 Molecular-Driven Multi-View Hypergraph Contrastive Learning for Drug-Drug Interaction Prediction
abstract
Recent concerns have arisen over adverse reactions caused by drug combinations, and drug-drug interaction (DDI) prediction helps identify potential risks by forecasting interactions between drugs. Previous methods have primarily explored drug interactions from the superficial level of drug molecules, often overlooking the internal structural information of the molecules. To this end, we propose Mol-HCL, a multi-view hypergraph contrastive learning framework based on molecular view. In this framework, we construct the molecular view to learn the internal information of drug molecules and, based on this, develop structural view and semantic view. These three views collaboratively learn both intra-molecular and inter-molecular information. Subsequently, we incorporate hypernodes into the structural view and design a novel hyperchain, integrating it into the semantic view to capture latent neighbor drug node structural relationships and long-range DDI chain semantic information. After that, contrastive learning is performed between the structural hypergraph and the molecular view, as well as between the semantic hypergraph and the molecular view, to enhance the representations learned from the molecular view. Finally, we conduct experiments on two real-world scientific datasets. The experimental results demonstrate a significant improvement of Mol-HCL over existing methods, showcasing its effectiveness and advantages in DDI prediction.
Shikai Guo, Hui Li 0014, Qian Ma 0003
IEEE Trans. Comput. Biol. Bioinform.6
2026 EnsDiffAD: Ensemble Diffusion Models for Multivariate Time Series Anomaly Detection
Qian Ma 0003, Yanyang Li, Mei Bai, Xite Wang, Shikai Guo, Yu Gu 0002, Ge Yu 0001
IEEE Trans. Knowl. Data Eng.1
2026 Augmenting Automated Spectrum-Based Multi-Fault Localization via Borderline Confident Learning
abstract
Software fault localization aims to identify faulty elements in a program by analyzing program information and the execution data of test cases. This process plays a crucial role in improving development efficiency, reducing debugging costs, and ensuring reliable software operation. However, in practical scenarios, due to the large scale of programs and the relatively small proportion of faulty statements, the number of failing test cases in a test suite is lower than that of passing test cases to some extent, leading to an imbalance in defect knowledge. Additionally, mutual interference and masking among multiple faults in a program can introduce characterizing noise into the test suite, further increasing the difficulty of fault localization. To address these limitations, we propose a novel method called SA-BCL, which aims to tackle the issues of defect knowledge imbalance and characterizing noise. SA-BCL comprises four key components: the data processing component, the defect knowledge balancing component, the confident learning denoising component, and the program spectrum reducing component. Specifically, the data processing component constructs the program spectrum by analyzing program behaviors and test results during the execution of test cases. The defect knowledge balancing component mitigates the imbalance of defect knowledge by employing boundary identification and oversampling techniques to augment the spectrum corresponding to failing test cases located at the boundary of the test suite. The confident learning denoising component leverages confident learning techniques to identify and eliminate characterizing noise in the program spectrum. The program spectrum reducing component iteratively computes the suspiciousness scores of program elements and reduces the spectrum to facilitate fault localization. We conduct a series of experiments on both the synthetic multi-fault dataset and the real multi-fault dataset, comparing SA-BCL with baseline approaches. The experimental results clearly demonstrate that SA-BCL outperforms the state-of-the-art methods, achieving improvements of up to 37.5% in average wasted effort, 29.2% in precision, and 22.5% in recall. In addition, SA-BCL maintains comparable time costs within the same order of magnitude as baseline methods, and most of the improvements are statistically significant, demonstrating its robustness.
Wenhui Du, Shouyu Yin, Haorui Lin, Furui Zhan, Qian Ma 0003
ACM Trans. Internet Techn.8
2025 LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph Completion
abstract
Multi-modal Knowledge Graph Completion (MMKGC) aims to predict missing entities, relations, or attributes in knowledge graphs by collaboratively modeling the triple structure and multimodal information (e.g., text, images, videos) associated with entities. This approach facilitates the automatic discovery of previously unobserved factual knowledge. However, existing MMKGC methods encounter several critical challenges: (i) the imbalance of inter-entity information across different modalities; (ii) the heterogeneity of intra-entity multimodal information; and (iii) for a given entity, the informational contributions of different modalities are inconsistent across contexts. In this paper, we propose a novel **L**arge model-driven **B**alanced **M**ultimodal **K**nowledge **G**raph **C**ompletion framework, termed LBMKGC. Subsequently, to bridge the semantic gap between heterogeneous modalities, LBMKGC aligns the multimodal embeddings of entities semantically by using the CLIP (Contrastive Language-Image Pre-Training) model. Furthermore, LBMKGC adaptively fuses multimodal embeddings with relational guidance by distinguishing between the perceptual and conceptual attributes of triples. Finally, extensive experiments conducted against 21 state-of-the-art baselines demonstrate that LBMKGC achieves superior performance across diverse datasets and scenarios while maintaining efficiency and generalizability. Our code and data are publicly available at: https://github.com/guoynow/LBMKGC.
Qian Ma 0003, Hui Li 0014, Furui Zhan, Yu Gu 0002, Ge Yu 0001, Shikai Guo
NeurIPS2
2025 ATTD and ATDS detecting abnormal trajectory detection for urban traffic data
Xite Wang, Xiao-Yue Liao, Mei Bai, Qian Ma 0003
Appl. Intell.5
2025 Generative imputation of incomplete images: Leveraging multimodal information for missing pixel
Qian Ma 0003, Jinlei Zhang, Shikai Guo, Bo Ning 0002, Yu Gu 0002, Ge Yu 0001
Inf. Sci.1
2025 AUD-YOLO: A Lightweight Object Detection Algorithm Model Incorporating Dynamic Upsampling for Remote Sensing Images
abstract
In the field of remote sensing image processing, remote sensing image object detection is a crucial undertaking. However, the existing target detection algorithms have a considerable number of model parameters, which results in a slow detection speed that is not conducive to application deployment and real-time detection. Additionally, due to the background complexity of remote sensing images and the large number of small target objects, the performance of existing algorithmic models applied directly with remote sensing images is unsatisfactory. To address the aforementioned issues, this paper proposes an efficient and lightweight architecture based on YOLOv8: AUD-YOLO. In particular, a convolution module, GEConv, based on improved efficient multi-scale attention (EMA), is proposed to replace the original ordinary convolution at the end of the backbone part. This can enhance the accuracy of target detection in remote sensing images while maintaining a lightweight operation. The utilization of dynamic upsampling operators serves to perform upsampling operations with greater accuracy and to more effectively extract image features pertaining to small target objects. Finally, the issue of sample imbalance is addressed by utilising Wise-iou as the position coordinate loss function of the neural network. In the experiments, we validate the effectiveness and robustness of the method using the public remote sensing datasets NWPU VHR-10 and DIOR. The experimental results demonstrate that the AUD-YOLO model proposed in this paper achieves mAP values of 91.1% and 82.7%, which are 1.6% and 1.5% higher than the mAP indexes of the baseline model, and the number of parameters is reduced by 4.7% compared with the original model. We comprehensively consider the number of model parameters and model detection accuracy to make the model more adept at detecting remote sensing images.
Xite Wang, Mei Bai, Qian Ma 0003
IEEE Trans. Geosci. Remote. Sens.4
2025 CAFormer: a connectivity-aware vision transformer for road extraction from remote sensing images
Xite Wang, Changsheng Qin, Mei Bai, Qian Ma 0003
Vis. Comput.4
2024 S_IDS: An efficient skyline query algorithm over incomplete data streams
Mei Bai, Yuxue Han, Xite Wang, Bo Ning 0002, Qian Ma 0003
Data Knowl. Eng.7
2024 Structuring Meaningful Code Review Automation in Developer Community
Zhenzhen Cao, Sijia Lv, Hui Li 0014, Qian Ma 0003, Cheng Guo 0001, Shikai Guo
Eng. Appl. Artif. Intell.5
2023 MIVAE: Multiple Imputation based on Variational Auto-Encoder
Qian Ma 0003, Mei Bai, Xite Wang, Bo Ning 0002
Eng. Appl. Artif. Intell.1
2022 HTD: heterogeneous throughput-driven task scheduling algorithm in MapReduce
Xite Wang, Chaojin Wang, Mei Bai, Qian Ma 0003
Distributed Parallel Databases4
2021 AONet: Active Offset Network for crowd flow prediction
Dafeng Wang, Qian Ma 0003, Naiyao Wang, Xuanzhe Fan, Mingyu Lu, Hongbo Liu 0001
Eng. Appl. Artif. Intell.2
2020 MIDIA: exploring denoising autoencoders for missing data imputation
Qian Ma 0003, Wang-Chien Lee, Tao-Yang Fu, Yu Gu 0002, Ge Yu 0001
Data Min. Knowl. Discov.1
2020 REMIAN: Real-Time and Error-Tolerant Missing Value Imputation
abstract
Missing value (MV) imputation is a critical preprocessing means for data mining. Nevertheless, existing MV imputation methods are mostly designed for batch processing, and thus are not applicable to streaming data, especially those with poor quality. In this article, we propose a framework, called Real-time and Error-tolerant Missing vAlue ImputatioN (REMAIN), to impute MVs in poor-quality streaming data. Instead of imputing MVs based on all the observed data, REMAIN first initializes the MV imputation model based on a-RANSAC which is capable of detecting and rejecting anomalies in an efficient manner, and then incrementally updates the model parameters upon the arrival of new data to support real-time MV imputation. As the correlations among attributes of the data may change over time in unforseenable ways, we devise a deterioration detection mechanism to capture the deterioration of the imputation model to further improve the imputation accuracy. Finally, we conduct an extensive evaluation on the proposed algorithms using real-world and synthetic datasets. Experimental results demonstrate that REMAIN achieves significantly higher imputation accuracy over existing solutions. Meanwhile, REMAIN improves up to one order of magnitude in time cost compared with existing approaches.
Qian Ma 0003, Yu Gu 0002, Wang-Chien Lee, Ge Yu 0001, Hongbo Liu 0001, Xindong Wu 0001
ACM Trans. Knowl. Discov. Data1
2019 Order-Sensitive Imputation for Clustered Missing Values (Extended Abstract)
abstract
To study the issue of missing values (MVs), we propose the Order-Sensitive Imputation for Clustered Missing values (OSICM) framework, in which missing values are imputed sequentially such that the values filled earlier in the process are also used for later imputation of other MVs. Obviously, the order of imputations is critical to the effectiveness and efficiency of OSICM framework. We formulate the searching of the optimal imputation order as an optimization problem, and show its NP-hardness. Furthermore, we devise an algorithm to find the exact optimal solution and propose two approximate/heuristic algorithms to trade off effectiveness for efficiency. Finally, we conduct extensive experiments on real and synthetic datasets to demonstrate the superiority of our OSICM framework.
Qian Ma 0003, Yu Gu 0002, Wang-Chien Lee, Ge Yu 0001
ICDE1
2019 Order-Sensitive Imputation for Clustered Missing Values
abstract
The issue of missing values (MVs) has appeared widely in real-world datasets and hindered the use of many statistical or machine learning algorithms for data analytics due to their incompetence in handling incomplete datasets. To address this issue, several MV imputation algorithms have been developed. However, these approaches do not perform well when most of the incomplete tuples are clustered with each other, coined here as the Clustered Missing Values Phenomenon, which attributes to the lack of sufficient complete tuples near an MV for imputation. In this paper, we propose the Order-Sensitive Imputation for Clustered Missing values (OSICM) framework, in which missing values are imputed sequentially such that the values filled earlier in the process are also used for later imputation of other MVs. Obviously, the order of imputations is critical to the effectiveness and efficiency of OSICM framework. We formulate the searching of the optimal imputation order as an optimization problem, and show its NP-hardness. Furthermore, we devise an algorithm to find the exact optimal solution and propose two approximate/heuristic algorithms to trade off effectiveness for efficiency. Finally, we conduct extensive experiments on real and synthetic datasets to demonstrate the superiority of our OSICM framework.
Qian Ma 0003, Yu Gu 0002, Wang-Chien Lee, Ge Yu 0001
IEEE Trans. Knowl. Data Eng.1
2017 An Effective and Efficient Truth Discovery Framework over Data Streams
Tianyi Li 0005, Yu Gu 0002, Xiangmin Zhou, Qian Ma 0003, Ge Yu 0001
EDBT4