Xinyi Gong

dblp:267/1148 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Time series and sequential data · 46% Vision and language · 39% Language models and text generation · 15%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 32% Bioinformatics and computational biology · 32% Computing education · 28%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language model
1.722025
DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup · ICCV 2025
Bayesian Prompt Flow Learning for Zero-Shot Anomaly Detection · CVPR 2025
Machine learning › Time series and sequential data › anomaly detection
anomaly segmentation
1.622025
DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup · ICCV 2025
VCP-CLIP: A Visual Context Prompting Model for Zero-Shot Anomaly Segmentation · ECCV (69) 2024
Medical and health informatics
drug safety
1.012026
Informative relational learning for adverse reaction prediction with enhanced generalization to novel drugs · Bioinform. 2026
Bioinformatics and computational biology › drug discovery
drug side effect prediction
1.012026
Informative relational learning for adverse reaction prediction with enhanced generalization to novel drugs · Bioinform. 2026
Machine learning › Time series and sequential data
anomaly detection
0.912025
Bayesian Prompt Flow Learning for Zero-Shot Anomaly Detection · CVPR 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs · NeurIPS 2025
Computer vision › Vision and language › vision-language model
prompt learning
0.912025
Bayesian Prompt Flow Learning for Zero-Shot Anomaly Detection · CVPR 2025
Machine learning › Time series and sequential data › anomaly detection
zero-shot anomaly detection
0.912025
Bayesian Prompt Flow Learning for Zero-Shot Anomaly Detection · CVPR 2025
Computing education › STEM education
engineering design education
0.912025
Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs · NeurIPS 2025
Natural language and speech › Language models and text generation › prompt tuning
prompt distribution learning
0.312025
Bayesian Prompt Flow Learning for Zero-Shot Anomaly Detection · CVPR 2025
Computer vision › Vision and language
visual prompting
0.212024
VCP-CLIP: A Visual Context Prompting Model for Zero-Shot Anomaly Segmentation · ECCV (69) 2024

Methods — techniques the papers use, named apart from their topics

simulation-based evaluation · 1.7relational graph convolutional network · 1.0mixture of experts · 1.0knowledge transfer · 1.0conditional domain adversarial network · 1.0self-supervised learning · 0.9residual cross-model attention · 0.9prompt learning · 0.9contrastive learning · 0.9bayesian prompt flow learning · 0.9visual context prompting · 0.8CLIP · 0.8
YearPublicationVenuePosition
2026 Informative relational learning for adverse reaction prediction with enhanced generalization to novel drugs
abstract
MOTIVATION: Accurate prediction of adverse drug reactions (ADRs) is essential for drug safety surveillance, and recent advances in machine learning with heterogeneous biomedical information have improved predictive performance. However, two challenges remain: current methods often learn inadequate ADR representations that fail to capture dependencies among ADRs, and generalize poorly to novel drugs. RESULTS: To obtain informative ADR embeddings, we construct a multi-source, multi-relational ADR graph that integrates hierarchical structure and empirical ADR co-occurrence, and apply a relational graph convolutional network (R-GCN) to learn relation-aware ADR representations. To enhance generalization to novel drugs, we exploit the hierarchical structure of the Anatomical Therapeutic Chemical (ATC) classification to link drugs via shared higher-level categories for effective knowledge transfer and model these relations with an R-GCN. We further introduce a Conditional Domain Adversarial Network (CDAN) to reduce distribution shifts between known and novel drugs by aligning features conditioned on predicted ADR labels, learning domain-invariant yet task-relevant representations. Additionally, to exploit similar ADR patterns among related drugs, we introduce a dual-branch mixture-of-experts (Dual-MoE) module where each expert captures ADR commonalities within a drug category in one branch, while a separate branch models global patterns. Extensive experiments show that our method consistently outperforms seven baselines, achieving F1 improvements of 4.3% and 4.7% over the best baseline on two datasets, respectively, with more balanced precision-recall trade-offs. It also improves AUC on uncommon ADRs by 7% more than on common ADRs, and remains more robust under data sparsity, with more gradual performance degradation as training data decreases. AVAILABILITY AND IMPLEMENTATION: The code of our model is available at https://github.com/fzsdb/Knowledge-guided-ADR-prediction.git.
Shuge Sun, Dalin Zhang 0001, Hongjun Chu, Xinyi Gong
Bioinform.4
2026 ONIR: Object-Noted Tagging for Aerial Image Captioning generation
abstract
Automated captioning for remote sensing imagery often struggles to balance the high descriptive power of large models with the deployment feasibility of smaller ones. To bridge this gap, this paper introduces ONIR, a LLM-efficient, tag-guided framework that empowers compact language models (1-3B parameters) to achieve state-of-the-art captioning accuracy. Specifically, the proposed approach synthesizes a large-scale pseudo-caption dataset by leveraging GPT-4O on existing segmentation benchmarks. Explicit semantic tags are then extracted to train a multi-label Contrastive Language-Image Pre-Training (CLIP) encoder, providing interpretable visual guidance. To maintain parameter efficiency, the architecture incorporates a simple Multilayer Perceptron (MLP) bridge and a two-stage LoRA fine-tuning strategy. Extensive experiments on standard benchmark dataset, such as UCM and Sydney Captions, demonstrate that ONIR significantly outperforms models up to four times its size (7-13B). By combining superior performance with computational efficiency and tag-based controllability, ONIR offers a highly practical solution for real-world remote sensing applications.
Xing Zi, Tengjun Ni, Xianjing Fan, Xian Tao, Xinyi Gong, Jun Li 0010, Ali Braytee, Mukesh Prasad
J. Vis. Commun. Image Represent.5
2026 SAM-IAD: Injecting specific knowledge into SAM for industrial anomaly detection
Yichi Chen 0002, Bin Chen 0022, Weizhi Xian, Xinyi Gong, Jianwen Han, Xian Tao
Knowl. Based Syst.5
2025 Bayesian Prompt Flow Learning for Zero-Shot Anomaly Detection
abstract
Recently, vision-language models (e.g. CLIP) have demonstrated remarkable performance in zero-shot anomaly detection (ZSAD). By leveraging auxiliary data during training, these models can directly perform cross-category anomaly detection on target datasets, such as detecting defects on industrial product surfaces or identifying tumors in organ tissues. Existing approaches typically construct text prompts through either manual design or the optimization of learnable prompt vectors. However, these methods face several challenges: 1) handcrafted prompts require extensive expert knowledge and trial-and-error; 2) single-form learnable prompts struggle to capture complex anomaly semantics; and 3) an unconstrained prompt space limits generalization to unseen categories. To address these issues, we propose Bayesian Prompt Flow Learning (Bayes-PFL), which models the prompt space as a learnable probability distribution from a Bayesian perspective. Specifically, a prompt flow module is designed to learn both image-specific and image-agnostic distributions, which are jointly utilized to regularize the text prompt space and improve the model's generalization on unseen categories. These learned distributions are then sampled to generate diverse text prompts, effectively covering the prompt space. Additionally, a residual cross-model attention (RCA) module is introduced to better align dynamic text embeddings with fine-grained image features. Extensive experiments on 15 industrial and medical datasets demonstrate our method's superior performance. The code is available at https://github.com/xiaozhen228/Bayes-PFL.
Zhen Qu, Xian Tao, Xinyi Gong, Shichen Qu, Qiyu Chen 0002, Zhengtao Zhang, Guiguang Ding
CVPR3
2025 DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup
abstract
Recent vision-language models (e.g., CLIP) have demonstrated remarkable class-generalizable ability to unseen classes in few-shot anomaly segmentation (FSAS), leveraging supervised prompt learning or fine-tuning on seen classes. However, their cross-category generalization largely depends on prior knowledge of real seen anomaly samples. In this paper, we propose a novel framework, namely DictAS, which enables a unified model to detect visual anomalies in unseen object categories without any retraining on the target data, only employing a few normal reference images as visual prompts. The insight behind DictAS is to transfer dictionary lookup capabilities to the FSAS task for unseen classes via self-supervised learning, instead of merely memorizing the normal and abnormal feature patterns from the training set. Specifically, DictAS mainly consists of three components: (1) Dictionary Construction - to simulate the index and content of a real dictionary using features from normal reference images. (2) Dictionary Lookup - to retrieve queried region features from the dictionary via a sparse lookup strategy. When a query feature cannot be retrieved, it is classified as an anomaly. (3) Query Discrimination Regularization - to enhance anomaly discrimination by making abnormal features harder to retrieve from the dictionary. To achieve this, Contrastive Query Constraint and Text Alignment Constraint are further proposed. Extensive experiments on seven public industrial and medical datasets demonstrate that DictAS consistently outperforms state-of-the-art FSAS methods.
Zhen Qu, Xian Tao, Xinyi Gong, Shichen Qu, Fei Shen 0002, Zhengtao Zhang, Mukesh Prasad, Guiguang Ding
ICCV3
2025 BCCIC3: Batch Clause Construction Enhanced Generalization in IC3
Xinyi Gong, Liangze Yin, Ji Wang 0001, Ting Wang 0009
ICFEM2
2025 Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
abstract
Modern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. However, engineering design poses a fundamentally different challenge for large language models (LLMs) compared with traditional textbook-style problem solving or factual question answering. Although existing benchmarks have driven progress in areas such as language understanding, code synthesis, and scientific problem solving, real-world engineering design demands the synthesis of domain knowledge, navigation of complex trade-offs, and management of the tedious processes that consume much of practicing engineers' time. Despite these shared challenges across engineering disciplines, no benchmark currently captures the unique demands of engineering design work. In this work, we introduce EngDesign, an Engineering Design benchmark that evaluates LLMs' abilities to perform practical design tasks across nine engineering domains. Unlike existing benchmarks that focus on factual recall or question answering, EngDesign uniquely emphasizes LLMs' ability to synthesize domain knowledge, reason under constraints, and generate functional, objective-oriented engineering designs. Each task in EngDesign represents a real-world engineering design problem, accompanied by a detailed task description specifying design goals, constraints, and performance requirements. EngDesign pioneers a simulation-based evaluation paradigm that moves beyond textbook knowledge to assess genuine engineering design capabilities and shifts evaluation from static answer checking to dynamic, simulation-driven functional verification, marking a crucial step toward realizing the vision of engineering Artificial General Intelligence (AGI).
Xingang Guo, Xiangyi Kong, Yilan Jiang, Xiayu Zhao, Zhihua Gong, Daixuan Li, Tianle Sang, Beixiao Zhu, Gregory Jun, Yingbing Huang, Yuqi Xue, Rahul Dev Kundu, Qi Jian Lim, Luke Alexander Granger, Mohamed Badr Younis, Darioush Keivan, Nippun Sabharwal, Shreyanka Sinha, Prakhar Agarwal, Kojo Vandyck, Hanlin Mai, Aditya Venkatesh, Ayush Barik, Jiankun Yang, Chongying Yue, Jingjie He, Licheng Xu, Liujun Xu, Rushabh Shetty, Ziheng Guo, Dahui Song, Manvi Jha, Weijie Liang, Weiman Yan, Bryan Zhang, Sahil Bhandary Karnoor, Rutva Pandya, Xinyi Gong, Mithesh Ballae Ganesh, Feize Shi, Ruiling Xu, Yanfeng Ouyang, Lianhui Qin, Elyse Rosenbaum, Corey Snyder, Peter J. Seiler, Geir E. Dullerud, Xiaojia Shelly Zhang, Zuofu Cheng, Pavan Kumar Hanumolu, Mayank Kulkarni, Mahdi Namazifar, Bin Hu 0002
NeurIPS47
2025 HVASR: Enhancing 360-degree video delivery with viewport-aware super resolution
Pingping Dong, Xinyi Gong, Lianming Zhang
Inf. Sci.3
2025 Robust Multi-UAV Cooperative Maritime Object Recognition Under Dynamic Aerial Perspectives via Conflict-Modulated Generative Continual Learning Framework
abstract
Multi-unmanned aerial vehicle (UAV) cooperative maritime object recognition aims to maintain high accuracy under dynamic aerial perspectives for maritime search and rescue missions. Existing continual learning methods retain knowledge by approximating the global distribution of prior data but fail to address cross-perspective knowledge conflicts caused by distribution shifts across aerial perspectives, leading to gradient perturbations that harm consistency and accuracy in dynamic maritime environments. To address the issues, we propose a conflict-modulated generative continual learning (ConMod) framework, comprising generative perspective-robust conflict estimation and conflict-modulated continual learning modules. The generative perspective-robust conflict estimation employs a perspective-aware scene generator that embeds maritime knowledge priors as perspective constraints to augment the data distribution, thereby facilitating explainable cross-perspective conflict association and promoting robust conflict index estimation. It also incorporates a dual-modal conflict index estimator that integrates geometric distortion and environmental variation branches to estimate conflict indices by associating simulated scenes with perspective-robust data distributions. Furthermore, conflict-modulated continual learning introduces a perspective-specific triplet loss to regularize consistency by aligning geometric and environmental features within perspective-specific representation space. Additionally, a conflict-modulated loss treats conflict indices as modulation weights to identify and reinforce representative conflict experiences from a cross-perspective memory buffer, guiding gradient updates toward global optimization. Results on the SeaDronesSee-CL and SeaDronesSee-CL-v2 datasets show that ConMod effectively improves multi-UAV cooperative maritime object recognition by identifying and mitigating cross-perspective knowledge conflicts.
Dalin Zhang 0001, Xinyi Gong
IEEE Trans. Geosci. Remote. Sens.3
2024 CAGK: Collaborative Aspect Graph Enhanced Knowledge-based Recommendation
abstract
Auxiliary information, such as knowledge graph (KG), has become increasingly crucial in recommender systems. However, the current KG-based recommendation still has some limitations: (1) low link rates between items and KG entities, (2) redundant knowledge in KG. In this paper, we introduce the aspect, which refers to keywords describing item attributes in reviews, to KG-based recommendation, and propose a new model, Collaborative Aspect Graph enhanced Knowledge-based Network (CAGK). Firstly, CAGK builds a Collaborative Aspect Graph (CAG) with user-item interactions, aspects and KG, where aspects can fill most of the sparsity. Secondly, we leverage interactive information and aspect features to generate aspect-aware guidance signals to customize knowledge extraction and eliminate redundant knowledge. Lastly, we utilize low ratings and negative aspect sentiment to capture features of that users dislike to prevent repetitive recommendations of disliked items. Experimental results on two widely used benchmark datasets, Amazon-book and Yelp2018, confirm the superiority of CAGK.
Xiaotong Song, Jiatao Zhu, Xinyi Gong
LREC/COLING4
2024 VCP-CLIP: A Visual Context Prompting Model for Zero-Shot Anomaly Segmentation
Zhen Qu, Xian Tao, Mukesh Prasad, Fei Shen 0002, Zhengtao Zhang, Xinyi Gong, Guiguang Ding
ECCV (69)6
2024 ALMRR: Anomaly Localization Mamba on Industrial Textured Surface with Feature Reconstruction and Refinement
Shichen Qu, Xian Tao, Zhen Qu, Xinyi Gong, Zhengtao Zhang, Mukesh Prasad
PRCV (9)4
2021 CADN: A weakly supervised learning-based category-aware object detection network for surface defect detection
Jiabin Zhang, Hu Su, Xinyi Gong, Zhengtao Zhang, Fei Shen 0002
Pattern Recognit.4
2021 Quality Inspection Based on Quadrangular Object Detection for Deep Aperture Component
abstract
This article focuses on automatic inspection for the commonly used component called spring-wire socket. An automatic inspection system is built that adopts an endoscope to improve the imaging quality. To detect the low contrast targets in complex background, we adopt the pipeline of Faster R-CNN but with several improvements. The improved network specifies the targets in the form of quadrangular bounding box as opposed to previous methods that specify them by rectangular bounding box. With the quadrangular representation, additional shape and pose information is provided and unexpected overlaps and background disturbance could be avoided. In the network, an eight-dimensional (8-D) vector is designed to represent the quadrangular bounding box followed by the improved anchor mechanism in the region proposal network. And also, a novel overlap score calculation method is proposed. On the basis of the detection result, rules arisen from expertise are provided to determine the quality of the component in which way automatic quality inspection could be accomplished. The superiority of our detection network over existing ones is sufficiently demonstrated with the detection result of irregular targets in images captured in the industrial scenario. Meanwhile, successful inspection result proves that the system meets the industrial requirements in terms of both accuracy and speed and thus is of practical significance to industrial applications.
Jiabin Zhang, Zhengtao Zhang, Hu Su, Xinyi Gong, Feng Zhang 0006
IEEE Trans. Syst. Man Cybern. Syst.5