Shuyan Li

dblp:12/3189 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semantic-Assisted Object Clustering for Multi-Modal Referring Video Segmentation
abstract
This paper concentrates on Multi-modal Referring Video Segmentation task, where a well optimized model is able to recognize and segment the target objects referred by the given guidance signals, e.g., language description. Early approaches model this task as a sequence prediction problem. The lack of a global view of video content leads to difficulties in effectively utilizing inter-frame relationships. Some recent works propose to perform temporal modeling with vanilla attention mechanism. However, the condensed visual representation tends to be messy about target information due to occlusion or motion blur. Unlimited non-local operation would spread such noise to all the sequences and interfere with the extraction of global representations. To address the above issue, we present Semantic-assisted Object Cluster network (SOC) and the improved SOC++ in this paper. Our method unifies temporally selective interaction and cross-modal alignment to achieve video-level understanding. In SOC++, a proxy-assisted multi-modal fusion module is introduced to perform preliminary bidirectional activation. Then a semantic integration module with progressive frame-to-video structure facilitates joint space learning across modalities and time steps. Considering that potential noisy visual embeddings would impair the overall representation of target objects in unconstrained inter-frame interactions, we propose to perform tendentious video aggregation through emphasizing the indicative role of the informative frames with lower entropy in this part. A multi-modal query contrastive supervision is also utilized to help construct well-aligned joint space at the video level. Moreover, to integrate the advantage of high-level video information and the low-level details of each frame, we introduce a dynamic query fusion module that performs joint updating of these embeddings. We conduct extensive experiments on popular referring video segmentation benchmarks, and our method outperforms state-of-the-art competitors on all benchmarks by a remarkable margin. Besides, the emphasis on temporal coherence enhances the segmentation stability and adaptability of our method in processing text expressions with temporal variations..
Yong Liu 0033, Zhuoyan Luo, Yicheng Xiao, Shuyan Li, Xiu Li 0001, Yujiu Yang 0001, Yansong Tang
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 EventVAD: Training-Free Event-Aware Video Anomaly Detection
abstract
Video Anomaly Detection (VAD) focuses on identifying anomalies within videos. Supervised methods require an amount of in-domain training data and often struggle to generalize to unseen anomalies. In contrast, training-free methods leverage the intrinsic world knowledge of large language models (LLMs) to detect anomalies but face challenges in localizing fine-grained visual transitions and diverse events. Therefore, we propose EventVAD, an event-aware video anomaly detection framework that combines tailored dynamic graph architectures and multimodal LLMs to perform fine-grained temporal-event reasoning. Specifically, EventVAD first employs dynamic spatiotemporal graph modeling with time-decay constraints to capture event-aware video features. Then, it performs adaptive noise filtering and uses signal ratio thresholding to detect event boundaries via unsupervised statistical features. Finally, it utilizes a hierarchical prompting strategy to guide MLLMs in performing reasoning and making final decisions. We conducted extensive experiments on the UCF-Crime and XD-Violence datasets. The results demonstrate that EventVAD with a 7B MLLM achieves state-of-the-art (SOTA) in training-free settings, outperforming strong baselines that use 7B or larger MLLMs. The code is available at https://github.com/YihuaJerry/EventVAD.
Yihua Shao, Haojin He, Siyu Chen 0021, Xinwei Long, Fanhu Zeng, Yuxuan Fan, Muyang Zhang, Ziyang Yan, Ao Ma 0005, Hao Tang 0005, Yan Wang 0105, Shuyan Li
ACM Multimedia14
2025 Three trustworthiness challenges in large language model-based financial systems: real-world examples and mitigation strategies
abstract
大语言模型 (LLM) 在金融应用中的集成展现出显著潜力, 可提升决策流程、实现操作自动化并提供个性化服务。 然而, 金融系统的高风险特性要求极高的可信度, 而当前LLM往往难以满足这一要求。 本研究识别并探讨了基于LLM的金融系统中的3大可信度挑战: (1) 逃逸式提示——利用模型对齐漏洞生成有害或违规响应; (2) 幻觉现象——模型产出事实错误的输出误导金融决策; (3) 偏见与公平性问题——LLM内嵌的人口统计或制度偏见可能导致个体或区域遭受不公平对待。 为具体呈现这些风险, 我们设计了3项金融相关测试, 并对涵盖专有与开源家族的主流LLM进行评估。 在所有模型中, 每项测试至少出现一次风险行为。 基于这些发现, 系统性地总结了现有风险缓解策略。 我们认为, 解决这些问题不仅对确保金融领域人工智能的负责任使用至关重要, 更是实现其安全可扩展部署的关键所在。
Shurui Xu, Shuyan Li, Mengzhen Fan, Zhongtian Sun
Frontiers Inf. Technol. Electron. Eng.3
2025 Construction of a novel five-dimensional Hamiltonian conservative hyperchaotic system and its application in image encryption
Minxiu Yan, Shuyan Li
Soft Comput.2
2024 AV-GAN: Attention-Based Varifocal Generative Adversarial Network for Uneven Medical Image Translation
abstract
Different types of staining highlight different structures in organs, thereby assisting in diagnosis. However, due to the impossibility of repeated staining, we cannot obtain different types of stained slides of the same tissue area. Translating the slide that is easy to obtain (e.g., H&E) to slides of staining types difficult to obtain (e.g., MT, PAS) is a promising way to solve this problem. However, some regions are closely connected to other regions, and to maintain this connection, they often have complex structures and are difficult to translate, which may lead to wrong translations. In this paper, we propose the Attention-Based Varifocal Generative Adversarial Network (AV-GAN), which solves multiple problems in pathologic image translation tasks, such as uneven translation difficulty in different regions, mutual interference of multiple resolution information, and nuclear deformation. Specifically, we develop an Attention-Based Key Region Selection Module, which can attend to regions with higher translation difficulty. We then develop a Varifocal Module to translate these regions at multiple resolutions. Experimental results show that our proposed AV-GAN outperforms existing image translation methods with two virtual kidney tissue staining tasks and improves FID values by 15.9 and 4.16 respectively in the H&E-MT and H&E-PAS tasks.
Yiyang Lin, Zijie Fang, Shuyan Li, Xiu Li 0001
IJCNN4
2024 Financial FAQ Question-Answering System Based on Question Semantic Similarity
Wenxing Hong, Shuyan Li
KSEM (3)3
2023 Adversarial Alignment for Source Free Object Detection
abstract
Source-free object detection (SFOD) aims to transfer a detector pre-trained on a label-rich source domain to an unlabeled target domain without seeing source data. While most existing SFOD methods generate pseudo labels via a source-pretrained model to guide training, these pseudo labels usually contain high noises due to heavy domain discrepancy. In order to obtain better pseudo supervisions, we divide the target domain into source-similar and source-dissimilar parts and align them in the feature space by adversarial learning.Specifically, we design a detection variance-based criterion to divide the target domain. This criterion is motivated by a finding that larger detection variances denote higher recall and larger similarity to the source domain. Then we incorporate an adversarial module into a mean teacher framework to drive the feature spaces of these two subsets indistinguishable. Extensive experiments on multiple cross-domain object detection datasets demonstrate that our proposed method consistently outperforms the compared SFOD methods. Our implementation is available at https://github.com/ChuQiaosong.
Qiaosong Chu, Shuyan Li, Guangyi Chen 0002, Kai Li 0012, Xiu Li 0001
AAAI2
2023 SSGD: A Smartphone Screen Glass Dataset for Defect Detection
abstract
Interactive devices with touch screen have become commonly used in various aspects of daily life, which raises the demand for high production quality of touch screen glass. While it is desirable to develop effective defect detection technologies to optimize the automatic touch screen production lines, the development of these technologies suffers from the lack of publicly available datasets. To address this issue, we in this paper propose a dedicated touch screen glass defect dataset which includes seven types of defects and consists of 2504 images captured in various scenarios. All data are captured with professional acquisition equipment on the fixed workstation. Additionally, we benchmark the CNN- and Transformer-based object detection frameworks on the proposed dataset to demonstrate the challenges of defect detection on high-resolution images. Dataset and related code will be available at https://github.com/VincentHancoder/SSGD.
Haonan Han, Rui Yang 0040, Shuyan Li, Runze Hu, Xiu Li 0001
ICASSP3
2023 Towards Realizing the Value of Labeled Target Samples: A Two-Stage Approach for Semi-Supervised Domain Adaptation
abstract
Semi-Supervised Domain Adaptation (SSDA) is a recently emerging research topic that extends from the widely-investigated Unsupervised Domain Adaptation (UDA) by further having a few target samples labeled, i.e., the model is trained with labeled source samples, unlabeled target samples as well as a few labeled target samples. Compared with UDA, the key to SSDA lies how to most effectively utilize the few labeled target samples. Existing SSDA approaches simply merge the few precious labeled target samples into vast labeled source samples or further align them, which dilutes the value of labeled target samples and thus still obtains a biased model. To remedy this, in this paper, we propose to decouple SSDA as an UDA problem and a semi-supervised learning problem where we first learn an UDA model using labeled source and unlabeled target samples and then adapt the learned UDA model in a semi-supervised way using labeled and unlabeled target samples. By utilizing the labeled source samples and target samples separately, the bias problem can be well mitigated. We further propose a consistency learning based mean teacher model to effectively adapt the learned UDA model using labeled and unlabeled target samples. Experiments show our approach outperforms existing methods.
Mengqun Jin, Kai Li 0012, Shuyan Li, Chunming He, Xiu Li 0001
ICASSP3
2023 SemanticAC: Semantics-Assisted Framework for Audio Classification
abstract
In this paper, we propose SemanticAC, a semantics-assisted framework for Audio Classification to better leverage the semantic information. Unlike conventional audio classification methods that treat class labels as discrete vectors, we employ a language model to extract abundant semantics from labels and optimize the semantic consistency between audio signals and their labels. We verify that simple textual information from labels and advanced pretraining models enable more abundant semantic supervision for better performance. Specifically, we design a text encoder to capture the semantic information from the text extension of labels. Then we map the audio signals to align with the semantics of corresponding class labels via an audio encoder and a similarity calculation module so as to enforce the semantic consistency. Extensive experiments on two audio datasets, ESC-50 and US8K demonstrate that our proposed method consistently outperforms the compared audio classification methods.
Yicheng Xiao, Yue Ma 0016, Shuyan Li, Hantao Zhou, Ran Liao, Xiu Li 0001
ICASSP3
2023 SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation
abstract
This paper studies referring video object segmentation (RVOS) by boosting video-level visual-linguistic alignment. Recent approaches model the RVOS task as a sequence prediction problem and perform multi-modal interaction as well as segmentation for each frame separately. However, the lack of a global view of video content leads to difficulties in effectively utilizing inter-frame relationships and understanding textual descriptions of object temporal variations. To address this issue, we propose Semantic-assisted Object Cluster (SOC), which aggregates video content and textual guidance for unified temporal modeling and cross-modal alignment. By associating a group of frame-level object embeddings with language tokens, SOC facilitates joint space learning across modalities and time steps. Moreover, we present multi-modal contrastive supervision to help construct well-aligned joint space at the video level. We conduct extensive experiments on popular RVOS benchmarks, and our method outperforms state-of-the-art competitors on all benchmarks by a remarkable margin. Besides, the emphasis on temporal coherence enhances the segmentation stability and adaptability of our method in processing text expressions with temporal variations. Code is available at https://github.com/RobertLuo1/NeurIPS2023_SOC.
Zhuoyan Luo, Yicheng Xiao, Yong Liu 0033, Shuyan Li, Yansong Tang, Xiu Li 0001, Yujiu Yang 0001
NeurIPS4
2022 Structure-Adaptive Neighborhood Preserving Hashing for Scalable Video Search
abstract
In this paper, we propose a Structure-adaptive Neighborhood Preserving Hashing (SNPH) method for unsupervised scalable video search. Unlike most existing hashing methods which equally encode an entire video into a binary feature vector, we propose a neighborhood attention mechanism which encodes the neighborhood-relevant content of a video to better preserve the neighborhood relationships among videos. Motivated by the fact that a video usually contains multiple shots and each shot depicts a different activity, we further develop a structure-adaptive encoder to model the hierarchical structure of the video. Specifically, the encoder adaptively divides each video into multiple segments via detecting temporal boundaries across frames and encodes these segments as a compact binary vector to capture rich structural information. We integrate the neighborhood attention mechanism into the structure-adaptive encoder to learn hash functions that jointly preserve the neighborhood relationships among videos and exploit the hierarchical structure in a video. Experimental results on three widely used benchmark datasets show that our proposed method consistently outperforms state-of-the-art unsupervised video hashing methods.
Shuyan Li, Xiu Li 0001, Jiwen Lu, Jie Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 Self-Supervised Video Hashing via Bidirectional Transformers
abstract
Most existing unsupervised video hashing methods are built on unidirectional models with less reliable training objectives, which underuse the correlations among frames and the similarity structure between videos. To enable efficient scalable video retrieval, we propose a self-supervised video Hashing method based on Bidirectional Transformers (BTH). Based on the encoder-decoder structure of transformers, we design a visual cloze task to fully exploit the bidirectional correlations between frames. To unveil the similarity structure between unlabeled video data, we further develop a similarity reconstruction task by establishing reliable and effective similarity connections in the video space. Furthermore, we develop a cluster assignment task to exploit the structural statistics of the whole dataset such that more discriminative binary codes can be learned. Extensive experiments implemented on three public benchmark datasets, FCVID, ActivityNet and YFCC, demonstrate the superiority of our proposed approach.
Shuyan Li, Xiu Li 0001, Jiwen Lu, Jie Zhou 0001
CVPR1
2021 Improving Relation Extraction by Knowledge Representation Learning
abstract
Relation extraction is an important NLP task to extract the semantic relationship between two entities. Recently, large-scale pre-training language models have achieved excellent performance in many NLP applications. Most of the existing relation extraction models mainly rely on context information, but entity information is also very important for relation extraction, especially domain knowledge of entity and the direction between entity pairs. In this paper, based on the pre-trained BERT model, we propose a multi-task joint relation extraction model incorporating knowledge representation learning(KRL). The experimental results on the SemEval 2010 task 8 dataset and the KBP37 dataset show that our proposed model outperforms most of state-of-the-art methods. The results on the larger dataset FewRel80 refined from FewRel also indicate that increasing the knowledge representation learning as an auxiliary objective is helpful for the relation extraction task.
Wenxing Hong, Shuyan Li, Abdur Rasool, Qingshan Jiang, Yang Weng
ICTAI2
2021 GCdiscrimination: identification of gastric cancer based on a milliliter of blood
abstract
Gastric cancer (GC) continues to be one of the major causes of cancer deaths worldwide. Meanwhile, liquid biopsies have received extensive attention in the screening and detection of cancer along with better understanding and clinical practice of biomarkers. In this work, 58 routine blood biochemical indices were tentatively used as integrated markers, which further expanded the scope of liquid biopsies and a discrimination system for GC consisting of 17 top-ranked indices, elaborated by random forest method was constructed to assist in preliminary assessment prior to histological and gastroscopic diagnosis based on the test data of a total of 2951 samples. The selected indices are composed of eight routine blood indices (MO%, IG#, IG%, EO%, P-LCR, RDW-SD, HCT and RDW-CV) and nine blood biochemical indices (TP, AMY, GLO, CK, CHO, CK-MB, TG, ALB and γ-GGT). The system presented a robust classification performance, which can quickly distinguish GC from other stomach diseases, different cancers and healthy people with sensitivity, specificity, total accuracy and area under the curve of 0.9067, 0.9216, 0.9138 and 0.9720 for the cross-validation set, respectively. Besides, this system can not only provide an innovative strategy to facilitate rapid and real-time GC identification, but also reveal the remote correlation between GC and these routine blood biochemical parameters, which helped to unravel the hidden association of these parameters with GC and serve as the basis for subsequent studies of the clinical value in prevention program and surveillance management for GC. The identification system, called GC discrimination, is now available online at http://lishuyan.lzu.edu.cn/GC/.
Jiangpeng Wu, Lili Xi, Xiaoying Xu, Dekui Zhang, Shuyan Li
Briefings Bioinform.10
2021 ABCModeller: an automatic data mining tool based on a consistent voting method with a user-friendly graphical interface
abstract
In order to extract useful information from a huge amount of biological data nowadays, simple and convenient tools are urgently needed for data analysis and modeling. In this paper, an automatic data mining tool, termed as ABCModeller (Automatic Binary Classification Modeller), with a user-friendly graphical interface was developed here, which includes automated functions as data preprocessing, significant feature extraction, classification modeling, model evaluation and prediction. In order to enhance the generalization ability of the final model, a consistent voting method was built here in this tool with the utilization of three popular machine-learning algorithms, as artificial neural network, support vector machine and random forest. Besides, Fibonacci search and orthogonal experimental design methods were also employed here to automatically select significant features in the data space and optimal hyperparameters of the three algorithms to achieve the best model. The reliability of this tool has been verified through multiple benchmark data sets. In addition, with the advantage of a user-friendly graphical interface of this tool, users without any programming skills can easily obtain reliable models directly from original data, which can reduce the complexity of modeling and data mining, and contribute to the development of related research including but not limited to biology. The excitable file of this tool can be downloaded from http://lishuyan.lzu.edu.cn/ABCModeller.rar.
Jiangpeng Wu, Honglin Zhai, Shuyan Li
Briefings Bioinform.4
2020 Unsupervised Variational Video Hashing With 1D-CNN-LSTM Networks
abstract
Most existing unsupervised video hashing methods generate binary codes by using RNNs in a deterministic manner, which fails to capture the dominant latent variation of videos. In addition, RNN-based video hashing methods suffer the content forgetting of early input frames due to the sequential processing inherency of RNNs, which is detrimental to global information capturing. In this work, we propose an unsupervised variational video hashing (UVVH) method for scalable video retrieval. Our UVVH method aims to capture the salient and global information in a video. Specifically, we introduce a variational autoencoder to learn a probabilistic latent representation of the salient factors of video variations. To better exploit the global information of videos, we design a 1D-CNN-LSTM model. The 1D-CNN-LSTM model processes long frame sequences in a parallel and hierarchical way, and exploits the correlations between frames to reconstruct the frame-level features. As a consequence, the learned hash functions can produce reliable binary codes for video retrieval. We conduct extensive experiments on three widely used benchmark datasets, FCVID, ActivityNet and YFCC to validate the effectiveness of our proposed approach.
Shuyan Li, Zhixiang Chen 0003, Xiu Li 0001, Jiwen Lu, Jie Zhou 0001
IEEE Trans. Multim.1
2019 Neighborhood Preserving Hashing for Scalable Video Retrieval
abstract
In this paper, we propose a Neighborhood Preserving Hashing (NPH) method for scalable video retrieval in an unsupervised manner. Unlike most existing deep video hashing methods which indiscriminately compress an entire video into a binary code, we embed the spatial-temporal neighborhood information into the encoding network such that the neighborhood-relevant visual content of a video can be preferentially encoded into a binary code under the guidance of the neighborhood information. Specifically, we propose a neighborhood attention mechanism which focuses on partial useful content of each input frame conditioned on the neighborhood information. We then integrate the neighborhood attention mechanism into an RNN-based reconstruction scheme to encourage the binary codes to capture the spatial-temporal structure in a video which is consistent with that in the neighborhood. As a consequence, the learned hashing functions can map similar videos to similar binary codes. Extensive experiments on three widely-used benchmark datasets validate the effectiveness of our proposed approach.
Shuyan Li, Zhixiang Chen 0003, Jiwen Lu, Xiu Li 0001, Jie Zhou 0001
ICCV1
2018 A type of energy-efficient data gathering method based on single sink moving along fixed points
Chao Sha, Jian-mei Qiu, Shuyan Li, Meng-ye Qiang, Ruchuan Wang 0001
Peer-to-Peer Netw. Appl.3
2014 Study of Effect of Urban Green Land on Thermal Environment of Surrounding Buildings: A Case Study in Beijing, China
Qingzu Luan, Caihua Ye, Shuyan Li
ICCSA (4)4