VLDB 2026 Research / reviewers in the wild / expert
Xinyan Yang
dblp:23/3516
· DBLP profile ↗
15ranked-venue papers
2as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automated irrelevant individuals recognition algorithm in video via motion trajectoriesabstractAbstract The privacy of irrelevant individuals appearing in social media videos is often compromised, which presents a challenging issue for privacy protection. Current approaches to privacy protection primarily rely on manually identifying irrelevant individuals, which is both time-consuming and labor-intensive. Therefore, this paper proposes an automatic algorithm for identifying irrelevant individuals to protect their privacy efficiently. The method employs multi-object tracking to extract spatiotemporal motion features, enabling the automated recognition of irrelevant individuals. In addition, a trajectory association algorithm is employed to improve the precision of tracking from occlusion and blurring during motion. Additionally, as the field of irrelevant people privacy protection lacks data support, this paper constructs a dataset for further research, consisting of 101 685 irrelevant faces and 40 247 relevant faces. Through experimental validation on various datasets, the proposed method shows significant improvements in various metrics compared with the state-of-the-art approaches, with MOTA increasing by 3.2%, HOTA increasing by 13%, and the accuracy of individual irrelevant recognition reaching 96.78%. Yixin Ma, Jiaping Lin, Junhao Zeng, Xinyan Yang, Kui Zhao, Gang Liang |
Comput. J. | 4 |
| 2025 | Image forgery localization integrating multi-scale and boundary featuresabstractAbstract Image forgery localization identifies tampered regions within an image by extracting distinctive forgery features. Current methods mainly use convolutional neural networks (CNNs) to extract features. However, CNNs’ limited receptive field emphasizes local features, impeding the global modeling of crucial lower-level features like edges and textures, leading to decreased precision. Moreover, prior methods use pyramid networks for multi-scale feature extraction but show deficiencies in multi-scale and interlayer modeling, leading to inadequate multi-scale information representation and limiting flexibility to tampered regions of varying sizes. To address these issues, this paper proposes a Transformer-based model integrating multi-scale and boundary features. The model employs a Pyramid Vision Transformer as the encoder, using self-attention over convolution to enhance global context modeling. Building on this, the model incorporates a multi-scale feature enhancement module that enriches forgery features by paralleling various convolutional layers. Features at various encoder stages are integrated through a cross-stage interaction module, enabling multi-level feature correlation for a strong feature representation. Furthermore, the model includes a forgery boundary information-guided branch, which focuses precisely on tampered region structures without introducing irrelevant noise. Experiments demonstrate that our model surpasses previous methods in localization accuracy, with F1 and AUC improving by 8.5% and 2.2% in pre-training, respectively. Xinyan Yang, Rongchuan Zhang, Shao Li, Gang Liang |
Comput. J. | 1 |
| 2025 | Forward Secure Equality Test for Secure Data Sharing in Healthcare SystemsabstractIn healthcare systems, the protection of the sensitive data of patient during sharing across multiple entities is crucial. Encrypting these data before outsourcing to the healthcare server is an effective method to protect data privacy. As a solution to protect data sharing within healthcare systems, identity-based encryption with equality test (IBEET) is an effective strategy that allows testing whether two ciphertexts encrypted from the same plaintext without certificate management. However, current IBEET schemes often lack of control over the trapdoors’ lifetime, i.e. the malicious server could use trapdoor to test whether two ciphertext generated after the trapdoor encrypted by the same plaintext all the time, potentially leaking sensitive data which the user are unwilling to see. To address this problem, we introduce a forward secure identity-based encryption with equality test (FS-IBEET) scheme. This new scheme prevents the malicious server from accessing useful information from newly encrypted files. We perform a thorough security analysis using the random oracle model under the q-BDHI assumption. Our experiments demonstrate that the performance of our scheme in comparison to other schemes. Jianchang Lai, Jinguang Han, Liquan Chen, Xinyan Yang |
IEEE Internet Things J. | 5 |
| 2025 | Cosine-Initialized MAE for Cross-Domain Few-Shot Recognition in Distributed Fiber-Optic Vibration Sensing Systems
Xiankun Wang, Zhengxian Zhou, Dawei Zhang 0009, Jun Qu, Jianping Shi, Yashuai Han, Xinyan Yang, Songlin Zhuang |
IEEE Internet Things J. | 7 |
| 2025 | Blockchain-based synchronized data transmission with dynamic device management for digital twin in IIoT
Xinyan Yang, Kai Zhang 0016, Zhimei Sui, Jian Zhao 0023 |
Peer Peer Netw. Appl. | 2 |
| 2024 | Video Editing Chatbot: Language-Driven Video Compositing System
Xinyan Yang, Aiqi Wang, Jianglin Zeng, Shaofei Liu |
ACM Multimedia | 2 |
| 2024 | Keyframe-guided Video Swin Transformer with Multi-path Excitation for Violence DetectionabstractAbstract Violence detection is a critical task aimed at identifying violent behavior in video by extracting frames and applying classification models. However, the complexity of video data and the suddenness of violent events present significant hurdles in accurately pinpointing instances of violence, making the extraction of frames that indicate violence a challenging endeavor. Furthermore, designing and applying high-performance models for violence detection remains an open problem. Traditional models embed extracted spatial features from sampled frames directly into a temporal sequence, which ignores the spatio-temporal characteristics of video and limits the ability to express continuous changes between adjacent frames. To address the existing challenges, this paper proposes a novel framework called ACTION-VST. First, a keyframe extraction algorithm is developed to select frames that are most likely to represent violent scenes in videos. To transform visual sequences into spatio-temporal feature maps, a multi-path excitation module is proposed to activate spatio-temporal, channel and motion features. Next, an advanced Video Swin Transformer-based network is employed for both global and local spatio-temporal modeling, which enables comprehensive feature extraction and representation of violence. The proposed method was validated on two large-scale datasets, RLVS and RWF-2000, achieving accuracies of over 98 and 93%, respectively, surpassing the state of the art. Xinyan Yang, Gang Liang |
Comput. J. | 2 |
| 2024 | Pornographic video detection based on semantic and image enhancementabstractAbstract Pornographic video detection is of significant importance in curbing the proliferation of pornographic information on online video platforms. However, existing works often employ generic frame extraction methods that ignore the low-latency requirements of detection scenarios and the characteristics of pornographic videos. Additionally, existing detection methods have difficulties in detail characterization and semantic understanding, resulting in low accuracy. Therefore, this paper proposes an efficient pornographic video detection framework based on semantic and image enhancement. Firstly, a keyframe extraction method tailored for pornographic video detection is proposed to select representative frames. Secondly, a light enhancement method is introduced to facilitate accurate capture of pornographic visual cues. Moreover, a compression-reconstruction network is employed to eliminate adversarial perturbations, enabling models to obtain reliable features. Subsequently, YOLOv5 is introduced to locate and crop human targets in keyframes, reducing background interference and enhancing the expression of human semantic information. Finally, MobileNetV3 is employed to determine if the human targets contain pornographic content. The proposed framework is validated on the publicly available NPDI dataset, achieving an accuracy of 95.9%, surpassing existing baseline methods. Junhao Zeng, Gang Liang, Yixin Ma, Xinyan Yang |
Comput. J. | 4 |
| 2024 | Improving disentanglement in variational auto-encoders via feature imbalance-informed dimension weighting
Zhenyao Yu, Zitu Liu, Ziyi Yu, Xinyan Yang, Xingyue Li, Yike Guo, Qun Liu 0005, Guoyin Wang 0001 |
Knowl. Based Syst. | 5 |
| 2024 | A Dataset and Benchmark for 3D Scene Plausibility AssessmentabstractThe surge in popularity of 3D scene synthesis has driven the development of diverse methods for assessing the quality of synthesized scenes. While subjective assessment methods are widespread, their time-consuming and labor-intensive nature prompts exploration into more efficient objective alternatives. This paper introduces an objective approach to evaluating scene plausibility, aiming to overcome the limitations associated with subjective methods. To underpin our objective evaluation, we present the 3D-SPAD dataset, comprising plausibility scores for 3000 scenes across 46 object categories. Leveraging this dataset, we propose a graph attention-based network designed to accurately estimate scene plausibility. A comprehensive evaluation of our network is conducted through a series of experiments, showcasing its feasibility and reliability. For additional details and access to the code, please refer to our GitHub repository athttps://github.com/Mayibo-cuc/3D-SPAN. Yibo Ma, Wei Zhong 0001, Long Ye, Xinyan Yang, Qin Zhang 0009 |
IEEE Trans. Multim. | 5 |
| 2024 | Nowhere to Hide: Online Rumor Detection Based on Retweeting Graph Neural NetworksabstractOnline rumor detection is crucial for a healthier online environment. Traditional methods mainly rely on content understanding. However, these contents can be easily adjusted to avoid such supervision and are insufficient to improve the detection result. Compared with the content, information propagation patterns are more informative to support further performance promotion. Unfortunately, learning the propagation patterns is difficult, since the retweeting tree is more topologically complicated than linear sequences or binary trees. In light of this, we propose a novel rumor detection framework based on structure-aware retweeting graph neural networks. To capture the propagation patterns, we first design a novel conversion method to transform the complex retweeting tree as more tractable binary tree without losing the reconstruction information. Then, we serialize the retweeting tree as a corpus of meta-tree paths, where each meta-tree can preserve a basic substructure. A deep neural network is then designed to integrate all meta-trees and to generate the global structural embeddings. Furthermore, we propose to integrate content, users, and propagation patterns to enhance more reliable performance. To this end, we propose a novel self-attention-based retweeting neural network to learn individual features from both content and users. We then fuse the node-level features with our global structural embeddings via a mutual attention unit. In this way, we can generate more comprehensive representations for rumor detection. Extensive evaluations on two real-world datasets show remarkable superiorities of our model compared with existing methods. Bo Liu 0004, Xiangguo Sun, Qing Meng, Xinyan Yang, Yang Lee, Jiuxin Cao, Junzhou Luo, Roy Ka-Wei Lee |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Multi-stage dynamic disinformation detection with graph entropy guidance
Xiaorong Hao, Bo Liu 0004, Xinyan Yang, Xiangguo Sun, Qing Meng, Jiuxin Cao |
World Wide Web (WWW) | 3 |
| 2021 | Text to Scene: A System of Configurable 3D Indoor Scene SynthesisabstractIn this work, we show the Text to Scene system, which can configure 3D indoor scene from natural language. Given a text, the system will organize inclusive semantic message to a graph template, complete the graph with a novel graph-based contextual completion method Contextual ConvE(CConvE) and visulize the graph by arranging 3D models under an object location protocol. In the experiments, qualitative results obtained by the Text to Scene(T2S) system and quantitative evaluation of CConvE compared with other state-of-the-art approaches are reported. Xinyan Yang, Long Ye |
ACM Multimedia | 1 |
| 2019 | M2-VISD: A Visual Intelligence Evaluation Dataset Based on Virtual SceneabstractMost of the existing visual intelligence evaluation datasets like ImageNet or ActivityNet share a common property, namely consisting of many images or videos captured from the real world. This property makes this evaluation datasets suitable for actual applications, but has limitations in achieving the diversity and interactivity of evaluation environments. The currently available datasets do not systematically provide data of different scales and different angles. In order to solve the above problems, this paper constructs a multi-angle and multi-scale data set based on UE4 platform, which can evaluate the performance of the algorithm from both scale and angle in an all-round way. Our experiments show that the algorithms detection performance varies greatly under different scale and different angle. Particularly, in scale-data, when the distance between the camera and the object is less than 50cm or greater than 3200cm, algorithm performance is poor, and in angle-data, the camera’s pitch angle is also poor when it is 18 degree and 0 degree. This result has a guiding role for the correct evaluation of the algorithm performance and provides help for the innovation and optimization of the algorithm. Yaning Tan, Xinyan Yang, Tianyi Feng, Jianbiao Wang, Long Ye |
ICIS | 2 |
| 2018 | Echo State Network with Multiple Loops Reservoir and its Application in Network Traffic PredictionabstractEcho state network (ESN), which was proposed as a novel recurrent neural network (RNN), has already been proved to exhibit better prediction ability than traditional neural networks in dealing with time series prediction. However, ESN's randomly generated reservoir structure is of high complexity and can not guarantee the stability of the prediction. In this paper, we propose a novel ESN with deterministic multiple loops reservoir structure (MLR) to avoid the randomness of the reservoir in the classic ESN. In addition, compared with the adjacent-feedback loop reservoir structure (ALR), the novel reservoir structure strengthens the connection of neurons in the reservoir and improves the nonlinear approximation ability of ESN. To test its performance, our MLR-based ESN is applied to network traffic prediction. Extensive simulation results regarding to different prediction steps demonstrate that MLR can achieve higher prediction accuracy, and outperform existing prediction models. Furthermore, we also analyze the influence of parameters of MLR on the prediction accuracy, such as the neuronal interval of multiple loops and the number of loops. Xinyan Yang, Jian Zhou 0009, Fu Xiao 0001 |
CSCWD | 2 |