Yangyang Jiang

dblp:257/9698 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models
abstract
Multimodal Large Language Models (MLLMs) have become widely deployed, yet their safety alignment remains fragile under adversarial inputs. Previous work has shown that increasing inference steps can disrupt safety mechanisms and lead MLLMs to generate attacker-desired harmful content. However, most existing attacks focus on increasing the complexity of the modified visual task itself and do not explicitly leverage the model’s own reasoning incentives. This leads to them underperforming on reasoning models (Models with Chain-of-Thoughts) compared to non-reasoning ones (Models without Chain-of-Thoughts). If a model can think like a human, can we influence its cognitive-stage decisions so that it proactively completes a jailbreak? To validate this idea, we propose GAMBIT (Gamified Adversarial Multimodal Breakout via Instructional Traps), a novel multimodal jailbreak framework that decomposes and reassembles harmful visual semantics, then constructs a gamified scene that drives the model to explore, reconstruct intent, and answer as part of winning the game. The resulting structured reasoning chain increases task complexity in both vision and text, positioning the model as a participant whose goal pursuit reduces safety attention and induces it to answer the reconstructed malicious query. Extensive experiments on popular reasoning and non-reasoning MLLMs demonstrate that GAMBIT achieves high Attack Success Rates (ASR), reaching 92.13% on Gemini 2.5 Flash, 91.20% on QvQ-MAX, and 85.87% on GPT-4o, significantly outperforming baselines. Warning: This paper contains unsafe and offensive examples.
Xiangdong Hu, Yangyang Jiang, Qin Hu 0001, Xiaojun Jia
ACL (1)2
2024 Multi-Band Speech Tensor Decomposition for Interactive Feature Extraction in Early Dysphagia Screening
abstract
Dysphagia is a prevalent symptom in numerous neurological disorders among older adults. Current dysphagia diagnostic systems either involve invasive procedures or necessitate the ingestion of liquids. Some researchers have devised automatic dysphagia detection methods based on vowels that are easy to collect and sensitive to vocal cord states. These methods extract features from each vowel separately and fuse them to train models. Nonetheless, they neglect potential interrelations among different vowels. Vowels collected from the same speaker could share subspaces since they are produced from the same vocal system. In this study, we introduce a tensor-based method that can simultaneously extract interactive information from all vowels across different modes. This method designs multi-band speech tensors and core-pruned tensor networks to investigate crucial frequency bands and connections for dysphagia screening. Experimental results show our model exceeds previous methods by approximately 10 percentage points in the classification accuracy.
Yipeng Liu 0001, Da Shen, Yangyang Jiang, Ce Zhu
ICASSP4
2024 Fault Detection and Location of Transmission Lines Based on Convolutional Neural Network
abstract
Transmission lines are an important part of the power system, and the normal operation of the transmission line is a key step to ensure the stable operation of the power system. The failure of transmission line can cause interruption of power supply and cause serious economic losses. This work uses Convolutional Neural Networks (CNN) to detect and locate the transmission line faults. In the model, a time window is used to perform sliding sampling on the fault data to generate the data set, which not only speeds up the training process, but also retains the temporal correlation of the data. Two CNNs are used to achieve fault detection and location, respectively. In addition, Gaussian white noise is added to the test set to evaluate the anti-interference ability of the model. The results show that the model has good performance and high stability for transmission line fault detection and location.
Yangyang Jiang, Yongxiang Xia, Haicheng Tu, Chunshan Liu
ISCAS1
2023 A cluster reputation-based hierarchical consensus model in blockchain
Yangyang Jiang
Peer Peer Netw. Appl.1
2022 Dysphagia diagnosis system with integrated speech analysis from throat vibration
Hengling Zhao, Yangyang Jiang, Shenghan Wang, Fangzhou Ren, Ce Zhu, Jirong Yue, Yipeng Liu 0001
Expert Syst. Appl.2
2021 Smart Dysphagia Detection System with Adaptive Boosting Analysis of Throat Signals
abstract
Dysphagia is a symptom of many neurological disorders. Existing diagnosis systems are either invasive or require swallowing liquids, which are costly and harmful to humans. In this work, we design a smart dysphagia detection system based on speech signals. Rather than the voice data acquired by traditional microphones, we apply a bone conduction headset for vibration signal acquisition from the throat to get cleaner speech signals. After speech feature extraction, under-sampling is performed to deal with the imbalanced data problem, and principal component analysis is used for dimensionality reduction. In this paper, we construct an ensemble adaptive boosting classifier to detect the dysphagia patient. Experimental results show that the testing classification accuracy of the proposed system reaches 71.2%. Sensitivity and specificity can reach 66.6 % and 76 %, respectively.
Shenghan Wang, Yangyang Jiang, Hengling Zhao, Ce Zhu, Yipeng Liu 0001
ISCAS2
2021 Deep Learning Based Gait Analysis for Contactless Dementia Detection System from Video Camera
abstract
Dementia is a neurodegenerative disease with a high incidence in the elderly. However, there is no effective treatment for this disease, and early intervention has a great effect to slow the deterioration. Currently, the detection of dementia is mainly achieved using questionnaire-like neuropsychological tests. Such ways usually cost a lot of time. To this end, we design a contactless dementia detection system based on gait analysis from surveillance video, and it can serve as a home-based healthcare system. This system applies a Kinect 2.0 camera to capture the human video and extract the skeleton joints at a rate of 15 frames per second. Two different gaits are collected for detection, namely single-task gait and dual-task gait. In this paper, we design a convolutional neural network based classifier to extract features in a data-driven way from these two groups of videos, but not take hand-crafted features. Experimental results show that we achieve a sensitivity of 74.10% on the test set using this system, and the processing only takes several minutes for early dementia detection.
Yangyang Jiang, Xingyu Cao, Ce Zhu, Yipeng Liu 0001
ISCAS2