Jingyi Zheng

dblp:166/3143 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
11since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 2 first-authorSecurity and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2027 Two-stage distributionally robust decoding for reliability-oriented natural language generation
Guangnan He, Lei La, Keyu Gao, Jingyi Zheng
Inf. Process. Manag.4
2026 Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe Adaptation
abstract
Tianyi Hu, Andrea Morales-Garzón, Jingyi Zheng, Maria Maistro, Daniel Hershcovich. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Andrea Morales-Garzón, Jingyi Zheng, Maria Maistro, Daniel Hershcovich
ACL (1)3
2025 CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
abstract
Backdoor attacks significantly compromise the security of large language models by triggering them to output specific and controlled content. Currently, triggers for textual backdoor attacks fall into two categories: fixed-token triggers and sentence-pattern triggers. However, the former are typically easy to identify and filter, while the latter, such as syntax and style, do not apply to all original samples and may lead to semantic shifts. In this paper, inspired by cross-lingual (CL) prompts of LLMs in real-world scenarios, we propose a higher-dimensional trigger method at the paragraph level, namely CL-Attack. CL-Attack injects the backdoor by using texts with specific structures that incorporate multiple languages, thereby offering greater stealthiness and universality compared to existing backdoor attack techniques. Extensive experiments on different tasks and model architectures demonstrate that CL-Attack can achieve nearly 100 percents attack success rate with a low poisoning rate in both classification and generation tasks. We also empirically show that CL-Attack is more robust against current major defense methods compared to baseline backdoor attacks. Additionally, in response to CL-Attack, we further develop a new defense called TranslateDefense, which can partially mitigate the impact of CL-Attack.
Jingyi Zheng, Tianshuo Cong, Xinlei He 0001
AAAI1
2025 On the Generalization and Adaptation Ability of Machine-Generated Text Detectors in Academic Writing
abstract
The rising popularity of large language models (LLMs) has raised concerns about potential abuse and harmful content. As a result, developing a highly generalizable and adaptable machine-generated text (MGT) detection system has become an urgent priority. Given that LLMs are most commonly misused in academic writing, this work investigates the generalization and adaptation capabilities of MGT detectors in three key aspects specific to academic writing: First, we construct MGT-Academic, a large-scale dataset comprising over 336M tokens and 749K samples. MGT-Academic focuses on academic writing, featuring human-written texts (HWTs) and MGTs across STEM, Humanities, and Social Sciences, paired with an extensible code framework for efficient benchmarking. Second, we benchmark the performance of various detectors for binary classification and text attribution tasks in both in-domain and cross-domain settings. This benchmark reveals the often-overlooked challenges of text attribution tasks. Third, we introduce a novel text attribution task in which models must adapt to new classes over time, with little or no access to prior training data, spanning both few-shot and many-shot scenarios. We implement a range of adaptation techniques to enhance performance across these settings. Our findings provide new insights into the generalization ability of MGT detectors and lay the foundation for building robust, adaptive detection systems. The code framework is available at https://github.com/Y-L-LIU/MGTBench-2.0.
Yule Liu, Zhiyuan Zhong, Zhen Sun 0001, Jingyi Zheng, Jiaheng Wei, Qingyuan Gong, Fenghua Tong, Yang Chen 0001, Yang Zhang 0016, Xinlei He 0001
KDD (2)5
2025 TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
abstract
As Large Language Models (LLMs) advance, Machine-Generated Texts (MGTs) have become increasingly fluent, high-quality, and informative. Existing wide-range MGT detectors are designed to identify MGTs to prevent the spread of plagiarism and misinformation. However, adversaries attempt to humanize MGTs to evade detection (named evading attacks), which requires only minor modifications to bypass MGT detectors. Unfortunately, existing attacks generally lack a unified and comprehensive evaluation framework, as they are assessed using different experimental settings, model architectures, and datasets. To fill this gap, we introduce the Text-Humanization Benchmark (TH-Bench), the first comprehensive benchmark to evaluate evading attacks against MGT detectors. TH-Bench evaluate attacks across three key dimensions: evading effectiveness, text quality, and computational overhead. Our extensive experiments evaluate 6 state-of-the-art attacks against 13 MGT detectors across 6 datasets, spanning 19 domains and generated by 11 widely used LLMs. Our findings reveal that no single evading attack excels across all three dimensions. Through in-depth analysis, we highlight the strengths and limitations of different attacks. More importantly, we identify a trade-off among three dimensions and propose two optimization insights. Through preliminary experiments, we validate their correctness and effectiveness, offering potential directions for future research.
Jingyi Zheng, Zhen Sun 0001, Wenhan Dong, Yule Liu, Xinlei He 0001
KDD (2)1
2025 CHASM: Unveiling Covert Advertisements on Chinese Social Media
abstract
Current benchmarks for evaluating large language models (LLMs) in social media moderation completely overlook a serious threat: covert advertisements, which disguise themselves as regular posts to deceive and mislead consumers into making purchases, leading to significant ethical and legal concerns. In this paper, we present the CHASM, a first-of-its-kind dataset designed to evaluate the capability of Multimodal Large Language Models (MLLMs) in detecting covert advertisements on social media. CHASM is a high-quality, anonymized, manually curated dataset consisting of 4,992 instances, based on real-world scenarios from the Chinese social media platform Rednote. The dataset was collected and annotated under strict privacy protection and quality control protocols. It includes many product experience sharing posts that closely resemble covert advertisements, making the dataset particularly challenging.The results show that under both zero-shot and in-context learning settings, none of the current MLLMs are sufficiently reliable for detecting covert advertisements.Our further experiments revealed that fine-tuning open-source MLLMs on our dataset yielded noticeable performance gains. However, significant challenges persist, such as detecting subtle cues in comments and differences in visual and textual structures.We provide in-depth error analysis and outline future research directions. We hope our study can serve as a call for the research community and platform moderators to develop more precise defenses against this emerging threat.
Jingyi Zheng, Yule Liu, Zhen Sun 0001, Zongmin Zhang, Zifan Peng, Wenhan Dong, Xinlei He 0001
NeurIPS1
2025 Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
Zeren Luo, Zifan Peng, Yule Liu, Zhen Sun 0001, Jingyi Zheng, Xinlei He 0001
USENIX Security Symposium6
2024 Volume-Optimal Persistence Homological Scaffolds of Hemodynamic Networks Covary with MEG Theta-Alpha Aperiodic Dynamics
Nghi Nguyen, Enrico Amico, Jingyi Zheng, Huajun Huang, Alan D. Kaplan, Giovanni Petri, Joaquín Goñi, Ralph Kaufmann, Yize Zhao, Duy Duong-Tran, Li Shen 0001
MICCAI (3)4
2024 A Two-Layer Blockchain Sharding Protocol Leveraging Safety and Liveness for Enhanced Performance
Yibin Xu, Jingyi Zheng, Boris Düdder, Tijs Slaats, Yongluan Zhou
NDSS2
2022 Time-Frequency Analysis of Scalp EEG With Hilbert-Huang Transform and Deep Learning
abstract
Electroencephalography (EEG) is a brain imaging approach that has been widely used in neuroscience and clinical settings. The conventional EEG analyses usually require pre-defined frequency bands when characterizing neural oscillations and extracting features for classifying EEG signals. However, neural responses are naturally heterogeneous by showing variations in frequency bands of brainwaves and peak frequencies of oscillatory modes across individuals. Fail to account for such variations might result in information loss and classifiers with low accuracy but high variation across individuals. To address these issues, we present a systematic time-frequency analysis approach for analyzing scalp EEG signals. In particular, we propose a data-driven method to compute the subject-specific frequency bands for brain oscillations via Hilbert-Huang Transform, lifting the restriction of using fixed frequency bands for all subjects. Then, we propose two novel metrics to quantify the power and frequency aspects of brainwaves represented by sub-signals decomposed from the EEG signals. The effectiveness of the proposed metrics are tested on two scalp EEG datasets and compared with four commonly used features sets extracted from wavelet and Hilbert-Huang Transform. The validation results show that the proposed metrics are more discriminatory than other features leading to accuracies in the range of 94.93% to 99.84%. Besides classification, the proposed metrics show great potential in quantification of neural oscillations and serving as biomarkers in the neuroscience research.
Jingyi Zheng, Mingli Liang, Sujata Sinha, Linqiang Ge, Wei Yu 0002, Arne D. Ekstrom, Fushing Hsieh
IEEE J. Biomed. Health Informatics1
2022 Unsupervised Adversarial Network Alignment with Reinforcement Learning
abstract
Network alignment, which aims at learning a matching between the same entities across multiple information networks, often suffers challenges from feature inconsistency, high-dimensional features, to unstable alignment results. This article presents a novel network alignment framework, Unsupervised Adversarial learning based Network Alignment(UANA), that combines generative adversarial network (GAN) and reinforcement learning (RL) techniques to tackle the above critical challenges. First, we propose a bidirectional adversarial network distribution matching model to perform the bidirectional cross-network alignment translations between two networks, such that the distributions of real and translated networks completely overlap together. In addition, two cross-network alignment translation cycles are constructed for training the unsupervised alignment without the need of prior alignment knowledge. Second, in order to address the feature inconsistency issue, we integrate a dual adversarial autoencoder module with an adversarial binary classification model together to project two copies of the same vertices with high-dimensional inconsistent features into the same low-dimensional embedding space. This facilitates the translations of the distributions of two networks in the adversarial network distribution matching model. Finally, we develop an RL based optimization approach to solve the vertex matching problem in the discrete space of the GAN model, i.e., directly select the vertices in target networks most relevant to the vertices in source networks, without unstable similarity computation that is sensitive to discriminative features and similarity metrics. Extensive evaluation on real-world graph datasets demonstrates the outstanding capability of UANA to address the unsupervised network alignment problem, in terms of both effectiveness and scalability.
Yang Zhou 0001, Jiaxiang Ren 0001, Ruoming Jin, Zijie Zhang 0001, Jingyi Zheng, Zhe Jiang 0001, Da Yan 0001, Dejing Dou
ACM Trans. Knowl. Discov. Data5
2020 Assessing the Impact of Government Interventions on the Spread of COVID-19 with Dynamic Epidemic Models: A case study of Texas
abstract
COVID-19 has been rapidly spreading and causing hundreds of thousands of casualties across the world. This disease originated in Wuhan, China in 2019, and it has since proliferated into a pandemic, causing the public's daily lives to change drastically around the world. Despite the evolution of COVID-19, remarkably little is known about the efficacy of proposed protective safety orders. Thus, our paper is intended to study the effectiveness of the local and state government restrictions and closures in limiting the spread of COVID-19. To mathematically model the spread of COVID-19, we propose a time-dependent SIR model together with Lasso to monitor the trajectories of the transmission and recover rates in relation to the government closures and restrictions, and further predict the number of cases. To validate our model, we conduct both simulation and state level data analysis. Since the government orders vary among different states, we use Texas as an example state to illustrate our algorithm for assessing the government intervention. Our one-day prediction error for the confirmed cases is around 2.47% and less than 1% for the recovered cases. We also find that there are many intervention methods that corresponded to changes in infection rate and recovery rate of the population.
Layla S. Araiinejad, Yuexin Li, Jacqueline R. Carlton, Jingyi Zheng
BIBM4
2020 Robust Meta Network Embedding against Adversarial Attacks
abstract
Recent studies have shown that graph mining models are vulnerable to adversarial attacks. This paper proposes a robust meta network embedding framework, RoMNE, which improves the robustness of multiple network embedding on adversarial noisy networks while preserving the utility on original clean ones. First, we propose a generic meta learning based multiple network embedding model that can quickly adapt it to new embedding tasks on a variety of network data with only a small number of parameter and training updates. Second, Gumbel estimator and Gaussian smoothing techniques are introduced to implement differentiable approximation for optimizing non-differential objective of effective adversarial attacks. Last but not least, the adversarial attack and defense models are integrated into a dynamic adversarial training model. The competition of two models helps the latter be robust to adversarial attacks.
Yang Zhou 0001, Jiaxiang Ren 0001, Dejing Dou, Ruoming Jin, Jingyi Zheng, Kisung Lee
ICDM5
2020 Automated Semantic Segmentation of Cardiac Magnetic Resonance Images with Deep Learning
abstract
Machine learning algorithms, especially deep learning architectures, have demonstrated immense potential for biomedical segmentation, often surpassing expert-level performance. For cardiac magnetic resonance (CMR) imaging, semantic segmentation is critical to deriving clinical measures such as myocardial mass and volume. However, challenges still exist. Manual delineation by domain experts is time-consuming and subject to human errors. With semi-automated segmentation techniques, it is challenging to analyze images of the same subject twice, end-diastole and end-systole of the cardiac cycle. To address these challenges, we propose a deep learning-based end-to-end analytical pipeline for automated segmentation of short-axis CMR imaging. The automated pipeline successfully avoids the problem of human subjectivity and achieves expert-level segmentation accuracy. With a large heterogeneous data inclusive of subjects with varying conditions, our model overcomes the data-homogeneity and achieves 99.9% dice similarity score, which outperforms the current state-of-art work.
Sujata Sinha, Thomas S. Denney Jr., Yang Zhou 0001, Jingyi Zheng
ICMLA4
2020 A Data-Driven Approach to Predict and Classify Epileptic Seizures from Brain-Wide Calcium Imaging Video Data
abstract
The prediction of epileptic seizures has been an essential problem of epilepsy study. The calcium imaging video data images the whole brain-wide neurons activities with electrical discharge recorded by calcium fluorescence intensity (CFI). In this paper, using the zebrafish's brain-wide calcium image video data, we propose a data-driven approach to effectively detect the systemic change-point, and further predict the epileptic seizures. Our approach includes two phases: offline training and online testing. Specifically, during offline training, we extract features and confirm the existence of systemic change-point, then estimate the ratio of unchanged system duration to interictal period duration. For online testing, we implement a statistical model to estimate the change-point, and then predict the onset of epileptic seizure. The testing results show that our proposed approach could effectively predict the time range of future epileptic seizure. Furthermore, we explore the macroscopic patterns of epileptic and control cases, and extract features based on the pattern difference, then implement and compare the classification performance from four machine learning models. Based on the data structure, we also propose a new method to discretize related features, and combine with hierarchical clustering to better visualize and explain the pattern difference between epileptic and control cases.
Jingyi Zheng, Fushing Hsieh, Linqiang Ge
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 A Reconfigurable Architecture for Discrete Cosine Transform in Video Coding
abstract
Discrete cosine transform (DCT) is an indispensable module in video codecs and is a major part in many video coding standards including the latest high efficiency video coding (HEVC). As the video resolution increases, both transform sizes and the number of transforms increase continuously which poses challenges to the reusability design especially in hardware implementation. This paper presents reconfigurable transform architecture to flexibly support the reusability of different transform sizes. The proposed architecture maximally reuses the hardware resources by rearranging the order of input data for different transform sizes while still exploiting the butterfly property. Furthermore, this architecture supports reconfigurable throughput according to different hardware resource requirements. By applying the proposed architecture to the field-programmable gate array (FPGA) design of HEVC core transform matrices, the synthesis results show much lower consumption of hardware resources comparing to existing methods in the literature. The implementation in Altera's Stratix III FPGA can operate at 139 MHz and supports real-time processing of 3840×2160 ultrahigh definition video at a minimum of 45 f/s and up to 359 f/s for different DCT sizes.
Mingkui Zheng, Jingyi Zheng, Linhuang Wu, Xiuzhi Yang, Nam Ling
IEEE Trans. Circuits Syst. Video Technol.2
2018 A Novel Method to Generate Frequent Itemsets in Distributed Environment
abstract
Frequent itemset mining (FIM) is an important topic in data mining, which extracts knowledge of the relationships among items in a transaction dataset. Apriori algorithm and its variants, apriori-like algorithms, are widely used FIM algorithms. However, in a big data environment, these algorithms are inefficient. Due to the iterative calculation and modification of intermediate results, if an apriori-like algorithm is applied on a high-dimension or large-scale dataset, the memory requirement is unacceptable for a single machine. Although parallel and distributed programming could be a solution to deal with big data problems, apriori-like algorithms are not quite suitable for parallel computing because they need extra time overhead of communication to update intermediate results iteratively in cluster memories. To solve this problem, we propose a novel FIM algorithm, Distributed Apriori Based on Itemset-Encoding (DABIE). Different from existing methods, DABIE has two main advantages. Firstly, it stores intermediate results encoded in the form of 0 and 1 to reduce memory usage. Secondly, generating frequent itemsets is based on logical operation of encoding to reduce modification of data in cluster memories. These two advantages make DABIE more friendly to cluster computing. We apply DABIE on datasets with different scales. Compared with other distributed apriori-like algorithms, the results of our experiments show that DABIE can efficiently improve the multi-iterative FIM in big data environment.
Jingyi Zheng, Xiaoheng Deng, Honggang Zhang 0003
IPCCC1
2018 On Association Study of Scalp EEG Data Channels Under Different Circumstances
Jingyi Zheng, Mingli Liang, Arne D. Ekstrom, Linqiang Ge, Wei Yu 0002, Fushing Hsieh
WASA1
2015 A fast AGC method for multimode zero-IF/sliding-IF WPAN/BAN receivers
abstract
This paper presents a fast mixed-signal automatic gain control (AGC) method for zero-IF/sliding-IF receivers used in wireless personal and body-area networks (WPAN/BAN). The preamble defined in Bluetooth low energy (BLE)/802.15.4 /802.15.6 specifications are as low as 1 Byte. It is a tough challenge for the zero-IF/sliding-IF receivers to perform AGC training, frequency synchronization and symbol timing estimation. By detecting the RF input of the quadrature mixer, IF input of analog filter and ADC output, the proposed method can achieve a desired IF amplitude within 1 bit when interference signal is small and 3 bits when carrier-to-interference ratio (CIR) is -27dB.
Jingjing Dong, Hanjun Jiang, Zhaoyang Weng, Jingyi Zheng, Chun Zhang 0001, Zhihua Wang 0001
ISCAS4