Junyuan Zhang

dblp:135/9113 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2027 BGA: A noise-immune neural distillation framework for malicious signature extraction in high-entropy encrypted flows
Junyuan Zhang, Ruijian Jiao
Inf. Process. Manag.4
2026 JARVIS or Ultron? A Survey on the Safety and Security Threats of Computer-Using Agents
abstract
Ada Chen, Yongjiang Wu, Junyuan Zhang, Jingyu Xiao, Shu Yang, Jen-tse Huang, Kun Wang, Wenxuan Wang, Shuai Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ada Chen, Yongjiang Wu, Junyuan Zhang, Jingyu Xiao, Shu Yang 0010, Jen-tse Huang 0001, Kun Wang 0056, Wenxuan Wang 0001, Shuai Wang 0011
ACL (1)3
2025 Stop Looking for "Important Tokens" in Multimodal Language Models: Duplication Matters More
abstract
Vision tokens in multimodal large language models often dominate huge computational overhead due to their excessive length compared to linguistic modality. Abundant recent methods aim to solve this problem with token pruning, which first defines an importance criterion for tokens and then prunes the unimportant vision tokens during inference. However, in this paper, we show that the importance is not an ideal indicator to decide whether a token should be pruned. Surprisingly, it usually results in inferior performance than random token pruning and leading to incompatibility to efficient attention computation operators. Instead, we propose DART (Duplication-Aware Reduction of Tokens), which prunes tokens based on its duplication with other tokens, leading to significant and training-free acceleration. Concretely, DART selects a small subset of pivot tokens and then retains the tokens with low duplication to the pivots, ensuring minimal information loss during token pruning. Experiments demonstrate that DART can prune 88.9% vision tokens while maintaining comparable performance, leading to a 1.99\times and 2.99\times speed-up in total time and prefilling stage, respectively, with good compatibility to efficient attention operators.
Zichen Wen, Yifeng Gao 0003, Shaobo Wang 0001, Junyuan Zhang, Qintong Zhang, Conghui He, Linfeng Zhang 0001
EMNLP4
2025 OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
abstract
Retrieval-augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external knowledge to reduce hallucinations and incorporate up-to-date information without retraining. As an essential part of RAG, external knowledge bases are commonly built by extracting structured data from unstructured PDF documents using Optical Character Recognition (OCR). However, given the imperfect prediction of OCR and the inherent non-uniform representation of structured data, knowledge bases inevitably contain various OCR noises. In this paper, we introduce OHRBench, the first benchmark for understanding the cascading impact of OCR on RAG systems. OHRBench includes 8,561 carefully selected unstructured document images from seven real-world RAG application domains, along with 8,498 Q&A pairs derived from multimodal elements in documents, challenging existing OCR solutions used for RAG. To better understand OCR's impact on RAG systems, we identify two primary types of OCR noise: Semantic Noise and Formatting Noise and apply perturbation to generate a set of structured data with varying degrees of each OCR noise. Using OHRBench, we first conduct a comprehensive evaluation of current OCR solutions and reveal that none is competent for constructing high-quality knowledge bases for RAG systems. We then systematically evaluate the impact of these two noise types and demonstrate the trend relationship between the degree of OCR noise and RAG performance. Our OHRBench, including PDF documents, Q&As, and the ground truth structured data are released at: https://github.com/opendatalab/OHR-Bench
Junyuan Zhang, Qintong Zhang, Bin Wang 0065, Linke Ouyang, Zichen Wen, Ka-Ho Chow 0001, Conghui He, Wentao Zhang 0001
ICCV1
2025 Metamorphic Testing for Audio Content Moderation Software
abstract
The rapid growth of audio-centric platforms and applications such as Whatsapp and Twitter has transformed the way people communicate and share audio content in modern society. However, these platforms are increasingly misused to disseminate harmful audio content, such as hate speech, deceptive advertisements, and explicit material, which can have significant negative consequences (e.g., detrimental effects on mental health). In response, researchers and practitioners have been actively developing and deploying audio content moderation tools to tackle this issue. Despite these efforts, malicious actors can bypass moderation systems by making subtle alterations to audio content, such as modifying pitch or inserting noise. Moreover, the effectiveness of modern audio moderation tools against such adversarial inputs remains insufficiently studied. To address these challenges, we propose MTAM, a Metamorphic Testing framework for Audio content Moderation software. Specifically, we conduct a pilot study on 2000 audio clips and define 14 metamorphic relations across two perturbation categories: Audio Features-Based and Heuristic perturbations. MTAM applies these metamorphic relations to toxic audio content to generate test cases that remain harmful while being more likely to evade detection. In our evaluation, we employ MTAM to test five commercial textual content moderation software and an academic model against three kinds of toxic content. The results show that MTAM achieves up to 38.6%, 18.3%, 35.1%, 16.7%, and 51.1% error finding rates (EFR) when testing commercial moderation software provided by Gladia, Assembly AI, Baidu, Nextdata, and Tencent respectively, and it obtains up to 45.7% EFR when testing the state-of-the-art algorithms from the academy. In addition, we leverage the test cases generated by MTAM to retrain the model we explored, which largely improves model robustness (nearly 0% EFR) while maintaining the accuracy on the original test set. We release the code and experiment data to facilitate future research1.
Wenxuan Wang 0001, Yongjiang Wu, Junyuan Zhang, Shuqing Li 0001, Yun Peng 0003, Wenting Chen, Shuai Wang 0011, Michael R. Lyu
ASE3
2025 Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
abstract
Visual tokens consume substantial computational resources in multi-modal large models (MLLMs), significantly compromising their efficiency. Recent works have attempted to improve efficiency by compressing visual tokens during training, either through modifications to model components or by introducing additional parameters. However, they often overlook the increased learning difficulty caused by such compression, as the model’s parameter space struggles to quickly adapt to the substantial perturbations in the feature space induced by token compression. In this work, we propose to develop Efficient MLLMs via Progressive Consistency Distillation (EPIC), a progressive learning framework. Specifically, by decomposing the feature space perturbations introduced by token compression along the token-wise and layer-wise dimensions, we introduce token consistency distillation and layer consistency distillation, respectively, aiming to reduce the training difficulty by leveraging guidance from a teacher model and following a progressive learning trajectory. Extensive experiments demonstrate the superior effectiveness, robustness, and generalization capabilities of our proposed framework.
Zichen Wen, Shaobo Wang 0001, Yufa Zhou 0001, Junyuan Zhang, Qintong Zhang, Yifeng Gao 0003, Zhaorun Chen, Bin Wang 0065, Conghui He, Linfeng Zhang 0001
NeurIPS4
2025 An explainable machine learning method for predicting and designing crashworthiness of multi-cell tubes under oblique load
Junyuan Zhang, Zheng Dou, Mengge Chang
Eng. Appl. Artif. Intell.2
2024 FLHetBench: Benchmarking Device and State Heterogeneity in Federated Learning
abstract
Federated learning (FL) is a powerful technology that enables collaborative training of machine learning models without sharing private data among clients. The fundamental challenge in FL lies in learning over extremely hetero-geneous data distributions, device capacities, and device state availabilities, all of which adversely impact performance and communication efficiency. While data hetero-geneity has been well-studied in the literature, this paper introduces FLHetBench, the first FL benchmark targeted toward understanding device and state heterogeneity. FL-HetBench comprises two new sampling methods to generate real-world device and state databases with varying het-erogeneity and new metrics for quantifying the success of FL methods under these real-world constraints. Using FL-HetBench, we conduct a comprehensive evaluation of existing methods and find that they struggle under these settings, which inspires us to propose BiasPrompt+, a new method employing staleness-aware aggregation and fast weights to tackle these new heterogeneity challenges. Experiments on various FL tasks and datasets validate the effectiveness of our BiasPrompt+ method and highlight the value of FLHet-Bench in fostering the development of more efficient and robust FL solutions under real-world device and state constraints.
Junyuan Zhang, Shuang Zeng, Miao Zhang 0030, Runxi Wang, Yuyin Zhou, Paul Pu Liang, Liangqiong Qu
CVPR1
2024 One-shot Federated Learning via Synthetic Distiller-Distillate Communication
abstract
One-shot Federated learning (FL) is a powerful technology facilitating collaborative training of machine learning models in a single round of communication. While its superiority lies in communication efficiency and privacy preservation compared to iterative FL, one-shot FL often compromises model performance. Prior research has primarily focused on employing data-free knowledge distillation to optimize data generators and ensemble models for better aggregating local knowledge into the server model. However, these methods typically struggle with data heterogeneity, where inconsistent local data distributions can cause teachers to provide misleading knowledge. Additionally, they may encounter scalability issues with complex datasets due to inherent two-step information loss: first, during local training (from data to model), and second, when transferring knowledge to the server model (from model to inversed data). In this paper, we propose FedSD2C, a novel and practical one-shot FL framework designed to address these challenges. FedSD2C introduces a distiller to synthesize informative distillates directly from local data to reduce information loss and proposes sharing synthetic distillates instead of inconsistent local models to tackle data heterogeneity. Our empirical results demonstrate that FedSD2C consistently outperforms other one-shot FL methods with more complex and real datasets, achieving up to 2.6 $\times$ the performance of the best baseline. Code: https://github.com/Carkham/FedSD2C
Junyuan Zhang, Songhua Liu, Xinchao Wang
NeurIPS1
2023 Fast instance selection method for SVM training based on fuzzy distance metric
Junyuan Zhang
Appl. Intell.1
2023 QUATRE-EMS: QUATRE algorithm with novel adaptation of evolution matrix and selection operation for numerical optimization
Zhenyu Meng, Junyuan Zhang
Inf. Sci.2
2022 GAN-Based Fusion Adversarial Training
Ying Lin 0004, Shengfu Ning, Huan Pi, Junyuan Zhang, Jianpeng Hu
KSEM (3)5
2022 PPBR-FL: A Privacy-Preserving and Byzantine-Robust Federated Learning System
Ying Lin 0004, Shengfu Ning, Jianpeng Hu, Jiansong Liu, Junyuan Zhang, Huan Pi
KSEM (3)6