Jiacheng Cui

dblp:268/3316 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0005-4048-709XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LLMSurgeon: Diagnosing Data Mixture of Large Language Models
abstract
Yaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang, Jiacheng Liu, Xinyue Bi, Zhaoyi Li, Zhiqiang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang, Xinyue Bi
ACL (1)2
2026 FloodGuard: A Prediction-Control Closed Loop for Mitigating Cold-Start Floods in Cloud Services
Jiacheng Cui, Junyu Xue, Guoming Tang
ICDCS1
2026 BFNet: A real-time edge-deployable dual-stream boundary-aware network for defect detection of aquatic photovoltaic systems
abstract
In the industrial inspection of aquatic photovoltaic (PV) systems, semantic segmentation faces two critical challenges: the extremely low proportion of defect pixels relative to the overall image and the restrictive edge deployment conditions, which necessitate real-time inference on resource-limited devices. This paper introduces BFNet, a boundary-aware architecture employing a dual-stream design to distinctly separate boundary features from semantic representations, thereby addressing the gradient dominance issue caused by majority classes. The main contributions include: (1) a learnable boundary gating mechanism for adaptive edge enhancement; (2) Dilated Dense Blocks that significantly reduce parameter volume while preserving receptive field coverage; and (3) a multi-task training approach weighted by inverse square root frequency. Experimental evaluations demonstrate that BFNet achieves defect detection accuracy statistically comparable to heavyweight methods, yet with substantially fewer parameters, enabling real-time deployment on battery-powered unmanned surface vessels.
Jiacheng Cui, Zizhen Zhao, Erchao Fang, Shinan Zhao, Jianlin Gao, Yujiang Hong
Appl. Intell.2
2025 Augmented Reality-Enabled Interaction of Intrinsic Global Physical Information for Aircraft Assembly
abstract
Enhanced transparency of the assembly process is critical for intelligent aerospace manufacturing. While augmented reality (AR) technology provides valuable support for assisted assembly, existing systems are limited in their ability to deliver real-time, intuitive visualization of global, multi-dimensional physical states—such as stress distribution, dynamic strain fields, and multi-component coupled deformations. To overcome these challenges, this study proposes an AR-assisted system for real-time global physical information interaction and multi-dimensional assembly guidance. By integrating multi-view visual measurement with finite element analysis, and leveraging model order reduction and optimal sensor placement strategies, the system achieves real-time reconstruction of global displacement, stress, and strain fields. These reconstructed physical fields are overlaid onto the actual assembly environment through AR, establishing a closed-loop "measurement–perception–feedback" interaction. Experiments on aerospace composite thin-walled structures validate the system’s capability for real-time perception and interactive visualization. Compared with conventional AR systems featuring unidirectional data flow, the proposed approach significantly enhances process transparency and controllability, offering a new paradigm for intelligent assembly.
Yulin Jin, Jiacheng Cui, Yongkang Lu, Yang Zhang 0011
INDIN3
2025 FADRM: Fast and Accurate Data Residual Matching for Dataset Distillation
abstract
Residual connection has been extensively studied and widely applied at the model architecture level. However, its potential in the more challenging data-centric approaches remains unexplored. In this work, we introduce the concept of ***Data Residual Matching*** for the first time, leveraging data-level skip connections to facilitate data generation and mitigate data information vanishing. This approach maintains a balance between newly acquired knowledge through pixel space optimization and existing core local information identification within raw data modalities, specifically for the dataset distillation task. Furthermore, by incorporating training-time refinements, our method significantly improves computational efficiency, achieving superior performance while reducing training time and peak GPU memory usage by 50\%. Consequently, the proposed method **F**ast and **A**ccurate **D**ata **R**esidual **M**atching for Dataset Distillation (**FADRM**) establishes a new state-of-the-art, demonstrating substantial improvements over existing methods across multiple dataset benchmarks in both efficiency and effectiveness. For instance, with ResNet-18 as the student model and a 0.8\% compression ratio on ImageNet-1K, the method achieves 48.4\% test accuracy in single-model dataset distillation and 50.9\% in multi-model dataset distillation, surpassing RDED by +6.4\% and outperforming state-of-the-art multi-model approaches, EDC and CV-DD, by +2.3\% and +4.9\%.
Jiacheng Cui, Xinyue Bi, Yaxin Luo, Xiaohan Zhao
NeurIPS1
2025 A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1
abstract
Despite promising performance on open-source large vision-language models (LVLMs), transfer-based targeted attacks often fail against closed-source commercial LVLMs. Analyzing failed adversarial perturbations reveals that the learned perturbations typically originate from a uniform distribution and lack clear semantic details, resulting in unintended responses. This critical absence of semantic information leads commercial black-box LVLMs to either ignore the perturbation entirely or misinterpret its embedded semantics, thereby causing the attack to fail. To overcome these issues, we propose to refine semantic clarity by encoding explicit semantic details within local regions, thus ensuring the capture of finer-grained features and inter-model transferability, and by concentrating modifications on semantically rich areas rather than applying them uniformly. To achieve this, we propose *a simple yet highly effective baseline*: at each optimization step, the adversarial image is cropped randomly by a controlled aspect ratio and scale, resized, and then aligned with the target image in the embedding space. While the naive source-target matching method has been utilized before in the literature, we are the first to provide a tight analysis, which establishes a close connection between perturbation optimization and semantics. Experimental results confirm our hypothesis. Our adversarial examples crafted with local-aggregated perturbations focused on crucial regions exhibit surprisingly good transferability to commercial LVLMs, including GPT-4.5, GPT-4o, Gemini-2.0-flash, Claude-3.5/3.7-sonnet, and even reasoning models like o1, Claude-3.7-thinking and Gemini-2.0-flash-thinking. Our approach achieves success rates exceeding 90\% on GPT-4.5, 4o, and o1, significantly outperforming all prior state-of-the-art attack methods with lower $\ell_1/\ell_2$ perturbations. Our optimized adversarial examples under different configurations are available at https://huggingface.co/datasets/MBZUAI-LLM/M-Attack_AdvSamples and our training code at https://github.com/VILA-Lab/M-Attack.
Xiaohan Zhao, Dong-Dong Wu, Jiacheng Cui
NeurIPS4
2025 Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
abstract
CAPTCHAs have been a critical bottleneck for deploying web agents in real-world applications, often blocking them from completing end-to-end automation tasks. While modern multimodal LLM agents have demonstrated impressive performance in static perception tasks, their ability to handle interactive, multi-step reasoning challenges like CAPTCHAs is largely untested. To address this gap, we introduce Open CaptchaWorld, the first web-based benchmark and platform specifically designed to evaluate the visual reasoning and interaction capabilities of MLLM-powered agents through diverse and dynamic CAPTCHA puzzles. Our benchmark spans 20 modern CAPTCHA types, totaling 225 CAPTCHAs, annotated with a new metric we propose: CAPTCHA Reasoning Depth, which quantifies the number of cognitive and motor steps required to solve each puzzle. Experimental results show that humans consistently achieve near-perfect scores, state-of-the-art MLLM agents struggle significantly, with success rates at most 40.0\% by Browser-Use Openai-o3, far below human-level performance,93.3\%. This highlights Open CaptchaWorld as a vital benchmark for diagnosing the limits of current multimodal agents and guiding the development of more robust multimodal reasoning systems.
Yaxin Luo, Jiacheng Cui, Xiaohan Zhao
NeurIPS4
2023 SMRSTORE: A Storage Engine for Cloud Object Storage on HM-SMR Drives
Erci Xu, Jiacheng Cui, Wanyu Fu, Yingni Wang, Shouqu Sun, Xianfei Wang, Biyun Zhu, Weikang Kong, Linyan Liu, Zhongjie Wu, Qingchao Luo, Jiesheng Wu
FAST5
2022 Spectral knowledge-based regression for laser-induced breakdown spectroscopy quantitative analysis
Weiran Song, Muhammad Sher Afgan, Yong-Huan Yun, Hui Wang 0001, Jiacheng Cui, Weilun Gu, Zongyu Hou
Expert Syst. Appl.5