Jiwen Zhang

dblp:87/2590 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
abstract
Evaluating multimodal large language models (MLLMs) is becoming increasingly expensive as benchmarks grow in scale and crossmodality complexity.Inspired by structuralism in cognitive psychology, we tackle this difficulty with an adaptive evaluation framework for efficient benchmarking, namely AutoJudger.Instead of passively scoring on a fixed test set, AutoJudger treats evaluation as an interviewlike process by keeping a hypothesized ability structure of the evaluated model and actively selecting the informative questions so as to refine these ability boundaries.Specifically, Auto-Judger has three core components: ability decomposition to organize evaluation along meaningful capability dimensions, ability estimation to maintain an up-to-date quantitative profile of the model competence, and adaptive question selection to choose the most informative questions.To operationalize this paradigm, we introduce A 2 -Judger, a novel MLLM-based Agentic instantiation of AutoJudger equipped with semantic-aware retrieval and dynamic memory.Experiments on four representative multimodal benchmarks show that A 2 -Judger significantly improves sample efficiency while maintaining reliable evaluation results.
Xuanwen Ding, Chengjun Pan, Jiwen Zhang, Zhongyu Wei
ACL (1)4
2026 MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution
abstract
Mobile GUI agents powered by large foundation models enable autonomous task execution in applications, but frequent updates that alter UI appearance and reorganize workflows cause agents trained on historical data to fail. Despite these surface changes, we observe that functional semantics and task intents remain fundamentally stable. Building on this insight, we introduce MAGNET, a memory-driven adaptive agent framework with dual-level memory: stationary memory that links diverse visual features to stable functional semantics for robust action grounding and procedural memory that captures stable task intents across varying workflows. Furthermore, we propose a dynamic memory evolution mechanism that continuously refines both memories by prioritizing frequently accessed knowledge. Evaluations on the online benchmark AndroidWorld demonstrate substantial improvements over memory-augmented baselines, while offline benchmarks confirm consistent gains under distribution shifts. These results validate that leveraging stable structures across interface changes improves agent performance and generalization in evolving software environments.
Jiwen Zhang, Zhongyu Wei
ACL (1)2
2026 A Robust and Compliant Robotic Assembly Control Strategy for Batch Precision Assembly Task
abstract
In many high-precision industrial applications, robots are deployed to perform precision peg-in-hole assembly on large batches of manufactured pegs and holes. When the nominal design adopts a transition fit, machining errors may cause each peg–hole pair to exhibit either a clearance fit or a slight interference fit, with an unknown and continuously varying fit amount. This paper addresses robotic batch precision assembly tasks under uncertain fit types and fit amounts, and proposes a novel multi-stage framework for systematically constructing a robust assembly control strategy. The overall batch precision assembly task is first decomposed into multiple deterministic subtasks characterized by different but fixed fit amounts. For these subtasks, a force-vision fusion controller-driven reinforcement learning method combined with a multi-task reinforcement learning training framework (FVFC-MTRL) is developed to jointly learn multiple compliance control strategies. Subsequently, a multi-teacher policy distillation scheme is designed to integrate multiple trained strategies into a single student strategy, yielding a robust control strategy. Real-world experiments demonstrate that the proposed method successfully constructs a robust control strategy for high-precision assembly tasks with varying fit types and fit amounts. Moreover, the MTRL framework significantly improves training efficiency by 50%, and the final control strategy achieves superior force compliance and a higher success rate compared with several existing methods.
Bin Wang 0096, Jiwen Zhang, Dan Wu 0008
IEEE Trans Autom. Sci. Eng.2
2025 UI-Hawk: Unleashing the Screen Stream Understanding for Mobile GUI Agents
abstract
Graphical User Interface (GUI) agents are expected to precisely operate on the screens of digital devices. Existing GUI agents merely depend on current visual observations and plain-text action history, ignoring the significance of history screens. To mitigate this issue, we propose UI-Hawk, a multi-modal GUI agent specially designed to process screen streams encountered during GUI navigation. UI-Hawk incorporates a history-aware visual encoder to handle the screen sequences. To acquire a better understanding of screen streams, we select four fundamental tasks—UI grounding, UI referring, screen question answering, and screen summarization. We further propose a curriculum learning strategy to subsequently guide the model from fundamental tasks to advanced screen-stream comprehension.Along with the efforts above, we have also created a benchmark FunUI to quantitatively evaluate the fundamental screen understanding ability of MLLMs. Extensive experiments on FunUI and GUI navigation benchmarks consistently validate that screen stream understanding is essential for GUI tasks.Our code and data are now available at https://github.com/IMNearth/UIHawk.
Jiwen Zhang, Ya-Qi Yu, Minghui Liao, WenTao Li, Jihao Wu, Zhongyu Wei
EMNLP1
2025 VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
abstract
Zejun Li, Ruipu Luo, Jiwen Zhang, Minghui Qiu, Xuanjing Huang, Zhongyu Wei. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ruipu Luo, Jiwen Zhang, Minghui Qiu, Xuanjing Huang 0001, Zhongyu Wei
NAACL (Long Papers)3
2024 DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive Learning
abstract
Vision-and-Language navigation (VLN) requires an agent to navigate in unseen environment by following natural language instruction. For task completion, the agent needs to align and integrate various navigation modalities, including instruction, observation and navigation history. Existing works primarily concentrate on cross-modal attention at the fusion stage to achieve this objective. Nevertheless, modality features generated by disparate uni-encoders reside in their own spaces, leading to a decline in the quality of cross-modal fusion and decision. To address this problem, we propose a Dual-levEL AligNment (DELAN) framework by cross-modal contrastive learning. This framework is designed to align various navigation-related modalities before fusion, thereby enhancing cross-modal interaction and action decision-making. Specifically, we divide the pre-fusion alignment into dual levels: instruction-history level and landmark-observation level according to their semantic correlations. We also reconstruct a dual-level instruction for adaptation to the dual-level alignment. As the training signals for pre-fusion alignment are extremely limited, self-supervised contrastive learning strategies are employed to enforce the matching between different modalities. Our approach seamlessly integrates with the majority of existing models, resulting in improved navigation performance on various VLN benchmarks, including R2R, R4R, RxR and CVDN.
Mengfei Du, Binhao Wu, Jiwen Zhang, Zhihao Fan, Ruipu Luo, Xuanjing Huang 0001, Zhongyu Wei
LREC/COLING3
2024 ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented Benchmarks
abstract
Recent years have witnessed remarkable progress in the development of large vision-language models (LVLMs). Benefiting from the strong language backbones and efficient cross-modal alignment strategies, LVLMs exhibit surprising capabilities to perceive visual signals and perform visually grounded reasoning. However, the capabilities of LVLMs have not been comprehensively and quantitatively evaluated. Most existing multi-modal benchmarks require task-oriented input-output formats, posing great challenges to automatically assess the free-form text output of LVLMs. To effectively leverage the annotations available and reduce the manual efforts required for constructing new benchmarks, we propose to re-formulate existing benchmarks into unified LVLM-compatible formats. Through systematic data collection and reformulation, we present ReForm-Eval benchmark, offering substantial data for evaluating various capabilities of LVLMs. Through extensive experiments and analysis in ReForm-Eval, we demonstrate the comprehensiveness and reliability of ReForm-Eval in assessing various LVLMs. Our benchmark and evaluation framework is now available at https://github.com/FudanDISC/ReForm-Eval
Mengfei Du, Qingwen Liu 0002, Binhao Wu, Jiwen Zhang, Chengxing Zhou, Zhihao Fan, Jie Fu 0001, Jingjing Chen 0001, Zhongyu Wei, Xuanjing Huang 0001
ACM Multimedia6
2024 An optimization-based motion planner for dual-arm manipulation of the soft deformable linear objects with nonnegligible gravity
Shirui Wu, Jiwen Zhang, Dan Wu 0008
Adv. Eng. Informatics2
2024 Robotic assembly control reconfiguration based on transfer reinforcement learning for objects with different geometric features
Yuhang Gai, Bin Wang 0096, Jiwen Zhang, Dan Wu 0008, Ken Chen 0002
Eng. Appl. Artif. Intell.3
2024 Local connection reinforcement learning method for efficient robotic peg-in-hole assembly
Yuhang Gai, Jiwen Zhang, Dan Wu 0008, Ken Chen 0002
Eng. Appl. Artif. Intell.2
2024 Learning from demonstration for 7-DOF anthropomorphic manipulators without offset via analytical inverse kinematics
Kui Hu, Jiwen Zhang
Neurocomputing2
2023 Kernelized gradient descent method for learning from demonstration
Kui Hu, Jiwen Zhang
Neurocomputing2
2021 Curriculum Learning for Vision-and-Language Navigation
abstract
Vision-and-Language Navigation (VLN) is a task where an agent navigates in an embodied indoor environment under human instructions. Previous works ignore the distribution of sample difficulty and we argue that this potentially degrade their agent performance. To tackle this issue, we propose a novel curriculum- based training paradigm for VLN tasks that can balance human prior knowledge and agent learning progress about training samples. We develop the principle of curriculum design and re-arrange the benchmark Room-to-Room (R2R) dataset to make it suitable for curriculum training. Experiments show that our method is model-agnostic and can significantly improve the performance, the generalizability, and the training efficiency of current state-of-the-art navigation agents without increasing model complexity.
Jiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie Peng
NeurIPS1
2020 Self-Supervised Learning for Specified Latent Representation
abstract
Current latent representation methods using unsupervised learning have no semantic meaning; thus, it is difficult to directly express their physical task in the real world. To this end, this paper attempts to propose a specified latent representation with physical semantic meaning. First, a few labeled samples are used to generate the framework of the latent space, and these labeled samples are mapped to framework nodes in the latent space. Second, a self-learning method using structured unlabeled samples is proposed to shape the free space between the framework nodes in the latent space. The proposed specified latent representation therefore possesses the advantages provided by both supervised and unsupervised learning. The proposed method is verified by numerical simulations and real-world experiments.
Chicheng Liu, Libin Song, Jiwen Zhang, Ken Chen 0002, Jing Xu 0011
IEEE Trans. Fuzzy Syst.3
2005 Unifying C-curves and H-curves by extending the calculation to complex numbers
Jiwen Zhang, Frank L. Krause, Huaiyu Zhang
Comput. Aided Geom. Des.1
2005 Extending cubic uniform B-splines by unified trigonometric and hyperbolic basis
Jiwen Zhang, Frank L. Krause
Graph. Model.1
1999 C-Bézier Curves and Surfaces
Jiwen Zhang
Graph. Model. Image Process.1
1997 Two different forms of C-B-splines
Jiwen Zhang
Comput. Aided Geom. Des.1