Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhennan Shen

dblp:355/4561 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 40% Language models and text generation · 36% Efficient and distributed learning · 18%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
multimodal in-context learning
1.012026
CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning · AAAI 2026
Machine learning › Efficient and distributed learning › model compression
token pruning
1.012026
CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning · AAAI 2026
Natural language and speech › Language models and text generation › LLM agents
computer-use agent
0.912025
OpenCUA: Open Foundations for Computer-Use Agents · NeurIPS 2025
Computer vision › Vision and language
vision-language model
0.912025
OpenCUA: Open Foundations for Computer-Use Agents · NeurIPS 2025
Human-AI interaction
GUI agent
0.912025
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials · ICLR 2025
Computing education
large language model evaluation
0.812024
SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research · AAAI 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312026
CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning · AAAI 2026
Knowledge, reasoning and agents › Multi-agent systems › multimodal agent
vision-language model agent
0.312025
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials · ICLR 2025
Natural language and speech › Language models and text generation
large language model
0.212024
SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research · AAAI 2024

Methods — techniques the papers use, named apart from their topics

guided replay · 1.7VLM-based evaluation · 1.7dynamic question generation · 1.5bloom's taxonomy · 1.5token pruning · 1.0cross-modal interaction · 1.0supervised fine-tuning · 0.9chain-of-thought reasoning · 0.9
YearPublicationVenuePosition
2026 CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
abstract
Modern large vision-language models (LVLMs) convert each input image into a large set of tokens that far outnumber the text tokens. Although this improves visual perception, it also introduces severe image token redundancy. Because image tokens contain sparse information, many contribute little to reasoning but greatly increase inference cost. Recent image token pruning methods address this issue by identifying important tokens and removing the rest. These methods improve efficiency with only small performance drops. However, most of them focus on single-image tasks and overlook multimodal in-context learning (ICL), where redundancy is higher and efficiency is more important. Redundant tokens weaken the advantage of multimodal ICL for rapid domain adaptation and lead to unstable performance. When existing pruning methods are applied in this setting, they cause large accuracy drops, which exposes a clear gap and the need for new approaches. To address this, we propose Contextually Adaptive Token Pruning (CATP), a training-free pruning method designed for multimodal ICL. CATP uses two stages of progressive pruning that fully reflect the complex cross-modal interactions in the input sequence. After removing 77.8% of the image tokens, CATP achieves an average performance gain of 0.6% over the vanilla model on four LVLMs and eight benchmarks, clearly outperforming all baselines. At the same time, it improves efficiency by reducing inference latency by an average of 10.78%. CATP strengthens the practical value of multimodal ICL and lays the foundation for future progress in interleaved image-text settings.
Yanshu Li, Jianjiang Yang, Zhennan Shen, Ligong Han, Haoyan Xu, Ruixiang Tang
AAAI3
2025 AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
abstract
Graphical User Interface (GUI) agents hold great potential for automating complex tasks across diverse digital environments, from web applications to desktop software. However, the development of such agents is hindered by the lack of high-quality, multi-step trajectory data required for effective training. Existing approaches rely on expensive and labor-intensive human annotation, making them unsustainable at scale. To address this challenge, we propose AgentTrek, a scalable data synthesis pipeline that generates high-quality web agent trajectories by leveraging web tutorials. Our method automatically gathers tutorial-like texts from the internet, transforms them into task goals with step-by-step instructions, and employs a visual-language model (VLM) agent to simulate their execution in a real digital environment. A VLM-based evaluator ensures the correctness of the generated trajectories. We demonstrate that training GUI agents with these synthesized trajectories significantly improves their grounding and planning performance over the current models. Moreover, our approach is more cost-efficient compared to traditional human annotation methods. This work underscores the potential of guided replay with web tutorials as a viable strategy for large-scale GUI agent training, paving the way for more capable and autonomous digital agents.
Yiheng Xu, Dunjie Lu, Zhennan Shen, Caiming Xiong, Tao Yu 0009
ICLR3
2025 OpenCUA: Open Foundations for Computer-Use Agents
abstract
Vision-language models have demonstrated impressive capabilities as computer-use agents (CUAs) capable of automating diverse computer tasks. As their commercial potential grows, critical details of the most capable CUA systems remain closed. As these agents will increasingly mediate digital interactions and execute consequential decisions on our behalf, the research community needs access to open CUA frameworks to study their capabilities, limitations, and risks. To bridge this gap, we propose OpenCUA, a comprehensive open-source framework for scaling CUA data and foundation models. Our framework consists of: (1) an annotation infrastructure that seamlessly captures human computer-use demonstrations; (2) AgentNet, the first large-scale computer-use task dataset spanning 3 operating systems and 200+ applications and websites; (3) a scalable pipeline that transforms demonstrations into state–action pairs with reflective long Chain-of-Thought reasoning that sustain robust performance gains as data scales. Our end-to-end agent models demonstrate strong performance across CUA benchmarks. In particular, OpenCUA-72B achieves an average success rate of 45.0% on OSWorld‑Verified, establishing a new state-of-the-art (SOTA) among open-source models. Further analysis confirms that our approach generalizes well across domains and benefits significantly from increased test-time computation. We release our annotation tool, datasets, code, and models to build open foundations for further CUA research.
Xinyuan Wang 0010, Dunjie Lu, Junlin Yang, Tianbao Xie, Xiaole Guo, Yiheng Xu, Chen Henry Wu, Zhennan Shen, Zhuokai Li, Xiaochuan Li 0003, Junda Chen, Boyuan Zheng 0001, Peihang Li, Fangyu Lei, Ruisheng Cao, Yeqiao Fu, Dongchan Shin, Martin Shin, Jiarui Hu 0006, Jixuan Chen, Yuxiao Ye, Yipu Wang, Diyi Yang, Victor Zhong, Y. Charles, Tao Yu 0009
NeurIPS11
2024 SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research
abstract
Recently, there has been growing interest in using Large Language Models (LLMs) for scientific research. Numerous benchmarks have been proposed to evaluate the ability of LLMs for scientific research. However, current benchmarks are mostly based on pre-collected objective questions. This design suffers from data leakage problem and lacks the evaluation of subjective Q/A ability. In this paper, we propose SciEval, a comprehensive and multi-disciplinary evaluation benchmark to address these issues. Based on Bloom's taxonomy, SciEval covers four dimensions to systematically evaluate scientific research ability. In particular, we design a "dynamic" subset based on scientific principles to prevent evaluation from potential data leakage. Both objective and subjective questions are included in SciEval. These characteristics make SciEval a more effective benchmark for scientific research ability evaluation of LLMs. Comprehensive experiments on most advanced LLMs show that, although GPT-4 achieves SOTA performance compared to other LLMs, there is still substantial room for improvement, especially for dynamic questions. The codes and data are publicly available on https://github.com/OpenDFM/SciEval.
Liangtai Sun, Yang Han 0007, Zihan Zhao 0001, Zhennan Shen, Baocai Chen, Lu Chen 0002, Kai Yu 0004
AAAI5