Zhaokai Wang

dblp:232/7086 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
abstract
Diffusion Transformers (DiTs) are a powerful yet underexplored class of generative models compared to U-Net-based diffusion architectures. We propose TIDE—Temporal-aware sparse autoencoders for Interpretable Diffusion transformErs—a framework designed to extract sparse, interpretable activation features across timesteps in DiTs. TIDE effectively captures temporally-varying representations and reveals that DiTs naturally learn hierarchical semantics (e.g., 3D structure, object class, and fine-grained concepts) during large-scale pretraining. Experiments show that TIDE enhances interpretability and controllability while maintaining reasonable generation quality, enabling applications such as safe image editing and style transfer.
Victor Shea-Jay Huang, Le Zhuo, Yi Xin 0003, Zhaokai Wang, Fu-Yun Wang, Yuchi Wang, Renrui Zhang, Peng Gao 0007, Hongsheng Li 0001
AAAI4
2026 A Needle Biopsy-Inspired Method for Rapid Mechanical Characterization of Soft Tissue
Zhaokai Wang, Syed Ali Raza Bukhari, Tiffany Yu, Diancheng Li, Matthew A. Robertson, Andrew Giles, Jamie Purzner, Yong Jun Lai, Xian Wang 0001
IEEE Trans Autom. Sci. Eng.1
2026 A Magnetic Capsule for Navigation and Multitargeted Sampling in the Gastrointestinal Tract
abstract
Untethered capsules are capable of entering the gastrointestinal (GI) tract and collecting fluid samples containing microbial communities from specific locations, facilitating the study of chronic diseases. However, existing sampling capsules are designed for single-site sampling, making it challenging to gather samples from multiple targets. This paper reports a magnetic-driven capsule for multiple sampling within the GI tract and an on-demand magnetic-triggered fluid sampling strategy. The capsule consists of a body, a magnetic-triggered negative pressure unit, and a reservoir unit. Composed of an elastic membrane and Magnet I, the negative pressure unit controls pressure change inside the capsule cavity on demand to pump the sample by switching the magnetic field, while the embedded Magnet I also enables real-time magnetic localization for regional targeting and position tracking. The reservoir unit integrates three sampling papers for fluid absorption, two waterproof layers that maintain contamination levels below 25% to ensure reliable multi-site sampling, and a rotating arm embedded with Magnet II for posture adjustment of the sampling paper. The pumping and storage performance of the capsule was systematically evaluated and optimized. Meanwhile, the capsule, actuated by an external magnetic field, was evaluated for its active locomotion performance. Finally, the feasibility of using the capsule to perform active navigation and multi-target sampling in a porcine intestine was validated viaex vivoexperiments.
Huayang Ren, Zhaokai Wang, Jingfang Han, Jiaqing Xie, Ruicheng Li, Chunyun Wei, Tao Yue 0001, Yue Wang 0110, Yan Peng 0001, Jiangfan Yu, Xian Wang 0001, Na Liu 0004, Yu Sun 0001
IEEE Trans. Robotics3
2025 OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use
abstract
Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shenzhi Wang, Xinchen Xu, Shuofei Qiao, Zhaokai Wang, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang, Keting Yin, Zhou Zhao, Hongxia Yang, Fan Wu, Shengyu Zhang, Fei Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xueyu Hu, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen 0004, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao 0001, Yuhuai Li, Shengze Xu, Shenzhi Wang, Shuofei Qiao, Zhaokai Wang, Kun Kuang 0001, Tieyong Zeng, Liang Wang 0001, Jiwei Li 0001, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang 0002, Keting Yin, Zhou Zhao 0001, Hongxia Yang, Fan Wu 0006, Shengyu Zhang 0001, Fei Wu 0001
ACL (1)16
2025 SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
abstract
The remarkable success of Large Language Models (LLMs) has extended to the multimodal domain, achieving outstanding performance in image understanding and generation. Recent efforts to develop unified Multimodal Large Language Models (MLLMs) that integrate these capabilities have shown promising results. However, existing approaches often involve complex designs in model architecture or training pipeline, increasing the difficulty of model training and scaling. In this paper, we propose SynerGen-VL, a simple yet powerful encoder-free MLLM capable of both image understanding and generation. To address challenges identified in existing encoder-free unified MLLMs, we introduce the token folding mechanism and the vision-expert-based progressive alignment pretraining strategy, which effectively support high-resolution image understanding while reducing training complexity. After being trained on large-scale mixed image-text data with a unified next-token prediction objective, SynerGen-VL achieves or surpasses the performance of existing encoder-free unified MLLMs with comparable or smaller parameter sizes, and narrows the gap with task-specific state-of-the-art models, highlighting a promising path toward future unified MLLMs. Our code and models are released at https://github.com/cpsxhao/SynerGen-VL.
Hao Li 0069, Changyao Tian, Xizhou Zhu, Zhaokai Wang, Jinguo Zhu, Wenhan Dou, Xiaogang Wang 0001, Hongsheng Li 0001, Lewei Lu, Jifeng Dai
CVPR5
2025 Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
abstract
In this paper, we focus on monolithic Multimodal Large Language Models (MLLMs) that integrate visual encoding and language decoding into a single LLM. In particular, we identify that existing pre-training strategies for monolithic MLLMs often suffer from unstable optimization or catastrophic forgetting. To address this issue, our core idea is to embed a new visual parameter space into a pre-trained LLM, thereby stably learning visual knowledge from noisy data while freezing the LLM. Based on this principle, we present Mono-InternVL, a novel monolithic MLLM that seamlessly integrates a set of visual experts via a multimodal mixture-of-experts structure. Moreover, we propose an innovative pre-training strategy to maximize the visual capability of Mono-InternVL, namely Endogenous Visual Pre-training (EViP). In particular, EViP is designed as a progressive learning process for visual experts, which aims to fully exploit the visual knowledge from noisy data to high-quality data. To validate our approach, we conduct extensive experiments on 16 benchmarks. Experimental results confirm the superior performance of Mono-InternVL than existing monolithic MLLMs on 13 of 16 multimodal benchmarks, e.g., +80 points over Emu3 on OCRBench. Compared to the modular baseline, i.e., InternVL-1.5, Mono-InternVL still retains comparable multimodal performance while reducing up to 67% first token latency. Our project is available at https://internvl.github.io/blog/2024-10-10Mono-InternVL/.
Gen Luo, Xue Yang 0005, Wenhan Dou, Zhaokai Wang, Jifeng Dai, Yu Qiao 0001, Xizhou Zhu
CVPR4
2025 Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has recently demonstrated notable success in enhancing the reasoning performance of large language models (LLMs), particularly in mathematics and programming tasks. It is widely believed that, similar to how traditional RL helps agents to explore and learn new strategies, RLVR enables LLMs to continuously self-improve, thus acquiring novel reasoning abilities that exceed the capacity of the corresponding base models. In this study, we take a critical look at \textit{the current state of RLVR} by systematically probing the reasoning capability boundaries of RLVR-trained LLMs across diverse model families, RL algorithms, and math/coding/visual reasoning benchmarks, using pass@\textit{k} at large \textit{k} values as the evaluation metric. While RLVR improves sampling efficiency towards the correct path, we surprisingly find that current training does \emph{not} elicit fundamentally new reasoning patterns. We observe that while RLVR-trained models outperform their base models at smaller values of $k$ (\eg, $k$=1), base models achieve higher pass@$k$ score when $k$ is large. Moreover, we observe that the reasoning capability boundary of LLMs often narrows as RLVR training progresses. Further coverage and perplexity analysis shows that the reasoning paths generated by RLVR models are already included in the base models' sampling distribution, suggesting that their reasoning abilities originate from and are \textit{bounded} by the base model. From this perspective, treating the base model as an upper bound, our quantitative analysis shows that six popular RLVR algorithms perform similarly and remain far from optimal in fully leveraging the potential of the base model. In contrast, we find that distillation can introduce new reasoning patterns from the teacher and genuinely expand the model’s reasoning capabilities. Taken together, our findings suggest that current RLVR methods have not fully realized the potential of RL to elicit genuinely novel reasoning abilities in LLMs. This underscores the need for improved RL paradigms—such as continual scaling and multi-turn agent-environment interaction—to unlock this potential.
Rui Lu 0001, Andrew Zhao, Zhaokai Wang, Shiji Song, Gao Huang 0001
NeurIPS4
2025 Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
abstract
Image pyramids are widely adopted in top-performing methods to obtain multi-scale features for precise visual perception and understanding. However, current image pyramids use the same large-scale model to process multiple resolutions of images, leading to significant computational cost. To address this challenge, we propose a novel network architecture, called Parameter-Inverted Image Pyramid Networks (PIIP). Specifically, PIIP uses pretrained models (ViTs or CNNs) as branches to process multi-scale images, where images of higher resolutions are processed by smaller network branches to balance computational cost and performance. To integrate information from different spatial scales, we further propose a novel cross-branch feature interaction mechanism. To validate PIIP, we apply it to various perception models and a representative multimodal large language model called LLaVA, and conduct extensive experiments on various tasks such as object detection, segmentation, image classification and multimodal understanding. PIIP achieves superior performance compared to single-branch and existing multi-resolution approaches with lower computational cost. When applied to InternViT-6B, a large-scale vision foundation model, PIIP can improve its performance by 1%-2% on detection and segmentation with only 40%-60% of the original computation, finally achieving 60.0 box AP on MS COCO and 59.7 mIoU on ADE20 K. For multimodal understanding, our PIIP-LLaVA achieves 73.0% accuracy on TextVQA and 74.5% on MMBench with only 2.8 M training data.
Zhaokai Wang, Xizhou Zhu, Xue Yang 0005, Gen Luo, Hao Li 0069, Changyao Tian, Wenhan Dou, Junqi Ge, Lewei Lu, Yu Qiao 0001, Jifeng Dai
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
abstract
Many reinforcement learning environments (e.g., Minecraft) provide only sparse rewards that indicate task completion or failure with binary values. The challenge in exploration efficiency in such environments makes it difficult for reinforcement-learning-based agents to learn complex tasks. To address this, this paper introduces an advanced learning system, named Auto MC-Reward, that leverages Large Language Models (LLMs) to automatically design dense reward functions, thereby enhancing the learning efficiency. Auto MC-Reward consists of three important components: Reward Designer, Reward Critic, and Trajectory Analyzer. Given the environment information and task descriptions, the Reward Designer first design the reward function by coding an executable Python function with predefined observation inputs. Then, our Reward Critic will be responsible for verifying the code, checking whether the code is self-consistent and free of syntax and semantic errors. Further, the Trajectory Analyzer summarizes possible failure causes and provides refinement suggestions according to collected trajectories. In the next round, Reward Designer will further refine and iterate the dense reward function based on feedback. Experiments demonstrate a significant improvement in the success rate and learning efficiency of our agents in complex tasks in Minecraft, such as obtaining diamond with the efficient ability to avoid lava, and efficiently explore trees and animals that are sparse in the plains biome.
Hao Li 0069, Xue Yang 0005, Zhaokai Wang, Xizhou Zhu, Jie Zhou 0001, Yu Qiao 0001, Xiaogang Wang 0001, Hongsheng Li 0001, Lewei Lu, Jifeng Dai
CVPR3
2024 Parameter-Inverted Image Pyramid Networks
abstract
Image pyramids are commonly used in modern computer vision tasks to obtain multi-scale features for precise understanding of images. However, image pyramids process multiple resolutions of images using the same large-scale model, which requires significant computational cost. To overcome this issue, we propose a novel network architecture known as the Parameter-Inverted Image Pyramid Networks (PIIP). Our core idea is to use models with different parameter sizes to process different resolution levels of the image pyramid, thereby balancing computational efficiency and performance. Specifically, the input to PIIP is a set of multi-scale images, where higher resolution images are processed by smaller networks. We further propose a feature interaction mechanism to allow features of different resolutions to complement each other and effectively integrate information from different spatial scales. Extensive experiments demonstrate that the PIIP achieves superior performance in tasks such as object detection, segmentation, and image classification, compared to traditional image pyramid methods and single-branch networks, while reducing computational cost. Notably, when applying our method on a large-scale vision foundation model InternViT-6B, we improve its performance by 1\%-2\% on detection and segmentation with only 40\%-60\% of the original computation. These results validate the effectiveness of the PIIP approach and provide a new technical direction for future vision computing tasks.
Xizhou Zhu, Xue Yang 0005, Zhaokai Wang, Hao Li 0069, Wenhan Dou, Junqi Ge, Lewei Lu, Yu Qiao 0001, Jifeng Dai
NeurIPS3
2023 Video Background Music Generation: Dataset, Method and Evaluation
abstract
Music is essential when editing videos, but selecting music manually is difficult and time-consuming. Thus, we seek to automatically generate background music tracks given video input. This is a challenging task since it requires music-video datasets, efficient architectures for video-to-music generation, and reasonable metrics, none of which currently exist. To close this gap, we introduce a complete recipe including dataset, benchmark model, and evaluation metric for video background music generation. We present SymMV, a video and symbolic music dataset with various musical annotations. To the best of our knowledge, it is the first video-music dataset with rich musical annotations. We also propose a benchmark video background music generation framework named V-MusProd, which utilizes music priors of chords, melody, and accompaniment along with video-music relations of semantic, color, and motion features. To address the lack of objective metrics for video-music correspondence, we design a retrieval-based metric VMCP built upon a powerful video-music representation learning model. Experiments show that with our dataset, V-MusProd outperforms the state-of-the-art method in both music quality and correspondence with videos. We believe our dataset, benchmark model, and evaluation metric will boost the development of video background music generation. Our dataset and code are available at https://github.com/zhuole1025/SymMV.
Le Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao, Chenxi Bao, Stanley Peng, Songhao Han, Aixi Zhang, Fei Fang 0002, Si Liu 0001
ICCV2
2021 Confidence-aware Non-repetitive Multimodal Transformers for TextCaps
abstract
When describing an image, reading text in the visual scene is crucial to understand the key information. Recent work explores the TextCaps task, i.e. image captioning with reading Optical Character Recognition (OCR) tokens, which requires models to read text and cover them in generated captions. Existing approaches fail to generate accurate descriptions because of their (1) poor reading ability; (2) inability to choose the crucial words among all extracted OCR tokens; (3) repetition of words in predicted captions. To this end, we propose a Confidence-aware Non-repetitive Multimodal Transformers (CNMT) to tackle the above challenges. Our CNMT consists of a reading, a reasoning and a generation modules, in which Reading Module employs better OCR systems to enhance text reading ability and a confidence embedding to select the most noteworthy tokens. To address the issue of word redundancy in captions, our Generation Module includes a repetition mask to avoid predicting repeated word in captions. Our model outperforms state-of-the-art models on TextCaps dataset, improving from 81.0 to 93.0 in CIDEr. Our source code is publicly available.
Zhaokai Wang, Renda Bao, Qi Wu 0001, Si Liu 0001
AAAI1
2021 Video Background Music Generation with Controllable Music Transformer
abstract
In this work, we address the task of video background music generation. Some previous works achieve effective music generation but are unable to generate melodious music specifically for a given video, and none of them considers the video-music rhythmic consistency. To generate the background music that matches the given video, we first establish the rhythmic relationships between video and background music. In particular, we connect timing, motion speed, and motion saliency from video with beat, simu-note density, and simu-note strength from music, respectively. We then propose CMT, a Controllable Music Transformer that enables the local control of the aforementioned rhythmic features, as well as the global control of the music genre and the used instrument specified by users. Objective and subjective evaluations show that the generated background music has achieved satisfactory compatibility with the input videos, and at the same time, impressive music quality.
Shangzhe Di, Zeren Jiang, Si Liu 0001, Zhaokai Wang, Leyan Zhu, Zexin He, Shuicheng Yan
ACM Multimedia4
2019 LogGOPSC: A Parallel Computation Model Extending Network Contention into LogGOPS
abstract
Benefits from the simplicity, fastness and accuracy, the LogP model family is widely used to predict the parallel application communication performance, especial for the large-scale parallel application prediction or online prediction. However, this type of methods usually lacks consideration of modeling the network contention effect. This hinders their performance in predicting some real-world parallel applications. We thus propose a new parallel computation model via extending the LogGOPS mode with a parameter C. The additional parameter C is the additional time overhead caused by the network contention and it is predicted by a designed BP neural network. The experimental results show that LogGOPSC is more accurate than LogGOPS when there occur network contentions. Compared to the LogGOPS model, LogGOPSC gains a 93.25% average accuracy improvement for predicting 8mb point-to-point message passing. Furthermore, the average error of predicting two communication patterns is as low as 14.50% on the TianHe-2 HPC system.
Baicheng Yan, Jiantong Huo, Zhaokai Wang
CLUSTER5
2019 Deeper Monocular Depth Prediction via Long and Short Skip Connection
abstract
This paper presents a fully convolutional neural network to tackle the mapping between single view RGB images and depth maps. To regress the depth maps from monocular images, we leverage deep short skip connections in residual learning for extracting features rather than using hand-crafted features. We further propose long skip connections in up-sampling stage to reuse the feature maps which is proved to enhance the result experimentally. To show the impact of loss functions in monocular depth map predictions, we train our model with kind of loss functions and compare the results qualitatively and quantitatively. The proposed model outperforms all current state-of-the-art results with less training data as well as less than half of training epochs in two standard benchmark data sets without any post-processing procedures or other refinement steps.
Zhaokai Wang, Rongbin Xu, Shubin Su, Shupan Li
IJCNN1
2019 An adaptive template matching-based single object tracking algorithm with parallel acceleration
Baicheng Yan, Daliang Xu, Zhaokai Wang
J. Vis. Commun. Image Represent.6