EDBT 2026 Demo / reviewers in the wild / expert
Bowen Zhang 0002
dblp:85/7433-2
· DBLP profile ↗
23ranked-venue papers
6as first author
15since 2021 · last 2025
0000-0002-4971-4878ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improve Vision Language Model Chain-of-thought ReasoningabstractRuohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang, Zhiqing Sun, Zhe Gan, Yinfei Yang, Ruoming Pang, Yiming Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruohong Zhang, Bowen Zhang 0002, Yanghao Li, Haotian Zhang 0005, Zhiqing Sun, Zhe Gan, Yinfei Yang, Ruoming Pang, Yiming Yang 0002 |
ACL (1) | 2 |
| 2025 | STIV: Scalable Text and Image Conditioned Video GenerationabstractThe field of video generation has made remarkable advancements, yet there remains a pressing need for a clear, systematic recipe that can guide the development of robust and scalable models. In this work, we present a comprehensive study that systematically explores the interplay of model architectures, training recipes, and data curation strategies, culminating in a simple and scalable text-image-conditioned video generation method, named STIV. Our framework integrates image condition into a Diffusion Transformer (DiT) through frame replacement, while incorporating text conditioning via a joint image-text conditional classifier-free guidance. This design enables STIV to perform both text-to-video (T2V) and text-image-to-video (TI2V) tasks simultaneously. Additionally, STIV can be easily extended to various applications, such as video prediction, frame interpolation, multi-view generation, and long video generation, etc. With comprehensive ablation studies on T2I, T2V, and TI2V, STIV demonstrate strong performance, despite its simple design. An 8.7B model with 512 resolution achieves 83.1 on VBench T2V, surpassing both leading open and closed-source models like CogVideoX-5B, Pika, Kling, and Gen-3. The same-sized model also achieves a state-of-the-art result of 90.1 on VBench I2V task at 512 resolution. By providing a transparent and extensible recipe for building cutting-edge video generation models, we aim to empower future research and accelerate progress toward more versatile and reliable video generation solutions. Zongyu Lin, Chen Chen 0005, Jiasen Lu, Wenze Hu, Tsu-Jui Fu, Jesse Allardice, Zhengfeng Lai, Liangchen Song, Bowen Zhang 0002, Cha Chen, Yiran Fei, Lezhi Li, Yinfei Yang, Yizhou Sun, Kai-Wei Chang 0001 |
ICCV | 10 |
| 2025 | Gaussian Variation Field Diffusion for High-Fidelity Video-to-4D SynthesisabstractIn this paper, we present a novel framework for video-to-4D generation that creates high-quality dynamic 3D content from single video inputs. Direct 4D diffusion modeling is extremely challenging due to costly data construction and the high-dimensional nature of jointly representing 3D shape, appearance, and motion. We address these challenges by introducing a Direct 4DMesh-to-GS Variation Field VAE that directly encodes canonical Gaussian Splats (GS) and their temporal variations from 3D animation data without per-instance fitting, and compresses high-dimensional animations into a compact latent space. Building upon this efficient representation, we train a Gaussian Variation Field diffusion model with temporal-aware Diffusion Transformer conditioned on input videos and canonical GS. Trained on carefully-curated animatable 3D objects from the Objaverse dataset, our model demonstrates superior generation quality compared to existing methods. It also exhibits remarkable generalization to in-the-wild video inputs despite being trained exclusively on synthetic data, paving the way for generating high-quality animated 3D content. Project page: https://gvfdiffusion.github.io/. Bowen Zhang 0002, Sicheng Xu, Chuxin Wang, Jiaolong Yang, Feng Zhao 0004, Dong Chen 0003, Baining Guo |
ICCV | 1 |
| 2025 | Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation ModelsabstractRecent advancements in multimodal models highlight the value of rewritten captions for improving performance, yet key challenges remain. For example, while synthetic captions often provide superior quality and image-text alignment, it is not clear whether they can fully replace AltTexts: the role of synthetic captions and their interaction with original web-crawled AltTexts in pre-training is still not well understood. Moreover, different multimodal foundation models may have unique preferences for specific caption formats, but efforts to identify the optimal captions for each model remain limited. In this work, we propose a novel, controllable, and scalable captioning pipeline designed to generate diverse caption formats tailored to various multimodal models. By examining short synthetic captions (SSC) and descriptive synthetic captions (DSC) as case studies, we systematically explore their effects and interactions with AltTexts across models such as CLIP, multimodal LLMs, and diffusion models. Our findings reveal that a hybrid approach that keeps both synthetic captions and AltTexts can outperform the use of synthetic captions alone, improving both alignment and performance, with each model demonstrating preferences for particular caption formats. This comprehensive analysis provides valuable insights into optimizing captioning strategies, thereby advancing the pre-training of multimodal foundation models. Zhengfeng Lai, Vasileios Saveris, Chen Chen 0005, Hong-You Chen, Haotian Zhang 0005, Bowen Zhang 0002, Wenze Hu, Juan Lao Tebar, Zhe Gan, Peter Grasch, Yinfei Yang |
ICLR | 6 |
| 2025 | EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice RoutingabstractDiffusion transformers have been widely adopted for text-to-image synthesis. While scaling these models up to billions of parameters shows promise, the effectiveness of scaling beyond current sizes remains underexplored and challenging. By explicitly exploiting the computational heterogeneity of image generations, we develop a new family of Mixture-of-Experts (MoE) models (EC-DIT) for diffusion transformers with expert-choice routing. EC-DIT learns to adaptively optimize the compute allocated to understand the input texts and generate the respective image patches, enabling heterogeneous computation aligned with varying text-image complexities. This heterogeneity provides an efficient way of scaling EC-DIT up to 97 billion parameters and achieving significant improvements in training convergence, text-to-image alignment, and overall generation quality over dense models and conventional MoE models. Through extensive ablations, we show that EC-DIT demonstrates superior scalability and adaptive compute allocation by recognizing varying textual importance through end-to-end training. Notably, in text-to-image alignment evaluation, our largest models achieve a state-of-the-art GenEval score of 71.68% and still maintain competitive inference speed with intuitive interpretability. Tao Lei 0001, Bowen Zhang 0002, Yanghao Li, Haoshuo Huang, Ruoming Pang, Bo Dai 0001, Nan Du 0002 |
ICLR | 3 |
| 2025 | MMEgo: Towards Building Egocentric Multimodal LLMs for Video QAabstractThis research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding.
To achieve this goal, we work on three fronts.
First, as there is a lack of QA data for egocentric video understanding, we automatically generate 7M high-quality QA samples for egocentric videos ranging from 30 seconds to one hour long in Ego4D based on human-annotated data.
This is one of the largest egocentric QA datasets.
Second, we contribute a challenging egocentric QA benchmark with 629 videos and 7,026 questions to evaluate the models' ability in recognizing and memorizing visual details across videos of varying lengths. We introduce a new de-biasing evaluation method to help mitigate the unavoidable language bias present in the models being evaluated.
Third, we propose a specialized multimodal architecture featuring a novel ``Memory Pointer Prompting" mechanism. This design includes a global glimpse step to gain an overarching understanding of the entire video and identify key visual information, followed by a fallback step that utilizes the key visual information to generate responses. This enables the model to more effectively comprehend extended video content.
With the data, benchmark, and model, we build MM-Ego, an egocentric multimodal LLM that shows powerful performance on egocentric video understanding. Hanrong Ye, Haotian Zhang 0005, Erik A. Daxberger, Lin Chen 0010, Zongyu Lin, Yanghao Li, Bowen Zhang 0002, Haoxuan You, Dan Xu 0002, Zhe Gan, Jiasen Lu, Yinfei Yang |
ICLR | 7 |
| 2025 | MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuningabstractWe present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture, MM1.5 adopts a data-centric approach to model training, systematically exploring the impact of diverse data mixtures across the entire model training lifecycle. This includes high-quality OCR data and synthetic captions for continual pre-training, as well as an optimized visual instruction-tuning data mixture for supervised fine-tuning. Our models range from 1B to 30B parameters, encompassing both dense and mixture-of-experts (MoE) variants, and demonstrate that careful data curation and training strategies can yield strong performance even at small scales (1B and 3B). Additionally, we introduce two specialized variants: MM1.5-Video, designed for video understanding, and MM1.5-UI, tailored for mobile UI understanding. Through extensive empirical studies and ablations, we provide detailed insights into the training processes and decisions that inform our final designs, offering valuable guidance for future research in MLLM development. Haotian Zhang 0005, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang 0002, Yanghao Li, Sam Dodge, Keen You, Aleksei Timofeev, Hong-You Chen, Jean-Philippe Fauconnier, Zhengfeng Lai, Haoxuan You |
ICLR | 9 |
| 2025 | Contrastive Localized Language-Image Pre-TrainingabstractCLIP has been a celebrated method for training vision encoders to generate image/text representations facilitating various applications. Recently, it has been widely adopted as the vision backbone of multimodal large language models (MLLMs). The success of CLIP relies on aligning web-crawled noisy text annotations at image levels. However, such criteria may be insufficient for downstream tasks in need of fine-grained vision representations, especially when understanding region-level is demanding for MLLMs. We improve the localization capability of CLIP with several advances. Our proposed pre-training method, Contrastive Localized Language-Image Pre-training (CLOC), complements CLIP with region-text contrastive loss and modules. We formulate a new concept, promptable embeddings, of which the encoder produces image embeddings easy to transform into region representations given spatial hints. To support large-scale pre-training, we design a visually-enriched and spatially-localized captioning framework to effectively generate region-text labels. By scaling up to billions of annotated images, CLOC enables high-quality regional embeddings for recognition and retrieval tasks, and can be a drop-in replacement of CLIP to enhance MLLMs, especially on referring and grounding tasks. Hong-You Chen, Zhengfeng Lai, Haotian Zhang 0005, Xinze Wang, Marcin Eichner, Keen You, Bowen Zhang 0002, Yinfei Yang, Zhe Gan |
ICML | 8 |
| 2024 | VeCLIP: Improving CLIP Training via Visual-Enriched Captions
Zhengfeng Lai, Haotian Zhang 0005, Bowen Zhang 0002, Haoping Bai, Aleksei Timofeev, Xianzhi Du, Zhe Gan, Jiulong Shan, Chen-Nee Chuah, Yinfei Yang |
ECCV (42) | 3 |
| 2024 | MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang 0002, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Anton Belyi, Haotian Zhang 0005, Karanjeet Singh 0003, Doug Kang, Hongyu Hè, Max Schwarzer, Tom Gunter, Xiang Kong, Aonan Zhang, Nan Du 0002, Tao Lei 0001, Sam Wiseman, Mark Lee 0003, Ruoming Pang, Peter Grasch, Alexander Toshev, Yinfei Yang |
ECCV (29) | 5 |
| 2024 | MOFI: Learning Image Representations from Noisy Entity Annotated ImagesabstractWe present MOFI, Manifold OF Images, a new vision foundation model designed to learn image representations from noisy entity annotated images. MOFI differs from previous work in two key aspects: 1. pre-training data, and 2. training recipe. Regarding data, we introduce a new approach to automatically assign entity labels to images from noisy image-text pairs. Our approach involves employing a named entity recognition model to extract entities from the alt-text, and then using a CLIP model to select the correct entities as labels of the paired image. It's a simple, cost-effective method that can scale to handle billions of web-mined image-text pairs. Through this method, we have created Image-to-Entities (I2E), a new dataset with 1 billion images and 2 million distinct entities, covering rich visual concepts in the wild. Building upon the I2E dataset, we study different training recipes like supervised pre-training, contrastive pre-training, and multi-task learning. For constrastive pre-training, we treat entity names as free-form text, and further enrich them with entity descriptions. Experiments show that supervised pre-training with large-scale fine-grained entity labels is highly effective for image retrieval tasks, and multi-task training further improves the performance. The final MOFI model achieves 86.66\% mAP on the challenging GPR1200 dataset, surpassing the previous state-of-the-art performance of 72.19% from OpenAI's CLIP model. Further experiments on zero-shot and linear probe image classification also show that MOFI outperforms a CLIP model trained on the original image-text data, demonstrating the effectiveness of the I2E dataset in learning strong image representations. We release our code and model weights at https://github.com/apple/ml-mofi. Aleksei Timofeev, Chen Chen 0005, Bowen Zhang 0002, Kun Duan, Shuangning Liu, Yantao Zheng, Jonathon Shlens, Xianzhi Du, Yinfei Yang |
ICLR | 4 |
| 2024 | Ferret: Refer and Ground Anything Anywhere at Any GranularityabstractWe introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hybrid region representation that integrates discrete coordinates and continuous features jointly to represent a region in the image. To extract the continuous features of versatile regions, we propose a spatial-aware visual sampler, adept at handling varying sparsity across different shapes. Consequently, Ferret can accept diverse region inputs, such as points, bounding boxes, and free-form shapes. To bolster the desired capability of Ferret, we curate GRIT, a comprehensive refer-and-ground instruction tuning dataset including 1.1M samples that contain rich hierarchical spatial knowledge, with an additional 130K hard negative data to promote model robustness. The resulting model not only achieves superior performance in classical referring and grounding tasks, but also greatly outperforms existing MLLMs in region-based and localization-demanded multimodal chatting. Our evaluations also reveal a significantly improved capability of describing image details and a remarkable alleviation in object hallucination. Haoxuan You, Haotian Zhang 0005, Zhe Gan, Xianzhi Du, Bowen Zhang 0002, Liangliang Cao, Shih-Fu Chang, Yinfei Yang |
ICLR | 5 |
| 2023 | STAIR: Learning Sparse Text and Image Representation in Grounded TokensabstractChen Chen, Bowen Zhang, Liangliang Cao, Jiguang Shen, Tom Gunter, Albin Jose, Alexander Toshev, Yantao Zheng, Jonathon Shlens, Ruoming Pang, Yinfei Yang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Chen Chen 0005, Bowen Zhang 0002, Liangliang Cao, Jiguang Shen, Tom Gunter, Albin Madappally Jose, Alexander Toshev, Yantao Zheng, Jonathon Shlens, Ruoming Pang, Yinfei Yang |
EMNLP | 2 |
| 2021 | Systematic Generalization on gSCAN: What is Nearly Solved and What is Next?abstractWe analyze the grounded SCAN (gSCAN) benchmark, which was recently proposed to study systematic generalization for grounded language understanding.First, we study which aspects of the original benchmark can be solved by commonly used methods in multimodal research.We find that a generalpurpose Transformer-based model with crossmodal attention achieves strong performance on a majority of the gSCAN splits, surprisingly outperforming more specialized approaches from prior work.Furthermore, our analysis suggests that many of the remaining errors reveal the same fundamental challenge in systematic generalization of linguistic constructs regardless of visual context.Second, inspired by this finding, we propose challenging new tasks for gSCAN by generating data to incorporate relations between objects in the visual environment.Finally, we find that current models are surprisingly data inefficient given the narrow scope of commands in gSCAN, suggesting another challenge for future work. Linlu Qiu, Hexiang Hu, Bowen Zhang 0002, Peter Shaw 0004, Fei Sha |
EMNLP (1) | 3 |
| 2021 | CREATe: Clinical Report Extraction and Annotation TechnologyabstractClinical case reports are written descriptions of the unique aspects of a particular clinical case, playing an essential role in sharing clinical experiences about atypical disease phenotypes and new therapies. However, to our knowledge, there has been no attempt to develop an end-to-end system to annotate, index, or otherwise curate these reports. In this paper, we propose a novel computational resource platform, CREATe, for extracting, indexing, and querying the contents of clinical case reports. CREATe fosters an environment of sustainable resource support and discovery, enabling researchers to overcome the challenges of information science. An online video of the demonstration can be viewed at https://youtu.be/Q8owBQYTjDc. Yichao Zhou 0001, Bowen Zhang 0002, J. Harry Caufield, Kai-Wei Chang 0001, Yizhou Sun, Peipei Ping, Wei Wang 0010 |
ICDE | 3 |
| 2020 | Learning to Represent Image and Text with Denotation GraphabstractLearning to fuse vision and language information and representing them is an important research problem with many applications.Recent progresses have leveraged the ideas of pretraining (from language modeling) and attention layers in Transformers to learn representation from datasets containing images aligned with linguistic expressions that describe the images.In this paper, we propose learning representations from a set of implied, visually grounded expressions between image and text, automatically mined from those datasets.In particular, we use denotation graphs to represent how specific concepts (such as sentences describing images) can be linked to abstract and generic concepts (such as short phrases) that are also visually grounded.This type of generic-to-specific relations can be discovered using linguistic analysis tools.We propose methods to incorporate such relations into learning representation.We show that state-of-the-art multimodal learning models can be further improved by leveraging automatically harvested structural relations.The representations lead to stronger empirical results on downstream tasks of cross-modal image retrieval, referring expression, and compositional attribute-object recognition.Both our codes and the extracted denotation graphs on the Flickr30K and the COCO datasets are publically available on Bowen Zhang 0002, Hexiang Hu, Vihan Jain, Eugene Ie, Fei Sha |
EMNLP (1) | 1 |
| 2018 | Cross-Modal and Hierarchical Modeling of Video and Text
Bowen Zhang 0002, Hexiang Hu, Fei Sha |
ECCV (13) | 1 |
| 2018 | A Probabilistic Model for Joint Learning of Word Embeddings from Texts and ImagesabstractSeveral recent studies have shown the benefits of combining language and perception to infer word embeddings.These multimodal approaches either simply combine pre-trained textual and visual representations (e.g.features extracted from convolutional neural networks), or use the latter to bias the learning of textual word embeddings.In this work, we propose a novel probabilistic model to formalize how linguistic and perceptual inputs can work in concert to explain the observed word-context pairs in a text corpus.Our approach learns textual and visual representations jointly: latent visual factors couple together a skip-gram model for co-occurrence in linguistic data and a generative latent variable model for visual data.Extensive experimental studies validate the proposed model.Concretely, on the tasks of assessing pairwise word similarity and image/caption retrieval, our approach attains equally competitive or stronger results when compared to other state-of-the-art multimodal models. Melissa Ailem, Bowen Zhang 0002, Aurélien Bellet, Pascal Denis, Fei Sha |
EMNLP | 2 |
| 2018 | Real-Time Action Recognition With Deeply Transferred Motion Vector CNNsabstractThe two-stream CNNs prove very successful for video based action recognition. However the classical two-stream CNNs are time costly, mainly due to the bottleneck of calculating optical flows. In this paper, we propose a two-stream based real-time action recognition approach by using motion vector to replace optical flow. Motion vectors are encoded in video stream and can be extracted directly without extra calculation. However directly training CNN with motion vectors degrades accuracy severely due to the noise and the lack of fine details in motion vectors. In order to relieve this problem, we propose four training strategies which leverage the knowledge learned from optical flow CNN to enhance the accuracy of motion vector CNN. Our insight is that motion vector and optical flow share inherent similar structures which allows us to transfer knowledge from one domain to another. To fully utilize the knowledge learned in optical flow domain, we develop deeply transferred motion vector CNN. Experimental results on various datasets show the effectiveness of our training strategies. Our approach is significantly faster than optical flow based approaches and achieves processing speed of 390.7 frames per second, surpassing real-time requirement. We release our model and code to facilitate further research. Bowen Zhang 0002, Limin Wang 0002, Zhe Wang 0013, Yu Qiao 0001, Hanli Wang |
IEEE Trans. Image Process. | 1 |
| 2017 | Learning correlations for human action recognition in videos
Yun Yi, Hanli Wang, Bowen Zhang 0002 |
Multim. Tools Appl. | 3 |
| 2017 | Weakly Supervised PatchNets: Describing and Aggregating Local Patches for Scene RecognitionabstractTraditional feature encoding scheme (e.g., Fisher vector) with local descriptors (e.g., SIFT) and recent convolutional neural networks (CNNs) are two classes of successful methods for image recognition. In this paper, we propose a hybrid representation, which leverages the discriminative capacity of CNNs and the simplicity of descriptor encoding schema for image recognition, with a focus on scene recognition. To this end, we make three main contributions from the following aspects. First, we propose a patch-level and end-to-end architecture to model the appearance of local patches, called PatchNet. PatchNet is essentially a customized network trained in a weakly supervised manner, which uses the image-level supervision to guide the patch-level feature extraction. Second, we present a hybrid visual representation, called VSAD, by utilizing the robust feature representations of PatchNet to describe local patches and exploiting the semantic probabilities of PatchNet to aggregate these local patches into a global representation. Third, based on the proposed VSAD representation, we propose a new state-of-the-art scene recognition approach, which achieves an excellent performance on two standard benchmarks: MIT Indoor67 (86.2%) and SUN397 (73.0%). Zhe Wang 0013, Limin Wang 0002, Yali Wang 0001, Bowen Zhang 0002, Yu Qiao 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Real-Time Action Recognition with Enhanced Motion Vector CNNsabstractThe deep two-stream architecture [23] exhibited excellent performance on video based action recognition. The most computationally expensive step in this approach comes from the calculation of optical flow which prevents it to be real-time. This paper accelerates this architecture by replacing optical flow with motion vector which can be obtained directly from compressed videos without extra calculation. However, motion vector lacks fine structures, and contains noisy and inaccurate motion patterns, leading to the evident degradation of recognition performance. Our key insight for relieving this problem is that optical flow and motion vector are inherent correlated. Transferring the knowledge learned with optical flow CNN to motion vector CNN can significantly boost the performance of the latter. Specifically, we introduce three strategies for this, initialization transfer, supervision transfer and their combination. Experimental results show that our method achieves comparable recognition performance to the state-of-the-art, while our method can process 390.7 frames per second, which is 27 times faster than the original two-stream method. Bowen Zhang 0002, Limin Wang 0002, Zhe Wang 0013, Yu Qiao 0001, Hanli Wang |
CVPR | 1 |
| 2015 | Encoding scale into fisher vector for human action recognitionabstractIn this paper, a new kind of Fisher Vector (FV) model, named Scale FV (ScaleFV), is proposed to ameliorate visual feature encoding for human action recognition. Although several researches have been proposed for feature encoding, the temporal scale information is almost ignored. Similar to the spatial scale information which has shown to be important in extracting and encoding visual features, the temporal scale information also plays an important role in video content analysis based on our investigation. To demonstrate this, a definition of temporal scale in videos is given, and it is presented that both of the spatial and temporal scale information can be encoded into the FV model by slightly modifying the underlying Gaussian Mixture Models (GMM). Furthermore, an enhanced FV model termed as Combined FV (CombFV) is designed to capture both position and scale information for human action recognition. Comparative experiments are carried out to demonstrate the superior performance of the proposed methods. Bowen Zhang 0002, Hanli Wang |
VCIP | 1 |