Zhenlin Xu

dblp:66/5350 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
abstract
Large language models (LLMs) are typically trained on enormous quantities of unlicensed text, a practice that has led to scrutiny due to possible intellectual property infringement and ethical concerns. Training LLMs on openly licensed text presents a first step towards addressing these issues, but prior data collection efforts have yielded datasets too small or low-quality to produce performant LLMs. To address this gap, we collect, curate, and release the Common Pile v0.1, an eight terabyte collection of openly licensed text designed for LLM pretraining. The Common Pile comprises content from 30 sources that span diverse domains including research papers, code, books, encyclopedias, educational materials, audio transcripts, and more. Crucially, we validate our efforts by training two 7 billion parameter LLMs on text from the Common Pile: Comma v0.1-1T and Comma v0.1-2T, trained on 1 and 2 trillion tokens respectively. Both models attain competitive performance to LLMs trained on unlicensed text with similar computational budgets, such as Llama 1 and 2 7B. In addition to releasing the Common Pile v0.1 itself, we also release the code used in its creation as well as the training mixture and checkpoints for the Comma v0.1 models.
Nikhil Kandpal, Brian Lester, Colin Raffel, Sebastian Majstorovic, Stella Biderman, Baber Abbasi, Luca Soldaini, Enrico Shippole, A. Feder Cooper, Aviya Skowron, Shayne Longpre, Lintang Sutawika, Alon Albalak, Zhenlin Xu, Guilherme Penedo, Loubna Ben Allal, Elie Bakouch, John David Pressman, Honglu Fan, Dashiell Stander, Guangyu Song, Aaron Gokaslan, John Kirchenbauer, Tom Goldstein, Brian R. Bartoldson, Bhavya Kailkhura, Tyler Murray
NeurIPS14
2025 GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-Grained Video-Language Learning
abstract
In various video-language learning tasks, the challenge of achieving cross-modality alignment with multi-grained data persists. We propose a method to tackle this challenge from two crucial perspectives: data and modeling. Given the absence of a multi-grained video-text pretraining dataset, we introduce a Granularity EXpansion (GEX) method with Integration and Compression operations to expand the granularity of a single-grained dataset. To better model multi-grained data, we introduce an Iterative Approximation Module (IAM), which embeds multi-grained videos and texts into a unified, low-dimensional semantic space while preserving essential information for cross-modal alignment. Furthermore, GEXIA is highly scalable with no restrictions on the number of video-text granularities for alignment. We evaluate our work on three categories of video tasks across seven benchmark datasets, showcasing state-of-the-art or comparable performance. Remarkably, our model excels in tasks involving long-form video understanding, even though the pretraining dataset only contains short video clips.
Jue Wang 0010, David Fan 0001, Zhenlin Xu, Linda Liu, Vimal Bhat, Xinyu Li 0003
WACV5
2025 DT Assisted Task Offloading for C-V2X Networks With Imperfect DT Prediction Conditions
abstract
The development of intelligent transportation has generated many ultra reliable low latency communication (URLLC) tasks, which require sufficient communication and computation resources for task offloading and processing. Although mobile edge computing (MEC) provides a promising solution, its efficiency is subject to the limited knowledge and analysis capability on the physical networks. Therefore, in this paper, we propose a digital twin (DT) empowered MEC framework to strengthen the MEC task offloading efficiency in cellular vehicle-to-everything (C-V2X) networks. Our proposed DT is constructed through a hybrid data-driven and model-driven approach to capture the realistic transportation network features. Then, DT leverages the metric of time to collision to predict vehicular safety levels and estimates the corresponding URLLC task requirements of future time slots. The prediction results are further utilized to make decisions on the URLLC resource reservation. Different from conventional studies, we consider the influence of DT’s inaccurate predictions (i.e., the prediction with error) on the resource allocations. Specifically, the inaccurate DT prediction results are considered as uncertain constraints of the resource reservation problem. A robust parameter from the robust optimization is adopted to adjust the tradeoff between the problem uncertainty and solution optimality degree. Further, we leverage the optimized resource reservation results to construct the task offloading problem. The problem is decoupled into two sub-problems of channel resource allocation and computation resource allocation, respectively. And a two-stage matching algorithm is developed to solve each sub-problem based on the resource reservation constraints. Finally, realistic road information is mapped into DT for simulations. Simulation results validate the advantages of our proposed approach by comparing with existing schemes.
Bo Fan 0003, Zhenlin Xu, Zhidu Li, Yuan Wu 0001, Yan Zhang 0002
IEEE Trans. Intell. Transp. Syst.2
2025 Does Another Pedestrian Matter? A Virtual Reality Study on the Interaction Between Multiple Pedestrians and Autonomous Vehicles in Shared Space
abstract
This study utilized Virtual Reality (VR) experiments to investigate pedestrian-autonomous vehicle interaction in shared spaces. In the VR experiment, pedestrians attempt to cross the road under different conditions, including the presence of another pedestrian, different external Human-Machine-Interfaces, AV driving styles, and road conditions. We employed an innovative VR setup that enabled two pedestrians to interact in real time with physical movements within an immersive VR environment. Overall, we found that the presence of multiple pedestrians significantly influenced pedestrian movement dynamics during road crossing. Additionally, the relative standing position had a significant impact on the distant pedestrians regarding time before crossing and vehicle-gazing behavior. While previous studies predominantly focused on pedestrian-AV interaction with a single pedestrian, this study takes an important step forward in terms of theory, methods, and relevance by considering interactions between multiple pedestrians and AVs. The findings establish a basis for further exploration of pedestrian-AV interaction in shared space.
Zhenlin Xu, Haneen Farah, Bart van Arem
IEEE Trans. Intell. Transp. Syst.2
2024 Self-Supervised Multi-Object Tracking with Path Consistency
abstract
In this paper, we propose a novel concept of path consis-tency to learn robust object matching without using manual object identity supervision. Our key idea is that, to track a object through frames, we can obtain multiple different as-sociation results from a model by varying the frames it can observe, i.e., skipping frames in observation. As the differ-ences in observations do not alter the identities of objects, the obtained association results should be consistent. Based on this rationale, we generate multiple observation paths, each specifying a different set of frames to be skipped, and formulate the Path Consistency Loss that enforces the as-sociation results are consistent across different observation paths. We use the proposed loss to train our object matching model with only self-supervision. By extensive experiments on three tracking datasets (MOT17, PersonPath22, KITTI), we demonstrate that our method outperforms existing unsu-pervised methods with consistent margins on various eval-uation metrics, and even achieves performance close to su-pervised methods.
Zijia Lu, Bing Shuai, Yanbei Chen, Zhenlin Xu, Davide Modolo
CVPR4
2023 ScaleDet: A Scalable Multi-Dataset Object Detector
abstract
Multi-dataset training provides a viable solution for exploiting heterogeneous large-scale datasets without extra annotation cost. In this work, we propose a scalable multi-dataset detector (ScaleDet) that can scale up its generalization across datasets when increasing the number of training datasets. Unlike existing multi-dataset learners that mostly rely on manual relabelling efforts or sophisticated optimizations to unify labels across datasets, we introduce a simple yet scalable formulation to derive a unified semantic label space for multi-dataset training. ScaleDet is trained by visual-textual alignment to learn the label assignment with label semantic similarities across datasets. Once trained, ScaleDet can generalize well on any given upstream and downstream datasets with seen and unseen classes. We conduct extensive experiments using LVIS, COCO, Objects365, OpenImages as upstream datasets, and 13 datasets from Object Detection in the Wild (ODinW) as downstream datasets. Our results show that ScaleDet achieves compelling strong model performance with an mAP of 50.7 on LVIS, 58.8 on COCO, 46.8 on Objects365, 76.2 on OpenImages, and 71.8 on ODinW, surpassing state-of-the-art detectors with the same backbone.
Yanbei Chen, Manchen Wang, Abhay Mittal, Zhenlin Xu, Paolo Favaro, Joseph Tighe, Davide Modolo
CVPR4
2023 SimpleClick: Interactive Image Segmentation with Simple Vision Transformers
abstract
Click-based interactive image segmentation aims at extracting objects with a limited user clicking. A hierarchical backbone is the de-facto architecture for current methods. Recently, the plain, non-hierarchical Vision Transformer (ViT) has emerged as a competitive backbone for dense prediction tasks. This design allows the original ViT to be a foundation model that can be finetuned for downstream tasks without redesigning a hierarchical backbone for pretraining. Although this design is simple and has been proven effective, it has not yet been explored for interactive image segmentation. To fill this gap, we propose SimpleClick, the first interactive segmentation method that leverages a plain backbone. Based on the plain backbone, we introduce a symmetric patch embedding layer that encodes clicks into the backbone with minor modifications to the backbone itself. With the plain backbone pretrained as a masked autoencoder (MAE), SimpleClick achieves state-of-the-art performance. Remarkably, our method achieves 4.15 NoC@90 on SBD, improving 21.8% over the previous best result. Extensive evaluation on medical images demonstrates the generalizability of our method. We provide a detailed computational analysis, highlighting the suitability of our method as a practical annotation tool.
Qin Liu 0008, Zhenlin Xu, Gedas Bertasius, Marc Niethammer
ICCV2
2022 iSegFormer: Interactive Segmentation via Transformers with Application to 3D Knee MR Images
Qin Liu 0008, Zhenlin Xu, Yining Jiao, Marc Niethammer
MICCAI (5)2
2022 Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language
abstract
Deep learning models struggle with compositional generalization, i.e. the ability to recognize or generate novel combinations of observed elementary concepts. In hopes of enabling compositional generalization, various unsupervised learning algorithms have been proposed with inductive biases that aim to induce compositional structure in learned representations (e.g. disentangled representation and emergent language learning). In this work, we evaluate these unsupervised learning algorithms in terms of how well they enable \textit{compositional generalization}. Specifically, our evaluation protocol focuses on whether or not it is easy to train a simple model on top of the learned representation that generalizes to new combinations of compositional factors. We systematically study three unsupervised representation learning algorithms - $\beta$-VAE, $\beta$-TCVAE, and emergent language (EL) autoencoders - on two datasets that allow directly testing compositional generalization. We find that directly using the bottleneck representation with simple models and few labels may lead to worse generalization than using representations from layers before or after the learned representation itself. In addition, we find that the previously proposed metrics for evaluating the levels of compositionality are not correlated with actual compositional generalization in our framework. Surprisingly, we find that increasing pressure to produce a disentangled representation (e.g. increasing $\beta$ in the $\beta$-VAE) produces representations with worse generalization, while representations from EL models show strong compositional generalization. Motivated by this observation, we further investigate the advantages of using EL to induce compositional structure in unsupervised representation learning, finding that it shows consistently stronger generalization than disentanglement models, especially when using less unlabeled data for unsupervised learning and fewer labels for downstream tasks. Taken together, our results shed new light onto the compositional generalization behavior of different unsupervised learning algorithms with a new setting to rigorously test this behavior, and suggest the potential benefits of developing EL learning algorithms for more generalizable representations. Our code is publicly available at https://github.com/wildphoton/Compositional-Generalization .
Zhenlin Xu, Marc Niethammer, Colin Raffel
NeurIPS1
2022 DADP: Dynamic abnormality detection and progression for longitudinal knee magnetic resonance images from the Osteoarthritis Initiative
Chao Huang 0005, Zhenlin Xu, Zhengyang Shen, Tianyou Luo, Tengfei Li 0001, Daniel Nissman, Amanda Nelson, Yvonne Golightly, Marc Niethammer, Hongtu Zhu
Medical Image Anal.2
2021 Robust and Generalizable Visual Representation Learning via Random Convolutions
Zhenlin Xu, Deyi Liu, Junlin Yang, Colin Raffel, Marc Niethammer
ICLR1
2020 Adversarial Data Augmentation via Deformation Statistics
Sahin Olut, Zhengyang Shen, Zhenlin Xu, Samuel Gerber, Marc Niethammer
ECCV (29)3
2020 Anatomical Data Augmentation via Fluid-Based Image Registration
Zhengyang Shen, Zhenlin Xu, Sahin Olut, Marc Niethammer
MICCAI (3)2
2019 Networks for Joint Affine and Non-Parametric Image Registration
abstract
We introduce an end-to-end deep-learning framework for 3D medical image registration. In contrast to existing approaches, our framework combines two registration methods: an affine registration and a vector momentum-parameterized stationary velocity field (vSVF) model. Specifically, it consists of three stages. In the first stage, a multi-step affine network predicts affine transform parameters. In the second stage, we use a U-Net-like network to generate a momentum, from which a velocity field can be computed via smoothing. Finally, in the third stage, we employ a self-iterable map-based vSVF component to provide a non-parametric refinement based on the current estimate of the transformation map. Once the model is trained, a registration is completed in one forward pass. To evaluate the performance, we conducted longitudinal and cross-subject experiments on 3D magnetic resonance images (MRI) of the knee of the Osteoarthritis Initiative (OAI) dataset. Results show that our framework achieves comparable performance to state-of-the-art medical image registration approaches, but it is much faster, with a better control of transformation regularity including the ability to produce approximately symmetric transformations, and combining affine as well as non-parametric registration.
Zhengyang Shen, Xu Han 0009, Zhenlin Xu, Marc Niethammer
CVPR3
2019 DeepAtlas: Joint Semi-supervised Learning of Image Registration and Segmentation
Zhenlin Xu, Marc Niethammer
MICCAI (2)1
2018 Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras
abstract
We propose a new approach for 3D reconstruction of dynamic indoor and outdoor scenes in everyday environments, leveraging only cameras worn by a user. This approach allows 3D reconstruction of experiences at any location and virtual tours from anywhere. The key innovation of the proposed ego-centric reconstruction system is to capture the wearer's body pose and facial expression from near-body views, e.g. cameras on the user's glasses, and to capture the surrounding environment using outward-facing views. The main challenge of the ego-centric reconstruction, however, is the poor coverage of the near-body views - that is, the user's body and face are observed from vantage points that are convenient for wear but inconvenient for capture. To overcome these challenges, we propose a parametric-model-based approach to user motion estimation. This approach utilizes convolutional neural networks (CNNs) for near-view body pose estimation, and we introduce a CNN-based approach for facial expression estimation that combines audio and video. For each time-point during capture, the intermediate model-based reconstructions from these systems are used to re-target a high-fidelity pre-scanned model of the user. We demonstrate that the proposed self-sufficient, head-worn capture system is capable of reconstructing the wearer's movements and their surrounding environment in both indoor and outdoor situations without any additional views. As a proof of concept, we show how the resulting 3D-plus-time reconstruction can be immersively experienced within a virtual reality system (e.g., the HTC Vive). We expect that the size of the proposed egocentric capture-and-reconstruction system will eventually be reduced to fit within future AR glasses, and will be widely useful for immersive 3D telepresence, virtual tours, and general use-anywhere 3D content creation.
Young-Woon Cha, True Price, Xinran Lu, Nicholas Rewkowski, Rohan Chabra, Zihe Qin, Hyounghun Kim, Zhaoqi Su, Yebin Liu, Adrian Ilie, Andrei State, Zhenlin Xu, Jan-Michael Frahm, Henry Fuchs
IEEE Trans. Vis. Comput. Graph.13