Zhengzhong Tu

dblp:218/1473 · DBLP profile ↗
← Back
42ranked-venue papers
12as first author
39since 2021 · last 2026
0000-0002-7594-2292ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 12 first-author · 28 since 2021Artificial intelligence and machine learning · 26 · 3 first-author · 26 since 2021Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2026 3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation
abstract
We introduce DreamLand, an interactive 4D visualization framework that integrates WebGL with Supersplat rendering. It transforms static images and text into coherent 4D scenes through four core modules and employs a foveated rendering strategy for efficient, real-time multi-modal interaction. This framework enables adaptive, user-driven exploration of complex 4D environments.
Yunhong He, Zhengqing Yuan, Zhengzhong Tu, Yanfang Ye 0001, Lichao Sun 0001
AAAI3
2026 CyPortQA: Benchmarking Multimodal Large Language Models for Cyclone Preparedness in Port Operation
abstract
As tropical cyclones intensify and track forecasts become increasingly uncertain, U.S. ports face heightened supply-chain risk under extreme weather conditions. Port operators need to rapidly synthesize diverse multimodal forecast products, such as probabilistic wind maps, track cones, and official advisories, into clear, actionable guidance as cyclones approach. Multimodal large language models (MLLMs) offer a powerful means to integrate these heterogeneous data sources alongside broader contextual knowledge, yet their accuracy and reliability in the specific context of port cyclone preparedness have not been rigorously evaluated. To fill this gap, we introduce CyPortQA, the first multimodal benchmark tailored to port operations under cyclone threat. CyPortQA assembles 2,917 real-world disruption scenarios from 2015 through 2023, spanning 145 U.S. principal ports and 90 named storms. Each scenario fuses multi-source data (i.e., tropical cyclone products, port operational impact records, and port condition bulletins) and is expanded through an automated pipeline into 117,178 structured question–answer pairs. Using this benchmark, we conduct extensive experiments on diverse MLLMs, including both open-source and proprietary model. MLLMs demonstrate great potential in situation understanding but still face considerable challenges in reasoning tasks, including potential impact estimation and decision reasoning.
Chenchen Kuai, Yang Zhou 0019, Xiubin Bruce Wang, Tianbao Yang, Zhengzhong Tu
AAAI6
2026 Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models
abstract
Recent studies posit that Reinforcement Learning with Verifiable Rewards (RLVR) primarily amplifies behaviors inherent to the pre-training distribution rather than inducing new capabilities, but these insights are predominantly limited to language-only domains, leaving the dynamics of visual-centric spatial reasoning under-explored. To examine the impact of RLVR on the capability boundaries of Vision-Language Models (VLMs), we introduce \textbf{Ariadne}, a controlled framework based on synthetic maze navigation where the reasoning difficulty is precisely regulated by path length and the number of turns. We demonstrate that applying RLVR extends the spatial reasoning boundary, achieving success on problems where the base policy VLM consistently attains $0\%$ accuracy despite increasing pass@k sampling budgets, indicating that the optimized policy successfully navigates search spaces that were effectively unreachable by the base distribution. Furthermore, despite being trained exclusively on synthetic mazes, we evaluate the model on two real-world navigation benchmarks (MapBench and ReasonMap) in a zero-shot setting. The observed improvements in these out-of-domain tasks suggest genuine spatial reasoning capability expansion rather than mere sampling efficiency.
Minghe Shen, Zhuo Zhi, Chonghan Liu, Shuo Xing, Zhengzhong Tu
ACL (1)5
2026 Edge-based multimodal sensor data fusion with Vision-Language-Action (VLA) model for real-time autonomous vehicle accident avoidance
Fengze Yang, Yang Zhou 0019, Xuewen Luo, Zhengzhong Tu
Eng. Appl. Artif. Intell.5
2025 DPU: Dynamic Prototype Updating for Multimodal Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection is crucial for ensuring the robustness of machine learning models by identifying samples that deviate from the training distribution. While traditional OOD detection has predominantly focused on single-modality inputs, such as images, recent advancements in multimodal models have shown the potential of utilizing multiple modalities (e.g., video, optical flow, audio) to improve detection performance. However, existing approaches often neglect intra-class variability within in-distribution (ID) data, assuming that samples of the same class are perfectly cohesive and consistent. This assumption can lead to performance degradation, especially when prediction discrepancies are indiscriminately amplified across all samples. To address this issue, we propose Dynamic Prototype Updating (DPU), a novel plug-and-play framework for multimodal OOD detection that accounts for intra-class variations. Our method dynamically updates class center representations for each class by measuring the variance of similar samples within each batch, enabling tailored adjustments. This approach allows us to intensify prediction discrepancies based on the updated class centers, thereby enhancing the model’s robustness and generalization across different modalities. Extensive experiments on two tasks, five datasets, and nine base OOD algorithms demonstrate that DPU significantly improves OOD detection performances, setting a new state-of-the-art in multimodal OOD detection, including improvements up to 80% in Far-OOD detection. To improve accessibility and reproducibility, our code is released at https://github.com/lili0415/DPU-OOD-Detection.
Shawn Li, Huixian Gong, Hao Dong 0011, Tiankai Yang 0001, Zhengzhong Tu, Yue Zhao 0016
CVPR5
2025 SleeperMark: Towards Robust Watermark against Fine-Tuning Text-to-image Diffusion Models
abstract
Recent advances in large-scale text-to-image (T2I) diffusion models have enabled a variety of downstream applications. As T2I models require extensive resources for training, they constitute highly valued intellectual property (IP) for their legitimate owners, yet making them incentive targets for unauthorized fine-tuning by adversaries seeking to leverage these models for customized, usually profitable applications. Existing IP protection methods for diffusion models generally involve embedding watermark patterns and then verifying ownership through generated outputs examination, or inspecting the model’s feature space. However, these techniques are inherently ineffective in practical scenarios when the watermarked model undergoes fine-tuning, and the feature space is inaccessible during verification (i.e., black-box setting). The model is prone to forgetting the previously learned watermark knowledge when it adapts to a new task. To address this challenge, we propose SleeperMark, a novel framework designed to embed resilient watermarks into T2I diffusion models. SleeperMark explicitly guides the model to disentangle the watermark information from the semantic concepts it learns, allowing the model to retain the embedded watermark while continuing to be adapted to new downstream tasks. Our extensive experiments demonstrate the effectiveness of SleeperMark across various types of diffusion models, including latent diffusion models (e.g., Stable Diffusion) and pixel diffusion models (e.g., DeepFloyd-IF), showing robustness against downstream fine-tuning and various attacks at both the image and model levels, with minimal impact on the model’s generative capability. The code is available at https://github.com/taco-group/SleeperMark.
Zilan Wang, Yiming Li 0004, Heng Huang 0001, Muhao Chen 0001, Zhengzhong Tu
CVPR7
2025 Edit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative Editing
abstract
Recent advancements in diffusion models have made generative image editing more accessible than ever. While these developments allow users to generate creative edits with ease, they also raise significant ethical concerns, particularly regarding malicious edits to human portraits that threaten individuals’ privacy and identity security. Existing general-purpose image protection methods primarily focus on generating adversarial perturbations to nullify edit effects. However, these approaches often exhibit instability to protect against diverse editing requests. In this work, we introduce a novel perspective to personal human portrait protection against malicious editing. Unlike traditional methods aiming to prevent edits from taking effect, our method, FaceLock, optimizes adversarial perturbations to ensure that original biometric information—such as facial features—is either destroyed or substantially altered post-editing, rendering the subject in the edited output biometrically unrecognizable. Our approach innovatively integrates facial recognition and visual perception factors into the perturbation optimization process, ensuring robust protection against a variety of editing attempts. Besides, we shed light on several critical issues with commonly used evaluation metrics in image editing and reveal cheating methods by which they can be easily manipulated, leading to deceptive assessments of protection. Through extensive experiments, we demonstrate that FaceLock significantly outperforms all baselines in defense performance against a wide range of malicious edits. Moreover, our method also exhibits strong robustness against purification techniques. Comprehensive ablation studies confirm the stability and broad applicability of our method across diverse diffusion-based editing algorithms. Our work not only advances the state-of-the-art in biometric defense but also sets the foundation for more secure and privacy-preserving practices in image editing. The code is publicly available at: https://github.com/taco-group/FaceLock.
Hanhui Wang, Ruizheng Bai, Yue Zhao 0016, Sijia Liu 0001, Zhengzhong Tu
CVPR6
2025 Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
abstract
Shuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai, Yueqi Wang, Chan-Wei Hu, Chengxuan Qian, Huaxiu Yao, Zhengzhong Tu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Shuo Xing, Ruizheng Bai, Chan-Wei Hu, Chengxuan Qian, Huaxiu Yao, Zhengzhong Tu
EMNLP9
2025 Secure On-Device Video OOD Detection without Backpropagation
Shawn Li, Peilin Cai, Yuxiao Zhou 0006, Zhiyu Ni, Renjie Liang, You Qin, Yi Nian, Zhengzhong Tu, Xiyang Hu, Yue Zhao 0016
ICCV8
2025 Uniocc: a Unified Benchmark for Occupancy Forecasting and Prediction in Autonomous Driving
Xiangyu Huang, Xiaokang Sun, Mingxuan Yan, Shuo Xing, Zhengzhong Tu, Jiachen Li 0001
ICCV6
2025 STAMP: Scalable Task- And Model-agnostic Collaborative Perception
abstract
Perception is a crucial component of autonomous driving systems. However, single-agent setups often face limitations due to sensor constraints, especially under challenging conditions like severe occlusion, adverse weather, and long-range object detection. Multi-agent collaborative perception (CP) offers a promising solution that enables communication and information sharing between connected vehicles. Yet, the heterogeneity among agents—in terms of sensors, models, and tasks—significantly hinders effective and efficient cross-agent collaboration. To address these challenges, we propose STAMP, a scalable task- and model-agnostic collaborative perception framework tailored for heterogeneous agents. STAMP utilizes lightweight adapter-reverter pairs to transform Bird's Eye View (BEV) features between agent-specific domains and a shared protocol domain, facilitating efficient feature sharing and fusion while minimizing computational overhead. Moreover, our approach enhances scalability, preserves model security, and accommodates a diverse range of agents. Extensive experiments on both simulated (OPV2V) and real-world (V2V4Real) datasets demonstrate that STAMP achieves comparable or superior accuracy to state-of-the-art models with significantly reduced computational costs. As the first-of-its-kind task- and model-agnostic collaborative perception framework, STAMP aims to advance research in scalable and secure mobility systems, bringing us closer to Level 5 autonomy. Our project page is at https://xiangbogaobarry.github.io/STAMP and the code is available at https://github.com/taco-group/STAMP.
Xiangbo Gao, Runsheng Xu, Jiachen Li 0001, Ziran Wang, Zhiwen Fan, Zhengzhong Tu
ICLR6
2025 4K4DGen: Panoramic 4D Generation at 4K Resolution
abstract
The blooming of virtual reality and augmented reality (VR/AR) technologies has driven an increasing demand for the creation of high-quality, immersive, and dynamic environments. However, existing generative techniques either focus solely on dynamic objects or perform outpainting from a single perspective image, failing to meet the requirements of VR/AR applications that need free-viewpoint, 360$^{\circ}$ virtual views where users can move in all directions. In this work, we tackle the challenging task of elevating a single panorama to an immersive 4D experience. For the first time, we demonstrate the capability to generate omnidirectional dynamic scenes with 360$^{\circ}$ views at 4K (4096 $\times$ 2048) resolution, thereby providing an immersive user experience. Our method introduces a pipeline that facilitates natural scene animations and optimizes a set of 3D Gaussians using efficient splatting techniques for real-time exploration. To overcome the lack of scene-scale annotated 4D data and models, especially in panoramic formats, we propose a novel Panoramic Denoiser that adapts generic 2D diffusion priors to animate consistently in 360$^{\circ}$ images, transforming them into panoramic videos with dynamic scenes at targeted regions. Subsequently, we propose Dynamic Panoramic Lifting to elevate the panoramic video into a 4D immersive environment while preserving spatial and temporal consistency. By transferring prior knowledge from 2D models in the perspective domain to the panoramic domain and the 4D lifting with spatial appearance and geometry regularization, we achieve high-quality Panorama-to-4D generation at a resolution of 4K for the first time. Project page: https://4k4dgen.github.io/.
Renjie Li 0003, Panwang Pan, Bangbang Yang, Dejia Xu, Shijie Zhou 0003, Xuanyang Zhang, Achuta Kadambi, Zhangyang Wang, Zhengzhong Tu, Zhiwen Fan
ICLR10
2025 V2X-DGW: Domain Generalization for Multi-Agent Perception Under Adverse Weather Conditions
abstract
Current LiDAR-based Vehicle-to-Everything (V2X) multi-agent perception systems have shown the significant success on 3D object detection. While these models perform well in the trained clean weather, they struggle in unseen adverse weather conditions with the domain gap. In this paper, we propose a Domain Generalization based approach, named V2X-DGW, for LiDAR-based 3D object detection on multi-agent perception system under adverse weather conditions. Our research aims to not only maintain favorable multi-agent performance in the clean weather but also promote the performance in the unseen adverse weather conditions by learning only on the clean weather data. To realize the Domain Generalization, we first introduce the Adaptive Weather Augmentation (AWA) to mimic the unseen adverse weather conditions, and then propose two alignments for generalizable representation learning: Trust-region Weatherinvariant Alignment (TWA) and Agent-aware Contrastive Alignment (ACA). To evaluate this research, we add Fog, Rain, Snow conditions on two publicized multi-agent datasets based on physics-based models, resulting in two new datasets: OPV2V-w and V2XSet-w. Extensive experiments demonstrate that our V2X-DGW achieved significant improvements in the unseen adverse weathers. The code is available at https://github.com/Baolu1998/V2X-DGW.
Xinyu Liu 0009, Runsheng Xu, Zhengzhong Tu, Jiacheng Guo, Qin Zou 0001, Xiaopeng Li 0020, Hongkai Yu
ICRA5
2025 CoMamba: Real-time Cooperative Perception Unlocked with State-Space Models
abstract
Cooperative perception systems play a vital role in enhancing the safety and efficiency of vehicular autonomy. Although recent studies have highlighted the efficacy of vehicle-to-everything (V2X) communication techniques in autonomous driving, a significant challenge persists: how to efficiently integrate multiple high-bandwidth features across an expanding network of connected agents such as vehicles and infrastructure. In this paper, we introduce CoMamba, a novel cooperative 3D detection framework designed to leverage state-space models for real-time onboard vehicle perception. Compared to prior state-of-the-art transformer-based models, CoMamba enjoys being a more scalable 3D model using bidirectional state space models, bypassing the quadratic complexity pain-point of attention mechanisms. Through extensive experimentation on V2X/V2V datasets, CoMamba achieves superior performance compared to existing methods while maintaining real-time processing capabilities. The proposed framework not only enhances object detection accuracy but also significantly reduces processing time, making it a promising solution for next-generation cooperative perception systems in intelligent transportation networks.
Xinyu Liu 0009, Runsheng Xu, Jiachen Li 0001, Hongkai Yu, Zhengzhong Tu
IROS7
2025 CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
abstract
Multi-agent collaborative perception enhances each agent’s perceptual capabilities by sharing sensing information to cooperatively perform robot perception tasks. This approach has proven effective in addressing challenges such as sensor deficiencies, occlusions, and long-range perception. However, existing representative collaborative perception systems transmit intermediate feature maps, such as bird’s-eye view (BEV) representations, which contain a significant amount of non-critical information, leading to high communication bandwidth requirements. To enhance communication efficiency while preserving perception capability, we introduce CoCMT, an object-query-based collaboration framework that optimizes communication bandwidth by selectively extracting and transmitting essential features. Within CoCMT, we introduce the Efficient Query Transformer (EQFormer) to effectively fuse multi-agent object queries and implement a synergistic deep supervision to enhance the positive reinforcement between stages, leading to improved overall performance. Experiments on OPV2V and V2V4Real datasets show CoCMT outperforms state-of-the-art methods while drastically reducing communication needs. On V2V4Real, our model (Top-50 object queries) requires only 0.416 Mb bandwidth—83 times less than SOTA methods—while improving AP@70 by 1.1%. This efficiency breakthrough enables practical collaborative perception deployment in bandwidth-constrained environments without sacrificing detection accuracy. The code and models are open-sourced through the following link: https://github.com/taco-group/COCMT.
Rujia Wang, Xiangbo Gao, Hao Xiang 0001, Runsheng Xu, Zhengzhong Tu
IROS5
2025 DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization
abstract
The recent success and openness of DeepSeek-R1 have brought widespread attention to Group Relative Policy Optimization (GRPO) as a reinforcement learning method for large reasoning models (LRMs). In this work, we analyze the GRPO objective under a binary reward setting and reveal an inherent limitation of question-level difficulty bias arising from its group relative advantage function. We also identify a connection between GRPO and traditional discriminative methods in supervised learning. Motivated by these insights, we introduce a new **Discriminative Constrained Optimization (DisCO)** framework for reinforcing LRMs, grounded in the principle of discriminative learning: increasing the scores of positive answers while decreasing those of negative ones. The main differences between DisCO and GRPO and its recent variants are: (1) it replaces the group relative objective with a discriminative objective defined by a scoring function; (2) it abandons clipping-based surrogates in favor of non-clipping RL surrogate objectives used as scoring functions; (3) it employs a simple yet effective constrained optimization approach to enforce the KL divergence constraint. As a result, DisCO offers notable advantages over GRPO and its variants: (i) it completely eliminates difficulty bias by adopting discriminative objectives; (ii) it addresses the entropy instability in GRPO and its variants through the use of non-clipping scoring functions and a constrained optimization approach, yielding long and stable training dynamics; (iii) it allows the incorporation of advanced discriminative learning techniques to address data imbalance, where a significant number of questions have more negative than positive generated answers during training. Our experiments on enhancing the mathematical reasoning capabilities of SFT-finetuned models show that DisCO significantly outperforms GRPO and its improved variants such as DAPO, achieving average gains of 7\% over GRPO and 6\% over DAPO across six benchmark tasks for a 1.5B model.
Gang Li 0048, Ming Lin 0002, Tomer Galanti, Zhengzhong Tu, Tianbao Yang
NeurIPS4
2025 4KAgent: Agentic Any Image to 4K Super-Resolution
abstract
We present 4KAgent, a unified agentic super-resolution generalist system designed to universally upscale any image to 4K resolution (and even higher, if applied iteratively). Our system can transform images from extremely low resolutions with severe degradations, for example, highly distorted inputs at $256\times 256$, into crystal-clear, photorealistic 4K outputs. 4KAgent comprises three core components: (1) Profiling, a module that customizes the 4KAgent pipeline based on bespoke use cases; (2) A Perception Agent, which leverages vision-language models alongside image quality assessment experts to analyze the input image and make a tailored restoration plan; and (3) A Restoration Agent, which executes the plan, following a recursive execution-reflection paradigm, guided by a quality-driven mixture-of-experts policy to select the optimal output for each step. Additionally, 4KAgent embeds a specialized face restoration pipeline, significantly enhancing facial details in portrait and selfie photos. We rigorously evaluate our 4KAgent across 11 distinct task categories encompassing a total of 26 diverse benchmarks, setting new state-of-the-art on a broad spectrum of imaging domains. Our evaluations cover natural images, portrait photos, AI-generated content, satellite imagery, fluorescence microscopy, and medical imaging like fundoscopy, ultrasound, and X-ray, demonstrating superior performance in terms of both perceptual (e.g., NIQE, MUSIQ) and fidelity (e.g., PSNR) metrics. By establishing a novel agentic paradigm for low-level vision tasks, we aim to catalyze broader interest and innovation within vision-centric autonomous agents across diverse research communities. We release all the code, models, and results at: https://4kagent.github.io.
Yushen Zuo, Qi Zheng 0004, Renjie Li 0003, Jian Wang 0100, Yide Zhang, Gengchen Mai, Lihong V. Wang, James Zou 0001, Ming-Hsuan Yang 0001, Zhengzhong Tu
NeurIPS13
2025 V2X-ViTv2: Improved Vision Transformers for Vehicle-to-Everything Cooperative Perception
abstract
In this paper, we study the application of Vehicle-to-Everything (V2X) communication to improve the perception performance of autonomous vehicles. We present V2X-ViTs, a robust cooperative perception framework with V2X communication using novel vision Transformer models. First, we present V2X-ViTv1 containing holistic attention modules that can effectively fuse information across on-road agents (i.e., vehicles and infrastructure). Specifically, V2X-ViTv1 consists of alternating layers of heterogeneous multi-agent self-attention and multi-scale window self-attention, which captures inter-agent interaction and per-agent spatial relationships. These key modules are designed in a unified Transformer architecture to handle common V2X challenges, including asynchronous information sharing, pose errors, and heterogeneity of V2X components. Second, we propose an advanced architecture, V2X-ViTv2, that enjoys increased ability for multi-scale perception. We also propose advanced data augmentation techniques tailored for V2X applications to improve performance. We construct a large-scale V2X perception dataset using CARLA and OpenCDA to validate our approach. Extensive experimental results on both synthetic and real-world datasets show that V2X-ViTs achieve state-of-the-art performance for 3D object detection and are robust even under harsh, noisy environments.
Runsheng Xu, Chia-Ju Chen, Zhengzhong Tu, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Understanding, detecting, and removing perceptual banding artifacts in compressed videos
abstract
Banding artifacts, or false contouring, are a common compression impairment that often appears on large smooth regions of encoded videos and images. These staircase-like color bands can be very noticeable and annoying, even on otherwise high-quality videos, especially when displayed on high-definition screens. Yet, relatively little attention has been applied to this problem. Here we study this artifact, by first analyzing the perceptual and encoding aspects of banding artifacts, then propose a new distortion-specific no-reference video quality algorithm for predicting banding artifacts, inspired by perceptual models. The proposed banding detector can generate a pixel-wise banding visibility map, and output overall banding severity scores at both the frame and video levels. Furthermore, we propose a deep learning based approach to improve the overall perceptual quality of compressed videos by joint debanding and compression artifact removal. Our experimental results show that the proposed banding detector delivers better consistency with subjective evaluations, and is able to detect different perceptual severity levels of bands. The debanding experiments also show that the proposed algorithm outperforms recent debanding models both visually and quantitatively. The code is available at https://github.com/google/bband-adaband and https://github.com/vztu/DebandingNet .
Zhengzhong Tu, Chia-Ju Chen, Jessie Lin, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik
Signal Process. Image Commun.1
2025 Subjective and Objective Quality Assessment of Banding Artifacts on Compressed Videos
abstract
Although there have been notable advancements in video compression technologies in recent years, banding artifacts remain a serious issue affecting the quality of compressed videos, particularly on smooth regions of high-definition videos. Noticeable banding artifacts can severely impact the perceptual quality of videos viewed on a high-end HDTV or high-resolution screen. Hence, there is a pressing need for a systematic investigation of the banding video quality assessment problem for advanced video codecs. Given that the existing publicly available datasets for studying banding artifacts are limited to still picture data only, which cannot account for temporal banding dynamics, we have created a first-of-a-kind open video dataset, dubbed LIVE-YT-Banding, which consists of 160 videos generated by four different compression parameters using the AV1 video codec. A total of 7,200 subjective opinions are collected from a cohort of 45 human subjects. To demonstrate the value of this new resources, we tested and compared a variety of models that detect banding occurrences, and measure their impact on perceived quality. Among these, we introduce an effective and efficient new no-reference (NR) video quality evaluator which we call CBAND. CBAND leverages the properties of the learned statistics of natural images expressed in the embeddings of deep neural networks. Our experimental results show that the perceptual banding prediction performance of CBAND significantly exceeds that of previous state-of-the-art models, and is also orders of magnitude faster. Moreover, CBAND can be employed as a differentiable loss function to optimize video debanding models. The LIVE-YT-Banding database, code, and pre-trained model are all publically available at https://github.com/uniqzheng/CBAND.
Qi Zheng 0004, Li-Heng Chen, Chenlong He, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik, Yibo Fan, Zhengzhong Tu
IEEE Trans. Image Process.9
2024 Light the Night: A Multi-Condition Diffusion Framework for Unpaired Low-Light Enhancement in Autonomous Driving
abstract
Vision-centric perception systems for autonomous driving have gained considerable attention recently due to their cost-effectiveness and scalability, especially compared to LiDAR-based systems. However, these systems often struggle in low-light conditions, potentially compromising their performance and safety. To address this, our paper introduces LightDiff, a domain-tailored framework designed to enhance the low-light image quality for autonomous driving applications. Specifically, we employ a multi-condition controlled diffusion model. LightDiff works without any human-collected paired data, leveraging a dynamic data degradation process instead. It incorporates a novel multi-condition adapter that adaptively controls the input weights from different modalities, including depth maps, RGB images, and text captions, to effectively illuminate dark scenes while maintaining context consistency. Furthermore, to align the enhanced images with the detection model's knowledge, LightDiff employs perception-specific scores as rewards to guide the diffusion training process through reinforcement learning. Extensive experiments on the nuScenes datasets demonstrate that LightDiff can significantly improve the performance of several state-of-the-art 3D detectors in night-time conditions while achieving high visual quality scores, highlighting its potential to safeguard autonomous driving.
Zhengzhong Tu, Xinyu Liu 0009, Qing Guo 0005, Felix Juefei-Xu, Runsheng Xu, Hongkai Yu
CVPR3
2024 CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image Generation
abstract
Large generative diffusion models have revolution-ized text-to-image generation and offer immense po-tential for conditional generation tasks such as im-age enhancement, restoration, editing, and compositing. However, their widespread adoption is hindered by the high computational cost, which limits their real-time application. To address this challenge, we in-troduce a novel method dubbed CoDi, that adapts a pre-trained latent diffusion model to accept additional image conditioning inputs while significantly reducing the sampling steps required to achieve high-quality results. Our method can leverage architectures such as ControlNet to incorporate conditioning inputs with-out compromising the model's prior knowledge gained during large scale pre-training. Additionally, a con-ditional consistency loss enforces consistent predictions across diffusion steps, effectively compelling the model to generate high-quality images with conditions in a few steps. Our conditional-task learning and distil-lation approach outperforms previous distillation meth-ods, achieving a new state-of-the-art in producing high-quality images with very few steps (e.g., 1–4) across multiple tasks, including super-resolution, text-guided image editing, and depth-to-image generation.
Kangfu Mei, Mauricio Delbracio, Hossein Talebi, Zhengzhong Tu, Vishal M. Patel, Peyman Milanfar
CVPR4
2024 SPIRE: Semantic Prompt-Driven Image Restoration
Zhengzhong Tu, Keren Ye, Mauricio Delbracio, Peyman Milanfar, Qifeng Chen 0001, Hossein Talebi
ECCV (40)2
2024 FAVER: Blind quality prediction of variable frame rate videos
Qi Zheng 0004, Zhengzhong Tu, Pavan C. Madhusudana, Xiaoyang Zeng, Alan C. Bovik, Yibo Fan
Signal Process. Image Commun.2
2024 MWFormer: Multi-Weather Image Restoration Using Degradation-Aware Transformers
abstract
Restoring images captured under adverse weather conditions is a fundamental task for many computer vision applications. However, most existing weather restoration approaches are only capable of handling a specific type of degradation, which is often insufficient in real-world scenarios, such as rainy-snowy or rainy-hazy weather. Towards being able to address these situations, we propose a multi-weather Transformer, or MWFormer for short, which is a holistic vision Transformer that aims to solve multiple weather-induced degradations using a single, unified architecture. MWFormer uses hyper-networks and feature-wise linear modulation blocks to restore images degraded by various weather types using the same set of learned parameters. We first employ contrastive learning to train an auxiliary network that extracts content-independent, distortion-aware feature embeddings that efficiently represent predicted weather types, of which more than one may occur. Guided by these weather-informed predictions, the image restoration Transformer adaptively modulates its parameters to conduct both local and global feature processing, in response to multiple possible weather. Moreover, MWFormer allows for a novel way of tuning, during application, to either a single type of weather restoration or to hybrid weather restoration without any retraining, offering greater controllability than existing methods. Our experimental results on multi-weather restoration benchmarks show that MWFormer achieves significant performance improvements compared to existing state-of-the-art methods, without requiring much computational cost. Moreover, we demonstrate that our methodology of using hyper-networks can be integrated into various network architectures to further boost their performance. The code is available at: https://github.com/taco-group/MWFormer.
Ruoxi Zhu, Zhengzhong Tu, Alan C. Bovik, Yibo Fan
IEEE Trans. Image Process.2
2023 V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception
abstract
Modern perception systems of autonomous vehicles are known to be sensitive to occlusions and lack the capability of long perceiving range. It has been one of the key bottlenecks that prevents Level 5 autonomy. Recent research has demonstrated that the Vehicle-to-Vehicle (V2V) cooperative perception system has great potential to revolutionize the autonomous driving industry. However, the lack of a real-world dataset hinders the progress of this field. To facilitate the development of cooperative perception, we present V2V4Real, the first large-scale real-world multi-modal dataset for V2V perception. The data is collected by two vehicles equipped with multi-modal sensors driving together through diverse scenarios. Our V2V4Real dataset covers a driving area of 410 km, comprising 20K LiDAR frames, 40K RGB frames, 240K annotated 3D bounding boxes for 5 classes, and HDMaps that cover all the driving routes. V2V4Real introduces three perception tasks, including cooperative 3D object detection, cooperative 3D object tracking, and Sim2Real domain adaptation for cooperative perception. We provide comprehensive benchmarks of recent cooperative perception algorithms on three tasks. The V2V4Real dataset can be found at research.seas.ucla.edu/mobility-lab/v2v4real/.
Runsheng Xu, Xin Xia 0007, Hanzhao Li, Zhengzhong Tu, Zonglin Meng, Hao Xiang 0001, Rui Song 0007, Hongkai Yu, Bolei Zhou, Jiaqi Ma 0003
CVPR6
2023 MULLER: Multilayer Laplacian Resizer for Vision
abstract
Image resizing operation is a fundamental preprocessing module in modern computer vision. Throughout the deep learning revolution, researchers have overlooked the potential of alternative resizing methods beyond the commonly used resizers that are readily available, such as nearest-neighbors, bilinear, and bicubic. The key question of our interest is whether the front-end resizer affects the performance of deep vision models? In this paper, we present an extremely lightweight multilayer Laplacian resizer with only a handful of trainable parameters, dubbed MULLER resizer. MULLER has a bandpass nature in that it learns to boost details in certain frequency subbands that benefit the downstream recognition models. We show that MULLER can be easily plugged into various training pipelines, and it effectively boosts the performance of the underlying vision task with little to no extra cost. Specifically, we select a state-of-the-art vision Transformer, MaxViT [50], as the baseline, and show that, if trained with MULLER, MaxViT gains up to 0.6% top-1 accuracy, and meanwhile enjoys 36% inference cost saving to achieve similar top-1 accuracy on ImageNet-1k, as compared to the standard training scheme. Notably, MULLER’s performance also scales with model size and training data size such as ImageNet-21k and JFT, and it is widely applicable to multiple vision tasks, including image classification, object detection and segmentation, as well as image quality assessment. The code is available at https://github.com/google-research/google-research/tree/master/muller.
Zhengzhong Tu, Peyman Milanfar, Hossein Talebi
ICCV1
2023 Pik-Fix: Restoring and Colorizing Old Photos
abstract
Restoring and inpainting the visual memories that are present, but often impaired, in old photos remains an intriguing but unsolved research topic. Decades-old photos often suffer from severe and commingled degradation such as cracks, defocus, and color-fading, which are difficult to treat individually and harder to repair when they interact. Deep learning presents a plausible avenue, but the lack of large-scale datasets of old photos makes addressing this restoration task very challenging. Here we present a novel reference-based end-to-end learning framework that is able to both repair and colorize old, degraded pictures. Our proposed framework consists of three modules: a restoration sub-network that conducts restoration from degradations, a similarity network that performs color histogram matching and color transfer, and a colorization subnet that learns to predict the chroma elements of images conditioned on chromatic reference signals. The overall system makes uses of color histogram priors from reference images, which greatly reduces the need for large-scale training data. We have also created a first-of-a-kind public dataset of real old photos that are paired with ground truth "pristine" photos that have been manually restored by PhotoShop experts. We conducted extensive experiments on this dataset and synthetic datasets, and found that our method significantly outperforms previous state-of-the-art models using both qualitative comparisons and quantitative measurements. The code is available at https://github.com/DerrickXuNu/Pik-Fix.
Runsheng Xu, Zhengzhong Tu, Yuanqi Du, Zibo Meng, Jiaqi Ma 0003, Alan C. Bovik, Hongkai Yu
WACV2
2022 MAXIM: Multi-Axis MLP for Image Processing
abstract
Recent progress on Transformers and multilayer perceptron (MLP) models provide new network architectural designs for computer vision tasks. Although these models proved to be effective in many vision tasks such as image recognition, there remain challenges in adapting them for lowlevel vision. The inflexibility to support high-resolution images and limitations of local attention are perhaps the main bottlenecks. In this work, we present a multi-axis MLP based architecture called MAXIM, that can serve as an efficient and flexible general-purpose vision backbone for image processing tasks. MAXIM uses a UNet-shaped hierarchical structure and supports long-range interactions enabled by spatially-gated MLPs. Specifically, MAXIM contains two MLP-based building blocks: a multi-axis gated MLP that allows for efficient and scalable spatial mixing of local and global visual cues, and a cross-gating block, an alternative to cross-attention, which accounts for cross-feature conditioning. Both these modules are exclusively based on MLPs, but also benefit from being both global and ‘fully-convolutional’, two properties that are desirable for image processing. Our extensive experimental results show that the proposed MAXIM model achieves state-of-the-art performance on more than ten benchmarks across a range of image processing tasks, including denoising, deblurring, de raining, dehazing, and enhancement while requiring fewer or comparable numbers of parameters and FLOPs than competitive models. The source code and trained models will be available at https://github.com/google-research/maxim.
Zhengzhong Tu, Hossein Talebi, Han Zhang 0010, Feng Yang 0008, Peyman Milanfar, Alan C. Bovik, Yinxiao Li
CVPR1
2022 MaxViT: Multi-axis Vision Transformer
Zhengzhong Tu, Hossein Talebi, Han Zhang 0010, Feng Yang 0008, Peyman Milanfar, Alan C. Bovik, Yinxiao Li
ECCV (24)1
2022 V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision Transformer
Runsheng Xu, Hao Xiang 0001, Zhengzhong Tu, Xin Xia 0007, Ming-Hsuan Yang 0001, Jiaqi Ma 0003
ECCV (39)3
2022 No-Reference Quality Assessment of Variable Frame-Rate Videos Using Temporal Bandpass Statistics
abstract
Recent advances in mobile devices and cloud computing techniques have made it possible to capture, process, and share high resolution, high frame rate (HFR) videos across the Internet nearly instantaneously. Being able to monitor and control the quality of these streamed videos can enable the de-livery of many enjoyable content and perceptually optimized rate control. However, the development of no-reference (NR) VQA algorithms targeting frame rate variations has been little studied. Here, we propose a first-of-a-kind blind VQA model for evaluating HFR videos, which we dub the Framerate-Aware Videos Evaluator w/o Reference (FAVER). FAVER uses extended models of spatial natural scene statistics that encompass space-time wavelet-decomposed video signals, to conduct efficient frame rate sensitive quality prediction. Our extensive experiments on several HFR video quality datasets show that FAVER outperforms other blind VQA algorithms at a reasonable computational cost. The code will be released on https://github.com/uniqzheng/HFR-BVQA.
Qi Zheng 0004, Zhengzhong Tu, Yibo Fan, Xiaoyang Zeng, Alan C. Bovik
ICASSP2
2022 Blind Video Quality Assessment via Space-Time Slice Statistics
abstract
User-generated contents (UGC) have gained increased attention in the video quality community recently. Perceptual video quality assessment (VQA) of UGC videos is of great significance for content providers to monitor, process, and deliver massive numbers of UGC videos. Blind video quality prediction of UGC videos is challenging since complex mixtures of spatial and temporal distortions contribute to the overall perceptual quality. In this paper, we develop a simple, effective, and efficient blind VQA framework (STS-QA) based on the statistical analysis of space-time slices (STS) of videos. Specifically, we extract spatio-temporal statistical features along different orientations of video STS, that capture directional global motion, then train a shallow quality predictor. The proposed framework can be used to easily extend any existing video/image quality model to account for temporal or motion regularities. Our experimental results on three publicly available UGC databases demonstrate that our proposed STS-QA model can significantly boost prediction performance compared to baselines. The code will be released at: https://github.com/uniqzheng/STS_BVQA.
Qi Zheng 0004, Zhengzhong Tu, Zhijian Hao, Xiaoyang Zeng, Alan C. Bovik, Yibo Fan
ICIP2
2022 Completely Blind Video Quality Evaluator
abstract
Automatic video quality assessment of user-generated content (UGC) has gained increased interest recently, due to the ubiquity of shared video clips uploaded and circulated on social media platforms across the globe. Most existing video quality models developed for this vast content are trained on large numbers of samples labeled during large-scale subjective studies, which are often fail to exhibit adequate generalization abilities on unseen data. Moreover, large labeled video quality datasets are not always available for every scenario, and may not address the coincident evaluation of social videos and the distortions that afflict them. Because of this, it is also desirable to develop opinion-unaware, “completely blind” video quality models, that are free of training, yet can compete with existing learning-based models. Here we propose such a model called VIQE (VIdeo Quality Evaluator), which we designed based on a comprehensive analysis of patch- and frame-wise video statistics, as well as of space-time statistical regularities of videos. The statistical features desired from the analysis capture complementary predictive aspects of perceptual quality, which are aggregated to obtain final video quality scores. Extensive experiments on recent large-scale video quality databases demonstrate that VIQE is even competitive with state-of-the-art opinion-aware models. The source code is being made available athttps://github.com/uniqzheng/Complete-Blind-VQA.
Qi Zheng 0004, Zhengzhong Tu, Xiaoyang Zeng, Alan C. Bovik, Yibo Fan
IEEE Signal Process. Lett.2
2021 Regression or classification? New methods to evaluate no-reference picture and video quality models
abstract
Video and image quality assessment has long been projected as a regression problem, which requires predicting a continuous quality score given an input stimulus. However, recent efforts have shown that accurate quality score regression on real-world user-generated content (UGC) is a very challenging task. To make the problem more tractable, we propose two new methods - binary, and ordinal classification - as alternatives to evaluate and compare no-reference quality models at coarser levels. Moreover, the proposed new tasks convey more practical meaning on perceptually optimized UGC transcoding, or for preprocessing on media processing platforms. We conduct a comprehensive benchmark experiment of popular no-reference quality models on recent in-the-wild picture and video quality datasets, providing reliable baselines for both evaluation methods to support further studies. We hope this work promotes coarse-grained perceptual modeling and its applications to efficient UGC processing.
Zhengzhong Tu, Chia-Ju Chen, Li-Heng Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik
ICASSP1
2021 Video Quality Assessment of User Generated Content: A Benchmark Study and a New Model
abstract
Recent years have witnessed an explosion of user-generated content (UGC) shared and streamed over the Internet. Accordingly, there is a great need for accurate video quality assessment (VQA) models for consumer or UGC videos to monitor, control, and optimize this vast content. Here we contribute to advancing the UGC-VQA problem by conducting a comprehensive evaluation of leading blind VQA (BVQA) models. Besides, we also created a new fusion-based BVQA model, which we dub the VIDeo quality EVALuator (VIDEVAL), that effectively balances the trade-off between performance and efficiency. Our experimental results show that VIDEVAL achieves state-of-the-art performance at a lower computational cost. We believe our reliable and reproducible benchmark will facilitate further research on deep learning-based BVQA modeling. An implementation of VIDEVAL has been made available online1.1https://github.com/vztu/VIDEVAL_release
Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik
ICIP1
2021 A Temporal Statistics Model For UGC Video Quality Prediction
abstract
Blind video quality assessment of user-generated content (UGC) has become a trending and challenging problem. Previous studies have shown the efficacy of natural scene statistics for capturing spatial distortions. The exploration of temporal video statistics on UGC, however, is relatively limited. Here we propose the first general, effective and efficient temporal statistics model accounting for temporal- or motion-related distortions for UGC video quality assessment, by analyzing regularities in the temporal bandpass domain. The proposed temporal model can serve as a plug-in module to boost existing no-reference video quality predictors that lack motion-relevant features. Our experimental results on recent large-scale UGC video databases show that the proposed model can significantly improve the performances of existing methods, at a very reasonable computational expense.
Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik
ICIP1
2021 Efficient User-Generated Video Quality Prediction
abstract
Blind video quality assessment of user-generated content (UGC) has become a trending, challenging, unsolved problem. Accurate and efficient video quality predictors suitable for this content are thus in great demand to achieve intelligent analysis and processing of UGC videos. However, previous video quality models are either incapable or inefficient for predicting the quality of complex, diverse UGC videos in practical applications. Here we introduce an effective and efficient video quality model for UGC content, which we dub the Rapid and Accurate Video Quality Evaluator (RAPIQUE), which we show performs comparably to state-of-the-art models but with orders-of-magnitude faster runtime. Our experimental results on recent large-scale UGC video quality databases show that RAPIQUE delivers top performances on all datasets at a considerably lower computational expense. An implementation of RAPIQUE is online: https://github.com/vztu/RAPIQUE.
Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik
PCS1
2021 UGC-VQA: Benchmarking Blind Video Quality Assessment for User Generated Content
abstract
Recent years have witnessed an explosion of user-generated content (UGC) videos shared and streamed over the Internet, thanks to the evolution of affordable and reliable consumer capture devices, and the tremendous popularity of social media platforms. Accordingly, there is a great need for accurate video quality assessment (VQA) models for UGC/consumer videos to monitor, control, and optimize this vast content. Blind quality prediction of in-the-wild videos is quite challenging, since the quality degradations of UGC videos are unpredictable, complicated, and often commingled. Here we contribute to advancing the UGC-VQA problem by conducting a comprehensive evaluation of leading no-reference/blind VQA (BVQA) features and models on a fixed evaluation architecture, yielding new empirical insights on both subjective video quality studies and objective VQA model design. By employing a feature selection strategy on top of efficient BVQA models, we are able to extract 60 out of 763 statistical features used in existing methods to create a new fusion-based model, which we dub the VIDeo quality EVALuator (VIDEVAL), that effectively balances the trade-off between VQA performance and efficiency. Our experimental results show that VIDEVAL achieves state-of-the-art performance at considerably lower computational cost than other leading models. Our study protocol also defines a reliable benchmark for the UGC-VQA problem, which we believe will facilitate further research on deep learning-based VQA modeling, as well as perceptually-optimized efficient UGC video processing, transcoding, and streaming. To promote reproducible research and public evaluation, an implementation of VIDEVAL has been made available online: https://github.com/vztu/VIDEVAL.
Zhengzhong Tu, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik
IEEE Trans. Image Process.1
2020 BBAND INDEX: A NO-REFERENCE BANDING ARTIFACT PREDICTOR
abstract
Banding artifact, or false contouring, is a common video compression impairment that tends to appear on large flat regions in encoded videos. These staircase-shaped color bands can be very noticeable in high-definition videos. Here we study this artifact, and propose a new distortion-specific no-reference video quality model for predicting banding artifacts, called the Blind BANding Detector (BBAND index). BBAND is inspired by human visual models. The proposed detector can generate a pixel-wise banding visibility map and output a banding severity score at both the frame and video levels. Experimental results show that our proposed method outperforms state-of-the-art banding detection algorithms and delivers better consistency with subjective evaluations.
Zhengzhong Tu, Jessie Lin, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik
ICASSP1
2020 A Comparative Evaluation Of Temporal Pooling Methods For Blind Video Quality Assessment
abstract
Many objective video quality assessment (VQA) algorithms include a key step of temporal pooling of frame-level quality scores. However, less attention has been paid to studying the relative efficiencies of different pooling methods on noreference (blind) VQA. Here we conduct a large-scale comparative evaluation to assess the capabilities and limitations of multiple temporal pooling strategies on blind VQA of usergenerated videos. The study yields insights and general guidance regarding the application and selection of temporal pooling models. In addition, we also propose an ensemble pooling model built on top of high-performing temporal pooling models. Our experimental results demonstrate the relative efficacies of the evaluated temporal pooling models, using several popular VQA algorithms evaluated on two recent largescale natural video quality databases. Conclusively, we also provide an empirical recipe for applying temporal pooling of frame-based quality predictions.
Zhengzhong Tu, Chia-Ju Chen, Li-Heng Chen, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik
ICIP1
2020 Adaptive Debanding Filter
abstract
Banding artifacts, which manifest as staircase-like color bands on pictures or video frames, is a common distortion caused by compression of low-textured smooth regions. These false contours can be very noticeable even on high-quality videos, especially when displayed on high-definition screens. Yet, relatively little attention has been applied to this problem. Here we consider banding artifact removal as a visual enhancement problem, and accordingly, we solve it by applying a form of content-adaptive smoothing filtering followed by dithered quantization, as a post-processing module. The proposed debanding filter is able to adaptively smooth banded regions while preserving image edges and details, yielding perceptually enhanced gradient rendering with limited bit-depths. Experimental results show that our proposed debanding filter outperforms state-of-the-art false contour removing algorithms both visually and quantitatively.
Zhengzhong Tu, Jessie Lin, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik
IEEE Signal Process. Lett.1