Qihang Zhang

dblp:282/1036 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 World-consistent Video Diffusion with Explicit 3D Modeling
abstract
Recent advancements in diffusion models have set new benchmarks in image and video generation, enabling realistic visual synthesis across single- and multi-frame contexts. However, these models still struggle with efficiently and explicitly generating 3D-consistent content. To address this, we propose World-consistent Video Diffusion (WVD), a novel framework that incorporates explicit 3D supervision using XYZ images, which encode global 3D coordinates for each image pixel. More specifically, we train a diffusion transformer to learn the joint distribution of RGB and XYZ frames. This approach supports multi-task adaptability via a flexible inpainting strategy. For example, WVD can estimate XYZ frames from ground-truth RGB or generate novel RGB frames using XYZ projections along a specified camera trajectory. In doing so, WVD unifies tasks like singleimage-to-3D generation, multi-view stereo, and camera-controlled video generation. Our approach demonstrates competitive performance across multiple benchmarks, providing a scalable solution for 3D-consistent video and image generation with a single pretrained model. Our project website is at https://zqh0253.github.io/wvd.
Qihang Zhang, Shuangfei Zhai, Miguel Ángel Bautista Martin, Kevin Miao, Alexander Toshev, Joshua M. Susskind, Jiatao Gu
CVPR1
2025 Denoising Autoregressive Transformers for Scalable Text-to-Image Generation
abstract
Diffusion models have become the dominant approach for visual generation. They are trained by denoising a Markovian process which gradually adds noise to the input. We argue that the Markovian property limits the model’s ability to fully utilize the generation trajectory, leading to inefficiencies during training and inference. In this paper, we propose DART, a transformer-based model that unifies autoregressive (AR) and diffusion within a non-Markovian framework. DART iteratively denoises image patches spatially and spectrally using an AR model that has the same architecture as standard language models. DART does not rely on image quantization, which enables more effective image modeling while maintaining flexibility. Furthermore, DART seamlessly trains with both text and image data in a unified model. Our approach demonstrates competitive performance on class-conditioned and text-to-image generation tasks, offering a scalable, efficient alternative to traditional diffusion models. Through this unified framework, DART sets a new benchmark for scalable, high-quality image synthesis.
Jiatao Gu, Yizhe Zhang 0002, Qihang Zhang, Dinghuai Zhang, Navdeep Jaitly, Joshua M. Susskind, Shuangfei Zhai
ICLR4
2025 3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian Splatting
abstract
Scene image editing is crucial for entertainment, photography, and advertising design. Existing methods solely focus on either 2D individual object or 3D global scene editing. This results in a lack of a unified approach to effectively control and manipulate scenes at the 3D level with different levels of granularity. In this work, we propose 3DitScene, a novel and unified scene editing framework leveraging language-guided disentangled Gaussian Splatting that enables seamless editing from 2D to 3D, allowing precise control over scene composition and individual objects. We first incorporate 3D Gaussians that are refined through generative priors and optimization techniques. Language features from CLIP then introduce semantics into 3D geometry for object disentanglement. With the disentangled Gaussians, 3DitScene allows for manipulation at both the global and individual levels, revolutionizing creative expression and empowering control over scenes and objects. Experimental results demonstrate the effectiveness and versatility of 3DitScene in scene image editing.
Qihang Zhang, Yinghao Xu 0001, Chaoyang Wang 0001, Hsin-Ying Lee 0001, Gordon Wetzstein, Bolei Zhou, Ceyuan Yang
ICLR1
2025 A Multi-Granularity Clustering Approach for Federated Backdoor Defense with the Adam Optimizer
abstract
Federated learning is vulnerable to backdoor attacks due to its distributed nature and the inability to access local datasets. Meanwhile, the heterogeneity of distributed data further complicates the detection of such attacks. However, existing defense strategies often overlook the presence of non-stationary objectives and noisy gradients across multiple clients, making it challenging to accurately and efficiently identify malicious participants. To address these challenges, we propose a backdoor defense method for Federated Learning with Adam optimizer and multi-granularity Clustering (FLAC), incorporating both coarse-grained and fine-grained clustering mechanisms to neutralize backdoor attacks. First, the Adam optimizer accelerates the learning process by mitigating the impact of noisy gradients and addressing the non-stationary objectives posed by different clients under attack. Second, a multi-granularity clustering process is considered to differentiate between benign clients and potential attackers. This is followed by an adaptive clipping strategy to further alleviate the influence of malicious attackers. Our theoretical analysis demonstrates the consistent convergence of Adam in a federated backdoor defense environment. Extensive experimental results validate the effectiveness of our defense approach.
Jidong Yuan, Qihang Zhang, Naiyue Chen, Shengbo Chen, Baomin Xu
IJCAI2
2025 Syn-rPPG: Improving unsupervised remote photoplethysmography extraction with synthesized videos using generative models
Hanguang Xiao, Yisha Sun, Kun Zuo, Qihang Zhang, Feizhong Zhou
Eng. Appl. Artif. Intell.5
2024 Towards Text-guided 3D Scene Composition
abstract
We are witnessing significant breakthroughs in the tech-nology for generating 3D objects from text. Existing approaches either leverage large text-to-image models to optimize a 3D representation or train 3D generators on object-centric datasets. Generating entire scenes, however, remains very challenging as a scene contains multiple 3D objects, diverse and scattered. In this work, we introduce SceneWiz3D - a novel approach to synthesize high-fidelity 3D scenes from text. We marry the locality of objects with globality of scenes by introducing a hybrid 3D representation - explicit for objects and implicit for scenes. Remarkably, an object, being represented explicitly, can be either generated from text using conventional text-to-3D approaches, or provided by users. To configure the layout of the scene and automatically place objects, we apply the Particle Swarm Optimization technique during the optimization process. Furthermore, it is difficult for certain parts of the scene (e.g., corners, occlusion) to receive multi-view supervision, leading to inferior geometry. We incor-porate an RGBD panorama diffusion model to mitigate it, resulting in high-quality geometry. Extensive evaluation supports that our approach achieves superior quality over previous approaches, enabling the generation of detailed and view-consistent 3D scenes. Our project website is at https://zqh0253.github.io/SceneWiz3D/.
Qihang Zhang, Chaoyang Wang 0001, Aliaksandr Siarohin, Peiye Zhuang, Yinghao Xu 0001, Ceyuan Yang, Dahua Lin, Bolei Zhou, Sergey Tulyakov, Hsin-Ying Lee 0001
CVPR1
2024 BerfScene: Bev-conditioned Equivariant Radiance Fields for Infinite 3D Scene Generation
abstract
Generating large-scale 3D scenes cannot simply apply existing 3D object synthesis technique since 3D scenes usually hold complex spatial configurations and consist of a number of objects at varying scales. We thus propose a practical and efficient 3D representation that incorporates an equivariant radiance field with the guidance of a bird's-eye view (BEV) map. Concretely, objects of synthesized 3D scenes could be easily manipulated through steering the corresponding BEV maps. Moreover, by adequately incorporating positional encoding and low-pass filters into the generator, the representation becomes equivariant to the given BEV map. Such equivariance allows us to produce large-scale, even infinite-scale, 3D scenes via synthesizing local scenes and then stitching them with smooth consistency. Extensive experiments on 3D scene datasets demonstrate the effectiveness of our approach. Our project web-site is at: https://zqh0253.github.io/BerfScene/.
Qihang Zhang, Yinghao Xu 0001, Yujun Shen, Bo Dai 0002, Bolei Zhou, Ceyuan Yang
CVPR1
2024 A High Efficiency, Low EMI Non-inverting Buck-Boost Converter in Wireless Power and Data Transfer System for Brain Computer Interface
abstract
This paper presents a high efficiency, fixed- frequency waveform Buck-Boost converter which converts a 2.7V to 4.5V Li-ion battery voltage to the regulated voltage featured high efficiency and low electromagnetic interference (EMI) for brain computer interface applications. A smoothly switching peak-valley current control mode, enabling this converter to switch freely from valley-current-mode control in Buck mode to peak-current-mode control in Boost mode, is adopted to avoid sub-harmonic oscillations. And a fixed- frequency control technique based on cycle alternation in Buck-Boost mode, which fixes the frequency of the inductor current as opposed to random-mode-control, is adopted to obtain high efficiency and low EMI. The proposed Buck-Boost converter is implemented in 180nm BCD process. It can achieve and keep the efficiency of above 92.33%, of which the peak efficiency is 98.97%, over the load current ranging from 100mA to 400mA and the input voltage ranging from 2.7V to 4.5V.
Xiangsheng Xu, Qihang Zhang, Songping Mai
ISCAS2
2024 Hundreds Guide Millions: Adaptive Offline Reinforcement Learning With Expert Guidance
abstract
Offline reinforcement learning (RL) optimizes the policy on a previously collected dataset without any interactions with the environment, yet usually suffers from the distributional shift problem. To mitigate this issue, a typical solution is to impose a policy constraint on a policy improvement objective. However, existing methods generally adopt a "one-size-fits-all" practice, i.e., keeping only a single improvement-constraint balance for all the samples in a mini-batch or even the entire offline dataset. In this work, we argue that different samples should be treated with different policy constraint intensities. Based on this idea, a novel plug-in approach named guided offline RL (GORL) is proposed. GORL employs a guiding network, along with only a few expert demonstrations, to adaptively determine the relative importance of the policy improvement and policy constraint for every sample. We theoretically prove that the guidance provided by our method is rational and near-optimal. Extensive experiments on various environments suggest that GORL can be easily installed on most offline RL algorithms with statistically significant performance improvements.
Qisen Yang, Shenzhi Wang, Qihang Zhang, Gao Huang 0001, Shiji Song
IEEE Trans. Neural Networks Learn. Syst.3
2023 GeoMIM: Towards Better 3D Knowledge Transfer via Masked Image Modeling for Multi-view 3D Understanding
abstract
Multi-view camera-based 3D detection is a challenging problem in computer vision. Recent works leverage a pretrained LiDAR detection model to transfer knowledge to a camera-based student network. However, we argue that there is a major domain gap between the LiDAR BEV features and the camera-based BEV features, as they have different characteristics and are derived from different sources. In this paper, we propose Geometry Enhanced Masked Image Modeling (GeoMIM) to transfer the knowledge of the LiDAR model in a pretrain-finetune paradigm for improving the multi-view camera-based 3D detection. GeoMIM is a multi-camera vision transformer with Cross-View Attention (CVA) blocks that uses LiDAR BEV features encoded by the pretrained BEV model as learning targets. During pretraining, GeoMIM’s decoder has a semantic branch completing dense perspective-view features and the other geometry branch reconstructing dense perspective-view depth maps. The depth branch is designed to be camera-aware by inputting the camera’s parameters for better transfer capability. Extensive results demonstrate that GeoMIM outperforms existing methods on nuScenes benchmark, achieving state-of-the-art performance for camera-based 3D object detection and 3D segmentation.
Jihao Liu, Boxiao Liu, Qihang Zhang, Yu Liu 0015, Hongsheng Li 0001
ICCV4
2023 StPrformer: A Stock Price Prediction Model Based on Convolutional Attention Mechanism
Zhaoguo Liu, Qihang Zhang
ICIC (5)2
2023 CWA-LSTM: A Stock Price Prediction Model Based on Causal Weight Adjustment
Qihang Zhang, Zhaoguo Liu, Zhuoer Wen
ICIC (5)1
2023 Towards Smooth Video Composition
Qihang Zhang, Ceyuan Yang, Yujun Shen, Yinghao Xu 0001, Bolei Zhou
ICLR1
2023 Learning Modulated Transformation in GANs
abstract
The success of style-based generators largely benefits from style modulation, which helps take care of the cross-instance variation within data. However, the instance-wise stochasticity is typically introduced via regular convolution, where kernels interact with features at some fixed locations, limiting its capacity for modeling geometric variation. To alleviate this problem, we equip the generator in generative adversarial networks (GANs) with a plug-and-play module, termed as modulated transformation module (MTM). This module predicts spatial offsets under the control of latent codes, based on which the convolution operation can be applied at variable locations for different instances, and hence offers the model an additional degree of freedom to handle geometry deformation. Extensive experiments suggest that our approach can be faithfully generalized to various generative tasks, including image generation, 3D-aware image synthesis, and video generation, and get compatible with state-of-the-art frameworks without any hyper-parameter tuning. It is noteworthy that, towards human generation on the challenging TaiChi dataset, we improve the FID of StyleGAN3 from 21.36 to 13.60, demonstrating the efficacy of learning modulated geometry transformation. Code and models are available at https://github.com/limbo0000/mtm.
Ceyuan Yang, Qihang Zhang, Yinghao Xu 0001, Jiapeng Zhu 0001, Yujun Shen, Bo Dai 0002
NeurIPS2
2023 MetaDrive: Composing Diverse Driving Scenarios for Generalizable Reinforcement Learning
abstract
Driving safely requires multiple capabilities from human and intelligent agents, such as the generalizability to unseen environments, the safety awareness of the surrounding traffic, and the decision-making in complex multi-agent settings. Despite the great success of Reinforcement Learning (RL), most of the RL research works investigate each capability separately due to the lack of integrated environments. In this work, we develop a new driving simulation platform called MetaDrive to support the research of generalizable reinforcement learning algorithms for machine autonomy. MetaDrive is highly compositional, which can generate an infinite number of diverse driving scenarios from both the procedural generation and the real data importing. Based on MetaDrive, we construct a variety of RL tasks and baselines in both single-agent and multi-agent settings, including benchmarking generalizability across unseen scenes, safe exploration, and learning multi-agent traffic. The generalization experiments conducted on both procedurally generated scenarios and real-world scenarios show that increasing the diversity and the size of the training set leads to the improvement of the RL agent's generalizability. We further evaluate various safe reinforcement learning and multi-agent reinforcement learning algorithms in MetaDrive environments and provide the benchmarks. Source code, documentation, and demo video are available at https://metadriverse.github.io/metadrive.
Quanyi Li, Zhenghao Peng, Lan Feng, Qihang Zhang, Zhenghai Xue, Bolei Zhou
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Sublessor: A Cost-Saving Internet Transit Mechanism for Cooperative MEC Providers in Industrial Internet of Things
abstract
Mobile edge computing (MEC) is becoming increasingly popular due to its remarkable computing capacities in close proximity to end users or devices. With the widespread use of Industrial Internet of Things, more and more cloud service providers move their services to the edge of the network for a better quality of service and become MEC providers. These MEC providers require to rent wide area network (WAN) connections to transfer industrial data, which is a considerable expense. In this article, we propose a framework calledSublessorto reduce the WAN transmission cost for a group of cooperative MEC providers. The key idea ofSublessoris allowing some specific MEC providers to act as Internet transit brokers, transmitting not only their own network traffic but also the traffic of their partners under a reasonable reselling price. This article formulates the problem as a mixed-integer programming and finds the most suitable broker number and corresponding reselling price without damaging the profit of both brokers and partners by a deep-reinforcement-learning-based algorithm. Experimental results show that our algorithm can significantly reduce the traffic transmission cost by up to 35%.
Sheng Chen 0015, Qihang Zhang, Xiaodong Dong, Xiaoyi Tao, Keqiu Li, Tie Qiu 0001, Ivan Lee 0001
IEEE Trans. Ind. Informatics2
2022 Learning to Drive by Watching YouTube Videos: Action-Conditioned Contrastive Policy Pretraining
Qihang Zhang, Zhenghao Peng, Bolei Zhou
ECCV (26)1
2022 Secure RFID Handwriting Recognition-Attacker Can Hear but Cannot Understand
Qihang Zhang, Jiuwu Zhang, Xiulong Liu 0001, Xinyu Tong 0001, Keqiu Li
WASA (1)1
2021 F³A-GAN: Facial Flow for Face Animation With Generative Adversarial Networks
abstract
Formulated as a conditional generation problem, face animation aims at synthesizing continuous face images from a single source image driven by a set of conditional face motion. Previous works mainly model the face motion as conditions with 1D or 2D representation (e.g., action units, emotion codes, landmark), which often leads to low-quality results in some complicated scenarios such as continuous generation and large-pose transformation. To tackle this problem, the conditions are supposed to meet two requirements, i.e., motion information preserving and geometric continuity. To this end, we propose a novel representation based on a 3D geometric flow, termed facial flow, to represent the natural motion of the human face at any pose. Compared with other previous conditions, the proposed facial flow well controls the continuous changes to the face. After that, in order to utilize the facial flow for face editing, we build a synthesis framework generating continuous images with conditional facial flows. To fully take advantage of the motion information of facial flows, a hierarchical conditional framework is designed to combine the extracted multi-scale appearance features from images and motion features from flows in a hierarchical manner. The framework then decodes multiple fused features back to images progressively. Experimental results demonstrate the effectiveness of our method compared to other state-of-the-art methods.
Xintian Wu, Qihang Zhang, Yiming Wu 0005, Lingyun Sun, Xi Li 0001
IEEE Trans. Image Process.2