Yezhen Wang

dblp:249/2195 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
11since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DTR: Dynamic Tree-Ring Watermarking Framework for Diffusion-Based Video Generation
abstract
The growing capabilities of diffusion-based text-to-video models have raised significant concerns about copyright protection and the traceability of synthetic video content. To address these concerns, existing watermarking techniques have been developed to invisibly embed information within video content. However, many of these techniques lack resilience when faced with different distortions. This makes existing methods unsuitable for real-world applications, where videos frequently undergo compression, scaling, and other transformations. In response to this challenge, we propose DTR, a training-free Dynamic Tree-Ring watermarking framework designed to enhance the robustness of video watermarks while preserving video quality. Our DTR embeds the predefined user-specific key into the Fourier domain of the initial latent variables. Then, the invertible Fourier transformation ensures the invisible nature of such an insertion. To further improve video quality, the key is divided and embedded as discrete segments across different initial latents, thereby minimizing distribution shifts caused by the watermark insertion. Comprehensive experiments demonstrate the superiority of the DTR framework compared to traditional approaches in terms of robustness, especially under the H.264 and H.265 compression standards, while maintaining high video quality and inference efficiency.
Shunyang Zeng, Yezhen Wang
ICASSP4
2025 Empirical Analysis of Energy Drift in Battery Energy Storage Systems on Supporting Grid Frequency Stability
abstract
Battery energy storage systems (BESS) are crucial for maintaining grid frequency stability, particularly with the increasing integration of intermittent renewable energy sources. However, energy delivered from BESS drifts over time due to asymmetric charging and discharging inefficiencies, posing significant challenges for effective frequency support. In this work, simulations revealed that energy drift accelerated early crossing of the state of charge (SOC) boundaries. A sensitivity analysis of asymmetric inefficiencies showed that improving charging and discharging consistency reduced energy drift and decreased frequency of energy transactions. C-rate constraints highlighted a trade-off between BESS safety and the rapidity of frequency support responses. Simplified dynamic droop control was shown to stabilize early-stage grid operations but required degradation-informed BESS control to mitigate the energy drift. The sizing of BESS showcased a trade-off between energy drift and cost. These findings offer valuable insights into designing safe and robust SOC management strategies for grid-connected BESS. Broadly, this work enhances fundamental understanding of energy drift arising from inherent battery degradation and its impact on the reliable energy storage systems for the stable and sustainable power grid.
Shengyu Tao, Yezhen Wang, Hanyang Lin, Scott J. Moura, Hongbin Sun 0002, Qiuwei Wu, Xuan Zhang 0004
IECON2
2024 Understanding and Improving Training-free Loss-based Diffusion Guidance
abstract
Adding additional guidance to pretrained diffusion models has become an increasingly popular research area, with extensive applications in computer vision, reinforcement learning, and AI for science. Recently, several studies have proposed training-free loss-based guidance by using off-the-shelf networks pretrained on clean images. This approach enables zero-shot conditional generation for universal control formats, which appears to offer a free lunch in diffusion guidance. In this paper, we aim to develop a deeper understanding of training-free guidance, as well as overcome its limitations. We offer a theoretical analysis that supports training-free guidance from the perspective of optimization, distinguishing it from classifier-based (or classifier-free) guidance. To elucidate their drawbacks, we theoretically demonstrate that training-free guidance is more susceptible to misaligned gradients and exhibits slower convergence rates compared to classifier guidance. We then introduce a collection of techniques designed to overcome the limitations, accompanied by theoretical rationale and empirical evidence. Our experiments in image and motion generation confirm the efficacy of these techniques.
Yifei Shen 0004, Xinyang Jiang, Yifan Yang 0004, Yezhen Wang, Dongsheng Li 0002
NeurIPS4
2024 Memory-Efficient Gradient Unrolling for Large-Scale Bi-level Optimization
abstract
Bi-level optimizaiton (BO) has become a fundamental mathematical framework for addressing hierarchical machine learning problems. As deep learning models continue to grow in size, the demand for scalable bi-level optimization has become increasingly critical. Traditional gradient-based bi-level optimizaiton algorithms, due to their inherent characteristics, are ill-suited to meet the demands of large-scale applications. In this paper, we introduce **F**orward **G**radient **U**nrolling with **F**orward **G**radient, abbreviated as **$($FG$)^2$U**, which achieves an unbiased stochastic approximation of the meta gradient for bi-level optimizaiton. $($FG$)^2$U circumvents the memory and approximation issues associated with classical bi-level optimizaiton approaches, and delivers significantly more accurate gradient estimates than existing large-scale bi-level optimizaiton approaches. Additionally, $($FG$)^2$U is inherently designed to support parallel computing, enabling it to effectively leverage large-scale distributed computing systems to achieve significant computational efficiency. In practice, $($FG$)^2$U and other methods can be strategically placed at different stages of the training process to achieve a more cost-effective two-phase paradigm. Further, $($FG$)^2$U is easy to implement within popular deep learning frameworks, and can be conveniently adapted to address more challenging zeroth-order bi-level optimizaiton scenarios. We provide a thorough convergence analysis and a comprehensive practical discussion for $($FG$)^2$U, complemented by extensive empirical evaluations, showcasing its superior performance in diverse large-scale bi-level optimizaiton tasks.
Qianli Shen, Yezhen Wang, Zhouhao Yang, Jonathan Scarlett, Zhanxing Zhu, Kenji Kawaguchi
NeurIPS2
2023 Sparse Mixture-of-Experts are Domain Generalizable Learners
Bo Li 0080, Yifei Shen 0004, Yezhen Wang, Jiawei Ren 0001, Tong Che, Jun Zhang 0004, Ziwei Liu 0002
ICLR4
2022 Invariant Information Bottleneck for Domain Generalization
abstract
Invariant risk minimization (IRM) has recently emerged as a promising alternative for domain generalization. Nevertheless, the loss function is difficult to optimize for nonlinear classifiers and the original optimization objective could fail when pseudo-invariant features and geometric skews exist. Inspired by IRM, in this paper we propose a novel formulation for domain generalization, dubbed invariant information bottleneck (IIB). IIB aims at minimizing invariant risks for nonlinear classifiers and simultaneously mitigating the impact of pseudo-invariant features and geometric skews. Specifically, we first present a novel formulation for invariant causal prediction via mutual information. Then we adopt the variational formulation of the mutual information to develop a tractable loss function for nonlinear classifiers. To overcome the failure modes of IRM, we propose to minimize the mutual information between the inputs and the corresponding representations. IIB significantly outperforms IRM on synthetic datasets, where the pseudo-invariant features and geometric skews occur, showing the effectiveness of proposed formulation in overcoming failure modes of IRM. Furthermore, experiments on DomainBed show that IIB outperforms 13 baselines by 0.9% on average across 7 real datasets.
Bo Li 0080, Yifei Shen 0004, Yezhen Wang, Wenzhen Zhu, Colorado Reed, Dongsheng Li 0002, Kurt Keutzer, Han Zhao 0002
AAAI3
2022 SPE: Symmetrical Prompt Enhancement for Fact Probing
abstract
Pretrained language models (PLMs) have been shown to accumulate factual knowledge during pretraining (Petroni et al., 2019).Recent works probe PLMs for the extent of this knowledge through prompts either in discrete or continuous forms.However, these methods do not consider symmetry of the task: object prediction and subject prediction.In this work, we propose Symmetrical Prompt Enhancement (SPE), a continuous prompt-based method for factual probing in PLMs that leverages the symmetry of the task by constructing symmetrical prompts for subject and object prediction.Our results on a popular factual probing dataset, LAMA, show significant improvement of SPE over previous probing methods.
Yiyuan Li, Tong Che, Yezhen Wang, Zhengbao Jiang, Caiming Xiong, Snigdha Chaturvedi
EMNLP3
2021 ePointDA: An End-to-End Simulation-to-Real Domain Adaptation Framework for LiDAR Point Cloud Segmentation
abstract
Due to its robust and precise distance measurements, LiDAR plays an important role in scene understanding for autonomous driving. Training deep neural networks (DNNs) on LiDAR data requires large-scale point-wise annotations, which are time-consuming and expensive to obtain. Instead, simulation-to-real domain adaptation (SRDA) trains a DNN using unlimited synthetic data with automatically generated labels and transfers the learned model to real scenarios. Existing SRDA methods for LiDAR point cloud segmentation mainly employ a multi-stage pipeline and focus on feature-level alignment. They require prior knowledge of real-world statistics and ignore the pixel-level dropout noise gap and the spatial feature gap between different domains. In this paper, we propose a novel end-to-end framework, named ePointDA, to address the above issues. Specifically, ePointDA consists of three modules: self-supervised dropout noise rendering, statistics-invariant and spatially-adaptive feature alignment, and transferable segmentation learning. The joint optimization enables ePointDA to bridge the domain shift at the pixel-level by explicitly rendering dropout noise for synthetic LiDAR and at the feature-level by spatially aligning the features between different domains, without requiring the real-world statistics. Extensive experiments adapting from synthetic GTA-LiDAR to real KITTI and SemanticKITTI demonstrate the superiority of ePointDA for LiDAR point cloud segmentation.
Sicheng Zhao, Yezhen Wang, Bo Li 0080, Bichen Wu, Yang Gao 0029, Pengfei Xu 0013, Trevor Darrell, Kurt Keutzer
AAAI2
2021 Learning Invariant Representations and Risks for Semi-Supervised Domain Adaptation
abstract
The success of supervised learning hinges on the assumption that the training and test data come from the same underlying distribution, which is often not valid in practice due to potential distribution shift. In light of this, most existing methods for unsupervised domain adaptation focus on achieving domain-invariant representations and small source domain error. However, recent works have shown that this is not sufficient to guarantee good generalization on the target domain, and in fact, is provably detrimental under label distribution shift. Furthermore, in many real-world applications it is often feasible to obtain a small amount of labeled data from the target domain and use them to facilitate model training with source data. Inspired by the above observations, in this paper we propose the first method that aims to simultaneously learn invariant representations and risks under the setting of semi-supervised domain adaptation (Semi-DA). First, we provide a finite sample bound for both classification and regression problems under Semi-DA. The bound suggests a principled way to obtain target generalization, i.e., by aligning both the marginal and conditional distributions across domains in feature space. Motivated by this, we then introduce the LIRR algorithm for jointly Learning Invariant Representations and Risks. Finally, extensive experiments are conducted on both classification and regression tasks, which demonstrate that LIRR consistently achieves state-of-the-art performance and significant improvements compared with the methods that only learn invariant representations or invariant risks. Our code will be released at LIRR@github
Bo Li 0080, Yezhen Wang, Shanghang Zhang, Dongsheng Li 0002, Kurt Keutzer, Trevor Darrell, Han Zhao 0002
CVPR2
2021 Dual Metric Discriminator for Open Set Video Domain Adaptation
abstract
Existing video domain adaptation methods focus on addressing closed set problems. However, it is nearly impossible to guarantee different domains share exactly the same set of categories in realistic scenarios. Hence, open set video domain adaptation (OSVDA) problem, which involves unknown categories, has achieved increasingly close attention. In this paper, we propose a seminal framework, which involves spatial and temporal information to address OSVDA problem. Besides, we design a novel discrimination module, i.e., Dual Metric Discriminator (DMD), to separate known and unknown categories based on implicit and explicit similarity metrics. We conduct comprehensive experiments on several benchmarks and achieve state-of-the-art performance with 40.4%, 33.7%, and 79.2% accuracy on UCF to HMDB, HMDB to UCF, and Kinetics to UCF scenarios respectively.
Yatian Wang, Yezhen Wang, Pengfei Xu 0013, Runbo Hu
ICASSP3
2021 Energy-Based Open-World Uncertainty Modeling for Confidence Calibration
abstract
Confidence calibration is of great importance to the reliability of decisions made by machine learning systems. However, discriminative classifiers based on deep neural networks are often criticized for producing overconfident predictions that fail to reflect the true correctness likelihood of classification accuracy. We argue that such an inability to model uncertainty is mainly caused by the closed-world nature in softmax: a model trained by the cross-entropy loss will be forced to classify input into one of K pre-defined categories with high probability. To address this problem, we for the first time propose a novel K+1-way softmax formulation, which incorporates the modeling of open-world uncertainty as the extra dimension. To unify the learning of the original K-way classification task and the extra dimension that models uncertainty, we 1) propose a novel energy-based objective function, and moreover, 2) theoretically prove that optimizing such an objective essentially forces the extra dimension to capture the marginal data distribution. Extensive experiments show that our approach, Energy-based Open-World Softmax (EOW-Softmax), is superior to existing state-of-the-art methods in improving confidence calibration.
Yezhen Wang, Bo Li 0080, Tong Che, Kaiyang Zhou, Ziwei Liu 0002, Dongsheng Li 0002
ICCV1
2019 Perspective-Guided Convolution Networks for Crowd Counting
abstract
In this paper, we propose a novel perspective-guided convolution (PGC) for convolutional neural network (CNN) based crowd counting (i.e. PGCNet), which aims to overcome the dramatic intra-scene scale variations of people due to the perspective effect. While most state-of-the-arts adopt multi-scale or multi-column architectures to address such issue, they generally fail in modeling continuous scale variations since only discrete representative scales are considered. PGCNet, on the other hand, utilizes perspective information to guide the spatially variant smoothing of feature maps before feeding them to the successive convolutions. An effective perspective estimation branch is also introduced to PGCNet, which can be trained in either supervised setting or weakly-supervised setting when the branch has been pre-trained. Our PGCNet is single-column with moderate increase in computation, and extensive experimental results on four benchmark datasets show the improvements of our method against the state-of-the-arts. Additionally, we also introduce Crowd Surveillance, a large scale dataset for crowd counting that contains 13,000+ high-resolution images with challenging scenarios. Code is available at https://github.com/Zhaoyi-Yan/PGCNet.
Zhaoyi Yan, Yuchen Yuan, Wangmeng Zuo, Xiao Tan 0001, Yezhen Wang, Shilei Wen, Errui Ding
ICCV5