VLDB 2026 Research / reviewers in the wild / expert
Sheng-Yu Wang
dblp:30/11438
· DBLP profile ↗
12ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0003-4000-2046ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fast Data Attribution for Text-to-Image ModelsabstractData attribution for text-to-image models aims to identify the training images that most significantly influenced a generated output. Existing attribution methods involve considerable computational resources for each query, making them impractical for real-world applications. We propose a novel approach for scalable and efficient data attribution. Our key idea is to distill a slow, unlearning-based attribution method to a feature embedding space for efficient retrieval of highly influential training images. During deployment, combined with efficient indexing and search methods, our method successfully finds highly influential images without running expensive attribution algorithms.
We show extensive results on both medium-scale models trained on MSCOCO and large-scale Stable Diffusion models trained on LAION, demonstrating that our method can achieve better or competitive performance in a few seconds, faster than existing methods by 2,500x - 400,000x. Our work represents a meaningful step towards the large-scale application of data attribution methods on real-world models such as Stable Diffusion. Sheng-Yu Wang, Aaron Hertzmann, Alexei A. Efros, Richard Zhang 0001, Jun-Yan Zhu |
NeurIPS | 1 |
| 2024 | Data Attribution for Text-to-Image Models by Unlearning Synthesized ImagesabstractThe goal of data attribution for text-to-image models is to identify the training images that most influence the generation of a new image. Influence is defined such that, for a given output, if a model is retrained from scratch without the most influential images, the model would fail to reproduce the same output. Unfortunately, directly searching for these influential images is computationally infeasible, since it would require repeatedly retraining models from scratch. In our work, we propose an efficient data attribution method by simulating unlearning the synthesized image. We achieve this by increasing the training loss on the output image, without catastrophic forgetting of other, unrelated concepts. We then identify training images with significant loss deviations after the unlearning process and label these as influential. We evaluate our method with a computationally intensive but "gold-standard" retraining from scratch and demonstrate our method's advantages over previous methods. Sheng-Yu Wang, Aaron Hertzmann, Alexei A. Efros, Jun-Yan Zhu, Richard Zhang 0001 |
NeurIPS | 1 |
| 2024 | Customizing Text-to-Image Models with a Single Image Pair
Maxwell Jones, Sheng-Yu Wang, Nupur Kumari, David Bau, Jun-Yan Zhu |
SIGGRAPH Asia | 2 |
| 2023 | Ablating Concepts in Text-to-Image Diffusion ModelsabstractLarge-scale text-to-image diffusion models can generate high-fidelity images with powerful compositional ability. However, these models are typically trained on an enormous amount of Internet data, often containing copyrighted material, licensed images, and personal photos. Furthermore, they have been found to replicate the style of various living artists or memorize exact training samples. How can we remove such copyrighted concepts or images without retraining the model from scratch? To achieve this goal, we propose an efficient method of ablating concepts in the pretrained model, i.e., preventing the generation of a target concept. Our algorithm learns to match the image distribution for a target style, instance, or text prompt we wish to ablate to the distribution corresponding to an anchor concept. This prevents the model from generating target concepts given its text condition. Extensive experiments show that our method can successfully prevent the generation of the ablated concept while preserving closely related concepts in the model. Nupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman, Richard Zhang 0001, Jun-Yan Zhu |
ICCV | 3 |
| 2023 | Evaluating Data Attribution for Text-to-Image ModelsabstractWhile large text-to-image models are able to synthesize "novel" images, these images are necessarily a reflection of the training data. The problem of data attribution in such models – which of the images in the training set are most responsible for the appearance of a given generated image – is a difficult yet important one. As an initial step toward this problem, we evaluate attribution through "customization" methods, which tune an existing large-scale model toward a given exemplar object or style. Our key insight is that this allow us to efficiently create synthetic images that are computationally influenced by the exemplar by construction. With our new dataset of such exemplar-influenced images, we are able to evaluate various data attribution algorithms and different possible feature spaces. Furthermore, by training on our dataset, we can tune standard models, such as DINO, CLIP, and ViT, toward the attribution problem. Even though the procedure is tuned towards small exemplar sets, we show generalization to larger sets. Finally, by taking into account the inherent uncertainty of the problem, we can assign soft attribution scores over a set of training images. Sheng-Yu Wang, Alexei A. Efros, Jun-Yan Zhu, Richard Zhang 0001 |
ICCV | 1 |
| 2023 | A Ripple-Based Constant On-Time Controlled DC-DC Buck Converter with Inductor Current Sensing TechniqueabstractThis letter implements a ripple-based constant on-time (RBCOT) buck converter with a fast transient response fabricated in the TSMC$0.18 \mu \mathrm{m}$CMOS process. The steady-state measurement shows that this chip can regulate output voltage from 0.9V to 1.8V while the input voltage is set at 3.3V and the output load current varies from 0.1A to 1A. The load transient response shows that when the output voltage is 0.9V, the undershoot voltage is 78m V and the overshoot voltage is 126m V. The settling time is$2.8\mu \mathrm{s}$for a step-up load and$2.5\mu \mathrm{s}$for a step-down load. The chip area is 0.922 mm2, The maximum efficiency is 92.32%. This work improves the traditional inductor current ramp compensation technique. Here, we utilize an accurate transconductance amplifier to amplify the inductor current, increasing system stability and efficiency. By adopting negative inductor current feedback to improve the transient response. Through the V2controlled dual loop structures to eliminate the output dc voltage offset issues. Sheng-Jen Cheng, Chieh-Ju Tsai, Sheng-Yu Wang, Charlie Chung-Ping Chen |
ISCAS | 3 |
| 2023 | Content-based Search for Deep Generative ModelsabstractThe growing proliferation of customized and pretrained generative models has made it infeasible for a user to be fully cognizant of every model in existence. To address this need, we introduce the task of content-based model search: given a query and a large set of generative models, finding the models that best match the query. As each generative model produces a distribution of images, we formulate the search task as an optimization problem to select the model with the highest probability of generating similar content as the query. We introduce a formulation to approximate this probability given the query from different modalities, e.g., image, sketch, and text. Furthermore, we propose a contrastive learning framework for model retrieval, which learns to adapt features for various query modalities. We demonstrate that our method outperforms several baselines on Generative Model Zoo, a new benchmark we create for the model retrieval task. Daohan Lu, Sheng-Yu Wang, Nupur Kumari, Rohan Agarwal, Mia Tang, David Bau, Jun-Yan Zhu |
SIGGRAPH Asia | 2 |
| 2022 | Rewriting geometric rules of a GANabstractDeep generative models make visual content creation more accessible to novice users by automating the synthesis of diverse, realistic content based on a collected dataset. However, the current machine learning approaches miss a key element of the creative process - the ability to synthesize things that go far beyond the data distribution and everyday experience. To begin to address this issue, we enable a user to "warp" a given model by editing just a handful of original model outputs with desired geometric changes. Our method applies a low-rank update to a single model layer to reconstruct edited examples. Furthermore, to combat overfitting, we propose a latent space augmentation method based on style-mixing. Our method allows a user to create a model that synthesizes endless objects with defined geometric changes, enabling the creation of a new generative model without the burden of curating a large-scale dataset. We also demonstrate that edited models can be composed to achieve aggregated effects, and we present an interactive interface to enable users to create new models through composition. Empirical measurements on multiple test cases suggest the advantage of our method against recent GAN fine-tuning methods. Finally, we showcase several applications using the edited models, including latent space interpolation and image editing. Sheng-Yu Wang, David Bau, Jun-Yan Zhu |
ACM Trans. Graph. | 1 |
| 2021 | Sketch Your Own GANabstractCan a user create a deep generative model by sketching a single example? Traditionally, creating a GAN model has required the collection of a large-scale dataset of exemplars and specialized knowledge in deep learning. In contrast, sketching is possibly the most universally accessible way to convey a visual concept. In this work, we present a method, GAN Sketching, for rewriting GANs with one or more sketches, to make GANs training easier for novice users. In particular, we change the weights of an original GAN model according to user sketches. We encourage the model’s output to match the user sketches through a crossdomain adversarial loss. Furthermore, we explore different regularization methods to preserve the original model’s diversity and image quality. Experiments have shown that our method can mold GANs to match shapes and poses specified by sketches while maintaining realism and diversity. Finally, we demonstrate a few applications of the resulting GAN, including latent space interpolation and image editing. Sheng-Yu Wang, David Bau, Jun-Yan Zhu |
ICCV | 1 |
| 2020 | CNN-Generated Images Are Surprisingly Easy to Spot... for NowabstractIn this work we ask whether it is possible to create a "universal" detector for telling apart real images from these generated by a CNN, regardless of architecture or dataset used. To test this, we collect a dataset consisting of fake images generated by 11 different CNN-based image generator models, chosen to span the space of commonly used architectures today (ProGAN, StyleGAN, BigGAN, CycleGAN, StarGAN, GauGAN, DeepFakes, cascaded refinement networks, implicit maximum likelihood estimation, second-order attention super-resolution, seeing-in-the-dark). We demonstrate that, with careful pre- and post-processing and data augmentation, a standard image classifier trained on only one specific CNN generator (ProGAN) is able to generalize surprisingly well to unseen architectures, datasets, and training methods (including the just released StyleGAN2). Our findings suggest the intriguing possibility that today's CNN-generated images share some common systematic flaws, preventing them from achieving realistic image synthesis. Sheng-Yu Wang, Oliver Wang, Richard Zhang 0001, Andrew Owens, Alexei A. Efros |
CVPR | 1 |
| 2019 | Detecting Photoshopped Faces by Scripting PhotoshopabstractMost malicious photo manipulations are created using standard image editing tools, such as Adobe Photoshop. We present a method for detecting one very popular Photoshop manipulation -- image warping applied to human faces -- using a model trained entirely using fake images that were automatically generated by scripting Photoshop itself. We show that our model outperforms humans at the task of recognizing manipulated images, can predict the specific location of edits, and in some cases can be used to "undo" a manipulation to reconstruct the original, unedited image. We demonstrate that the system can be successfully applied to artist-created image manipulations. Sheng-Yu Wang, Oliver Wang, Richard Zhang 0001, Andrew Owens, Alexei A. Efros |
ICCV | 1 |
| 2011 | Human and car identification using motion vector in H.264 compressed videoabstractThis paper presents a novice method for human and car identification in H.264/AVC compressed video domain. By analyzing the shape and motion vector homogeneity of the segmented objects, we can identify car and human. Our system consists of three main processes: (1) Moving object segmentation based on clustering MVs and Markov Random Field (MRF) iteration, (2) Feature Extraction based on motion analysis to obtain the difference of MVs direction (dMVD) and shape analysis to find the number of MBs (nMB) of an object, and (3) Object classification using Bayesian Classifier. In the experiments, we show that the recognition rate of car and human are 88% and 98% respectively. Quan-Xi Yang, Ke-Wei Lin, Sheng-Yu Wang, Chung-Lin Huang |
VCIP | 4 |