VLDB 2026 Research / reviewers in the wild / expert
Kunpeng Song
dblp:194/1391
· DBLP profile ↗
10ranked-venue papers
4as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
Kunpeng Song, Yizhe Zhu, Ahmed M. Elgammal |
ECCV (40) | 1 |
| 2024 | Multi-Dimensional Data Collection Under Personalized Local Differential PrivacyabstractThis work addresses the problem of collecting multidimensional data while adhering to personalized local differential privacy. In this local context, each user possesses a data record containing multiple attributes. Due to variations in attribute sensitivity, privacy requirements of the attributes differ. Balancing personalized privacy needs of attributes with maximizing the accuracy of statistical results presents a significant challenge. The accuracy of estimation results is closely tied to the allocation of privacy budgets and user grouping. Consequently, we define an optimization problem: determining the allocation of privacy budgets to each attribute within the group and identifying the corresponding number of users in the group to ensure optimal estimation result accuracy while meeting privacy constraints. To simplify the complexity of this optimization problem, we introduce a progressive optimization method. This method initially groups attributes and subsequently fine-tunes and optimizes them. Practically, our approach meets the personalized privacy protection requirements of attributes. Experimental results indicate that our approach achieves significant improvements in data utility. Kunpeng Song, Mingzhang Sun, Kui Zhou, Peng Tang 0002, Ning Wang 0026, Shanqing Guo |
TrustCom | 1 |
| 2024 | ProxEdit: Improving Tuning-Free Real Image Editing with Proximal GuidanceabstractDDIM inversion has revealed the remarkable potential of real image editing within diffusion-based methods. However, the accuracy of DDIM reconstruction degrades as larger classifier-free guidance (CFG) scales being used for enhanced editing. Null-text inversion (NTI) optimizes null embeddings to align the reconstruction and inversion trajectories with larger CFG scales, enabling real image editing with cross-attention control. Negative-prompt inversion (NPI) further offers a training-free closed-form solution of NTI. However, it may introduce artifacts and is still constrained by DDIM reconstruction quality. To overcome these limitations, we propose proximal guidance and incorporate it to NPI with cross-attention control. We enhance NPI with a regularization term and inversion guidance, which reduces artifacts while capitalizing on its training-free nature. Additionally, we extend the concepts to incorporate mutual self-attention control, enabling geometry and layout alterations in the editing process. Our method provides an efficient and straightforward approach, effectively addressing real image editing tasks with minimal computational overhead. Ligong Han, Song Wen 0001, Kunpeng Song, Mengwei Ren, Ruijiang Gao, Anastasis Stathopoulos, Xiaoxiao He, Yuxiao Chen 0002, Di Liu 0003, Qilong Zhangli, Jindong Jiang, Zhaoyang Xia, Akash Srivastava, Dimitris N. Metaxas |
WACV | 5 |
| 2024 | StyleGAN-Fusion: Diffusion Guided Domain Adaptation of Image GeneratorsabstractCan a text-to-image diffusion model be used as a training objective for adapting a GAN generator to another domain? In this paper, we show that the classifier-free guidance can be leveraged as a critic and enable generators to distill knowledge from large-scale text-to-image diffusion models. Generators can be efficiently shifted into new domains indicated by text prompts without access to groundtruth samples from target domains. We demonstrate the effectiveness and controllability of our method through extensive experiments. Although not trained to minimize CLIP loss, our model achieves equally high CLIP scores and significantly lower FID than prior work on short prompts, and outperforms the baseline qualitatively and quantitatively on long and complicated prompts. To our best knowledge, the proposed method is the first attempt at incorporating large-scale pre-trained diffusion models and distillation sampling for text-driven image generator domain adaptation and gives a quality previously beyond possible. Moreover, we extend our work to 3D-aware style-based generators and DreamBooth guidance. For code and more visual samples, please visit our Project Webpage. Kunpeng Song, Ligong Han, Dimitris N. Metaxas, Ahmed M. Elgammal |
WACV | 1 |
| 2023 | Using IRS to Improve the Secrecy Rate of Millimeter Wave Communication SystemabstractWith the development of 6G, millimeter wave communication has received extensive attention. Due to the characteristics of wireless transmission, information secrecy transmission is facing significant challenges. This paper uses the physical layer security (PLS) to explore the information secrecy transmission. Specifically, we use Intelligent Reflecting Surface (IRS) to control the wireless propagation environment and improve the secrecy rate of millimeter wave communication. The active beamforming matrix of the base station and the passive beamforming matrix of the IRS are optimized to achieve the maximum secrecy rate. We deduce the closed form solution of active beamforming and the approximate optimal solution of passive beamforming. An alternating optimization (AO) algorithm is applied to solve the non-convex optimization problem. Simulation results verify the convergence and effectiveness of the algorithm, which can obtain nearly twice the secrecy rate gain of the benchmark algorithm. Kunpeng Song, Fangshu Ma, Zexian Chen, Yong Shang, Yuxin Cheng |
VTC2023-Spring | 1 |
| 2021 | TIME: Text and Image Mutual-Translation Adversarial NetworksabstractFocusing on text-to-image (T2I) generation, we propose Text and Image Mutual-Translation Adversarial Networks (TIME), a lightweight but effective model that jointly learns a T2I generator G and an image captioning discriminator D under the Generative Adversarial Network framework. While previous methods tackle the T2I problem as a uni-directional task and use pre-trained language models to enforce the image--text consistency, TIME requires neither extra modules nor pre-training. We show that the performance of G can be boosted substantially by training it jointly with D as a language model. Specifically, we adopt Transformers to model the cross-modal connections between the image features and word embeddings, and design an annealing conditional hinge loss that dynamically balances the adversarial learning. In our experiments, TIME achieves state-of-the-art (SOTA) performance on the CUB dataset (Inception Score of 4.91 and Fréchet Inception Distance of 14.3 on CUB), and shows promising performance on MS-COCO dataset on image captioning and downstream vision-language tasks. Kunpeng Song, Yizhe Zhu, Gerard de Melo, Ahmed M. Elgammal |
AAAI | 2 |
| 2021 | Self-Supervised Sketch-to-Image SynthesisabstractImagining a colored realistic image from an arbitrary-drawn sketch is one of human capabilities that we eager machines to mimic. Unlike previous methods that either require the sketch-image pairs or utilize low-quantity detected edges as sketches, we study the exemplar-based sketch-to-image (s2i) synthesis task in a self-supervised learning manner, eliminating the necessity of the paired sketch data. To this end, we first propose an unsupervised method to efficiently synthesize line-sketches for general RGB-only datasets. With the synthetic paired-data, we then present a self-supervised Auto-Encoder (AE) to decouple the content/style features from sketches and RGB-images, and synthesize images both content-faithful to the sketches and style-consistent to the RGB-images. While prior works employ either the cycle-consistence loss or dedicated attentional modules to enforce the content/style fidelity, we show AE's superior performance with pure self-supervisions. To further improve the synthesis quality in high resolution, we also leverage an adversarial network to refine the details of synthetic images. Extensive experiments on $1024^2$ resolution demonstrate a new state-of-art-art performance of the proposed model on CelebA-HQ and Wiki-Art datasets. Moreover, with the proposed sketch generator, the model shows a promising performance on style mixing and style transfer, which the synthesized images are not only style-consistent but also semantically meaningful. Yizhe Zhu, Kunpeng Song, Ahmed M. Elgammal |
AAAI | 3 |
| 2021 | Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image Synthesis
Yizhe Zhu, Kunpeng Song, Ahmed M. Elgammal |
ICLR | 3 |
| 2020 | Sketch-to-Art: Synthesizing Stylized Art Images from Sketches
Kunpeng Song, Yizhe Zhu, Ahmed M. Elgammal |
ACCV (6) | 2 |
| 2017 | BestConfig: tapping the performance potential of systems via automatic configuration tuningabstractAn ever increasing number of configuration parameters are provided to system users. But many users have used one configuration setting across different workloads, leaving untapped the performance potential of systems. A good configuration setting can greatly improve the performance of a deployed system under certain workloads. But with tens or hundreds of parameters, it becomes a highly costly task to decide which configuration setting leads to the best performance. While such task requires the strong expertise in both the system and the application, users commonly lack such expertise. Yuqing Zhu 0001, Jianxun Liu 0006, Mengying Guo, Yungang Bao, Wenlong Ma 0001, Zhuoyue Liu, Kunpeng Song, Yingchun Yang |
SoCC | 7 |