VLDB 2026 Research / reviewers in the wild / expert
Yiwei Yang 0007
dblp:233/9195-7
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0001-8531-0574ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreeQR: Free Lunch for Aesthetic QR Codes Emerging From the Latent Space in Diffusion ModelsabstractIn the modern digital age, Quick Response (QR) codes serve as a critical interface for bridging the physical and virtual worlds, widely utilized in multimedia applications. However, traditional binary QR codes often lack the visual appeal desired in contexts. Aesthetic QR codes address this limitation by enabling the customization of QR code patterns to enhance visual attractiveness while retaining compatibility with standard QR decoders. Previous works have explored the use of diffusion models for generating such codes but often require extensive training of ControlNets and face challenges in maintaining scannability. To address these issues, we present FreeQR, a streamlined and effective approach that enables the stable generation of QR code images with diffusion models. Our methodology involves the strategic fusion between the specific channel in the latent space of the denoising process with the noised latent representations of the QR blueprint image at corresponding timesteps. This ensures that the generated images adhere to the brightness distribution required for effective scanning while achieving a balance between aesthetics and functionality. Additionally, we introduce gradient guidance based on scanning errors directly in the latent space, enabling the generation of scannable QR codes in seconds without additional model parameters. Experimental results demonstrate that FreeQR significantly enhances the aesthetics and scannability of QR codes compared to existing methods, making it a lightweight and efficient solution for multimedia applications. Yiwei Yang 0007, Jun Jia, Zheyuan Liu 0011, Zhongpai Gao, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Multim. | 1 |
| 2025 | DiffDeid: High-Quality Face De-identification and Recovery via Diffusion InversionabstractNowadays, personal privacy protection is extremely emphasised. Face de-identification is considered as an effective way to protect the visual privacy through disguising or replacing identity attributes. Existing methods compromise either high fidelity or reversibility. To address these issues, this paper proposes DiffDeid, the first diffusion-based face de-identification and recovery method. Leveraging recent diffusion inversion and control technicques, DiffDeid achieves both high quality imperceptible de-identification and exact recovery with passwords. DiffDeid has three attractions: (1) It can generate de-identified faces with the superior fidelity while maintaining other non-identity attributes for visual tasks. (2) The correct password is powerful enough to restore facial images with extreme details. Meanwhile, incorrect passwords can lead to vastly different decryption results. (3) DiffDeid demands minimal computing resources and instant training time compared to others. We conducted experiments on various face datasets to showcase the superiority of our proposed method. Additional experiments show that DiffDeid is powerful with diverse text prompts and control instructions even beyond human faces. Codes are available at project page. Zheyuan Liu 0011, Jun Jia, Hongyi Miao, Yiwei Yang 0007, Yanwei Jiang, Yingjie Zhou 0003, Zhi Liu 0004, Guangtao Zhai |
ICME | 4 |
| 2025 | CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-stepabstractCurrent text-to-image (T2I) generation models struggle to align spatial composition with the input text, especially in complex scenes.
Even layout-based approaches yield suboptimal spatial control, as their generation process is decoupled from layout planning, making it difficult to refine the layout during synthesis.
We present CoT-Diff, a framework that brings step-by-step CoT-style reasoning into T2I generation by tightly integrating Multimodal Large Language Model (MLLM)-driven 3D layout planning with the diffusion process.
CoT-Diff enables layout-aware reasoning inline within a single diffusion round: at each denoising step, the MLLM evaluates intermediate predictions, dynamically updates the 3D scene layout, and continuously guides the generation process.
The updated layout is converted into semantic conditions and depth maps, which are fused into the diffusion model via a condition-aware attention mechanism, enabling precise spatial control and semantic injection.
Experiments on 3D Scene benchmarks show that CoT-Diff significantly improves spatial alignment and compositional fidelity, and outperforms the state-of-the-art method by 34.7% in complex scene spatial accuracy, thereby validating the effectiveness of this entangled generation paradigm. Zheyuan Liu 0012, Munan Ning, Qihui Zhang, Yiwei Yang 0007, Yibing Song, Fan Wang 0019, Li Yuan 0007 |
NeurIPS | 6 |
| 2024 | DiffStega: Towards Universal Training-Free Coverless Image Steganography with Diffusion Models
Yiwei Yang 0007, Zheyuan Liu 0011, Jun Jia, Zhongpai Gao, Wei Sun 0029, Xiaohong Liu 0001, Guangtao Zhai |
IJCAI | 1 |
| 2024 | Hidden Barcode in Sub-Images with Invisible Locating MarkerabstractThe prevalence of the Internet of Things (IoT) has led to the widespread adoption of 2D barcodes as a means of offline-to-online communication. Whereas, 2D barcodes are not ideal for publicity materials, due to their space-consuming nature. Recent works have proposed 2D image barcodes that contain invisible codes or hyperlinks to transmit hidden information from offline to online. However, these methods undermine the purpose of the codes being invisible, due to the the requirement of markers to locate them. The conference version of this work has presents a novel imperceptible information embedding framework for display or print-camera scenarios, which includes not only hiding and recvoery but also locating and correcting. With the assistance of learned invisible markers, hidden codes can be rendered truly imperceptible. A highly effective multi-stage training scheme is proposed to achieve high visual fidelity and retrieval resiliency, wherein information is concealed in a sub-region rather than the entire image. However, our conference version does not address the optimal sub-region for hiding, which is crucial when dealing with local region concealment problems. In this paper extension, we consider human perceptual characteristics and introduce an optimal hiding region recommendation algorithm that comprehensively incorporates Just Noticeable Difference (JND) and visual saliency factors into consideration. Extensive experiments demonstrate superior visual quality and robustness compared to state-of-the-art methods. With the assistance of our proposed hiding region recommendation algorithm, concealed information becomes even less visible than the results of our conference version without compromising robustness. Jun Jia, Zhongpai Gao, Yiwei Yang 0007, Wei Sun 0029, Dandan Zhu 0001, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Viewing Behavior Supported Visual Saliency Predictor for 360 Degree VideosabstractIn virtual reality (VR), correct and precise estimations of user’s visual fixations and head movements can enhance the quality of experience by allocating more computation resources for analysing and rendering on the areas of interest. However, there is insufficient research about understanding the visual exploration of users when modeling VR visual attention. To bridge the gap between the saliency prediction for traditional 2D content and omnidirectional content, we construct the visual attention dataset and propose the visual saliency prediction framework for panoramic videos. Around the instantaneous viewing behavior, we propose a traditional method to adapt 2D saliency models and design a CNN-based model to better predict visual saliency. In the proposed traditional model, mechanism of visual attention and viewing behaviors are considered in the computation of edge weights on graphs which are interpreted as Markov chains. The fraction of the visual attention that is diverted to each high-clarity vision (HCV) area is estimated through equilibrium distribution of this chain. We also propose the Graph-Based CNN model. The RGB channel and optical flow form the spatial-temporal units of HCVs, from which node feature vectors are extracted. Graph convolution is used to learn the mutual information between node feature vectors of HCVs and retain geometric information. Then feature vectors are aligned according to geometry structure of equirectangular format, and the feature decoder maps the aligned feature maps to the data distribution. We also construct the dynamic omnidirectional monocular (DOM) saliency dataset with 64 diverse videos evaluated by 28 people. The subjective results show that the instantaneous viewing behavior is important in the VR experience. Extensive experiments are conducted on the dataset and the results demonstrate the effectiveness of the proposed framework. The dataset will be released to facilitate the future studies related to visual saliency prediction for 360-degree contents. Yucheng Zhu, Guangtao Zhai, Yiwei Yang 0007, Huiyu Duan, Xiongkuo Min, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | SalGFCN: Graph Based Fully Convolutional Network for Panoramic Saliency PredictionabstractThe saliency prediction of panoramic images is dramatically affected by the distortion caused by non-Euclidean geometry characteristic. Traditional CNN based saliency pre-diction algorithms for 2D images are no longer suitable for 360-degree images. Intuitively, we propose a graph based fully convolutional network for saliency prediction of 360-degree images, which can reasonably map panoramic pixels to spherical graph data structures for representation. The saliency prediction network is based on residual U-Net architecture, with dilated graph convolutions and attention mechanism in the bottleneck. Furthermore, we design a fully convolutional layer for graph pooling and unpooling operations in spherical graph space to retain node-to-node features. Experimental results show that our proposed method outperforms other state-of-the-art saliency models on the large-scale dataset. Yiwei Yang 0007, Yucheng Zhu, Zhongpai Gao, Guangtao Zhai |
VCIP | 1 |