VLDB 2026 Research / reviewers in the wild / expert
Yitong Yang
dblp:199/7545
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FantasyStyle: Controllable Stylized Distillation for 3D Gaussian SplattingabstractThe success of 3DGS in generative and editing applications has sparked growing interest in 3DGS-based style transfer. However, current methods still face two major challenges: (1) multi-view inconsistency often leads to style conflicts, resulting in appearance smoothing and distortion; and (2) heavy reliance on VGG features, which struggle to disentangle style and content from style images, often causing content leakage and excessive stylization. To tackle these issues, we introduce FantasyStyle, a 3DGS-based style transfer framework, and the first to rely entirely on diffusion model distillation. It comprises two key components: (1) Multi-View Frequency Consistency. We enhance cross-view consistency by applying a 3D filter to multi-view noisy latent, selectively reducing low-frequency components to mitigate stylized prior conflicts. (2) Controllable Stylized Distillation. To suppress content leakage from style images, we introduce negative guidance to exclude undesired content. In addition, we identify the limitations of Score Distillation Sampling and Delta Denoising Score in 3D style transfer and remove the reconstruction term accordingly. Building on these insights, we propose a controllable stylized distillation that leverages negative guidance to more effectively optimize the 3D Gaussians. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art approaches, achieving higher stylization quality and visual realism across various scenes and styles. Yitong Yang, Changshuo Wang 0001, Huajie Wang, Shuting He |
AAAI | 1 |
| 2026 | LCFusion: Infrared and visible image fusion network based on local contour enhancementabstractInfrared and visible light image fusion aims to generate integrated representations that synergistically preserve salient thermal targets in the infrared modality and high-resolution textural details in the visible light modality. However, existing methods face two core challenges: First, high-frequency noise in visible images, such as sensor noise and nonuniform illumination artifacts, is often highly coupled with effective textures. Traditional fusion paradigms readily amplify noise interference while enhancing details, leading to structural distortion and visual graininess in fusion results. Second, mainstream approaches predominantly rely on simple aggregation operations like feature stitching or linear weighting, lacking deep modeling of cross-modal semantic correlations. This prevents adaptive interaction and collaborative enhancement of complementary information between modalities, creating a significant trade-off between target saliency and detail preservation. To address these challenges, we propose a dual-branch fusion network based on local contour enhancement. Specifically, it distinguishes and enhances meaningful contour details in a learnable manner while suppressing meaningless noise, thereby purifying the detail information used for fusion at its source. Cross-attention weights are computed based on feature representations extracted from different modal branches, enabling a feature selection mechanism that facilitates dynamic cross-modal interaction between infrared and visible light information. We evaluate our method against 11 state-of-the-art deep learning-based fusion approaches across four benchmark datasets using both subjective assessments and objective metrics. The experimental results demonstrate superior performance on public datasets. Furthermore, YOLOv12-based detection tests reveal that our method achieves higher confidence scores and better overall detection performance compared to other fusion techniques. Yitong Yang, Xinyang Yao |
Image Vis. Comput. | 1 |
| 2025 | Prompt-Softbox-Prompt: A Free-Text Embedding Control for Image EditingabstractWhile text-driven diffusion models demonstrate remarkable performance in image editing, the critical components of their text embeddings remain underexplored. The ambiguity and entanglement of these embeddings pose challenges for precise editing. In this paper, we provide a comprehensive analysis of text embeddings in Stable Diffusion XL, offering three key insights: (1) aug embedding ~. aug embedding is obtained by combining the pooled output of the final text encoder with the timestep embeddings. https://github.com/huggingface/diffusers retains complete textual semantics but contributes minimally to image generation as it is only fused via the ResBlocks. More text information weakens its local semantics while preserving most global semantics. (2) BOS and padding embedding do not contain any semantic information. (3) EOS holds the semantic information of all words and stylistic information. Each word embedding is important and does not interfere with the semantic injection of other embeddings. Based on these insights, we propose PSP (Prompt-Softbox-Prompt), a training-free image editing method that leverages free-text embedding. PSP enables precise image editing by modifying text embeddings within the cross-attention layers and using Softbox to control the specific area for semantic injection. This technique enables the addition and replacement of objects without affecting other areas of the image. Additionally, PSP can achieve style transfer by simply replacing text embeddings. Extensive experiments show that PSP performs remarkably well in tasks such as object replacement, object addition, and style transfer. Our code is available at https://github.com/yangyt46/PSP. Yitong Yang, Jing Wang 0128, Shuting He |
ACM Multimedia | 1 |
| 2024 | Single image deraining using scale constraint iterative update network
Yitong Yang, Yongjun Zhang 0007, Zhongwei Cui, Haoliang Zhao, Ting Ouyang |
Expert Syst. Appl. | 1 |
| 2024 | A multi-color and multistage collaborative network guided by refined transmission prior for underwater image enhancement
Ting Ouyang, Yongjun Zhang 0007, Haoliang Zhao, Zhongwei Cui, Yitong Yang |
Vis. Comput. | 5 |
| 2023 | High-Frequency Stereo Matching NetworkabstractIn the field of binocular stereo matching, remarkable progress has been made by iterative methods like RAFT-Stereo and CREStereo. However, most of these methods lose information during the iterative process, making it difficult to generate more detailed difference maps that take full advantage of high-frequency information. We propose the Decouple module to alleviate the problem of data coupling and allow features containing subtle details to transfer across the iterations which proves to alleviate the problem significantly in the ablations. To further capture high-frequency details, we propose a Normalization Refinement module that unifies the disparities as a proportion of the disparities over the width of the image, which address the problem of module failure in cross-domain scenarios. Further, with the above improvements, the ResNet-like feature extractor that has not been changed for years becomes a bottleneck. Towards this end, we proposed a multi-scale and multi-stage feature extractor that introduces the channel-wise self-attention mechanism which greatly addresses this bottleneck. Our method (DLNR) ranks 1st on the Middlebury leaderboard, significantly outperforming the next best method by 13.04%. Our method also achieves SOTA performance on the KITTI-2015 benchmark for D1-fg. Code and demos are available at: https://github.com/David-Zhao-1997/High-frequency-Stereo-Matching-Network. Haoliang Zhao, Huizhou Zhou, Yongjun Zhang 0007, Yitong Yang, Yong Zhao 0001 |
CVPR | 5 |
| 2023 | DGRN: Image super-resolution with dual gradient regression guidance
Heliang Yang, Yongjun Zhang 0007, Zhongwei Cui, Yitong Yang |
Comput. Graph. | 5 |
| 2022 | EAI-Stereo: Error Aware Iterative Network for Stereo Matching
Haoliang Zhao, Huizhou Zhou, Yongjun Zhang 0007, Yong Zhao 0010, Yitong Yang, Ting Ouyang |
ACCV (1) | 5 |
| 2022 | Multi-scale dehazing network via high-frequency feature fusion
Yongjun Zhang 0007, Zhi Li 0012, Zhongwei Cui, Yitong Yang |
Comput. Graph. | 5 |
| 2022 | Single image deraining using multi-stage and multi-scale joint channel coordinate attention fusion networkabstractRain streaks can seriously degrade the visual quality of an image and are detrimental to subsequent algorithms such as object detection and semantic segmentation. Therefore, removing rain streaks is a very important task. The deraining task has two main limitations: the first is to encode information about rain streaks in different densities and directions, the second is to keep the background details of the image while removing the rain streak. To address these limitations, we propose an effective algorithm, called multi-stage and multi-scale joint channel coordinate attention fusion network (MMAFN). We mainly propose a two-stage network structure, both of which use an encoder-decoder network to extract features. The first-stage network extracts coarse features and the second-stage network integrates the features of the former to further refine features. We design the joint channel coordinate attention block to encode features of rain streaks in different directions and densities. In addition, to better fuse features of different scales and enhance the generalization performance of the network, the inception attention branch block and the multi-level feature fusion block are designed. Extensive experiments substantiate the superiority of the proposed network and prove that our method outperforms the recent state-of-the-art method. The average PSNR of the five test sets is improved by 0.2dB. On the Test100 test set, the PSNR is increased by 0.93dB at most. Yitong Yang, Yongjun Zhang 0007, Zhongwei Cui, Zhi Li 0012, Haoliang Zhao, Yangtin Ou, Heliang Yang, Xihe Wang |
Int. J. Intell. Syst. | 1 |