Sung-Lin Tsai

dblp:34/895 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0002-1859-2406ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 91% Representation and self-supervised learning · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model › guided diffusion
training-free guidance
0.912025
Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation · ACM Multimedia 2025
Machine learning › Representation and self-supervised learning
text embedding
0.312025
Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

text embedding manipulation · 0.9large language model · 0.9CIELab color space · 0.9
YearPublicationVenuePosition
2026 FreeCond: Free Lunch in the Input Conditions of Text-Guided Inpainting
abstract
Text-to-image inpainting models often exhibit an unpredictable balance among image coherence and prompt adherence. This rigidity limits their adaptability across diverse scenarios, including coarse masks, non-object, and interaction prompts. Recognizing this instability as an indicator of learned generation diversity, we aim to control model behavior for given objective. We propose Empirical Feature Intervention (EFI), a metric-agnostic framework that precomputes how feature interventions influence evaluation metrics—such as CLIP, Human Preference Score (HPS), and Image Reward (IR). Building on EFI, we introduce FreeCond, a free-of-cost framework that applies two simple input interventions (Image Frequency and Mask Value Modulation), these interventions can be further optimized via Surrogate Intervention Optimization (SIO) based on a surrogate model regressed with precomputed EFI data. FreeCond enables real-time, user-interactive control of pretrained models without retraining or architectural modifications. Also, to benchmark performance on challenging settings, we present FCIBench. Experiments on EditBench, BrushBench, and FCIBench demonstrate that FreeCond substantially improves CLIP, HPS, and IR metrics by up to 22%, 8%, and 54%, respectively.
Teng-Fang Hsiao, Bo-Kai Ruan, Sung-Lin Tsai, Yi-Lun Wu, Hong-Han Shuai
WACV3
2025 Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion Generation
abstract
Accurate color alignment in text-to-image (T2I) generation is critical for applications such as fashion, product visualization, and interior design, yet current diffusion models struggle with nuanced and compound color terms (e.g., Tiffany blue, baby pink), often producing images that are misaligned with human intent. Existing approaches rely on cross-attention manipulation, reference images, or fine-tuning but fail to systematically resolve ambiguous color descriptions. To precisely render colors under prompt ambiguity, we propose a training-free framework that enhances color fidelity by leveraging a large language model (LLM) to disambiguate color-related prompts and guiding color blending operations directly in the text embedding space. Our method first employs a large language model (LLM) to resolve ambiguous color terms in the text prompt, and then refines the text embeddings based on the spatial relationships of the resulting color terms in the CIELab color space. Unlike prior methods, our approach improves color accuracy without requiring additional training or external reference images. Experimental results demonstrate that our framework improves color alignment without compromising image quality, bridging the gap between text semantics and visual generation. All supplementary materials are available at https://Sung-Lin.github.io/TintBench/.
Sung-Lin Tsai, Bo-Lun Huang, Yu-Ting Shen, Cheng-Yu Yeo, Chiang Tseng, Bo-Kai Ruan, Wen-Sheng Lien, Hong-Han Shuai
ACM Multimedia1
2008 Automatic image authentication and recovery using fractal code embedding and image inpainting
Shuenn-Shyang Wang, Sung-Lin Tsai
Pattern Recognit.2