Tianyi Liang 0002

dblp:137/6983-2 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0001-8372-8379ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Visual content generation and editing · 56% Multimedia analysis and retrieval · 26% Audio and music processing · 18%
Artificial intelligence
3 papers
Generative modeling · 44% Efficient and distributed learning · 34% 3D vision · 12%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 100%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 9 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › 3d content generation
text-to-3d generation
1.012026
Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation · IEEE Trans. Vis. Comput. Graph. 2026
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation · ICML 2025
Data integration and cleaning
data quality
0.912025
CritiQ: Mining Data Quality Criteria from Human Preferences · ACL (1) 2025
Visual content generation and editing › graphic design
color palette generation
0.912025
Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation Learning · ACM Multimedia 2025
Visual content generation and editing
image editing
0.912025
TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation · ICML 2025
Audio and music processing › music information retrieval
music emotion recognition
0.912025
Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation Learning · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model
cross-attention control
0.312025
TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation · ICML 2025
Computer vision › Vision and language › cross-modal alignment
image-text alignment
0.312025
TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation · ICML 2025
Multimedia analysis and retrieval › multimodal learning
multimodal representation learning
0.312025
Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation Learning · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 3.0multi-view hybrid scoring · 3.0MLLM · 3.0multimodal large language model · 2.0modulation-based fusion · 2.0direct preference optimization · 2.0spatial attention constraint · 1.7preference mining · 1.7force-directed graph · 1.7cross-modal representation learning · 0.9cross-attention maps · 0.9cross-attention map · 0.9
YearPublicationVenuePosition
2026 MPJudge: Towards Perceptual Assessment of Music-Induced Paintings
abstract
Music-induced painting is a unique artistic practice, where visual artworks are created under the influence of music. Evaluating whether a painting faithfully reflects the music that inspired it poses a challenging perceptual assessment task. Existing methods primarily rely on emotion recognition models to assess the similarity between music and painting, but such models introduce considerable noise and overlook broader perceptual cues beyond emotion. To address these limitations, we propose a novel framework for music-induced painting assessment that directly models perceptual coherence between music and visual art. We introduce MPD, the first large-scale dataset of music–painting pairs annotated by domain experts based on perceptual coherence. To better handle ambiguous cases, we further collect pairwise preference annotations. Building on this dataset, we present MPJudge, a model that integrates music features into a visual encoder via a modulation-based fusion mechanism. To effectively learn from ambiguous cases, we adopt Direct Preference Optimization for training. Extensive experiments demonstrate that our method outperforms existing approaches. Qualitative results further show that our model more accurately identifies music-relevant regions in paintings.
Shiqi Jiang 0001, Tianyi Liang 0002, Huayuan Ye, Changbo Wang, Chenhui Li 0001
AAAI2
2026 Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
abstract
Text-to-3D (T23D) generation has transformed digital content creation, yet remains bottlenecked by blind trial-and-error prompting processes that yield unpredictable results. While visual prompt engineering has advanced in text-to-image domains, its application to 3D generation presents unique challenges requiring multi-view consistency evaluation and spatial understanding. We present Sel3DCraft, a visual prompt engineering system for T23D that transforms unstructured exploration into a guided visual process. Our approach introduces three key innovations: a dual-branch structure combining retrieval and generation for diverse candidate exploration; a multi-view hybrid scoring approach that leverages MLLMs with innovative high-level metrics to assess 3D models with human-expert consistency; and a prompt-driven visual analytics suite that enables intuitive defect identification and refinement. Extensive testing and a user study demonstrate that Sel3DCraft surpasses other T23D systems in supporting creativity for designers.
Tianyi Liang 0002, Haiwen Huang, Shiqi Jiang 0001, Yifei Huang 0006, Liangyu Chen 0001, Changbo Wang, Chenhui Li 0001
IEEE Trans. Vis. Comput. Graph.2
2025 CritiQ: Mining Data Quality Criteria from Human Preferences
abstract
Honglin Guo, Kai Lv, Qipeng Guo, Tianyi Liang, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun, Kai Chen, Xipeng Qiu, Tao Gui. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Honglin Guo, Kai Lv 0001, Qipeng Guo, Tianyi Liang 0002, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun 0031, Kai Chen 0026, Xipeng Qiu, Tao Gui
ACL (1)4
2025 TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation
abstract
Text-to-image (T2I) generation has made remarkable progress in producing high-quality images, but a fundamental challenge remains: creating backgrounds that naturally accommodate text placement without compromising image quality. This capability is non-trivial for real-world applications like graphic design, where clear visual hierarchy between content and text is essential. Prior work has primarily focused on arranging layouts within existing static images, leaving unexplored the potential of T2I models for generating text-friendly backgrounds. We present TextCenGen, a training-free approach that actively relocates objects before optimizing text regions, rather than directly reducing cross-attention which degrades image quality. Our method introduces: (1) a force-directed graph approach that detects conflicting objects and guides them relocation using cross-attention maps, and (2) a spatial attention constraint that ensures smooth background generation in text regions. Our method is plug-and-play, requiring no additional training while well balancing both semantic fidelity and visual quality. Evaluated on our proposed text-friendly T2I benchmark of 27,000 images across three seed datasets, TextCenGen outperforms existing methods by achieving 23\% lower saliency overlap in text regions while maintaining 98\% of the original semantic fidelity measured by CLIP score and our proposed Visual-Textual Concordance Metric (VTCM).
Tianyi Liang 0002, Jiangqi Liu, Yifei Huang 0006, Shiqi Jiang 0001, Jianshen Shi, Changbo Wang, Chenhui Li 0001
ICML1
2025 Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation Learning
abstract
Emotion alignment between music and palettes is crucial for effective multimedia content, yet misalignment creates confusion that weakens the intended message. However, existing methods often generate only a single dominant color, missing emotion variation. Others rely on indirect mappings through text or images, resulting in the loss of crucial emotion details. To address these challenges, we present Music2Palette, a novel method for emotion-aligned color palette generation via cross-modal representation learning. We first construct MuCED, a dataset of 2,634 expert-validated music-palette pairs aligned through Russell-based emotion vectors. To directly translate music into palettes, we propose a cross-modal representation learning framework with a music encoder and color decoder. We further propose a multi-objective optimization approach that jointly enhances emotion alignment, color diversity, and palette coherence. Extensive experiments demonstrate that our method outperforms current methods in interpreting music emotion and generating attractive and diverse color palettes. Our approach enables applications like music-driven image recoloring, video generating, and data visualization, bridging the gap between auditory and visual emotion experiences.
Jiayun Hu, Yueyi He, Tianyi Liang 0002, Changbo Wang, Chenhui Li 0001
ACM Multimedia3
2023 Enhancing Visual Understanding by Removing Dithering with Global and Self-Conditioned Transformation
abstract
PNG-8 images are commonly used on the web due to their small size, but their limited color palette often leads to dithering artifacts. Unfortunately, restoring these images using a conventional convolutional neural network (CNN) often results in suboptimal performance since the spatial distribution of dithering is not uniform across the image. This is because the convolutional operator is spatially consistent, meaning it applies the same kernel to all pixels, which we refer to as a global transformation. To address this issue, we propose PNG8IRNet, one approach that combines global and self-conditioned transformations to remove dithering artifacts. Our method incorporates a multilayer perceptron (MLP) to generate diverse kernels for each pixel, taking into account the spatial non-uniformity of dithering, which we define as a self-conditioned transformation. PNG8IRNet demonstrates its performance on multiple datasets, substantially enhancing visual comprehension through a comprehensive set of experiments.
Yifei Huang 0006, Chenhui Li 0001, Risheng Liu, Tianyi Liang 0002, Changbo Wang
VINCI4