Shiqi Jiang 0001

dblp:07/10820-1 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-8175-874XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MPJudge: Towards Perceptual Assessment of Music-Induced Paintings
abstract
Music-induced painting is a unique artistic practice, where visual artworks are created under the influence of music. Evaluating whether a painting faithfully reflects the music that inspired it poses a challenging perceptual assessment task. Existing methods primarily rely on emotion recognition models to assess the similarity between music and painting, but such models introduce considerable noise and overlook broader perceptual cues beyond emotion. To address these limitations, we propose a novel framework for music-induced painting assessment that directly models perceptual coherence between music and visual art. We introduce MPD, the first large-scale dataset of music–painting pairs annotated by domain experts based on perceptual coherence. To better handle ambiguous cases, we further collect pairwise preference annotations. Building on this dataset, we present MPJudge, a model that integrates music features into a visual encoder via a modulation-based fusion mechanism. To effectively learn from ambiguous cases, we adopt Direct Preference Optimization for training. Extensive experiments demonstrate that our method outperforms existing approaches. Qualitative results further show that our model more accurately identifies music-relevant regions in paintings.
Shiqi Jiang 0001, Tianyi Liang 0002, Huayuan Ye, Changbo Wang, Chenhui Li 0001
AAAI1
2026 NewsVis: GenAI-Based Visual Storytelling for Corporate Financial News
abstract
Corporate financial news is pivotal for market decisions, but often overwhelms general audiences. While data videos effectively bridge this comprehension gap, their production remains a bottleneck for journalists. We present NewsVis, an authoring tool powered by Generative Artificial Intelligence (GenAI) that automates the transformation of unstructured narratives and raw financial datasets into professional data videos. Unlike generic models, our pipeline ensures factual accuracy through a domain-specific taxonomy of financial attributes and optimizes visual information presentation via a multimodal layout algorithm. Additionally, a human-in-the-loop interface empowers journalists to audit and calibrate generative outputs. Comprehensive quantitative and qualitative evaluations demonstrate that NewsVis significantly reduces production barriers while enhancing information accessibility for viewers.
Jia Bu, Mingwei Jiang, Tong Lyu, Lumeng Wu, Shiqi Jiang 0001, Boyuan Huangfu, Changbo Wang, Chenhui Li 0001
IEEE Trans. Vis. Comput. Graph.6
2026 Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
abstract
Text-to-3D (T23D) generation has transformed digital content creation, yet remains bottlenecked by blind trial-and-error prompting processes that yield unpredictable results. While visual prompt engineering has advanced in text-to-image domains, its application to 3D generation presents unique challenges requiring multi-view consistency evaluation and spatial understanding. We present Sel3DCraft, a visual prompt engineering system for T23D that transforms unstructured exploration into a guided visual process. Our approach introduces three key innovations: a dual-branch structure combining retrieval and generation for diverse candidate exploration; a multi-view hybrid scoring approach that leverages MLLMs with innovative high-level metrics to assess 3D models with human-expert consistency; and a prompt-driven visual analytics suite that enables intuitive defect identification and refinement. Extensive testing and a user study demonstrate that Sel3DCraft surpasses other T23D systems in supporting creativity for designers.
Tianyi Liang 0002, Haiwen Huang, Shiqi Jiang 0001, Yifei Huang 0006, Liangyu Chen 0001, Changbo Wang, Chenhui Li 0001
IEEE Trans. Vis. Comput. Graph.4
2025 Robust Message Embedding via Attention Flow-Based Steganography
abstract
Image steganography can hide information in a host image and obtain a stego image that is perceptually indistinguishable from the original one. This technique has tremendous potential in scenarios like copyright protection and information retrospection. Some previous studies have proposed to enhance the robustness of the methods against image disturbances to increase their applicability. However, they generally cannot achieve a satisfying balance between the steganography quality and robustness. Instead of image-in-image steganography, we focus on the issue of message-in-image embedding that is robust to various real- world image distortions. This task aims to embed information into a natural image and the decoding result is required to be completely accurate, which increases the difficulty of data concealing and revealing. Inspired by the recent developments in transformer-based vision models, we discover that the tokenized representation of image is naturally suitable for steganography task. In this paper, we propose a novel message embedding framework, called Robust Message Steganography (RMSteg), which is competent to hide message via QR Code in a host image based on an normalizing flow-based model. The stego image derived by our method has imperceptible changes and the encoded message can be accurately restored even if the image is printed out and photographed. To our best knowledge, this is the first work that integrates the advantages of transformer models into normalizing flow. The code is available at https://github.com/huayuan4396/RMSteg.
Huayuan Ye, Shenzhuo Zhang, Shiqi Jiang 0001, Jing Liao 0001, Shuhang Gu, Dejun Zheng, Changbo Wang, Chenhui Li 0001
CVPR3
2025 TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation
abstract
Text-to-image (T2I) generation has made remarkable progress in producing high-quality images, but a fundamental challenge remains: creating backgrounds that naturally accommodate text placement without compromising image quality. This capability is non-trivial for real-world applications like graphic design, where clear visual hierarchy between content and text is essential. Prior work has primarily focused on arranging layouts within existing static images, leaving unexplored the potential of T2I models for generating text-friendly backgrounds. We present TextCenGen, a training-free approach that actively relocates objects before optimizing text regions, rather than directly reducing cross-attention which degrades image quality. Our method introduces: (1) a force-directed graph approach that detects conflicting objects and guides them relocation using cross-attention maps, and (2) a spatial attention constraint that ensures smooth background generation in text regions. Our method is plug-and-play, requiring no additional training while well balancing both semantic fidelity and visual quality. Evaluated on our proposed text-friendly T2I benchmark of 27,000 images across three seed datasets, TextCenGen outperforms existing methods by achieving 23\% lower saliency overlap in text regions while maintaining 98\% of the original semantic fidelity measured by CLIP score and our proposed Visual-Textual Concordance Metric (VTCM).
Tianyi Liang 0002, Jiangqi Liu, Yifei Huang 0006, Shiqi Jiang 0001, Jianshen Shi, Changbo Wang, Chenhui Li 0001
ICML4
2025 PPJudge: Towards Human-Aligned Assessment of Artistic Painting Process
abstract
Artistic image assessment has become a prominent research area in computer vision. In recent years, the field has witnessed a proliferation of datasets and methods designed to evaluate the aesthetic quality of paintings. However, most existing approaches focus solely on static final images, overlooking the dynamic and multi-stage nature of the artistic painting process. To address this gap, we propose a novel framework for human-aligned assessment of painting processes. Specifically, we introduce the Painting Process Assessment Dataset (PPAD)-the first large-scale dataset comprising real and synthetic painting process images, annotated by domain experts across eight detailed attributes. Furthermore, we present PPJudge (Painting Process Judge), a Transformer-based model enhanced with temporally-aware positional encoding and a heterogeneous mixture-of-experts architecture, enabling effective assessment of the painting process. Experimental results demonstrate that our method outperforms existing baselines in accuracy, robustness, and alignment with human judgment, offering new insights into computational creativity and art education.
Shiqi Jiang 0001, Xinpeng Li 0002, Xi Mao, Changbo Wang, Chenhui Li 0001
ACM Multimedia1
2025 Prompt2Color: A prompt-based framework for image-derived color generation and visualization optimization
Jiayun Hu, Shiqi Jiang 0001, Haiwen Huang, Changbo Wang, Chenhui Li 0001
Comput. Graph.2
2024 AACP: Aesthetics Assessment of Children's Paintings Based on Self-Supervised Learning
abstract
The Aesthetics Assessment of Children's Paintings (AACP) is an important branch of the image aesthetics assessment (IAA), playing a significant role in children's education. This task presents unique challenges, such as limited available data and the requirement for evaluation metrics from multiple perspectives. However, previous approaches have relied on training large datasets and subsequently providing an aesthetics score to the image, which is not applicable to AACP. To solve this problem, we construct an aesthetics assessment dataset of children's paintings and a model based on self-supervised learning. 1) We build a novel dataset composed of two parts: the first part contains more than 20k unlabeled images of children's paintings; the second part contains 1.2k images of children's paintings, and each image contains eight attributes labeled by multiple design experts. 2) We design a pipeline that includes a feature extraction module, perception modules and a disentangled evaluation module. 3) We conduct both qualitative and quantitative experiments to compare our model's performance with five other methods using the AACP dataset. Our experiments reveal that our method can accurately capture aesthetic features and achieve state-of-the-art performance.
Shiqi Jiang 0001, Changbo Wang, Chenhui Li 0001
AAAI1
2024 DoodleTunes: Interactive Visual Analysis of Music-Inspired Children Doodles with Automated Feature Annotation
abstract
Music and visual arts are essential in children’s arts education, and their integration has garnered significant attention. Existing data analysis methods for exploring audio-visual correlations are limited. Yet, relevant research is necessary for innovating and promoting arts integration courses. In our work, we collected substantial volumes of music-inspired doodles created by children and interviewed education experts to comprehend the challenges they encountered in the relevant analysis. Based on the insights we obtained, we designed and constructed an interactive visualization system DoodleTunes. DoodleTunes integrates deep learning-driven methods for automatically annotating several types of data features. The visual designs of the system are based on a four-level analysis structure to construct a progressive workflow, facilitating data exploration and insight discovery between doodle images and corresponding music pieces. We evaluated the accuracy of our feature prediction results and collected usage feedback on DoodleTunes from five domain experts.
Jia Bu, Huayuan Ye, Juntong Chen, Shiqi Jiang 0001, Mingtian Tao, Changbo Wang, Chenhui Li 0001
CHI5
2023 Estimating Market Value of Companies Based on Finance Statement through Data Fusion
abstract
The evaluation of a company's value can serve as a guide for investors to assess the company and make informed investment decisions. However, conventional valuation techniques are not applicable to Initial Public Offering (IPO) companies in China, mainly due to the absence of historical market performance. In contrast, a company's finance statement provides a periodic overview of the company's operational and production activities, which is linked to its market performance. Traditional methods often rely on the selection of a limited number of financial indicators from the finance statement and the application of regression analysis. These approaches fail to fully exploit the comprehensive data available in the finance statement. This study proposes a comprehensive method that leverages all relevant information contained in the finance statement, including industry interconnections, financial indices, and additional insights obtained from the report. The structured data is analyzed through tree models, while the interrelationships between different companies are modeled through graph neural networks. Our approach offers a multi-perspective evaluation of IPO companies. The results of our experiments demonstrate that our method can effectively utilize the valuable information in finance statements and improve outcomes.
Shiqi Jiang 0001, Yaxuan Zheng, Wenli Xiong, Yanpeng Hu, Changbo Wang, Chenhui Li 0001
IJCNN2
2021 CoPaint: Guiding Sketch Painting with Consistent Color and Coherent Generative Adversarial Networks
Shiqi Jiang 0001, Chenhui Li 0001, Changbo Wang
CGI1