Ruixiang Jiang

dblp:35/1259 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0001-8666-6767ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021Computer networks · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
abstract
The rapid technical progress of generative art (GenArt) has democratized the creation of visually appealing imagery. However, achieving genuine artistic impact - the kind that resonates with viewers on a deeper, more meaningful level - remains formidable as it requires a sophisticated aesthetic sensibility. This sensibility involves a multifaceted cognitive process extending beyond mere visual appeal, which is often overlooked by current computational methods. This paper pioneers an approach to capture this complex process by investigating how the reasoning capabilities of Multimodal LLMs (MLLMs) can be effectively elicited to perform aesthetic judgment. Our analysis reveals a critical challenge: MLLMs exhibit a tendency towards hallucinations during aesthetic reasoning, characterized by subjective opinions and unsubstantiated artistic interpretations. We further demonstrate that these hallucinations can be suppressed by employing an evidence-based and objective reasoning process, as substantiated by our proposed baseline, ArtCoT. MLLMs prompted by this principle produce multifaceted, in-depth aesthetic reasoning that aligns significantly better with human judgment. These findings have direct applications in areas such as AI art tutoring and as reward models for image generation. Ultimately, we hope this work paves the way for AI systems that can truly understand, appreciate, and contribute to art that aligns with human aesthetic values. Project homepage: https://github.com/songrise/MLLM4Art.
Ruixiang Jiang, Chang Wen Chen
ACM Multimedia1
2025 DiffArtist: Towards Structure and Appearance Controllable Image Stylization
Ruixiang Jiang, Chang Wen Chen
ACM Multimedia1
2024 NeRF-Art: Text-Driven Neural Radiance Fields Stylization
abstract
As a powerful representation of 3D scenes, the neural radiance field (NeRF) enables high-quality novel view synthesis from multi-view images. Stylizing NeRF, however, remains challenging, especially in simulating a text-guided style with both the appearance and the geometry altered simultaneously. In this paper, we present NeRF-Art, a text-guided NeRF stylization approach that manipulates the style of a pre-trained NeRF model with a simple text prompt. Unlike previous approaches that either lack sufficient geometry deformations and texture details or require meshes to guide the stylization, our method can shift a 3D scene to the target style characterized by desired geometry and appearance variations without any mesh guidance. This is achieved by introducing a novel global-local contrastive learning strategy, combined with the directional constraint to simultaneously control both the trajectory and the strength of the target style. Moreover, we adopt a weight regularization method to effectively suppress cloudy artifacts and geometry noises which arise easily when the density field is transformed during geometry stylization. Through extensive experiments on various styles, we demonstrate that our method is effective and robust regarding both single-view stylization quality and cross-view consistency.
Can Wang 0007, Ruixiang Jiang, Menglei Chai, Mingming He, Dongdong Chen 0001, Jing Liao 0001
IEEE Trans. Vis. Comput. Graph.2
2023 AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose Control
abstract
Neural implicit fields are powerful for representing 3D scenes and generating high-quality novel views, but it remains challenging to use such implicit representations for creating a 3D human avatar with a specific identity and artistic style that can be easily animated. Our proposed method, AvatarCraft, addresses this challenge by using diffusion models to guide the learning of geometry and texture for a neural avatar based on a single text prompt. We carefully design the optimization framework of neural implicit fields, including a coarse-to-fine multi-bounding box training strategy, shape regularization, and diffusion-based constraints, to produce high-quality geometry and texture. Additionally, we make the human avatar animatable by deforming the neural implicit field with an explicit warping field that maps the target human mesh to a template human mesh, both represented using parametric human models. This simplifies animation and reshaping of the generated avatar by controlling pose and shape parameters. Extensive experiments on various text descriptions show that AvatarCraft is effective and robust in creating human avatars and rendering novel views, poses, and shapes. Our project page is: https://avatar-craft.github.io/.
Ruixiang Jiang, Can Wang 0007, Jingbo Zhang 0002, Menglei Chai, Mingming He, Dongdong Chen 0001, Jing Liao 0001
ICCV1
2023 CLIP-Count: Towards Text-Guided Zero-Shot Object Counting
abstract
Recent advances in visual-language models have shown remarkable zero-shot text-image matching ability that is transferable to downstream tasks such as object detection and segmentation. Adapting these models for object counting, however, remains a formidable challenge. In this study, we first investigate transferring vision-language models (VLMs) for class-agnostic object counting. Specifically, we propose CLIP-Count, the first end-to-end pipeline that estimates density maps for open-vocabulary objects with text guidance in a zero-shot manner. To align the text embedding with dense visual features, we introduce a patch-text contrastive loss that guides the model to learn informative patch-level visual representations for dense prediction. Moreover, we design a hierarchical patch-text interaction module to propagate semantic information across different resolution levels of visual features. Benefiting from the full exploitation of the rich image-text alignment knowledge of pretrained VLMs, our method effectively generates high-quality density maps for objects-of-interest. Extensive experiments on FSC-147, CARPK, and ShanghaiTech crowd counting datasets demonstrate state-of-the-art accuracy and generalizability of the proposed method. Code is available: https://github.com/songrise/CLIP-Count. https://github.com/songrise/CLIP-Count.
Ruixiang Jiang, Lingbo Liu, Chang Wen Chen
ACM Multimedia1
2007 Design and performance evaluation of a multi-agent-based dynamic lifetime security scheme for AODV routing protocol
Hongsong Chen, Zhenzhou Ji, Mingzeng Hu, Zhongchuan Fu, Ruixiang Jiang
J. Netw. Comput. Appl.5
2005 Fusion of censored decisions in wireless sensor networks
abstract
Sensor censoring has been introduced for reduced communication rate in a decentralized detection system where decisions made at peripheral nodes need to be communicated to a fusion center. In this letter, the fusion of decisions from censoring sensors transmitted over wireless fading channels is investigated. The knowledge of fading channels, either in the form of instantaneous channel envelopes or the fading statistics, is integrated in the optimum and suboptimum fusion rule design. The sensor censoring and the ensuing fusion rule design have two major advantages compared with the previous work. 1) Communication overhead is dramatically reduced. 2) It allows incoherent detection, hence, the phase information of transmission channels is no longer required. As such, it is particularly suitable for wireless sensor network applications with severe resource constraints.
Ruixiang Jiang, Biao Chen 0001
IEEE Trans. Wirel. Commun.1
2004 Decision fusion with censored sensors
abstract
Motivated by the sensor censoring idea, we consider a canonical decentralized detection problem with each sensor employing an on/off signaling scheme. Novel to the current work is the integration of fading transmission channels in the fusion algorithm design. The on/off signaling, in addition to its low communication overhead which is crucial to bandwidth limited applications, also enables the decision fusion to be carried out without the knowledge of channel phase. Resorting to incoherent detection schemes, we develop optimal fusion rules for the following two scenarios: 1) when the fading channel envelopes are available at the fusion center; 2) when only the fading statistic is available. Under the low signal-to-noise ratio regime, we further reduce the optimal fusion rule into simple nonlinearities that are both easy to implement and are not subject to prior knowledge constraints.
Ruixiang Jiang
ICASSP (2)1