Jia Chen 0012

dblp:99/6879-12 · DBLP profile ↗
← Back
21ranked-venue papers
13as first author
19since 2021 · last 2025
0000-0003-2421-6968ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 10 first-author · 16 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FusionCraft: A New Paradigm for Fine-Grained Multimodal Fashion Design Generation
Jia Chen 0012, Haohan Gui, Xinrong Hu, Jinxing Liang, Shuchen Ju, Zhaoyong Li
CGI (3)1
2025 D2-Diff: Controllable Fashion Image Generation with Disentangled Style and Content
Ruhan He, Jia Chen 0012, Xinrong Hu
CGI (3)5
2025 STCGen: Sketch-based Text-to-Clothing Image Generation with Contour and Style Consistency
abstract
In modern fashion design field, it is a mainstream practice to generate clothing images by combining sketch and text. However, the image quality generated by existing multimodal methods combining sketches and text descriptions is suboptimal, as the clothing in the generated image often lacks contour accuracy and stylistic coherence. In this paper, we present STCGen, an advanced multimodal framework that uses both sketches and text to generate clothing images with improved contours and more consistent style. First, we introduce the sketch prior embedding module, which processes sketches to extract key structural features and ensure the consistency of contours, thereby enhancing image details. Second, we propose a cross space attention mechanism to address the issue of text information loss and ensure stylistic consistency, thereby enhancing overall image coherence. Finally, we propose a network simplification scheme to reduce complexity without compromising the quality of resulting images. Experimental results demonstrate that our method excels in generating high-fidelity clothing images.
Chunxia Xiao, Ruhan He, Jia Chen 0012, Mingfu Xiong, Tao Peng 0006, Xinrong Hu
MMAsia6
2025 Physics-Aware Lighting Gaussian-Embedded-Mesh Avatars from Monocular Video
Zhihong Peng, Xinrong Hu, Saishang Zhong, Jinxing Liang, Li Li 0094, Jia Chen 0012
PRCV (10)6
2025 CasRPN: Cascade region proposal network for visual tracking
abstract
Trackers based on the region proposal network (RPN) have garnered extensive attention within object tracking. However, current RPN trackers utilizing the ResNet as the feature extraction network only employ its local convolution operation to extract the image features, thereby constraining the model’s comprehension of the global contextual information. To address this issue, this paper employs a Vision Transformer (ViT) to construct global relations across the entire sequence features, enhancing the model’s capability to represent global features. Additionally, traditional single-stage RPN trackers generate candidate boxes through a coarse regression process. Thus, the candidate boxes may only partially cover the object or include excessive background information, diminishing the model’s accuracy in object localization. Consequently, we use a multi-stage RPN to adjust the anchor boxes in a cascade RPN manner. Moreover, traditional multi-stage RPN exists the misalignment problem between the anchor boxes and image features. To further optimize the multi-stage RPN, this paper employs adaptive convolution to align features with anchor boxes. Our tracking method achieves state-of-the-art results on five common tracking benchmark datasets. Specifically, the CasRPN achieves an AUC score of 85.3% on the large-scale Trackingnet dataset.
Jia Chen 0012, Youkang Yuan, Xinrong Hu, Tao Peng 0006
Expert Syst. Appl.1
2025 ViT-BF: vision transformer with border-aware features for visual tracking
Ping Li 0016, Jinxing Liang, Tao Peng 0006, Jia Chen 0012, Li Li 0094, Xinrong Hu, Junping Liu
Vis. Comput.6
2024 DesignGAN: Generation of Hand-Drawn Garment Sketches
Xinrong Hu, Jiwei Huang, Tao Peng 0006, Feng Yu 0017, Jia Chen 0012
CGI (1)7
2024 SGM: A Dataset for 3D Garment Reconstruction from Single Hand-Drawn Sketch
abstract
High-fidelity garment reconstruction is essential for various applications such as garment design and virtual try-on. While image-based reconstruction methods have made significant progress with deep generative models, generating 3D models from hand-drawn sketches to meet design intentions remains challenging. One of the main obstacles is the limited availability of large-scale 3D garment models accompanied by corresponding sketches. To address this issue, we propose SGM, a comprehensive dataset comprising 656 garment models categorized into short and long sleeves. Each garment model in SGM is accompanied by four types of rendered images and a series of UDF values. Furthermore, we introduce a novel baseline approach for sketch-based garment reconstruction using an end-to-end generative network capable of generating garment models from single hand-drawn sketches. Extensive experimental results highlight the significance and value of our proposed dataset and method. We plan to make SGM publicly available upon publication.
Jia Chen 0012, Jinlong Qin, Saishang Zhong, Xinrong Hu, Tao Peng 0006
ICASSP1
2023 MARANet: Multi-scale Adaptive Region Attention Network for Few-Shot Learning
Jia Chen 0012, Xiyang Li, Yangjun Ou, Xinrong Hu, Tao Peng 0006
CGI (1)1
2023 FoldGEN: Multimodal Transformer for Garment Sketch-to-Photo Generation
Jia Chen 0012, Yanfang Wen, Xinrong Hu, Tao Peng 0006
CGI1
2023 Cross-cycle Transformer-based Stitching Method for Low-resolution Borehole Images
abstract
The stitching of borehole images has an important predictive role in safety analysis in the field of geotechnical engineering and intelligent geological exploration. Applying traditional image stitching methods that designed specifically for high-resolution images to low-resolution images will lead to blurred stitching results, stitching seams, fewer matched feature points and difficulties in massive image stitching. To address these problems, we propose an autoencoder-based coarse-to-fine feature extraction network, which can extract image features with high semantic and improves the accuracy of the feature point matching. Besides, we design a cross-cycle Transformer-based image stitching framework, which increase the number of matching feature points by Cross-QuadTree attention and stitch image by affine transformation. Experimental results show that the proposed method can effectively stitch low-resolution geotechnical borehole images with satisfactory visual quality.
Jia Chen 0012, Zhenpeng Fu, Mingfu Xiong, Xinrong Hu, Tao Peng 0006
ICME1
2023 Unsupervised Fashion Style Learning by Solving Fashion Jigsaw Puzzles
abstract
Fashion style learning is the basis for many tasks in fashion AI, such as clothing recommendations, fashion trend analysis and popularity prediction. Most of the existing methods rely on the quality and quantity of the annotations. This paper proposes an efficient two-step unsupervised fashion style learning framework with "Fashion Jigsaw" task and centroid-based density clustering algorithm. First, we design the "Fashion Jigsaw" unsupervised learning task according to the distribution of fashion elements in full-body fashion images. By splitting and recovering fashion images, we pre-train a model that can extract both intra-image and inter-image information. Second, we propose a centroid-based density clustering algorithm and introduce the concept of "centroid" to cluster fashion image features and represent fashion styles. Meanwhile, we keep the noise features to discover the newly sprouted fashion styles. Experiment results demonstrate the effectiveness of our proposed method.
Jia Chen 0012, Haidongqing Yuan, Tao Peng 0006, Xinrong Hu
ICME1
2023 Fashion Trend Forecasting Based on Multivariate Attention Fusion
Jia Chen 0012, Saishang Zhong, Xinrong Hu
ICONIP (8)1
2023 Graphormer-Based Contextual Reasoning Network for Small Object Detection
Jia Chen 0012, Xiyang Li, Yangjun Ou, Xinrong Hu, Tao Peng 0006
PRCV (9)1
2023 BovdGFE: buffer overflow vulnerability detection based on graph feature extraction
Xinghang Lv, Tao Peng 0006, Jia Chen 0012, Junping Liu, Xinrong Hu, Ruhan He, Minghua Jiang, Wenli Cao
Appl. Intell.3
2022 Few-Shot Detection Based on an Enhanced Prototype for Outdoor Small Forbidden Objects
Jia Chen 0012, Xinzhou Chen, Xinrong Hu, Tao Peng 0006
CGI1
2022 Boosting vision transformer for low-resolution borehole image stitching through algebraic multigrid
Jia Chen 0012, Zhenpeng Fu, Xinrong Hu, Tao Peng 0006
Vis. Comput.1
2021 A Triplet Appearance Parsing Network for Person Re-Identification
abstract
As one of the specific vision tasks, person re-identification has become a prevalent research topic in the field of multimedia and computer vision. However, existing feature extraction methods, originating from the quality of the bounding boxes which could cause the inhomogeneity and incoherence of person representation for cluttered backgrounds, are difficult to adapt the challenges of the harsh real-world scenarios. This study develops a Triplet person Appearances Parsing Framework (TAPF) which eliminates the surrounding interference factors of bounding boxes for person re-identification. The framework consists of a triplet person parsing network and an integration mechanism for person local and global appearance information. Concretely, the triplet parsing network includes a channel parsing module, a position parsing module and a color parsing module, which are used to extract the person channel parsing descriptor, regional descriptor and color perception descriptor, respectively. Then, a local and global flatten gaussian operations are performed to integrate the person appearance parsing descriptors to obtain more discriminative features for the person representation. The experimental results have been conducted to validate our proposed algorithm can achieve a better performance for person re-identification on several public datasets, i.e., VIPeR and Market-1501, respectively.
Mingfu Xiong, Zhongyuan Wang 0001, Ruhan He, Xinrong Hu, Xiao Qin 0001, Jia Chen 0012
ICASSP7
2021 A Structured Feature Learning Model for Clothing Keypoints Localization
Ruhan He, Yuyi Su, Tao Peng 0006, Jia Chen 0012, Xinrong Hu
MMM (1)4
2020 HybridGAN: hybrid generative adversarial networks for MR image synthesis
Jia Chen 0012, Mingfu Xiong, Tao Peng 0006, Minghua Jiang, Xiao Qin 0001
Multim. Tools Appl.1
2018 RAPID: Measuring Deformation of Biological Tissues from MR Images Through the Riemannian Pseudo Kernel
abstract
Due to the nonlinear deformation of nonrigid and nonuniform tissues, it is challenging to accurately measure the displacements of feature points distributed on the inner parts, boundaries, and separatrices of tissue layers. To address this challenge, we propose a feature point matching technique called RAPID to measure MR 2D slice deformation of nonuniform and nonrigid biological tissues. We propose to use the covariance of several neighboring point statistics computed around a keypoint, as the keypoint descriptor. Inspired by the kernel methods, we advocate adopting a Riemannian pseudo kernel to map SPD matrices to a high dimensional Hilbert space, where the Euclidean geometry applies. We compare our RAPID with two existing schemes (i.e., SIFT and SURF). Our experimental results show that our RAPID is superior to SIFT and SURF, because the benefits offered by RAPID are two-fold. First, our RAPID increases the number of matched data points. Second, RAPID substantially improves the key-point matching accuracy of SIFT and SURF.
Jia Chen 0012, Ruhan He, Xinrong Hu, Xiao Qin 0001
Int. J. Pattern Recognit. Artif. Intell.1