Jianfang Li 0001

dblp:162/2981-1 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0003-2716-1886ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2025 SemTalk: Holistic Co-Speech Motion Generation with Frame-Level Semantic Emphasis
abstract
Co-speech gesture generation must carefully integrate common rhythmic motion with rare yet essential semantic gestures. In this work, we propose SemTalk for holistic co-speech gesture generation with frame-level semantic emphasis. Our key insight is to separately learn base motions and sparse motions, and then adaptively fuse them. In particular, coarse2fine cross-attention module and rhythmic consistency learning are explored to establish rhythm-related base motion, ensuring a coherent foundation that synchronizes gestures with the speech rhythm. Subsequently, semantic emphasis learning is designed to generate semantic-aware sparse motion, focusing on frame-level semantic cues. Finally, to integrate sparse motion into the base motion and generate semantic-emphasized co-speech gestures, we further leverage a learned semantic score for adaptive synthesis. Qualitative and quantitative comparisons on two public datasets demonstrate that our method outperforms the state-of-the-art, delivering high-quality co-speech motion with enhanced semantic richness over a stable base motion.
Xiangyue Zhang, Jianfang Li 0001, Ziqiang Dang, Jianqiang Ren, Liefeng Bo, Zhigang Tu 0001
ICCV2
2025 EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
Xiangyue Zhang, Jianfang Li 0001, Jianqiang Ren, Liefeng Bo, Zhigang Tu 0001
ACM Multimedia2
2025 Cascaded Dual Vision Transformer for Accurate Facial Landmark Detection
abstract
Facial landmark detection is a fundamental problem in computer vision for many downstream applications. This paper introduces a new facial landmark detector based on vision transformers, which consists of two unique designs: Dual Vision Transformer (D-ViT) and Long Skip Connections (LSC). Based on the observation that the channel dimension of feature maps essentially represents the linear bases of the heatmap space, we propose learning the inter-connections between these linear bases to model the inherent geometric relations among landmarks via channel-split ViT. We integrate such channel-split ViT into the standard vision transformer (i.e., spatial-split ViT),forming our Dual Vision Transformer to constitute the prediction blocks. We also suggest using long skip connections to deliver low-level image features to all prediction blocks, thereby preventing useful information from being discarded by intermediate supervision. Extensive experiments are conducted to evaluate the performance of our proposal on the widely used benchmarks, i.e., WFLW [45], COFW [3], and 300W [34], demonstrating that our model outperforms the previous SOTAs across all three benchmarks.
Ziqiang Dang, Jianfang Li 0001
WACV2
2023 DG3D: Generating High Quality 3D Textured Shapes by Learning to Discriminate Multi-Modal Diffusion-Renderings
abstract
Many virtual reality applications require massive 3D content, which impels the need for low-cost and efficient modeling tools in terms of quality and quantity. In this paper, we present a Diffusion-augmented Generative model to generate high-fidelity 3D textured meshes that can be directly used in modern graphics engines. Challenges in directly generating textured mesh arise from the instability and texture incompleteness of a hybrid framework which contains conversion between 2D features and 3D space. To alleviate these difficulties, DG3D incorporates a diffusion-based augmentation module into the min-max game between the 3D tetrahedral mesh generator and 2D renderings discriminators, which stabilizes network optimization and prevents mode collapse in vanilla GANs. We also suggest using multi-modal renderings in discrimination to further increase the aesthetics and completeness of generated textures. Extensive experiments on the public benchmark and real scans show that our proposed DG3D outperforms existing state-of-the-art methods by a large margin, i.e., 5% ∼ 40% in FID-3D score and 5%∼10% in geometry-related metrics. Code is available at https://github.com/seakforzq/DG3D.
Qi Zuo, Jianfang Li 0001, Liefeng Bo
ICCV3
2022 GAT-CADNet: Graph Attention Network for Panoptic Symbol Spotting in CAD Drawings
abstract
Spotting graphical symbols from the computer-aided design (CAD) drawings is essential to many industrial applications. Different from raster images, CAD drawings are vector graphics consisting of geometric primitives such as segments, arcs, and circles. By treating each CAD drawing as a graph, we propose a novel graph attention network GAT-CADNet to solve the panoptic symbol spotting problem: vertex features derived from the GAT branch are mapped to semantic labels, while their attention scores are cascaded and mapped to instance prediction. Our key contributions are three-fold: 1) the instance symbol spotting task is formulated as a subgraph detection problem and solved by predicting the adjacency matrix; 2) a relative spatial encoding (RSE) module explicitly encodes the relative positional and geometric relation among vertices to enhance the vertex attention; 3) a cascaded edge encoding (CEE) module extracts vertex attentions from multiple stages of GAT and treats them as edge encoding to predict the adjacency matrix. The proposed GAT-CADNet is intuitive yet effective and manages to solve the panoptic symbol spotting problem in one consolidated network. Extensive experiments and ablation studies on the public benchmark show that our graph-based approach surpasses existing state-of-the-art methods by a large margin.
Zhaohua Zheng, Jianfang Li 0001, Lingjie Zhu, Honghua Li, Frank Petzold, Ping Tan 0002
CVPR2
2015 Boundary-dominant flower blooming simulation
abstract
Abstract This paper presents a new physics‐based simulation method for flower blossom, which is based on biological observations that flower opening is usually driven by a boundary‐dominant morphological transition in a curved petal. We use an elastic triangular mesh representing a flower petal and adopt in‐plane expansion to induce global bending. Out‐of‐plane curl plays an auxiliary role in reducing the curvatures of cross‐sections. We also propose to adapt semi‐implicit Euler time integrator for fast simulation results, which has intrinsic damping and at least one order precision. Our system allows users to control the blossoming process by simply specifying a growth curve, which is easy to design because of the boundary‐dominant property. Experimental results show that our physics‐based system runs faster and generates more realistic and convincing blossom results than the existing simulation methods. Copyright © 2015 John Wiley & Sons, Ltd.
Jianfang Li 0001, Weiwei Xu 0003, Haiyi Liang
Comput. Animat. Virtual Worlds1