Yanzhi Song

dblp:198/1197 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0002-7324-3797ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Text2Scenes: Language-Guided Synthesis of Complex Indoor Scenes
Xintong Dong, Chuanyang Li, Zhouwang Yang, Yanzhi Song
Int. J. Comput. Vis.5
2026 LuBan: Constructing Hierarchical Graphs for CAD Sketch Generation via Transformer Intermediate Outputs
abstract
Computer-Aided Design (CAD) sketches, composed of geometric primitives and constraints, are fundamental to CAD models and play a critical role in industrial design and manufacturing. Leveraging artificial intelligence to convert hand-drawn sketches and rendered images into CAD sketches has the potential to streamline and accelerate the design process. Existing approaches predominantly focus on separately learning primitives and constraints, often employing two-stage methods or learnable tokens to model these elements independently. However, such methods fail to fully exploit the intrinsic relationships between primitives and constraints. In this paper, we propose LuBan, a lightweight, end-to-end model for CAD sketch generation that eliminates the need for separate constraint models or tokens. LuBan leverages the DEtection TRansformer (DETR) architecture for primitive modeling and distinguishes between parametric and non-parametric features. By deriving sub-primitive features from the intermediate outputs of the primitive model, LuBan facilitates constraint prediction, effectively capturing the inherent relationships between primitives and constraints. This enables the generation of CAD sketches as hierarchical graphs. Qualitative and quantitative experiments on both precise and hand-drawn renderings demonstrate that LuBan achieves state-of-the-art performance. Ablation studies further confirm its superiority over independently trained primitive models, validating its effectiveness. Additionally, LuBan embodies the principle of "what you draw is what you get," offering significant enhancements to the design process for designers.
Chuanyang Li, Chuqi Han, Yanzhi Song, Zhouwang Yang
IEEE Trans. Vis. Comput. Graph.3
2025 Image Hashing Based on Hamming Ball Spacing
abstract
Image hashing has been widely used in large-scale image retrieval due to its enhancement of storage space and retrieval speed. Recently, methods based on hash centers have achieved impressive retrieval performance, aiming to assign mutually separated hash codes as center points for each category and learn compact binary codes by minimizing intra-class variance. However, current methods tend to generate codebooks in advance, which are then used as pre-defined hash centers, leading to compromised quality of hash centers in instance-level datasets with excessive categories. To address this problem, we designed a training strategy that learns hash centers by constraining the margin of category clusters in Hamming space, during which the Gilbert-Varshamov bound and the upper bound of Hamming ball packing from coding theory are utilized to determine the range of margins. Finally, we jointly optimize hash encoder and hash centers to improve retrieval performance. Extensive experiments on more challenging large-scale instance-level datasets demonstrate that our method effectively overcomes the limitation of retrieval performance imposed by the number of categories. Code is at: https://github.com/GrimmAI/Hamming-Ball-Hashing.git.
Zhaomeng Ma, Hailong Shen, Yanzhi Song
CIKM4
2025 NLPA-AD: normal label propagation algorithm for zero-shot texture anomaly detection
Yanzhi Song, Zhouwang Yang, Chencheng Wang
Appl. Intell.2
2025 Distribution line inspection method using multi-scale information augmentation and ensemble learning
Yihao Liang, Liangwu Wei, Yanzhi Song, Zhouwang Yang
Eng. Appl. Artif. Intell.3
2025 GDViT: Group-level decorrelation-based vision transformer for domain generalization
Wenqiang Tang, Zhouwang Yang, Yanzhi Song
Neurocomputing3
2025 Cross-hierarchical bidirectional consistency learning for fine-grained visual classification
Pengxiang Gao, Yihao Liang, Yanzhi Song, Zhouwang Yang
Inf. Sci.3
2025 DFTGL: Domain Filtered and Target Guided Learning for few-shot anomaly detection
Yanzhi Song, Zhouwang Yang
J. Vis. Commun. Image Represent.2
2025 DC-AD: A Divide-and-Conquer Method for Few-Shot Anomaly Detection
Zhouwang Yang, Yanzhi Song
Pattern Recognit.3
2024 CLCP: Realtime Text-Image Retrieval for Retailing via Pre-trained Clustering and Priority Queue
abstract
Real-time matching between customer demands and product information via text-image retrieval remains a fundamental problem in intelligent retailing. However, this process involves challenges covering data quality, multi-modal retrieval strategies and performing efficiency. To alleviate the case, we propose a cross-modality retrieval pipeline leveraging contrastive loss and a novel sampling strategy. We also address text-image retrieval as a two-stage process, involving unsupervised clustering and contrastive feature representation. Additionally, we create an image-caption matching dataset by expanding the Grocery Store Dataset using a fundamental visual-language model. Our experiments demonstrate the effectiveness of our method on both an expanded new dataset and the well-known cross-modality retrieval benchmark, Flicker30k.
Liangwu Wei, Yuntao Wei, Yanzhi Song
ICMR5
2024 SGIR: Star Graph-Based Interaction for Efficient and Robust Multimodal Representation
abstract
Multimodal representation aims to integrate information from multiple modalities to improve overall performance. Recent works utilizing pairwise interactions have been proposed to deal with the long-range inter-modal and intra-modal dependencies in modeling multimodal data. However, these works usually feature high model complexity, and they are not robust to noisy multimodal data. To address these problems, we propose a novel multimodal representation method that learns private and hub representations of modalities. These representations and their connections form a star graph, a basis for Star Graph-based Interaction (SGI). SGI not only captures the long-range dependencies in multimodal data but also has two natural properties. Firstly, the number of modal interactions increases linearly with the number of modalities, which is computationally efficient compared with the square increase rate of pairwise interactions in previous works. Secondly, the indirect modal interactions through the hub representation in SGI (rather than the direct pairwise interactions between modalities) ensure the model's robustness to noisy modalities. Experiments on five benchmark datasets demonstrate that our new SGI representation (SGIR) achieves state-of-the-art performance on various multimodal tasks, and our qualitative and quantitative analyses show the excellent generalization ability of SGIR. Further experiments reveal that SGIR still outperforms widely used baseline models when modalities are corrupted by low levels of noise.
Zhiwei Ding, Guilin Lan, Yanzhi Song, Zhouwang Yang
IEEE Trans. Multim.3
2023 A two-stage anomaly detection framework: Towards low omission rate in industrial vision applications
Jianyu Liu, Zhouwang Yang, Yanzhi Song
Adv. Eng. Informatics3
2023 Disease-grading networks with ordinal regularization for medical imaging
Wenqiang Tang, Zhouwang Yang, Yanzhi Song
Neurocomputing3
2023 Selective interactive networks with knowledge graphs for image classification
Wenqiang Tang, Zhouwang Yang, Yanzhi Song
Knowl. Based Syst.3
2023 A Natural Threshold Model for Ordinal Regression
Yanzhi Song, Zhouwang Yang
Neural Process. Lett.2
2022 Vectorized instance segmentation using periodic B-splines based on cascade architecture
Fangjun Wang, Yanzhi Song, Zhangjin Huang, Zhouwang Yang
Comput. Graph.2
2018 Function representation based slicer for 3D printing
Yanzhi Song, Zhouwang Yang, Yuan Liu 0025, Jiansong Deng
Comput. Aided Geom. Des.1
2017 Implicit surface reconstruction with total variation regularization
Yuan Liu 0025, Yanzhi Song, Zhouwang Yang, Jiansong Deng
Comput. Aided Geom. Des.2