Hongyu Yan

dblp:254/8227 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CTR3D: Cross-View Token Reduction for Dense Multi-View Generation
abstract
Recent multi-view diffusion (MVD) methods have utilized the generative capabilities of 2D image diffusion models to produce multi-view images from a single-view input. However, existing approaches often depend on dense crossview attention layers, which hinder scalability and fidelity due to their high computational costs. In this paper, we propose CTR3D, a novel method that incorporates token reduction in multi-view attention layers to efficiently generate dense, high-resolution multi-view images without restricting the camera viewpoints of the generated views. Our approach is designed into three key steps: redundancy removal, attention interaction, and token recovery. These steps leverage lightweight, projection-based techniques for multi-view token reduction and recovery, significantly improving the computational efficiency of MVD. By reducing the number of tokens in attention layers while preserving multi-view consistency, our model achieves state-of-the-art performance in novel view synthesis and 3D reconstruction while keeping efficiency for generation of dense high-resolution images and normals. Experimental results demonstrate that our method surpasses existing approaches, providing a more efficient and effective solution for multi-view generation. https://github.com/HKUST-SAIL/CTR3D
Kunming Luo, Hongyu Yan, Yuan Liu 0025, Manyuan Zhang, Wenping Wang 0001, Ping Tan 0002
3DV2
2026 A Novel Ray-Tracing Channel Model with Full-Wave Simulations for THz Communication Systems
Yihan Zhao, Songjiang Yang, Yinghua Wang, Junling Li, Hongyu Yan, Cheng-Xiang Wang 0001
WCNC5
2025 SymmCompletion: High-Fidelity and High-Consistency Point Cloud Completion with Symmetry Guidance
abstract
Point cloud completion aims to recover a complete point shape from a partial point cloud. Although existing methods can form satisfactory point clouds in global completeness, they often lose the original geometry details and face the problem of geometric inconsistency between existing point clouds and reconstructed missing parts. To tackle this problem, we introduce SymmCompletion, a highly effective completion method based on symmetry guidance. Our method comprises two primary components: a Local Symmetry Transformation Network (LSTNet) and a Symmetry-Guidance Transformer (SGFormer). First, LSTNet efficiently estimates point-wise local symmetry transformation to transform key geometries of partial inputs into missing regions, thereby generating geometry-align partial-missing pairs and initial point clouds. Second, SGFormer leverages the geometric features of partial-missing pairs as the explicit symmetric guidance that can constrain the refinement process for initial point clouds. As a result, SGFormer can exploit provided priors to form high-fidelity and geometry-consistency final point clouds. Qualitative and quantitative evaluations on several benchmark datasets demonstrate that our method outperforms state-of-the-art completion networks.
Hongyu Yan, Kunming Luo
AAAI1
2025 Neural Assembler: Learning to Generate Fine-Grained Robotic Assembly Instructions from Multi-View Images
abstract
Image-guided object assembly represents a burgeoning research topic in computer vision. This paper introduces a novel task: translating multi-view images of a structural 3D model (for example, one constructed with building blocks drawn from a 3D-object library) into a detailed sequence of assembly instructions executable by a robotic arm. Fed with multi-view images of the target 3D model for replication, the model designed for this task must address several sub-tasks, including recognizing individual components used in constructing the 3D model, estimating the geometric pose of each component, and deducing a feasible assembly order adhering to physical rules. Establishing accurate 2D-3D correspondence between multi-view images and 3D objects is technically challenging. To tackle this, we propose an end-to-end model known as the Neural Assembler. This model learns an object graph where each vertex represents recognized components from the images, and the edges specify the topology of the 3D model, enabling the derivation of an assembly plan. We establish benchmarks for this task and conduct comprehensive empirical evaluations of Neural Assembler and alternative solutions. Our experiments clearly demonstrate the superiority of Neural Assembler.
Hongyu Yan, Yadong Mu
AAAI1
2025 CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry Refiner
abstract
We present a novel generative 3D modeling system, coined CraftsMan3D, which can generate high-fidelity 3D geometries with highly varied shapes, detailed surfaces, and, notably, allows for refining the geometry in an interactive manner. Despite the significant advancements in 3D generation, existing methods still struggle with lengthy optimization processes, self-occlusion, irregular mesh topologies, and difficulties in accommodating user editing, consequently impeding their widespread adoption and implementation in 3D modeling softwares. Our work is inspired by the craftsman, who usually roughs out the holistic figure of the work first and elaborates the surface details subsequently. Specifically, we first introduce a robust data preprocessing pipeline that utilizes visibility check and winding mumber to maximize the use of existing 3D data. Leveraging this data, we employ a 3D-native DiT model that directly models the distribution of 3D data in latent space, generating coarse geometries in seconds. Subsequently, a normal-based geometry refiner enhances local surface details, which can be applied automatically or interactively with user input. Extensive experiments demonstrate that our method achieves high efficacy in producing superior quality 3D meshes compared to existing methods.
Jiarui Liu 0003, Hongyu Yan, Yixun Liang, Xuelin Chen, Ping Tan 0002, Xiaoxiao Long
CVPR3
2025 HT-STNet: a hierarchical Tucker decomposition and spatio-temporal LSTM network for accurate and efficient shared mobility demand forecasting on sparse data
Hongyu Yan, Benjia Chu, Zhihao Xu 0002
Appl. Intell.1
2025 TerraCraft: City-scale generative procedural modeling with natural languages
abstract
Automated generation of large-scale 3D scenes presents a significant challenge due to the resource-intensive training and datasets required. This is in sharp contrast to the 2D counterparts that have become readily available due to their superior speed and quality. However, prior work in 3D procedural modeling has demonstrated promise in generating high-quality assets using the combination of algorithms and user-defined rules. To leverage the best of both 2D generative models and procedural modeling tools, we present TerraCraft, a novel framework for generating geometrically high-quality 3D city-scale scenes. By utilizing Large Language Models (LLMs), TerraCraft can generate city-scale 3D scenes from natural text descriptions. With its intuitive operation and powerful capabilities, TerraCraft enables users to easily create geometrically high-quality scenes readily for various applications, such as virtual reality and game design. We validate TerraCraft’s effectiveness through extensive experiments and user studies, showing its superior performance compared to existing baselines.
Zhihao Yao 0004, Zi-Qi Lu, Hongyu Yan, Tai-Jiang Mu, Qun-Ce Xu
Graph. Model.5
2025 Shared mobility demand prediction via A fast spatiotemporal tensor autoregression
Hongyu Yan, Zhiqiang Lv, Benjia Chu, Zhihao Xu 0002
Eng. Appl. Artif. Intell.1
2022 FBNet: Feedback Network for Point Cloud Completion
Hongyu Yan, Jingjing Wang 0005, Di Xie, Shiliang Pu
ECCV (2)2
2022 Low-Level Graph Convolution Network for Point Cloud Processing
Hongyu Yan
ICANN (2)1
2020 End-to-end video subtitle recognition via a deep Residual Neural Network
Hongyu Yan, Xin Xu 0007
Pattern Recognit. Lett.1
2019 Adversarial Training Based Cross-Lingual Emotion Cause Extraction
Hongyu Yan, Qinghong Gao, Jiachen Du, Binyang Li, Ruifeng Xu 0001
CICLing (2)1
2019 A Knowledge Regularized Hierarchical Approach for Emotion Cause Analysis
abstract
Chuang Fan, Hongyu Yan, Jiachen Du, Lin Gui, Lidong Bing, Min Yang, Ruifeng Xu, Ruibin Mao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Chuang Fan, Hongyu Yan, Jiachen Du, Lin Gui 0003, Lidong Bing, Min Yang 0007, Ruifeng Xu 0001, Ruibin Mao
EMNLP/IJCNLP (1)2