Honghui Yuan

dblp:290/8343 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SceneTextStylizer: a training-free diffusion framework for scene text style transfer
Honghui Yuan, Keiji Yanai
Multim. Syst.1
2025 Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program Tuning
abstract
Tensor program tuning is essential for the efficient deployment of deep neural networks. Search-based approaches have demonstrated scalability and effectiveness in automatically finding high-performance programs for specific hardware. However, the search process is often inefficient, taking hours or even days to discover optimal programs due to the exploration mechanisms guided by an accurate but slow-learned cost model. Meanwhile, the learned cost model trained on one platform cannot seamlessly adapt online to another, which we call cross-platform online unawareness. In this work, we propose Pruner and MoA-Pruner. Pruner is a ''Draft-then-Verify'' exploration mechanism that accelerates the schedule search process. Instead of applying the complex learned cost model to all explored candidates, Pruner drafts small-scale potential candidates by introducing a naive Symbol-based Analyzer (draft model), then identifies the best candidates by the learned cost model. MoA-Pruner introduces a Momentum online Adaptation strategy to address the cross-platform online unawareness.
Jun Shi 0007, Minfan Zhao, Junshi Chen 0003, Hong An, Xulong Tang, Honghui Yuan
ASPLOS (2)12
2025 Japanese Kuzushiji Font Generation Employing Differentiable Renderer
Honghui Yuan, Junwen Chen 0003, Keiji Yanai
ICDAR (5)1
2025 CIExplorer: Microarchitecture-Aware Exploration for Tightly Integrated Custom Instruction
abstract
Extending existing architectures with customized instruction extensions is emerging to achieve high performance and energy efficiency for specific applications.Automated discovery of custom instructions (CIs) is well-studied nowadays, which requires exploring combinations of different types and quantities of operations, resulting in a vast search space.However, previous works typically use microarchitectureagnostic cost models, leading to suboptimal CIs that may degrade performance.They leverage graph isomorphism to reduce area overhead, but few of them consider its potential to benefit performance-oriented exploration.To this end, we present CIExplorer, a framework for adaptive CI exploration.
Qingcai Jiang, Jun Shi 0007, Junshi Chen 0003, Hong An, Xulong Tang, Honghui Yuan
ICS10
2025 KuzushijiGen: A Real-Time Few-Shot Japanese Kuzushiji Generator via Differentiable Rendering
abstract
Kuzushiji is an ancient form of Japanese script that has deteriorated over time, making many characters unrecognizable. Recent font generation methods have mainly focused on modern font generation using large-scale raster images, making it difficult to apply them to ancient handwritten scripts like Kuzushiji. To address this challenge, we propose a web-based real-time application for generating Kuzushiji fonts using vector images without requiring any model training, which incorporates a differentiable rasterizer, a conditional diffusion model, a discriminator, and a Multi-Stroke Encoder for stroke-level text control. Multiple loss functions are employed to optimize font parameters. This application allows users to freely input modern Japanese characters to generate corresponding Kuzushiji images. Our results are resolution-independent and are more effective for preserving historical text. Supplementary slides: http://bit.ly/41YOyug
Honghui Yuan, Keiji Yanai
MMAsia1
2025 KuzushijiDiffuser: Japanese Kuzushiji Font Generation with FontDiffuser
Honghui Yuan, Keiji Yanai
MMM (2)1
2025 KuzushijiFontDiff: Diffusion Model for Japanese Kuzushiji Font Generation
Honghui Yuan, Keiji Yanai
MMM (5)1
2025 SceneTextStyler: Editing Text with Style Transformation
Honghui Yuan, Keiji Yanai
MMM (5)1
2024 Font Style Translation in Scene Text Images with CLIPstyler
Honghui Yuan, Keiji Yanai
ICPR (19)1
2021 Ascend: a Scalable and Unified Architecture for Ubiquitous Deep Neural Network Computing : Industry Track Paper
abstract
Deep neural networks (DNNs) have been successfully applied to a great variety of applications, ranging from small IoT devices to large scale services in a data center. In order to improve the efficiency of processing these DNN models, dedicated hardware accelerators are required for all these scenarios. Theoretically, there exists an optimized acceleration architecture for each application. However, considering the cost of chip design and corresponding tool-chain development, researchers need to trade off between efficiency and generality. In this work, we demonstrate that it is practical to use a unified architecture, called Ascend, to support those applications, ranging from IoT devices to data-center services. We provide a lot of design details to explain that the success of Ascend relies on contributions from different levels. First, heterogeneous computing units are employed to support various DNN models. And the datapath is adapted according to the requirement of computing and data access. Second, when scaling the Ascend architecture from a single core to a cluster containing thousands of cores, it involves design efforts, such as memory hierarchy and system level integration. Third, a multi-tier compiler, which provides flexible choices for developers, is the last critical piece. Experimental results show that using accelerators based on the Ascend architecture can achieve comparable or even better performance in different applications. In addition, various chips based on the Ascend architecture have been successfully commercialized. More than 100 million chips have been used in real products.
Heng Liao, Jiajin Tu, Xiping Zhou, Honghui Yuan, Yuxing Hu
HPCA6