Hao-Yu Wu

dblp:117/6315 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Generative Model Based Standard Cell Timing Library Characterization
abstract
Accurate cell timing characterization is essential, on which static timing analysis relies to verify timing performance and ensure design robustness across various PVT conditions (corners). The corner explosion in modern design amplifies the efficiency and scalability challenge for accurate characterization. However, the conventional characterization approach of SPICE simulation alone becomes prohibitively expensive due to the increasing computational complexity and the amount of characterized data. In this paper, we view the characterization problem from a generative modeling perspective to tackle the efficiency and scalability challenge. With a hybrid of generative adversarial network (GAN) and autoencoder, our generative model learns and generalizes among various timing arcs and corners. Experimental results demonstrate that the proposed framework achieves high accuracy and extensibility while reducing the runtime significantly.
Hao-Yu Wu, Hsin-Tzu Chang, Shiuan-Yun Ding, Iris Hui-Ru Jiang, Benson Tsao, Vinson Wu, Wei-Kai Shih
DAC1
2025 How Does Human-Robot Collaboration Affect Hotel Employees' Proactive Behavior?
abstract
Based on Conservation of Resources (COR) theory, this study collected the data of 342 front-line employees of five-star hotels in northeast China by means of longitudinal data collection. SPSS 25.0 and Mplus 8.0 were used for data analysis to test the hypotheses proposed in this study. This study found that workload and harmonious passion play a mediating role between human-robot collaboration and proactive behavior. Robot trust not only moderates the relationship between human-robot collaboration on workload and harmonious passion, but also moderates the mediating effect of workload and harmonious passion. This study not only examines the effects of human-robot collaboration on the proactive behavior of hotel employees through cognitive and emotional pathways, but verifies the boundary conditions of robot trust. Otherwise, our study validates and extends COR theory to some extent.
Jia-Min Li, Ke-Xi Liu, Ji-Fei Xie, Hao-Yu Wu
Int. J. Hum. Comput. Interact.4
2025 Graceful Register Clustering and Rebanking for Power and Timing Balancing
abstract
As dynamic power has become the bottleneck to achieving low power, clock power reduction is crucial in modern IC design. Register clustering can effectively save clock power because of significantly reducing the number of clock sinks and register pin capacitance, clock routed wirelength, and the number of clock buffers. In this article, we propose effective mean shift to naturally form clusters according to register distribution without placement disruption. Effective mean shift fulfills the requirements to be a good register clustering algorithm because it needs no prespecified number of clusters, is insensitive to initializations, is robust to outliers, is tolerant of various register distributions, is efficient and scalable, and balances clock power reduction against timing degradation. Furthermore, we devise a rebanking mechanism to legalize the clustering result according to the adopted multibit register library. Experimental results show that our approach achieves superior power and timing balancing.
Tung-Wei Lin, Ya-Chu Chang, Hao-Yu Wu, Iris Hui-Ru Jiang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 Slack Redistributed Register Clustering with Mixed-Driving Strength Multi-bit Flip-Flops
abstract
Register clustering is an effective technique for suppressing the increasing dynamic power ratio in modern IC design. By clustering registers (flip-flops) into multi-bit flip-flops (MBFFs), clock circuitry can be shared, and the number of clock sinks and buffers can be lowered, thereby reducing power consumption. Recently, the use of mixed-driving strength MBFFs has provided more flexibility for power and timing optimization. Nevertheless, existing register clustering methods usually employ evenly distributed and invariant path slack strategies. Unlike them, in this work, we propose a register clustering algorithm with slack redistribution at the post-placement stage. Our approach allows registers to borrow slack from connected paths, creates the possibility to cluster with neighboring maximal cliques, and releases extra slack. An adaptive interval graph based on the red-black tree is developed to efficiently adapt timing feasible regions of flip-flops for slack redistribution. An attraction-repulsion force model is tailored to wisely select flip-flops to be included in each MBFF. Experimental results show that our approach outperforms state-of-the-art work in terms of clock power reduction, timing balancing, and runtime.
Hao-Yu Wu, Iris Hui-Ru Jiang, Cheng-Hong Tsai, Chien-Cheng Wu
ISPD2
2022 Billion-Scale Pretraining with Vision Transformers for Multi-Task Visual Representations
abstract
Large-scale pretraining of visual representations has led to state-of-the-art performance on a range of benchmark computer vision tasks, yet the benefits of these techniques at extreme scale in complex production systems has been relatively unexplored. We consider the case of a popular visual discovery product, where these representations are trained with multi-task learning, from use-case specific visual understanding (e.g. skin tone classification) to general representation learning for all visual content (e.g. embeddings for retrieval). In this work, we describe how we (1) generate a dataset with over a billion images via large weakly-supervised pretraining to improve the performance of these visual representations, and (2) leverage Transformers to replace the traditional convolutional backbone, with insights into both system and performance improvements, especially at 1B+ image scale. To support this backbone model, we detail a systematic approach to deriving weakly-supervised image annotations from heterogenous text signals, demonstrating the benefits of clustering techniques to handle the long-tail distribution of image labels. Through a comprehensive study of offline and online evaluation, we show that large-scale Transformer-based pretraining provides significant benefits to industry computer vision applications. The model is deployed in a production visual shopping system, with 36% improvement in top-1 relevance and 23% improvement in click-through volume. We conduct extensive experiments to better understand the empirical relationships between Transformer-based architectures, dataset scale, and the performance of production vision systems.
Josh Beal, Hao-Yu Wu, Dong Huk Park, Andrew Zhai, Dmitry Kislyuk
WACV2
2020 Shop The Look: Building a Large Scale Visual Shopping System at Pinterest
abstract
As online content becomes ever more visual, the demand for searching by visual queries grows correspondingly stronger. Shop The Look is an online shopping discovery service at Pinterest, leveraging visual search to enable users to find and buy products within an image. In this work, we provide a holistic view of how we built Shop The Look, a shopping oriented visual search system, along with lessons learned from addressing shopping needs. We discuss topics including core technology across object detection and visual embeddings, serving infrastructure for realtime inference, and data labeling methodology for training/evaluation data collection and human evaluation. The user-facing impacts of our system design choices are measured through offline evaluations, human relevance judgements, and online A/B experiments. The collective improvements amount to cumulative relative gains of over 160% in end-to-end human relevance judgements and over 80% in engagement. Shop The Look is deployed in production at Pinterest.
Raymond Shiau, Hao-Yu Wu, Eric Kim, Yue Li Du, Anqi Guo, Eileen Li, Kunlong Gu, Charles Rosenberg 0001, Andrew Zhai
KDD2
2019 Classification is a Strong Baseline for Deep Metric Learning
Andrew Zhai, Hao-Yu Wu
BMVC2
2019 Learning a Unified Embedding for Visual Search at Pinterest
abstract
At Pinterest, we utilize image embeddings throughout our search and recommendation systems to help our users navigate through visual content by powering experiences like browsing of related content and searching for exact products for shopping. In this work we describe a multi-task deep metric learning system to learn a single unified image embedding which can be used to power our multiple visual search products. The solution we present not only allows us to train for multiple application objectives in a single deep neural network architecture, but takes advantage of correlated information in the combination of all training data from each application to generate a unified embedding that outperforms all specialized embeddings previously deployed for each product.
Andrew Zhai, Hao-Yu Wu, Eric Tzeng, Dong Huk Park, Charles Rosenberg 0001
KDD2
2012 Eulerian video magnification for revealing subtle changes in the world
abstract
Our goal is to reveal temporal variations in videos that are difficult or impossible to see with the naked eye and display them in an indicative manner. Our method, which we call Eulerian Video Magnification, takes a standard video sequence as input, and applies spatial decomposition, followed by temporal filtering to the frames. The resulting signal is then amplified to reveal hidden information. Using our method, we are able to visualize the flow of blood as it fills the face and also to amplify and reveal small motions. Our technique can run in real time to show phenomena occurring at the temporal frequencies selected by the user.
Hao-Yu Wu, Michael Rubinstein, Eugene Shih, John V. Guttag, Frédo Durand, William T. Freeman
ACM Trans. Graph.1